mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-27 07:01:54 +02:00
* fix: acknowledge seeded plans before invoking review skills * fix: distinguish current plan input from conversation history * fix: keep hermetic plan reviews on manual permissions * fix: distinguish tool discovery from file permission ownership * fix: preserve initial plan mode in observation tests * fix: wait for scope decisions before writing review findings * fix: carry autoplan decisions consistently into review artifacts * test: retain native failure context in periodic assertions * fix: advance active file permissions before queued questions * fix: finish red-team attempts before retry and cleanup * fix: finalize plan format captures and judges before retry * fix: cancel setup-gbrain SDK attempts before fixture cleanup * test: select periodic consumers of the bounded attempt helper * fix native Bash permission cards and queued questions * fix: preserve independent decisions and review scope Keep CEO approach, engineering scope and outside-review choices from approving independent remedies together. Carry declared contracts through DX polish and resolve new gaps before editing the plan. Regenerate every host and retain existing stop boundaries. Validation: 654 focused tests passed across nine files; all-host generation passed. Full free and periodic validation pending. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: require approval before design plan amendments Align the Design review philosophy and rating recipe with its section protocol: resolve one proposed fix, then apply only that approved decision and retain honest scores for declined fixes. Validation: 469 focused tests passed across four files; all-host generation passed. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: observe native question completion before transcript persistence Match owned completion hooks to submitted choices, reject conflicting or late answers, and retain bounded failure evidence. * test: recognize review posture in acknowledged native questions Require the selected mode acknowledgement, a completed follow-up question, and its current decoded display while preserving existing posture assertions. * fix: preserve settled CEO choices and isolate pending remedies Resolve established approach gates with cited authority and keep independent fixes out of unrelated option commitments and plan amendments. * fix: carry approved DX choices through later review steps Choose documentation approaches within the accepted scope and map resolved confusion points without reopening them through a bulk menu. * test: handle native settings-file edit prompts Keep one-time owned-file approvals and retain the actual sampled Autoplan permission frame with its matching barrier state. * test: accept standard CEO reply directives with tuning footers Recognize the exact trailing preference footer and letter-list directive while preserving current-display and exact acknowledgement checks. * test: scope split reviewers to their generated plan artifacts * test: observe native Bash permissions and invocation results * test: handle owned Bash prompts during mode preference checks * test: preserve synchronous subprocess rejection in Codex fixture * Fix periodic review handoff navigation Recognize review-first and explicit manual-next-step labels while preserving exact action families, manual preference, and ambiguous-menu rejection. Co-authored-by: OpenAI Codex <noreply@openai.com> * Bind pending file permissions to distinct current targets Allow one captured file request to own the complete current dialog while unrelated file work is pending. Preserve same-path ambiguity, exact input ownership, and one-time grant checks. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make paired CEO verification choices genuinely unresolved Start the positive control with proposed manual checks so its unchanged oracle measures two new coverage decisions. Preserve runtime contracts, targets, count bounds, and all assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO review options and verification within approved scope Audit every offered option for independent add-ons and keep new verification depth pending until accepted. Preserve already requested coverage and trace plan changes to the actual decision. Co-authored-by: OpenAI Codex <noreply@openai.com> * Assemble DX review artifacts before appending the final report Keep early DX evidence above decisions, update artifact sections in place, and append the report using the actual current file suffix. Re-read after deleting an existing report before choosing the append anchor. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep outside plan reviews exclusive and invocation-owned Follow one preflight-selected backend, terminate failed Codex work before fallback, and allocate extra prompt/output files uniquely. Consume only the current invocation’s completed output. Co-authored-by: OpenAI Codex <noreply@openai.com> * Select periodic completion evaluations for report writer changes Register the shared review resolver for eight missing consumers and regress selection for all nine completion cases without changing their IDs or tiers. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep permission ambiguity fixtures on the same normalized target Use distinct raw spellings of one target in the four negative fixtures so they exercise the normalized duplicate-owner guard after exact current-file disambiguation. Preserve the existing exception, no-input, diagnostic and cleanup assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Clarify preserved contracts in engineering review fixture Co-authored-by: OpenAI Codex <noreply@openai.com> * Recognize the offered DX follow-up handoff Co-authored-by: OpenAI Codex <noreply@openai.com> * Check independent commitments before presenting review options Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep Codex review output and status in one shell invocation Co-authored-by: OpenAI Codex <noreply@openai.com> * Distinguish seeded plans from reports written by a test attempt Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Autoplan file approvals with bounded viewport resizing Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Bash approvals before binding the complete command Co-authored-by: OpenAI Codex <noreply@openai.com> * Isolate setup message tests from the shared checkout Run the real installer in a temporary payload with private config, require successful completion, and guard source and binary contents and mtimes. Co-authored-by: OpenAI Codex <noreply@openai.com> * Fix periodic native permission and report completion handling Match the pinned CLI's soft wraps and clipped headings without granting from incomplete frames. Retire completed file requests, retain mode annotations, and ask section captures for a short final acknowledgement after their full report is saved. Co-authored-by: OpenAI Codex <noreply@openai.com> * Preserve review approvals and validate DX comparison artifacts Keep independent remedies and approved amendments explicit. Give the synthetic DX review its existing documentation and validate peer comparison as required analysis alongside four native decisions. Add positive and negative semantic calibrations while preserving review counts, model budgets and prompt size limits. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make the five-finding CEO fixture's application boundary explicit Materialize the request adapter and service composition used by the synthetic payment application. Explicitly declare the revised unregistered-event and mail-telemetry assumptions while preserving uncaught handler errors, the original invoice path and all five unresolved findings. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO state-path checks scoped to directory preparation Co-authored-by: OpenAI Codex <noreply@openai.com> * Use checked ports and bounded cleanup in pair-agent tests Discover the daemon port from its owned state file, retain startup diagnostics, and await failed-start cleanup. Add occupied-port, early-exit, deadline, and foreign-state regressions while preserving the existing HTTP assertions and hook budgets. Co-authored-by: Codex <noreply@openai.com> * Preserve queued edit identity and recover clipped Bash permissions Distinguish separately queued unfinished edits from mutation of one native tool ID. Keep grants bound to an exact owned request and reject reused IDs, ambiguous inputs, and competing owners. Support the pinned renderer's literal em dash and request a repaint when only the Bash card's top rule is clipped. Grants still require the complete fresh card and an exact native acknowledgment. Validation: 413 integrated parser/event tests passed; private repaint controls and joint source review passed. Full canonical suite and native periodic rerun remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep periodic reviews within their approved contracts and deliverables Carry exact approvals through engineering review, preserve declared contracts when amending CEO plans, and keep prioritization at the requested decision level. Materialize the revised synthetic SDK reference contract while retaining the five original documentation gaps. Accept the observed semicolon in the finite DX handoff menu and register the direct source dependencies used by the engineering cases. Regenerate canonical review documents without changing model budgets, retries, count bands, or native completion assertions. Validation: all-host generation and 275 review, fixture, selection and parity tests passed. Full free-suite and native periodic validation remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep Eng approval cadence and independence guards explicit * Accept ordinary punctuation in manual review handoffs * Recover file permissions alongside queued Bash calls * Carry approved DX work through later review findings * Clarify the synthetic auth internal failure decision * Bound the periodic DX fixture to onboarding changes * Recognize native Design review handoff labels * Hold scope in the integration-choice review fixture * Carry approved Design decisions through review evidence * Capture listener state when feedback reload fails * Exclude workspace caches before checking deprecated flags * Verify Design UI scope against a seeded review plan * Clarify plan review decisions and outside-voice approval flow * Reject setup menus in the Design UI gate * docs: require focused repair validation before final acceptance * fix: separate review commitments within existing prompt budgets * docs: align generation and contributor validation guidance * fix: advance native review prompts and count acknowledged findings * chore: bump version and changelog (v1.87.1.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: enforce cheap checks and side-effect-free validation previews * fix: handle owned Fetch permissions and oversized native cards * test: ground review fixtures in independent executable contracts * fix: preserve review decisions and verify reports before completion * test: construct the synthetic credential URL without a scanner false positive * test: materialize DX examples and verify their actual local behavior * fix: clarify CEO review decisions and execution order * fix: clarify review workflow ordering and select Design quality checks * Fix review decision gates and incomplete evaluation fixtures Persist CEO and engineering commitment ledgers before menus, preserve exact approvals, and distinguish implementation structure from feature scope. Route Autoplan through the canonical CEO Step 0 ordering. Classify DX findings before requesting approval and ground runtime claims in actual evidence. Complete neutral non-target fixture contracts and accept the captured Design handoff purpose without relaxing its ownership or acknowledgment checks. Record runtime-capability verification in AGENTS.md validation discipline. Validation: 1,335 focused tests passed across 21 files; build, all-host freshness, skill validation (647 artifacts / 107 tracked), and credential checks passed. Prior paid failures are preserved; behavioral acceptance remains pending. * Fix review decision boundaries and owned Read prompts Preserve exact approvals across review options, compare consistent DX milestones, and keep proposed implementation separate from review evidence. Bind modern Read prompts to one immutable native request and wait for its result. Retain captured regression verdicts, correct fixture error names, improve import probe diagnostics, and record focused-first validation discipline in AGENTS.md. * Clarify CEO and engineering review decisions Use explicit decision steps, one engineering ledger, and clear scope/write transitions. Preserve exact approvals and distinguish pending test requirements. Keep unrelated generated content unchanged. * Fix review decision ordering and native evaluation interactions * Clarify engineering decisions and test artifact order * Clarify pending choices and approvals in CEO reviews * Make CEO review phases sequential and clarify completion * Fix Design board submission intent matching * Seed an existing browser test baseline for Autoplan * Document decision-log payloads before state initialization * Preserve exact review scope and decide one change before drafting options * Require input identity before repeating passing model judges * Honor permitted storage throughout CEO review completion * Match complete native permission text within the pinned renderer contract * Align review approvals, independent choices, and bounded validation * fix: preserve reopened approvals and declare fixture interfaces * fix: isolate review artifacts and audit complete questions * fix: match detector artifact permissions to configured storage * fix: complete native permissions and review fixture workflows * fix: order CEO review work and separate engineering guarantees * fix: preserve native validation and separate review choices * fix: clarify review decisions and judge complete report context * fix: constrain review judgments and retain parse failures * fix: compare each affected value before review decisions * fix: make engineering review decisions and completion order explicit * fix: give the complete Autoplan evaluation a bounded chain budget * fix(cso): diagnose forbidden Docker endpoints before tool lookup * fix(reviews): reconcile workflow contracts and generated artifacts after main integration * fix(evals): migrate retained regressions to the native review harness * fix(tests): close native harness and workflow integration regressions * fix(evals): preserve complete permission context and native menu contracts * fix(tests): capture synchronous command output without pipe drain stalls * fix(reviews): clarify decision and completion ordering * fix(reviews): separate decision readiness from final completion checks * refactor(reviews): consolidate decision rules and completion branches * fix(plan-eng-review): order preparation and clarify decision routing * fix(plan-eng-review): restore size and question-format guard parity * fix(plan-eng-review): clarify scope phases and blocked completion * fix(plan-eng-review): unify review flow and report destination * fix(plan-eng-review): define bootstrap and question stage ownership * fix(plan-eng-review): clarify review structure and design lookup * fix(plan-eng-review): render report examples and show saved decisions * fix: consolidate Eng review decisions and select their evaluations * test: cover overlapping terminal attachments and clean merged runner type * fix: preserve Office Hours relationship closings during review updates * fix: retain pasted review targets across slash invocations * docs: preserve validation traces and correct release scope * test: cover pasted targets in both review skills * fix: validate report artifacts before recording success * fix: redact source roots at CSO report boundaries * fix: bind native Design questions before answering * test: select report privacy and native recovery regressions * test: bind rejection predicate in extracted observers * fix: bind complete boxed native questions * test: keep the Design UI fixture on native review * fix: preserve review decisions and evaluation completion outcomes * fix: clarify CEO approval and report completion order * fix: align native review evaluation ownership and completion * fix: bind review evaluators to native decisions and owned artifacts * fix: validate review decisions against native outcomes * fix: preserve review evidence and Autoplan phase handoffs * test: bind review evidence to owned decisions and completion * fix: retain owned native history across compaction * fix(evals): validate current review decisions and setup choices * fix: bind Autoplan reviews and phase completion to current amended input * fix: reconcile native review evidence and close Autoplan phases * test: recognize owned whole-candidate complexity decisions * test: preserve report freshness for approved investigation handoffs * fix: recognize scoped review findings and isolate dual voice fixtures * fix: make review handoffs and question dispatch self-contained * test: recognize complete CEO decisions and procedural pauses * fix: bind current CEO comparison options and risk intervals * test: bind engineering decisions and completion to owned evidence * fix: publish Autoplan phase reports before continuing tools * test: verify actual Autoplan dual-review dispatch evidence * test: select dual review when shared evidence fixtures change * fix: clarify plan review decisions and completion gates * fix: make CEO review decisions and return paths explicit * test: keep Autoplan prompt files inside attempt state * test: preserve source whitespace across permission dialog wraps * fix: publish Autoplan phase reports before continuing * test: recognize current CEO comparisons and reject inactive records * fix: reconcile engineering decision states before completion * test: recognize complete Design decisions and reports * test: verify current engineering decisions before navigation * Recognize source-owned component reduction choices * fix: recognize current CEO ledger and commitment grids * test: supply RequestPolicy context to Eng count fixture * fix: save complete engineering decisions before asking * fix: bind Autoplan publication to the complete phase readback * chore: prepare 1.87.5.0 reliability release * fix: clarify engineering review completion and preserve log failures * fix: bind CEO saved choices and current section ancestry * fix(evals): bind review execution and completion evidence * fix(plan-ceo-review): verify complete decisions before asking * fix(evals): preserve complete engineering choice records * fix(evals): preserve complete review outcomes and bounded fixtures * fix(autoplan): publish phase reports before advancing * fix(plan-ceo-review): validate option fields before asking * fix(plan-eng-review): verify current decisions after answers * fix(evals): bind review decisions and bound fixture scope * fix(plan-ceo-review): verify decision rows and edit saved checkpoints * fix(evals): bind review evidence and scope document lookup * fix(plan-eng-review): update resolution state with its answer * fix(reviews): preserve complete questions through dispatch * fix(evals): recognize completed mode declarations * fix(evals): define cache consistency at wrapper completion * fix(evals): validate owned initial scope and completed review handoffs * fix: assemble complete CEO decision fields before saving * fix: authenticate automatic mode decisions without guessing selectors * fix: bind engineering coverage to approved regression contracts * fix(evals): supply review helpers to native Eng capture * fix(plan-eng-review): preserve the full selected option scope * fix(evals): recognize owned engineering seed and regression evidence * fix(evals): bind engineering retry reports to native approvals * docs: clarify release guarantees (v1.87.5.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix(evals): recognize owned engineering decisions and handoffs * fix(evals): bind engineering decisions and completion evidence * fix(tests): align review contracts and selection fixtures * fix(skills): restore review prompt size limits * fix(plan-eng-review): clarify review execution and completion * fix(evals): preserve configured retries through all supervision layers * Clarify Engineering decisions and report completion * Keep native decision assertions within their source boundary * fix: recognize owned engineering decisions and completed navigation * fix: bind completed auto decisions to their current review * fix: recognize explicit CEO source attribution * fix: dispatch verified CEO decisions without recomposing fields * test: expose existing execution deadlines to review actors * fix: distinguish CEO decision records from incidental headings * test: bind split-scope choices to the registered native actor * test: connect reviewed regressions to required evaluation coverage * Clarify CEO decision routing and completion stages * test: expose existing section review deadlines to fixture actors * test: recognize complete native CEO pacing inventories * test: exclude answered history from current CEO payloads * test: detect phase entry through owned skill HOME aliases * test: validate native review completion and owned report permissions * fix: make Autoplan close packets carry the parent handoff steps * test: assess source-bound HOLD decisions within the existing deadline * fix: keep CEO native decision fields under one formatting authority * test: register integrated review and permission dependencies * test: align native review adapters and finding coverage Preserve explicit AUTO decisions, apply native single-select defaults, and bind complete cropped questions and report permissions to their owned requests. Require seeded review findings instead of crediting setup menus. Keep captured failure controls and additive selection dependencies. The integrated candidate passed 3,099 focused tests across 65 files; affected paid validation remains required before publication. * fix(autoplan): require phase reports before advancing * fix(evals): bind setup and evidence to complete attempts * fix(evals): bind native answers and pending writes to fixture scope Preserve complete option rows when native descriptions wrap, retain current owned Write arguments before journal publication, and keep engineering and DX answers within their declared fixture interfaces. Add captured free regressions without increasing model budgets or relaxing completion checks. * fix(autoplan): verify phase reports across native tool paths Guard owned methodology reads and reviewer dispatches, detect complete driver loads through Bash, and distinguish report-only edits from implementation changes. Follow authenticated native UUID ancestry when journal writes arrive out of order and verify earlier native content for cached phase reads. Keep current close acknowledgment and parent publication in order, require CEO entry before later phases, and register captured failure regressions. * fix(evals): honor native input and collection lifecycles Match complete native Edit panes and truncated question borders, reject stderr close before EOF, and stop the CEO split fixture once its acknowledged scope decisions are collected. Keep semantic validation, process failures, report requirements, and absolute deadlines authoritative. Add captured-event and real-process regressions with selection dependencies. Focused checks pass; final integrated paid and full-suite acceptance remain pending. * fix(autoplan): retain native session ownership across directory changes Recover missed native UUID ancestry through the existing strict graph while preserving ordinary event order and legacy scoping. Bind publication hooks to Claude's original project directory while retaining current cwd for requested file paths. Captured public-event regressions, existing caller checks, and a pinned native CLI loopback verify both fixes. Preserve failed attempts and require fresh paid and final full-suite acceptance. * docs: align evaluation limits and completion version * fix(autoplan): allow authenticated phase reads during journal streaming * fix(evals): bind clipped native questions and owned edit dialogs * fix: preserve overlay retries and bounded cleanup * fix: recognize owned planning preludes in native questions * docs: explain overlay scheduling and cleanup guarantees * fix: require fresh publication after Autoplan phase reruns * Release gstack 1.87.6 * fix: preserve CI paths, process identity, and test deadlines * fix: keep informational setup commands independent of install probes * fix: clarify plan review decisions and bound source audit reports * Fix remaining Windows identity and native path CI failures * Clarify CEO review decision and reviewer-result routing * test: accept no-install planner in retry supervision * fix(ceo-review): make review decisions and report completion explicit * perf(test): add fast PR gates, input-keyed judge reuse and isolated free shards * fix(test): start isolated CEO smoke from its existing project plan * fix(test): repair CI fixture races and preserve retry evidence * fix(ceo-review): clarify approvals, depth and saved completion --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
3695 lines
188 KiB
TypeScript
3695 lines
188 KiB
TypeScript
/**
|
||
* Deterministic unit tests for claude-pty-runner.ts behavior changes.
|
||
*
|
||
* Free-tier (no EVALS=1 needed). Runs in <1s on every `bun test`. Catches
|
||
* harness plumbing bugs before stochastic PTY runs surface them.
|
||
*
|
||
* Two surface areas tested:
|
||
*
|
||
* 1. Permission-dialog short-circuit in 'asked' classification: a TTY frame
|
||
* that matches BOTH isPermissionDialogVisible AND isNumberedOptionListVisible
|
||
* must NOT be classified as a skill question — permission dialogs render
|
||
* as numbered lists too, but they're not what we're guarding.
|
||
*
|
||
* 2. Env passthrough surface: runPlanSkillObservation accepts an `env`
|
||
* option and threads it to launchClaudePty. We can't fully exercise the
|
||
* spawn pipeline without paying for a PTY session, but we CAN verify the
|
||
* option exists in the type signature and that calling without env still
|
||
* works (no regression).
|
||
*
|
||
* The PTY test (skill-e2e-plan-ceo-plan-mode.test.ts) is the integration
|
||
* check; this file is the cheap deterministic guard for the harness primitives
|
||
* those tests stand on.
|
||
*/
|
||
|
||
import { describe, test, expect } from 'bun:test';
|
||
import { readFileSync } from 'node:fs';
|
||
import {
|
||
isPermissionDialogVisible,
|
||
isNumberedOptionListVisible,
|
||
isProseAUQVisible,
|
||
isScopeGateQuestionVisible,
|
||
isScopeGateAutoSelectVisible,
|
||
isPlanReadyVisible,
|
||
isAutoDecidedVisible,
|
||
parseNumberedOptions,
|
||
classifyVisible,
|
||
TAIL_SCAN_BYTES,
|
||
optionsSignature,
|
||
parseQuestionPrompt,
|
||
stripAnsi,
|
||
auqFingerprint,
|
||
COMPLETION_SUMMARY_RE,
|
||
MODE_RE,
|
||
findModeOption,
|
||
classifyPlanCountFrame,
|
||
capturePlanCountQuestion,
|
||
matchesNativePlanQuestion,
|
||
createPlanCountPermissionGuard,
|
||
planCountPrerequisitePick,
|
||
planCountSubmissionInput,
|
||
assertReviewReportAtBottom,
|
||
ceoStep0Boundary,
|
||
engStep0Boundary,
|
||
engSetupAUQ,
|
||
engFirstReviewAUQ,
|
||
designStep0Boundary,
|
||
designFirstReviewAUQ,
|
||
planCountQuestionPhase,
|
||
nativePlanCallFingerprint,
|
||
devexStep0Boundary,
|
||
type ClaudePtyOptions,
|
||
type AskUserQuestionFingerprint,
|
||
} from './claude-pty-runner';
|
||
|
||
describe('saved preference annotation', () => {
|
||
test('recognizes the explicit preference attribution from the timed-out CEO capture', () => {
|
||
const visible = 'Now I have a clear picture of the branch. Let me proceed with the full review. ' +
|
||
'Mode is HOLD SCOPE (auto-decided from plan-tune preference).';
|
||
expect(isAutoDecidedVisible(visible)).toBe(true);
|
||
expect(classifyVisible(visible)?.outcome).toBe('auto_decided');
|
||
expect(classifyVisible(visible.replace(/\s+/g, ''))?.outcome).toBe('auto_decided');
|
||
});
|
||
|
||
test('retains the canonical annotation and its precedence over plan-ready', () => {
|
||
const visible = 'Auto-decided review mode → HOLD SCOPE (your preference). Change with /plan-tune.\nReady to execute?';
|
||
expect(classifyVisible(visible)?.outcome).toBe('auto_decided');
|
||
});
|
||
|
||
test('does not equate an unrequested choice or plan-tune advice with a saved preference', () => {
|
||
for (const visible of [
|
||
'Mode is HOLD SCOPE (AUTO_DECIDED).',
|
||
'I auto-decided HOLD SCOPE because this is a refactor.',
|
||
'I auto-decided HOLD SCOPE. You can set a plan-tune preference later.',
|
||
'Mode is HOLD SCOPE (not auto-decided from plan-tune preference).',
|
||
'Mode is HOLD SCOPE (will be auto-decided from plan-tune preference).',
|
||
]) expect(isAutoDecidedVisible(visible)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('mode option rendering', () => {
|
||
test('letter-prefixed native mode labels retain their actual target indices', () => {
|
||
const options = ['C — HOLD SCOPE (Recommended)', 'B — SELECTIVE EXPANSION', 'A — SCOPE EXPANSION', 'D — SCOPE REDUCTION']
|
||
.map((label, i) => ({ index: i + 1, label }));
|
||
expect(options.every(option => MODE_RE.test(option.label))).toBe(true);
|
||
for (const [mode, index] of [['HOLD SCOPE', 1], ['SELECTIVE EXPANSION', 2], ['SCOPE EXPANSION', 3], ['SCOPE REDUCTION', 4]] as const) {
|
||
expect(findModeOption(options, mode)?.index).toBe(index);
|
||
}
|
||
expect(findModeOption(options.filter(option => option.index !== 3), 'SCOPE EXPANSION')).toBeUndefined();
|
||
});
|
||
test('letter-prefixed matching excludes prose, unrelated choices and unsupported framing', () => {
|
||
for (const label of ['Choose C — HOLD SCOPE', 'Approach C — HOLD SCOPE', 'C — Keep this approach\nHOLD SCOPE',
|
||
'B — Ideal Architecture (Recommended)', 'A — Fix-Only (Minimal Viable)', 'CC — HOLD SCOPE', 'E — HOLD SCOPE',
|
||
'1 — HOLD SCOPE', 'C: HOLD SCOPE', 'C - HOLD SCOPE', 'C — SCOPE EXPANSIONIST']) {
|
||
expect(MODE_RE.test(label), label).toBe(false);
|
||
expect(findModeOption([{ index: 1, label }], 'HOLD SCOPE'), label).toBeUndefined();
|
||
}
|
||
expect(findModeOption([{ index: 1, label: 'C — HOLD SCOPE\nPrefer this over A — SCOPE EXPANSION.' }], 'SCOPE EXPANSION')).toBeUndefined();
|
||
});
|
||
test('parenthesized native modes preserve capture group and actual target indices', () => {
|
||
const options = ['A) SCOPE EXPANSION', 'B) SELECTIVE EXPANSION (recommended)', 'C) HOLD SCOPE', 'D) SCOPE REDUCTION']
|
||
.map((label, i) => ({ index: i + 1, label }));
|
||
for (const [mode, index] of [['SCOPE EXPANSION', 1], ['SELECTIVE EXPANSION', 2], ['HOLD SCOPE', 3], ['SCOPE REDUCTION', 4]] as const) {
|
||
expect(MODE_RE.exec(options[index - 1]!.label)?.[1]).toBe(mode);
|
||
expect(findModeOption(options, mode)?.index).toBe(index);
|
||
}
|
||
expect(findModeOption(options.filter(option => option.index !== 1), 'SCOPE EXPANSION')).toBeUndefined();
|
||
expect(findModeOption([{ index: 3, label: '**C) HOLD SCOPE**' }], 'HOLD SCOPE')?.index).toBe(3);
|
||
});
|
||
test('parenthesized mode recognition rejects prose, descriptions and unrelated framing', () => {
|
||
for (const label of ['Discuss C) HOLD SCOPE', 'Approach C) HOLD SCOPE', 'C) Keep this approach\nHOLD SCOPE',
|
||
'B) Ideal Architecture (Recommended)', 'A) Fix-Only (Minimal Viable)', 'CC) HOLD SCOPE', 'E) HOLD SCOPE',
|
||
'1) HOLD SCOPE', '(C) HOLD SCOPE', 'C: HOLD SCOPE', 'C - HOLD SCOPE', 'C) SCOPE EXPANSIONIST']) {
|
||
expect(MODE_RE.test(label), label).toBe(false);
|
||
expect(findModeOption([{ index: 1, label }], 'HOLD SCOPE'), label).toBeUndefined();
|
||
}
|
||
expect(findModeOption([{ index: 1, label: 'C) HOLD SCOPE\nPrefer this over A) SCOPE EXPANSION.' }], 'SCOPE EXPANSION')).toBeUndefined();
|
||
});
|
||
test('selects the actual collapsed-space mode from the failed periodic menu', () => {
|
||
const options = [
|
||
{ index: 1, label: 'HOLDSCOPE(recommended)\rMake the reliability wave bulletproof.' },
|
||
{ index: 2, label: 'SELECTIVEEXPANSION\rKeep the current scope as the baseline.' },
|
||
{ index: 3, label: 'SCOPEREDUCTION\rFind the minimum subset.' },
|
||
{ index: 4, label: 'SCOPEEXPANSION\rDream up adjacent reliability improvements.' },
|
||
{ index: 5, label: 'Type something.' },
|
||
{ index: 6, label: 'Chat about this\rUser answered → HOLD SCOPE (recommended)' },
|
||
];
|
||
expect(options.slice(0, 4).every(option => MODE_RE.test(option.label))).toBe(true);
|
||
expect(findModeOption(options, 'SCOPE EXPANSION')?.index).toBe(4);
|
||
expect(findModeOption(options, 'HOLD SCOPE')?.index).toBe(1);
|
||
});
|
||
|
||
test('retains ordinary, wrapped, and emphasized mode labels', () => {
|
||
for (const label of ['SCOPE EXPANSION (Recommended)', 'scope\t expansion', 'SCOPE\r\nEXPANSION', '**SCOPE EXPANSION**']) {
|
||
expect(findModeOption([{ index: 2, label }], 'SCOPE EXPANSION')?.index).toBe(2);
|
||
}
|
||
});
|
||
|
||
test('an omitted target remains missing, including when another mode mentions it', () => {
|
||
const options = [
|
||
{ index: 1, label: 'HOLD SCOPE\rPrefer this over SCOPE EXPANSION.' },
|
||
{ index: 2, label: 'SELECTIVE EXPANSION' },
|
||
{ index: 3, label: 'SCOPE REDUCTION' },
|
||
{ index: 4, label: 'Chat about this\rSCOPEEXPANSION (old screen)' },
|
||
];
|
||
expect(findModeOption(options, 'SCOPE EXPANSION')).toBeUndefined();
|
||
expect(MODE_RE.test(options[3]!.label)).toBe(false);
|
||
expect(MODE_RE.test('Scope expansionist')).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('isPermissionDialogVisible', () => {
|
||
test('matches "Bash command requires permission" prompts', () => {
|
||
const sample = `
|
||
Some preamble output
|
||
|
||
Bash command \`gstack-config get telemetry\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. Yes, and always allow
|
||
3. No, abort
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches "allow all edits" file-edit prompts', () => {
|
||
// Isolated to the "allow all edits" clause only — no overlapping
|
||
// "Do you want to proceed?" co-trigger, so this asserts the clause works.
|
||
const sample = `
|
||
Edit to ~/.gstack/config.yaml
|
||
|
||
❯ 1. Yes
|
||
2. Yes, allow all edits during this session
|
||
3. No
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the "Do you want to proceed?" file-edit confirmation by itself', () => {
|
||
// Separate fixture so weakening this clause is detected by a dedicated test.
|
||
const sample = `
|
||
Edit to ~/.gstack/config.yaml
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches workspace-trust "always allow access to" prompt', () => {
|
||
const sample = `
|
||
Do you trust the files in this folder?
|
||
|
||
❯ 1. Yes, proceed
|
||
2. Yes, and always allow access to /Users/me/repo
|
||
3. No, exit
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('recognizes the captured collapsed native overwrite confirmation', () => {
|
||
const sample = [
|
||
'Doyouwanttooverwritegstack-test-plan-design.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
expect(isPermissionDialogVisible(sample.replace('Esctocancel·Tabtoamend', 'Enter to select'))).toBe(false);
|
||
});
|
||
|
||
test('the captured paired-CEO Edit grant is permission, not another review finding', () => {
|
||
const sample = [
|
||
'Do youwt to makehis dittogstack-test-plan-ceo-paired.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/gstack-paid-shard-EbUl9j/tmp/gstack-e2e-plan-ceo-paired-gkjAd5forthissession',
|
||
'(shift+tab)', '3.No', 'Esctocancel·Tabtoamend',
|
||
].join('\r');
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
expect(classifyPlanCountFrame(sample)).toBe('permission');
|
||
});
|
||
|
||
test('recognizes permission labels whose cursor-positioning spaces disappeared', () => {
|
||
expect(isPermissionDialogVisible('Yes,andalwaysallowaccessto/tmp/fixtureforthissession')).toBe(true);
|
||
expect(isPermissionDialogVisible('Yes,allowalleditsduringthissession')).toBe(true);
|
||
expect(isPermissionDialogVisible('Bashcommandrequirespermission')).toBe(true);
|
||
});
|
||
|
||
test('does NOT match a skill AskUserQuestion list', () => {
|
||
const sample = `
|
||
D1 — Premise challenge: do users actually want this?
|
||
|
||
❯ 1. Yes, validated
|
||
2. No, premise is wrong
|
||
3. Need more info
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT match a plan-ready confirmation', () => {
|
||
const sample = `
|
||
Ready to execute the plan?
|
||
|
||
❯ 1. Yes
|
||
2. No, keep planning
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT match a skill question that contains the bare phrase "Do you want to proceed?"', () => {
|
||
// Co-trigger requirement: "Do you want to proceed?" alone is not enough.
|
||
// It must appear with "Edit to <path>" or "Write to <path>" to count as
|
||
// a permission dialog. This guards against a skill question like
|
||
// "Do you want to proceed with HOLD SCOPE?" being mis-classified.
|
||
const sample = `
|
||
Choose your scope mode for this review.
|
||
Do you want to proceed?
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
3. SELECTIVE EXPANSION
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT mis-match when adversarial prose includes "Edit to <path>" alongside the bare proceed phrase', () => {
|
||
// Adversarial fixture: a skill question whose body legitimately mentions
|
||
// "Edit to <path>" in prose AND ends with "Do you want to proceed?". The
|
||
// current co-trigger regex would mis-classify this as a permission
|
||
// dialog. We DO want this test to fail until the regex is tightened
|
||
// further (e.g., proximity constraint, or anchoring "Edit to" to a
|
||
// line-start). For now this is documented as a known limitation: a
|
||
// skill question that talks about "Edit to" in prose IS still treated
|
||
// as a permission dialog. The test asserts the current behavior so a
|
||
// future fix can flip it intentionally.
|
||
const sample = `
|
||
Plan: I will Edit to ./plan.md to capture the decision.
|
||
Do you want to proceed?
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
// KNOWN LIMITATION: the co-trigger fires here. Documented as a
|
||
// post-merge follow-up. Flip this assertion once the regex tightens.
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
});
|
||
|
||
describe('isNumberedOptionListVisible', () => {
|
||
test('matches a basic ❯ 1. + 2. cursor list', () => {
|
||
const sample = `
|
||
❯ 1. Option one
|
||
2. Option two
|
||
3. Option three
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on a single-option prompt', () => {
|
||
const sample = `
|
||
❯ 1. Only option
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('returns false when no cursor renders', () => {
|
||
const sample = `
|
||
Just some prose with 1. a numbered point and 2. another.
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('overlaps permission dialogs (this is why D5 short-circuits)', () => {
|
||
// The whole point of D5: this string matches BOTH classifiers, so the
|
||
// runner must consult isPermissionDialogVisible to disambiguate.
|
||
const sample = `
|
||
Bash command \`do-thing\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(true);
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
});
|
||
|
||
describe('scope-gate render detectors', () => {
|
||
// The verbatim announcement string from the plan-eng/plan-design SKILL.md
|
||
// templates. If the template rewording drifts, THIS fixture fails first —
|
||
// before the paid plan-mode smokes silently degrade to vacuous asserts.
|
||
const TEMPLATE_ANNOUNCEMENT =
|
||
'Scope gate: plan mode — auto-selected B (reviewing <target>).';
|
||
|
||
describe('isScopeGateQuestionVisible', () => {
|
||
test('matches the clean prose gate render (question + option bodies)', () => {
|
||
const sample = `
|
||
What should I review?
|
||
A) The current branch diff — the work in progress on this branch.
|
||
B) A plan or design doc I'll paste or point you to.
|
||
C) A specific file, directory, or path.
|
||
Recommendation: A when a branch diff exists, otherwise B.
|
||
`;
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the native numbered render (no lettered markers)', () => {
|
||
const sample = `
|
||
What should I review?
|
||
|
||
❯ 1. The current branch diff — the work in progress on this branch.
|
||
2. A plan or design doc I'll paste or point you to.
|
||
3. A specific file, directory, or path.
|
||
`;
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the PTY-collapsed render (stripAnsi squished spaces)', () => {
|
||
const sample = 'WhatshouldIreview?A)Thecurrentbranchdiff—theworkinprogress';
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('stays false on narration quoting only the question', () => {
|
||
const sample =
|
||
"Normally I'd ask 'What should I review?' but plan mode is active, so I'm proceeding.";
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on unrelated review prose', () => {
|
||
const sample = 'I will review the current branch diff and report findings.';
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('isScopeGateAutoSelectVisible', () => {
|
||
test('matches the verbatim template announcement', () => {
|
||
expect(isScopeGateAutoSelectVisible(TEMPLATE_ANNOUNCEMENT)).toBe(true);
|
||
});
|
||
|
||
test('matches a real announcement with a concrete target', () => {
|
||
const sample =
|
||
'Scope gate: plan mode — auto-selected B (reviewing ~/.claude/plans/my-feature.md). Running the Design Doc Check next.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the PTY-collapsed announcement', () => {
|
||
const sample = 'Scopegate:planmode—auto-selectedB(reviewingPLAN.md).';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('stays false on narration about the behavior', () => {
|
||
const sample = "In plan mode I'd auto-select B and review the active plan.";
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on a VERBATIM QUOTE of the announcement (negation narration)', () => {
|
||
// The exact announcement line sits quoted in the skill context, so a
|
||
// model explaining why it is NOT firing it can reproduce it byte-exact
|
||
// inside quotes — that must not trip a must-stay-false assert.
|
||
const sample =
|
||
'Not in plan mode, so I won\'t announce "Scope gate: plan mode — auto-selected B (reviewing <target>)." and will ask instead.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('a later real render still matches after an earlier quoted mention', () => {
|
||
const sample =
|
||
'Earlier I said I would render "Scope gate: plan mode — auto-selected B (…)" and now:\n' +
|
||
'Scope gate: plan mode — auto-selected B (reviewing PLAN.md).';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches tense paraphrases WITH the announcement prefix (auto-selecting / auto-selects)', () => {
|
||
expect(
|
||
isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selecting B (reviewing the drafted plan).'),
|
||
).toBe(true);
|
||
expect(isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selects B.')).toBe(true);
|
||
});
|
||
|
||
test('stays false on tense paraphrases WITHOUT the announcement prefix', () => {
|
||
expect(isScopeGateAutoSelectVisible('Auto-selecting B since we are in plan mode.')).toBe(false);
|
||
});
|
||
|
||
test('stays false on AUTO_DECIDE preamble output', () => {
|
||
const sample = 'Auto-decided scope question → B (your preference). Change with /plan-tune.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on a bare "selected B" without the announcement prefix', () => {
|
||
const sample = 'I selected B as the review target.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('isProseAUQVisible', () => {
|
||
test('matches 4 lettered options A) B) C) D) at line starts (plan-eng prose AUQ shape)', () => {
|
||
const sample = `
|
||
What would you like me to review? Options:
|
||
A) Point me at an existing design doc or plan file (path).
|
||
B) Describe new work you're planning — I'll explore the codebase.
|
||
C) You meant /review for the diff already on this branch.
|
||
D) Something else (tell me).
|
||
Recommendation: A if you have a doc in mind, otherwise B.
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches 2 lettered options (minimum threshold)', () => {
|
||
const sample = `
|
||
A) First option
|
||
B) Second option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches 3 numbered options 1. 2. 3. without ❯ 1. cursor (autoplan prose AUQ shape)', () => {
|
||
const sample = `
|
||
What's the task? A few options:
|
||
1. You have a plan idea in mind — describe it.
|
||
2. You want to review an existing plan elsewhere.
|
||
3. You meant a different command — /plan-ceo-review etc.
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false when ❯ 1. cursor is present in the recent tail (native UI handled by isNumberedOptionListVisible)', () => {
|
||
const sample = `
|
||
❯ 1. First option
|
||
2. Second option
|
||
3. Third option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT suppress numbered-prose detection when ❯ 1. is only in early scrollback (trust dialog)', () => {
|
||
// Boot trust dialog rendered ❯ 1. Yes at startup, then a long body of
|
||
// model output, then prose-rendered numbered options now. The historic
|
||
// ❯ 1. is in the full buffer but NOT in the recent tail. Should detect
|
||
// the prose AUQ.
|
||
const trustHeader = '❯ 1. Yes, trust\n 2. No\n';
|
||
const filler = 'x'.repeat(5000); // pushes trust dialog out of last 4KB tail
|
||
const proseAUQ = `\n 1. Review the docs\n 2. Investigate the code\n 3. Defer to next session\n❯ \n`;
|
||
const sample = trustHeader + filler + proseAUQ;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on single lettered option', () => {
|
||
const sample = `
|
||
A) Only one option mentioned in passing.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('matches 2 numbered options (threshold matches lettered branch — tails miss option 1)', () => {
|
||
const sample = `
|
||
1. First note.
|
||
2. Second note.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on a single numbered option', () => {
|
||
const sample = `
|
||
1. Only one option mentioned.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does not match mid-prose lettered text like "(see option B) above"', () => {
|
||
const sample = `
|
||
This refers to (see option B) above and also to point A) earlier.
|
||
`;
|
||
// The B) and A) markers are mid-line, not at line starts, so they don't count.
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('matches with leading whitespace and ❯ prefix on options', () => {
|
||
const sample = `
|
||
A) Option with whitespace prefix
|
||
❯ B) Option with cursor prefix
|
||
C) Another option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on plain text with no option markers', () => {
|
||
expect(isProseAUQVisible('Just some plain text output from the model.')).toBe(false);
|
||
expect(isProseAUQVisible('')).toBe(false);
|
||
});
|
||
|
||
// Pattern 3: markdown bold-bullet options — office-hours renders its mode
|
||
// question this way under --disallowedTools, with no letter/number marker.
|
||
test('matches office-hours markdown bold-bullet mode question (Pattern 3)', () => {
|
||
const sample = `
|
||
> Before we dig in — what's your goal with this?
|
||
>
|
||
> - **Building a startup** (or thinking about it)
|
||
> - **Intrapreneurship** — internal project at a company, need to ship fast
|
||
> - **Hackathon / demo** — time-boxed, need to impress
|
||
> - **Open source / research** — building for a community
|
||
> - **Learning** — teaching yourself to code
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('bold-bullets require a preceding interrogative — no "?" => false', () => {
|
||
// 3+ bold bullets but no question stem: this is a feature list, not an AUQ.
|
||
const sample = `
|
||
Here is what shipped:
|
||
- **Faster builds** via caching
|
||
- **Smaller binaries** through tree-shaking
|
||
- **Better errors** with source maps
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('a question with fewer than 3 bold bullets stays false (guard)', () => {
|
||
const sample = `
|
||
Which approach do you prefer?
|
||
- **Option one** is simpler
|
||
- **Option two** is faster
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('plain (non-bold) bullets after a question do not trigger Pattern 3', () => {
|
||
// Only bold bullets count — plain "- text" prose lists are too common.
|
||
const sample = `
|
||
What should we do about this?
|
||
- run the tests
|
||
- ship the fix
|
||
- file a follow-up
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('Pattern 3 still defers to a live native cursor list (❯ 1.)', () => {
|
||
const sample = `
|
||
> What's your goal?
|
||
❯ 1. **Building a startup**
|
||
2. **Intrapreneurship**
|
||
3. **Hackathon**
|
||
`;
|
||
// The ❯1. cursor gate fires first — native list handling owns this.
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
// Pattern 4/5: collapsed-form prose AUQ. stripAnsi destroys the newlines +
|
||
// inter-word spaces, so a real prose AUQ arrives collapsed and defeats the
|
||
// line-anchored Patterns 1-3. These are the dominant Shape-B render mode in
|
||
// the plan-design smoke + floor timeouts — verbatim de-spinnered bytes from
|
||
// the real failing runs (bdm3sucql.output).
|
||
test('matches the real collapsed floor render (colon-delimited, Pattern 4/5)', () => {
|
||
const sample =
|
||
'The review is blocked on D1—reply withA, B, r Cabovetocontinue:' +
|
||
'- A(recommended): Spec thefull P1AskUserQuestioncopy in this review' +
|
||
'-B:LeaveP1copytotheimplementerwithstructuralrequirements' +
|
||
'C: Add a placeholder template to the plan';
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the real collapsed plan-mode render (Recommendation + collapsed A)/B), Pattern 4/5)', () => {
|
||
const sample =
|
||
'Recommendation:A—writethecopynow.(recommended)A) Writ the fullcopy in thisdesign review— now.' +
|
||
'(recommended) Completeness:10/10 B) Leveit to theimplemente — task spec is enough.' +
|
||
'Reply withA (write the copy now)orB(leavetoimplementer)';
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('collapsed-form requires BOTH signals — single B) + word "recommendation" stays false', () => {
|
||
// Only one punctuated letter marker: the two-signal contract is not met.
|
||
const sample =
|
||
'We should consider option B) here. My recommendation is to do it now.';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('collapsed-form requires letter punctuation — comma-only "ReplywithA,B,orC" stays false', () => {
|
||
// Reply-instruction present, but the letters carry no ) : or ( punctuation,
|
||
// so they could be incidental enumerations in running prose. Stays false.
|
||
const sample = 'ReplywithA,B,orC';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('collapsed-form does not regress the existing FP guard (see option B) ... point A))', () => {
|
||
// The classic citation FP: a model referencing prior options in prose.
|
||
// No reply-instruction / recommendation marker on its own line, so the
|
||
// collapsed-form signal does not fire either.
|
||
const sample =
|
||
'As noted (see option B) above, and the earlier point A) we discussed, this is fine.';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('classifyVisible (runtime path through the runner classifier)', () => {
|
||
// These tests call the actual classifier so a future contributor who
|
||
// reorders branches (e.g. moves the permission short-circuit before
|
||
// isPlanReadyVisible) is caught deterministically.
|
||
|
||
test('skill question → returns asked', () => {
|
||
const visible = `
|
||
D1 — Choose your scope mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
3. SELECTIVE EXPANSION
|
||
4. SCOPE REDUCTION
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('permission dialog (Bash) → returns null (skip, keep polling)', () => {
|
||
const visible = `
|
||
Bash command \`gstack-update-check\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isNumberedOptionListVisible(visible)).toBe(true); // pre-filter
|
||
expect(classifyVisible(visible)).toBeNull(); // post-filter
|
||
});
|
||
|
||
test('plan-ready confirmation → returns plan_ready (wins over asked)', () => {
|
||
const visible = `
|
||
Ready to execute the plan?
|
||
|
||
❯ 1. Yes, proceed
|
||
2. No, keep planning
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('plan_ready');
|
||
});
|
||
|
||
test('silent write to unsanctioned path → returns silent_write', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/dangerous-write.ts)
|
||
⎿ Wrote 42 lines
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('silent_write');
|
||
expect(result?.summary).toContain('src/app/dangerous-write.ts');
|
||
});
|
||
|
||
test('write to sanctioned path (.claude/plans) → returns null (allowed)', () => {
|
||
const visible = `
|
||
⏺ Write(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Wrote 42 lines
|
||
`;
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('write while a permission dialog is on screen → returns null (gated, not silent, not asked)', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/edit-with-permission.ts)
|
||
|
||
Edit to src/app/edit-with-permission.ts
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
// The numbered prompt is a permission dialog (Edit to + Do you want to proceed?);
|
||
// silent_write is suppressed because a numbered prompt is visible, AND
|
||
// 'asked' is suppressed because the prompt is a permission dialog.
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('write while a real skill question is on screen → returns asked (write is captured but not silent)', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/foo.ts)
|
||
|
||
D1 — Choose your scope mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
// The numbered prompt is a skill question, not a permission dialog;
|
||
// silent_write is suppressed (numbered prompt is visible) and the
|
||
// outcome is 'asked' — Step 0 fired.
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('idle / no signals → returns null', () => {
|
||
const visible = `
|
||
Some prose without any classifier signals.
|
||
`;
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('TAIL_SCAN_BYTES is exported as 1500', () => {
|
||
// Shared between runner and routing test; a regression that desyncs the
|
||
// recent-tail window would surface here.
|
||
expect(TAIL_SCAN_BYTES).toBe(1500);
|
||
});
|
||
|
||
// D4-B: strictPlanWrites detector. Catches the transcript bug where the
|
||
// model writes findings to the plan file before any AskUserQuestion fires.
|
||
test('strictPlanWrites: plan write before any AUQ → wrote_findings_before_asking', () => {
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('wrote_findings_before_asking');
|
||
expect(result?.summary).toContain('.claude/plans/some-plan.md');
|
||
});
|
||
|
||
test('strictPlanWrites: plan write AFTER an AUQ render → not flagged', () => {
|
||
// AUQ renders first, then the model writes the plan post-answer. This is
|
||
// the legitimate end-of-workflow flow and must NOT trigger the detector.
|
||
const visible = `
|
||
D1 — Some scope question
|
||
|
||
❯ 1. Option A
|
||
2. Option B
|
||
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
// Outcome is 'asked' (the numbered list rendered); the post-AUQ plan
|
||
// write is ignored by the detector.
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('strictPlanWrites: AUQ first then plan write — write_pos > auq_pos → not flagged', () => {
|
||
// Same scenario, more explicit ordering: the regex finds the write at a
|
||
// position AFTER the numbered list. Detector lets it through.
|
||
const visible = [
|
||
'D1 — Choose your approach',
|
||
'',
|
||
'❯ 1. Approach A',
|
||
' 2. Approach B',
|
||
'',
|
||
'⏺ Write(/Users/me/.claude/plans/draft.md)',
|
||
'⎿ Wrote 42 lines',
|
||
].join('\n');
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('strictPlanWrites: only a permission dialog visible → plan write still flagged', () => {
|
||
// A permission dialog ❯ 1./2. is NOT an AUQ; pre-AUQ plan writes still
|
||
// hit the detector even when a permission prompt is on screen.
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
|
||
Edit to /Users/me/.claude/plans/some-plan.md
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('wrote_findings_before_asking');
|
||
});
|
||
|
||
test('strictPlanWrites OFF: plan write before AUQ → returns null (legacy behavior preserved)', () => {
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
// Without strictPlanWrites, the sanctioned-path list lets this through.
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('parseNumberedOptions', () => {
|
||
test('does not combine an old AUQ prompt with the later ordinary test-case list', () => {
|
||
// B CEO retry, 2026-09-08: the old prompt cursor slid outside the
|
||
// option parser's 4KB window. Its prose fallback then supplied a new
|
||
// five-item test list while the prompt parser retained the old AUQ.
|
||
const visible = '☐Stripe event types\nWhich event should the handler accept?\n' +
|
||
'❯1.Specify one canonical event\n2.Accept all events\n' + '·'.repeat(4200) + '\n' +
|
||
'Minimum required test cases (all must be specified in the plan):\n' +
|
||
'1.Happypath:validcanonicalevent,knownuser→userupdated,emailsent\n' +
|
||
'2.Email failure:emailthrows→userupdated,errorlogged,HTTP200\n' +
|
||
'3.DB timeout: DB throws onuser update →exceptin ropagates, non-200\n' +
|
||
'4.Unkown event typ: non-canonical event→ HTTP200,nouserupdate\n' +
|
||
'5.Unknown user: valid event, usernotinDB→existingguard→HTTP200\n❯1\n';
|
||
const seen = new Set<string>();
|
||
expect(capturePlanCountQuestion(visible, seen, 0, false)).toBeNull();
|
||
expect(seen.size).toBe(0);
|
||
});
|
||
|
||
test('extracts options from a clean cursor list', () => {
|
||
const visible = `
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
const opts = parseNumberedOptions(visible);
|
||
expect(opts).toHaveLength(2);
|
||
expect(opts[0]).toEqual({ index: 1, label: 'HOLD SCOPE' });
|
||
expect(opts[1]).toEqual({ index: 2, label: 'SCOPE EXPANSION' });
|
||
});
|
||
|
||
test('returns empty array on prose-with-numbers (no cursor)', () => {
|
||
expect(parseNumberedOptions('text 1. one 2. two')).toEqual([]);
|
||
});
|
||
|
||
test('extracts options when the cursor is INLINE with prompt header (box-layout)', () => {
|
||
// Real /plan-ceo-review rendering: the TTY's cursor-positioning escapes
|
||
// collapse divider + header + prompt + cursor onto one logical line.
|
||
// Subsequent options (2..7) still start their own lines.
|
||
const visible = [
|
||
'────────────────────────────────────────',
|
||
'☐ Review scope What scope do you want me to CEO-review? ❯ 1. The branch\'s diff vs main',
|
||
' Review the full branch: ~10K LOC.',
|
||
'2. A specific plan file or design doc',
|
||
' You point me at a file (path) and I review that.',
|
||
'3. An idea you\'ll describe inline',
|
||
'4. Cancel — wrong skill',
|
||
'5. Type something.',
|
||
'────────────────────────────────────────',
|
||
'6. Chat about this',
|
||
'7. Skip interview and plan immediately',
|
||
].join('\n');
|
||
const opts = parseNumberedOptions(visible);
|
||
expect(opts).toHaveLength(7);
|
||
expect(opts[0]).toEqual({ index: 1, label: "The branch's diff vs main" });
|
||
expect(opts[1]?.index).toBe(2);
|
||
expect(opts[6]?.index).toBe(7);
|
||
expect(opts[6]?.label).toBe('Skip interview and plan immediately');
|
||
});
|
||
|
||
test('inline-cursor and start-of-line cursor both produce 7 options for the box-layout case', () => {
|
||
// The inline path captures option 1 from the cursor line itself; the
|
||
// subsequent-lines path captures 2..7 with the existing optionRe.
|
||
const inlineLayout = [
|
||
'header text ❯ 1. first option',
|
||
'2. second',
|
||
'3. third',
|
||
].join('\n');
|
||
expect(parseNumberedOptions(inlineLayout)).toEqual([
|
||
{ index: 1, label: 'first option' },
|
||
{ index: 2, label: 'second' },
|
||
{ index: 3, label: 'third' },
|
||
]);
|
||
|
||
const cleanLayout = [
|
||
' ❯ 1. first option',
|
||
' 2. second',
|
||
' 3. third',
|
||
].join('\n');
|
||
expect(parseNumberedOptions(cleanLayout)).toEqual([
|
||
{ index: 1, label: 'first option' },
|
||
{ index: 2, label: 'second' },
|
||
{ index: 3, label: 'third' },
|
||
]);
|
||
});
|
||
});
|
||
|
||
describe('pending native question on a damaged option render', () => {
|
||
// Exact final B CEO Test scope shape. The native call had been read in
|
||
// an in-progress snapshot, but option 2's missing dot prevented input.
|
||
const frame = [
|
||
'☐Test scope',
|
||
'│Section 6 (Tests) — Theplanhasnotestsforanewpaymentprocessingcodepath.Theexistingintegrationsuitehas',
|
||
'│never seen this handlerand cannotcatchregressionsinit.Minimumviabletestplanforminimalpatch:5unittests',
|
||
'│(happy path, mal failur, DB timeout, unknowneventtype,unknownuser).Shouldtheplanalsoincludeanintegration',
|
||
'│testhittingthefullwebhookstack?<gstack-qid:plan-ceo-test-scope>',
|
||
'❯1.Unittestsonlyfornow(recommended)',
|
||
'5 unit tests covering the criticalpaths. No integration stin v1.',
|
||
'2Uni tsts + one integration test',
|
||
'5 uit tsts + on ed-to-end integrationtestsendiga signe Stripeevent.',
|
||
'3.Integrationtestonly',
|
||
'4.Typesomething.',
|
||
'5. Chataboutthis',
|
||
'Enter to select · ↑/↓ to navigate · Esc to cancel',
|
||
'❯1',
|
||
].join('\n');
|
||
const pending = {
|
||
sessionId: '66fb6218-4a68-4f1a-a729-6407f14fd6b8',
|
||
toolUseId: 'toolu_017DicePqWNVyDsLCd2Y2MCi', answered: false,
|
||
questions: [{ header: 'Test scope', question: 'Should the plan also include an integration test hitting the full webhook stack? <gstack-qid:plan-ceo-test-scope>',
|
||
options: ['Unit tests only for now (recommended)', 'Unit tests + one integration test', 'Integration test only'].map(label => ({ label })) }],
|
||
};
|
||
|
||
test('uses lossless pending options after a positively matched native question has rendered', () => {
|
||
const seen = new Set<string>();
|
||
const captured = capturePlanCountQuestion(frame, seen, 0, false, pending);
|
||
expect(captured?.nativeCall).toBe(pending);
|
||
expect(captured?.options).toEqual(pending.questions[0].options.map((o, i) => ({ index: i + 1, label: o.label })));
|
||
expect(capturePlanCountQuestion(frame, seen, 1, false, pending)).toBeNull();
|
||
// A corrected redraw is still the same pending native question.
|
||
expect(capturePlanCountQuestion(frame.replace('2Uni tsts', '2.Unit tests'), seen, 2, false, pending)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame.replace('2Uni tsts', '2.Unit tests'), seen, 3, false)).toBeNull();
|
||
expect(seen.has(captured!.signature)).toBe(true);
|
||
});
|
||
|
||
test('binds delayed native metadata to the already-answered visible question', () => {
|
||
const seen = new Set<string>();
|
||
const clean = frame.replace('2Uni tsts', '2.Unit tests');
|
||
expect(capturePlanCountQuestion(clean, seen, 0, false)).not.toBeNull();
|
||
expect(capturePlanCountQuestion(clean, seen, 1, false, pending)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, seen, 2, false, pending)).toBeNull();
|
||
});
|
||
|
||
test('requires pending single-question metadata, matching current header, cursor, and navigation footer', () => {
|
||
for (const call of [undefined, { ...pending, answered: true }, { ...pending, failed: true },
|
||
{ ...pending, questions: [...pending.questions, ...pending.questions] },
|
||
{ ...pending, questions: [{ ...pending.questions[0], header: 'Prior decision' }] },
|
||
{ ...pending, questions: [{ ...pending.questions[0], question: 'Different issue <gstack-qid:plan-ceo-different-test-scope>' }] }]) {
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, call)).toBeNull();
|
||
}
|
||
for (const altered of [frame.replace('☐Test scope', 'Test scope'), frame.replace('❯1.', '1.'),
|
||
frame.replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
|
||
frame + '\n☐Different question\n❯1.Waiting for its choices']) {
|
||
expect(capturePlanCountQuestion(altered, new Set(), 0, false, pending)).toBeNull();
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('runPlanSkillObservation env passthrough surface', () => {
|
||
test('ClaudePtyOptions exposes env: Record<string, string>', () => {
|
||
// Type-level guard: this file would fail to compile if the env field
|
||
// were removed or its shape regressed. The actual env merge happens in
|
||
// launchClaudePty's spawn call (`env: { ...process.env, ...opts.env }`),
|
||
// so a regression where `env: opts.env` gets dropped from the
|
||
// runPlanSkillObservation -> launchClaudePty handoff is only caught by
|
||
// the live PTY test, not here.
|
||
const opts: ClaudePtyOptions = {
|
||
env: { QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' },
|
||
};
|
||
expect(opts.env).toEqual({ QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' });
|
||
});
|
||
});
|
||
|
||
describe('launchClaudePty model pin (static tripwire)', () => {
|
||
// Why static-grep, not a behavioral assert: the spawn fires immediately
|
||
// inside launchClaudePty, so asserting the built args array would require
|
||
// extracting an arg-builder seam — which rewrites the exact region kyoto-v5's
|
||
// hermetic --strict-mcp-config insertion edits, reintroducing a merge
|
||
// conflict the placement deliberately avoids. The end-to-end behavioral proof
|
||
// is the live PTY smoke (skill-e2e-plan-*-plan-mode.test.ts) running under the
|
||
// pinned model. These grep-level guards stop a refactor from silently
|
||
// dropping the pin or reordering it past extraArgs.
|
||
const src = readFileSync(new URL('./claude-pty-runner.ts', import.meta.url), 'utf-8');
|
||
|
||
test('ClaudePtyOptions exposes model?: string', () => {
|
||
const opts: ClaudePtyOptions = { model: 'claude-sonnet-4-6' };
|
||
expect(opts.model).toBe('claude-sonnet-4-6');
|
||
});
|
||
|
||
test('spawn args push --model from the EVALS_MODEL fallback chain', () => {
|
||
expect(src).toContain("args.push('--model', model)");
|
||
// opts.model -> EVALS_MODEL -> resolveEvalModel('capture') (mirrors session-runner.ts)
|
||
expect(src).toMatch(
|
||
/opts\.model\s*\?\?\s*process\.env\.EVALS_MODEL\s*\?\?\s*resolveEvalModel\('capture'\)/,
|
||
);
|
||
});
|
||
|
||
test('--model is pushed BEFORE extraArgs so a per-test --model override wins', () => {
|
||
const modelPush = src.indexOf("args.push('--model', model)");
|
||
const extraArgsPush = src.indexOf('if (opts.extraArgs) args.push(...opts.extraArgs)');
|
||
expect(modelPush).toBeGreaterThan(-1);
|
||
expect(extraArgsPush).toBeGreaterThan(-1);
|
||
expect(modelPush).toBeLessThan(extraArgsPush);
|
||
});
|
||
|
||
test('all three plan-skill wrappers forward model to launchClaudePty', () => {
|
||
// Count must match the number of wrappers (observation, counting, floor).
|
||
const forwards = src.match(/^\s*model: opts\.model,$/gm) ?? [];
|
||
expect(forwards.length).toBe(3);
|
||
});
|
||
});
|
||
|
||
// ────────────────────────────────────────────────────────────────────────────
|
||
// Per-finding count primitives — Section 3 unit tests #1–#5, #7, #12.
|
||
// ────────────────────────────────────────────────────────────────────────────
|
||
|
||
describe('optionsSignature', () => {
|
||
test('returns a "|"-joined `index:label` string for a clean list', () => {
|
||
const sig = optionsSignature([
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE EXPANSION' },
|
||
]);
|
||
expect(sig).toBe('1:HOLD SCOPE|2:SCOPE EXPANSION');
|
||
});
|
||
|
||
test('order-independent: shuffled inputs produce the same signature', () => {
|
||
// parseNumberedOptions already returns sorted, but defensive sort means
|
||
// a future caller that hands us shuffled input still produces a stable
|
||
// dedupe signature.
|
||
const a = optionsSignature([
|
||
{ index: 2, label: 'B' },
|
||
{ index: 1, label: 'A' },
|
||
{ index: 3, label: 'C' },
|
||
]);
|
||
const b = optionsSignature([
|
||
{ index: 1, label: 'A' },
|
||
{ index: 2, label: 'B' },
|
||
{ index: 3, label: 'C' },
|
||
]);
|
||
expect(a).toBe(b);
|
||
});
|
||
|
||
test('empty list returns empty string', () => {
|
||
expect(optionsSignature([])).toBe('');
|
||
});
|
||
|
||
test('single-item list returns just that entry', () => {
|
||
expect(optionsSignature([{ index: 1, label: 'Only' }])).toBe('1:Only');
|
||
});
|
||
});
|
||
|
||
describe('parseQuestionPrompt', () => {
|
||
test('keeps the captured boxed learnings header across native CR and blank borders', () => {
|
||
// Exact active-menu bytes from the targeted-a engineering batching run.
|
||
// Its answered setup AUQ lost the title at the standalone box border,
|
||
// leaving every later finding classified as preReview.
|
||
const raw = "☐ Learnings\u001b[K\r\u001b[1B\u001b[K\r\u001b[1B│ D1 — Cross-project learnings scope <gstack-qid:learnings-cross-project>\u001b[K\r\u001b[1B│\u001b[3G\u001b[K\r\r\n│\u001b[3Ggstack\u001b[10Gcan\u001b[14Gsearch\u001b[21Glearnings\u001b[31Gfrom\u001b[36Gyour\u001b[41Gother\u001b[47Gprojects\u001b[56Gon\u001b[59Gthis\u001b[64Gmachine\u001b[72Gto\u001b[75Gfind\u001b[80Gpatterns\u001b[89Gthat\u001b[94Gmight\u001b[100Gapply\u001b[106Ghere.\u001b[112GThis\r\r\n│\u001b[3Gstays\u001b[9Glocal\u001b[15G—\u001b[17Gno\u001b[20Gdata\u001b[25Gleaves\u001b[32Gyour\u001b[37Gmachine.\u001b[46GRecommended\u001b[58Gfor\u001b[62Gsolo\u001b[67Gdevelopers.\u001b[79GSkip\u001b[84Gif\u001b[87Gyou\u001b[91Gwork\u001b[96Gon\u001b[99Gmultiple\u001b[108Gclient\r\r\n│\u001b[3Gcodebases\u001b[13Gwhere\u001b[19Gcross-contamination\u001b[39Gwould\u001b[45Gbe\u001b[48Ga\u001b[50Gconcern.\r\r\n\r\r\n❯\u001b[3G1.\u001b[6GEnable\u001b[13Gcross-project\u001b[27Glearnings\u001b[37G(Recommended)\r\r\n\u001b[6GSearch\u001b[13Glearnings\u001b[23Gfrom\u001b[28Gall\u001b[32Gprojects\u001b[41Gon\u001b[44Gthis\u001b[49Gmachine\u001b[57G—\u001b[59Gsurfaces\u001b[68Gpatterns\u001b[77Gand\u001b[81Gpitfalls\u001b[90Gfrom\u001b[95Gprior\u001b[101Gsessions.\r\r\n\u001b[3G2.\u001b[6GKeep\u001b[11Glearnings\u001b[21Gproject-scoped\u001b[36Gonly\r\r\n\u001b[6GOnly\u001b[11Guse\u001b[15Glearnings\u001b[25Gfrom\u001b[30Gthis\u001b[35Gproject.\u001b[44GSafe\u001b[49Gfor\u001b[53Gmulti-client\u001b[66Genvironments.\r\r\n\u001b[3G3.\u001b[6GType\u001b[11Gsomething.\r\r\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r\r\n\u001b[3G4.\u001b[6GChat\u001b[11Gabout\u001b[17Gthis\r\r\n\r\r\nEnter\u001b[7Gto\u001b[10Gselect\u001b[17G·\u001b[19G↑/↓\u001b[23Gto\u001b[26Gnavigate\u001b[35G·\u001b[37GEsc\u001b[41Gto\u001b[44Gcancel";
|
||
const visible = stripAnsi(raw);
|
||
const question = capturePlanCountQuestion(visible, new Set(), 0, true)!;
|
||
expect(question.promptSnippet).toStartWith('Learnings D1 — Cross-project learnings scope');
|
||
expect(question.promptSnippet).toContain('<gstack-qid:learnings-cross-project>');
|
||
expect(engStep0Boundary(question)).toBe(true);
|
||
const phase = planCountQuestionPhase(question, false, engStep0Boundary);
|
||
expect(phase).toEqual({ preReview: true, reviewStarted: true });
|
||
});
|
||
|
||
test('keeps a long boxed question identity instead of its closing recommendation', () => {
|
||
const frame = [
|
||
'Planning: /tmp/hermetic/.claude/plans/review.md',
|
||
'─'.repeat(120),
|
||
'☐ Architecture',
|
||
'│ D2 — Architecture: custom retry scheduler vs library built-in <gstack-qid:arch-custom-retry-vs-library>',
|
||
'│',
|
||
...Array.from({ length: 12 }, (_, i) => `│ Review context line ${i}: the proposed retry behavior and its tradeoffs.`),
|
||
'│',
|
||
'│ Net: If the library hook is configurable, use the existing implementation.',
|
||
'❯1.Use library built-in (Recommended)',
|
||
'2.Extract shared retry envelope',
|
||
].join('\r\r\n');
|
||
const seen = new Set<string>();
|
||
const question = capturePlanCountQuestion(frame, seen, 0, false)!;
|
||
expect(question.promptSnippet).toStartWith('Architecture D2 — Architecture: custom retry scheduler');
|
||
expect(question.promptSnippet).toContain('<gstack-qid:arch-custom-retry-vs-library>');
|
||
expect(question.promptSnippet).not.toContain('Planning:');
|
||
expect(question.promptSnippet.length).toBeLessThanOrEqual(240);
|
||
expect(capturePlanCountQuestion(frame + '\n' + '·'.repeat(6000), seen, 1, false)).toBeNull();
|
||
});
|
||
|
||
test('does not reuse an old boxed header for a later unboxed menu', () => {
|
||
const visible = [
|
||
'☐ Old setup',
|
||
'D1 — Cross-project learnings scope',
|
||
'❯1.Enable',
|
||
'2.Skip',
|
||
'Planning: /tmp/hermetic/.claude/plans/review.md',
|
||
'D2 — Choose the retry behavior',
|
||
'❯1.Use library built-in',
|
||
'2.Extract shared retry envelope',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toBe('D2 — Choose the retry behavior');
|
||
expect(prompt).not.toContain('Old setup');
|
||
});
|
||
|
||
test('captures 1-line prompt above the cursor', () => {
|
||
const visible = `
|
||
D1 — Pick a mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toBe('D1 — Pick a mode');
|
||
});
|
||
|
||
test('captures multi-line prompt above the cursor', () => {
|
||
const visible = `
|
||
D2 — Approach selection
|
||
|
||
Which architecture should we follow?
|
||
|
||
❯ 1. Bypass existing helper
|
||
2. Reuse existing helper
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Multi-line prompts get joined with single spaces.
|
||
expect(prompt).toContain('D2 — Approach selection');
|
||
expect(prompt).toContain('Which architecture should we follow?');
|
||
});
|
||
|
||
test('returns "" when no cursor is rendered', () => {
|
||
expect(parseQuestionPrompt('Just some prose.\nNo cursor.')).toBe('');
|
||
});
|
||
|
||
test('truncates to 240 chars', () => {
|
||
const longPrompt = 'A'.repeat(500);
|
||
const visible = `${longPrompt}\n\n ❯ 1. yes\n 2. no`;
|
||
expect(parseQuestionPrompt(visible).length).toBeLessThanOrEqual(240);
|
||
});
|
||
|
||
test('does not pull text from a previous numbered list above', () => {
|
||
const visible = `
|
||
❯ 1. previous answered question
|
||
2. previous option two
|
||
|
||
D2 — A new question text
|
||
|
||
❯ 1. fresh option A
|
||
2. fresh option B
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Stops at the previous numbered-list line; should NOT contain "previous answered question".
|
||
expect(prompt).toContain('D2 — A new question text');
|
||
expect(prompt).not.toContain('previous answered question');
|
||
});
|
||
|
||
test('normalizes whitespace (collapses runs of spaces and tabs)', () => {
|
||
const visible = `D1 — Spaced out
|
||
|
||
❯ 1. yes
|
||
2. no`;
|
||
expect(parseQuestionPrompt(visible)).toBe('D1 — Spaced out');
|
||
});
|
||
|
||
test('inline-cursor box-layout: extracts prompt text BEFORE ❯1. on the cursor line', () => {
|
||
// Real /plan-ceo-review rendering: divider + ☐ header + prompt text +
|
||
// cursor are all on one logical line because TTY cursor-positioning
|
||
// escapes collapse the box layout under stripAnsi.
|
||
const visible = [
|
||
'──────────────────',
|
||
'☐ Review scope What scope do you want me to CEO-review? ❯ 1. The branch\'s diff vs main',
|
||
'2. A specific plan file',
|
||
'3. An idea inline',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Should extract "Review scope" and the prompt text, dropping the ☐ box-drawing sigil.
|
||
expect(prompt).toContain('Review scope');
|
||
expect(prompt).toContain('What scope do you want me to CEO-review?');
|
||
expect(prompt).not.toContain('❯');
|
||
expect(prompt).not.toMatch(/^☐/);
|
||
});
|
||
|
||
test('keeps the captured design scope prompt ahead of long Planning chrome', () => {
|
||
// The first failed live attempt fingerprinted only the divider/Planning
|
||
// path. Its actual AUQ was later on the active cursor line.
|
||
const visible = [
|
||
'─'.repeat(120),
|
||
`Planning: /tmp/hermetic/.claude/plans/${'long-path-'.repeat(24)}plan.md`,
|
||
'─'.repeat(120),
|
||
"☐Reviewfocus I've rated this Settings Page UI redesign plan 2/10 on design completeness. Want me to focus on specific areas? ❯1.All7passes(Recommended)",
|
||
'2.All7passesbutskipmockups',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toStartWith('Reviewfocus');
|
||
expect(prompt).toContain('design completeness');
|
||
expect(prompt).not.toContain('Planning:');
|
||
expect(designStep0Boundary({
|
||
signature: 'captured-design-scope', promptSnippet: prompt,
|
||
options: parseNumberedOptions(visible), observedAtMs: 0, preReview: true,
|
||
})).toBe(true);
|
||
});
|
||
|
||
test('keeps the captured devex persona header when cursor spacing collapses', () => {
|
||
const visible = [
|
||
`Planning: /tmp/hermetic/.claude/plans/${'long-path-'.repeat(24)}plan.md`,
|
||
'─'.repeat(120),
|
||
'☐Targetpersona D2—WhoistheprimarydeveloperthisSDKtargets? ❯1.AIappbuilder/startupfounder(Recommended)',
|
||
'2.Backend/platformengineer',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toStartWith('Targetpersona');
|
||
expect(devexStep0Boundary({
|
||
signature: 'captured-devex-persona', promptSnippet: prompt,
|
||
options: parseNumberedOptions(visible), observedAtMs: 0, preReview: true,
|
||
})).toBe(true);
|
||
});
|
||
|
||
test('retains a multiline question while excluding the preceding CLI divider', () => {
|
||
const visible = [
|
||
'Planning: /tmp/hermetic/.claude/plans/plan.md',
|
||
'─'.repeat(120),
|
||
'☐ Review focus',
|
||
'This plan is 2/10 on design completeness.',
|
||
'Want me to focus on specific areas? ❯1.All 7 passes',
|
||
'2.Skip mockups',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toContain('Review focus');
|
||
expect(prompt).toContain('design completeness');
|
||
expect(prompt).toContain('specific areas?');
|
||
expect(prompt).not.toContain('Planning:');
|
||
});
|
||
});
|
||
|
||
describe('auqFingerprint', () => {
|
||
test('returns the same fingerprint for identical inputs', () => {
|
||
const opts = [
|
||
{ index: 1, label: 'A' },
|
||
{ index: 2, label: 'B' },
|
||
];
|
||
expect(auqFingerprint('hello', opts)).toBe(auqFingerprint('hello', opts));
|
||
});
|
||
|
||
test('different prompts with shared option labels produce DIFFERENT fingerprints', () => {
|
||
// The collision regression Codex F1 caught: option-label-only fingerprints
|
||
// collapsed multiple distinct findings into one when they shared menu shape.
|
||
const sharedOpts = [
|
||
{ index: 1, label: 'Add to plan' },
|
||
{ index: 2, label: 'Defer' },
|
||
{ index: 3, label: 'Build now' },
|
||
];
|
||
const fpFinding1 = auqFingerprint('D5 — Architecture: bypass helper?', sharedOpts);
|
||
const fpFinding2 = auqFingerprint('D6 — Tests: zero coverage?', sharedOpts);
|
||
expect(fpFinding1).not.toBe(fpFinding2);
|
||
});
|
||
|
||
test('same prompt with different options produces DIFFERENT fingerprints', () => {
|
||
const prompt = 'D1 — Pick a mode';
|
||
const fpA = auqFingerprint(prompt, [
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE EXPANSION' },
|
||
]);
|
||
const fpB = auqFingerprint(prompt, [
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE REDUCTION' },
|
||
]);
|
||
expect(fpA).not.toBe(fpB);
|
||
});
|
||
|
||
test('whitespace-only differences in prompt do NOT change the fingerprint', () => {
|
||
// Same content, different rendering whitespace (TTY redraw artifact)
|
||
// must produce the same fingerprint so dedupe survives reflow.
|
||
const opts = [{ index: 1, label: 'A' }, { index: 2, label: 'B' }];
|
||
const fpA = auqFingerprint('Pick a mode', opts);
|
||
const fpB = auqFingerprint('Pick a mode', opts);
|
||
expect(fpA).toBe(fpB);
|
||
});
|
||
|
||
test('empty prompt + same options collide (caller must guard against this)', () => {
|
||
// Documents the contract: empty-prompt fingerprints WILL collide if the
|
||
// caller fingerprints them. runPlanSkillCounting must skip empty-prompt
|
||
// AUQs and re-poll instead.
|
||
const opts = [{ index: 1, label: 'A' }];
|
||
expect(auqFingerprint('', opts)).toBe(auqFingerprint('', opts));
|
||
});
|
||
});
|
||
|
||
describe('capturePlanCountQuestion replay', () => {
|
||
test('keeps captured CEO/eng fingerprints stable as later output trims the trailing window', () => {
|
||
// Exact prompt/option fields from the 07:30 corrected paid attempts.
|
||
// Both counted an answered Step0 question again as a review finding
|
||
// once the moving tail omitted the beginning of its prompt.
|
||
const captures = [
|
||
{
|
||
prompt: '☐ RevewMode Which review mode should I use for the remaining sections?',
|
||
labels: [
|
||
'HOLD SCOPE — make it ┌┐',
|
||
'SELECTIVEEXPANSION—│Focus:catcheverylandmineinApproachA│',
|
||
'SCOPEREDUCTION—strip│Tests:whatmustbecovered│',
|
||
'SCOPEEXPANSION—think│Observability:whatlogs/metricsareneeded│',
|
||
],
|
||
},
|
||
{
|
||
prompt: '☐ Scope cut │ D2 — Scope reduction proposal: drop TokenStore and RequestPolicy as standalone classes, inject AuthCache rather than │ exportitglobally.Acceptthisreductionbeforethesection-by-sectionreviewbegins? │ <gstack-qid:plan-eng-review-',
|
||
labels: [
|
||
'Acceptscopereduction┌───────────────────────────────────────────────────┐',
|
||
'Proceedfullscopeas-is│AuthBroker│',
|
||
],
|
||
},
|
||
];
|
||
for (const capture of captures) {
|
||
const options = capture.labels.map((label, i) => `${i === 0 ? '❯' : ''}${i + 1}.${label}`).join('\n');
|
||
const frame = `${capture.prompt}\n${options}`;
|
||
const seen = new Set<string>();
|
||
const first = capturePlanCountQuestion(frame, seen, 0, true)!;
|
||
expect(first).not.toBeNull();
|
||
// Leave the original menu within the trailing4KB, but move the
|
||
// start of that window into its question text, twice in succession.
|
||
const paddingLength = 4096 - options.length - 30;
|
||
for (const extra of [0, 15]) {
|
||
const advanced = frame + '\n' + '·'.repeat(paddingLength + extra - 1);
|
||
expect(advanced.slice(-4096)).not.toContain(capture.prompt);
|
||
expect(parseNumberedOptions(advanced)).toEqual(first.options);
|
||
expect(parseQuestionPrompt(advanced)).toBe(first.promptSnippet);
|
||
expect(auqFingerprint(parseQuestionPrompt(advanced), parseNumberedOptions(advanced))).toBe(first.signature);
|
||
expect(capturePlanCountQuestion(advanced, seen, extra + 1, false)).toBeNull();
|
||
}
|
||
const next = `${frame}\n${'·'.repeat(paddingLength)}\n☐ Next decision Should the revised plan use these same choices?\n${options}`;
|
||
const distinct = capturePlanCountQuestion(next, seen, 20, false)!;
|
||
expect(distinct).not.toBeNull();
|
||
expect(distinct.signature).not.toBe(first.signature);
|
||
expect(distinct.preReview).toBe(false);
|
||
expect(seen.size).toBe(2);
|
||
}
|
||
});
|
||
|
||
test('counts consecutive findings with identical choices and ignores redraws', () => {
|
||
const options = '\n❯1.Add to plan\n2.Defer\n3.Skip';
|
||
const seen = new Set<string>();
|
||
const frames = [
|
||
`D5 — SQL: interpolate the request parameter?${options}`,
|
||
`D5 — SQL: interpolate the request parameter?${options}`,
|
||
`D6 — Tests: no coverage for the webhook?${options}`,
|
||
`D6 — Tests: no coverage for the webhook?${options}`,
|
||
];
|
||
const captured = frames.map((frame, i) => capturePlanCountQuestion(frame, seen, i, false));
|
||
expect(captured.map((question) => question !== null)).toEqual([true, false, true, false]);
|
||
expect(captured[0]?.signature).not.toBe(captured[2]?.signature);
|
||
expect(captured[2]?.promptSnippet).toContain('Tests: no coverage');
|
||
});
|
||
|
||
test('does not consume an incomplete frame before its prompt arrives', () => {
|
||
const seen = new Set<string>();
|
||
const options = '❯1.Add to plan\n2.Defer';
|
||
expect(capturePlanCountQuestion(options, seen, 0, true)).toBeNull();
|
||
expect(capturePlanCountQuestion(`D1 — Pick an approach\n${options}`, seen, 1, true)).not.toBeNull();
|
||
});
|
||
|
||
test('answers the captured CEO retry question with a numeric-leading first label', () => {
|
||
// The live timeout sat on this question because the first label begins
|
||
// with "1retryattempt"; it was incorrectly rejected as a decimal token.
|
||
const frame = [
|
||
' ☐ Retry spec',
|
||
"│ Section 5/6 finding: 'retry-with-backoff fires once, then fails clean' is ambiguous.",
|
||
"│ What does 'fires once' mean?",
|
||
'❯1.1retryattempt—Stripecalledexactly2timestotal(Recommended)',
|
||
'Themostnaturalreading:1originalattempt+1retry=2totalStripecalls.',
|
||
'2.Addaclarifyingcommenttotheplan—lettheimplementerdecide',
|
||
'3.Theretrymechanismhandlesit—justassertfailureisreturned',
|
||
'4.Typesomething.',
|
||
'5.Chataboutthis',
|
||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||
].join('\r\r');
|
||
const question = capturePlanCountQuestion(frame, new Set(), 0, false);
|
||
expect(question?.options.map(({ index }) => index)).toEqual([1, 2, 3, 4, 5]);
|
||
expect(question?.options[0]?.label).toBe('1retryattempt—Stripecalledexactly2timestotal(Recommended)');
|
||
expect(question?.promptSnippet).toContain('Section 5/6 finding');
|
||
expect(question?.promptSnippet).not.toContain('Planning:');
|
||
});
|
||
|
||
test('still ignores decimal numbers inside option labels', () => {
|
||
const frame = 'Choose the retry delay\r❯1.1.5 seconds\r2.Wait 2.5 seconds\r3.No retry';
|
||
expect(parseNumberedOptions(frame)).toEqual([
|
||
{ index: 1, label: '1.5 seconds' },
|
||
{ index: 2, label: 'Wait 2.5 seconds' },
|
||
{ index: 3, label: 'No retry' },
|
||
]);
|
||
});
|
||
});
|
||
|
||
describe('planCountPrerequisitePick replay', () => {
|
||
test('declines captured office-hours prerequisite menus by label in either order', () => {
|
||
// Captured 2026-09-08 CEO/Devex prerequisite surfaces: the default index
|
||
// sometimes starts office-hours, changing the seeded review's input.
|
||
const captures = [
|
||
{
|
||
prompt: 'No design doc found for this branch. `/office-hours` produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input. Run it now, or skip and proceed with standard review?',
|
||
labels: ['Skip — proceed with standard review (Recommended)', 'Run /office-hours first'],
|
||
},
|
||
{
|
||
prompt: 'D2 — No design doc found. Run /office-hours first? <gstack-qid:plan-ceo-prereq-office-hours>',
|
||
labels: ['Skip — standard review (recommended)', 'Run /office-hours now'],
|
||
},
|
||
{
|
||
prompt: 'D3 — Run /office-hours first to produce a design doc for sharper input?',
|
||
labels: ['Skip — proceed with standard review (recommended)', 'Run /office-hours now'],
|
||
},
|
||
];
|
||
for (const { prompt, labels } of captures) {
|
||
for (const reversed of [false, true]) {
|
||
for (const collapsed of [false, true]) {
|
||
const ordered = reversed ? [...labels].reverse() : labels;
|
||
const text = ['☐ Prerequisite', prompt, `❯1.${ordered[0]}`, `2.${ordered[1]}`, '3.Type something.', '4.Chat about this'].join('\r');
|
||
const frame = collapsed ? text.replace(/ /g, '') : text;
|
||
const fp = capturePlanCountQuestion(frame, new Set(), 0, true)!;
|
||
expect(fp).not.toBeNull();
|
||
expect(planCountPrerequisitePick(fp)).toBe(reversed ? 2 : 1);
|
||
expect(planCountPrerequisitePick({ ...fp, preReview: false })).toBeNull();
|
||
}
|
||
}
|
||
}
|
||
});
|
||
|
||
test('keeps existing answers for incomplete, unrelated, and ambiguous menus', () => {
|
||
const fp = capturePlanCountQuestion(
|
||
'☐ Prerequisite\rNo design doc found. Run /office-hours first?\r❯1.Run /office-hours now\r2.Skip — proceed with standard review',
|
||
new Set(), 0, true,
|
||
)!;
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'No design doc found.' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'Should /office-hours skip the required SDK validation finding?' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'Want a second opinion from /office-hours?' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [{ index: 1, label: 'Run /office-hours now' }, { index: 2, label: 'Skip' }] })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [{ index: 1, label: 'Add to plan' }, fp.options[1]] })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [...fp.options, { index: 3, label: 'Skip — standard review' }] })).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('COMPLETION_SUMMARY_RE', () => {
|
||
test('matches GSTACK REVIEW REPORT heading', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('## GSTACK REVIEW REPORT')).toBe(true);
|
||
});
|
||
|
||
test('matches Completion Summary heading (ceo + eng)', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('## Completion Summary')).toBe(true);
|
||
expect(COMPLETION_SUMMARY_RE.test('## Completion summary')).toBe(true);
|
||
});
|
||
|
||
test('matches Status: clean (CEO review-log shape)', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('Status: clean')).toBe(true);
|
||
expect(COMPLETION_SUMMARY_RE.test('Status: issues_open')).toBe(true);
|
||
});
|
||
|
||
test('matches VERDICT: line', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('VERDICT: CLEARED — Eng Review passed')).toBe(true);
|
||
});
|
||
|
||
test('does NOT match prose mentions of "verdict" mid-line', () => {
|
||
// VERDICT must be at the start of a line to count.
|
||
expect(COMPLETION_SUMMARY_RE.test('the final verdict: undecided')).toBe(false);
|
||
});
|
||
|
||
test('does NOT treat source or proposed diff rows as assistant completion', () => {
|
||
for (const line of [
|
||
'409 +## GSTACK REVIEW REPORT',
|
||
'419 +**VERDICT:** Design Review complete — 8 decisions made.',
|
||
'+## GSTACK REVIEW REPORT',
|
||
'409→## GSTACK REVIEW REPORT',
|
||
'The plan must end with ## GSTACK REVIEW REPORT.',
|
||
]) expect(COMPLETION_SUMMARY_RE.test(line)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('classifyPlanCountFrame replay', () => {
|
||
test('waits through proposed Write approval and tool output, then accepts the actual report', () => {
|
||
// Sanitized rows and native prompt from the failed design-count attempt.
|
||
const proposedDiff = [
|
||
'409 +## GSTACK REVIEW REPORT',
|
||
'416 +| Design Review | 1 | issues_open | score: 2/10 → 8/10, 8 decisions |',
|
||
'419 +**VERDICT:** Design Review complete — 8 decisions made.',
|
||
].join('\n');
|
||
const permission = [
|
||
'Doyouwanttooverwritegstack-test-plan-design.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/fixtureforthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
const frames = [
|
||
`${proposedDiff}\n${permission}`,
|
||
`${proposedDiff}\n⏺ Updated gstack-test-plan-design.md`,
|
||
`${proposedDiff}\n⏺ ## GSTACK REVIEW REPORT\nDesign Review complete — 8 decisions made.`,
|
||
];
|
||
expect(frames.map(classifyPlanCountFrame)).toEqual(['permission', null, 'completion_summary']);
|
||
});
|
||
|
||
test('a pending native permission beats even an unnumbered report heading', () => {
|
||
const visible = '## GSTACK REVIEW REPORT\nDoyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
expect(classifyPlanCountFrame(visible)).toBe('permission');
|
||
});
|
||
|
||
test('a later report supersedes the granted menu still in short scrollback', () => {
|
||
const permission = 'Doyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
expect(classifyPlanCountFrame(permission)).toBe('permission');
|
||
expect(classifyPlanCountFrame(`${permission}\n● ## GSTACK REVIEW REPORT`)).toBe('completion_summary');
|
||
});
|
||
|
||
test('an active question after a prior report keeps the counter running', () => {
|
||
expect(classifyPlanCountFrame('## GSTACK REVIEW REPORT\nOne more choice\n❯1.Add to plan\n2.Defer')).toBeNull();
|
||
});
|
||
|
||
test('an active AUQ supersedes a granted permission menu in short scrollback', () => {
|
||
const permission = 'Doyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
const question = '☐ Error handling\nWhich failure path should we test?\n❯1.Timeout\n2.Refusal';
|
||
expect(classifyPlanCountFrame(`${permission}\n${question}`)).toBeNull();
|
||
});
|
||
|
||
test('preserves actual report variants and the native plan-ready terminal', () => {
|
||
for (const report of [
|
||
'## GSTACK REVIEW REPORT', '⏺##GSTACKREVIEWREPORT', '●GSTACKREVIEWREPORT',
|
||
'## Completion Summary', '● ## Completion Summary', 'Status: clean', 'Status: issues_open',
|
||
'VERDICT: CLEARED — Eng Review passed', '**VERDICT:** Design Review complete.',
|
||
]) expect(classifyPlanCountFrame(report)).toBe('completion_summary');
|
||
expect(classifyPlanCountFrame('Ready to execute the plan?\n❯1.Yes\n2.No, keep planning')).toBe('plan_ready');
|
||
});
|
||
});
|
||
|
||
describe('planCountSubmissionInput replay', () => {
|
||
test('uses the captured DevEx panel anchors when the Submit button label is damaged', () => {
|
||
const captured = [
|
||
'← ☒ Routing setup ☐ Cross-project ✔ Submit →',
|
||
'Review your answers',
|
||
'⚠You have not answere all questions',
|
||
' │ ●D1 — Shouldgstack add skill routingrulestothisproject\'sCLAUDE.md?<gstack-qid:routing-injection>',
|
||
'→dd routing rules (Recmmeded)',
|
||
'Ready to submit your answers?',
|
||
'❯1.Sbmi answers',
|
||
'2Cancel',
|
||
].join('\r\r');
|
||
expect(planCountSubmissionInput(captured)).toBe('\x1b[Z');
|
||
expect(planCountSubmissionInput(captured.replace('☐ Cross-project', '☒ Cross-project'))).toBe('\r');
|
||
expect(planCountSubmissionInput(captured + '\r☐ Retry spec\rRetry once?\r❯1.Yes\r2.No')).toBeNull();
|
||
expect(planCountSubmissionInput(captured + '\r☐ Proposal\rSend this proposal?\r❯1.Submit proposal\r2.Keep editing')).toBeNull();
|
||
expect(planCountSubmissionInput(captured + '\r☐ Retry spec\rRetry once?\r❯2.No\r3.Other')).toBeNull();
|
||
expect(planCountSubmissionInput(captured.replace('Review your answers', 'Review context'))).toBeNull();
|
||
expect(planCountSubmissionInput(captured.replace('Ready to submit your answers?', 'Read the proposed answers.'))).toBeNull();
|
||
});
|
||
|
||
test('the captured mode Submit panel with a damaged caption and dotless cursor returns to its unanswered tab', () => {
|
||
// Exact final active panel from targeted-a's SCOPE EXPANSION retry.
|
||
const captured = [
|
||
'← ☒ Routing rule ☐ Design doc ✔ Submit →',
|
||
'',
|
||
'Review your answrs',
|
||
'⚠ You hvenot answered all questions',
|
||
" ● Add gstack skill routing rules tothisproject'sCLAUDE.md?",
|
||
'→dd routing rues (Recommnded)',
|
||
'',
|
||
'Ready to submit your answers?',
|
||
'',
|
||
'❯1Submit answers',
|
||
' 2. Cancel',
|
||
].join('\r');
|
||
expect(planCountSubmissionInput(captured)).toBe('\x1b[Z');
|
||
const answered = captured.replace('☐ Design doc', '☒ Design doc').replace('⚠ You hvenot answered all questions', '');
|
||
expect(planCountSubmissionInput(answered)).toBe('\r');
|
||
expect(planCountSubmissionInput(captured + '\r☐ Design doc\rRun office hours?\r❯1Run now\r2.Skip')).toBeNull();
|
||
});
|
||
|
||
const incomplete = [
|
||
'← ☒ Learnings scope ☐ Approach ✔ Submit →',
|
||
'Review your answers',
|
||
'⚠You have not answered all questions',
|
||
' │ ●D1 — Cross-project learnings: Enable searching learnings from your other local projects?',
|
||
'→Enable cross-project (Recommended)',
|
||
'Ready t submit your answers?',
|
||
'❯1.Submit aswers',
|
||
'2Cancel',
|
||
].join('\r\r');
|
||
|
||
test('returns to the unanswered tab, then submits only after both answers', () => {
|
||
expect(planCountSubmissionInput(incomplete)).toBe('\x1b[Z');
|
||
const nextQuestion = [
|
||
'← ☒ Learnings scope ☐ Approach ✔ Submit →',
|
||
'│ Which approach should this plan use?',
|
||
'❯1.Extend the existing dispatcher',
|
||
'2.Add a separate handler',
|
||
].join('\r\r');
|
||
expect(planCountSubmissionInput(`${incomplete}\r${nextQuestion}`)).toBeNull();
|
||
const question = capturePlanCountQuestion(nextQuestion, new Set(), 0, true);
|
||
expect(question?.promptSnippet).toContain('Which approach');
|
||
expect(question?.options).toHaveLength(2);
|
||
const answered = incomplete.replace('☐ Approach', '☒ Approach').replace('⚠You have not answered all questions', '');
|
||
expect(planCountSubmissionInput(answered)).toBe('\r');
|
||
});
|
||
|
||
test('navigates to the first unanswered tab when more than one remains', () => {
|
||
const frame = incomplete.replace('☒ Learnings scope ☐ Approach', '☐ Learnings scope ☐ Approach ☒ Mode');
|
||
expect(planCountSubmissionInput(frame)).toBe('\x1b[Z\x1b[Z\x1b[Z');
|
||
});
|
||
|
||
test('does not revisit a stale submit panel when a later single question is active', () => {
|
||
expect(planCountSubmissionInput(`${incomplete}\r☐ Retry spec\rRetry once?\r❯1.Yes\r2.No`)).toBeNull();
|
||
expect(planCountSubmissionInput('Ready to submit the plan?\n❯1.Submit\n2.Cancel')).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('assertReviewReportAtBottom', () => {
|
||
test('passes when REVIEW REPORT is the only/last ## heading', () => {
|
||
const content = `# Plan
|
||
|
||
## Context
|
||
stuff
|
||
|
||
## Approach
|
||
more stuff
|
||
|
||
## GSTACK REVIEW REPORT
|
||
|
||
| col | col |
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(true);
|
||
});
|
||
|
||
test('fails when REVIEW REPORT is missing', () => {
|
||
const content = `# Plan
|
||
|
||
## Context
|
||
stuff
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.reason).toMatch(/no GSTACK REVIEW REPORT/);
|
||
});
|
||
|
||
test('fails when REVIEW REPORT exists but a ## heading follows it', () => {
|
||
const content = `# Plan
|
||
|
||
## GSTACK REVIEW REPORT
|
||
|
||
| col | col |
|
||
|
||
## Late Section
|
||
oops
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.reason).toMatch(/trailing ## heading/);
|
||
expect(r.trailingHeadings).toEqual(['## Late Section']);
|
||
});
|
||
|
||
test('passes when only ### subheadings follow REVIEW REPORT (deeper nesting allowed)', () => {
|
||
const content = `## GSTACK REVIEW REPORT
|
||
|
||
### Cross-model tension
|
||
- F1: resolved
|
||
- F2: resolved
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(true);
|
||
});
|
||
|
||
test('fails with multiple trailing ## headings reported', () => {
|
||
const content = `## GSTACK REVIEW REPORT
|
||
|
||
## First trailing
|
||
|
||
## Second trailing
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.trailingHeadings).toHaveLength(2);
|
||
});
|
||
});
|
||
|
||
describe('Step0BoundaryPredicate per-skill', () => {
|
||
// Helper to build a synthetic fingerprint for predicate tests.
|
||
function fp(promptSnippet: string, optionLabels: string[]): AskUserQuestionFingerprint {
|
||
const options = optionLabels.map((label, i) => ({ index: i + 1, label }));
|
||
return {
|
||
signature: auqFingerprint(promptSnippet, options),
|
||
promptSnippet,
|
||
options,
|
||
observedAtMs: 0,
|
||
preReview: true,
|
||
};
|
||
}
|
||
|
||
describe('native Cross-project onboarding boundary', () => {
|
||
// Fresh paid run D, 2026-09-08: native question stems, labels and
|
||
// successful answers. Long explanatory paragraphs are omitted; they
|
||
// must not determine this structural setup boundary.
|
||
const captured = [
|
||
{
|
||
"header": "Routing rules",
|
||
"question": "Should I add gstack skill routing rules to your project's CLAUDE.md? (Note: we're in plan mode — if you pick A, I'll make the edit after we exit plan mode.)",
|
||
"options": [
|
||
"Add routing rules (recommended)",
|
||
"Skip — invoke manually"
|
||
],
|
||
"answer": "Add routing rules (recommended)"
|
||
},
|
||
{
|
||
"header": "Scope challenge",
|
||
"question": "D2 — The plan introduces 4 new classes across 12 files. Should I flag scope reduction as a primary recommendation in the review, or accept the 4-class design and focus findings on quality issues?",
|
||
"options": [
|
||
"Accept 4-class design, focus on quality",
|
||
"Flag scope reduction as primary finding (recommended)"
|
||
],
|
||
"answer": "Accept 4-class design, focus on quality"
|
||
},
|
||
{
|
||
"header": "Cross-project",
|
||
"question": "D3 — Should gstack search learnings from your other projects on this machine when reviewing?",
|
||
"options": [
|
||
"Enable cross-project learnings (recommended)",
|
||
"Keep learnings project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project learnings (recommended)"
|
||
},
|
||
{
|
||
"header": "AuthCache race",
|
||
"question": "D4 — AuthCache is shared mutable state mutated by two services with no serialization. How should we fix it?",
|
||
"options": [
|
||
"Single-writer: AuthBroker owns all writes (recommended)",
|
||
"Immutable cache + versioned replace",
|
||
"Accept and document the race"
|
||
],
|
||
"answer": "Single-writer: AuthBroker owns all writes (recommended)"
|
||
},
|
||
{
|
||
"header": "Double-cache risk",
|
||
"question": "D5 — The plan introduces a new AuthCache class but doesn't say what happens to the existing cache adapter. Are they running in parallel?",
|
||
"options": [
|
||
"AuthCache replaces the adapter — add migration to plan (recommended)",
|
||
"AuthCache wraps the adapter — adapter stays as storage layer",
|
||
"Leave ambiguous — clarify in implementation"
|
||
],
|
||
"answer": "AuthCache replaces the adapter — add migration to plan (recommended)"
|
||
},
|
||
{
|
||
"header": "Error swallowing",
|
||
"question": "D6 — validateAndDispatch() swallows three different error classes across nested try/catch blocks. How should this be resolved in the plan?",
|
||
"options": [
|
||
"Decompose + typed error results (recommended)",
|
||
"Keep structure, add logging + rethrow",
|
||
"Leave as-is — document that swallowing is intentional"
|
||
],
|
||
"answer": "Decompose + typed error results (recommended)"
|
||
},
|
||
{
|
||
"header": "Invalidation tests",
|
||
"question": "D7 — When AuthCache replaces the existing adapter (per D5), the existing invalidation tests (logout, revocation, tenant suspension) become dead — they're testing a retired object. Should the plan explicitly require migrating them?",
|
||
"options": [
|
||
"Migrate invalidation tests to AuthCache — add to plan (recommended)",
|
||
"Scope to new component tests only — leave invalidation as follow-up",
|
||
"Assume existing tests cover it — no explicit migration step"
|
||
],
|
||
"answer": "Migrate invalidation tests to AuthCache — add to plan (recommended)"
|
||
},
|
||
{
|
||
"header": "IDP parallelization",
|
||
"question": "D8 — The plan identifies 5 sequential IDP calls that are independent and could be parallelized with Promise.all. Should we include the fix in this PR or defer it?",
|
||
"options": [
|
||
"Parallelize with Promise.all in this PR (recommended)",
|
||
"Defer to TODOS.md",
|
||
"Leave sequential — document as known limitation"
|
||
],
|
||
"answer": "Parallelize with Promise.all in this PR (recommended)"
|
||
}
|
||
];
|
||
const fingerprint = (index: number) => {
|
||
const row = captured[index];
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'fresh-eng-cross-project', toolUseId: `call-${index}`, answered: true,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
};
|
||
|
||
test('keeps D3 as setup and counts each following actual review call', () => {
|
||
let started = false;
|
||
const phases = captured.map((_, i) => {
|
||
const phase = planCountQuestionPhase(fingerprint(i), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.slice(0, 3).map(p => p.preReview)).toEqual([true, true, true]);
|
||
expect(phases[2].reviewStarted).toBe(true);
|
||
expect(phases.slice(3).map(p => p.preReview)).toEqual([false, false, false, false, false]);
|
||
});
|
||
|
||
test('the captured no-qid scope decision stays setup after Learnings scope', () => {
|
||
let started = false;
|
||
const phases = [0, 2, 1, 3, 4, 5, 6, 7].map(i => {
|
||
const phase = planCountQuestionPhase(fingerprint(i), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.slice(0, 3).map(p => p.preReview)).toEqual([true, true, true]);
|
||
expect(phases.slice(3).map(p => p.preReview)).toEqual([false, false, false, false, false]);
|
||
const call = structuredClone(fingerprint(1).nativeCall!);
|
||
call.questions[0].header = 'Scope complexity';
|
||
call.questions[0].options.reverse();
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
// Even with opposed scope labels, an ordinary per-issue finding lacks
|
||
// the captured whole-plan classes/files identity.
|
||
call.questions[0].question = 'How should we reduce shared mutable cache complexity?';
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
});
|
||
|
||
test('uses the native opposed scope choices without depending on a prompt suffix', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
const q = fp.nativeCall!.questions[0];
|
||
q.question = 'Should local lessons from other repositories be included during reviews?';
|
||
q.options.reverse();
|
||
fp.nativeCall!.answers = { [q.question]: q.options[0].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(fp.nativeCall!, 0, true))).toBe(true);
|
||
});
|
||
|
||
test("an unanswered Cross-project tab cannot borrow another tab's answer", () => {
|
||
const call = structuredClone(fingerprint(2).nativeCall!);
|
||
const other = { header: 'Routing rules', question: 'Add routing rules?', options: [{ label: 'Add rules' }, { label: 'Skip' }] };
|
||
call.questions.push(other);
|
||
call.answers = { [other.question]: 'Add rules' };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[1].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(true);
|
||
});
|
||
|
||
test('pending, failed, missing native metadata and ordinary review questions are not this gate', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engStep0Boundary({ ...fp, nativeCall: undefined })).toBe(false);
|
||
for (const alter of [
|
||
(call: any) => { call.answered = false; },
|
||
(call: any) => { call.failed = true; },
|
||
(call: any) => { call.questions[0].header = 'Architecture issue'; call.questions[0].question = 'D1 — Architecture issue: should the cross-project feature use shared storage?'; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project search' }, { label: 'Disable all search' }]; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project search with project-scoped storage' }, { label: 'Discuss later' }]; },
|
||
(call: any) => { call.questions[0].question = 'How should concurrent AuthCache writes across projects be serialized?';
|
||
call.questions[0].options = [{ label: 'Serialize mutations' }, { label: 'Use project-scoped locks' }]; },
|
||
]) {
|
||
const call = structuredClone(fp.nativeCall!);
|
||
alter(call);
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('ceoStep0Boundary', () => {
|
||
test('FIRES on retained letter-prefixed mode labels, not letter-prefixed architecture', () => {
|
||
expect(ceoStep0Boundary(fp('D3 — Which review mode should this CEO review run in?', [
|
||
'C — HOLD SCOPE (Recommended)', 'B — SELECTIVE EXPANSION', 'A — SCOPE EXPANSION', 'D — SCOPE REDUCTION',
|
||
]))).toBe(true);
|
||
expect(ceoStep0Boundary(fp('D2 — Which implementation approach should this plan follow?', [
|
||
'B — Ideal Architecture (Recommended)', 'A — Fix-Only (Minimal Viable)',
|
||
]))).toBe(false);
|
||
expect(ceoStep0Boundary(fp('Prefer HOLD SCOPE for this decision?', ['C — Keep the dispatcher', 'A — Replace it']))).toBe(false);
|
||
});
|
||
test('parenthesized mode labels end setup, while approach labels and mode mentions do not', () => {
|
||
expect(ceoStep0Boundary(fp('D1 — Which CEO review mode should I run?', [
|
||
'A) SCOPE EXPANSION', 'B) SELECTIVE EXPANSION (recommended)', 'C) HOLD SCOPE', 'D) SCOPE REDUCTION',
|
||
]))).toBe(true);
|
||
expect(ceoStep0Boundary(fp('D2 — Which implementation approach?', ['B) Ideal Architecture', 'A) Fix-Only']))).toBe(false);
|
||
expect(ceoStep0Boundary(fp('Prefer HOLD SCOPE?', ['Discuss C) HOLD SCOPE', 'A) Replace it']))).toBe(false);
|
||
});
|
||
test('FIRES on Step 0F mode-pick AUQ (HOLD SCOPE in options)', () => {
|
||
const f = fp('Pick a mode', ['HOLD SCOPE', 'SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'SCOPE REDUCTION']);
|
||
expect(ceoStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on collapsed mode labels captured from the 2026-09-08 paid controls', () => {
|
||
// Test each captured label independently: a spaced sibling option can
|
||
// otherwise conceal the mismatch and leave every review AUQ in Step 0.
|
||
const labels = [
|
||
'HOLDSCOPE—makeitbulletproof(Recommended)',
|
||
'SELECTIVEEXPANSION┌────────────────────────────────────────────────────────────────────────────────────┐\r (ecommnded) │SELECTIVEEXPANSION│',
|
||
'SCOPEEXPANSION│Neutralposture:presentopportunities,stateeffort,youdecide.│\r │ Good for: substantialfeaturewithsolidfoundation,shippedbeforescopelock.│\r└────────────────────────────────────────────────────────────────────┘',
|
||
'SCOPEREDUCTION—findtheminimalversion',
|
||
];
|
||
for (const label of labels) {
|
||
expect(ceoStep0Boundary(fp('Pick a mode', [label, 'Type something.']))).toBe(true);
|
||
}
|
||
});
|
||
|
||
test('FIRES on scope-selection AUQ with "Skip interview" option (skip-interview path)', () => {
|
||
// After calibration run 1: plan-ceo's first AUQ is scope-selection,
|
||
// and we route via "Skip interview and plan immediately" to bypass
|
||
// Step 0 entirely. Boundary must fire on this AUQ so subsequent
|
||
// AUQs go to reviewCount.
|
||
const f = fp(
|
||
'What scope do you want me to CEO-review?',
|
||
[
|
||
"The branch's diff vs main",
|
||
'A specific plan file',
|
||
"An idea you'll describe inline",
|
||
'Cancel — wrong skill',
|
||
'Type something.',
|
||
'Chat about this',
|
||
'Skip interview and plan immediately',
|
||
],
|
||
);
|
||
expect(ceoStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on premise challenge AUQs', () => {
|
||
const f = fp('D1 — Premise check: is this the right problem?', ['Yes', 'No', 'Other']);
|
||
expect(ceoStep0Boundary(f)).toBe(false);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Architecture: bypass helper?', ['Reuse existing', 'Roll new', 'Defer']);
|
||
expect(ceoStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('engStep0Boundary', () => {
|
||
// Captured native question text, labels and answers from targeted-b's
|
||
// engineering retry. Descriptions are immaterial to the phase boundary.
|
||
const captured = [
|
||
{
|
||
"header": "Scope",
|
||
"question": "D1 — Multi-tenant Auth Refactor complexity check: 12 files + 4 new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-complexity>",
|
||
"options": [
|
||
"Proceed as-is",
|
||
"Reduce: cut TokenStore + RequestPolicy",
|
||
"Reduce: single-pass strangler"
|
||
],
|
||
"answer": "Proceed as-is"
|
||
},
|
||
{
|
||
"header": "Shared Cache",
|
||
"question": "D2 — Arch issue 1: AuthBroker and SessionMint both mutate a global shared AuthCache via module-level export. How should this be fixed? <gstack-qid:plan-eng-shared-mutable-cache>",
|
||
"options": [
|
||
"Inject AuthCache as a dependency (recommended)",
|
||
"Make mutations go through a single owner",
|
||
"Accept the risk for now, document it"
|
||
],
|
||
"answer": "Inject AuthCache as a dependency (recommended)"
|
||
},
|
||
{
|
||
"header": "TOCTOU",
|
||
"question": "D3 — Arch issue 2: TOCTOU window during tenant suspension. The plan says AuthCache invalidates entries on tenant suspension, but with two services mutating the cache, a token validation begun before suspension completes may still succeed after the tenant is suspended. How should this be addressed? <gstack-qid:plan-eng-toctou-suspension>",
|
||
"options": [
|
||
"Add suspension check at session issuance boundary (recommended)",
|
||
"Add invalidation ordering guarantee to the plan",
|
||
"Accept the window, note it as an edge case"
|
||
],
|
||
"answer": "Add suspension check at session issuance boundary (recommended)"
|
||
},
|
||
{
|
||
"header": "Error handling",
|
||
"question": "D4 — Code quality issue 1: validateAndDispatch() swallows three distinct error classes in nested catch blocks with no re-throw, logging, or metrics. Errors disappear silently. How should this be handled? <gstack-qid:plan-eng-error-swallowing>",
|
||
"options": [
|
||
"Refactor to flat error handling with explicit re-throw or typed result (recommended)",
|
||
"Add logging inside each catch, keep structure",
|
||
"Leave it, add a lint rule to catch new instances"
|
||
],
|
||
"answer": "Refactor to flat error handling with explicit re-throw or typed result (recommended)"
|
||
},
|
||
{
|
||
"header": "Test coverage",
|
||
"question": "D5 — Test issue 1: 0/18 code paths covered in the plan. The plan scopes tests to 'new components and their success/error paths' but omits: cache invalidation edge cases, all three catch blocks in validateAndDispatch(), and the 5 IDP call failure modes. Should the test scope be expanded? <gstack-qid:plan-eng-test-coverage>",
|
||
"options": [
|
||
"Expand test scope to cover all 18 paths (recommended)",
|
||
"Cover new paths only, defer legacy and edge cases",
|
||
"Accept current test scope as stated in the plan"
|
||
],
|
||
"answer": "Expand test scope to cover all 18 paths (recommended)"
|
||
},
|
||
{
|
||
"header": "IDP calls",
|
||
"question": "D6 — Performance issue 1: token validation makes 5 sequential IDP API calls. The plan acknowledges they are independent and could be parallelized via Promise.all. Should this be fixed in this PR or deferred? <gstack-qid:plan-eng-idp-sequential-calls>",
|
||
"options": [
|
||
"Parallelize now with Promise.all (recommended)",
|
||
"Defer to a follow-up PR, add a TODO",
|
||
"Add a concurrency cap via Promise.all with limit"
|
||
],
|
||
"answer": "Parallelize now with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"header": "Cache bounds",
|
||
"question": "D7 — Performance issue 2 (medium confidence): AuthCache evicts on token expiry but the plan doesn't mention a max-size bound. In a high-tenant deployment, long-lived non-expiring tokens could grow the cache without bound. Is there already a size cap, or should one be added? <gstack-qid:plan-eng-cache-unbounded>",
|
||
"options": [
|
||
"Verify existing cap exists and document it in the plan",
|
||
"Add explicit max-size eviction policy to AuthCache (recommended)",
|
||
"Defer, this is a scaling concern not a correctness one"
|
||
],
|
||
"answer": "Verify existing cap exists and document it in the plan"
|
||
},
|
||
{
|
||
"header": "TODO",
|
||
"question": "D8 — TODO candidate: Auth failure observability. The plan replaces silently-swallowed errors with typed errors, but adds no metrics, logs, or alerts for auth failure patterns. This gap won't surface until production incidents occur. Add a TODO? <gstack-qid:plan-eng-todo-observability>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Build it now in this PR instead of deferring",
|
||
"Skip — not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
}
|
||
];
|
||
function nativeScopeFingerprint(index = 0): AskUserQuestionFingerprint {
|
||
const row = captured[index];
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'captured-eng-retry', toolUseId: `call-${index}`, answered: true,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
}
|
||
|
||
test('keeps captured scope-complexity setup and counts the six following findings', () => {
|
||
let started = false;
|
||
const phases = captured.map((_, index) => {
|
||
const question = nativeScopeFingerprint(index);
|
||
const phase = planCountQuestionPhase(question, started, engStep0Boundary);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases[0]).toEqual({ preReview: true, reviewStarted: true });
|
||
expect(phases.slice(1, 7).filter(phase => !phase.preReview)).toHaveLength(6);
|
||
// The later answered observability TODO retains the existing phase policy.
|
||
expect(phases.filter(phase => !phase.preReview)).toHaveLength(7);
|
||
});
|
||
|
||
test('requires an answered native scope decision with its opposed scope choices', () => {
|
||
const original = nativeScopeFingerprint();
|
||
expect(engStep0Boundary(original)).toBe(true);
|
||
expect(engStep0Boundary({ ...original, nativeCall: undefined })).toBe(false);
|
||
const pending = structuredClone(original);
|
||
pending.nativeCall!.answered = false;
|
||
delete pending.nativeCall!.answers;
|
||
expect(engStep0Boundary(pending)).toBe(false);
|
||
const unansweredScope = structuredClone(original);
|
||
unansweredScope.nativeCall!.answers = { 'Separate answered setup question': 'Continue' };
|
||
expect(engStep0Boundary(unansweredScope)).toBe(false);
|
||
for (const alter of [
|
||
(q: any) => { q.question = 'D1 — Architecture issue: the plan touches 12 files and introduces 4 classes, but AuthCache has a race. Reduce cache scope? <gstack-qid:plan-eng-cache-complexity>'; },
|
||
(q: any) => { q.options = [{ label: 'Change cache size' }, { label: 'Keep cache size' }]; },
|
||
]) {
|
||
const unrelated = structuredClone(original);
|
||
alter(unrelated.nativeCall!.questions[0]);
|
||
unrelated.nativeCall!.answers = { [unrelated.nativeCall!.questions[0].question]: captured[0].answer };
|
||
expect(engStep0Boundary(unrelated)).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('FIRES on cross-project learnings prompt', () => {
|
||
const f = fp('Enable cross-project learnings on this machine?', ['Yes', 'No']);
|
||
expect(engStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('recognizes the captured cross-project gate after cursor spacing collapses', () => {
|
||
const frame = [
|
||
'☐Cross-project gstackcansearchlearningsfromyourotherprojectsonthismachinetofindpatternsthatmightapplytothisreview.Enablecross-projectlearnings?',
|
||
'❯1.Enablecross-projectlearnings(Recommended)',
|
||
'2.Keeplearningsproject-scopedonly',
|
||
].join('\r\r');
|
||
const question = capturePlanCountQuestion(frame, new Set(), 0, true)!;
|
||
expect(question).not.toBeNull();
|
||
expect(engStep0Boundary(question)).toBe(true);
|
||
expect(engStep0Boundary(fp('Scopereductionrecommendation:cuttoMVP?', ['Reduce', 'Proceed']))).toBe(true);
|
||
});
|
||
|
||
test('FIRES on scope reduction recommendation', () => {
|
||
const f = fp('Scope reduction recommendation: cut to MVP?', ['Reduce', 'Proceed', 'Modify']);
|
||
expect(engStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Architecture: shared mutable state?', ['Refactor', 'Defer', 'Skip']);
|
||
expect(engStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('designStep0Boundary', () => {
|
||
const focusTemplate = readFileSync(new URL('../../plan-design-review/SKILL.md.tmpl', import.meta.url), 'utf8')
|
||
.match(/### 0D\. Focus Areas\nAskUserQuestion: "([^\n]+)"/)?.[1] ?? '';
|
||
const focusQuestion = (gaps: string) => focusTemplate.replace('{N}', '4').replace('{X, Y, Z}', gaps);
|
||
const focusOptions = ['Review all 7 dimensions', 'Focus on specific areas'];
|
||
const nativeFocus = (question: string): AskUserQuestionFingerprint => {
|
||
const fingerprint = nativePlanCallFingerprint({
|
||
sessionId: 'design-focus-session', toolUseId: 'toolu-design-focus',
|
||
answered: true, failed: false, answers: { [question]: focusOptions[0]! },
|
||
unansweredQuestionIndices: [],
|
||
questions: [{ question, header: 'Focus areas', multiSelect: false,
|
||
options: focusOptions.map(label => ({ label, description: label })) }],
|
||
}, 0, true);
|
||
fingerprint.promptSnippet = question.slice(0, 240);
|
||
return fingerprint;
|
||
};
|
||
|
||
test('FIRES on the current template Step 0D focus-area question', () => {
|
||
expect(focusTemplate).toContain('Want me to focus on specific areas instead of all 7?');
|
||
const question = focusQuestion('hierarchy, spacing, contrast');
|
||
expect(question.length).toBeLessThanOrEqual(240);
|
||
expect(designStep0Boundary(fp(question, focusOptions))).toBe(true);
|
||
});
|
||
|
||
test('reads the owned full focus question when its gap list exceeds the diagnostic snippet', () => {
|
||
const question = focusQuestion('primary-action hierarchy, inconsistent vertical rhythm, inaccessible error contrast, label-size drift, absent loading feedback, and missing recovery states');
|
||
const fingerprint = nativeFocus(question);
|
||
expect(fingerprint.promptSnippet).not.toContain('Want me to focus');
|
||
expect(question.length).toBeGreaterThan(240);
|
||
expect(designStep0Boundary(fingerprint)).toBe(true);
|
||
});
|
||
|
||
test.each([
|
||
"I've rated this plan 4/10 on design completeness. Should we add a loading state?",
|
||
'Want me to focus on specific areas instead of all 7?',
|
||
"I've rated the error message 4/10 on design completeness. Want me to focus on specific areas instead of all 7?",
|
||
"I've rated this plan 4/10 on design completeness. Should we focus on correcting error contrast?",
|
||
])('does NOT turn a later finding into setup from a partial focus match: %s', question => {
|
||
expect(designStep0Boundary(nativeFocus(question))).toBe(false);
|
||
});
|
||
|
||
test('does NOT combine partial focus matches across separate native question tabs', () => {
|
||
const fingerprint = nativeFocus("I've rated this plan 4/10 on design completeness. Should we add a loading state?");
|
||
const call = fingerprint.nativeCall!;
|
||
const secondQuestion = 'Want me to focus on specific areas instead of all 7?';
|
||
call.questions.push({ ...call.questions[0]!, question: secondQuestion });
|
||
call.answers![secondQuestion] = focusOptions[0]!;
|
||
expect(designStep0Boundary(fingerprint)).toBe(false);
|
||
});
|
||
|
||
test('FIRES on design system / posture mention', () => {
|
||
const f = fp('Pick a design posture for this review', ['Polish', 'Triage', 'Expansion']);
|
||
expect(designStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on first-dimension prompt', () => {
|
||
const f = fp('First dimension: visual hierarchy. Score?', ['7', '8', '9']);
|
||
expect(designStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on later dimension AUQs', () => {
|
||
const f = fp('Spacing dimension score?', ['7', '8', '9']);
|
||
expect(designStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('design review begins without an optional focus question', () => {
|
||
// Captured in the second07:53 paid attempt: real D1-D7 questions were
|
||
// all incorrectly marked preReview, producing reviewCount=0 at completion.
|
||
const questions = [
|
||
'☐Buttonstyle │D1—Howshouldthe4headerbuttons(Save,Reset,Cancel,Export)bevisuallydifferentiated? │<gstack-qid:plan-design-review-butn-hierarchy>',
|
||
'☐ Loading UX │D2—Whatloadingindicatorshouldappearduringthe2-5secondSaveoperation? │<gtack-qid:plan-esign-review-loadig-indicator>',
|
||
'☐ Spacing │D3—Whichspacingscaleshouldthesettingspagestandardizeon?<gstack-qid:plan-design-review-spacing-scale>',
|
||
'☐Typography │D4—Which2-sizetypographysystemshouldthesettingspageuse?<gstack-qid:plan-design-review-type-system>',
|
||
'☐Mobile layout │D5 — On obile (<768px), how should the4-buton header behave?<gstack-qid:plan-design-review-mobile-header>',
|
||
'☐DEIGN.md TODO │D6 — TODO: Create a DESIGN.md file codifying the5 decisions mdein this revew <gstack-qid:plan-design-review-todo-designmd>',
|
||
'☐PartialfailTODO │D7—TODO:Specifythepartial-failurestate—whatdoestheuserseeifSavesucceedsforsomefieldsbutfailsfor others? <gstack-qid:plan-design-review-todo-partialfail>',
|
||
];
|
||
|
||
test('counts the first captured finding and every subsequent finding', () => {
|
||
let reviewStarted = false;
|
||
const phases = questions.map(question => {
|
||
const phase = planCountQuestionPhase(fp(question, ['Apply', 'Defer']), reviewStarted,
|
||
designStep0Boundary, designFirstReviewAUQ);
|
||
reviewStarted = phase.reviewStarted;
|
||
return phase.preReview;
|
||
});
|
||
expect(phases).toEqual([false, false, false, false, false, false, false]);
|
||
});
|
||
|
||
test('keeps the observed focus gate separate when it is emitted', () => {
|
||
const focus = fp("☐ Focus areas │ I've rated this plan2/10 on design completeness. Review all7 dimensions?", ['All7dimensions', 'Priority gaps']);
|
||
const setup = planCountQuestionPhase(focus, false, designStep0Boundary, designFirstReviewAUQ);
|
||
expect(setup).toEqual({ preReview: true, reviewStarted: true });
|
||
expect(planCountQuestionPhase(fp(questions[0], ['Apply', 'Defer']), setup.reviewStarted,
|
||
designStep0Boundary, designFirstReviewAUQ)).toEqual({ preReview: false, reviewStarted: true });
|
||
});
|
||
|
||
test('requires review identity, not just a D1 label or setup question ID', () => {
|
||
for (const question of [
|
||
'☐ Setup │D1—Enable cross-project learnings?',
|
||
'☐ Review target │D1—Which plan should I review?<gstack-qid:plan-design-review-scope>',
|
||
'☐ Focus │D1—What should this design review focus on?<gstack-qid:plan-design-review-focus-areas>',
|
||
'☐ Scope │I will review Pass1 through Pass7 after setup. Proceed?',
|
||
]) expect(designFirstReviewAUQ(fp(question, ['Yes', 'No']))).toBe(false);
|
||
expect(designFirstReviewAUQ(fp('☐ Page structure │ Pass1 — Information Architecture: what page structure should this use?', ['Standard', 'Sidebar']))).toBe(true);
|
||
});
|
||
|
||
test('leaves callers without a first-review predicate unchanged', () => {
|
||
expect(planCountQuestionPhase(fp(questions[0], ['Apply', 'Defer']), false, designStep0Boundary))
|
||
.toEqual({ preReview: true, reviewStarted: false });
|
||
});
|
||
});
|
||
|
||
describe('devexStep0Boundary', () => {
|
||
test('FIRES on developer persona selection', () => {
|
||
const f = fp('Pick the target persona for this review', ['Senior backend', 'Junior frontend', 'Other']);
|
||
expect(devexStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on TTHW target prompt', () => {
|
||
const f = fp('What is the TTHW target for first run?', ['<5 min', '<15 min', '<30 min']);
|
||
expect(devexStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Friction point: 5-min CI wait. Address?', ['Now', 'Defer', 'Skip']);
|
||
expect(devexStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
});
|
||
|
||
|
||
describe('file permission lifecycle replay', () => {
|
||
const permission = (file = 'gstack-test-plan-design.md') => [
|
||
`Do you want to make this edit to ${file}?`,
|
||
'❯ 1. Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/fixtureforthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
|
||
test('ignores the granted menu and its redraw until a new request follows file-tool completion', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission();
|
||
expect(guard(first)).toBe('grant');
|
||
expect(guard(first)).toBe('handled');
|
||
const redraw = first + '\n' + permission();
|
||
expect(guard(redraw)).toBe('handled');
|
||
const completed = redraw + '\n●Write(/tmp/fixture/gstack-test-plan-design.md)\n' +
|
||
'⎿ Wrote320linesto../fixture/gstack-test-plan-design.md\n' + '·'.repeat(1600);
|
||
expect(classifyPlanCountFrame(completed)).toBeNull();
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + permission())).toBe('grant');
|
||
});
|
||
|
||
test('singular native Write/Edit results release a fresh identical permission', () => {
|
||
for (const result of ['⎿ Added1line,removed1line', '⎿ Wrote1lineto../fixture/plan.md', '⎿ Removed1line', '⎿\u00a0Wrote320linesto../fixture/plan.md']) {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission();
|
||
expect(guard(first)).toBe('grant');
|
||
const completed = first + '\n' + result;
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + permission())).toBe('grant');
|
||
}
|
||
});
|
||
|
||
test('the captured active Edit menu remains a permission behind a long diff repaint', () => {
|
||
const visible = permission('gstack-test-plan-ceo.md') + '\n' +
|
||
' 89 +The plan adds StripePaymentWebhookHandler outside WebhookDispatcher.\n'.repeat(40);
|
||
expect(visible.length).toBeLessThan(4096);
|
||
expect(classifyPlanCountFrame(visible)).toBeNull(); // The old 1.5 KB scan misses it.
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(guard(visible)).toBe('grant');
|
||
expect(guard(visible)).toBe('handled');
|
||
});
|
||
|
||
test('a completed Write invalidates an old menu even if polling missed the original grant', () => {
|
||
const visible = permission() + '\n⎿ Wrote320linesto../fixture/gstack-test-plan-design.md';
|
||
expect(createPlanCountPermissionGuard()(visible)).toBe('handled');
|
||
});
|
||
|
||
test('proposed results and tool headers do not release the same pending permission', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
let visible = permission();
|
||
expect(guard(visible)).toBe('grant');
|
||
for (const line of ['320 +⎿ Wrote320lines', '●Write(/tmp/fixture/plan.md)', '⎿ Tip: use /btw', '⎿ Error: denied']) {
|
||
visible += '\n' + line + '\n' + permission();
|
||
expect(guard(visible)).toBe('handled');
|
||
}
|
||
});
|
||
|
||
test('a different file and a genuine native file-policy question retain their own input', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission('first.md');
|
||
expect(guard(first)).toBe('grant');
|
||
expect(guard(first + '\n' + permission('FIRST.md'))).toBe('grant'); // Targets remain case-sensitive.
|
||
expect(guard(first + '\n' + permission('second.md'))).toBe('grant');
|
||
const question = '\n☐ File policy\nDo you want to create first.md?\n❯1.Yes\n2.No\n' +
|
||
'Enter to select · ↑/↓ to navigate · Esc to cancel';
|
||
expect(guard(first + question)).toBeNull();
|
||
expect(capturePlanCountQuestion(first + question, new Set(), 0, true)?.promptSnippet).toContain('File policy');
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng setup ordering (captured F)', () => {
|
||
// Exact native question stems and choice labels: two setup calls followed
|
||
// by five real findings. Descriptions do not establish phase identity.
|
||
const rows = [
|
||
{
|
||
"header": "Learnings scope",
|
||
"question": "D1 \u2014 Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>",
|
||
"options": [
|
||
"Enable cross-project (Recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project (Recommended)"
|
||
},
|
||
{
|
||
"header": "Scope complexity",
|
||
"question": "D2 \u2014 The plan introduces 4 new classes across 12 files. That's above the complexity threshold (>2 classes / >8 files). Should we reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-complexity>",
|
||
"options": [
|
||
"Reduce: merge to 2 classes (Recommended)",
|
||
"Proceed as-is \u2014 4 classes, 12 files",
|
||
"Reduce further: 1 new class only"
|
||
],
|
||
"answer": "Reduce: merge to 2 classes (Recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: shared state",
|
||
"question": "D3 \u2014 Architecture Issue 1: AuthCache is a global mutable singleton exported at module level; both AuthBroker and (previously) SessionMint mutate it. This creates concurrent-mutation risk across tenant requests and makes the services untestable in isolation. <gstack-qid:plan-eng-arch-global-cache>",
|
||
"options": [
|
||
"Inject AuthCache via constructor (Recommended)",
|
||
"Keep global, add locking",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Inject AuthCache via constructor (Recommended)"
|
||
},
|
||
{
|
||
"header": "Code quality",
|
||
"question": "D4 \u2014 Code Quality Issue 1: validateAndDispatch() is 60 lines with three nested try/catch blocks, each swallowing a different error class. Swallowed errors mean callers can't distinguish an IDP timeout from a policy rejection from a token parse failure \u2014 all three silently return the same result. <gstack-qid:plan-eng-quality-error-handling>",
|
||
"options": [
|
||
"Extract + typed error union (Recommended)",
|
||
"Add logging to each catch, keep structure",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Extract + typed error union (Recommended)"
|
||
},
|
||
{
|
||
"header": "Test: regression",
|
||
"question": "D5 \u2014 Test Issue 1 (IRON RULE): legacyAuthFlow() is being rewritten with no regression test for its prior behavior. The plan explicitly says coverage 'does not exercise legacyAuthFlow() or assert compatibility with its prior behavior.' A rewrite without a behavioral snapshot means any regression is invisible until production. <gstack-qid:plan-eng-test-legacy-regression>",
|
||
"options": [
|
||
"Add characterization tests before rewrite (Recommended)",
|
||
"Document expected behavior, manual verify",
|
||
"Skip regression coverage"
|
||
],
|
||
"answer": "Add characterization tests before rewrite (Recommended)"
|
||
},
|
||
{
|
||
"header": "Test: isolation",
|
||
"question": "D6 \u2014 Test Issue 2: The plan says 'unit and integration coverage is planned for success/error paths' but makes no mention of cross-tenant isolation tests. The AuthCache key includes tenant ID, issuer, audience, and policy version \u2014 a key-construction bug would let Tenant A read Tenant B's cached tokens. This is the highest-severity failure mode in a multi-tenant auth system. <gstack-qid:plan-eng-test-tenant-isolation>",
|
||
"options": [
|
||
"Add explicit cross-tenant isolation tests (Recommended)",
|
||
"Cover via integration tests only",
|
||
"Proceed with existing test plan"
|
||
],
|
||
"answer": "Add explicit cross-tenant isolation tests (Recommended)"
|
||
},
|
||
{
|
||
"header": "Perf: IDP calls",
|
||
"question": "D7 \u2014 Performance Issue 1: Token validation makes 5 sequential API calls to the IDP. The plan itself notes they are independent and could be parallelized via Promise.all 'trivially.' Sequential calls add latency proportional to IDP round-trip time x5 on every auth request. At p99 IDP latency of 100ms, that's 500ms of unnecessary serialization per login. <gstack-qid:plan-eng-perf-parallel-idp>",
|
||
"options": [
|
||
"Parallelize with Promise.all in this PR (Recommended)",
|
||
"Defer to follow-up ticket",
|
||
"Defer with in-code TODO comment"
|
||
],
|
||
"answer": "Parallelize with Promise.all in this PR (Recommended)"
|
||
}
|
||
];
|
||
const fingerprint = (index: number): AskUserQuestionFingerprint => {
|
||
const row = rows[index]!;
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'captured-f-eng', toolUseId: `f-${index}`, answered: true, failed: false,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
};
|
||
const phasesFor = (indices: number[]) => {
|
||
let started = false;
|
||
return indices.map(index => {
|
||
const phase = planCountQuestionPhase(fingerprint(index), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
};
|
||
|
||
test('both setup orders exclude setup and include the first of five actual findings', () => {
|
||
for (const order of [[0, 1], [1, 0]]) {
|
||
const phases = phasesFor([...order, 2, 3, 4, 5, 6]);
|
||
expect(phases.slice(0, 2).map(p => p.preReview)).toEqual([true, true]);
|
||
expect(phases[1]!.reviewStarted).toBe(true);
|
||
expect(phases[2]!.preReview).toBe(false);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(5);
|
||
}
|
||
});
|
||
|
||
test('setup IDs plus opposed answered choices survive header and question-body variants', () => {
|
||
for (const index of [0, 1]) {
|
||
const original = fingerprint(index).nativeCall!;
|
||
const id = index === 0 ? 'cross-project-learnings' : 'plan-eng-scope-complexity';
|
||
for (const header of ['Learnings scope', 'Scope complexity', 'Architecture', '']) {
|
||
const call = structuredClone(original);
|
||
const q = call.questions[0]!;
|
||
q.header = header;
|
||
q.question = `Choose the setup scope. <gstack-qid:${id}>`;
|
||
q.options.reverse();
|
||
call.answers = { [q.question]: q.options[0]!.label };
|
||
const fp = nativePlanCallFingerprint(call, 0, true);
|
||
expect(engSetupAUQ(fp)).toBe(true);
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('repeated setup does not become a finding after the boundary opens', () => {
|
||
const phases = phasesFor([0, 0, 1, 1, 2, 3, 4, 5, 6]);
|
||
expect(phases.slice(0, 4).every(p => p.preReview)).toBe(true);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(5);
|
||
});
|
||
|
||
test('only successful answered setup metadata can exclude a call', () => {
|
||
for (const index of [0, 1]) {
|
||
const fp = fingerprint(index);
|
||
expect(engSetupAUQ({ ...fp, nativeCall: undefined })).toBe(false);
|
||
for (const alter of [
|
||
(call: any) => { call.answered = false; },
|
||
(call: any) => { call.failed = true; },
|
||
(call: any) => { call.answers = {}; },
|
||
(call: any) => { call.questions[0].question = 'Review cache isolation. <gstack-qid:plan-eng-cache-complexity>'; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Apply fix' }, { label: 'Defer finding' }]; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project with project-scoped storage; proceed as-is or reduce' }, { label: 'Discuss' }]; },
|
||
]) {
|
||
const call = structuredClone(fp.nativeCall!);
|
||
alter(call);
|
||
// Keep a successful answer after question/option mutations, so those
|
||
// controls exercise identity/actions rather than an absent answer key.
|
||
if (Object.keys(call.answers ?? {}).length) {
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
}
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('an answered sibling cannot turn an unanswered setup tab or mixed finding packet into setup', () => {
|
||
const call = structuredClone(fingerprint(0).nativeCall!);
|
||
const issue = fingerprint(2).nativeCall!.questions[0]!;
|
||
call.questions.push(issue);
|
||
call.answers = { [issue.question]: issue.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
call.answers[call.questions[0]!.question] = call.questions[0]!.options[0]!.label;
|
||
const fp = nativePlanCallFingerprint(call, 0, false);
|
||
expect(engSetupAUQ(fp)).toBe(false);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
});
|
||
|
||
test('the optional predicate leaves other callers and substantive Eng qids unchanged', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engSetupAUQ(fp)).toBe(false);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
// Historical boundary detection can also fire on actual review qids;
|
||
// those must never be reused as the late-setup exclusion predicate.
|
||
fp.promptSnippet += ' <gstack-qid:plan-eng-review-global-cache>';
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
const setup = fingerprint(0);
|
||
expect(planCountQuestionPhase(setup, true, engStep0Boundary).preReview).toBe(false);
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng first packet and registry identities', () => {
|
||
const scope = {
|
||
header: 'Scope complexity',
|
||
question: 'D1 — Choose the whole-plan scope. <gstack-qid:plan-eng-review-scope-reduce>',
|
||
options: [{ label: 'Reduce: merge two classes (Recommended)' }, { label: 'Proceed as-is: four classes' }],
|
||
};
|
||
const finding = {
|
||
header: 'Arch: shared state',
|
||
question: 'D3 — Architecture Issue 1: AuthCache is a global mutable singleton exported at module level. <gstack-qid:plan-eng-arch-global-cache>',
|
||
options: [{ label: 'Inject an owned cache' }, { label: 'Keep the global cache' }],
|
||
};
|
||
const callWith = (questions: typeof scope[], answers: Record<string, string>) => ({
|
||
sessionId: 'eng-first-packet', toolUseId: 'mixed-call', answered: true, failed: false,
|
||
questions, answers,
|
||
});
|
||
const phase = (call: ReturnType<typeof callWith>) => planCountQuestionPhase(
|
||
nativePlanCallFingerprint(call, 0, true), false, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ,
|
||
);
|
||
|
||
test('registry setup qids stay setup and require opposed scope actions', () => {
|
||
const call = callWith([scope], { [scope.question]: scope.options[0]!.label });
|
||
const fp = nativePlanCallFingerprint(call, 0, true);
|
||
expect(engSetupAUQ(fp)).toBe(true);
|
||
expect(engFirstReviewAUQ(fp)).toBe(false);
|
||
expect(phase(call)).toEqual({ preReview: true, reviewStarted: true });
|
||
call.questions = [{ ...scope, options: [{ label: 'Change cache size' }, { label: 'Keep cache size' }] }];
|
||
call.answers = { [scope.question]: 'Change cache size' };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
const learnings = {
|
||
header: 'Learnings scope', question: 'Choose local learning scope. <gstack-qid:preamble-cross-project-learnings>',
|
||
options: [{ label: 'Enable cross-project learnings' }, { label: 'Keep project-scoped only' }],
|
||
};
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(callWith([learnings], {
|
||
[learnings.question]: learnings.options[0]!.label,
|
||
}), 0, true))).toBe(true);
|
||
});
|
||
|
||
test('a first packet with answered setup and a real finding counts as one review call', () => {
|
||
const call = callWith([scope, finding], {
|
||
[scope.question]: scope.options[0]!.label,
|
||
[finding.question]: finding.options[0]!.label,
|
||
});
|
||
expect(phase(call)).toEqual({ preReview: false, reviewStarted: true });
|
||
// The packet is one call even with two answered tabs; the reader/counter
|
||
// dedups the unchanged session/tool-use ID rather than counting each tab.
|
||
expect(nativePlanCallFingerprint(call, 0, true).signature).toBe('eng-first-packet:mixed-call');
|
||
for (const questions of [[scope, finding], [finding, scope]]) {
|
||
expect(phase({ ...call, questions }).preReview).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('an unanswered or failed finding cannot start review from a setup packet', () => {
|
||
const call = callWith([scope, finding], { [scope.question]: scope.options[0]!.label });
|
||
expect(phase(call)).toEqual({ preReview: true, reviewStarted: true });
|
||
call.answers = { [finding.question]: finding.options[0]!.label };
|
||
expect(phase(call).preReview).toBe(false);
|
||
for (const invalid of [{ ...call, answered: false }, { ...call, failed: true }, { ...call, answers: {} }]) {
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(invalid, 0, true))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('only positive substantive native question identity starts review', () => {
|
||
for (const qid of ['plan-eng-review-scope-reduce', 'plan-eng-scope-complexity', 'cross-project-learnings', 'plan-eng-review-next-steps']) {
|
||
const q = { ...finding, question: finding.question.replace('plan-eng-arch-global-cache', qid) };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
}
|
||
for (const qid of ['plan-eng-review-arch-finding', 'plan-eng-review-test-gap']) {
|
||
const q = { ...finding, question: finding.question.replace('plan-eng-arch-global-cache', qid) };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(true);
|
||
}
|
||
for (const qid of ['plan-eng-arch-focus', 'plan-eng-test-focus', 'plan-eng-quality-mode', 'plan-eng-perf-next-steps']) {
|
||
const q = { ...finding, header: 'Review setup', question: `Choose which issue to review first. <gstack-qid:${qid}>` };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
}
|
||
const testScope = { ...finding, header: 'Test scope', question: 'D3 — Test Issue: no coverage of the rewritten legacy flow is specified. <gstack-qid:plan-eng-test-scope>' };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([testScope], { [testScope.question]: testScope.options[0]!.label }), 0, true))).toBe(true);
|
||
const q = { ...finding, header: 'Architecture', question: 'Choose a review focus. <gstack-qid:plan-eng-review-arch-finding>' };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng setup semantics (captured G)', () => {
|
||
// Exact native question stems, offered actions and answers. Five findings
|
||
// and both substantive TODO decisions must remain in review coverage.
|
||
const rows = [
|
||
{
|
||
"header": "Cross-project",
|
||
"question": "gstack can search learnings from your other projects on this machine to find patterns that might apply to this auth refactor review. This stays local \u2014 no data leaves your machine. Enable cross-project learnings? <gstack-qid:cross-project-learnings>",
|
||
"options": [
|
||
"Enable (recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable (recommended)"
|
||
},
|
||
{
|
||
"header": "Scope",
|
||
"question": "D1 \u2014 Scope challenge: this plan touches 12 files and introduces 4 new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-challenge>",
|
||
"options": [
|
||
"Reduce: phase it (recommended)",
|
||
"Proceed as-is",
|
||
"Investigate first"
|
||
],
|
||
"answer": "Reduce: phase it (recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: cache wiring",
|
||
"question": "D2 \u2014 Architecture issue 1: AuthBroker and SessionMint share a global mutable AuthCache via module-level export. Module-level singletons prevent test isolation and break in multi-process deployments (cluster, serverless, worker threads). How should AuthCache be wired? <gstack-qid:plan-eng-arch-cache-wiring>",
|
||
"options": [
|
||
"Dependency injection (recommended)",
|
||
"Keep module-level export"
|
||
],
|
||
"answer": "Dependency injection (recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: concurrency",
|
||
"question": "D3 \u2014 Architecture issue 2: the plan explicitly states mutations to AuthCache are not serialized. With two services writing concurrently (e.g., AuthBroker evicting a token while SessionMint reads-then-writes it), you get classic check-then-act races. How should this be handled? <gstack-qid:plan-eng-arch-mutation-serialization>",
|
||
"options": [
|
||
"Make mutations idempotent + last-write-wins (recommended)",
|
||
"Add an async mutex per cache key",
|
||
"Document and defer"
|
||
],
|
||
"answer": "Make mutations idempotent + last-write-wins (recommended)"
|
||
},
|
||
{
|
||
"header": "Code quality",
|
||
"question": "D4 \u2014 Code quality issue 1: validateAndDispatch() is 60 lines with three nested try/catch blocks, each silently swallowing a different error class. Swallowed errors mean silent failures in production \u2014 a token validation error looks identical to a dispatch error from the outside. Fix approach? <gstack-qid:plan-eng-cq-error-swallowing>",
|
||
"options": [
|
||
"Extract + typed errors (recommended)",
|
||
"Add structured logging before swallowing",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Extract + typed errors (recommended)"
|
||
},
|
||
{
|
||
"header": "Tests",
|
||
"question": "D5 \u2014 Test issue: all 13 new code paths in AuthCache and AuthBroker are untested (0% coverage planned). The plan mentions unit and integration coverage for success/error paths, but specifics are absent. Add explicit test requirements to the plan now? <gstack-qid:plan-eng-test-coverage>",
|
||
"options": [
|
||
"Add explicit test plan (recommended)",
|
||
"Keep plan vague, trust implementation"
|
||
],
|
||
"answer": "Add explicit test plan (recommended)"
|
||
},
|
||
{
|
||
"header": "Performance",
|
||
"question": "D6 \u2014 Performance issue: token validation makes 5 sequential IDP API calls. The plan notes they are independent and could be parallelized via Promise.all trivially. Fix in this PR or defer? <gstack-qid:plan-eng-perf-idp-parallelization>",
|
||
"options": [
|
||
"Fix in this PR with Promise.all (recommended)",
|
||
"Defer to follow-up TODO"
|
||
],
|
||
"answer": "Fix in this PR with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"header": "TODO: PR 2",
|
||
"question": "D7 \u2014 TODO: capture PR 2 scope (SessionMint, TokenStore, RequestPolicy) in TODOS.md so it doesn\u2019t get lost after PR 1 ships. Add it? <gstack-qid:plan-eng-todo-pr2-scope>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Skip \u2014 not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
},
|
||
{
|
||
"header": "TODO: IDP retry",
|
||
"question": "D8 \u2014 TODO: IDP circuit breaker. Token validation makes 5 IDP calls (now parallelized). If the IDP is degraded, all 5 fail together \u2014 no retry, no fallback, no circuit breaker in the plan. Add a TODO to add circuit breaker / exponential retry logic around IDP calls? <gstack-qid:plan-eng-todo-idp-circuit-breaker>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Build it now in this PR",
|
||
"Skip \u2014 not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
}
|
||
];
|
||
const callAt = (index: number) => {
|
||
const row = rows[index]!;
|
||
return { sessionId: 'captured-g-eng', toolUseId: `g-${index}`, answered: true, failed: false,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer } };
|
||
};
|
||
const setup = (call: ReturnType<typeof callAt>) => engSetupAUQ(nativePlanCallFingerprint(call, 0, false));
|
||
const change = (index: number, question: string, labels?: string[]) => {
|
||
const call = callAt(index); const q = call.questions[0]!;
|
||
q.question = question;
|
||
if (labels) q.options = labels.map(label => ({ label }));
|
||
call.answers = { [question]: q.options[0]!.label };
|
||
return call;
|
||
};
|
||
test('two setup calls precede five findings and two substantive TODO calls', () => {
|
||
for (const order of [[0, 1], [1, 0]]) {
|
||
let started = false;
|
||
const phases = [...order, 2, 3, 4, 5, 6, 7, 8].map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted; return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
}
|
||
expect(setup(callAt(0))).toBe(true);
|
||
expect(setup(callAt(1))).toBe(true);
|
||
for (const index of [2, 3, 4, 5, 6, 7, 8]) expect(setup(callAt(index))).toBe(false);
|
||
});
|
||
test('whole-plan scope uses its premise and opposed actions, not a model-chosen qid or header', () => {
|
||
for (const question of [
|
||
'D1 — Scope challenge: this plan touches 12 files and introduces 4 new classes. Reduce scope or proceed as-is?',
|
||
'The plan introduces 4 new services. Reduce the scope or proceed as-is? <gstack-qid:plan-eng-unfamiliar-size-check>',
|
||
'Complexity check: the plan spans 12 files. Reduce scope or proceed as-is? <gstack-qid:another-scope-name>',
|
||
]) {
|
||
const call = change(1, question); call.questions[0]!.header = 'Review setup';
|
||
expect(setup(call)).toBe(true);
|
||
call.questions[0]!.options.reverse();
|
||
call.answers = { [question]: call.questions[0]!.options[0]!.label };
|
||
expect(setup(call)).toBe(true);
|
||
}
|
||
});
|
||
test('abbreviated enable requires explicit cross-project learnings and an opposed project-scoped action', () => {
|
||
for (const question of [
|
||
'Should gstack search learnings from your other projects on this machine?',
|
||
'Enable cross-project learnings for this review? <gstack-qid:model-selected-learning-scope>',
|
||
]) {
|
||
const call = change(0, question); call.questions[0]!.header = 'Review setup';
|
||
expect(setup(call)).toBe(true);
|
||
call.questions[0]!.options.reverse();
|
||
call.answers = { [question]: call.questions[0]!.options[0]!.label };
|
||
expect(setup(call)).toBe(true);
|
||
}
|
||
for (const question of ['Enable the new feature?', 'Choose setup. <gstack-qid:cross-project-learnings>']) {
|
||
expect(setup(change(0, question))).toBe(false);
|
||
}
|
||
expect(setup(change(0, rows[0]!.question, ['Enable (recommended)', 'Discuss later']))).toBe(false);
|
||
});
|
||
test('individual issues and TODOs cannot borrow whole-plan scope or cross-project words', () => {
|
||
for (const [index, question, header] of [
|
||
[1, 'D1 — Architecture issue: this plan touches 12 files and adds 4 classes, but the AuthCache has a race. Reduce scope or proceed as-is?', 'Architecture issue'],
|
||
[1, 'D1 — Scope challenge: this cache spans 12 files and introduces 4 classes. Reduce scope or proceed as-is?', 'Scope'],
|
||
[1, 'D1 — Test gap: the plan spans 12 files. Reduce test scope or proceed as-is?', 'Tests'],
|
||
[1, 'The plan spans 12 files. Reduce scope or proceed as-is?', 'TODO: deferred implementation'],
|
||
[0, 'D1 — Security issue: cross-project learnings leak client data. Enable the feature or use project-scoped storage?', 'Security issue'],
|
||
] as const) {
|
||
const call = change(index, question); call.questions[0]!.header = header;
|
||
expect(setup(call)).toBe(false);
|
||
}
|
||
expect(setup(change(1, rows[1]!.question, ['Reduce token scope', 'Investigate']))).toBe(false);
|
||
});
|
||
test('the captured retry scope heading and letter-prefixed native actions remain setup', () => {
|
||
const retry = {
|
||
"header": "Scope",
|
||
"question": "D2 \u2014 Scope reduction: 12 files + 4 new classes exceeds the complexity threshold\nProject/branch/task: Multi-tenant Auth Refactor on main\nELI10: The plan introduces 4 new classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files. That\u2019s a lot of new surface area at once. Auth refactors are already high-risk (broken auth = all users locked out). Adding 4 new abstractions simultaneously makes the blast radius of a mistake much larger. A simpler split \u2014 just AuthBroker + SessionMint, collapsing TokenStore into AuthCache and inlining RequestPolicy \u2014 would achieve the same goal with 2 new classes and fewer files touched.\nStakes if we pick wrong: With 4 new classes landing together, a single bug in any one of them could take down auth for all tenants simultaneously. Fewer classes = smaller blast radius, easier rollback, faster onboarding for the next engineer.\nRecommendation: B (proceed as-is) \u2014 the plan already has the existing cache adapter retained, and the 4-class split may reflect genuine domain separations the plan description doesn\u2019t fully explain. Review the design at full scope, flag individual issues per section.\nCompleteness: Note: options differ in kind, not coverage \u2014 no completeness score.\nPros / cons:\nA) Reduce scope \u2014 collapse TokenStore into AuthCache, inline RequestPolicy\n \u2705 Fewer moving parts: 2 new classes instead of 4, smaller blast radius if auth fails\n \u2705 Easier to review, test, and roll back each piece independently\n \u274c May discard intentional domain separation the plan author had in mind\n \u274c Requires re-planning before implementation can start\nB) Proceed as-is with full review (recommended)\n \u2705 Respects the planned architecture and lets the full review surface real issues per section\n \u2705 Faster path to implementation if the 4-class split turns out to be justified\n \u274c Higher blast radius: 4 simultaneous new classes touching 12 files is more fragile to ship\n \u274c The legacyAuthFlow rewrite without a regression test is a landmine that needs explicit attention\nNet: You\u2019re trading blast-radius safety (fewer classes) against re-planning delay. Recommend B \u2014 proceed at full scope, but treat each class boundary and the missing regression test as explicit issues in the review. <gstack-qid:plan-eng-review-scope-challenge>",
|
||
"options": [
|
||
"A) Reduce scope",
|
||
"B) Proceed as-is (Recommended)"
|
||
],
|
||
"answer": "A) Reduce scope"
|
||
};
|
||
const call = change(1, retry.question, retry.options);
|
||
call.questions[0]!.header = retry.header;
|
||
call.answers = { [retry.question]: retry.answer };
|
||
expect(setup(call)).toBe(true);
|
||
// A numerical component issue cannot borrow the whole-plan sentence in
|
||
// the explanatory body and the same Reduce/Proceed choices.
|
||
const issue = change(1, retry.question.replace('Scope reduction:', 'AuthCache issue:'), retry.options);
|
||
expect(setup(issue)).toBe(false);
|
||
const partial = change(1, retry.question.replace('ELI10: The plan introduces', 'ELI10: AuthCache introduces'), retry.options);
|
||
expect(setup(partial)).toBe(false);
|
||
const missingOpposition = change(1, retry.question, ['A) Reduce scope', 'B) Investigate']);
|
||
expect(setup(missingOpposition)).toBe(false);
|
||
});
|
||
test('late setup cannot hide earlier or later substantive findings', () => {
|
||
const order = [2, 1, 3, 0, 4, 5, 6, 7, 8];
|
||
let started = false;
|
||
const phases = order.map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted; return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([false, true, false, true, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
});
|
||
test('pending, failed, uncertain and mixed answered calls never become wholly setup', () => {
|
||
for (const index of [0, 1]) {
|
||
const original = callAt(index);
|
||
expect(engSetupAUQ({ ...nativePlanCallFingerprint(original, 0, false), nativeCall: undefined })).toBe(false);
|
||
for (const mutate of [
|
||
(c: ReturnType<typeof callAt>) => { c.answered = false; },
|
||
(c: ReturnType<typeof callAt>) => { c.failed = true; },
|
||
(c: ReturnType<typeof callAt>) => { c.answers = {}; },
|
||
(c: ReturnType<typeof callAt>) => { c.answers = { [c.questions[0]!.question]: 'Uncertain; I have not selected a choice' }; },
|
||
]) { const c = structuredClone(original); mutate(c); expect(setup(c)).toBe(false); }
|
||
const mixed = structuredClone(original); const finding = callAt(2).questions[0]!;
|
||
mixed.questions.push(finding);
|
||
mixed.answers[finding.question] = finding.options[0]!.label;
|
||
expect(setup(mixed)).toBe(false);
|
||
expect(planCountQuestionPhase(nativePlanCallFingerprint(mixed, 0, true), false,
|
||
engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
mixed.answers = { [finding.question]: finding.options[0]!.label };
|
||
expect(setup(mixed)).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('completed permission cannot become a queued review answer (captured G)', () => {
|
||
// Exact captured post-Write frame; the temporary repository path is sanitized.
|
||
// The damaged "wat" came from the CLI redraw, not the underlying question.
|
||
const captured = " real,specificgaps (Visual Hierachy, Spacing, Color, Typography,Motion)\r22-Itpreservesstrongaccessibilityandresponsivespecsfromtheexistingbehaviordescription\r23 -DESIGN.md exist andsuppliescorrectvaluesforall5gaps\r24\r 25 A 10/10would:\r26-Specifytheexactchangeforeachgap(concretetoken,before→after)\r27 -Add intection state table(loading,empty,error,success,partial)\r28-Storyboardtheuserjourney\r 29- Call out which DESIGN.md tokens each fix applies\r30-Resolveallimplementationambiguitiesbeforeadeveloperpicksitup\r31\r32 **DESIGN.md status:** Presnt and detailed. Spacing scale, clor system, typography,interacton patterns,andc\r omponentbehavior re all specified. This review calbrates against it.\r33\r34 **Exsting design lverage:**\r35 -Button, Field, InlineStatus,ErrorSummary,ConfirmationDialogcomponentsreusedas-is\r36-Focusring:2pxsolid#1d4ed8,offset2px(alreadyspecified)\r37 -8px bae spacingscalealreadyestablished\r38\r39 ---\r40\r41 ## Review Passes\r42\r43*(Populatedprogressivelyasthereviewproceeds)*\r44\r╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌\r Do you wat to create gstack-test-plan-design.md?\r❯ 1. Yes\r2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)for this session; Yes,an\r always allow access to /tmp/fixture for this session\r 3.Nohift+tab)\r\rEsc to cancl ·Tab to amend\u000f\u001b[?1000h\u001b[?1002h\u001b[?1003h\u001b[?1006h\rRan1shellcommand\r\r●Scopegate:planmodeactive—auto-selectingB(reviewingPLAN.mdintheworkingdirectory).\r\rSkippingroutinginjection—thisisareviewfixturerepo.Nowrunningthepre-reviewaudit.\r\rRead2 files, ran1 shell command\r\r●Now let me read the review sectons fil as required by theskill.\r\rRead file, ran 1shellcommand\r\r●Nopriorlearnings found. Cross-project config isnot set — will ask once Ihavetheuser'sattention.\r\rPreReview System Audit:\r- DESIGN.md:exists and comprehensive (spacing, typography, color, interactionpatternsallspecified)\r- UIscope:APPUIclassifier—account-settingsformwithProfile+Notificationssections\r- 5 gaps explicily listed in the plan; l have corresponding specin DESIGN.md\r-N rior design reviews inth reviewlog\r- NoTODOS.mdexists\r\rStep0 Assessment:\r\rInitial scor: 5/10. The plan correctly identifies5realgapsandpreservesstrongaccessibilityandresponsivespecs\rinheritedfromtheexistingbehaviordescription.A10/10wouldspecifyHOWeachgapisresolved(concretecomponent\rchanges,tokenvalues,before/afterlayout),notjustnametheproblem.TheDESIGN.mdalreadyhasallcorrectvalues;\rtheplanjustneedstoreferencethem.\r\rWritingtheinitialplanfilenowbeforeaskingStep0D.\r\r●Write(/tmp/fixture/gstack-test-plan-design.md)\r\r✢ Undulatig… (2m 0s ·↓ 5.4 okens)\r ⎿ Tip:Use/btwtoaskaquicksidequestionwithoutinterruptingClaude'scurrentwork\r ● high · /effort\r────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r❯ \r────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r ⏸ plan mod on(shift+tab tocycle)·esctointerrupt·←foragents\u001b[?25h\u001b[?25l\r●No prior earnings found.Cross-projectconfigisnotset—willaskonceIhavetheuser'sattention.\r Pre-Review System Audit:\r-DESIGN.md:existsandcomprehensive(spacing,typography,color,interactionpatternsallspecified)\r- UI scope: APP UI classifier — account-settings fomwith Prfile + Notifications sctions\r-5gapsexplicitlylistedintheplan;allhavecorrespondingspecinDESIGN.md\r- Noprior desigreviwsin the reviewlog\r-NoTODOS.mdexists\r\rStep0Assessment:\r\rInitialscore:5/10.Theplancorrectlyidentifies5realgapsandpreservesstrongaccessibilityandresponsivespecs\r inheted from theexisting behavir desription.A 10/10would specify HOWechgapi eolved (ccretecomponent\rchanges,tokenvalues,before/afterlayout),notjustnametheproblem.TheDESIGN.mdalreadyhasallcorrectvalues;\rth plan just neds to referencethem.\r\rWriting theinitial plan lenow before ask Step 0D.\r\r●Write(/tmp/fixture/gstack-test-plan-design.md)\r⎿ Wrote 44 lines";
|
||
|
||
test('a missed grant followed by native Write completion sends no stale answer', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(classifyPlanCountFrame(captured)).toBeNull();
|
||
expect(guard(captured)).toBe('handled');
|
||
expect(capturePlanCountQuestion(captured, new Set(), 0, true)).toBeNull();
|
||
});
|
||
|
||
test('the same damaged active file permission grants once and never counts as a finding', () => {
|
||
const menu = captured.slice(captured.indexOf('Do you wat'), captured.indexOf('Esc to cancl')) + 'Esc to cancl ·Tab to amend';
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(guard(menu)).toBe('grant');
|
||
expect(guard(menu)).toBe('handled');
|
||
expect(capturePlanCountQuestion(menu, new Set(), 0, false)).toBeNull();
|
||
const completed = menu + '\n⎿ Wrote 44 lines';
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + menu)).toBe('grant');
|
||
});
|
||
|
||
test('plain legacy file decisions remain questions without native permission controls', () => {
|
||
const frame = 'Do you want to create first.md?\n❯1.Create the reviewed file\n2.Keep the current layout';
|
||
expect(classifyPlanCountFrame(frame)).toBeNull();
|
||
expect(createPlanCountPermissionGuard()(frame)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)?.options).toHaveLength(2);
|
||
});
|
||
|
||
test('a matching native finding can discuss file permissions without being consumed', () => {
|
||
const question = {
|
||
header: 'File policy',
|
||
question: 'Should we create a file that documents always allow access to the project?',
|
||
options: [{ label: 'Create it' }, { label: 'Keep current policy' }],
|
||
};
|
||
const pending = { sessionId: 'file-policy', toolUseId: 'file-finding', answered: false, questions: [question] };
|
||
const frame = captured + '\n☐ ' + question.header + '\n' + question.question +
|
||
'\n❯1.Create it\n2.Keep current policy\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||
expect(classifyPlanCountFrame(frame)).toBeNull();
|
||
expect(createPlanCountPermissionGuard()(frame)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, pending)?.nativeCall).toBe(pending);
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)?.promptSnippet).toContain('File policy');
|
||
});
|
||
});
|
||
|
||
|
||
describe('native question identity outranks permission wording', () => {
|
||
const question = {
|
||
header: 'File policy',
|
||
question: 'D1 — Should we create a file that documents always allow access to the project? <gstack-qid:plan-eng-review-file-policy>',
|
||
options: [{ label: 'Create it' }, { label: 'Keep current policy' }],
|
||
};
|
||
const pending = { sessionId: 'file-policy', toolUseId: 'finding', answered: false, questions: [question] };
|
||
const frame = '☐ File policy\n' + question.question + '\n❯1.Create it\n2.Keep current policy\nEnter to select · ↑/↓ to navigte · Esc to cancel';
|
||
test('full native question and every option establish identity despite a damaged footer', () => {
|
||
// This is the formerly conflicting pure classifier result. The counting
|
||
// loop must consult native identity before taking its permission action.
|
||
expect(classifyPlanCountFrame(frame)).toBe('permission');
|
||
expect(matchesNativePlanQuestion(frame, pending)).toBe(true);
|
||
const seen = new Set<string>();
|
||
expect(capturePlanCountQuestion(frame, seen, 0, false, pending)?.nativeCall).toBe(pending);
|
||
expect(capturePlanCountQuestion(frame, seen, 1, false, pending)).toBeNull();
|
||
});
|
||
test('same header, changed choices, missing identity and an overlaid real permission cannot borrow a native call', () => {
|
||
for (const different of [
|
||
frame.replace('Should we create a file', 'Should we delete the file'),
|
||
frame.replace('2.Keep current policy', '2.Allow all edits'),
|
||
frame.replace(question.question, 'A different question with the same header?'),
|
||
frame + '\nDo you want to create actual.md?\n❯1.Yes\n2.Yes, and switch to accept edits (auto-approve file edits and common file commands) for this session (shift+tab)\n3.No\nEsc to cancel · Tab to amend',
|
||
]) expect(matchesNativePlanQuestion(different, pending)).toBe(false);
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, { ...pending, failed: true })).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)).toBeNull();
|
||
});
|
||
});
|
||
|
||
|
||
describe('native file/class complexity gate classification', () => {
|
||
const rows = [
|
||
{
|
||
"id": "toolu_01HJHA5nCKyCAEPRVnm84udk",
|
||
"header": "Prerequisite",
|
||
"question": "D1 — No design doc found for this branch. Run /office-hours first? <gstack-qid:plan-eng-review-office-hours-prereq>",
|
||
"labels": [
|
||
"Skip — proceed with review (Recommended)",
|
||
"Run /office-hours first"
|
||
],
|
||
"answer": "Skip — proceed with review (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01BF2LE9pqkt3Cs8PxC1uPaA",
|
||
"header": "Scope",
|
||
"question": "D2 — Complexity check triggered: 12 files, 4+ new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-review-complexity-check>",
|
||
"labels": [
|
||
"Proceed as-is",
|
||
"Reduce scope (Recommended)"
|
||
],
|
||
"answer": "Proceed as-is"
|
||
},
|
||
{
|
||
"id": "toolu_01J9SZHLFAuHjn19o9XFJvHo",
|
||
"header": "Learnings",
|
||
"question": "D3 — Cross-project learnings: search past sessions from other projects on this machine? <gstack-qid:plan-eng-review-cross-project-learnings>",
|
||
"labels": [
|
||
"Enable cross-project (Recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01CRbNCM44doMT1ubMAtcuBp",
|
||
"header": "Architecture",
|
||
"question": "D4 — Architecture A1: Two services mutate a shared AuthCache with no mutation serialization. How should this be resolved? <gstack-qid:plan-eng-review-arch-shared-mutable-cache>",
|
||
"labels": [
|
||
"Single write-coordinator (Recommended)",
|
||
"Serialize via mutex/queue",
|
||
"Accept the risk, add monitoring"
|
||
],
|
||
"answer": "Single write-coordinator (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01UwpB3rC52W99ZB6g6FMV6K",
|
||
"header": "Architecture",
|
||
"question": "D5 — Architecture A2: legacyAuthFlow() gets rewritten with no regression tests capturing prior behavior. Approach? <gstack-qid:plan-eng-review-arch-legacy-rewrite>",
|
||
"labels": [
|
||
"Regression tests first (Recommended)",
|
||
"Strangler fig",
|
||
"Big-bang rewrite as planned"
|
||
],
|
||
"answer": "Regression tests first (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_011biwzYyT1qCzVL8rn8gmTe",
|
||
"header": "Code Quality",
|
||
"question": "D6 — Code Quality CQ1: validateAndDispatch() has 3 nested try/catch blocks that each swallow a different error class. Fix? <gstack-qid:plan-eng-review-cq-validate-dispatch>",
|
||
"labels": [
|
||
"Decompose + surface errors (Recommended)",
|
||
"Flatten catch hierarchy only",
|
||
"Leave as-is"
|
||
],
|
||
"answer": "Decompose + surface errors (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01HbUJX6fD6CN3obeoyh8S5Z",
|
||
"header": "Code Quality",
|
||
"question": "D7 — Code Quality CQ2: AuthCache shared via module-level export (implicit global). Switch to dependency injection? <gstack-qid:plan-eng-review-cq-module-export>",
|
||
"labels": [
|
||
"Dependency injection (Recommended)",
|
||
"Keep module export"
|
||
],
|
||
"answer": "Dependency injection (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_019jNc2fhvrJhHkNZKTdztUm",
|
||
"header": "Tests",
|
||
"question": "D8 — Tests T1: No described test for coordinator ordering under concurrent mutations. Add concurrency tests to the plan? <gstack-qid:plan-eng-review-test-concurrency>",
|
||
"labels": [
|
||
"Add concurrency tests (Recommended)",
|
||
"Defer to code review"
|
||
],
|
||
"answer": "Add concurrency tests (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_0132EDXRcCffruHaV1E4kKRH",
|
||
"header": "Tests",
|
||
"question": "D9 — Tests T2: No E2E tests planned for core auth flows (cross-tenant validation, suspension, logout, revocation). Add them? <gstack-qid:plan-eng-review-test-e2e-auth>",
|
||
"labels": [
|
||
"Add E2E tests for auth flows (Recommended)",
|
||
"Unit/integration only as planned"
|
||
],
|
||
"answer": "Add E2E tests for auth flows (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01C2sC9k9YbvAvgyWHGZUTyb",
|
||
"header": "Performance",
|
||
"question": "D10 — Performance P1: 5 sequential IDP calls per token validation. Parallelize or eliminate? <gstack-qid:plan-eng-review-perf-idp-calls>",
|
||
"labels": [
|
||
"Local JWT validation (Recommended)",
|
||
"Promise.all parallelization",
|
||
"Leave sequential as-is"
|
||
],
|
||
"answer": "Local JWT validation (Recommended)"
|
||
}
|
||
];
|
||
|
||
function callAt(index: number) {
|
||
const row = rows[index]!;
|
||
return {
|
||
sessionId: '122ddb02-2346-4a38-9824-f04f9d5d8cae',
|
||
toolUseId: row.id,
|
||
answered: true,
|
||
failed: false,
|
||
questions: [{
|
||
header: row.header,
|
||
question: row.question,
|
||
options: row.labels.map(label => ({ label })),
|
||
multiSelect: false,
|
||
}],
|
||
answers: { [row.question]: row.answer },
|
||
unansweredQuestionIndices: [],
|
||
};
|
||
}
|
||
|
||
test('actual ten calls remain three setup and seven independent review decisions in either setup order', () => {
|
||
for (const order of [[0, 1, 2], [0, 2, 1]]) {
|
||
let started = false;
|
||
const phases = [...order, 3, 4, 5, 6, 7, 8, 9].map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
}
|
||
});
|
||
|
||
test('the answered numeric complexity gate does not depend on qid or action order', () => {
|
||
for (const qid of ['plan-eng-review-complexity-check', 'model-chosen-size-gate']) {
|
||
const call = callAt(1);
|
||
const q = call.questions[0]!;
|
||
q.question = q.question.replace('plan-eng-review-complexity-check', qid);
|
||
for (const reverse of [false, true]) {
|
||
if (reverse) q.options.reverse();
|
||
call.answers = { [q.question]: q.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('a component finding, TODO, missing count, or missing whole-scope opposition stays substantive', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.header = 'Architecture issue'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.header = 'TODO: scope'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.question = call.questions[0]!.question.replace('Complexity check triggered:', 'AuthCache complexity issue:'); },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.question = call.questions[0]!.question.replace('12 files, 4+ new classes', '12 cache entries'); },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.options[1]!.label = 'Reduce cache lock scope'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.options[0]!.label = 'Investigate the cache'; },
|
||
]) {
|
||
const call = callAt(1);
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('failed, pending, unoffered and mixed answered calls cannot be discarded as setup', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof callAt>) => { call.answered = false; },
|
||
(call: ReturnType<typeof callAt>) => { call.failed = true; },
|
||
(call: ReturnType<typeof callAt>) => { call.answers = {}; },
|
||
(call: ReturnType<typeof callAt>) => { call.answers = { [call.questions[0]!.question]: 'Add a new cache repair instead' }; },
|
||
(call: ReturnType<typeof callAt>) => {
|
||
const issue = callAt(3);
|
||
call.questions.push(issue.questions[0]!);
|
||
Object.assign(call.answers, issue.answers);
|
||
},
|
||
]) {
|
||
const call = callAt(1);
|
||
mutate(call);
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
|
||
describe('explicit whole-plan scope complexity premise', () => {
|
||
const calls = [
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01RY4Ec4BVKuX7JRycy5Y7zB",
|
||
"questions": [
|
||
{
|
||
"question": "D1 — The plan's scope (12 files, 4 new classes) triggers the complexity smell check. Proceed as-is or reduce scope first? <gstack-qid:plan-eng-complexity-check>",
|
||
"header": "Scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Proceed as-is (recommended)",
|
||
"description": "Review all 4 planned classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files. The plan already flags this as a smell worth investigating — we'll scrutinize every class boundary during the Architecture section and call out anything that should be merged or deferred. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Reduce scope first",
|
||
"description": "Propose a minimal-surface version before reviewing: collapse TokenStore into the existing cache adapter extension and defer RequestPolicy extraction (it's just a slice of validateAndDispatch()). This yields ~2 new classes and ~6-8 files. The Architecture section would then review the leaner design. Completeness: 7/10 — defers policy extraction"
|
||
},
|
||
{
|
||
"label": "Skip complexity and go straight to issues",
|
||
"description": "Treat scope as fixed, skip the scope challenge, and dive directly into the 4 review sections (Architecture, Code Quality, Tests, Performance). Fastest path but misses the structural smell the plan itself flagged. Completeness: 7/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D1 — The plan's scope (12 files, 4 new classes) triggers the complexity smell check. Proceed as-is or reduce scope first? <gstack-qid:plan-eng-complexity-check>": "Proceed as-is (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:58:36.032Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01Vz57KiddxkZg3v4NiWKzsC",
|
||
"questions": [
|
||
{
|
||
"question": "D2 — Architecture: shared global mutable AuthCache between AuthBroker and SessionMint, with no serialization, in a multi-tenant system. <gstack-qid:plan-eng-arch-shared-cache>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Inject via constructor (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan says: 'Two new services share a global mutable AuthCache instance via module-level export. Both services mutate it' and 'they do not serialize mutations.' In a multi-tenant system, concurrent mutations from AuthBroker (e.g., evict-on-logout) and SessionMint (e.g., write-on-mint) to the same cache can corrupt tenant isolation boundaries without any test catching it. Fix: pass AuthCache as a constructor argument to both services. Each test provides a fresh instance; production wires it once at startup. Zero mutations to the existing adapter. Effort: human ~1h / CC ~5min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Add a mutation coordinator",
|
||
"description": "Keep the module-level export but wrap all mutating calls in a queue or lock. Protects against races but does not fix the coupling (tests still share state, import side-effects are still global). Higher complexity than DI. Effort: human ~2h / CC ~10min. Completeness: 8/10 — misses the testability problem"
|
||
},
|
||
{
|
||
"label": "Accept as-is",
|
||
"description": "Trust the underlying adapter's existing invalidation hooks to prevent cross-tenant reads. This works only if the adapter already serializes all writes — the plan does not state this, and 'they do not serialize mutations' explicitly says it does not. Accepted risk of cross-tenant cache corruption on concurrent requests. Completeness: 5/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D2 — Architecture: shared global mutable AuthCache between AuthBroker and SessionMint, with no serialization, in a multi-tenant system. <gstack-qid:plan-eng-arch-shared-cache>": "Inject via constructor (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:59:24.191Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01MpiyqoPfnhoFmN1hBtNHdu",
|
||
"questions": [
|
||
{
|
||
"question": "D3 — Architecture: the plan describes 5 IDP API calls for token validation but says nothing about partial failure semantics. <gstack-qid:plan-eng-arch-idp-partial-failure>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add explicit failure semantics to the plan (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan: 'Token validation issues 5 sequential API calls to the IDP.' No text covers what happens when call 3 of 5 returns a 503. In an auth system, the fail-open vs fail-closed decision is load-bearing for security. Recommendation: add a plan section stating the rule explicitly — 'on any IDP call failure, token validation fails closed (reject the token, return 401, do not cache a partial result).' Effort: human ~30min / CC ~3min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Defer to implementation",
|
||
"description": "Leave the failure semantics undecided in the plan; let implementers decide per-call. Risk: two developers make different assumptions and one path fails open (accepts a partially-validated token). In auth systems, a silent fail-open is a security hole, not just a bug. Completeness: 5/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D3 — Architecture: the plan describes 5 IDP API calls for token validation but says nothing about partial failure semantics. <gstack-qid:plan-eng-arch-idp-partial-failure>": "Add explicit failure semantics to the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:59:52.277Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_019eGcLXXteu52dC67WNtJyL",
|
||
"questions": [
|
||
{
|
||
"question": "D4 — Architecture: TokenStore is listed as one of 4 new classes but the plan never describes its responsibility. The existing cache adapter already stores tokens keyed by tenant/issuer/audience/policy. <gstack-qid:plan-eng-arch-tokenstore-purpose>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Define TokenStore's boundary in the plan (recommended)",
|
||
"description": "[P2] (confidence: 8/10) Without a description, implementers may duplicate the existing adapter's logic inside TokenStore, producing two storage layers for the same data. The plan should state: what TokenStore does that the existing adapter does not, whether it holds in-memory tokens separately from the cache adapter, and how AuthCache (the facade) relates to TokenStore vs the adapter. Effort: human ~20min / CC ~2min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Keep it implicit",
|
||
"description": "Accept that TokenStore's boundary will be defined during implementation. Risk: two interpretations surface mid-sprint — one where TokenStore wraps the adapter (double indirection) and one where it holds state independently (split storage, double-write bugs). Completeness: 5/10"
|
||
},
|
||
{
|
||
"label": "Merge TokenStore into AuthCache",
|
||
"description": "If TokenStore is just a typed wrapper around what AuthCache already exposes, drop it and let AuthCache handle all token storage concerns. Reduces new-class count from 4 to 3 without losing capability. Completeness: 9/10 — assumes they genuinely overlap"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D4 — Architecture: TokenStore is listed as one of 4 new classes but the plan never describes its responsibility. The existing cache adapter already stores tokens keyed by tenant/issuer/audience/policy. <gstack-qid:plan-eng-arch-tokenstore-purpose>": "Define TokenStore's boundary in the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:00:16.358Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01WXRQF8z4JKkKvnwH7h6U6Z",
|
||
"questions": [
|
||
{
|
||
"question": "D5 — Code Quality: validateAndDispatch() is 60 lines with 3 nested try/catch blocks, each swallowing a different error class. <gstack-qid:plan-eng-cq-validate-dispatch>",
|
||
"header": "Code Quality",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Refactor as part of this PR (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan: 'each catch swallows a different error class.' In auth code, a swallowed error is a silent failure — a rejection that should surface as a 401 becomes invisible, or a partial validation looks like success. Since this PR already touches the function (RequestPolicy extraction), refactoring it is a same-diff change. Fix: extract each catch branch into a named handler, propagate errors explicitly via throw or typed Result<T,E>; add one log line per catch so failures appear in traces. Effort: human ~2h / CC ~10min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Add logging, keep structure",
|
||
"description": "Add a structured log at each catch site so failures are at least visible, but leave the nesting and swallowing in place. Reduces debuggability debt without restructuring. Still leaves the logical errors silently absorbed — a 401 that should have been thrown may still become a phantom pass. Effort: human ~30min / CC ~5min. Completeness: 7/10 — errors are visible but not propagated"
|
||
},
|
||
{
|
||
"label": "Defer to follow-up",
|
||
"description": "Capture as a TODO and address in a later PR. Risk: the refactor grows harder once 4 new classes depend on the current swallowing behavior. The 'right behavior' becomes ambiguous when callers have already been written against the current (broken) semantics. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D5 — Code Quality: validateAndDispatch() is 60 lines with 3 nested try/catch blocks, each swallowing a different error class. <gstack-qid:plan-eng-cq-validate-dispatch>": "Refactor as part of this PR (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:00:40.446Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01UwMPiH1hZutbzPCGGQ3sUZ",
|
||
"questions": [
|
||
{
|
||
"question": "D6 — Tests: IDP partial failure (calls 1-4 of 5 succeed, call 5 fails) has no test coverage in the plan. We just added the fail-closed semantic in D3. <gstack-qid:plan-eng-test-idp-partial>",
|
||
"header": "Tests",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to plan as required test (recommended)",
|
||
"description": "[P1] (confidence: 9/10) We added 'fail closed on any IDP error' as explicit plan language in D3. That semantic needs a test that stubs each IDP call position as the one that fails (5 separate test cases, or one parametrized one) and asserts the validator returns 401 and writes nothing to cache. Without this, the fail-closed rule is a comment in a doc, not a contract in code. Effort: human ~1h / CC ~8min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Cover only all-pass and all-fail",
|
||
"description": "Test the two extreme cases (all 5 calls succeed, all fail) and skip partial-failure positions. Simpler, but misses the case where calls 1-4 succeeded and call 5 fails — the most likely real-world scenario when an IDP endpoint degrades under load. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Defer",
|
||
"description": "Capture as TODO and address post-merge. Risk: the fail-closed rule we just added has no verification path; a future refactor could silently revert it and no test would catch it. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D6 — Tests: IDP partial failure (calls 1-4 of 5 succeed, call 5 fails) has no test coverage in the plan. We just added the fail-closed semantic in D3. <gstack-qid:plan-eng-test-idp-partial>": "Add to plan as required test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:00.740Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_0152qGvSSuBnRkqTe8ix2WTj",
|
||
"questions": [
|
||
{
|
||
"question": "D7 — Tests: the plan says coverage for 'new components and their success/error paths' is planned, but validateAndDispatch() is an existing function being refactored. Its 3 catch blocks that currently swallow errors need explicit tests for each error branch. <gstack-qid:plan-eng-test-validate-dispatch-errors>",
|
||
"header": "Tests",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add 3 explicit error-path tests to the plan (recommended)",
|
||
"description": "[P2] (confidence: 8/10) After the refactor from D5, each of the 3 catch blocks becomes a named handler. Each named handler needs a test: trigger error class A/B/C, assert the function returns the expected error response (not silently succeeds). Without these tests, the refactor from D5 has no verification that the new explicit error handling is correct. Effort: human ~1h / CC ~8min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Cover in integration tests only",
|
||
"description": "Rely on integration tests that exercise validateAndDispatch() indirectly through AuthBroker. Lower signal — integration tests that cover error paths often don't isolate which branch triggered, making regressions hard to pinpoint. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Defer",
|
||
"description": "Leave error-path test coverage for a follow-up. Risk: the refactored function's error behavior is unverified until the next sprint. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D7 — Tests: the plan says coverage for 'new components and their success/error paths' is planned, but validateAndDispatch() is an existing function being refactored. Its 3 catch blocks that currently swallow errors need explicit tests for each error branch. <gstack-qid:plan-eng-test-validate-dispatch-errors>": "Add 3 explicit error-path tests to the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:10.780Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_0194VcsnAMMDRFEafkx4dEsQ",
|
||
"questions": [
|
||
{
|
||
"question": "D8 — Performance: token validation makes 5 sequential IDP API calls that the plan itself identifies as trivially parallelizable. <gstack-qid:plan-eng-perf-idp-parallel>",
|
||
"header": "Performance",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Parallelize as part of this PR (recommended)",
|
||
"description": "[P2] (confidence: 9/10) Plan: 'Token validation issues 5 sequential API calls to the IDP; they could be parallelized via Promise.all trivially (calls are independent).' The plan identifies the fix itself. In a multi-tenant system, token validation is on the hot path for every authenticated request. Sequential IDP calls multiply latency by 5x unnecessarily. Fix: wrap the 5 calls in Promise.all(); the refactor from D5 (validateAndDispatch extraction) is the natural point to do this. Effort: human ~30min / CC ~5min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Defer to performance pass",
|
||
"description": "Merge first, profile in production, then optimize. Risk: the plan already identifies this as trivial to fix. Deferring a known, safe, low-effort improvement costs real latency on every auth request from day 1 of deployment. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Add to TODOS.md only",
|
||
"description": "Capture in TODOS.md and address when performance metrics confirm it matters. Lowest risk of introducing a bug, but leaves 5x unnecessary serial latency in prod indefinitely. Completeness: 4/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D8 — Performance: token validation makes 5 sequential IDP API calls that the plan itself identifies as trivially parallelizable. <gstack-qid:plan-eng-perf-idp-parallel>": "Parallelize as part of this PR (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:26.856Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01AHbQGxj8dc6F4MPYhsPo2q",
|
||
"questions": [
|
||
{
|
||
"question": "D9 — TODOS: the plan has no mention of IDP circuit breaker or timeout per call. With 5 calls now running in parallel (D8 decision), an IDP outage generates 5 concurrent timeouts per request. <gstack-qid:plan-eng-todo-idp-circuit-breaker>",
|
||
"header": "TODO",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to TODOS.md (recommended)",
|
||
"description": "Capture as a deferred item: add per-call timeout + a circuit breaker so that IDP degradation fails fast instead of hanging. Not blocking this PR, but the parallelization we added in D8 increases the concurrent timeout surface. Concrete TODO: implement exponential backoff + circuit breaker with configurable open/half-open thresholds."
|
||
},
|
||
{
|
||
"label": "Skip — not valuable enough",
|
||
"description": "Accept that IDP timeout handling is the IDP library's responsibility, or that network timeouts are set at the HTTP client level. No additional app-level circuit breaker needed."
|
||
},
|
||
{
|
||
"label": "Build it now in this PR",
|
||
"description": "Add circuit breaker logic as part of the validateAndDispatch refactor (D5). The refactor is already touching that function; adding a circuit breaker is incremental. Effort: human ~4h / CC ~20min. Higher complexity, but IDP outages in auth systems have severe user impact."
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D9 — TODOS: the plan has no mention of IDP circuit breaker or timeout per call. With 5 calls now running in parallel (D8 decision), an IDP outage generates 5 concurrent timeouts per request. <gstack-qid:plan-eng-todo-idp-circuit-breaker>": "Add to TODOS.md (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:03:23.081Z"
|
||
}
|
||
];
|
||
|
||
test('captured retry counts its scope as setup but keeps all eight substantive decisions above the original ceiling', () => {
|
||
let started = false;
|
||
const phases = calls.map(call => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(call, 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, false, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(8);
|
||
});
|
||
|
||
test('an individual component, missing numeric premise, or a finding cannot borrow a scope label', () => {
|
||
for (const mutate of [
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace("The plan's scope", "The cache's scope"); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace('12 files, 4 new classes', '12 cached tokens'); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace('triggers the complexity smell check', 'has a race in AuthCache'); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.header = 'Architecture finding'; },
|
||
]) {
|
||
const call = structuredClone(calls[0]!);
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
|
||
describe('explicit Step 0 complexity gate with size in native choices', () => {
|
||
const calls = [
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01UbyixE6kojrR1enUcpg1bT",
|
||
"questions": [
|
||
{
|
||
"question": "D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>",
|
||
"header": "Routing setup",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add routing rules (recommended)"
|
||
},
|
||
{
|
||
"label": "Skip — invoke manually"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>": "Add routing rules (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:51:39.918Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01JswH5HdmeQvdAGx8P6tApN",
|
||
"questions": [
|
||
{
|
||
"question": "D2 — Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>",
|
||
"header": "Learnings scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Enable cross-project learnings (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep learnings project-scoped"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D2 — Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>": "Enable cross-project learnings (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:51:59.954Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01Hqr5i8DN2iukjCGmj1DQro",
|
||
"questions": [
|
||
{
|
||
"question": "D3 — Step 0 Scope Challenge: This plan triggers the complexity gate. Proceed at full scope or reduce? <gstack-qid:plan-eng-step0-complexity>",
|
||
"header": "Step 0 scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Proceed at full scope (recommended)",
|
||
"description": "Review and implement as written: 4 new classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files in one pass. The plan already self-flags this as a scope smell, and the reviewer will dig into the specific overbuilt pieces in the Architecture section.",
|
||
"preview": "Scope: 4 classes, 12 files, 2 new services\nRisk: larger blast radius; harder to roll back\nENG REVIEW will call out any specific over-abstractions"
|
||
},
|
||
{
|
||
"label": "Reduce to two phases",
|
||
"description": "Split into Phase 1 (core auth services: AuthBroker + SessionMint with direct adapter calls) and Phase 2 (TokenStore + RequestPolicy as follow-up). AuthCache facade may be mergeable into the existing adapter. Smaller blast radius, easier review. (human: +1 day planning / CC: ~10 min)",
|
||
"preview": "Phase 1: AuthBroker + SessionMint (core flow)\nPhase 2: TokenStore + RequestPolicy\nRisk: sequencing overhead; Phase 2 may stall"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D3 — Step 0 Scope Challenge: This plan triggers the complexity gate. Proceed at full scope or reduce? <gstack-qid:plan-eng-step0-complexity>": "Proceed at full scope (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:52:54.081Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01LT7ctyJNE1UqyTvLdQkULD",
|
||
"questions": [
|
||
{
|
||
"question": "D4 — Architecture A1: Global mutable AuthCache shared via module-level export, no mutation serialization. How should this be fixed? <gstack-qid:plan-eng-arch-global-cache>",
|
||
"header": "Arch: global cache",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Inject AuthCache as constructor dependency (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep global export, add write serialization"
|
||
},
|
||
{
|
||
"label": "Accept as-is — verify adapter serializes internally"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D4 — Architecture A1: Global mutable AuthCache shared via module-level export, no mutation serialization. How should this be fixed? <gstack-qid:plan-eng-arch-global-cache>": "Inject AuthCache as constructor dependency (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:53:50.210Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01LTXTWs6w4qhdo7FVRXFKdZ",
|
||
"questions": [
|
||
{
|
||
"question": "D5 — Architecture A2: Cache adapter failure posture is unstated. What should happen when the cache adapter is unavailable? <gstack-qid:plan-eng-arch-cache-failure>",
|
||
"header": "Arch: cache failmode",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Fail closed (recommended) — state this explicitly in the plan"
|
||
},
|
||
{
|
||
"label": "Fail open with logging — state this explicitly in the plan"
|
||
},
|
||
{
|
||
"label": "Accept as-is — inherited from existing adapter"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D5 — Architecture A2: Cache adapter failure posture is unstated. What should happen when the cache adapter is unavailable? <gstack-qid:plan-eng-arch-cache-failure>": "Fail closed (recommended) — state this explicitly in the plan"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:54:24.301Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01MCM7KMWBJyKPBeD5D5xQXE",
|
||
"questions": [
|
||
{
|
||
"question": "D6 — Code Quality Q1: validateAndDispatch() — 60 lines, 3 nested try/catch blocks each swallowing a different error class. How do you want to address this? <gstack-qid:plan-eng-cq-validate-dispatch>",
|
||
"header": "CQ: swallowed errors",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Refactor — split function + typed error returns (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep structure, convert to catch-and-log"
|
||
},
|
||
{
|
||
"label": "Defer — leave for follow-up PR"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D6 — Code Quality Q1: validateAndDispatch() — 60 lines, 3 nested try/catch blocks each swallowing a different error class. How do you want to address this? <gstack-qid:plan-eng-cq-validate-dispatch>": "Refactor — split function + typed error returns (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:55:06.409Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01THi9A5uG6H6YGD4FQgL6vJ",
|
||
"questions": [
|
||
{
|
||
"question": "D7 — Tests T1: The plan says success/error paths will be covered, but doesn't include a test for the fail-closed cache behavior (cache adapter unavailable → auth rejected). Add it? <gstack-qid:plan-eng-test-fail-closed>",
|
||
"header": "Tests: fail-closed",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add explicit fail-closed test (recommended)"
|
||
},
|
||
{
|
||
"label": "Note as TODO — rely on adapter's existing coverage"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D7 — Tests T1: The plan says success/error paths will be covered, but doesn't include a test for the fail-closed cache behavior (cache adapter unavailable → auth rejected). Add it? <gstack-qid:plan-eng-test-fail-closed>": "Add explicit fail-closed test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:56:50.701Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01379PhfvZ4iX43gReT87gob",
|
||
"questions": [
|
||
{
|
||
"question": "D8 — Tests T2: Concurrent mint + revoke for the same tenant has no planned test. AuthBroker and SessionMint both write to the shared AuthCache instance. Add a concurrency test? <gstack-qid:plan-eng-test-concurrent>",
|
||
"header": "Tests: concurrency",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add concurrent access test (recommended)"
|
||
},
|
||
{
|
||
"label": "Defer — note in TODOS.md"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D8 — Tests T2: Concurrent mint + revoke for the same tenant has no planned test. AuthBroker and SessionMint both write to the shared AuthCache instance. Add a concurrency test? <gstack-qid:plan-eng-test-concurrent>": "Add concurrent access test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:57:14.766Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_0112YwENaKQsigFNvBDmZiUa",
|
||
"questions": [
|
||
{
|
||
"question": "D9 — Performance P1: 5 sequential IDP calls during token validation — the plan identifies this as trivially parallelizable. Address it in this PR? <gstack-qid:plan-eng-perf-idp-calls>",
|
||
"header": "Perf: IDP calls",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Parallelize in this PR with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"label": "Defer to a follow-up PR"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D9 — Performance P1: 5 sequential IDP calls during token validation — the plan identifies this as trivially parallelizable. Address it in this PR? <gstack-qid:plan-eng-perf-idp-calls>": "Parallelize in this PR with Promise.all (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:57:34.827Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01YFVEdfATVrYA8LbWUmzN72",
|
||
"questions": [
|
||
{
|
||
"question": "D10 — TODO: Add p99 latency metric for IDP calls before/after Promise.all parallelization. Add to TODOS.md? <gstack-qid:plan-eng-todo-idp-metrics>",
|
||
"header": "TODO: IDP metrics",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to TODOS.md (recommended)"
|
||
},
|
||
{
|
||
"label": "Skip — not valuable enough"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D10 — TODO: Add p99 latency metric for IDP calls before/after Promise.all parallelization. Add to TODOS.md? <gstack-qid:plan-eng-todo-idp-metrics>": "Add to TODOS.md (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:58:35.032Z"
|
||
}
|
||
];
|
||
const scopeCall = () => structuredClone(calls[2]!);
|
||
|
||
test('captured sequence keeps three setup calls and all seven substantive decisions', () => {
|
||
let started = false;
|
||
const phases = calls.map(call => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(call, 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(phase => phase.preReview)).toEqual([true, true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(phase => !phase.preReview)).toHaveLength(7);
|
||
expect(calls.at(-1)!.questions[0]!.header).toBe('TODO: IDP metrics');
|
||
});
|
||
|
||
test('native offered-answer binding does not depend on model qid or option order', () => {
|
||
for (const reverse of [false, true]) {
|
||
const call = scopeCall();
|
||
const question = call.questions[0]!;
|
||
question.question = question.question.replace('plan-eng-step0-complexity', 'different-model-id');
|
||
if (reverse) question.options.reverse();
|
||
call.answers = { [question.question]: question.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
}
|
||
});
|
||
|
||
test('component findings, TODOs and incomplete scope evidence remain review decisions', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.header = 'Architecture finding'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.header = 'TODO: scope'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.question = call.questions[0]!.question.replace('This plan triggers the complexity gate', 'This cache triggers the complexity gate'); },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.question = call.questions[0]!.question.replace('Step 0 Scope Challenge:', 'Architecture issue:'); },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.description = 'Inspect 12 files for a cache race.'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.description = 'Implement as written: 4 new classes for token validation.'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.label = 'Proceed with a cache lock'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[1]!.label = 'Investigate cache failures'; },
|
||
]) {
|
||
const call = scopeCall();
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('pending, failed, unoffered and mixed answered packets cannot be discarded as setup', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof scopeCall>) => { call.answered = false; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.failed = true; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.answers = {}; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.answers = { [call.questions[0]!.question]: 'Add a new cache repair' }; },
|
||
(call: ReturnType<typeof scopeCall>) => {
|
||
const finding = structuredClone(calls[3]!);
|
||
call.questions.push(finding.questions[0]!);
|
||
Object.assign(call.answers, finding.answers);
|
||
},
|
||
]) {
|
||
const call = scopeCall();
|
||
mutate(call);
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|