mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 18:05:31 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
3531 lines
178 KiB
TypeScript
3531 lines
178 KiB
TypeScript
/**
|
||
* Deterministic unit tests for claude-pty-runner.ts behavior changes.
|
||
*
|
||
* Free-tier (no EVALS=1 needed). Runs in <1s on every `bun test`. Catches
|
||
* harness plumbing bugs before stochastic PTY runs surface them.
|
||
*
|
||
* Two surface areas tested:
|
||
*
|
||
* 1. Permission-dialog short-circuit in 'asked' classification: a TTY frame
|
||
* that matches BOTH isPermissionDialogVisible AND isNumberedOptionListVisible
|
||
* must NOT be classified as a skill question — permission dialogs render
|
||
* as numbered lists too, but they're not what we're guarding.
|
||
*
|
||
* 2. Env passthrough surface: runPlanSkillObservation accepts an `env`
|
||
* option and threads it to launchClaudePty. We can't fully exercise the
|
||
* spawn pipeline without paying for a PTY session, but we CAN verify the
|
||
* option exists in the type signature and that calling without env still
|
||
* works (no regression).
|
||
*
|
||
* The PTY test (skill-e2e-plan-ceo-plan-mode.test.ts) is the integration
|
||
* check; this file is the cheap deterministic guard for the harness primitives
|
||
* those tests stand on.
|
||
*/
|
||
|
||
import { describe, test, expect } from 'bun:test';
|
||
import { readFileSync } from 'node:fs';
|
||
import {
|
||
isPermissionDialogVisible,
|
||
isNumberedOptionListVisible,
|
||
isProseAUQVisible,
|
||
isScopeGateQuestionVisible,
|
||
isScopeGateAutoSelectVisible,
|
||
isPlanReadyVisible,
|
||
parseNumberedOptions,
|
||
classifyVisible,
|
||
TAIL_SCAN_BYTES,
|
||
optionsSignature,
|
||
parseQuestionPrompt,
|
||
stripAnsi,
|
||
auqFingerprint,
|
||
COMPLETION_SUMMARY_RE,
|
||
classifyPlanCountFrame,
|
||
capturePlanCountQuestion,
|
||
matchesNativePlanQuestion,
|
||
createPlanCountPermissionGuard,
|
||
planCountPrerequisitePick,
|
||
planCountSubmissionInput,
|
||
assertReviewReportAtBottom,
|
||
ceoStep0Boundary,
|
||
engStep0Boundary,
|
||
engSetupAUQ,
|
||
engFirstReviewAUQ,
|
||
designStep0Boundary,
|
||
designFirstReviewAUQ,
|
||
planCountQuestionPhase,
|
||
nativePlanCallFingerprint,
|
||
devexStep0Boundary,
|
||
type ClaudePtyOptions,
|
||
type AskUserQuestionFingerprint,
|
||
} from './claude-pty-runner';
|
||
|
||
describe('isPermissionDialogVisible', () => {
|
||
test('matches "Bash command requires permission" prompts', () => {
|
||
const sample = `
|
||
Some preamble output
|
||
|
||
Bash command \`gstack-config get telemetry\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. Yes, and always allow
|
||
3. No, abort
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches "allow all edits" file-edit prompts', () => {
|
||
// Isolated to the "allow all edits" clause only — no overlapping
|
||
// "Do you want to proceed?" co-trigger, so this asserts the clause works.
|
||
const sample = `
|
||
Edit to ~/.gstack/config.yaml
|
||
|
||
❯ 1. Yes
|
||
2. Yes, allow all edits during this session
|
||
3. No
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the "Do you want to proceed?" file-edit confirmation by itself', () => {
|
||
// Separate fixture so weakening this clause is detected by a dedicated test.
|
||
const sample = `
|
||
Edit to ~/.gstack/config.yaml
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches workspace-trust "always allow access to" prompt', () => {
|
||
const sample = `
|
||
Do you trust the files in this folder?
|
||
|
||
❯ 1. Yes, proceed
|
||
2. Yes, and always allow access to /Users/me/repo
|
||
3. No, exit
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('recognizes the captured collapsed native overwrite confirmation', () => {
|
||
const sample = [
|
||
'Doyouwanttooverwritegstack-test-plan-design.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
expect(isPermissionDialogVisible(sample.replace('Esctocancel·Tabtoamend', 'Enter to select'))).toBe(false);
|
||
});
|
||
|
||
test('the captured paired-CEO Edit grant is permission, not another review finding', () => {
|
||
const sample = [
|
||
'Do youwt to makehis dittogstack-test-plan-ceo-paired.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/gstack-paid-shard-EbUl9j/tmp/gstack-e2e-plan-ceo-paired-gkjAd5forthissession',
|
||
'(shift+tab)', '3.No', 'Esctocancel·Tabtoamend',
|
||
].join('\r');
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
expect(classifyPlanCountFrame(sample)).toBe('permission');
|
||
});
|
||
|
||
test('recognizes permission labels whose cursor-positioning spaces disappeared', () => {
|
||
expect(isPermissionDialogVisible('Yes,andalwaysallowaccessto/tmp/fixtureforthissession')).toBe(true);
|
||
expect(isPermissionDialogVisible('Yes,allowalleditsduringthissession')).toBe(true);
|
||
expect(isPermissionDialogVisible('Bashcommandrequirespermission')).toBe(true);
|
||
});
|
||
|
||
test('does NOT match a skill AskUserQuestion list', () => {
|
||
const sample = `
|
||
D1 — Premise challenge: do users actually want this?
|
||
|
||
❯ 1. Yes, validated
|
||
2. No, premise is wrong
|
||
3. Need more info
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT match a plan-ready confirmation', () => {
|
||
const sample = `
|
||
Ready to execute the plan?
|
||
|
||
❯ 1. Yes
|
||
2. No, keep planning
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT match a skill question that contains the bare phrase "Do you want to proceed?"', () => {
|
||
// Co-trigger requirement: "Do you want to proceed?" alone is not enough.
|
||
// It must appear with "Edit to <path>" or "Write to <path>" to count as
|
||
// a permission dialog. This guards against a skill question like
|
||
// "Do you want to proceed with HOLD SCOPE?" being mis-classified.
|
||
const sample = `
|
||
Choose your scope mode for this review.
|
||
Do you want to proceed?
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
3. SELECTIVE EXPANSION
|
||
`;
|
||
expect(isPermissionDialogVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT mis-match when adversarial prose includes "Edit to <path>" alongside the bare proceed phrase', () => {
|
||
// Adversarial fixture: a skill question whose body legitimately mentions
|
||
// "Edit to <path>" in prose AND ends with "Do you want to proceed?". The
|
||
// current co-trigger regex would mis-classify this as a permission
|
||
// dialog. We DO want this test to fail until the regex is tightened
|
||
// further (e.g., proximity constraint, or anchoring "Edit to" to a
|
||
// line-start). For now this is documented as a known limitation: a
|
||
// skill question that talks about "Edit to" in prose IS still treated
|
||
// as a permission dialog. The test asserts the current behavior so a
|
||
// future fix can flip it intentionally.
|
||
const sample = `
|
||
Plan: I will Edit to ./plan.md to capture the decision.
|
||
Do you want to proceed?
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
// KNOWN LIMITATION: the co-trigger fires here. Documented as a
|
||
// post-merge follow-up. Flip this assertion once the regex tightens.
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
});
|
||
|
||
describe('isNumberedOptionListVisible', () => {
|
||
test('matches a basic ❯ 1. + 2. cursor list', () => {
|
||
const sample = `
|
||
❯ 1. Option one
|
||
2. Option two
|
||
3. Option three
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on a single-option prompt', () => {
|
||
const sample = `
|
||
❯ 1. Only option
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('returns false when no cursor renders', () => {
|
||
const sample = `
|
||
Just some prose with 1. a numbered point and 2. another.
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('overlaps permission dialogs (this is why D5 short-circuits)', () => {
|
||
// The whole point of D5: this string matches BOTH classifiers, so the
|
||
// runner must consult isPermissionDialogVisible to disambiguate.
|
||
const sample = `
|
||
Bash command \`do-thing\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isNumberedOptionListVisible(sample)).toBe(true);
|
||
expect(isPermissionDialogVisible(sample)).toBe(true);
|
||
});
|
||
});
|
||
|
||
describe('scope-gate render detectors', () => {
|
||
// The verbatim announcement string from the plan-eng/plan-design SKILL.md
|
||
// templates. If the template rewording drifts, THIS fixture fails first —
|
||
// before the paid plan-mode smokes silently degrade to vacuous asserts.
|
||
const TEMPLATE_ANNOUNCEMENT =
|
||
'Scope gate: plan mode — auto-selected B (reviewing <target>).';
|
||
|
||
describe('isScopeGateQuestionVisible', () => {
|
||
test('matches the clean prose gate render (question + option bodies)', () => {
|
||
const sample = `
|
||
What should I review?
|
||
A) The current branch diff — the work in progress on this branch.
|
||
B) A plan or design doc I'll paste or point you to.
|
||
C) A specific file, directory, or path.
|
||
Recommendation: A when a branch diff exists, otherwise B.
|
||
`;
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the native numbered render (no lettered markers)', () => {
|
||
const sample = `
|
||
What should I review?
|
||
|
||
❯ 1. The current branch diff — the work in progress on this branch.
|
||
2. A plan or design doc I'll paste or point you to.
|
||
3. A specific file, directory, or path.
|
||
`;
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the PTY-collapsed render (stripAnsi squished spaces)', () => {
|
||
const sample = 'WhatshouldIreview?A)Thecurrentbranchdiff—theworkinprogress';
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('stays false on narration quoting only the question', () => {
|
||
const sample =
|
||
"Normally I'd ask 'What should I review?' but plan mode is active, so I'm proceeding.";
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on unrelated review prose', () => {
|
||
const sample = 'I will review the current branch diff and report findings.';
|
||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('isScopeGateAutoSelectVisible', () => {
|
||
test('matches the verbatim template announcement', () => {
|
||
expect(isScopeGateAutoSelectVisible(TEMPLATE_ANNOUNCEMENT)).toBe(true);
|
||
});
|
||
|
||
test('matches a real announcement with a concrete target', () => {
|
||
const sample =
|
||
'Scope gate: plan mode — auto-selected B (reviewing ~/.claude/plans/my-feature.md). Running the Design Doc Check next.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the PTY-collapsed announcement', () => {
|
||
const sample = 'Scopegate:planmode—auto-selectedB(reviewingPLAN.md).';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('stays false on narration about the behavior', () => {
|
||
const sample = "In plan mode I'd auto-select B and review the active plan.";
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on a VERBATIM QUOTE of the announcement (negation narration)', () => {
|
||
// The exact announcement line sits quoted in the skill context, so a
|
||
// model explaining why it is NOT firing it can reproduce it byte-exact
|
||
// inside quotes — that must not trip a must-stay-false assert.
|
||
const sample =
|
||
'Not in plan mode, so I won\'t announce "Scope gate: plan mode — auto-selected B (reviewing <target>)." and will ask instead.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('a later real render still matches after an earlier quoted mention', () => {
|
||
const sample =
|
||
'Earlier I said I would render "Scope gate: plan mode — auto-selected B (…)" and now:\n' +
|
||
'Scope gate: plan mode — auto-selected B (reviewing PLAN.md).';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches tense paraphrases WITH the announcement prefix (auto-selecting / auto-selects)', () => {
|
||
expect(
|
||
isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selecting B (reviewing the drafted plan).'),
|
||
).toBe(true);
|
||
expect(isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selects B.')).toBe(true);
|
||
});
|
||
|
||
test('stays false on tense paraphrases WITHOUT the announcement prefix', () => {
|
||
expect(isScopeGateAutoSelectVisible('Auto-selecting B since we are in plan mode.')).toBe(false);
|
||
});
|
||
|
||
test('stays false on AUTO_DECIDE preamble output', () => {
|
||
const sample = 'Auto-decided scope question → B (your preference). Change with /plan-tune.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('stays false on a bare "selected B" without the announcement prefix', () => {
|
||
const sample = 'I selected B as the review target.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('isProseAUQVisible', () => {
|
||
test('matches 4 lettered options A) B) C) D) at line starts (plan-eng prose AUQ shape)', () => {
|
||
const sample = `
|
||
What would you like me to review? Options:
|
||
A) Point me at an existing design doc or plan file (path).
|
||
B) Describe new work you're planning — I'll explore the codebase.
|
||
C) You meant /review for the diff already on this branch.
|
||
D) Something else (tell me).
|
||
Recommendation: A if you have a doc in mind, otherwise B.
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches 2 lettered options (minimum threshold)', () => {
|
||
const sample = `
|
||
A) First option
|
||
B) Second option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches 3 numbered options 1. 2. 3. without ❯ 1. cursor (autoplan prose AUQ shape)', () => {
|
||
const sample = `
|
||
What's the task? A few options:
|
||
1. You have a plan idea in mind — describe it.
|
||
2. You want to review an existing plan elsewhere.
|
||
3. You meant a different command — /plan-ceo-review etc.
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false when ❯ 1. cursor is present in the recent tail (native UI handled by isNumberedOptionListVisible)', () => {
|
||
const sample = `
|
||
❯ 1. First option
|
||
2. Second option
|
||
3. Third option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does NOT suppress numbered-prose detection when ❯ 1. is only in early scrollback (trust dialog)', () => {
|
||
// Boot trust dialog rendered ❯ 1. Yes at startup, then a long body of
|
||
// model output, then prose-rendered numbered options now. The historic
|
||
// ❯ 1. is in the full buffer but NOT in the recent tail. Should detect
|
||
// the prose AUQ.
|
||
const trustHeader = '❯ 1. Yes, trust\n 2. No\n';
|
||
const filler = 'x'.repeat(5000); // pushes trust dialog out of last 4KB tail
|
||
const proseAUQ = `\n 1. Review the docs\n 2. Investigate the code\n 3. Defer to next session\n❯ \n`;
|
||
const sample = trustHeader + filler + proseAUQ;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on single lettered option', () => {
|
||
const sample = `
|
||
A) Only one option mentioned in passing.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('matches 2 numbered options (threshold matches lettered branch — tails miss option 1)', () => {
|
||
const sample = `
|
||
1. First note.
|
||
2. Second note.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on a single numbered option', () => {
|
||
const sample = `
|
||
1. Only one option mentioned.
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('does not match mid-prose lettered text like "(see option B) above"', () => {
|
||
const sample = `
|
||
This refers to (see option B) above and also to point A) earlier.
|
||
`;
|
||
// The B) and A) markers are mid-line, not at line starts, so they don't count.
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('matches with leading whitespace and ❯ prefix on options', () => {
|
||
const sample = `
|
||
A) Option with whitespace prefix
|
||
❯ B) Option with cursor prefix
|
||
C) Another option
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('returns false on plain text with no option markers', () => {
|
||
expect(isProseAUQVisible('Just some plain text output from the model.')).toBe(false);
|
||
expect(isProseAUQVisible('')).toBe(false);
|
||
});
|
||
|
||
// Pattern 3: markdown bold-bullet options — office-hours renders its mode
|
||
// question this way under --disallowedTools, with no letter/number marker.
|
||
test('matches office-hours markdown bold-bullet mode question (Pattern 3)', () => {
|
||
const sample = `
|
||
> Before we dig in — what's your goal with this?
|
||
>
|
||
> - **Building a startup** (or thinking about it)
|
||
> - **Intrapreneurship** — internal project at a company, need to ship fast
|
||
> - **Hackathon / demo** — time-boxed, need to impress
|
||
> - **Open source / research** — building for a community
|
||
> - **Learning** — teaching yourself to code
|
||
❯
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('bold-bullets require a preceding interrogative — no "?" => false', () => {
|
||
// 3+ bold bullets but no question stem: this is a feature list, not an AUQ.
|
||
const sample = `
|
||
Here is what shipped:
|
||
- **Faster builds** via caching
|
||
- **Smaller binaries** through tree-shaking
|
||
- **Better errors** with source maps
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('a question with fewer than 3 bold bullets stays false (guard)', () => {
|
||
const sample = `
|
||
Which approach do you prefer?
|
||
- **Option one** is simpler
|
||
- **Option two** is faster
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('plain (non-bold) bullets after a question do not trigger Pattern 3', () => {
|
||
// Only bold bullets count — plain "- text" prose lists are too common.
|
||
const sample = `
|
||
What should we do about this?
|
||
- run the tests
|
||
- ship the fix
|
||
- file a follow-up
|
||
`;
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('Pattern 3 still defers to a live native cursor list (❯ 1.)', () => {
|
||
const sample = `
|
||
> What's your goal?
|
||
❯ 1. **Building a startup**
|
||
2. **Intrapreneurship**
|
||
3. **Hackathon**
|
||
`;
|
||
// The ❯1. cursor gate fires first — native list handling owns this.
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
// Pattern 4/5: collapsed-form prose AUQ. stripAnsi destroys the newlines +
|
||
// inter-word spaces, so a real prose AUQ arrives collapsed and defeats the
|
||
// line-anchored Patterns 1-3. These are the dominant Shape-B render mode in
|
||
// the plan-design smoke + floor timeouts — verbatim de-spinnered bytes from
|
||
// the real failing runs (bdm3sucql.output).
|
||
test('matches the real collapsed floor render (colon-delimited, Pattern 4/5)', () => {
|
||
const sample =
|
||
'The review is blocked on D1—reply withA, B, r Cabovetocontinue:' +
|
||
'- A(recommended): Spec thefull P1AskUserQuestioncopy in this review' +
|
||
'-B:LeaveP1copytotheimplementerwithstructuralrequirements' +
|
||
'C: Add a placeholder template to the plan';
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('matches the real collapsed plan-mode render (Recommendation + collapsed A)/B), Pattern 4/5)', () => {
|
||
const sample =
|
||
'Recommendation:A—writethecopynow.(recommended)A) Writ the fullcopy in thisdesign review— now.' +
|
||
'(recommended) Completeness:10/10 B) Leveit to theimplemente — task spec is enough.' +
|
||
'Reply withA (write the copy now)orB(leavetoimplementer)';
|
||
expect(isProseAUQVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('collapsed-form requires BOTH signals — single B) + word "recommendation" stays false', () => {
|
||
// Only one punctuated letter marker: the two-signal contract is not met.
|
||
const sample =
|
||
'We should consider option B) here. My recommendation is to do it now.';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('collapsed-form requires letter punctuation — comma-only "ReplywithA,B,orC" stays false', () => {
|
||
// Reply-instruction present, but the letters carry no ) : or ( punctuation,
|
||
// so they could be incidental enumerations in running prose. Stays false.
|
||
const sample = 'ReplywithA,B,orC';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
|
||
test('collapsed-form does not regress the existing FP guard (see option B) ... point A))', () => {
|
||
// The classic citation FP: a model referencing prior options in prose.
|
||
// No reply-instruction / recommendation marker on its own line, so the
|
||
// collapsed-form signal does not fire either.
|
||
const sample =
|
||
'As noted (see option B) above, and the earlier point A) we discussed, this is fine.';
|
||
expect(isProseAUQVisible(sample)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('classifyVisible (runtime path through the runner classifier)', () => {
|
||
// These tests call the actual classifier so a future contributor who
|
||
// reorders branches (e.g. moves the permission short-circuit before
|
||
// isPlanReadyVisible) is caught deterministically.
|
||
|
||
test('skill question → returns asked', () => {
|
||
const visible = `
|
||
D1 — Choose your scope mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
3. SELECTIVE EXPANSION
|
||
4. SCOPE REDUCTION
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('permission dialog (Bash) → returns null (skip, keep polling)', () => {
|
||
const visible = `
|
||
Bash command \`gstack-update-check\` requires permission to run.
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
expect(isNumberedOptionListVisible(visible)).toBe(true); // pre-filter
|
||
expect(classifyVisible(visible)).toBeNull(); // post-filter
|
||
});
|
||
|
||
test('plan-ready confirmation → returns plan_ready (wins over asked)', () => {
|
||
const visible = `
|
||
Ready to execute the plan?
|
||
|
||
❯ 1. Yes, proceed
|
||
2. No, keep planning
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('plan_ready');
|
||
});
|
||
|
||
test('silent write to unsanctioned path → returns silent_write', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/dangerous-write.ts)
|
||
⎿ Wrote 42 lines
|
||
`;
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('silent_write');
|
||
expect(result?.summary).toContain('src/app/dangerous-write.ts');
|
||
});
|
||
|
||
test('write to sanctioned path (.claude/plans) → returns null (allowed)', () => {
|
||
const visible = `
|
||
⏺ Write(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Wrote 42 lines
|
||
`;
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('write while a permission dialog is on screen → returns null (gated, not silent, not asked)', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/edit-with-permission.ts)
|
||
|
||
Edit to src/app/edit-with-permission.ts
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
// The numbered prompt is a permission dialog (Edit to + Do you want to proceed?);
|
||
// silent_write is suppressed because a numbered prompt is visible, AND
|
||
// 'asked' is suppressed because the prompt is a permission dialog.
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('write while a real skill question is on screen → returns asked (write is captured but not silent)', () => {
|
||
const visible = `
|
||
⏺ Write(src/app/foo.ts)
|
||
|
||
D1 — Choose your scope mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
// The numbered prompt is a skill question, not a permission dialog;
|
||
// silent_write is suppressed (numbered prompt is visible) and the
|
||
// outcome is 'asked' — Step 0 fired.
|
||
const result = classifyVisible(visible);
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('idle / no signals → returns null', () => {
|
||
const visible = `
|
||
Some prose without any classifier signals.
|
||
`;
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
|
||
test('TAIL_SCAN_BYTES is exported as 1500', () => {
|
||
// Shared between runner and routing test; a regression that desyncs the
|
||
// recent-tail window would surface here.
|
||
expect(TAIL_SCAN_BYTES).toBe(1500);
|
||
});
|
||
|
||
// D4-B: strictPlanWrites detector. Catches the transcript bug where the
|
||
// model writes findings to the plan file before any AskUserQuestion fires.
|
||
test('strictPlanWrites: plan write before any AUQ → wrote_findings_before_asking', () => {
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('wrote_findings_before_asking');
|
||
expect(result?.summary).toContain('.claude/plans/some-plan.md');
|
||
});
|
||
|
||
test('strictPlanWrites: plan write AFTER an AUQ render → not flagged', () => {
|
||
// AUQ renders first, then the model writes the plan post-answer. This is
|
||
// the legitimate end-of-workflow flow and must NOT trigger the detector.
|
||
const visible = `
|
||
D1 — Some scope question
|
||
|
||
❯ 1. Option A
|
||
2. Option B
|
||
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
// Outcome is 'asked' (the numbered list rendered); the post-AUQ plan
|
||
// write is ignored by the detector.
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('strictPlanWrites: AUQ first then plan write — write_pos > auq_pos → not flagged', () => {
|
||
// Same scenario, more explicit ordering: the regex finds the write at a
|
||
// position AFTER the numbered list. Detector lets it through.
|
||
const visible = [
|
||
'D1 — Choose your approach',
|
||
'',
|
||
'❯ 1. Approach A',
|
||
' 2. Approach B',
|
||
'',
|
||
'⏺ Write(/Users/me/.claude/plans/draft.md)',
|
||
'⎿ Wrote 42 lines',
|
||
].join('\n');
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('asked');
|
||
});
|
||
|
||
test('strictPlanWrites: only a permission dialog visible → plan write still flagged', () => {
|
||
// A permission dialog ❯ 1./2. is NOT an AUQ; pre-AUQ plan writes still
|
||
// hit the detector even when a permission prompt is on screen.
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
|
||
Edit to /Users/me/.claude/plans/some-plan.md
|
||
|
||
Do you want to proceed?
|
||
|
||
❯ 1. Yes
|
||
2. No
|
||
`;
|
||
const result = classifyVisible(visible, { strictPlanWrites: true });
|
||
expect(result?.outcome).toBe('wrote_findings_before_asking');
|
||
});
|
||
|
||
test('strictPlanWrites OFF: plan write before AUQ → returns null (legacy behavior preserved)', () => {
|
||
const visible = `
|
||
⏺ Edit(/Users/me/.claude/plans/some-plan.md)
|
||
⎿ Updated 12 lines
|
||
`;
|
||
// Without strictPlanWrites, the sanctioned-path list lets this through.
|
||
expect(classifyVisible(visible)).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('parseNumberedOptions', () => {
|
||
test('does not combine an old AUQ prompt with the later ordinary test-case list', () => {
|
||
// B CEO retry, 2026-09-08: the old prompt cursor slid outside the
|
||
// option parser's 4KB window. Its prose fallback then supplied a new
|
||
// five-item test list while the prompt parser retained the old AUQ.
|
||
const visible = '☐Stripe event types\nWhich event should the handler accept?\n' +
|
||
'❯1.Specify one canonical event\n2.Accept all events\n' + '·'.repeat(4200) + '\n' +
|
||
'Minimum required test cases (all must be specified in the plan):\n' +
|
||
'1.Happypath:validcanonicalevent,knownuser→userupdated,emailsent\n' +
|
||
'2.Email failure:emailthrows→userupdated,errorlogged,HTTP200\n' +
|
||
'3.DB timeout: DB throws onuser update →exceptin ropagates, non-200\n' +
|
||
'4.Unkown event typ: non-canonical event→ HTTP200,nouserupdate\n' +
|
||
'5.Unknown user: valid event, usernotinDB→existingguard→HTTP200\n❯1\n';
|
||
const seen = new Set<string>();
|
||
expect(capturePlanCountQuestion(visible, seen, 0, false)).toBeNull();
|
||
expect(seen.size).toBe(0);
|
||
});
|
||
|
||
test('extracts options from a clean cursor list', () => {
|
||
const visible = `
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
const opts = parseNumberedOptions(visible);
|
||
expect(opts).toHaveLength(2);
|
||
expect(opts[0]).toEqual({ index: 1, label: 'HOLD SCOPE' });
|
||
expect(opts[1]).toEqual({ index: 2, label: 'SCOPE EXPANSION' });
|
||
});
|
||
|
||
test('returns empty array on prose-with-numbers (no cursor)', () => {
|
||
expect(parseNumberedOptions('text 1. one 2. two')).toEqual([]);
|
||
});
|
||
|
||
test('extracts options when the cursor is INLINE with prompt header (box-layout)', () => {
|
||
// Real /plan-ceo-review rendering: the TTY's cursor-positioning escapes
|
||
// collapse divider + header + prompt + cursor onto one logical line.
|
||
// Subsequent options (2..7) still start their own lines.
|
||
const visible = [
|
||
'────────────────────────────────────────',
|
||
'☐ Review scope What scope do you want me to CEO-review? ❯ 1. The branch\'s diff vs main',
|
||
' Review the full branch: ~10K LOC.',
|
||
'2. A specific plan file or design doc',
|
||
' You point me at a file (path) and I review that.',
|
||
'3. An idea you\'ll describe inline',
|
||
'4. Cancel — wrong skill',
|
||
'5. Type something.',
|
||
'────────────────────────────────────────',
|
||
'6. Chat about this',
|
||
'7. Skip interview and plan immediately',
|
||
].join('\n');
|
||
const opts = parseNumberedOptions(visible);
|
||
expect(opts).toHaveLength(7);
|
||
expect(opts[0]).toEqual({ index: 1, label: "The branch's diff vs main" });
|
||
expect(opts[1]?.index).toBe(2);
|
||
expect(opts[6]?.index).toBe(7);
|
||
expect(opts[6]?.label).toBe('Skip interview and plan immediately');
|
||
});
|
||
|
||
test('inline-cursor and start-of-line cursor both produce 7 options for the box-layout case', () => {
|
||
// The inline path captures option 1 from the cursor line itself; the
|
||
// subsequent-lines path captures 2..7 with the existing optionRe.
|
||
const inlineLayout = [
|
||
'header text ❯ 1. first option',
|
||
'2. second',
|
||
'3. third',
|
||
].join('\n');
|
||
expect(parseNumberedOptions(inlineLayout)).toEqual([
|
||
{ index: 1, label: 'first option' },
|
||
{ index: 2, label: 'second' },
|
||
{ index: 3, label: 'third' },
|
||
]);
|
||
|
||
const cleanLayout = [
|
||
' ❯ 1. first option',
|
||
' 2. second',
|
||
' 3. third',
|
||
].join('\n');
|
||
expect(parseNumberedOptions(cleanLayout)).toEqual([
|
||
{ index: 1, label: 'first option' },
|
||
{ index: 2, label: 'second' },
|
||
{ index: 3, label: 'third' },
|
||
]);
|
||
});
|
||
});
|
||
|
||
describe('pending native question on a damaged option render', () => {
|
||
// Exact final B CEO Test scope shape. The native call had been read in
|
||
// an in-progress snapshot, but option 2's missing dot prevented input.
|
||
const frame = [
|
||
'☐Test scope',
|
||
'│Section 6 (Tests) — Theplanhasnotestsforanewpaymentprocessingcodepath.Theexistingintegrationsuitehas',
|
||
'│never seen this handlerand cannotcatchregressionsinit.Minimumviabletestplanforminimalpatch:5unittests',
|
||
'│(happy path, mal failur, DB timeout, unknowneventtype,unknownuser).Shouldtheplanalsoincludeanintegration',
|
||
'│testhittingthefullwebhookstack?<gstack-qid:plan-ceo-test-scope>',
|
||
'❯1.Unittestsonlyfornow(recommended)',
|
||
'5 unit tests covering the criticalpaths. No integration stin v1.',
|
||
'2Uni tsts + one integration test',
|
||
'5 uit tsts + on ed-to-end integrationtestsendiga signe Stripeevent.',
|
||
'3.Integrationtestonly',
|
||
'4.Typesomething.',
|
||
'5. Chataboutthis',
|
||
'Enter to select · ↑/↓ to navigate · Esc to cancel',
|
||
'❯1',
|
||
].join('\n');
|
||
const pending = {
|
||
sessionId: '66fb6218-4a68-4f1a-a729-6407f14fd6b8',
|
||
toolUseId: 'toolu_017DicePqWNVyDsLCd2Y2MCi', answered: false,
|
||
questions: [{ header: 'Test scope', question: 'Should the plan also include an integration test hitting the full webhook stack? <gstack-qid:plan-ceo-test-scope>',
|
||
options: ['Unit tests only for now (recommended)', 'Unit tests + one integration test', 'Integration test only'].map(label => ({ label })) }],
|
||
};
|
||
|
||
test('uses lossless pending options after a positively matched native question has rendered', () => {
|
||
const seen = new Set<string>();
|
||
const captured = capturePlanCountQuestion(frame, seen, 0, false, pending);
|
||
expect(captured?.nativeCall).toBe(pending);
|
||
expect(captured?.options).toEqual(pending.questions[0].options.map((o, i) => ({ index: i + 1, label: o.label })));
|
||
expect(capturePlanCountQuestion(frame, seen, 1, false, pending)).toBeNull();
|
||
// A corrected redraw is still the same pending native question.
|
||
expect(capturePlanCountQuestion(frame.replace('2Uni tsts', '2.Unit tests'), seen, 2, false, pending)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame.replace('2Uni tsts', '2.Unit tests'), seen, 3, false)).toBeNull();
|
||
expect(seen.has(captured!.signature)).toBe(true);
|
||
});
|
||
|
||
test('binds delayed native metadata to the already-answered visible question', () => {
|
||
const seen = new Set<string>();
|
||
const clean = frame.replace('2Uni tsts', '2.Unit tests');
|
||
expect(capturePlanCountQuestion(clean, seen, 0, false)).not.toBeNull();
|
||
expect(capturePlanCountQuestion(clean, seen, 1, false, pending)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, seen, 2, false, pending)).toBeNull();
|
||
});
|
||
|
||
test('requires pending single-question metadata, matching current header, cursor, and navigation footer', () => {
|
||
for (const call of [undefined, { ...pending, answered: true }, { ...pending, failed: true },
|
||
{ ...pending, questions: [...pending.questions, ...pending.questions] },
|
||
{ ...pending, questions: [{ ...pending.questions[0], header: 'Prior decision' }] },
|
||
{ ...pending, questions: [{ ...pending.questions[0], question: 'Different issue <gstack-qid:plan-ceo-different-test-scope>' }] }]) {
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, call)).toBeNull();
|
||
}
|
||
for (const altered of [frame.replace('☐Test scope', 'Test scope'), frame.replace('❯1.', '1.'),
|
||
frame.replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
|
||
frame + '\n☐Different question\n❯1.Waiting for its choices']) {
|
||
expect(capturePlanCountQuestion(altered, new Set(), 0, false, pending)).toBeNull();
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('runPlanSkillObservation env passthrough surface', () => {
|
||
test('ClaudePtyOptions exposes env: Record<string, string>', () => {
|
||
// Type-level guard: this file would fail to compile if the env field
|
||
// were removed or its shape regressed. The actual env merge happens in
|
||
// launchClaudePty's spawn call (`env: { ...process.env, ...opts.env }`),
|
||
// so a regression where `env: opts.env` gets dropped from the
|
||
// runPlanSkillObservation -> launchClaudePty handoff is only caught by
|
||
// the live PTY test, not here.
|
||
const opts: ClaudePtyOptions = {
|
||
env: { QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' },
|
||
};
|
||
expect(opts.env).toEqual({ QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' });
|
||
});
|
||
});
|
||
|
||
describe('launchClaudePty model pin (static tripwire)', () => {
|
||
// Why static-grep, not a behavioral assert: the spawn fires immediately
|
||
// inside launchClaudePty, so asserting the built args array would require
|
||
// extracting an arg-builder seam — which rewrites the exact region kyoto-v5's
|
||
// hermetic --strict-mcp-config insertion edits, reintroducing a merge
|
||
// conflict the placement deliberately avoids. The end-to-end behavioral proof
|
||
// is the live PTY smoke (skill-e2e-plan-*-plan-mode.test.ts) running under the
|
||
// pinned model. These grep-level guards stop a refactor from silently
|
||
// dropping the pin or reordering it past extraArgs.
|
||
const src = readFileSync(new URL('./claude-pty-runner.ts', import.meta.url), 'utf-8');
|
||
|
||
test('ClaudePtyOptions exposes model?: string', () => {
|
||
const opts: ClaudePtyOptions = { model: 'claude-sonnet-4-6' };
|
||
expect(opts.model).toBe('claude-sonnet-4-6');
|
||
});
|
||
|
||
test('spawn args push --model from the EVALS_MODEL fallback chain', () => {
|
||
expect(src).toContain("args.push('--model', model)");
|
||
// opts.model -> EVALS_MODEL -> resolveEvalModel('capture') (mirrors session-runner.ts)
|
||
expect(src).toMatch(
|
||
/opts\.model\s*\?\?\s*process\.env\.EVALS_MODEL\s*\?\?\s*resolveEvalModel\('capture'\)/,
|
||
);
|
||
});
|
||
|
||
test('--model is pushed BEFORE extraArgs so a per-test --model override wins', () => {
|
||
const modelPush = src.indexOf("args.push('--model', model)");
|
||
const extraArgsPush = src.indexOf('if (opts.extraArgs) args.push(...opts.extraArgs)');
|
||
expect(modelPush).toBeGreaterThan(-1);
|
||
expect(extraArgsPush).toBeGreaterThan(-1);
|
||
expect(modelPush).toBeLessThan(extraArgsPush);
|
||
});
|
||
|
||
test('all three plan-skill wrappers forward model to launchClaudePty', () => {
|
||
// Count must match the number of wrappers (observation, counting, floor).
|
||
const forwards = src.match(/^\s*model: opts\.model,$/gm) ?? [];
|
||
expect(forwards.length).toBe(3);
|
||
});
|
||
});
|
||
|
||
// ────────────────────────────────────────────────────────────────────────────
|
||
// Per-finding count primitives — Section 3 unit tests #1–#5, #7, #12.
|
||
// ────────────────────────────────────────────────────────────────────────────
|
||
|
||
describe('optionsSignature', () => {
|
||
test('returns a "|"-joined `index:label` string for a clean list', () => {
|
||
const sig = optionsSignature([
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE EXPANSION' },
|
||
]);
|
||
expect(sig).toBe('1:HOLD SCOPE|2:SCOPE EXPANSION');
|
||
});
|
||
|
||
test('order-independent: shuffled inputs produce the same signature', () => {
|
||
// parseNumberedOptions already returns sorted, but defensive sort means
|
||
// a future caller that hands us shuffled input still produces a stable
|
||
// dedupe signature.
|
||
const a = optionsSignature([
|
||
{ index: 2, label: 'B' },
|
||
{ index: 1, label: 'A' },
|
||
{ index: 3, label: 'C' },
|
||
]);
|
||
const b = optionsSignature([
|
||
{ index: 1, label: 'A' },
|
||
{ index: 2, label: 'B' },
|
||
{ index: 3, label: 'C' },
|
||
]);
|
||
expect(a).toBe(b);
|
||
});
|
||
|
||
test('empty list returns empty string', () => {
|
||
expect(optionsSignature([])).toBe('');
|
||
});
|
||
|
||
test('single-item list returns just that entry', () => {
|
||
expect(optionsSignature([{ index: 1, label: 'Only' }])).toBe('1:Only');
|
||
});
|
||
});
|
||
|
||
describe('parseQuestionPrompt', () => {
|
||
test('keeps the captured boxed learnings header across native CR and blank borders', () => {
|
||
// Exact active-menu bytes from the targeted-a engineering batching run.
|
||
// Its answered setup AUQ lost the title at the standalone box border,
|
||
// leaving every later finding classified as preReview.
|
||
const raw = "☐ Learnings\u001b[K\r\u001b[1B\u001b[K\r\u001b[1B│ D1 — Cross-project learnings scope <gstack-qid:learnings-cross-project>\u001b[K\r\u001b[1B│\u001b[3G\u001b[K\r\r\n│\u001b[3Ggstack\u001b[10Gcan\u001b[14Gsearch\u001b[21Glearnings\u001b[31Gfrom\u001b[36Gyour\u001b[41Gother\u001b[47Gprojects\u001b[56Gon\u001b[59Gthis\u001b[64Gmachine\u001b[72Gto\u001b[75Gfind\u001b[80Gpatterns\u001b[89Gthat\u001b[94Gmight\u001b[100Gapply\u001b[106Ghere.\u001b[112GThis\r\r\n│\u001b[3Gstays\u001b[9Glocal\u001b[15G—\u001b[17Gno\u001b[20Gdata\u001b[25Gleaves\u001b[32Gyour\u001b[37Gmachine.\u001b[46GRecommended\u001b[58Gfor\u001b[62Gsolo\u001b[67Gdevelopers.\u001b[79GSkip\u001b[84Gif\u001b[87Gyou\u001b[91Gwork\u001b[96Gon\u001b[99Gmultiple\u001b[108Gclient\r\r\n│\u001b[3Gcodebases\u001b[13Gwhere\u001b[19Gcross-contamination\u001b[39Gwould\u001b[45Gbe\u001b[48Ga\u001b[50Gconcern.\r\r\n\r\r\n❯\u001b[3G1.\u001b[6GEnable\u001b[13Gcross-project\u001b[27Glearnings\u001b[37G(Recommended)\r\r\n\u001b[6GSearch\u001b[13Glearnings\u001b[23Gfrom\u001b[28Gall\u001b[32Gprojects\u001b[41Gon\u001b[44Gthis\u001b[49Gmachine\u001b[57G—\u001b[59Gsurfaces\u001b[68Gpatterns\u001b[77Gand\u001b[81Gpitfalls\u001b[90Gfrom\u001b[95Gprior\u001b[101Gsessions.\r\r\n\u001b[3G2.\u001b[6GKeep\u001b[11Glearnings\u001b[21Gproject-scoped\u001b[36Gonly\r\r\n\u001b[6GOnly\u001b[11Guse\u001b[15Glearnings\u001b[25Gfrom\u001b[30Gthis\u001b[35Gproject.\u001b[44GSafe\u001b[49Gfor\u001b[53Gmulti-client\u001b[66Genvironments.\r\r\n\u001b[3G3.\u001b[6GType\u001b[11Gsomething.\r\r\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r\r\n\u001b[3G4.\u001b[6GChat\u001b[11Gabout\u001b[17Gthis\r\r\n\r\r\nEnter\u001b[7Gto\u001b[10Gselect\u001b[17G·\u001b[19G↑/↓\u001b[23Gto\u001b[26Gnavigate\u001b[35G·\u001b[37GEsc\u001b[41Gto\u001b[44Gcancel";
|
||
const visible = stripAnsi(raw);
|
||
const question = capturePlanCountQuestion(visible, new Set(), 0, true)!;
|
||
expect(question.promptSnippet).toStartWith('Learnings D1 — Cross-project learnings scope');
|
||
expect(question.promptSnippet).toContain('<gstack-qid:learnings-cross-project>');
|
||
expect(engStep0Boundary(question)).toBe(true);
|
||
const phase = planCountQuestionPhase(question, false, engStep0Boundary);
|
||
expect(phase).toEqual({ preReview: true, reviewStarted: true });
|
||
});
|
||
|
||
test('keeps a long boxed question identity instead of its closing recommendation', () => {
|
||
const frame = [
|
||
'Planning: /tmp/hermetic/.claude/plans/review.md',
|
||
'─'.repeat(120),
|
||
'☐ Architecture',
|
||
'│ D2 — Architecture: custom retry scheduler vs library built-in <gstack-qid:arch-custom-retry-vs-library>',
|
||
'│',
|
||
...Array.from({ length: 12 }, (_, i) => `│ Review context line ${i}: the proposed retry behavior and its tradeoffs.`),
|
||
'│',
|
||
'│ Net: If the library hook is configurable, use the existing implementation.',
|
||
'❯1.Use library built-in (Recommended)',
|
||
'2.Extract shared retry envelope',
|
||
].join('\r\r\n');
|
||
const seen = new Set<string>();
|
||
const question = capturePlanCountQuestion(frame, seen, 0, false)!;
|
||
expect(question.promptSnippet).toStartWith('Architecture D2 — Architecture: custom retry scheduler');
|
||
expect(question.promptSnippet).toContain('<gstack-qid:arch-custom-retry-vs-library>');
|
||
expect(question.promptSnippet).not.toContain('Planning:');
|
||
expect(question.promptSnippet.length).toBeLessThanOrEqual(240);
|
||
expect(capturePlanCountQuestion(frame + '\n' + '·'.repeat(6000), seen, 1, false)).toBeNull();
|
||
});
|
||
|
||
test('does not reuse an old boxed header for a later unboxed menu', () => {
|
||
const visible = [
|
||
'☐ Old setup',
|
||
'D1 — Cross-project learnings scope',
|
||
'❯1.Enable',
|
||
'2.Skip',
|
||
'Planning: /tmp/hermetic/.claude/plans/review.md',
|
||
'D2 — Choose the retry behavior',
|
||
'❯1.Use library built-in',
|
||
'2.Extract shared retry envelope',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toBe('D2 — Choose the retry behavior');
|
||
expect(prompt).not.toContain('Old setup');
|
||
});
|
||
|
||
test('captures 1-line prompt above the cursor', () => {
|
||
const visible = `
|
||
D1 — Pick a mode
|
||
|
||
❯ 1. HOLD SCOPE
|
||
2. SCOPE EXPANSION
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toBe('D1 — Pick a mode');
|
||
});
|
||
|
||
test('captures multi-line prompt above the cursor', () => {
|
||
const visible = `
|
||
D2 — Approach selection
|
||
|
||
Which architecture should we follow?
|
||
|
||
❯ 1. Bypass existing helper
|
||
2. Reuse existing helper
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Multi-line prompts get joined with single spaces.
|
||
expect(prompt).toContain('D2 — Approach selection');
|
||
expect(prompt).toContain('Which architecture should we follow?');
|
||
});
|
||
|
||
test('returns "" when no cursor is rendered', () => {
|
||
expect(parseQuestionPrompt('Just some prose.\nNo cursor.')).toBe('');
|
||
});
|
||
|
||
test('truncates to 240 chars', () => {
|
||
const longPrompt = 'A'.repeat(500);
|
||
const visible = `${longPrompt}\n\n ❯ 1. yes\n 2. no`;
|
||
expect(parseQuestionPrompt(visible).length).toBeLessThanOrEqual(240);
|
||
});
|
||
|
||
test('does not pull text from a previous numbered list above', () => {
|
||
const visible = `
|
||
❯ 1. previous answered question
|
||
2. previous option two
|
||
|
||
D2 — A new question text
|
||
|
||
❯ 1. fresh option A
|
||
2. fresh option B
|
||
`;
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Stops at the previous numbered-list line; should NOT contain "previous answered question".
|
||
expect(prompt).toContain('D2 — A new question text');
|
||
expect(prompt).not.toContain('previous answered question');
|
||
});
|
||
|
||
test('normalizes whitespace (collapses runs of spaces and tabs)', () => {
|
||
const visible = `D1 — Spaced out
|
||
|
||
❯ 1. yes
|
||
2. no`;
|
||
expect(parseQuestionPrompt(visible)).toBe('D1 — Spaced out');
|
||
});
|
||
|
||
test('inline-cursor box-layout: extracts prompt text BEFORE ❯1. on the cursor line', () => {
|
||
// Real /plan-ceo-review rendering: divider + ☐ header + prompt text +
|
||
// cursor are all on one logical line because TTY cursor-positioning
|
||
// escapes collapse the box layout under stripAnsi.
|
||
const visible = [
|
||
'──────────────────',
|
||
'☐ Review scope What scope do you want me to CEO-review? ❯ 1. The branch\'s diff vs main',
|
||
'2. A specific plan file',
|
||
'3. An idea inline',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
// Should extract "Review scope" and the prompt text, dropping the ☐ box-drawing sigil.
|
||
expect(prompt).toContain('Review scope');
|
||
expect(prompt).toContain('What scope do you want me to CEO-review?');
|
||
expect(prompt).not.toContain('❯');
|
||
expect(prompt).not.toMatch(/^☐/);
|
||
});
|
||
|
||
test('keeps the captured design scope prompt ahead of long Planning chrome', () => {
|
||
// The first failed live attempt fingerprinted only the divider/Planning
|
||
// path. Its actual AUQ was later on the active cursor line.
|
||
const visible = [
|
||
'─'.repeat(120),
|
||
`Planning: /tmp/hermetic/.claude/plans/${'long-path-'.repeat(24)}plan.md`,
|
||
'─'.repeat(120),
|
||
"☐Reviewfocus I've rated this Settings Page UI redesign plan 2/10 on design completeness. Want me to focus on specific areas? ❯1.All7passes(Recommended)",
|
||
'2.All7passesbutskipmockups',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toStartWith('Reviewfocus');
|
||
expect(prompt).toContain('design completeness');
|
||
expect(prompt).not.toContain('Planning:');
|
||
expect(designStep0Boundary({
|
||
signature: 'captured-design-scope', promptSnippet: prompt,
|
||
options: parseNumberedOptions(visible), observedAtMs: 0, preReview: true,
|
||
})).toBe(true);
|
||
});
|
||
|
||
test('keeps the captured devex persona header when cursor spacing collapses', () => {
|
||
const visible = [
|
||
`Planning: /tmp/hermetic/.claude/plans/${'long-path-'.repeat(24)}plan.md`,
|
||
'─'.repeat(120),
|
||
'☐Targetpersona D2—WhoistheprimarydeveloperthisSDKtargets? ❯1.AIappbuilder/startupfounder(Recommended)',
|
||
'2.Backend/platformengineer',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toStartWith('Targetpersona');
|
||
expect(devexStep0Boundary({
|
||
signature: 'captured-devex-persona', promptSnippet: prompt,
|
||
options: parseNumberedOptions(visible), observedAtMs: 0, preReview: true,
|
||
})).toBe(true);
|
||
});
|
||
|
||
test('retains a multiline question while excluding the preceding CLI divider', () => {
|
||
const visible = [
|
||
'Planning: /tmp/hermetic/.claude/plans/plan.md',
|
||
'─'.repeat(120),
|
||
'☐ Review focus',
|
||
'This plan is 2/10 on design completeness.',
|
||
'Want me to focus on specific areas? ❯1.All 7 passes',
|
||
'2.Skip mockups',
|
||
].join('\n');
|
||
const prompt = parseQuestionPrompt(visible);
|
||
expect(prompt).toContain('Review focus');
|
||
expect(prompt).toContain('design completeness');
|
||
expect(prompt).toContain('specific areas?');
|
||
expect(prompt).not.toContain('Planning:');
|
||
});
|
||
});
|
||
|
||
describe('auqFingerprint', () => {
|
||
test('returns the same fingerprint for identical inputs', () => {
|
||
const opts = [
|
||
{ index: 1, label: 'A' },
|
||
{ index: 2, label: 'B' },
|
||
];
|
||
expect(auqFingerprint('hello', opts)).toBe(auqFingerprint('hello', opts));
|
||
});
|
||
|
||
test('different prompts with shared option labels produce DIFFERENT fingerprints', () => {
|
||
// The collision regression Codex F1 caught: option-label-only fingerprints
|
||
// collapsed multiple distinct findings into one when they shared menu shape.
|
||
const sharedOpts = [
|
||
{ index: 1, label: 'Add to plan' },
|
||
{ index: 2, label: 'Defer' },
|
||
{ index: 3, label: 'Build now' },
|
||
];
|
||
const fpFinding1 = auqFingerprint('D5 — Architecture: bypass helper?', sharedOpts);
|
||
const fpFinding2 = auqFingerprint('D6 — Tests: zero coverage?', sharedOpts);
|
||
expect(fpFinding1).not.toBe(fpFinding2);
|
||
});
|
||
|
||
test('same prompt with different options produces DIFFERENT fingerprints', () => {
|
||
const prompt = 'D1 — Pick a mode';
|
||
const fpA = auqFingerprint(prompt, [
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE EXPANSION' },
|
||
]);
|
||
const fpB = auqFingerprint(prompt, [
|
||
{ index: 1, label: 'HOLD SCOPE' },
|
||
{ index: 2, label: 'SCOPE REDUCTION' },
|
||
]);
|
||
expect(fpA).not.toBe(fpB);
|
||
});
|
||
|
||
test('whitespace-only differences in prompt do NOT change the fingerprint', () => {
|
||
// Same content, different rendering whitespace (TTY redraw artifact)
|
||
// must produce the same fingerprint so dedupe survives reflow.
|
||
const opts = [{ index: 1, label: 'A' }, { index: 2, label: 'B' }];
|
||
const fpA = auqFingerprint('Pick a mode', opts);
|
||
const fpB = auqFingerprint('Pick a mode', opts);
|
||
expect(fpA).toBe(fpB);
|
||
});
|
||
|
||
test('empty prompt + same options collide (caller must guard against this)', () => {
|
||
// Documents the contract: empty-prompt fingerprints WILL collide if the
|
||
// caller fingerprints them. runPlanSkillCounting must skip empty-prompt
|
||
// AUQs and re-poll instead.
|
||
const opts = [{ index: 1, label: 'A' }];
|
||
expect(auqFingerprint('', opts)).toBe(auqFingerprint('', opts));
|
||
});
|
||
});
|
||
|
||
describe('capturePlanCountQuestion replay', () => {
|
||
test('keeps captured CEO/eng fingerprints stable as later output trims the trailing window', () => {
|
||
// Exact prompt/option fields from the 07:30 corrected paid attempts.
|
||
// Both counted an answered Step0 question again as a review finding
|
||
// once the moving tail omitted the beginning of its prompt.
|
||
const captures = [
|
||
{
|
||
prompt: '☐ RevewMode Which review mode should I use for the remaining sections?',
|
||
labels: [
|
||
'HOLD SCOPE — make it ┌┐',
|
||
'SELECTIVEEXPANSION—│Focus:catcheverylandmineinApproachA│',
|
||
'SCOPEREDUCTION—strip│Tests:whatmustbecovered│',
|
||
'SCOPEEXPANSION—think│Observability:whatlogs/metricsareneeded│',
|
||
],
|
||
},
|
||
{
|
||
prompt: '☐ Scope cut │ D2 — Scope reduction proposal: drop TokenStore and RequestPolicy as standalone classes, inject AuthCache rather than │ exportitglobally.Acceptthisreductionbeforethesection-by-sectionreviewbegins? │ <gstack-qid:plan-eng-review-',
|
||
labels: [
|
||
'Acceptscopereduction┌───────────────────────────────────────────────────┐',
|
||
'Proceedfullscopeas-is│AuthBroker│',
|
||
],
|
||
},
|
||
];
|
||
for (const capture of captures) {
|
||
const options = capture.labels.map((label, i) => `${i === 0 ? '❯' : ''}${i + 1}.${label}`).join('\n');
|
||
const frame = `${capture.prompt}\n${options}`;
|
||
const seen = new Set<string>();
|
||
const first = capturePlanCountQuestion(frame, seen, 0, true)!;
|
||
expect(first).not.toBeNull();
|
||
// Leave the original menu within the trailing4KB, but move the
|
||
// start of that window into its question text, twice in succession.
|
||
const paddingLength = 4096 - options.length - 30;
|
||
for (const extra of [0, 15]) {
|
||
const advanced = frame + '\n' + '·'.repeat(paddingLength + extra - 1);
|
||
expect(advanced.slice(-4096)).not.toContain(capture.prompt);
|
||
expect(parseNumberedOptions(advanced)).toEqual(first.options);
|
||
expect(parseQuestionPrompt(advanced)).toBe(first.promptSnippet);
|
||
expect(auqFingerprint(parseQuestionPrompt(advanced), parseNumberedOptions(advanced))).toBe(first.signature);
|
||
expect(capturePlanCountQuestion(advanced, seen, extra + 1, false)).toBeNull();
|
||
}
|
||
const next = `${frame}\n${'·'.repeat(paddingLength)}\n☐ Next decision Should the revised plan use these same choices?\n${options}`;
|
||
const distinct = capturePlanCountQuestion(next, seen, 20, false)!;
|
||
expect(distinct).not.toBeNull();
|
||
expect(distinct.signature).not.toBe(first.signature);
|
||
expect(distinct.preReview).toBe(false);
|
||
expect(seen.size).toBe(2);
|
||
}
|
||
});
|
||
|
||
test('counts consecutive findings with identical choices and ignores redraws', () => {
|
||
const options = '\n❯1.Add to plan\n2.Defer\n3.Skip';
|
||
const seen = new Set<string>();
|
||
const frames = [
|
||
`D5 — SQL: interpolate the request parameter?${options}`,
|
||
`D5 — SQL: interpolate the request parameter?${options}`,
|
||
`D6 — Tests: no coverage for the webhook?${options}`,
|
||
`D6 — Tests: no coverage for the webhook?${options}`,
|
||
];
|
||
const captured = frames.map((frame, i) => capturePlanCountQuestion(frame, seen, i, false));
|
||
expect(captured.map((question) => question !== null)).toEqual([true, false, true, false]);
|
||
expect(captured[0]?.signature).not.toBe(captured[2]?.signature);
|
||
expect(captured[2]?.promptSnippet).toContain('Tests: no coverage');
|
||
});
|
||
|
||
test('does not consume an incomplete frame before its prompt arrives', () => {
|
||
const seen = new Set<string>();
|
||
const options = '❯1.Add to plan\n2.Defer';
|
||
expect(capturePlanCountQuestion(options, seen, 0, true)).toBeNull();
|
||
expect(capturePlanCountQuestion(`D1 — Pick an approach\n${options}`, seen, 1, true)).not.toBeNull();
|
||
});
|
||
|
||
test('answers the captured CEO retry question with a numeric-leading first label', () => {
|
||
// The live timeout sat on this question because the first label begins
|
||
// with "1retryattempt"; it was incorrectly rejected as a decimal token.
|
||
const frame = [
|
||
' ☐ Retry spec',
|
||
"│ Section 5/6 finding: 'retry-with-backoff fires once, then fails clean' is ambiguous.",
|
||
"│ What does 'fires once' mean?",
|
||
'❯1.1retryattempt—Stripecalledexactly2timestotal(Recommended)',
|
||
'Themostnaturalreading:1originalattempt+1retry=2totalStripecalls.',
|
||
'2.Addaclarifyingcommenttotheplan—lettheimplementerdecide',
|
||
'3.Theretrymechanismhandlesit—justassertfailureisreturned',
|
||
'4.Typesomething.',
|
||
'5.Chataboutthis',
|
||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||
].join('\r\r');
|
||
const question = capturePlanCountQuestion(frame, new Set(), 0, false);
|
||
expect(question?.options.map(({ index }) => index)).toEqual([1, 2, 3, 4, 5]);
|
||
expect(question?.options[0]?.label).toBe('1retryattempt—Stripecalledexactly2timestotal(Recommended)');
|
||
expect(question?.promptSnippet).toContain('Section 5/6 finding');
|
||
expect(question?.promptSnippet).not.toContain('Planning:');
|
||
});
|
||
|
||
test('still ignores decimal numbers inside option labels', () => {
|
||
const frame = 'Choose the retry delay\r❯1.1.5 seconds\r2.Wait 2.5 seconds\r3.No retry';
|
||
expect(parseNumberedOptions(frame)).toEqual([
|
||
{ index: 1, label: '1.5 seconds' },
|
||
{ index: 2, label: 'Wait 2.5 seconds' },
|
||
{ index: 3, label: 'No retry' },
|
||
]);
|
||
});
|
||
});
|
||
|
||
describe('planCountPrerequisitePick replay', () => {
|
||
test('declines captured office-hours prerequisite menus by label in either order', () => {
|
||
// Captured 2026-09-08 CEO/Devex prerequisite surfaces: the default index
|
||
// sometimes starts office-hours, changing the seeded review's input.
|
||
const captures = [
|
||
{
|
||
prompt: 'No design doc found for this branch. `/office-hours` produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input. Run it now, or skip and proceed with standard review?',
|
||
labels: ['Skip — proceed with standard review (Recommended)', 'Run /office-hours first'],
|
||
},
|
||
{
|
||
prompt: 'D2 — No design doc found. Run /office-hours first? <gstack-qid:plan-ceo-prereq-office-hours>',
|
||
labels: ['Skip — standard review (recommended)', 'Run /office-hours now'],
|
||
},
|
||
{
|
||
prompt: 'D3 — Run /office-hours first to produce a design doc for sharper input?',
|
||
labels: ['Skip — proceed with standard review (recommended)', 'Run /office-hours now'],
|
||
},
|
||
];
|
||
for (const { prompt, labels } of captures) {
|
||
for (const reversed of [false, true]) {
|
||
for (const collapsed of [false, true]) {
|
||
const ordered = reversed ? [...labels].reverse() : labels;
|
||
const text = ['☐ Prerequisite', prompt, `❯1.${ordered[0]}`, `2.${ordered[1]}`, '3.Type something.', '4.Chat about this'].join('\r');
|
||
const frame = collapsed ? text.replace(/ /g, '') : text;
|
||
const fp = capturePlanCountQuestion(frame, new Set(), 0, true)!;
|
||
expect(fp).not.toBeNull();
|
||
expect(planCountPrerequisitePick(fp)).toBe(reversed ? 2 : 1);
|
||
expect(planCountPrerequisitePick({ ...fp, preReview: false })).toBeNull();
|
||
}
|
||
}
|
||
}
|
||
});
|
||
|
||
test('keeps existing answers for incomplete, unrelated, and ambiguous menus', () => {
|
||
const fp = capturePlanCountQuestion(
|
||
'☐ Prerequisite\rNo design doc found. Run /office-hours first?\r❯1.Run /office-hours now\r2.Skip — proceed with standard review',
|
||
new Set(), 0, true,
|
||
)!;
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'No design doc found.' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'Should /office-hours skip the required SDK validation finding?' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, promptSnippet: 'Want a second opinion from /office-hours?' })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [{ index: 1, label: 'Run /office-hours now' }, { index: 2, label: 'Skip' }] })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [{ index: 1, label: 'Add to plan' }, fp.options[1]] })).toBeNull();
|
||
expect(planCountPrerequisitePick({ ...fp, options: [...fp.options, { index: 3, label: 'Skip — standard review' }] })).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('COMPLETION_SUMMARY_RE', () => {
|
||
test('matches GSTACK REVIEW REPORT heading', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('## GSTACK REVIEW REPORT')).toBe(true);
|
||
});
|
||
|
||
test('matches Completion Summary heading (ceo + eng)', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('## Completion Summary')).toBe(true);
|
||
expect(COMPLETION_SUMMARY_RE.test('## Completion summary')).toBe(true);
|
||
});
|
||
|
||
test('matches Status: clean (CEO review-log shape)', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('Status: clean')).toBe(true);
|
||
expect(COMPLETION_SUMMARY_RE.test('Status: issues_open')).toBe(true);
|
||
});
|
||
|
||
test('matches VERDICT: line', () => {
|
||
expect(COMPLETION_SUMMARY_RE.test('VERDICT: CLEARED — Eng Review passed')).toBe(true);
|
||
});
|
||
|
||
test('does NOT match prose mentions of "verdict" mid-line', () => {
|
||
// VERDICT must be at the start of a line to count.
|
||
expect(COMPLETION_SUMMARY_RE.test('the final verdict: undecided')).toBe(false);
|
||
});
|
||
|
||
test('does NOT treat source or proposed diff rows as assistant completion', () => {
|
||
for (const line of [
|
||
'409 +## GSTACK REVIEW REPORT',
|
||
'419 +**VERDICT:** Design Review complete — 8 decisions made.',
|
||
'+## GSTACK REVIEW REPORT',
|
||
'409→## GSTACK REVIEW REPORT',
|
||
'The plan must end with ## GSTACK REVIEW REPORT.',
|
||
]) expect(COMPLETION_SUMMARY_RE.test(line)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('classifyPlanCountFrame replay', () => {
|
||
test('waits through proposed Write approval and tool output, then accepts the actual report', () => {
|
||
// Sanitized rows and native prompt from the failed design-count attempt.
|
||
const proposedDiff = [
|
||
'409 +## GSTACK REVIEW REPORT',
|
||
'416 +| Design Review | 1 | issues_open | score: 2/10 → 8/10, 8 decisions |',
|
||
'419 +**VERDICT:** Design Review complete — 8 decisions made.',
|
||
].join('\n');
|
||
const permission = [
|
||
'Doyouwanttooverwritegstack-test-plan-design.md?',
|
||
'❯1.Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/fixtureforthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
const frames = [
|
||
`${proposedDiff}\n${permission}`,
|
||
`${proposedDiff}\n⏺ Updated gstack-test-plan-design.md`,
|
||
`${proposedDiff}\n⏺ ## GSTACK REVIEW REPORT\nDesign Review complete — 8 decisions made.`,
|
||
];
|
||
expect(frames.map(classifyPlanCountFrame)).toEqual(['permission', null, 'completion_summary']);
|
||
});
|
||
|
||
test('a pending native permission beats even an unnumbered report heading', () => {
|
||
const visible = '## GSTACK REVIEW REPORT\nDoyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
expect(classifyPlanCountFrame(visible)).toBe('permission');
|
||
});
|
||
|
||
test('a later report supersedes the granted menu still in short scrollback', () => {
|
||
const permission = 'Doyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
expect(classifyPlanCountFrame(permission)).toBe('permission');
|
||
expect(classifyPlanCountFrame(`${permission}\n● ## GSTACK REVIEW REPORT`)).toBe('completion_summary');
|
||
});
|
||
|
||
test('an active question after a prior report keeps the counter running', () => {
|
||
expect(classifyPlanCountFrame('## GSTACK REVIEW REPORT\nOne more choice\n❯1.Add to plan\n2.Defer')).toBeNull();
|
||
});
|
||
|
||
test('an active AUQ supersedes a granted permission menu in short scrollback', () => {
|
||
const permission = 'Doyouwanttooverwriteplan.md?\n❯1.Yes\n2.No\nEsctocancel·Tabtoamend';
|
||
const question = '☐ Error handling\nWhich failure path should we test?\n❯1.Timeout\n2.Refusal';
|
||
expect(classifyPlanCountFrame(`${permission}\n${question}`)).toBeNull();
|
||
});
|
||
|
||
test('preserves actual report variants and the native plan-ready terminal', () => {
|
||
for (const report of [
|
||
'## GSTACK REVIEW REPORT', '⏺##GSTACKREVIEWREPORT', '●GSTACKREVIEWREPORT',
|
||
'## Completion Summary', '● ## Completion Summary', 'Status: clean', 'Status: issues_open',
|
||
'VERDICT: CLEARED — Eng Review passed', '**VERDICT:** Design Review complete.',
|
||
]) expect(classifyPlanCountFrame(report)).toBe('completion_summary');
|
||
expect(classifyPlanCountFrame('Ready to execute the plan?\n❯1.Yes\n2.No, keep planning')).toBe('plan_ready');
|
||
});
|
||
});
|
||
|
||
describe('planCountSubmissionInput replay', () => {
|
||
test('uses the captured DevEx panel anchors when the Submit button label is damaged', () => {
|
||
const captured = [
|
||
'← ☒ Routing setup ☐ Cross-project ✔ Submit →',
|
||
'Review your answers',
|
||
'⚠You have not answere all questions',
|
||
' │ ●D1 — Shouldgstack add skill routingrulestothisproject\'sCLAUDE.md?<gstack-qid:routing-injection>',
|
||
'→dd routing rules (Recmmeded)',
|
||
'Ready to submit your answers?',
|
||
'❯1.Sbmi answers',
|
||
'2Cancel',
|
||
].join('\r\r');
|
||
expect(planCountSubmissionInput(captured)).toBe('\x1b[Z');
|
||
expect(planCountSubmissionInput(captured.replace('☐ Cross-project', '☒ Cross-project'))).toBe('\r');
|
||
expect(planCountSubmissionInput(captured + '\r☐ Retry spec\rRetry once?\r❯1.Yes\r2.No')).toBeNull();
|
||
expect(planCountSubmissionInput(captured + '\r☐ Proposal\rSend this proposal?\r❯1.Submit proposal\r2.Keep editing')).toBeNull();
|
||
expect(planCountSubmissionInput(captured + '\r☐ Retry spec\rRetry once?\r❯2.No\r3.Other')).toBeNull();
|
||
expect(planCountSubmissionInput(captured.replace('Review your answers', 'Review context'))).toBeNull();
|
||
expect(planCountSubmissionInput(captured.replace('Ready to submit your answers?', 'Read the proposed answers.'))).toBeNull();
|
||
});
|
||
|
||
test('the captured mode Submit panel with a damaged caption and dotless cursor returns to its unanswered tab', () => {
|
||
// Exact final active panel from targeted-a's SCOPE EXPANSION retry.
|
||
const captured = [
|
||
'← ☒ Routing rule ☐ Design doc ✔ Submit →',
|
||
'',
|
||
'Review your answrs',
|
||
'⚠ You hvenot answered all questions',
|
||
" ● Add gstack skill routing rules tothisproject'sCLAUDE.md?",
|
||
'→dd routing rues (Recommnded)',
|
||
'',
|
||
'Ready to submit your answers?',
|
||
'',
|
||
'❯1Submit answers',
|
||
' 2. Cancel',
|
||
].join('\r');
|
||
expect(planCountSubmissionInput(captured)).toBe('\x1b[Z');
|
||
const answered = captured.replace('☐ Design doc', '☒ Design doc').replace('⚠ You hvenot answered all questions', '');
|
||
expect(planCountSubmissionInput(answered)).toBe('\r');
|
||
expect(planCountSubmissionInput(captured + '\r☐ Design doc\rRun office hours?\r❯1Run now\r2.Skip')).toBeNull();
|
||
});
|
||
|
||
const incomplete = [
|
||
'← ☒ Learnings scope ☐ Approach ✔ Submit →',
|
||
'Review your answers',
|
||
'⚠You have not answered all questions',
|
||
' │ ●D1 — Cross-project learnings: Enable searching learnings from your other local projects?',
|
||
'→Enable cross-project (Recommended)',
|
||
'Ready t submit your answers?',
|
||
'❯1.Submit aswers',
|
||
'2Cancel',
|
||
].join('\r\r');
|
||
|
||
test('returns to the unanswered tab, then submits only after both answers', () => {
|
||
expect(planCountSubmissionInput(incomplete)).toBe('\x1b[Z');
|
||
const nextQuestion = [
|
||
'← ☒ Learnings scope ☐ Approach ✔ Submit →',
|
||
'│ Which approach should this plan use?',
|
||
'❯1.Extend the existing dispatcher',
|
||
'2.Add a separate handler',
|
||
].join('\r\r');
|
||
expect(planCountSubmissionInput(`${incomplete}\r${nextQuestion}`)).toBeNull();
|
||
const question = capturePlanCountQuestion(nextQuestion, new Set(), 0, true);
|
||
expect(question?.promptSnippet).toContain('Which approach');
|
||
expect(question?.options).toHaveLength(2);
|
||
const answered = incomplete.replace('☐ Approach', '☒ Approach').replace('⚠You have not answered all questions', '');
|
||
expect(planCountSubmissionInput(answered)).toBe('\r');
|
||
});
|
||
|
||
test('navigates to the first unanswered tab when more than one remains', () => {
|
||
const frame = incomplete.replace('☒ Learnings scope ☐ Approach', '☐ Learnings scope ☐ Approach ☒ Mode');
|
||
expect(planCountSubmissionInput(frame)).toBe('\x1b[Z\x1b[Z\x1b[Z');
|
||
});
|
||
|
||
test('does not revisit a stale submit panel when a later single question is active', () => {
|
||
expect(planCountSubmissionInput(`${incomplete}\r☐ Retry spec\rRetry once?\r❯1.Yes\r2.No`)).toBeNull();
|
||
expect(planCountSubmissionInput('Ready to submit the plan?\n❯1.Submit\n2.Cancel')).toBeNull();
|
||
});
|
||
});
|
||
|
||
describe('assertReviewReportAtBottom', () => {
|
||
test('passes when REVIEW REPORT is the only/last ## heading', () => {
|
||
const content = `# Plan
|
||
|
||
## Context
|
||
stuff
|
||
|
||
## Approach
|
||
more stuff
|
||
|
||
## GSTACK REVIEW REPORT
|
||
|
||
| col | col |
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(true);
|
||
});
|
||
|
||
test('fails when REVIEW REPORT is missing', () => {
|
||
const content = `# Plan
|
||
|
||
## Context
|
||
stuff
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.reason).toMatch(/no GSTACK REVIEW REPORT/);
|
||
});
|
||
|
||
test('fails when REVIEW REPORT exists but a ## heading follows it', () => {
|
||
const content = `# Plan
|
||
|
||
## GSTACK REVIEW REPORT
|
||
|
||
| col | col |
|
||
|
||
## Late Section
|
||
oops
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.reason).toMatch(/trailing ## heading/);
|
||
expect(r.trailingHeadings).toEqual(['## Late Section']);
|
||
});
|
||
|
||
test('passes when only ### subheadings follow REVIEW REPORT (deeper nesting allowed)', () => {
|
||
const content = `## GSTACK REVIEW REPORT
|
||
|
||
### Cross-model tension
|
||
- F1: resolved
|
||
- F2: resolved
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(true);
|
||
});
|
||
|
||
test('fails with multiple trailing ## headings reported', () => {
|
||
const content = `## GSTACK REVIEW REPORT
|
||
|
||
## First trailing
|
||
|
||
## Second trailing
|
||
`;
|
||
const r = assertReviewReportAtBottom(content);
|
||
expect(r.ok).toBe(false);
|
||
expect(r.trailingHeadings).toHaveLength(2);
|
||
});
|
||
});
|
||
|
||
describe('Step0BoundaryPredicate per-skill', () => {
|
||
// Helper to build a synthetic fingerprint for predicate tests.
|
||
function fp(promptSnippet: string, optionLabels: string[]): AskUserQuestionFingerprint {
|
||
const options = optionLabels.map((label, i) => ({ index: i + 1, label }));
|
||
return {
|
||
signature: auqFingerprint(promptSnippet, options),
|
||
promptSnippet,
|
||
options,
|
||
observedAtMs: 0,
|
||
preReview: true,
|
||
};
|
||
}
|
||
|
||
describe('native Cross-project onboarding boundary', () => {
|
||
// Fresh paid run D, 2026-09-08: native question stems, labels and
|
||
// successful answers. Long explanatory paragraphs are omitted; they
|
||
// must not determine this structural setup boundary.
|
||
const captured = [
|
||
{
|
||
"header": "Routing rules",
|
||
"question": "Should I add gstack skill routing rules to your project's CLAUDE.md? (Note: we're in plan mode — if you pick A, I'll make the edit after we exit plan mode.)",
|
||
"options": [
|
||
"Add routing rules (recommended)",
|
||
"Skip — invoke manually"
|
||
],
|
||
"answer": "Add routing rules (recommended)"
|
||
},
|
||
{
|
||
"header": "Scope challenge",
|
||
"question": "D2 — The plan introduces 4 new classes across 12 files. Should I flag scope reduction as a primary recommendation in the review, or accept the 4-class design and focus findings on quality issues?",
|
||
"options": [
|
||
"Accept 4-class design, focus on quality",
|
||
"Flag scope reduction as primary finding (recommended)"
|
||
],
|
||
"answer": "Accept 4-class design, focus on quality"
|
||
},
|
||
{
|
||
"header": "Cross-project",
|
||
"question": "D3 — Should gstack search learnings from your other projects on this machine when reviewing?",
|
||
"options": [
|
||
"Enable cross-project learnings (recommended)",
|
||
"Keep learnings project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project learnings (recommended)"
|
||
},
|
||
{
|
||
"header": "AuthCache race",
|
||
"question": "D4 — AuthCache is shared mutable state mutated by two services with no serialization. How should we fix it?",
|
||
"options": [
|
||
"Single-writer: AuthBroker owns all writes (recommended)",
|
||
"Immutable cache + versioned replace",
|
||
"Accept and document the race"
|
||
],
|
||
"answer": "Single-writer: AuthBroker owns all writes (recommended)"
|
||
},
|
||
{
|
||
"header": "Double-cache risk",
|
||
"question": "D5 — The plan introduces a new AuthCache class but doesn't say what happens to the existing cache adapter. Are they running in parallel?",
|
||
"options": [
|
||
"AuthCache replaces the adapter — add migration to plan (recommended)",
|
||
"AuthCache wraps the adapter — adapter stays as storage layer",
|
||
"Leave ambiguous — clarify in implementation"
|
||
],
|
||
"answer": "AuthCache replaces the adapter — add migration to plan (recommended)"
|
||
},
|
||
{
|
||
"header": "Error swallowing",
|
||
"question": "D6 — validateAndDispatch() swallows three different error classes across nested try/catch blocks. How should this be resolved in the plan?",
|
||
"options": [
|
||
"Decompose + typed error results (recommended)",
|
||
"Keep structure, add logging + rethrow",
|
||
"Leave as-is — document that swallowing is intentional"
|
||
],
|
||
"answer": "Decompose + typed error results (recommended)"
|
||
},
|
||
{
|
||
"header": "Invalidation tests",
|
||
"question": "D7 — When AuthCache replaces the existing adapter (per D5), the existing invalidation tests (logout, revocation, tenant suspension) become dead — they're testing a retired object. Should the plan explicitly require migrating them?",
|
||
"options": [
|
||
"Migrate invalidation tests to AuthCache — add to plan (recommended)",
|
||
"Scope to new component tests only — leave invalidation as follow-up",
|
||
"Assume existing tests cover it — no explicit migration step"
|
||
],
|
||
"answer": "Migrate invalidation tests to AuthCache — add to plan (recommended)"
|
||
},
|
||
{
|
||
"header": "IDP parallelization",
|
||
"question": "D8 — The plan identifies 5 sequential IDP calls that are independent and could be parallelized with Promise.all. Should we include the fix in this PR or defer it?",
|
||
"options": [
|
||
"Parallelize with Promise.all in this PR (recommended)",
|
||
"Defer to TODOS.md",
|
||
"Leave sequential — document as known limitation"
|
||
],
|
||
"answer": "Parallelize with Promise.all in this PR (recommended)"
|
||
}
|
||
];
|
||
const fingerprint = (index: number) => {
|
||
const row = captured[index];
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'fresh-eng-cross-project', toolUseId: `call-${index}`, answered: true,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
};
|
||
|
||
test('keeps D3 as setup and counts each following actual review call', () => {
|
||
let started = false;
|
||
const phases = captured.map((_, i) => {
|
||
const phase = planCountQuestionPhase(fingerprint(i), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.slice(0, 3).map(p => p.preReview)).toEqual([true, true, true]);
|
||
expect(phases[2].reviewStarted).toBe(true);
|
||
expect(phases.slice(3).map(p => p.preReview)).toEqual([false, false, false, false, false]);
|
||
});
|
||
|
||
test('the captured no-qid scope decision stays setup after Learnings scope', () => {
|
||
let started = false;
|
||
const phases = [0, 2, 1, 3, 4, 5, 6, 7].map(i => {
|
||
const phase = planCountQuestionPhase(fingerprint(i), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.slice(0, 3).map(p => p.preReview)).toEqual([true, true, true]);
|
||
expect(phases.slice(3).map(p => p.preReview)).toEqual([false, false, false, false, false]);
|
||
const call = structuredClone(fingerprint(1).nativeCall!);
|
||
call.questions[0].header = 'Scope complexity';
|
||
call.questions[0].options.reverse();
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
// Even with opposed scope labels, an ordinary per-issue finding lacks
|
||
// the captured whole-plan classes/files identity.
|
||
call.questions[0].question = 'How should we reduce shared mutable cache complexity?';
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
});
|
||
|
||
test('uses the native opposed scope choices without depending on a prompt suffix', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
const q = fp.nativeCall!.questions[0];
|
||
q.question = 'Should local lessons from other repositories be included during reviews?';
|
||
q.options.reverse();
|
||
fp.nativeCall!.answers = { [q.question]: q.options[0].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(fp.nativeCall!, 0, true))).toBe(true);
|
||
});
|
||
|
||
test("an unanswered Cross-project tab cannot borrow another tab's answer", () => {
|
||
const call = structuredClone(fingerprint(2).nativeCall!);
|
||
const other = { header: 'Routing rules', question: 'Add routing rules?', options: [{ label: 'Add rules' }, { label: 'Skip' }] };
|
||
call.questions.push(other);
|
||
call.answers = { [other.question]: 'Add rules' };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[1].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(true);
|
||
});
|
||
|
||
test('pending, failed, missing native metadata and ordinary review questions are not this gate', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engStep0Boundary({ ...fp, nativeCall: undefined })).toBe(false);
|
||
for (const alter of [
|
||
(call: any) => { call.answered = false; },
|
||
(call: any) => { call.failed = true; },
|
||
(call: any) => { call.questions[0].header = 'Architecture issue'; call.questions[0].question = 'D1 — Architecture issue: should the cross-project feature use shared storage?'; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project search' }, { label: 'Disable all search' }]; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project search with project-scoped storage' }, { label: 'Discuss later' }]; },
|
||
(call: any) => { call.questions[0].question = 'How should concurrent AuthCache writes across projects be serialized?';
|
||
call.questions[0].options = [{ label: 'Serialize mutations' }, { label: 'Use project-scoped locks' }]; },
|
||
]) {
|
||
const call = structuredClone(fp.nativeCall!);
|
||
alter(call);
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
expect(engStep0Boundary(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('ceoStep0Boundary', () => {
|
||
test('FIRES on Step 0F mode-pick AUQ (HOLD SCOPE in options)', () => {
|
||
const f = fp('Pick a mode', ['HOLD SCOPE', 'SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'SCOPE REDUCTION']);
|
||
expect(ceoStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on collapsed mode labels captured from the 2026-09-08 paid controls', () => {
|
||
// Test each captured label independently: a spaced sibling option can
|
||
// otherwise conceal the mismatch and leave every review AUQ in Step 0.
|
||
const labels = [
|
||
'HOLDSCOPE—makeitbulletproof(Recommended)',
|
||
'SELECTIVEEXPANSION┌────────────────────────────────────────────────────────────────────────────────────┐\r (ecommnded) │SELECTIVEEXPANSION│',
|
||
'SCOPEEXPANSION│Neutralposture:presentopportunities,stateeffort,youdecide.│\r │ Good for: substantialfeaturewithsolidfoundation,shippedbeforescopelock.│\r└────────────────────────────────────────────────────────────────────┘',
|
||
'SCOPEREDUCTION—findtheminimalversion',
|
||
];
|
||
for (const label of labels) {
|
||
expect(ceoStep0Boundary(fp('Pick a mode', [label, 'Type something.']))).toBe(true);
|
||
}
|
||
});
|
||
|
||
test('FIRES on scope-selection AUQ with "Skip interview" option (skip-interview path)', () => {
|
||
// After calibration run 1: plan-ceo's first AUQ is scope-selection,
|
||
// and we route via "Skip interview and plan immediately" to bypass
|
||
// Step 0 entirely. Boundary must fire on this AUQ so subsequent
|
||
// AUQs go to reviewCount.
|
||
const f = fp(
|
||
'What scope do you want me to CEO-review?',
|
||
[
|
||
"The branch's diff vs main",
|
||
'A specific plan file',
|
||
"An idea you'll describe inline",
|
||
'Cancel — wrong skill',
|
||
'Type something.',
|
||
'Chat about this',
|
||
'Skip interview and plan immediately',
|
||
],
|
||
);
|
||
expect(ceoStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on premise challenge AUQs', () => {
|
||
const f = fp('D1 — Premise check: is this the right problem?', ['Yes', 'No', 'Other']);
|
||
expect(ceoStep0Boundary(f)).toBe(false);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Architecture: bypass helper?', ['Reuse existing', 'Roll new', 'Defer']);
|
||
expect(ceoStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('engStep0Boundary', () => {
|
||
// Captured native question text, labels and answers from targeted-b's
|
||
// engineering retry. Descriptions are immaterial to the phase boundary.
|
||
const captured = [
|
||
{
|
||
"header": "Scope",
|
||
"question": "D1 — Multi-tenant Auth Refactor complexity check: 12 files + 4 new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-complexity>",
|
||
"options": [
|
||
"Proceed as-is",
|
||
"Reduce: cut TokenStore + RequestPolicy",
|
||
"Reduce: single-pass strangler"
|
||
],
|
||
"answer": "Proceed as-is"
|
||
},
|
||
{
|
||
"header": "Shared Cache",
|
||
"question": "D2 — Arch issue 1: AuthBroker and SessionMint both mutate a global shared AuthCache via module-level export. How should this be fixed? <gstack-qid:plan-eng-shared-mutable-cache>",
|
||
"options": [
|
||
"Inject AuthCache as a dependency (recommended)",
|
||
"Make mutations go through a single owner",
|
||
"Accept the risk for now, document it"
|
||
],
|
||
"answer": "Inject AuthCache as a dependency (recommended)"
|
||
},
|
||
{
|
||
"header": "TOCTOU",
|
||
"question": "D3 — Arch issue 2: TOCTOU window during tenant suspension. The plan says AuthCache invalidates entries on tenant suspension, but with two services mutating the cache, a token validation begun before suspension completes may still succeed after the tenant is suspended. How should this be addressed? <gstack-qid:plan-eng-toctou-suspension>",
|
||
"options": [
|
||
"Add suspension check at session issuance boundary (recommended)",
|
||
"Add invalidation ordering guarantee to the plan",
|
||
"Accept the window, note it as an edge case"
|
||
],
|
||
"answer": "Add suspension check at session issuance boundary (recommended)"
|
||
},
|
||
{
|
||
"header": "Error handling",
|
||
"question": "D4 — Code quality issue 1: validateAndDispatch() swallows three distinct error classes in nested catch blocks with no re-throw, logging, or metrics. Errors disappear silently. How should this be handled? <gstack-qid:plan-eng-error-swallowing>",
|
||
"options": [
|
||
"Refactor to flat error handling with explicit re-throw or typed result (recommended)",
|
||
"Add logging inside each catch, keep structure",
|
||
"Leave it, add a lint rule to catch new instances"
|
||
],
|
||
"answer": "Refactor to flat error handling with explicit re-throw or typed result (recommended)"
|
||
},
|
||
{
|
||
"header": "Test coverage",
|
||
"question": "D5 — Test issue 1: 0/18 code paths covered in the plan. The plan scopes tests to 'new components and their success/error paths' but omits: cache invalidation edge cases, all three catch blocks in validateAndDispatch(), and the 5 IDP call failure modes. Should the test scope be expanded? <gstack-qid:plan-eng-test-coverage>",
|
||
"options": [
|
||
"Expand test scope to cover all 18 paths (recommended)",
|
||
"Cover new paths only, defer legacy and edge cases",
|
||
"Accept current test scope as stated in the plan"
|
||
],
|
||
"answer": "Expand test scope to cover all 18 paths (recommended)"
|
||
},
|
||
{
|
||
"header": "IDP calls",
|
||
"question": "D6 — Performance issue 1: token validation makes 5 sequential IDP API calls. The plan acknowledges they are independent and could be parallelized via Promise.all. Should this be fixed in this PR or deferred? <gstack-qid:plan-eng-idp-sequential-calls>",
|
||
"options": [
|
||
"Parallelize now with Promise.all (recommended)",
|
||
"Defer to a follow-up PR, add a TODO",
|
||
"Add a concurrency cap via Promise.all with limit"
|
||
],
|
||
"answer": "Parallelize now with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"header": "Cache bounds",
|
||
"question": "D7 — Performance issue 2 (medium confidence): AuthCache evicts on token expiry but the plan doesn't mention a max-size bound. In a high-tenant deployment, long-lived non-expiring tokens could grow the cache without bound. Is there already a size cap, or should one be added? <gstack-qid:plan-eng-cache-unbounded>",
|
||
"options": [
|
||
"Verify existing cap exists and document it in the plan",
|
||
"Add explicit max-size eviction policy to AuthCache (recommended)",
|
||
"Defer, this is a scaling concern not a correctness one"
|
||
],
|
||
"answer": "Verify existing cap exists and document it in the plan"
|
||
},
|
||
{
|
||
"header": "TODO",
|
||
"question": "D8 — TODO candidate: Auth failure observability. The plan replaces silently-swallowed errors with typed errors, but adds no metrics, logs, or alerts for auth failure patterns. This gap won't surface until production incidents occur. Add a TODO? <gstack-qid:plan-eng-todo-observability>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Build it now in this PR instead of deferring",
|
||
"Skip — not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
}
|
||
];
|
||
function nativeScopeFingerprint(index = 0): AskUserQuestionFingerprint {
|
||
const row = captured[index];
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'captured-eng-retry', toolUseId: `call-${index}`, answered: true,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
}
|
||
|
||
test('keeps captured scope-complexity setup and counts the six following findings', () => {
|
||
let started = false;
|
||
const phases = captured.map((_, index) => {
|
||
const question = nativeScopeFingerprint(index);
|
||
const phase = planCountQuestionPhase(question, started, engStep0Boundary);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases[0]).toEqual({ preReview: true, reviewStarted: true });
|
||
expect(phases.slice(1, 7).filter(phase => !phase.preReview)).toHaveLength(6);
|
||
// The later answered observability TODO retains the existing phase policy.
|
||
expect(phases.filter(phase => !phase.preReview)).toHaveLength(7);
|
||
});
|
||
|
||
test('requires an answered native scope decision with its opposed scope choices', () => {
|
||
const original = nativeScopeFingerprint();
|
||
expect(engStep0Boundary(original)).toBe(true);
|
||
expect(engStep0Boundary({ ...original, nativeCall: undefined })).toBe(false);
|
||
const pending = structuredClone(original);
|
||
pending.nativeCall!.answered = false;
|
||
delete pending.nativeCall!.answers;
|
||
expect(engStep0Boundary(pending)).toBe(false);
|
||
const unansweredScope = structuredClone(original);
|
||
unansweredScope.nativeCall!.answers = { 'Separate answered setup question': 'Continue' };
|
||
expect(engStep0Boundary(unansweredScope)).toBe(false);
|
||
for (const alter of [
|
||
(q: any) => { q.question = 'D1 — Architecture issue: the plan touches 12 files and introduces 4 classes, but AuthCache has a race. Reduce cache scope? <gstack-qid:plan-eng-cache-complexity>'; },
|
||
(q: any) => { q.options = [{ label: 'Change cache size' }, { label: 'Keep cache size' }]; },
|
||
]) {
|
||
const unrelated = structuredClone(original);
|
||
alter(unrelated.nativeCall!.questions[0]);
|
||
unrelated.nativeCall!.answers = { [unrelated.nativeCall!.questions[0].question]: captured[0].answer };
|
||
expect(engStep0Boundary(unrelated)).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('FIRES on cross-project learnings prompt', () => {
|
||
const f = fp('Enable cross-project learnings on this machine?', ['Yes', 'No']);
|
||
expect(engStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('recognizes the captured cross-project gate after cursor spacing collapses', () => {
|
||
const frame = [
|
||
'☐Cross-project gstackcansearchlearningsfromyourotherprojectsonthismachinetofindpatternsthatmightapplytothisreview.Enablecross-projectlearnings?',
|
||
'❯1.Enablecross-projectlearnings(Recommended)',
|
||
'2.Keeplearningsproject-scopedonly',
|
||
].join('\r\r');
|
||
const question = capturePlanCountQuestion(frame, new Set(), 0, true)!;
|
||
expect(question).not.toBeNull();
|
||
expect(engStep0Boundary(question)).toBe(true);
|
||
expect(engStep0Boundary(fp('Scopereductionrecommendation:cuttoMVP?', ['Reduce', 'Proceed']))).toBe(true);
|
||
});
|
||
|
||
test('FIRES on scope reduction recommendation', () => {
|
||
const f = fp('Scope reduction recommendation: cut to MVP?', ['Reduce', 'Proceed', 'Modify']);
|
||
expect(engStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Architecture: shared mutable state?', ['Refactor', 'Defer', 'Skip']);
|
||
expect(engStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('designStep0Boundary', () => {
|
||
test('FIRES on design system / posture mention', () => {
|
||
const f = fp('Pick a design posture for this review', ['Polish', 'Triage', 'Expansion']);
|
||
expect(designStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on first-dimension prompt', () => {
|
||
const f = fp('First dimension: visual hierarchy. Score?', ['7', '8', '9']);
|
||
expect(designStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on later dimension AUQs', () => {
|
||
const f = fp('Spacing dimension score?', ['7', '8', '9']);
|
||
expect(designStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
|
||
describe('design review begins without an optional focus question', () => {
|
||
// Captured in the second07:53 paid attempt: real D1-D7 questions were
|
||
// all incorrectly marked preReview, producing reviewCount=0 at completion.
|
||
const questions = [
|
||
'☐Buttonstyle │D1—Howshouldthe4headerbuttons(Save,Reset,Cancel,Export)bevisuallydifferentiated? │<gstack-qid:plan-design-review-butn-hierarchy>',
|
||
'☐ Loading UX │D2—Whatloadingindicatorshouldappearduringthe2-5secondSaveoperation? │<gtack-qid:plan-esign-review-loadig-indicator>',
|
||
'☐ Spacing │D3—Whichspacingscaleshouldthesettingspagestandardizeon?<gstack-qid:plan-design-review-spacing-scale>',
|
||
'☐Typography │D4—Which2-sizetypographysystemshouldthesettingspageuse?<gstack-qid:plan-design-review-type-system>',
|
||
'☐Mobile layout │D5 — On obile (<768px), how should the4-buton header behave?<gstack-qid:plan-design-review-mobile-header>',
|
||
'☐DEIGN.md TODO │D6 — TODO: Create a DESIGN.md file codifying the5 decisions mdein this revew <gstack-qid:plan-design-review-todo-designmd>',
|
||
'☐PartialfailTODO │D7—TODO:Specifythepartial-failurestate—whatdoestheuserseeifSavesucceedsforsomefieldsbutfailsfor others? <gstack-qid:plan-design-review-todo-partialfail>',
|
||
];
|
||
|
||
test('counts the first captured finding and every subsequent finding', () => {
|
||
let reviewStarted = false;
|
||
const phases = questions.map(question => {
|
||
const phase = planCountQuestionPhase(fp(question, ['Apply', 'Defer']), reviewStarted,
|
||
designStep0Boundary, designFirstReviewAUQ);
|
||
reviewStarted = phase.reviewStarted;
|
||
return phase.preReview;
|
||
});
|
||
expect(phases).toEqual([false, false, false, false, false, false, false]);
|
||
});
|
||
|
||
test('keeps the observed focus gate separate when it is emitted', () => {
|
||
const focus = fp("☐ Focus areas │ I've rated this plan2/10 on design completeness. Review all7 dimensions?", ['All7dimensions', 'Priority gaps']);
|
||
const setup = planCountQuestionPhase(focus, false, designStep0Boundary, designFirstReviewAUQ);
|
||
expect(setup).toEqual({ preReview: true, reviewStarted: true });
|
||
expect(planCountQuestionPhase(fp(questions[0], ['Apply', 'Defer']), setup.reviewStarted,
|
||
designStep0Boundary, designFirstReviewAUQ)).toEqual({ preReview: false, reviewStarted: true });
|
||
});
|
||
|
||
test('requires review identity, not just a D1 label or setup question ID', () => {
|
||
for (const question of [
|
||
'☐ Setup │D1—Enable cross-project learnings?',
|
||
'☐ Review target │D1—Which plan should I review?<gstack-qid:plan-design-review-scope>',
|
||
'☐ Focus │D1—What should this design review focus on?<gstack-qid:plan-design-review-focus-areas>',
|
||
'☐ Scope │I will review Pass1 through Pass7 after setup. Proceed?',
|
||
]) expect(designFirstReviewAUQ(fp(question, ['Yes', 'No']))).toBe(false);
|
||
expect(designFirstReviewAUQ(fp('☐ Page structure │ Pass1 — Information Architecture: what page structure should this use?', ['Standard', 'Sidebar']))).toBe(true);
|
||
});
|
||
|
||
test('leaves callers without a first-review predicate unchanged', () => {
|
||
expect(planCountQuestionPhase(fp(questions[0], ['Apply', 'Defer']), false, designStep0Boundary))
|
||
.toEqual({ preReview: true, reviewStarted: false });
|
||
});
|
||
});
|
||
|
||
describe('devexStep0Boundary', () => {
|
||
test('FIRES on developer persona selection', () => {
|
||
const f = fp('Pick the target persona for this review', ['Senior backend', 'Junior frontend', 'Other']);
|
||
expect(devexStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('FIRES on TTHW target prompt', () => {
|
||
const f = fp('What is the TTHW target for first run?', ['<5 min', '<15 min', '<30 min']);
|
||
expect(devexStep0Boundary(f)).toBe(true);
|
||
});
|
||
|
||
test('does NOT fire on review-section AUQs', () => {
|
||
const f = fp('Friction point: 5-min CI wait. Address?', ['Now', 'Defer', 'Skip']);
|
||
expect(devexStep0Boundary(f)).toBe(false);
|
||
});
|
||
});
|
||
});
|
||
|
||
|
||
describe('file permission lifecycle replay', () => {
|
||
const permission = (file = 'gstack-test-plan-design.md') => [
|
||
`Do you want to make this edit to ${file}?`,
|
||
'❯ 1. Yes',
|
||
'2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)forthissession;Yes,and',
|
||
'alwaysallowaccessto/tmp/fixtureforthissession',
|
||
'3.No',
|
||
'Esctocancel·Tabtoamend',
|
||
].join('\n');
|
||
|
||
test('ignores the granted menu and its redraw until a new request follows file-tool completion', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission();
|
||
expect(guard(first)).toBe('grant');
|
||
expect(guard(first)).toBe('handled');
|
||
const redraw = first + '\n' + permission();
|
||
expect(guard(redraw)).toBe('handled');
|
||
const completed = redraw + '\n●Write(/tmp/fixture/gstack-test-plan-design.md)\n' +
|
||
'⎿ Wrote320linesto../fixture/gstack-test-plan-design.md\n' + '·'.repeat(1600);
|
||
expect(classifyPlanCountFrame(completed)).toBeNull();
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + permission())).toBe('grant');
|
||
});
|
||
|
||
test('singular native Write/Edit results release a fresh identical permission', () => {
|
||
for (const result of ['⎿ Added1line,removed1line', '⎿ Wrote1lineto../fixture/plan.md', '⎿ Removed1line', '⎿\u00a0Wrote320linesto../fixture/plan.md']) {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission();
|
||
expect(guard(first)).toBe('grant');
|
||
const completed = first + '\n' + result;
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + permission())).toBe('grant');
|
||
}
|
||
});
|
||
|
||
test('the captured active Edit menu remains a permission behind a long diff repaint', () => {
|
||
const visible = permission('gstack-test-plan-ceo.md') + '\n' +
|
||
' 89 +The plan adds StripePaymentWebhookHandler outside WebhookDispatcher.\n'.repeat(40);
|
||
expect(visible.length).toBeLessThan(4096);
|
||
expect(classifyPlanCountFrame(visible)).toBeNull(); // The old 1.5 KB scan misses it.
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(guard(visible)).toBe('grant');
|
||
expect(guard(visible)).toBe('handled');
|
||
});
|
||
|
||
test('a completed Write invalidates an old menu even if polling missed the original grant', () => {
|
||
const visible = permission() + '\n⎿ Wrote320linesto../fixture/gstack-test-plan-design.md';
|
||
expect(createPlanCountPermissionGuard()(visible)).toBe('handled');
|
||
});
|
||
|
||
test('proposed results and tool headers do not release the same pending permission', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
let visible = permission();
|
||
expect(guard(visible)).toBe('grant');
|
||
for (const line of ['320 +⎿ Wrote320lines', '●Write(/tmp/fixture/plan.md)', '⎿ Tip: use /btw', '⎿ Error: denied']) {
|
||
visible += '\n' + line + '\n' + permission();
|
||
expect(guard(visible)).toBe('handled');
|
||
}
|
||
});
|
||
|
||
test('a different file and a genuine native file-policy question retain their own input', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
const first = permission('first.md');
|
||
expect(guard(first)).toBe('grant');
|
||
expect(guard(first + '\n' + permission('FIRST.md'))).toBe('grant'); // Targets remain case-sensitive.
|
||
expect(guard(first + '\n' + permission('second.md'))).toBe('grant');
|
||
const question = '\n☐ File policy\nDo you want to create first.md?\n❯1.Yes\n2.No\n' +
|
||
'Enter to select · ↑/↓ to navigate · Esc to cancel';
|
||
expect(guard(first + question)).toBeNull();
|
||
expect(capturePlanCountQuestion(first + question, new Set(), 0, true)?.promptSnippet).toContain('File policy');
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng setup ordering (captured F)', () => {
|
||
// Exact native question stems and choice labels: two setup calls followed
|
||
// by five real findings. Descriptions do not establish phase identity.
|
||
const rows = [
|
||
{
|
||
"header": "Learnings scope",
|
||
"question": "D1 \u2014 Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>",
|
||
"options": [
|
||
"Enable cross-project (Recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project (Recommended)"
|
||
},
|
||
{
|
||
"header": "Scope complexity",
|
||
"question": "D2 \u2014 The plan introduces 4 new classes across 12 files. That's above the complexity threshold (>2 classes / >8 files). Should we reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-complexity>",
|
||
"options": [
|
||
"Reduce: merge to 2 classes (Recommended)",
|
||
"Proceed as-is \u2014 4 classes, 12 files",
|
||
"Reduce further: 1 new class only"
|
||
],
|
||
"answer": "Reduce: merge to 2 classes (Recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: shared state",
|
||
"question": "D3 \u2014 Architecture Issue 1: AuthCache is a global mutable singleton exported at module level; both AuthBroker and (previously) SessionMint mutate it. This creates concurrent-mutation risk across tenant requests and makes the services untestable in isolation. <gstack-qid:plan-eng-arch-global-cache>",
|
||
"options": [
|
||
"Inject AuthCache via constructor (Recommended)",
|
||
"Keep global, add locking",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Inject AuthCache via constructor (Recommended)"
|
||
},
|
||
{
|
||
"header": "Code quality",
|
||
"question": "D4 \u2014 Code Quality Issue 1: validateAndDispatch() is 60 lines with three nested try/catch blocks, each swallowing a different error class. Swallowed errors mean callers can't distinguish an IDP timeout from a policy rejection from a token parse failure \u2014 all three silently return the same result. <gstack-qid:plan-eng-quality-error-handling>",
|
||
"options": [
|
||
"Extract + typed error union (Recommended)",
|
||
"Add logging to each catch, keep structure",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Extract + typed error union (Recommended)"
|
||
},
|
||
{
|
||
"header": "Test: regression",
|
||
"question": "D5 \u2014 Test Issue 1 (IRON RULE): legacyAuthFlow() is being rewritten with no regression test for its prior behavior. The plan explicitly says coverage 'does not exercise legacyAuthFlow() or assert compatibility with its prior behavior.' A rewrite without a behavioral snapshot means any regression is invisible until production. <gstack-qid:plan-eng-test-legacy-regression>",
|
||
"options": [
|
||
"Add characterization tests before rewrite (Recommended)",
|
||
"Document expected behavior, manual verify",
|
||
"Skip regression coverage"
|
||
],
|
||
"answer": "Add characterization tests before rewrite (Recommended)"
|
||
},
|
||
{
|
||
"header": "Test: isolation",
|
||
"question": "D6 \u2014 Test Issue 2: The plan says 'unit and integration coverage is planned for success/error paths' but makes no mention of cross-tenant isolation tests. The AuthCache key includes tenant ID, issuer, audience, and policy version \u2014 a key-construction bug would let Tenant A read Tenant B's cached tokens. This is the highest-severity failure mode in a multi-tenant auth system. <gstack-qid:plan-eng-test-tenant-isolation>",
|
||
"options": [
|
||
"Add explicit cross-tenant isolation tests (Recommended)",
|
||
"Cover via integration tests only",
|
||
"Proceed with existing test plan"
|
||
],
|
||
"answer": "Add explicit cross-tenant isolation tests (Recommended)"
|
||
},
|
||
{
|
||
"header": "Perf: IDP calls",
|
||
"question": "D7 \u2014 Performance Issue 1: Token validation makes 5 sequential API calls to the IDP. The plan itself notes they are independent and could be parallelized via Promise.all 'trivially.' Sequential calls add latency proportional to IDP round-trip time x5 on every auth request. At p99 IDP latency of 100ms, that's 500ms of unnecessary serialization per login. <gstack-qid:plan-eng-perf-parallel-idp>",
|
||
"options": [
|
||
"Parallelize with Promise.all in this PR (Recommended)",
|
||
"Defer to follow-up ticket",
|
||
"Defer with in-code TODO comment"
|
||
],
|
||
"answer": "Parallelize with Promise.all in this PR (Recommended)"
|
||
}
|
||
];
|
||
const fingerprint = (index: number): AskUserQuestionFingerprint => {
|
||
const row = rows[index]!;
|
||
return nativePlanCallFingerprint({
|
||
sessionId: 'captured-f-eng', toolUseId: `f-${index}`, answered: true, failed: false,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer },
|
||
}, index, true);
|
||
};
|
||
const phasesFor = (indices: number[]) => {
|
||
let started = false;
|
||
return indices.map(index => {
|
||
const phase = planCountQuestionPhase(fingerprint(index), started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
};
|
||
|
||
test('both setup orders exclude setup and include the first of five actual findings', () => {
|
||
for (const order of [[0, 1], [1, 0]]) {
|
||
const phases = phasesFor([...order, 2, 3, 4, 5, 6]);
|
||
expect(phases.slice(0, 2).map(p => p.preReview)).toEqual([true, true]);
|
||
expect(phases[1]!.reviewStarted).toBe(true);
|
||
expect(phases[2]!.preReview).toBe(false);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(5);
|
||
}
|
||
});
|
||
|
||
test('setup IDs plus opposed answered choices survive header and question-body variants', () => {
|
||
for (const index of [0, 1]) {
|
||
const original = fingerprint(index).nativeCall!;
|
||
const id = index === 0 ? 'cross-project-learnings' : 'plan-eng-scope-complexity';
|
||
for (const header of ['Learnings scope', 'Scope complexity', 'Architecture', '']) {
|
||
const call = structuredClone(original);
|
||
const q = call.questions[0]!;
|
||
q.header = header;
|
||
q.question = `Choose the setup scope. <gstack-qid:${id}>`;
|
||
q.options.reverse();
|
||
call.answers = { [q.question]: q.options[0]!.label };
|
||
const fp = nativePlanCallFingerprint(call, 0, true);
|
||
expect(engSetupAUQ(fp)).toBe(true);
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('repeated setup does not become a finding after the boundary opens', () => {
|
||
const phases = phasesFor([0, 0, 1, 1, 2, 3, 4, 5, 6]);
|
||
expect(phases.slice(0, 4).every(p => p.preReview)).toBe(true);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(5);
|
||
});
|
||
|
||
test('only successful answered setup metadata can exclude a call', () => {
|
||
for (const index of [0, 1]) {
|
||
const fp = fingerprint(index);
|
||
expect(engSetupAUQ({ ...fp, nativeCall: undefined })).toBe(false);
|
||
for (const alter of [
|
||
(call: any) => { call.answered = false; },
|
||
(call: any) => { call.failed = true; },
|
||
(call: any) => { call.answers = {}; },
|
||
(call: any) => { call.questions[0].question = 'Review cache isolation. <gstack-qid:plan-eng-cache-complexity>'; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Apply fix' }, { label: 'Defer finding' }]; },
|
||
(call: any) => { call.questions[0].options = [{ label: 'Enable cross-project with project-scoped storage; proceed as-is or reduce' }, { label: 'Discuss' }]; },
|
||
]) {
|
||
const call = structuredClone(fp.nativeCall!);
|
||
alter(call);
|
||
// Keep a successful answer after question/option mutations, so those
|
||
// controls exercise identity/actions rather than an absent answer key.
|
||
if (Object.keys(call.answers ?? {}).length) {
|
||
call.answers = { [call.questions[0].question]: call.questions[0].options[0].label };
|
||
}
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('an answered sibling cannot turn an unanswered setup tab or mixed finding packet into setup', () => {
|
||
const call = structuredClone(fingerprint(0).nativeCall!);
|
||
const issue = fingerprint(2).nativeCall!.questions[0]!;
|
||
call.questions.push(issue);
|
||
call.answers = { [issue.question]: issue.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
call.answers[call.questions[0]!.question] = call.questions[0]!.options[0]!.label;
|
||
const fp = nativePlanCallFingerprint(call, 0, false);
|
||
expect(engSetupAUQ(fp)).toBe(false);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
});
|
||
|
||
test('the optional predicate leaves other callers and substantive Eng qids unchanged', () => {
|
||
const fp = fingerprint(2);
|
||
expect(engSetupAUQ(fp)).toBe(false);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
// Historical boundary detection can also fire on actual review qids;
|
||
// those must never be reused as the late-setup exclusion predicate.
|
||
fp.promptSnippet += ' <gstack-qid:plan-eng-review-global-cache>';
|
||
expect(engStep0Boundary(fp)).toBe(true);
|
||
expect(planCountQuestionPhase(fp, true, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
const setup = fingerprint(0);
|
||
expect(planCountQuestionPhase(setup, true, engStep0Boundary).preReview).toBe(false);
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng first packet and registry identities', () => {
|
||
const scope = {
|
||
header: 'Scope complexity',
|
||
question: 'D1 — Choose the whole-plan scope. <gstack-qid:plan-eng-review-scope-reduce>',
|
||
options: [{ label: 'Reduce: merge two classes (Recommended)' }, { label: 'Proceed as-is: four classes' }],
|
||
};
|
||
const finding = {
|
||
header: 'Arch: shared state',
|
||
question: 'D3 — Architecture Issue 1: AuthCache is a global mutable singleton exported at module level. <gstack-qid:plan-eng-arch-global-cache>',
|
||
options: [{ label: 'Inject an owned cache' }, { label: 'Keep the global cache' }],
|
||
};
|
||
const callWith = (questions: typeof scope[], answers: Record<string, string>) => ({
|
||
sessionId: 'eng-first-packet', toolUseId: 'mixed-call', answered: true, failed: false,
|
||
questions, answers,
|
||
});
|
||
const phase = (call: ReturnType<typeof callWith>) => planCountQuestionPhase(
|
||
nativePlanCallFingerprint(call, 0, true), false, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ,
|
||
);
|
||
|
||
test('registry setup qids stay setup and require opposed scope actions', () => {
|
||
const call = callWith([scope], { [scope.question]: scope.options[0]!.label });
|
||
const fp = nativePlanCallFingerprint(call, 0, true);
|
||
expect(engSetupAUQ(fp)).toBe(true);
|
||
expect(engFirstReviewAUQ(fp)).toBe(false);
|
||
expect(phase(call)).toEqual({ preReview: true, reviewStarted: true });
|
||
call.questions = [{ ...scope, options: [{ label: 'Change cache size' }, { label: 'Keep cache size' }] }];
|
||
call.answers = { [scope.question]: 'Change cache size' };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||
const learnings = {
|
||
header: 'Learnings scope', question: 'Choose local learning scope. <gstack-qid:preamble-cross-project-learnings>',
|
||
options: [{ label: 'Enable cross-project learnings' }, { label: 'Keep project-scoped only' }],
|
||
};
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(callWith([learnings], {
|
||
[learnings.question]: learnings.options[0]!.label,
|
||
}), 0, true))).toBe(true);
|
||
});
|
||
|
||
test('a first packet with answered setup and a real finding counts as one review call', () => {
|
||
const call = callWith([scope, finding], {
|
||
[scope.question]: scope.options[0]!.label,
|
||
[finding.question]: finding.options[0]!.label,
|
||
});
|
||
expect(phase(call)).toEqual({ preReview: false, reviewStarted: true });
|
||
// The packet is one call even with two answered tabs; the reader/counter
|
||
// dedups the unchanged session/tool-use ID rather than counting each tab.
|
||
expect(nativePlanCallFingerprint(call, 0, true).signature).toBe('eng-first-packet:mixed-call');
|
||
for (const questions of [[scope, finding], [finding, scope]]) {
|
||
expect(phase({ ...call, questions }).preReview).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('an unanswered or failed finding cannot start review from a setup packet', () => {
|
||
const call = callWith([scope, finding], { [scope.question]: scope.options[0]!.label });
|
||
expect(phase(call)).toEqual({ preReview: true, reviewStarted: true });
|
||
call.answers = { [finding.question]: finding.options[0]!.label };
|
||
expect(phase(call).preReview).toBe(false);
|
||
for (const invalid of [{ ...call, answered: false }, { ...call, failed: true }, { ...call, answers: {} }]) {
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(invalid, 0, true))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('only positive substantive native question identity starts review', () => {
|
||
for (const qid of ['plan-eng-review-scope-reduce', 'plan-eng-scope-complexity', 'cross-project-learnings', 'plan-eng-review-next-steps']) {
|
||
const q = { ...finding, question: finding.question.replace('plan-eng-arch-global-cache', qid) };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
}
|
||
for (const qid of ['plan-eng-review-arch-finding', 'plan-eng-review-test-gap']) {
|
||
const q = { ...finding, question: finding.question.replace('plan-eng-arch-global-cache', qid) };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(true);
|
||
}
|
||
for (const qid of ['plan-eng-arch-focus', 'plan-eng-test-focus', 'plan-eng-quality-mode', 'plan-eng-perf-next-steps']) {
|
||
const q = { ...finding, header: 'Review setup', question: `Choose which issue to review first. <gstack-qid:${qid}>` };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
}
|
||
const testScope = { ...finding, header: 'Test scope', question: 'D3 — Test Issue: no coverage of the rewritten legacy flow is specified. <gstack-qid:plan-eng-test-scope>' };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([testScope], { [testScope.question]: testScope.options[0]!.label }), 0, true))).toBe(true);
|
||
const q = { ...finding, header: 'Architecture', question: 'Choose a review focus. <gstack-qid:plan-eng-review-arch-finding>' };
|
||
expect(engFirstReviewAUQ(nativePlanCallFingerprint(callWith([q], { [q.question]: q.options[0]!.label }), 0, true))).toBe(false);
|
||
});
|
||
});
|
||
|
||
|
||
describe('native Eng setup semantics (captured G)', () => {
|
||
// Exact native question stems, offered actions and answers. Five findings
|
||
// and both substantive TODO decisions must remain in review coverage.
|
||
const rows = [
|
||
{
|
||
"header": "Cross-project",
|
||
"question": "gstack can search learnings from your other projects on this machine to find patterns that might apply to this auth refactor review. This stays local \u2014 no data leaves your machine. Enable cross-project learnings? <gstack-qid:cross-project-learnings>",
|
||
"options": [
|
||
"Enable (recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable (recommended)"
|
||
},
|
||
{
|
||
"header": "Scope",
|
||
"question": "D1 \u2014 Scope challenge: this plan touches 12 files and introduces 4 new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-scope-challenge>",
|
||
"options": [
|
||
"Reduce: phase it (recommended)",
|
||
"Proceed as-is",
|
||
"Investigate first"
|
||
],
|
||
"answer": "Reduce: phase it (recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: cache wiring",
|
||
"question": "D2 \u2014 Architecture issue 1: AuthBroker and SessionMint share a global mutable AuthCache via module-level export. Module-level singletons prevent test isolation and break in multi-process deployments (cluster, serverless, worker threads). How should AuthCache be wired? <gstack-qid:plan-eng-arch-cache-wiring>",
|
||
"options": [
|
||
"Dependency injection (recommended)",
|
||
"Keep module-level export"
|
||
],
|
||
"answer": "Dependency injection (recommended)"
|
||
},
|
||
{
|
||
"header": "Arch: concurrency",
|
||
"question": "D3 \u2014 Architecture issue 2: the plan explicitly states mutations to AuthCache are not serialized. With two services writing concurrently (e.g., AuthBroker evicting a token while SessionMint reads-then-writes it), you get classic check-then-act races. How should this be handled? <gstack-qid:plan-eng-arch-mutation-serialization>",
|
||
"options": [
|
||
"Make mutations idempotent + last-write-wins (recommended)",
|
||
"Add an async mutex per cache key",
|
||
"Document and defer"
|
||
],
|
||
"answer": "Make mutations idempotent + last-write-wins (recommended)"
|
||
},
|
||
{
|
||
"header": "Code quality",
|
||
"question": "D4 \u2014 Code quality issue 1: validateAndDispatch() is 60 lines with three nested try/catch blocks, each silently swallowing a different error class. Swallowed errors mean silent failures in production \u2014 a token validation error looks identical to a dispatch error from the outside. Fix approach? <gstack-qid:plan-eng-cq-error-swallowing>",
|
||
"options": [
|
||
"Extract + typed errors (recommended)",
|
||
"Add structured logging before swallowing",
|
||
"Proceed as-is"
|
||
],
|
||
"answer": "Extract + typed errors (recommended)"
|
||
},
|
||
{
|
||
"header": "Tests",
|
||
"question": "D5 \u2014 Test issue: all 13 new code paths in AuthCache and AuthBroker are untested (0% coverage planned). The plan mentions unit and integration coverage for success/error paths, but specifics are absent. Add explicit test requirements to the plan now? <gstack-qid:plan-eng-test-coverage>",
|
||
"options": [
|
||
"Add explicit test plan (recommended)",
|
||
"Keep plan vague, trust implementation"
|
||
],
|
||
"answer": "Add explicit test plan (recommended)"
|
||
},
|
||
{
|
||
"header": "Performance",
|
||
"question": "D6 \u2014 Performance issue: token validation makes 5 sequential IDP API calls. The plan notes they are independent and could be parallelized via Promise.all trivially. Fix in this PR or defer? <gstack-qid:plan-eng-perf-idp-parallelization>",
|
||
"options": [
|
||
"Fix in this PR with Promise.all (recommended)",
|
||
"Defer to follow-up TODO"
|
||
],
|
||
"answer": "Fix in this PR with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"header": "TODO: PR 2",
|
||
"question": "D7 \u2014 TODO: capture PR 2 scope (SessionMint, TokenStore, RequestPolicy) in TODOS.md so it doesn\u2019t get lost after PR 1 ships. Add it? <gstack-qid:plan-eng-todo-pr2-scope>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Skip \u2014 not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
},
|
||
{
|
||
"header": "TODO: IDP retry",
|
||
"question": "D8 \u2014 TODO: IDP circuit breaker. Token validation makes 5 IDP calls (now parallelized). If the IDP is degraded, all 5 fail together \u2014 no retry, no fallback, no circuit breaker in the plan. Add a TODO to add circuit breaker / exponential retry logic around IDP calls? <gstack-qid:plan-eng-todo-idp-circuit-breaker>",
|
||
"options": [
|
||
"Add to TODOS.md (recommended)",
|
||
"Build it now in this PR",
|
||
"Skip \u2014 not valuable enough"
|
||
],
|
||
"answer": "Add to TODOS.md (recommended)"
|
||
}
|
||
];
|
||
const callAt = (index: number) => {
|
||
const row = rows[index]!;
|
||
return { sessionId: 'captured-g-eng', toolUseId: `g-${index}`, answered: true, failed: false,
|
||
questions: [{ header: row.header, question: row.question, options: row.options.map(label => ({ label })) }],
|
||
answers: { [row.question]: row.answer } };
|
||
};
|
||
const setup = (call: ReturnType<typeof callAt>) => engSetupAUQ(nativePlanCallFingerprint(call, 0, false));
|
||
const change = (index: number, question: string, labels?: string[]) => {
|
||
const call = callAt(index); const q = call.questions[0]!;
|
||
q.question = question;
|
||
if (labels) q.options = labels.map(label => ({ label }));
|
||
call.answers = { [question]: q.options[0]!.label };
|
||
return call;
|
||
};
|
||
test('two setup calls precede five findings and two substantive TODO calls', () => {
|
||
for (const order of [[0, 1], [1, 0]]) {
|
||
let started = false;
|
||
const phases = [...order, 2, 3, 4, 5, 6, 7, 8].map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted; return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
}
|
||
expect(setup(callAt(0))).toBe(true);
|
||
expect(setup(callAt(1))).toBe(true);
|
||
for (const index of [2, 3, 4, 5, 6, 7, 8]) expect(setup(callAt(index))).toBe(false);
|
||
});
|
||
test('whole-plan scope uses its premise and opposed actions, not a model-chosen qid or header', () => {
|
||
for (const question of [
|
||
'D1 — Scope challenge: this plan touches 12 files and introduces 4 new classes. Reduce scope or proceed as-is?',
|
||
'The plan introduces 4 new services. Reduce the scope or proceed as-is? <gstack-qid:plan-eng-unfamiliar-size-check>',
|
||
'Complexity check: the plan spans 12 files. Reduce scope or proceed as-is? <gstack-qid:another-scope-name>',
|
||
]) {
|
||
const call = change(1, question); call.questions[0]!.header = 'Review setup';
|
||
expect(setup(call)).toBe(true);
|
||
call.questions[0]!.options.reverse();
|
||
call.answers = { [question]: call.questions[0]!.options[0]!.label };
|
||
expect(setup(call)).toBe(true);
|
||
}
|
||
});
|
||
test('abbreviated enable requires explicit cross-project learnings and an opposed project-scoped action', () => {
|
||
for (const question of [
|
||
'Should gstack search learnings from your other projects on this machine?',
|
||
'Enable cross-project learnings for this review? <gstack-qid:model-selected-learning-scope>',
|
||
]) {
|
||
const call = change(0, question); call.questions[0]!.header = 'Review setup';
|
||
expect(setup(call)).toBe(true);
|
||
call.questions[0]!.options.reverse();
|
||
call.answers = { [question]: call.questions[0]!.options[0]!.label };
|
||
expect(setup(call)).toBe(true);
|
||
}
|
||
for (const question of ['Enable the new feature?', 'Choose setup. <gstack-qid:cross-project-learnings>']) {
|
||
expect(setup(change(0, question))).toBe(false);
|
||
}
|
||
expect(setup(change(0, rows[0]!.question, ['Enable (recommended)', 'Discuss later']))).toBe(false);
|
||
});
|
||
test('individual issues and TODOs cannot borrow whole-plan scope or cross-project words', () => {
|
||
for (const [index, question, header] of [
|
||
[1, 'D1 — Architecture issue: this plan touches 12 files and adds 4 classes, but the AuthCache has a race. Reduce scope or proceed as-is?', 'Architecture issue'],
|
||
[1, 'D1 — Scope challenge: this cache spans 12 files and introduces 4 classes. Reduce scope or proceed as-is?', 'Scope'],
|
||
[1, 'D1 — Test gap: the plan spans 12 files. Reduce test scope or proceed as-is?', 'Tests'],
|
||
[1, 'The plan spans 12 files. Reduce scope or proceed as-is?', 'TODO: deferred implementation'],
|
||
[0, 'D1 — Security issue: cross-project learnings leak client data. Enable the feature or use project-scoped storage?', 'Security issue'],
|
||
] as const) {
|
||
const call = change(index, question); call.questions[0]!.header = header;
|
||
expect(setup(call)).toBe(false);
|
||
}
|
||
expect(setup(change(1, rows[1]!.question, ['Reduce token scope', 'Investigate']))).toBe(false);
|
||
});
|
||
test('the captured retry scope heading and letter-prefixed native actions remain setup', () => {
|
||
const retry = {
|
||
"header": "Scope",
|
||
"question": "D2 \u2014 Scope reduction: 12 files + 4 new classes exceeds the complexity threshold\nProject/branch/task: Multi-tenant Auth Refactor on main\nELI10: The plan introduces 4 new classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files. That\u2019s a lot of new surface area at once. Auth refactors are already high-risk (broken auth = all users locked out). Adding 4 new abstractions simultaneously makes the blast radius of a mistake much larger. A simpler split \u2014 just AuthBroker + SessionMint, collapsing TokenStore into AuthCache and inlining RequestPolicy \u2014 would achieve the same goal with 2 new classes and fewer files touched.\nStakes if we pick wrong: With 4 new classes landing together, a single bug in any one of them could take down auth for all tenants simultaneously. Fewer classes = smaller blast radius, easier rollback, faster onboarding for the next engineer.\nRecommendation: B (proceed as-is) \u2014 the plan already has the existing cache adapter retained, and the 4-class split may reflect genuine domain separations the plan description doesn\u2019t fully explain. Review the design at full scope, flag individual issues per section.\nCompleteness: Note: options differ in kind, not coverage \u2014 no completeness score.\nPros / cons:\nA) Reduce scope \u2014 collapse TokenStore into AuthCache, inline RequestPolicy\n \u2705 Fewer moving parts: 2 new classes instead of 4, smaller blast radius if auth fails\n \u2705 Easier to review, test, and roll back each piece independently\n \u274c May discard intentional domain separation the plan author had in mind\n \u274c Requires re-planning before implementation can start\nB) Proceed as-is with full review (recommended)\n \u2705 Respects the planned architecture and lets the full review surface real issues per section\n \u2705 Faster path to implementation if the 4-class split turns out to be justified\n \u274c Higher blast radius: 4 simultaneous new classes touching 12 files is more fragile to ship\n \u274c The legacyAuthFlow rewrite without a regression test is a landmine that needs explicit attention\nNet: You\u2019re trading blast-radius safety (fewer classes) against re-planning delay. Recommend B \u2014 proceed at full scope, but treat each class boundary and the missing regression test as explicit issues in the review. <gstack-qid:plan-eng-review-scope-challenge>",
|
||
"options": [
|
||
"A) Reduce scope",
|
||
"B) Proceed as-is (Recommended)"
|
||
],
|
||
"answer": "A) Reduce scope"
|
||
};
|
||
const call = change(1, retry.question, retry.options);
|
||
call.questions[0]!.header = retry.header;
|
||
call.answers = { [retry.question]: retry.answer };
|
||
expect(setup(call)).toBe(true);
|
||
// A numerical component issue cannot borrow the whole-plan sentence in
|
||
// the explanatory body and the same Reduce/Proceed choices.
|
||
const issue = change(1, retry.question.replace('Scope reduction:', 'AuthCache issue:'), retry.options);
|
||
expect(setup(issue)).toBe(false);
|
||
const partial = change(1, retry.question.replace('ELI10: The plan introduces', 'ELI10: AuthCache introduces'), retry.options);
|
||
expect(setup(partial)).toBe(false);
|
||
const missingOpposition = change(1, retry.question, ['A) Reduce scope', 'B) Investigate']);
|
||
expect(setup(missingOpposition)).toBe(false);
|
||
});
|
||
test('late setup cannot hide earlier or later substantive findings', () => {
|
||
const order = [2, 1, 3, 0, 4, 5, 6, 7, 8];
|
||
let started = false;
|
||
const phases = order.map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted; return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([false, true, false, true, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
});
|
||
test('pending, failed, uncertain and mixed answered calls never become wholly setup', () => {
|
||
for (const index of [0, 1]) {
|
||
const original = callAt(index);
|
||
expect(engSetupAUQ({ ...nativePlanCallFingerprint(original, 0, false), nativeCall: undefined })).toBe(false);
|
||
for (const mutate of [
|
||
(c: ReturnType<typeof callAt>) => { c.answered = false; },
|
||
(c: ReturnType<typeof callAt>) => { c.failed = true; },
|
||
(c: ReturnType<typeof callAt>) => { c.answers = {}; },
|
||
(c: ReturnType<typeof callAt>) => { c.answers = { [c.questions[0]!.question]: 'Uncertain; I have not selected a choice' }; },
|
||
]) { const c = structuredClone(original); mutate(c); expect(setup(c)).toBe(false); }
|
||
const mixed = structuredClone(original); const finding = callAt(2).questions[0]!;
|
||
mixed.questions.push(finding);
|
||
mixed.answers[finding.question] = finding.options[0]!.label;
|
||
expect(setup(mixed)).toBe(false);
|
||
expect(planCountQuestionPhase(nativePlanCallFingerprint(mixed, 0, true), false,
|
||
engStep0Boundary, engFirstReviewAUQ, engSetupAUQ).preReview).toBe(false);
|
||
mixed.answers = { [finding.question]: finding.options[0]!.label };
|
||
expect(setup(mixed)).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
describe('completed permission cannot become a queued review answer (captured G)', () => {
|
||
// Exact captured post-Write frame; the temporary repository path is sanitized.
|
||
// The damaged "wat" came from the CLI redraw, not the underlying question.
|
||
const captured = " real,specificgaps (Visual Hierachy, Spacing, Color, Typography,Motion)\r22-Itpreservesstrongaccessibilityandresponsivespecsfromtheexistingbehaviordescription\r23 -DESIGN.md exist andsuppliescorrectvaluesforall5gaps\r24\r 25 A 10/10would:\r26-Specifytheexactchangeforeachgap(concretetoken,before→after)\r27 -Add intection state table(loading,empty,error,success,partial)\r28-Storyboardtheuserjourney\r 29- Call out which DESIGN.md tokens each fix applies\r30-Resolveallimplementationambiguitiesbeforeadeveloperpicksitup\r31\r32 **DESIGN.md status:** Presnt and detailed. Spacing scale, clor system, typography,interacton patterns,andc\r omponentbehavior re all specified. This review calbrates against it.\r33\r34 **Exsting design lverage:**\r35 -Button, Field, InlineStatus,ErrorSummary,ConfirmationDialogcomponentsreusedas-is\r36-Focusring:2pxsolid#1d4ed8,offset2px(alreadyspecified)\r37 -8px bae spacingscalealreadyestablished\r38\r39 ---\r40\r41 ## Review Passes\r42\r43*(Populatedprogressivelyasthereviewproceeds)*\r44\r╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌\r Do you wat to create gstack-test-plan-design.md?\r❯ 1. Yes\r2.Yes,andswitchtoacceptedits(auto-approvefileeditsandcommonfilecommands)for this session; Yes,an\r always allow access to /tmp/fixture for this session\r 3.Nohift+tab)\r\rEsc to cancl ·Tab to amend\u000f\u001b[?1000h\u001b[?1002h\u001b[?1003h\u001b[?1006h\rRan1shellcommand\r\r●Scopegate:planmodeactive—auto-selectingB(reviewingPLAN.mdintheworkingdirectory).\r\rSkippingroutinginjection—thisisareviewfixturerepo.Nowrunningthepre-reviewaudit.\r\rRead2 files, ran1 shell command\r\r●Now let me read the review sectons fil as required by theskill.\r\rRead file, ran 1shellcommand\r\r●Nopriorlearnings found. Cross-project config isnot set — will ask once Ihavetheuser'sattention.\r\rPreReview System Audit:\r- DESIGN.md:exists and comprehensive (spacing, typography, color, interactionpatternsallspecified)\r- UIscope:APPUIclassifier—account-settingsformwithProfile+Notificationssections\r- 5 gaps explicily listed in the plan; l have corresponding specin DESIGN.md\r-N rior design reviews inth reviewlog\r- NoTODOS.mdexists\r\rStep0 Assessment:\r\rInitial scor: 5/10. The plan correctly identifies5realgapsandpreservesstrongaccessibilityandresponsivespecs\rinheritedfromtheexistingbehaviordescription.A10/10wouldspecifyHOWeachgapisresolved(concretecomponent\rchanges,tokenvalues,before/afterlayout),notjustnametheproblem.TheDESIGN.mdalreadyhasallcorrectvalues;\rtheplanjustneedstoreferencethem.\r\rWritingtheinitialplanfilenowbeforeaskingStep0D.\r\r●Write(/tmp/fixture/gstack-test-plan-design.md)\r\r✢ Undulatig… (2m 0s ·↓ 5.4 okens)\r ⎿ Tip:Use/btwtoaskaquicksidequestionwithoutinterruptingClaude'scurrentwork\r ● high · /effort\r────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r❯ \r────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\r ⏸ plan mod on(shift+tab tocycle)·esctointerrupt·←foragents\u001b[?25h\u001b[?25l\r●No prior earnings found.Cross-projectconfigisnotset—willaskonceIhavetheuser'sattention.\r Pre-Review System Audit:\r-DESIGN.md:existsandcomprehensive(spacing,typography,color,interactionpatternsallspecified)\r- UI scope: APP UI classifier — account-settings fomwith Prfile + Notifications sctions\r-5gapsexplicitlylistedintheplan;allhavecorrespondingspecinDESIGN.md\r- Noprior desigreviwsin the reviewlog\r-NoTODOS.mdexists\r\rStep0Assessment:\r\rInitialscore:5/10.Theplancorrectlyidentifies5realgapsandpreservesstrongaccessibilityandresponsivespecs\r inheted from theexisting behavir desription.A 10/10would specify HOWechgapi eolved (ccretecomponent\rchanges,tokenvalues,before/afterlayout),notjustnametheproblem.TheDESIGN.mdalreadyhasallcorrectvalues;\rth plan just neds to referencethem.\r\rWriting theinitial plan lenow before ask Step 0D.\r\r●Write(/tmp/fixture/gstack-test-plan-design.md)\r⎿ Wrote 44 lines";
|
||
|
||
test('a missed grant followed by native Write completion sends no stale answer', () => {
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(classifyPlanCountFrame(captured)).toBeNull();
|
||
expect(guard(captured)).toBe('handled');
|
||
expect(capturePlanCountQuestion(captured, new Set(), 0, true)).toBeNull();
|
||
});
|
||
|
||
test('the same damaged active file permission grants once and never counts as a finding', () => {
|
||
const menu = captured.slice(captured.indexOf('Do you wat'), captured.indexOf('Esc to cancl')) + 'Esc to cancl ·Tab to amend';
|
||
const guard = createPlanCountPermissionGuard();
|
||
expect(guard(menu)).toBe('grant');
|
||
expect(guard(menu)).toBe('handled');
|
||
expect(capturePlanCountQuestion(menu, new Set(), 0, false)).toBeNull();
|
||
const completed = menu + '\n⎿ Wrote 44 lines';
|
||
expect(guard(completed)).toBe('handled');
|
||
expect(guard(completed + '\n' + menu)).toBe('grant');
|
||
});
|
||
|
||
test('plain legacy file decisions remain questions without native permission controls', () => {
|
||
const frame = 'Do you want to create first.md?\n❯1.Create the reviewed file\n2.Keep the current layout';
|
||
expect(classifyPlanCountFrame(frame)).toBeNull();
|
||
expect(createPlanCountPermissionGuard()(frame)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)?.options).toHaveLength(2);
|
||
});
|
||
|
||
test('a matching native finding can discuss file permissions without being consumed', () => {
|
||
const question = {
|
||
header: 'File policy',
|
||
question: 'Should we create a file that documents always allow access to the project?',
|
||
options: [{ label: 'Create it' }, { label: 'Keep current policy' }],
|
||
};
|
||
const pending = { sessionId: 'file-policy', toolUseId: 'file-finding', answered: false, questions: [question] };
|
||
const frame = captured + '\n☐ ' + question.header + '\n' + question.question +
|
||
'\n❯1.Create it\n2.Keep current policy\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||
expect(classifyPlanCountFrame(frame)).toBeNull();
|
||
expect(createPlanCountPermissionGuard()(frame)).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, pending)?.nativeCall).toBe(pending);
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)?.promptSnippet).toContain('File policy');
|
||
});
|
||
});
|
||
|
||
|
||
describe('native question identity outranks permission wording', () => {
|
||
const question = {
|
||
header: 'File policy',
|
||
question: 'D1 — Should we create a file that documents always allow access to the project? <gstack-qid:plan-eng-review-file-policy>',
|
||
options: [{ label: 'Create it' }, { label: 'Keep current policy' }],
|
||
};
|
||
const pending = { sessionId: 'file-policy', toolUseId: 'finding', answered: false, questions: [question] };
|
||
const frame = '☐ File policy\n' + question.question + '\n❯1.Create it\n2.Keep current policy\nEnter to select · ↑/↓ to navigte · Esc to cancel';
|
||
test('full native question and every option establish identity despite a damaged footer', () => {
|
||
// This is the formerly conflicting pure classifier result. The counting
|
||
// loop must consult native identity before taking its permission action.
|
||
expect(classifyPlanCountFrame(frame)).toBe('permission');
|
||
expect(matchesNativePlanQuestion(frame, pending)).toBe(true);
|
||
const seen = new Set<string>();
|
||
expect(capturePlanCountQuestion(frame, seen, 0, false, pending)?.nativeCall).toBe(pending);
|
||
expect(capturePlanCountQuestion(frame, seen, 1, false, pending)).toBeNull();
|
||
});
|
||
test('same header, changed choices, missing identity and an overlaid real permission cannot borrow a native call', () => {
|
||
for (const different of [
|
||
frame.replace('Should we create a file', 'Should we delete the file'),
|
||
frame.replace('2.Keep current policy', '2.Allow all edits'),
|
||
frame.replace(question.question, 'A different question with the same header?'),
|
||
frame + '\nDo you want to create actual.md?\n❯1.Yes\n2.Yes, and switch to accept edits\n3.No\nEsc to cancel · Tab to amend',
|
||
]) expect(matchesNativePlanQuestion(different, pending)).toBe(false);
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false, { ...pending, failed: true })).toBeNull();
|
||
expect(capturePlanCountQuestion(frame, new Set(), 0, false)).toBeNull();
|
||
});
|
||
});
|
||
|
||
|
||
describe('native file/class complexity gate classification', () => {
|
||
const rows = [
|
||
{
|
||
"id": "toolu_01HJHA5nCKyCAEPRVnm84udk",
|
||
"header": "Prerequisite",
|
||
"question": "D1 — No design doc found for this branch. Run /office-hours first? <gstack-qid:plan-eng-review-office-hours-prereq>",
|
||
"labels": [
|
||
"Skip — proceed with review (Recommended)",
|
||
"Run /office-hours first"
|
||
],
|
||
"answer": "Skip — proceed with review (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01BF2LE9pqkt3Cs8PxC1uPaA",
|
||
"header": "Scope",
|
||
"question": "D2 — Complexity check triggered: 12 files, 4+ new classes. Reduce scope or proceed as-is? <gstack-qid:plan-eng-review-complexity-check>",
|
||
"labels": [
|
||
"Proceed as-is",
|
||
"Reduce scope (Recommended)"
|
||
],
|
||
"answer": "Proceed as-is"
|
||
},
|
||
{
|
||
"id": "toolu_01J9SZHLFAuHjn19o9XFJvHo",
|
||
"header": "Learnings",
|
||
"question": "D3 — Cross-project learnings: search past sessions from other projects on this machine? <gstack-qid:plan-eng-review-cross-project-learnings>",
|
||
"labels": [
|
||
"Enable cross-project (Recommended)",
|
||
"Project-scoped only"
|
||
],
|
||
"answer": "Enable cross-project (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01CRbNCM44doMT1ubMAtcuBp",
|
||
"header": "Architecture",
|
||
"question": "D4 — Architecture A1: Two services mutate a shared AuthCache with no mutation serialization. How should this be resolved? <gstack-qid:plan-eng-review-arch-shared-mutable-cache>",
|
||
"labels": [
|
||
"Single write-coordinator (Recommended)",
|
||
"Serialize via mutex/queue",
|
||
"Accept the risk, add monitoring"
|
||
],
|
||
"answer": "Single write-coordinator (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01UwpB3rC52W99ZB6g6FMV6K",
|
||
"header": "Architecture",
|
||
"question": "D5 — Architecture A2: legacyAuthFlow() gets rewritten with no regression tests capturing prior behavior. Approach? <gstack-qid:plan-eng-review-arch-legacy-rewrite>",
|
||
"labels": [
|
||
"Regression tests first (Recommended)",
|
||
"Strangler fig",
|
||
"Big-bang rewrite as planned"
|
||
],
|
||
"answer": "Regression tests first (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_011biwzYyT1qCzVL8rn8gmTe",
|
||
"header": "Code Quality",
|
||
"question": "D6 — Code Quality CQ1: validateAndDispatch() has 3 nested try/catch blocks that each swallow a different error class. Fix? <gstack-qid:plan-eng-review-cq-validate-dispatch>",
|
||
"labels": [
|
||
"Decompose + surface errors (Recommended)",
|
||
"Flatten catch hierarchy only",
|
||
"Leave as-is"
|
||
],
|
||
"answer": "Decompose + surface errors (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01HbUJX6fD6CN3obeoyh8S5Z",
|
||
"header": "Code Quality",
|
||
"question": "D7 — Code Quality CQ2: AuthCache shared via module-level export (implicit global). Switch to dependency injection? <gstack-qid:plan-eng-review-cq-module-export>",
|
||
"labels": [
|
||
"Dependency injection (Recommended)",
|
||
"Keep module export"
|
||
],
|
||
"answer": "Dependency injection (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_019jNc2fhvrJhHkNZKTdztUm",
|
||
"header": "Tests",
|
||
"question": "D8 — Tests T1: No described test for coordinator ordering under concurrent mutations. Add concurrency tests to the plan? <gstack-qid:plan-eng-review-test-concurrency>",
|
||
"labels": [
|
||
"Add concurrency tests (Recommended)",
|
||
"Defer to code review"
|
||
],
|
||
"answer": "Add concurrency tests (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_0132EDXRcCffruHaV1E4kKRH",
|
||
"header": "Tests",
|
||
"question": "D9 — Tests T2: No E2E tests planned for core auth flows (cross-tenant validation, suspension, logout, revocation). Add them? <gstack-qid:plan-eng-review-test-e2e-auth>",
|
||
"labels": [
|
||
"Add E2E tests for auth flows (Recommended)",
|
||
"Unit/integration only as planned"
|
||
],
|
||
"answer": "Add E2E tests for auth flows (Recommended)"
|
||
},
|
||
{
|
||
"id": "toolu_01C2sC9k9YbvAvgyWHGZUTyb",
|
||
"header": "Performance",
|
||
"question": "D10 — Performance P1: 5 sequential IDP calls per token validation. Parallelize or eliminate? <gstack-qid:plan-eng-review-perf-idp-calls>",
|
||
"labels": [
|
||
"Local JWT validation (Recommended)",
|
||
"Promise.all parallelization",
|
||
"Leave sequential as-is"
|
||
],
|
||
"answer": "Local JWT validation (Recommended)"
|
||
}
|
||
];
|
||
|
||
function callAt(index: number) {
|
||
const row = rows[index]!;
|
||
return {
|
||
sessionId: '122ddb02-2346-4a38-9824-f04f9d5d8cae',
|
||
toolUseId: row.id,
|
||
answered: true,
|
||
failed: false,
|
||
questions: [{
|
||
header: row.header,
|
||
question: row.question,
|
||
options: row.labels.map(label => ({ label })),
|
||
multiSelect: false,
|
||
}],
|
||
answers: { [row.question]: row.answer },
|
||
unansweredQuestionIndices: [],
|
||
};
|
||
}
|
||
|
||
test('actual ten calls remain three setup and seven independent review decisions in either setup order', () => {
|
||
for (const order of [[0, 1, 2], [0, 2, 1]]) {
|
||
let started = false;
|
||
const phases = [...order, 3, 4, 5, 6, 7, 8, 9].map(index => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(callAt(index), 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(7);
|
||
}
|
||
});
|
||
|
||
test('the answered numeric complexity gate does not depend on qid or action order', () => {
|
||
for (const qid of ['plan-eng-review-complexity-check', 'model-chosen-size-gate']) {
|
||
const call = callAt(1);
|
||
const q = call.questions[0]!;
|
||
q.question = q.question.replace('plan-eng-review-complexity-check', qid);
|
||
for (const reverse of [false, true]) {
|
||
if (reverse) q.options.reverse();
|
||
call.answers = { [q.question]: q.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
}
|
||
}
|
||
});
|
||
|
||
test('a component finding, TODO, missing count, or missing whole-scope opposition stays substantive', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.header = 'Architecture issue'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.header = 'TODO: scope'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.question = call.questions[0]!.question.replace('Complexity check triggered:', 'AuthCache complexity issue:'); },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.question = call.questions[0]!.question.replace('12 files, 4+ new classes', '12 cache entries'); },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.options[1]!.label = 'Reduce cache lock scope'; },
|
||
(call: ReturnType<typeof callAt>) => { call.questions[0]!.options[0]!.label = 'Investigate the cache'; },
|
||
]) {
|
||
const call = callAt(1);
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('failed, pending, unoffered and mixed answered calls cannot be discarded as setup', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof callAt>) => { call.answered = false; },
|
||
(call: ReturnType<typeof callAt>) => { call.failed = true; },
|
||
(call: ReturnType<typeof callAt>) => { call.answers = {}; },
|
||
(call: ReturnType<typeof callAt>) => { call.answers = { [call.questions[0]!.question]: 'Add a new cache repair instead' }; },
|
||
(call: ReturnType<typeof callAt>) => {
|
||
const issue = callAt(3);
|
||
call.questions.push(issue.questions[0]!);
|
||
Object.assign(call.answers, issue.answers);
|
||
},
|
||
]) {
|
||
const call = callAt(1);
|
||
mutate(call);
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
|
||
describe('explicit whole-plan scope complexity premise', () => {
|
||
const calls = [
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01RY4Ec4BVKuX7JRycy5Y7zB",
|
||
"questions": [
|
||
{
|
||
"question": "D1 — The plan's scope (12 files, 4 new classes) triggers the complexity smell check. Proceed as-is or reduce scope first? <gstack-qid:plan-eng-complexity-check>",
|
||
"header": "Scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Proceed as-is (recommended)",
|
||
"description": "Review all 4 planned classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files. The plan already flags this as a smell worth investigating — we'll scrutinize every class boundary during the Architecture section and call out anything that should be merged or deferred. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Reduce scope first",
|
||
"description": "Propose a minimal-surface version before reviewing: collapse TokenStore into the existing cache adapter extension and defer RequestPolicy extraction (it's just a slice of validateAndDispatch()). This yields ~2 new classes and ~6-8 files. The Architecture section would then review the leaner design. Completeness: 7/10 — defers policy extraction"
|
||
},
|
||
{
|
||
"label": "Skip complexity and go straight to issues",
|
||
"description": "Treat scope as fixed, skip the scope challenge, and dive directly into the 4 review sections (Architecture, Code Quality, Tests, Performance). Fastest path but misses the structural smell the plan itself flagged. Completeness: 7/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D1 — The plan's scope (12 files, 4 new classes) triggers the complexity smell check. Proceed as-is or reduce scope first? <gstack-qid:plan-eng-complexity-check>": "Proceed as-is (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:58:36.032Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01Vz57KiddxkZg3v4NiWKzsC",
|
||
"questions": [
|
||
{
|
||
"question": "D2 — Architecture: shared global mutable AuthCache between AuthBroker and SessionMint, with no serialization, in a multi-tenant system. <gstack-qid:plan-eng-arch-shared-cache>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Inject via constructor (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan says: 'Two new services share a global mutable AuthCache instance via module-level export. Both services mutate it' and 'they do not serialize mutations.' In a multi-tenant system, concurrent mutations from AuthBroker (e.g., evict-on-logout) and SessionMint (e.g., write-on-mint) to the same cache can corrupt tenant isolation boundaries without any test catching it. Fix: pass AuthCache as a constructor argument to both services. Each test provides a fresh instance; production wires it once at startup. Zero mutations to the existing adapter. Effort: human ~1h / CC ~5min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Add a mutation coordinator",
|
||
"description": "Keep the module-level export but wrap all mutating calls in a queue or lock. Protects against races but does not fix the coupling (tests still share state, import side-effects are still global). Higher complexity than DI. Effort: human ~2h / CC ~10min. Completeness: 8/10 — misses the testability problem"
|
||
},
|
||
{
|
||
"label": "Accept as-is",
|
||
"description": "Trust the underlying adapter's existing invalidation hooks to prevent cross-tenant reads. This works only if the adapter already serializes all writes — the plan does not state this, and 'they do not serialize mutations' explicitly says it does not. Accepted risk of cross-tenant cache corruption on concurrent requests. Completeness: 5/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D2 — Architecture: shared global mutable AuthCache between AuthBroker and SessionMint, with no serialization, in a multi-tenant system. <gstack-qid:plan-eng-arch-shared-cache>": "Inject via constructor (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:59:24.191Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01MpiyqoPfnhoFmN1hBtNHdu",
|
||
"questions": [
|
||
{
|
||
"question": "D3 — Architecture: the plan describes 5 IDP API calls for token validation but says nothing about partial failure semantics. <gstack-qid:plan-eng-arch-idp-partial-failure>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add explicit failure semantics to the plan (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan: 'Token validation issues 5 sequential API calls to the IDP.' No text covers what happens when call 3 of 5 returns a 503. In an auth system, the fail-open vs fail-closed decision is load-bearing for security. Recommendation: add a plan section stating the rule explicitly — 'on any IDP call failure, token validation fails closed (reject the token, return 401, do not cache a partial result).' Effort: human ~30min / CC ~3min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Defer to implementation",
|
||
"description": "Leave the failure semantics undecided in the plan; let implementers decide per-call. Risk: two developers make different assumptions and one path fails open (accepts a partially-validated token). In auth systems, a silent fail-open is a security hole, not just a bug. Completeness: 5/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D3 — Architecture: the plan describes 5 IDP API calls for token validation but says nothing about partial failure semantics. <gstack-qid:plan-eng-arch-idp-partial-failure>": "Add explicit failure semantics to the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T20:59:52.277Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_019eGcLXXteu52dC67WNtJyL",
|
||
"questions": [
|
||
{
|
||
"question": "D4 — Architecture: TokenStore is listed as one of 4 new classes but the plan never describes its responsibility. The existing cache adapter already stores tokens keyed by tenant/issuer/audience/policy. <gstack-qid:plan-eng-arch-tokenstore-purpose>",
|
||
"header": "Architecture",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Define TokenStore's boundary in the plan (recommended)",
|
||
"description": "[P2] (confidence: 8/10) Without a description, implementers may duplicate the existing adapter's logic inside TokenStore, producing two storage layers for the same data. The plan should state: what TokenStore does that the existing adapter does not, whether it holds in-memory tokens separately from the cache adapter, and how AuthCache (the facade) relates to TokenStore vs the adapter. Effort: human ~20min / CC ~2min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Keep it implicit",
|
||
"description": "Accept that TokenStore's boundary will be defined during implementation. Risk: two interpretations surface mid-sprint — one where TokenStore wraps the adapter (double indirection) and one where it holds state independently (split storage, double-write bugs). Completeness: 5/10"
|
||
},
|
||
{
|
||
"label": "Merge TokenStore into AuthCache",
|
||
"description": "If TokenStore is just a typed wrapper around what AuthCache already exposes, drop it and let AuthCache handle all token storage concerns. Reduces new-class count from 4 to 3 without losing capability. Completeness: 9/10 — assumes they genuinely overlap"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D4 — Architecture: TokenStore is listed as one of 4 new classes but the plan never describes its responsibility. The existing cache adapter already stores tokens keyed by tenant/issuer/audience/policy. <gstack-qid:plan-eng-arch-tokenstore-purpose>": "Define TokenStore's boundary in the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:00:16.358Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01WXRQF8z4JKkKvnwH7h6U6Z",
|
||
"questions": [
|
||
{
|
||
"question": "D5 — Code Quality: validateAndDispatch() is 60 lines with 3 nested try/catch blocks, each swallowing a different error class. <gstack-qid:plan-eng-cq-validate-dispatch>",
|
||
"header": "Code Quality",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Refactor as part of this PR (recommended)",
|
||
"description": "[P1] (confidence: 9/10) Plan: 'each catch swallows a different error class.' In auth code, a swallowed error is a silent failure — a rejection that should surface as a 401 becomes invisible, or a partial validation looks like success. Since this PR already touches the function (RequestPolicy extraction), refactoring it is a same-diff change. Fix: extract each catch branch into a named handler, propagate errors explicitly via throw or typed Result<T,E>; add one log line per catch so failures appear in traces. Effort: human ~2h / CC ~10min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Add logging, keep structure",
|
||
"description": "Add a structured log at each catch site so failures are at least visible, but leave the nesting and swallowing in place. Reduces debuggability debt without restructuring. Still leaves the logical errors silently absorbed — a 401 that should have been thrown may still become a phantom pass. Effort: human ~30min / CC ~5min. Completeness: 7/10 — errors are visible but not propagated"
|
||
},
|
||
{
|
||
"label": "Defer to follow-up",
|
||
"description": "Capture as a TODO and address in a later PR. Risk: the refactor grows harder once 4 new classes depend on the current swallowing behavior. The 'right behavior' becomes ambiguous when callers have already been written against the current (broken) semantics. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D5 — Code Quality: validateAndDispatch() is 60 lines with 3 nested try/catch blocks, each swallowing a different error class. <gstack-qid:plan-eng-cq-validate-dispatch>": "Refactor as part of this PR (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:00:40.446Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01UwMPiH1hZutbzPCGGQ3sUZ",
|
||
"questions": [
|
||
{
|
||
"question": "D6 — Tests: IDP partial failure (calls 1-4 of 5 succeed, call 5 fails) has no test coverage in the plan. We just added the fail-closed semantic in D3. <gstack-qid:plan-eng-test-idp-partial>",
|
||
"header": "Tests",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to plan as required test (recommended)",
|
||
"description": "[P1] (confidence: 9/10) We added 'fail closed on any IDP error' as explicit plan language in D3. That semantic needs a test that stubs each IDP call position as the one that fails (5 separate test cases, or one parametrized one) and asserts the validator returns 401 and writes nothing to cache. Without this, the fail-closed rule is a comment in a doc, not a contract in code. Effort: human ~1h / CC ~8min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Cover only all-pass and all-fail",
|
||
"description": "Test the two extreme cases (all 5 calls succeed, all fail) and skip partial-failure positions. Simpler, but misses the case where calls 1-4 succeeded and call 5 fails — the most likely real-world scenario when an IDP endpoint degrades under load. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Defer",
|
||
"description": "Capture as TODO and address post-merge. Risk: the fail-closed rule we just added has no verification path; a future refactor could silently revert it and no test would catch it. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D6 — Tests: IDP partial failure (calls 1-4 of 5 succeed, call 5 fails) has no test coverage in the plan. We just added the fail-closed semantic in D3. <gstack-qid:plan-eng-test-idp-partial>": "Add to plan as required test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:00.740Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_0152qGvSSuBnRkqTe8ix2WTj",
|
||
"questions": [
|
||
{
|
||
"question": "D7 — Tests: the plan says coverage for 'new components and their success/error paths' is planned, but validateAndDispatch() is an existing function being refactored. Its 3 catch blocks that currently swallow errors need explicit tests for each error branch. <gstack-qid:plan-eng-test-validate-dispatch-errors>",
|
||
"header": "Tests",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add 3 explicit error-path tests to the plan (recommended)",
|
||
"description": "[P2] (confidence: 8/10) After the refactor from D5, each of the 3 catch blocks becomes a named handler. Each named handler needs a test: trigger error class A/B/C, assert the function returns the expected error response (not silently succeeds). Without these tests, the refactor from D5 has no verification that the new explicit error handling is correct. Effort: human ~1h / CC ~8min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Cover in integration tests only",
|
||
"description": "Rely on integration tests that exercise validateAndDispatch() indirectly through AuthBroker. Lower signal — integration tests that cover error paths often don't isolate which branch triggered, making regressions hard to pinpoint. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Defer",
|
||
"description": "Leave error-path test coverage for a follow-up. Risk: the refactored function's error behavior is unverified until the next sprint. Completeness: 3/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D7 — Tests: the plan says coverage for 'new components and their success/error paths' is planned, but validateAndDispatch() is an existing function being refactored. Its 3 catch blocks that currently swallow errors need explicit tests for each error branch. <gstack-qid:plan-eng-test-validate-dispatch-errors>": "Add 3 explicit error-path tests to the plan (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:10.780Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_0194VcsnAMMDRFEafkx4dEsQ",
|
||
"questions": [
|
||
{
|
||
"question": "D8 — Performance: token validation makes 5 sequential IDP API calls that the plan itself identifies as trivially parallelizable. <gstack-qid:plan-eng-perf-idp-parallel>",
|
||
"header": "Performance",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Parallelize as part of this PR (recommended)",
|
||
"description": "[P2] (confidence: 9/10) Plan: 'Token validation issues 5 sequential API calls to the IDP; they could be parallelized via Promise.all trivially (calls are independent).' The plan identifies the fix itself. In a multi-tenant system, token validation is on the hot path for every authenticated request. Sequential IDP calls multiply latency by 5x unnecessarily. Fix: wrap the 5 calls in Promise.all(); the refactor from D5 (validateAndDispatch extraction) is the natural point to do this. Effort: human ~30min / CC ~5min. Completeness: 10/10"
|
||
},
|
||
{
|
||
"label": "Defer to performance pass",
|
||
"description": "Merge first, profile in production, then optimize. Risk: the plan already identifies this as trivial to fix. Deferring a known, safe, low-effort improvement costs real latency on every auth request from day 1 of deployment. Completeness: 6/10"
|
||
},
|
||
{
|
||
"label": "Add to TODOS.md only",
|
||
"description": "Capture in TODOS.md and address when performance metrics confirm it matters. Lowest risk of introducing a bug, but leaves 5x unnecessary serial latency in prod indefinitely. Completeness: 4/10"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D8 — Performance: token validation makes 5 sequential IDP API calls that the plan itself identifies as trivially parallelizable. <gstack-qid:plan-eng-perf-idp-parallel>": "Parallelize as part of this PR (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:02:26.856Z"
|
||
},
|
||
{
|
||
"sessionId": "3326e31d-a9e2-41ea-9489-f99013a81de8",
|
||
"toolUseId": "toolu_01AHbQGxj8dc6F4MPYhsPo2q",
|
||
"questions": [
|
||
{
|
||
"question": "D9 — TODOS: the plan has no mention of IDP circuit breaker or timeout per call. With 5 calls now running in parallel (D8 decision), an IDP outage generates 5 concurrent timeouts per request. <gstack-qid:plan-eng-todo-idp-circuit-breaker>",
|
||
"header": "TODO",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to TODOS.md (recommended)",
|
||
"description": "Capture as a deferred item: add per-call timeout + a circuit breaker so that IDP degradation fails fast instead of hanging. Not blocking this PR, but the parallelization we added in D8 increases the concurrent timeout surface. Concrete TODO: implement exponential backoff + circuit breaker with configurable open/half-open thresholds."
|
||
},
|
||
{
|
||
"label": "Skip — not valuable enough",
|
||
"description": "Accept that IDP timeout handling is the IDP library's responsibility, or that network timeouts are set at the HTTP client level. No additional app-level circuit breaker needed."
|
||
},
|
||
{
|
||
"label": "Build it now in this PR",
|
||
"description": "Add circuit breaker logic as part of the validateAndDispatch refactor (D5). The refactor is already touching that function; adding a circuit breaker is incremental. Effort: human ~4h / CC ~20min. Higher complexity, but IDP outages in auth systems have severe user impact."
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D9 — TODOS: the plan has no mention of IDP circuit breaker or timeout per call. With 5 calls now running in parallel (D8 decision), an IDP outage generates 5 concurrent timeouts per request. <gstack-qid:plan-eng-todo-idp-circuit-breaker>": "Add to TODOS.md (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:03:23.081Z"
|
||
}
|
||
];
|
||
|
||
test('captured retry counts its scope as setup but keeps all eight substantive decisions above the original ceiling', () => {
|
||
let started = false;
|
||
const phases = calls.map(call => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(call, 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(p => p.preReview)).toEqual([true, false, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(p => !p.preReview)).toHaveLength(8);
|
||
});
|
||
|
||
test('an individual component, missing numeric premise, or a finding cannot borrow a scope label', () => {
|
||
for (const mutate of [
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace("The plan's scope", "The cache's scope"); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace('12 files, 4 new classes', '12 cached tokens'); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.question = call.questions[0]!.question.replace('triggers the complexity smell check', 'has a race in AuthCache'); },
|
||
(call: typeof calls[number]) => { call.questions[0]!.header = 'Architecture finding'; },
|
||
]) {
|
||
const call = structuredClone(calls[0]!);
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|
||
|
||
|
||
describe('explicit Step 0 complexity gate with size in native choices', () => {
|
||
const calls = [
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01UbyixE6kojrR1enUcpg1bT",
|
||
"questions": [
|
||
{
|
||
"question": "D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>",
|
||
"header": "Routing setup",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add routing rules (recommended)"
|
||
},
|
||
{
|
||
"label": "Skip — invoke manually"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>": "Add routing rules (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:51:39.918Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01JswH5HdmeQvdAGx8P6tApN",
|
||
"questions": [
|
||
{
|
||
"question": "D2 — Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>",
|
||
"header": "Learnings scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Enable cross-project learnings (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep learnings project-scoped"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D2 — Should gstack search learnings from your other projects on this machine? <gstack-qid:cross-project-learnings>": "Enable cross-project learnings (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:51:59.954Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01Hqr5i8DN2iukjCGmj1DQro",
|
||
"questions": [
|
||
{
|
||
"question": "D3 — Step 0 Scope Challenge: This plan triggers the complexity gate. Proceed at full scope or reduce? <gstack-qid:plan-eng-step0-complexity>",
|
||
"header": "Step 0 scope",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Proceed at full scope (recommended)",
|
||
"description": "Review and implement as written: 4 new classes (TokenStore, SessionMint, AuthCache, RequestPolicy) across 12 files in one pass. The plan already self-flags this as a scope smell, and the reviewer will dig into the specific overbuilt pieces in the Architecture section.",
|
||
"preview": "Scope: 4 classes, 12 files, 2 new services\nRisk: larger blast radius; harder to roll back\nENG REVIEW will call out any specific over-abstractions"
|
||
},
|
||
{
|
||
"label": "Reduce to two phases",
|
||
"description": "Split into Phase 1 (core auth services: AuthBroker + SessionMint with direct adapter calls) and Phase 2 (TokenStore + RequestPolicy as follow-up). AuthCache facade may be mergeable into the existing adapter. Smaller blast radius, easier review. (human: +1 day planning / CC: ~10 min)",
|
||
"preview": "Phase 1: AuthBroker + SessionMint (core flow)\nPhase 2: TokenStore + RequestPolicy\nRisk: sequencing overhead; Phase 2 may stall"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D3 — Step 0 Scope Challenge: This plan triggers the complexity gate. Proceed at full scope or reduce? <gstack-qid:plan-eng-step0-complexity>": "Proceed at full scope (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:52:54.081Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01LT7ctyJNE1UqyTvLdQkULD",
|
||
"questions": [
|
||
{
|
||
"question": "D4 — Architecture A1: Global mutable AuthCache shared via module-level export, no mutation serialization. How should this be fixed? <gstack-qid:plan-eng-arch-global-cache>",
|
||
"header": "Arch: global cache",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Inject AuthCache as constructor dependency (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep global export, add write serialization"
|
||
},
|
||
{
|
||
"label": "Accept as-is — verify adapter serializes internally"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D4 — Architecture A1: Global mutable AuthCache shared via module-level export, no mutation serialization. How should this be fixed? <gstack-qid:plan-eng-arch-global-cache>": "Inject AuthCache as constructor dependency (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:53:50.210Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01LTXTWs6w4qhdo7FVRXFKdZ",
|
||
"questions": [
|
||
{
|
||
"question": "D5 — Architecture A2: Cache adapter failure posture is unstated. What should happen when the cache adapter is unavailable? <gstack-qid:plan-eng-arch-cache-failure>",
|
||
"header": "Arch: cache failmode",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Fail closed (recommended) — state this explicitly in the plan"
|
||
},
|
||
{
|
||
"label": "Fail open with logging — state this explicitly in the plan"
|
||
},
|
||
{
|
||
"label": "Accept as-is — inherited from existing adapter"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D5 — Architecture A2: Cache adapter failure posture is unstated. What should happen when the cache adapter is unavailable? <gstack-qid:plan-eng-arch-cache-failure>": "Fail closed (recommended) — state this explicitly in the plan"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:54:24.301Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01MCM7KMWBJyKPBeD5D5xQXE",
|
||
"questions": [
|
||
{
|
||
"question": "D6 — Code Quality Q1: validateAndDispatch() — 60 lines, 3 nested try/catch blocks each swallowing a different error class. How do you want to address this? <gstack-qid:plan-eng-cq-validate-dispatch>",
|
||
"header": "CQ: swallowed errors",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Refactor — split function + typed error returns (recommended)"
|
||
},
|
||
{
|
||
"label": "Keep structure, convert to catch-and-log"
|
||
},
|
||
{
|
||
"label": "Defer — leave for follow-up PR"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D6 — Code Quality Q1: validateAndDispatch() — 60 lines, 3 nested try/catch blocks each swallowing a different error class. How do you want to address this? <gstack-qid:plan-eng-cq-validate-dispatch>": "Refactor — split function + typed error returns (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:55:06.409Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01THi9A5uG6H6YGD4FQgL6vJ",
|
||
"questions": [
|
||
{
|
||
"question": "D7 — Tests T1: The plan says success/error paths will be covered, but doesn't include a test for the fail-closed cache behavior (cache adapter unavailable → auth rejected). Add it? <gstack-qid:plan-eng-test-fail-closed>",
|
||
"header": "Tests: fail-closed",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add explicit fail-closed test (recommended)"
|
||
},
|
||
{
|
||
"label": "Note as TODO — rely on adapter's existing coverage"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D7 — Tests T1: The plan says success/error paths will be covered, but doesn't include a test for the fail-closed cache behavior (cache adapter unavailable → auth rejected). Add it? <gstack-qid:plan-eng-test-fail-closed>": "Add explicit fail-closed test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:56:50.701Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01379PhfvZ4iX43gReT87gob",
|
||
"questions": [
|
||
{
|
||
"question": "D8 — Tests T2: Concurrent mint + revoke for the same tenant has no planned test. AuthBroker and SessionMint both write to the shared AuthCache instance. Add a concurrency test? <gstack-qid:plan-eng-test-concurrent>",
|
||
"header": "Tests: concurrency",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add concurrent access test (recommended)"
|
||
},
|
||
{
|
||
"label": "Defer — note in TODOS.md"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D8 — Tests T2: Concurrent mint + revoke for the same tenant has no planned test. AuthBroker and SessionMint both write to the shared AuthCache instance. Add a concurrency test? <gstack-qid:plan-eng-test-concurrent>": "Add concurrent access test (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:57:14.766Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_0112YwENaKQsigFNvBDmZiUa",
|
||
"questions": [
|
||
{
|
||
"question": "D9 — Performance P1: 5 sequential IDP calls during token validation — the plan identifies this as trivially parallelizable. Address it in this PR? <gstack-qid:plan-eng-perf-idp-calls>",
|
||
"header": "Perf: IDP calls",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Parallelize in this PR with Promise.all (recommended)"
|
||
},
|
||
{
|
||
"label": "Defer to a follow-up PR"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D9 — Performance P1: 5 sequential IDP calls during token validation — the plan identifies this as trivially parallelizable. Address it in this PR? <gstack-qid:plan-eng-perf-idp-calls>": "Parallelize in this PR with Promise.all (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:57:34.827Z"
|
||
},
|
||
{
|
||
"sessionId": "bb1c2502-5944-4872-aab4-0cc8a34296b2",
|
||
"toolUseId": "toolu_01YFVEdfATVrYA8LbWUmzN72",
|
||
"questions": [
|
||
{
|
||
"question": "D10 — TODO: Add p99 latency metric for IDP calls before/after Promise.all parallelization. Add to TODOS.md? <gstack-qid:plan-eng-todo-idp-metrics>",
|
||
"header": "TODO: IDP metrics",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to TODOS.md (recommended)"
|
||
},
|
||
{
|
||
"label": "Skip — not valuable enough"
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D10 — TODO: Add p99 latency metric for IDP calls before/after Promise.all parallelization. Add to TODOS.md? <gstack-qid:plan-eng-todo-idp-metrics>": "Add to TODOS.md (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-08T21:58:35.032Z"
|
||
}
|
||
];
|
||
const scopeCall = () => structuredClone(calls[2]!);
|
||
|
||
test('captured sequence keeps three setup calls and all seven substantive decisions', () => {
|
||
let started = false;
|
||
const phases = calls.map(call => {
|
||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(call, 0, !started),
|
||
started, engStep0Boundary, engFirstReviewAUQ, engSetupAUQ);
|
||
started = phase.reviewStarted;
|
||
return phase;
|
||
});
|
||
expect(phases.map(phase => phase.preReview)).toEqual([true, true, true, false, false, false, false, false, false, false]);
|
||
expect(phases.filter(phase => !phase.preReview)).toHaveLength(7);
|
||
expect(calls.at(-1)!.questions[0]!.header).toBe('TODO: IDP metrics');
|
||
});
|
||
|
||
test('native offered-answer binding does not depend on model qid or option order', () => {
|
||
for (const reverse of [false, true]) {
|
||
const call = scopeCall();
|
||
const question = call.questions[0]!;
|
||
question.question = question.question.replace('plan-eng-step0-complexity', 'different-model-id');
|
||
if (reverse) question.options.reverse();
|
||
call.answers = { [question.question]: question.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(true);
|
||
}
|
||
});
|
||
|
||
test('component findings, TODOs and incomplete scope evidence remain review decisions', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.header = 'Architecture finding'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.header = 'TODO: scope'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.question = call.questions[0]!.question.replace('This plan triggers the complexity gate', 'This cache triggers the complexity gate'); },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.question = call.questions[0]!.question.replace('Step 0 Scope Challenge:', 'Architecture issue:'); },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.description = 'Inspect 12 files for a cache race.'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.description = 'Implement as written: 4 new classes for token validation.'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[0]!.label = 'Proceed with a cache lock'; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.questions[0]!.options[1]!.label = 'Investigate cache failures'; },
|
||
]) {
|
||
const call = scopeCall();
|
||
mutate(call);
|
||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
|
||
test('pending, failed, unoffered and mixed answered packets cannot be discarded as setup', () => {
|
||
for (const mutate of [
|
||
(call: ReturnType<typeof scopeCall>) => { call.answered = false; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.failed = true; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.answers = {}; },
|
||
(call: ReturnType<typeof scopeCall>) => { call.answers = { [call.questions[0]!.question]: 'Add a new cache repair' }; },
|
||
(call: ReturnType<typeof scopeCall>) => {
|
||
const finding = structuredClone(calls[3]!);
|
||
call.questions.push(finding.questions[0]!);
|
||
Object.assign(call.answers, finding.answers);
|
||
},
|
||
]) {
|
||
const call = scopeCall();
|
||
mutate(call);
|
||
expect(engSetupAUQ(nativePlanCallFingerprint(call, 0, false))).toBe(false);
|
||
}
|
||
});
|
||
});
|