v1.86.0.0 feat: route outside reviews by harness (#2850)

* feat: add a restricted and supervised Claude Code runner

Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment.

* feat: route outside reviews by harness and migrate wrapper installs

Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage.

* test: recognize CEO mode labels without terminal spacing

The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions.

* test: isolate plan-count fixtures before starting review workflows

Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations.

* test: stabilize review fixtures and Claude eval startup

Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: classify collapsed review modes and isolate seeded findings

Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: isolate browser daemon state across free shards

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: stabilize native review counting and interactive navigation

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* chore: prepare v1.82.0.0 release

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: eliminate browser and process-cleanup test flakes

Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing
reused live sockets. Add an isolated GC/listener regression that fails on
Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime.

Check renderer cleanup against the render's own staging directory so
concurrent renders cannot invalidate the assertion. Make the no-pgrep
process-tree walk tolerate disappearing /proc entries, and synchronize
its test fixture through child readiness and pipe EOF instead of sleeps.

Validation: 9,157 passed, 31 skipped, zero failures across 556 files with
retries disabled. Build, all-host generation freshness, and skill checks
passed. All three races have failing-before/passing-after regressions.

* fix: count completed native review questions in evals

* fix: drive review navigation from confirmed native choices

* fix: require complete section-loading eval reports

* test: isolate telemetry HTTP transport from local assertions

* fix: keep review input on the active native question

* test: let tunnel revocation daemon choose an available port

* test: allocate available ports for pairing and watchdog fixtures

* fix: stabilize planning eval navigation and phase reporting

* test: isolate installed runtime paths in planning evals

* test: stabilize review evidence and concurrent refresh fixtures

* fix: resolve design findings before editing the plan

* fix: honor and persist disabled outside plan reviews

* fix: preserve planning decisions and terminal evidence

Load installed host reviews at autoplan phase entry and wait for completed
reviewers and saved artifacts. Reuse approved remedies while preserving
individual finding decisions.

Drive interactive evals from the current terminal viewport, bind native
questions across scrolling, and require complete native report evidence.
Cover captured stale menus, permission lifecycles, setup classification,
and disabled-review tool availability with deterministic regressions.

Advance release metadata and the upgrade migration to the unclaimed
1.83.0.0 slot.

* fix: drive native review questions and preserve current plans

Use the native single-choice keyboard protocol and current terminal viewport,
with per-question navigation inside packets and completed-call coverage.
Keep permissions, multi-select menus, and Submit controls distinct.

Send Autoplan reviewers the amended implementation plan, keep its review record
separate, and supply retained application contracts in the chain fixture.
Clarify individual DevEx decisions and complete CEO fix options; use one active
plan destination for the section-loading report.

* fix: preserve complete plan-review decisions

* fix: recognize native plan dialogs and reviewer controls

* fix: preserve review decisions and phase completion

* fix: recognize completed reviews without losing findings

* fix: preserve review continuity and native eval completion

* test: fix native review completion and eval retry isolation

* test: handle native review menus and complete eval fixtures

* test: fix native review setup, completion, and isolation failures

* test: limit native skill discovery to runtime assets

* fix: bind Autoplan reviews to full ordered phase inputs

* test: fix planning eval routing, counting, and timeout handling

* chore: advance queued release to v1.84.0.0

* fix: preserve complete review inputs and planning decisions

* fix: reconcile review approvals and preserve phase obligations

* fix: preserve review obligations and unblock eval permissions

Carry recorded Autoplan requirements into blind phase inputs, require Eng
review approvals before exit, and exercise combined asynchronous flows in
CEO reviews. Correct native finding and handoff classification and unblock
repeated report edits using scoped request identities.

* fix: retain plan requirements and complete native review dialogs

* fix: complete native review prompts and retain plan references

* fix: preserve review inputs and classify native eval evidence

* fix: check competing completion orders in CEO reviews

* fix: recognize review decisions and require phase methodology

Require the current phase methodology before Autoplan snapshots. Correct
substantive decision, closed handoff, and cache-finding classification, and
honor the recommended implementation approach in native review dialogs.

Add captured-transcript regressions without changing review thresholds,
provider models, retries, or deadlines.

* test: bind native review decisions and close completed handoffs

* fix: complete review dialogs and verify methodology delivery

* fix: preserve review evidence and unblock native eval prompts

* fix: handle native review question completions

* fix: recognize native review narration and controls

* fix: count native review decisions and isolate eval fixtures

* test: verify seeded review coverage and current artifact permissions

* test: isolate model and brain-aware skill renders

* fix: repair native workflow evaluation and clarify review steps

* fix: stabilize workflow eval evidence and review guidance

* test: repair native workflow observation and fixture isolation

* fix: recognize completed workflow evidence and owned skill reads

* test: repair seeded workflow delivery and completion evidence

* test: recognize current review evidence across native forms

* test: handle native review variants and permission redraws

* fix: honor review preferences and recognize native eval evidence

* test: recognize completed review decisions and queued permissions

* test: match current review contracts and partial-line edits

* test: recognize completed workflow evidence and bounded human waits

* fix: preserve review entry gates and native eval interactions

* fix: recognize native workflow evidence and preserve review gates

* test: recognize current review evidence and preconfigure workflow fixtures

* test: recognize completed review findings and scoped artifact permissions

* fix: stabilize native workflow review and permission evidence

* fix: recognize current review evidence and scoped edit confirmations

Clarify Design and engineering review entry instructions and Design scoring.
Recognize required legacy coverage and public Autoplan completion recaps.
Bind the pending Edit confirmation to its exact file, ordered digest, and
one-request approval when a preceding command display remains visible.
Keep reviews within their existing size limits and preserve scope gates
when extracting workflow fixtures from either supported preamble header.

Keep failure outcomes, review thresholds, provider choices, and eval budgets.

* fix: recover review workflow progress and eval evidence

* fix: recognize valid review evidence and scope selection

* test: fix review evidence parsing and repeated artifact prompts

* test: recognize valid review decisions and pending native cards

* fix(plan-eng-review): keep final navigation consistent with approved tasks

* test: recognize valid review evidence and bind legacy diff requests

* fix: stabilize review eval evidence and harness repair guidance

* docs: update project documentation for v1.85.0.0

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: fix Windows CI fixtures and credential scan

Rebase captured JSON values and filesystem evidence using the appropriate
path convention. Compile native fake CLIs on Windows and synchronize pipe
holder readiness, with cleanup retained when assertions fail.

Assemble synthetic credential fixtures at runtime so the added-line scan
keeps enforcing the same gate without flagging its own rejection controls.

Discover generated skills directly for the empty-find regression check,
avoiding a recursive scan through saved evaluation artifacts and dependencies.

* fix: preserve source renders on Windows

Compare canonical generator paths using native separators so an output
sidecar pointing at the source cannot overwrite its skill or metadata.
Keep the regression fixture isolated from the real checkout and expose
freshness diagnostics before asserting subprocess status.

Detach Windows drain-test pipe holders from the fake provider's automatic
child cleanup while preserving the enclosing runner job and its assertions.

* fix: clarify outside review fallback and CEO decisions

Render one applicable own-harness fallback path and retain native review,
disabled policy, and missing-coverage semantics. Align report field names
and mode labels, and make the existing per-cut scope approval explicit.

Regenerate skill outputs and keep the workflow judge's model, thresholds,
and retry policy unchanged.

* chore: move release to free version slot (v1.86.0.0)

PR #2852 now claims v1.85.0.0. Align the release metadata and
rename migration so upgrades from that version still receive it.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: include engineering review prerequisites and restore branch context

* fix: recognize coverage diagrams and clarify design review instructions

* fix: preserve file identities and join Windows test processes

---------

Co-authored-by: OpenAI Codex <noreply@openai.com>
This commit is contained in:
Garry Tan
2026-09-14 14:32:45 -07:00
committed by GitHub
co-authored by OpenAI Codex
parent 71f6048e8a
commit 9f81911136
767 changed files with 127420 additions and 3917 deletions
+51
View File
@@ -32,6 +32,7 @@ import {
} from '../test/helpers/agent-sdk-runner';
import {
validateFixtures,
OVERLAY_FIXTURES,
fanoutPass,
type OverlayFixture,
} from '../test/fixtures/overlay-nudges';
@@ -805,6 +806,56 @@ describe('validateFixtures', () => {
// fanoutPass predicate
// ---------------------------------------------------------------------------
describe('overlay first logical message metric', () => {
// Public SDK shape: separate assistant events share one message.id, and
// tool results may arrive between them. The initial empty public event
// carries no inspected private content. All IDs here are synthetic.
const fanout = OVERLAY_FIXTURES.filter(f => f.id.includes('-fanout-'));
function splitResponse(): AgentSdkResult {
const initial = systemInit();
const event = (messageId: string, id?: string) => {
const e = assistantTurn(id ? [{ type: 'tool_use', name: 'Read', input: {} }] : []) as any;
e.message.id = messageId;
if (id) e.message.content[0].id = id;
return e;
};
const turns = [event('first'), event('first', 'alpha'), event('first', 'beta'), event('first', 'gamma'), event('later', 'later-tool')];
const result = { type: 'user', session_id: 'test-session', parent_tool_use_id: null,
message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 'alpha', content: 'Alpha' }] } };
return { events: [initial, turns[0], turns[1], result, ...turns.slice(2)], assistantTurns: turns } as unknown as AgentSdkResult;
}
test('all fanout fixtures count one split first response across interleaved results', () => {
expect(fanout).toHaveLength(4);
for (const fixture of fanout) expect(fixture.metric(splitResponse())).toBe(3);
});
test('a combined message and repeated tool ID have the same count', () => {
for (const combined of [false, true]) {
const r = splitResponse();
if (combined) {
(r.assistantTurns[0]!.message.content as any[]).push(...r.assistantTurns.slice(1, 4).flatMap(e => e.message.content as any[]));
} else r.assistantTurns.splice(3, 0, structuredClone(r.assistantTurns[1]!));
for (const fixture of fanout) expect(fixture.metric(r)).toBe(3);
}
});
test('child, foreign-session and later-response tools cannot inflate the first response', () => {
const r = splitResponse();
const foreign = structuredClone(r.assistantTurns[1]!) as any;
foreign.session_id = 'other-session'; foreign.message.content[0].id = 'foreign';
const child = structuredClone(r.assistantTurns[1]!) as any;
child.parent_tool_use_id = 'agent-tool'; child.message.content[0].id = 'child';
r.assistantTurns.unshift(child, foreign);
for (const fixture of fanout) expect(fixture.metric(r)).toBe(3);
});
test('missing first-response identity cannot borrow a later response', () => {
for (const field of ['id', 'session_id']) {
const r = splitResponse();
if (field === 'id') (r.assistantTurns[0]!.message as any).id = '';
else (r.events[0] as any).session_id = '';
for (const fixture of fanout) expect(fixture.metric(r)).toBe(0);
}
});
});
describe('fanoutPass predicate', () => {
test('accepts mean lift >= 0.5 AND >=3/10 overlay trials >= 2', () => {
const overlay = [2, 2, 2, 2, 2, 2, 2, 2, 2, 2];
+71 -18
View File
@@ -6,7 +6,7 @@
* installed and open (macOS dev machines); the live fallback render runs
* wherever a browse binary resolves (Linux CI builds one via build:gates).
*/
import { describe, test, expect, beforeAll, afterAll, setDefaultTimeout } from 'bun:test';
import { describe, test, expect, beforeAll, afterAll, setDefaultTimeout, spyOn } from 'bun:test';
import * as fs from 'fs';
import * as os from 'os';
import * as path from 'path';
@@ -114,8 +114,7 @@ async function liveRoundTrip(engine: 'aside' | 'browse', renderFn: typeof render
],
timeoutMs: 90_000,
});
expect(out.error).toBeUndefined();
expect(out.ok).toBe(true);
expectOk(out);
expect(out.engine).toBe(engine);
expect(out.outputs).toEqual([path.join(dir, 'out.pdf'), path.join(dir, 'm.jpg'), path.join(dir, 'v.txt'), path.join(dir, 'bytes.bin')]);
expect(fs.readFileSync(path.join(dir, 'out.pdf')).subarray(0, 4).toString()).toBe('%PDF');
@@ -135,8 +134,7 @@ async function lateReadiness(engine: 'aside' | 'browse', renderFn: typeof render
fs.writeFileSync(path.join(dir, 'late.html'), '<!doctype html><title>Late</title><body><script>setTimeout(() => { window.later = { ok: true }; }, 800);</script></body>');
try {
const out = await renderFn({ file: path.join(dir, 'late.html'), waitFor: { expression: 'window.later.ok', timeoutMs: 10_000 }, steps: [{ kind: 'eval', expression: 'document.title' }], timeoutMs: 60_000 });
expect(out.error).toBeUndefined();
expect(out.ok).toBe(true);
expectOk(out);
expect(out.engine).toBe(engine);
expect(out.evals[0]).toBe('Late');
} finally {
@@ -318,7 +316,32 @@ const expectOk = (r: RenderResult): void => {
// not latency. Bun's 5s default once failed a CI run whose render was merely slow
// under a full six-shard load, so the budget is generous and hangs still fail.
setDefaultTimeout(30_000);
const browseWorkDirs = (): string[] => fs.readdirSync(SAFE_TMP_DIR).filter((n) => n.startsWith('gstack-render-browse-'));
/** Check this render's staging directory, regardless of other renders using /tmp. */
async function renderCheckingCleanup(spec: RenderSpec, bin: string, renderFn = renderWithBrowse): Promise<RenderResult> {
const workDirs: string[] = [];
const mkdtemp = fs.mkdtempSync;
// The renderer allocates before its first await. Observe that synchronous
// call, then restore immediately so unrelated async work is never captured.
const allocation = spyOn(fs, 'mkdtempSync').mockImplementation(((...args: Parameters<typeof fs.mkdtempSync>) => {
const dir = mkdtemp(...args);
if (args[0] === path.join(SAFE_TMP_DIR, 'gstack-render-browse-')) workDirs.push(dir.toString());
return dir;
}) as typeof fs.mkdtempSync);
let pending: Promise<RenderResult>;
try {
pending = renderFn(spec, bin);
} finally {
allocation.mockRestore();
}
try {
return await pending;
} finally {
// A changed allocation boundary must fail, not silently skip leak checks.
expect(workDirs, 'expected to observe this render\'s staging directory').toHaveLength(1);
for (const dir of workDirs) expect(fs.existsSync(dir), `render leaked staging directory: ${dir}`).toBe(false);
}
}
/** The subprocess driver: one job per process, so the module's engine cache and the spawn-time PATH are both under the test's control. */
function writeDriver(dir: string): string {
@@ -656,10 +679,47 @@ describe.skipIf(!HERMETIC)('aside-render: renderWithBrowse — daemon CLI contra
};
const T = '--tab-id 7';
test('cleanup remains verifiable when another render removes its staging directory', async () => {
const sibling = fs.mkdtempSync(path.join(SAFE_TMP_DIR, 'gstack-render-browse-'));
try {
// Synchronize the other render's cleanup with our newtab command, so
// this reproduces the shared-/tmp race without relying on timing.
const b = fake({ newtab: `rmdir '${sibling.replaceAll("'", "'\\''")}'\necho '{"tabId":7}'` });
const r = await renderCheckingCleanup({ file: doc, steps: [] }, b);
expectOk(r);
expect(fs.existsSync(sibling)).toBe(false);
} finally {
fs.rmSync(sibling, { recursive: true, force: true });
}
});
test('cleanup leaves another render\'s new staging directory intact', async () => {
const sibling = fs.mkdtempSync(path.join(SAFE_TMP_DIR, 'gstack-render-browse-'));
fs.rmdirSync(sibling);
try {
const b = fake({ newtab: `mkdir '${sibling.replaceAll("'", "'\\''")}'\necho '{"tabId":7}'` });
expectOk(await renderCheckingCleanup({ file: doc, steps: [] }, b));
expect(fs.existsSync(sibling)).toBe(true);
} finally {
fs.rmSync(sibling, { recursive: true, force: true });
}
});
test('cleanup check still rejects an owned staging directory leak', async () => {
let leaked: string | undefined;
try {
await expect(renderCheckingCleanup({ file: doc, steps: [] }, 'unused', async () => {
leaked = fs.mkdtempSync(path.join(SAFE_TMP_DIR, 'gstack-render-browse-'));
return { ok: true, engine: 'browse', outputs: [], evals: {}, stdout: '' };
})).rejects.toThrow('render leaked staging directory');
} finally {
if (leaked) fs.rmSync(leaked, { recursive: true, force: true });
}
});
test('happy path: newtab → goto <nonce URL> → per-step CLI calls → closetab; artifacts copied, evals inline, work dir and server released', async () => {
const before = browseWorkDirs();
const b = fake();
const r = await renderWithBrowse({
const r = await renderCheckingCleanup({
file: doc,
steps: [
{ kind: 'pdf', out: path.join(outDir, 'doc.pdf'), options: { paperWidth: 8.5, paperHeight: 11 } },
@@ -693,23 +753,19 @@ describe.skipIf(!HERMETIC)('aside-render: renderWithBrowse — daemon CLI contra
const payload = fs.readFileSync(`${log}.payloads`, 'utf8');
expect(payload).toContain('"width":"8.5in"');
expect(payload).toMatch(/"output":"\/tmp\/gstack-render-browse-[^"]+\/gstack-render-0\.pdf"/);
expect(browseWorkDirs()).toEqual(before); // /tmp staging dir removed
await expect(fetch(goto.slice('goto '.length, -` ${T}`.length))).rejects.toThrow(); // loopback server stopped
});
test('`newtab --json` without a tabId → the named error, no closetab, no staging dir left in /tmp', async () => {
const before = browseWorkDirs();
const r = await renderWithBrowse({ file: doc, steps: [{ kind: 'eval', expression: '1' }] }, fake({ newtab: `echo '{"ok":true}'` }));
const r = await renderCheckingCleanup({ file: doc, steps: [{ kind: 'eval', expression: '1' }] }, fake({ newtab: `echo '{"ok":true}'` }));
expect(r.ok).toBe(false);
expect(r.engine).toBe('browse');
expect(r.error).toBe('browse newtab --json returned no tabId');
expect(readLines(log)).toEqual(['newtab --json']);
expect(browseWorkDirs()).toEqual(before);
});
test('a failing goto → "browse goto failed: <first stderr line>", the tab is still closed, /tmp is left clean', async () => {
const before = browseWorkDirs();
const r = await renderWithBrowse({ file: doc, steps: [{ kind: 'pdf', out: path.join(outDir, 'x.pdf') }] }, fake({ goto: 'echo "net::ERR_CONNECTION_REFUSED at http://127.0.0.1" >&2; echo "second line" >&2; exit 1' }));
const r = await renderCheckingCleanup({ file: doc, steps: [{ kind: 'pdf', out: path.join(outDir, 'x.pdf') }] }, fake({ goto: 'echo "net::ERR_CONNECTION_REFUSED at http://127.0.0.1" >&2; echo "second line" >&2; exit 1' }));
expect(r.ok).toBe(false);
expect(r.error!.startsWith('browse goto failed:')).toBe(true);
expect(r.error).toContain('net::ERR_CONNECTION_REFUSED');
@@ -719,7 +775,6 @@ describe.skipIf(!HERMETIC)('aside-render: renderWithBrowse — daemon CLI contra
expect(lines.some((l) => l.startsWith('goto '))).toBe(true);
expect(lines.at(-1)).toBe('closetab 7');
expect(lines.some((l) => l.startsWith('pdf '))).toBe(false);
expect(browseWorkDirs()).toEqual(before);
expect(fs.existsSync(path.join(outDir, 'x.pdf'))).toBe(false);
});
@@ -803,19 +858,17 @@ describe.skipIf(!HERMETIC)('aside-render: renderWithBrowse — daemon CLI contra
// runProc is not exported: its timeout + kill path is observed through a hanging fake.
test('a CLI call that hangs past spec.timeoutMs is killed and reported as timed out — even when a grandchild keeps the pipes open', async () => {
const before = browseWorkDirs();
// `sleep` is a CHILD of the sh fake, so SIGTERM kills sh while sleep still holds stdout/stderr:
// the read must give up on its own (timeout + 10s) rather than wait for EOF. 14s (not 30s) so no orphan outlives this file.
const b = fake({ newtab: 'sleep 14' });
const t0 = Date.now();
const r = await renderWithBrowse({ file: doc, steps: [{ kind: 'eval', expression: '1' }], timeoutMs: 1_500 }, b);
const r = await renderCheckingCleanup({ file: doc, steps: [{ kind: 'eval', expression: '1' }], timeoutMs: 1_500 }, b);
const elapsed = Date.now() - t0;
expect(r.ok).toBe(false);
expect(r.error!.startsWith('browse newtab failed:')).toBe(true);
expect(r.error).toContain('timed out');
expect(elapsed).toBeLessThan(25_000);
expect(readLines(log)).toEqual(['newtab --json']); // no tab → nothing to close
expect(browseWorkDirs()).toEqual(before);
}, 40_000);
test('a hanging CLI that honours SIGTERM is reaped promptly at the budget', async () => {
+80
View File
@@ -0,0 +1,80 @@
import { expect, test } from 'bun:test';
import { findNativeAutoDecision } from './helpers/native-auto-decide';
import captured from './fixtures/auto-decide-saved-ai.json';
const clone=()=>structuredClone(captured) as any;
const decision=(f=clone())=>findNativeAutoDecision(f.transcript,f.tools,f.options);
const message=(f:any)=>f.transcript.assistantMessages.find((m:any)=>m.text.includes('Auto-decided'));
test('actual saved mode preference annotation is a completed native auto-decision',()=>{
const f=clone(), result=decision(f);
expect(result).not.toBeNull();
expect(result!.option).toBe('HOLD SCOPE');
expect(message(f).text).toContain(result!.annotation);
expect(f.transcript.calls).toEqual([]);
});
test('saved preference is bound to this invoked skill and an agreeing current mode',()=>{
for(const change of [
(s:string)=>s.replace('`plan-ceo-review-mode`','`plan-design-review-mode`'),
(s:string)=>s.replace('`plan-ceo-review-mode`','`plan-ceo-review-routing`'),
(s:string)=>s.replace('**Review mode: HOLD SCOPE.**','**Review mode: SCOPE EXPANSION.**'),
(s:string)=>s.replace('**Review mode: HOLD SCOPE.**\n\n',''),
(s:string)=>s.replace('"Select review mode"','"Select report folder"'),
(s:string)=>s.replace('(your saved preference on','(a proposed preference on'),
(s:string)=>s.replace('Change with /plan-tune.',''),
(s:string)=>s.replace('Auto-decided','I will auto-decide'),
]) {const f=clone();message(f).text=change(message(f).text);expect(decision(f)).toBeNull();}
});
test('prefixed examples, quotations and hypothetical notices do not assert a current choice',()=>{
for(const change of [
(s:string)=>'Example:\n\n'+s,
(s:string)=>'```text\n'+s+'\n```',
(s:string)=>s.split('\n').map(l=>'> '+l).join('\n'),
(s:string)=>s.replace('Heads-up from the preamble: unshipped work on this branch','Heads-up from the preamble: a hypothetical example'),
(s:string)=>s.replace('Auto-decided "Select',' Auto-decided "Select'),
(s:string)=>s.replace('Auto-decided "Select','If approved, Auto-decided "Select'),
]) {const f=clone();message(f).text=change(message(f).text);expect(decision(f)).toBeNull();}
});
test('failed loads, foreign sessions, actual questions and later withdrawals retain precedence',()=>{
for(const mutate of [
(f:any)=>{f.options.sessionId='foreign';},
(f:any)=>{const use=f.tools.find((t:any)=>t.kind==='use'&&t.name==='Skill');f.tools.find((t:any)=>t.kind==='result'&&t.toolUseId===use.toolUseId).isError=true;},
(f:any)=>{f.transcript.calls.push({sessionId:f.options.sessionId,toolUseId:'actual-question'});},
(f:any)=>{f.options.now=Date.parse(message(f).timestamp)-1;},
(f:any)=>{message(f).text+='\n\nCorrection: I withdraw this decision.';},
(f:any)=>{message(f).text+='\n\n**Review mode: SCOPE EXPANSION.**';},
]) {const f=clone();mutate(f);expect(decision(f)).toBeNull();}
});
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
test('new evidence inputs retain every native observation caller',()=>{
const owners=['plan-ceo-review-plan-mode','plan-eng-review-plan-mode','plan-design-review-plan-mode','plan-devex-review-plan-mode','plan-mode-no-op','auto-decide-preserved','conductor-prose'];
for(const file of ['test/auto-decide-saved-ai.test.ts','test/fixtures/auto-decide-saved-ai.json','test/fixtures/auto-decide-retry-ai.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(file)).map(([name])=>name)).toEqual(owners);
});
import retry from './fixtures/auto-decide-retry-ai.json';
test('actual retry mode-decision heading retains its own annotation, excluding prior foreign text',()=>{
const f:any=structuredClone(retry), actual=decision(f);
expect(actual).not.toBeNull();expect(actual!.sessionId).toBe(f.options.sessionId);
expect(actual!.option).toBe('HOLD SCOPE');
expect(actual!.annotation).toContain('(your preference)');
const own=f.transcript.assistantMessages.filter((m:any)=>m.sessionId===f.options.sessionId);
f.transcript.assistantMessages=f.transcript.assistantMessages.filter((m:any)=>m.sessionId!==f.options.sessionId);
expect(decision(f)).toBeNull();expect(own.length).toBeGreaterThan(0);
});
test('retry heading cannot supply a hypothetical, different decision, or withdrawn selection',()=>{
for(const change of [
(s:string)=>'Example:\n\n'+s,
(s:string)=>s.replace('Review mode for the deterministic','Review mode for the hypothetical'),
(s:string)=>s.replace('D1 — Review mode','D1 — Report destination'),
(s:string)=>s.replace('Auto-decided "Review mode:', 'Auto-decided "Report destination:'),
(s:string)=>s.replace('→ **HOLD SCOPE**','→ **Save a file**'),
(s:string)=>s+'\n\nCorrection: I withdraw this selection.',
(s:string)=>s+'\n\n**Review mode: SCOPE EXPANSION.**',
(s:string)=>s.replace('Heads-up from gstack: there is unshipped work on this branch','Heads-up from gstack: here is an example'),
]) {const f:any=structuredClone(retry),m=f.transcript.assistantMessages.find((m:any)=>m.sessionId===f.options.sessionId&&m.text.includes('Auto-decided'));m.text=change(m.text);expect(decision(f)).toBeNull();}
});
+214
View File
@@ -0,0 +1,214 @@
import { afterEach, describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import fixture from './fixtures/autoplan-artifact-permission-ad-v3.json';
import { autoplanArtifactPermissionInput } from './helpers/autoplan-artifact-permission';
import { isPermissionDialogVisible } from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, { recursive: true, force: true }); });
function replay(relative?: string) {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-artifact-permission-')); roots.push(root);
const cwd = path.join(root, path.basename(fixture.cwd)); fs.mkdirSync(cwd);
const ownedStateRoot = path.join(root, 'home', '.gstack');
const original = fixture.events.at(-1)!.input!.file_path;
const file = path.join(ownedStateRoot, 'projects', path.basename(cwd), relative ?? path.relative(
path.join(fixture.stateRoot, 'projects', path.basename(fixture.cwd)), original));
fs.mkdirSync(path.dirname(file), { recursive: true });
const publicTools = structuredClone(fixture.events) as NativePublicToolEvent[];
for (const event of publicTools) if (event.input?.file_path) event.input.file_path = file;
const lastWrite = publicTools.filter(event => event.name === 'Write').at(-1)!;
fs.writeFileSync(file, lastWrite.input!.content as string);
const context = { cwd, ownedStateRoot, commandStartedAt: fixture.commandStartedAt,
now: Date.parse('2026-09-09T20:36:27.729Z'), transcriptStatus: 'ready', publicTools };
const viewport = fixture.viewport.replaceAll(path.basename(original), path.basename(file));
return { root, file, context, viewport };
}
const pick = (r: ReturnType<typeof replay>, seen = new Set<string>()) =>
autoplanArtifactPermissionInput(r.viewport, r.context, seen);
describe('owned Autoplan artifact edit permission', () => {
test('captured cropped pane needs its pending identity; shared generic recognition stays unchanged', () => {
const r = replay();
expect(isPermissionDialogVisible(fixture.viewport)).toBe(false);
expect(pick(r)).toEqual({ input: '1\r', signature: `${fixture.sessionId}:${fixture.events.at(-1)!.toolUseId}`, file: r.file });
expect(pick(r, new Set([pick(r)!.signature]))).toBeNull();
});
test('a later same-file Edit has a new one-time epoch even when the footer is identical', () => {
const r = replay(); const first = pick(r)!; const edit = r.context.publicTools.at(-1)!;
fs.writeFileSync(r.file, fs.readFileSync(r.file, 'utf8').replace(edit.input!.old_string as string, edit.input!.new_string as string));
r.context.publicTools.push({ sessionId: fixture.sessionId, toolUseId: edit.toolUseId, kind: 'result',
timestamp: '2026-09-09T20:28:00.000Z', isError: false });
r.context.publicTools.push({ ...structuredClone(edit), toolUseId: 'next-owned-edit', timestamp: '2026-09-09T20:28:01.000Z' });
expect(pick(r, new Set([first.signature]))?.signature).toBe(`${fixture.sessionId}:next-owned-edit`);
});
test('a queued non-file tool cannot replace or grant the unique current Edit permission', () => {
const r = replay();
r.context.publicTools.push({ sessionId: fixture.sessionId, toolUseId: 'queued-bash', kind: 'use',
timestamp: '2026-09-09T20:27:41.541Z', name: 'Bash', input: { command: 'echo unrelated queued work' } });
expect(pick(r)?.input).toBe('1\r');
r.context.publicTools.push({ ...r.context.publicTools.at(-1)!, toolUseId: 'concurrent-write', name: 'Write',
input: { file_path: r.file, content: 'other mutation' } });
expect(pick(r)).toBeNull();
});
test('the two source-declared Eng test-plan layouts have the same bounded edit path', () => {
for (const file of ['test-main-eng-review-test-plan-20260909-203000.md', 'test-main-test-plan-20260909-203000.md'])
expect(pick(replay(file))?.input).toBe('1\r');
});
test('requires the exact owned project and known artifact filename; no broad state/home approval', () => {
for (const file of ['../sibling/ceo-plans/2026-09-09-user-dashboard.md', 'config.yaml', 'reviews.jsonl',
'main-autoplan-restore-20260909-200700.md', 'ceo-plans/archive/2026-09-09-user-dashboard.md',
'designs/screen-20260909/mockup.md', 'dx-plans/2026-09-09-plan.md', 'arbitrary.md'])
expect(pick(replay(file)), file).toBeNull();
const r = replay();
r.context.ownedStateRoot = undefined as any; expect(pick(r)).toBeNull();
r.context.ownedStateRoot = path.join(r.root, 'caller-GSTACK_HOME'); expect(pick(r)).toBeNull();
r.context.ownedStateRoot = path.join(r.root, 'home', '.gstack');
r.context.cwd = path.join(r.root, 'sibling'); expect(pick(r)).toBeNull();
});
test('regular current file and exact requested old/new text are mandatory', () => {
const r = replay(); const before = fs.readFileSync(r.file);
fs.writeFileSync(r.file, 'unrelated current content'); expect(pick(r)).toBeNull();
fs.writeFileSync(r.file, before);
r.context.publicTools.at(-1)!.input!.new_string = 'unrelated replacement'; expect(pick(r)).toBeNull();
fs.unlinkSync(r.file); fs.mkdirSync(r.file); expect(pick(r)).toBeNull();
});
test.skipIf(process.platform === 'win32')('rejects symlink escape and symlink aliases within the owned tree', () => {
const r = replay(); const other = path.join(r.root, 'external.md');
fs.renameSync(r.file, other); fs.symlinkSync(other, r.file); expect(pick(r)).toBeNull();
fs.unlinkSync(r.file); fs.renameSync(other, r.file);
const directory = path.dirname(r.file); const alias = directory + '-actual';
fs.renameSync(directory, alias); fs.symlinkSync(alias, directory); expect(pick(r)).toBeNull();
});
test.skipIf(process.platform === 'win32')('trusted temp-parent aliases preserve ownership without permitting a symlink state root', () => {
const r = replay(); const alias = path.join(r.root, 'temp-parent-alias');
fs.symlinkSync(path.join(r.root, 'home'), alias);
const target = path.join(alias, '.gstack', path.relative(r.context.ownedStateRoot, r.file));
r.context.ownedStateRoot = path.join(alias, '.gstack');
for (const event of r.context.publicTools) if (event.input?.file_path) event.input.file_path = target;
r.file = target;
expect(pick(r)?.input).toBe('1\r'); // e.g. macOS /var -> /private/var, above owned root
const stateAlias = path.join(r.root, 'state-alias');
fs.symlinkSync(r.context.ownedStateRoot, stateAlias);
const other = path.join(stateAlias, path.relative(r.context.ownedStateRoot, r.file));
r.context.ownedStateRoot = stateAlias;
for (const event of r.context.publicTools) if (event.input?.file_path) event.input.file_path = other;
r.file = other;
expect(pick(r)).toBeNull();
});
test('missing, stale, future, foreign, completed, failed, duplicate and concurrent identities stay closed', () => {
const mutations: Array<(r: ReturnType<typeof replay>) => void> = [
r => { r.context.transcriptStatus = 'error'; },
r => { r.context.publicTools = []; },
r => { r.context.commandStartedAt = r.context.now + 1; },
r => { r.context.commandStartedAt = Date.parse(r.context.publicTools.at(-1)!.timestamp) + 1; },
r => { r.context.publicTools.at(-1)!.timestamp = '2026-09-10T00:00:00.000Z'; },
r => { r.context.publicTools.at(-1)!.timestamp = 'invalid'; },
r => { r.context.publicTools.at(-1)!.sessionId = 'foreign'; },
r => { r.context.publicTools.at(-1)!.sessionId = ''; },
r => { r.context.publicTools.at(-1)!.toolUseId = ''; },
r => { r.context.publicTools.at(-1)!.name = 'Write'; },
r => { r.context.publicTools.at(-1)!.input!.replace_all = true; },
r => { r.context.publicTools.push({ ...r.context.publicTools.at(-1)!, kind: 'result', isError: false }); },
r => { r.context.publicTools.push({ ...r.context.publicTools.at(-1)!, kind: 'result', isError: true }); },
r => { r.context.publicTools.push(structuredClone(r.context.publicTools.at(-1)!)); },
r => { r.context.publicTools.splice(-1, 0, { ...structuredClone(r.context.publicTools.at(-1)!), toolUseId: 'other-pending-edit' }); },
r => { for (const event of r.context.publicTools) if (event.kind === 'result') event.isError = true; },
r => { for (const event of r.context.publicTools.slice(0, -1)) if (event.input) event.input.file_path = r.file + '-sibling'; },
r => { r.context.publicTools.reverse(); },
];
for (const mutate of mutations) { const r = replay(); mutate(r); expect(pick(r), mutate.toString()).toBeNull(); }
});
test('quotes, examples, unrelated diffs, malformed menus, extra options and broad selection are rejected', () => {
const mutations = [
(s: string) => 'Example:\n' + s, (s: string) => '```\n' + s + '\n```',
(s: string) => s.split('\n').map(line => '> ' + line).join('\n'),
(s: string) => s.replace('Success target made numeric', 'Unrelated line copied from another plan'),
(s: string) => s.replace('2026-09-09-user-dashboard.md?', 'sibling.md?'),
(s: string) => s.replace(' 1. Yes', ' 1. Yes').replace(' 2. Yes', '2. Yes'),
(s: string) => s.replace(' 1. Yes', ' 1. Yes, always allow'),
(s: string) => s.replace(' 3. No', ' 3. No\n 4. Change permission mode'),
(s: string) => s.replace('Esc to cancel · Tab to amend', 'Enter to select'),
(s: string) => s + '\nPlease choose the quoted example above.',
(s: string) => s.slice(s.indexOf(' Do you want')), // no bound diff
];
for (const mutate of mutations) { const r = replay(); r.viewport = mutate(r.viewport); expect(pick(r), mutate.toString()).toBeNull(); }
});
for (const deletion of [false, true]) test(`native ${deletion ? 'deletion' : 'replacement'} diff rows remain bound to the requested old/new text`, () => {
const r = replay(); const before = 'Old first\nOld second\nContext\n';
fs.writeFileSync(r.file, before);
r.context.publicTools.filter(event => event.name === 'Write').at(-1)!.input!.content = before;
const edit = r.context.publicTools.at(-1)!;
edit.input!.old_string = 'Old first\nOld second';
edit.input!.new_string = deletion ? '' : 'New first\nNew second';
const menu = r.viewport.slice(r.viewport.indexOf(' Do you want'));
// Existing native fixtures include 102-,103-,102+,103+ replacements,
// and deleted-only rows. These small controls are projected, not live panes.
r.viewport = ' 1 -Old first\n 2 -Old second\n' +
(deletion ? '' : ' 1 +New first\n 2 +New second\n') +
' 3 Context\n' + '╌'.repeat(20) + '\n' + menu;
expect(pick(r)?.input).toBe('1\r');
r.viewport = r.viewport.replace(' 2 -Old second', ' 2 -Context');
expect(pick(r)).toBeNull(); // Existing context is not part of the requested deletion.
});
// AZ's public line 116 wraps at column five, not the old fixed column four.
// These small panes exercise the same renderer rule without a transcript corpus.
for (const [line, numbered, continuation] of [
[7, ' 7 ', ' '], [17, ' 17 ', ' '],
[116, ' 116 ', ' '], [1024, ' 1024 ', ' '],
] as const) test(`wrapped line ${line} binds its own marker column and exact requested bytes`, () => {
const r = replay(), old = 'Old first portion kept together', replacement = 'New first portion kept together';
const before = Array.from({ length: line - 1 }, (_, n) => `Context ${n}`).concat(old, 'Context tail').join('\n');
fs.writeFileSync(r.file, before);
r.context.publicTools.filter(event => event.name === 'Write').at(-1)!.input!.content = before;
const edit = r.context.publicTools.at(-1)!;
edit.input!.old_string = old; edit.input!.new_string = replacement;
const menu = r.viewport.slice(r.viewport.indexOf(' Do you want'));
const rows = `${numbered}-Old first portion\n${continuation}- kept together\n` +
`${numbered}+New first portion\n${continuation}+ kept together\n`;
const pane = rows + '╌'.repeat(20) + '\n' + menu;
r.viewport = pane;
expect(pick(r)).toEqual({ input: '1\r', signature: `${edit.sessionId}:${edit.toolUseId}`, file: r.file });
expect(pick(r, new Set([pick(r)!.signature]))).toBeNull();
for (const invalid of [
pane.replaceAll(`\n${continuation}`, `\n${continuation.slice(1)}`), // left-shifted continuation
pane.replaceAll(`\n${continuation}`, `\n ${continuation}`), // right-shifted continuation
pane.replace(`${continuation}- kept`, `${continuation}+ kept`), // different kind
pane.replace(`${numbered}+New`, ` ${numbered}+New`), // mixed complete-row columns
`${continuation}- kept together\n` + pane, // no owning numbered row
pane.replace('New first portion', 'Foreign replacement'),
pane.replace(numbered, ' 0 '),
pane.replace(numbered, ' 01 '),
pane.replace(numbered, ' 9007199254740992 '),
]) { r.viewport = invalid; expect(pick(r), invalid).toBeNull(); }
});
test('an earlier unresolved mutation cannot make the latest completed Edit current', () => {
const r = replay(); const events = r.context.publicTools; const edit = events.at(-1)!;
events.splice(-1, 0, { ...structuredClone(edit), toolUseId: 'earlier-unresolved-edit',
input: { ...edit.input, file_path: r.file + '-other' } });
events.push({ sessionId: edit.sessionId, toolUseId: edit.toolUseId, kind: 'result',
timestamp: '2026-09-09T20:28:00.000Z', isError: false });
expect(pick(r)).toBeNull();
});
test('new helper, fixture, and regression select only the existing Autoplan paid case', () => {
for (const file of ['test/helpers/autoplan-artifact-permission.ts', 'test/autoplan-artifact-permission.test.ts',
'test/fixtures/autoplan-artifact-permission-ad-v3.json'])
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
});
});
+299
View File
@@ -0,0 +1,299 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { spawnSync } from 'node:child_process';
import { createAutoplanArtifactRecorder, recordAutoplanArtifact, readPendingAutoplanArtifact,
autoplanArtifactRecorderStatus, autoplanArtifactApprovalBoundary } from './helpers/autoplan-artifact-recorder';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
// Synthetic hook envelopes and owned temp paths. The live pending hook envelope
// was unpublished; these controls do not reconstruct it or provide paid coverage.
function fixture(approveEdits=false) {
const root = fs.mkdtempSync(path.join(os.tmpdir(), "artifact-' quote $()-"));
const cwd = path.join(root, 'repo with spaces'), config = path.join(root, 'config');
const stateRoot = path.join(root, 'home', '.gstack');
const project = path.join(config, 'projects', 'fixture');
const plans = path.join(stateRoot, 'projects', path.basename(cwd), 'ceo-plans');
fs.mkdirSync(cwd, {recursive:true}); fs.mkdirSync(project, {recursive:true}); fs.mkdirSync(plans, {recursive:true});
const session = 'synthetic-parent-session', transcript = path.join(project, session+'.jsonl');
fs.writeFileSync(transcript, '');
const artifact = path.join(plans, '2026-09-09-reviewed-plan.md');
fs.writeFileSync(artifact, 'Prior retained behavior.\n');
const startedAt = Date.now()-10;
const recorder = createAutoplanArtifactRecorder(cwd,config,stateRoot,approveEdits);
const event = (kind='PreToolUse',toolUseId='toolu_pending') => ({hook_event_name:kind,tool_name:'Edit',
session_id:session,tool_use_id:toolUseId,cwd,transcript_path:transcript,
tool_input:{file_path:artifact,old_string:'BODY_OLD_SENTINEL',new_string:'BODY_NEW_SENTINEL'}});
const write = (input:unknown) => recordAutoplanArtifact(typeof input==='string' ? input : JSON.stringify(input),recorder.file,cwd,config,stateRoot);
const history:NativePublicToolEvent[] = [
{kind:'use',sessionId:session,toolUseId:'toolu_prior',timestamp:new Date(startedAt).toISOString(),name:'Write',input:{file_path:artifact}},
{kind:'result',sessionId:session,toolUseId:'toolu_prior',timestamp:new Date(startedAt+1).toISOString(),isError:false},
];
const read = (tools=history,start=startedAt,now=Date.now()) => readPendingAutoplanArtifact(recorder.file,cwd,config,stateRoot,start,tools,now);
const status = () => autoplanArtifactRecorderStatus(recorder.file,cwd,config,stateRoot);
return {root,cwd,config,stateRoot,project,session,transcript,artifact,startedAt,recorder,event,write,history,read,status,
dispose:()=>{recorder.dispose();fs.rmSync(root,{recursive:true,force:true})}};
}
describe('owned Autoplan pending artifact metadata recorder',()=>{
test('pending authority is metadata only, with no body, result, success or phase evidence',()=>{
const f=fixture();try{
expect(f.status()).toEqual({status:'idle'});expect(f.read()).toBeUndefined();
const e=f.event() as ReturnType<typeof f.event> & {tool_response?:unknown};e.tool_response={content:'RESULT_SENTINEL'};f.write(e);
expect(f.read()).toEqual({source:'pre_tool_use',sessionId:f.session,toolUseId:'toolu_pending',tool:'Edit',file:f.artifact,timestamp:expect.any(String)});
const raw=fs.readFileSync(f.recorder.file,'utf8');
for(const secret of ['BODY_OLD_SENTINEL','BODY_NEW_SENTINEL','RESULT_SENTINEL','old_string','new_string','tool_response'])expect(raw).not.toContain(secret);
expect(Object.keys(JSON.parse(raw).pending).sort()).toEqual(['file','sessionId','source','timestamp','tool','toolUseId','transcriptPath'].sort());
expect(fs.statSync(f.recorder.file).mode&0o777).toBe(0o600);
}finally{f.dispose()}
});
test.each(['PostToolUse','PostToolUseFailure'])('%s closes and tombstones without creating successful native history',kind=>{
const f=fixture();try{
f.write(f.event());expect(f.read()).toBeDefined();f.write(f.event(kind));expect(f.status()).toEqual({status:'idle'});expect(f.read()).toBeUndefined();
f.write(f.event());expect(f.read()).toBeUndefined();f.write(f.event('PreToolUse','toolu_next'));expect(f.read()?.toolUseId).toBe('toolu_next');
expect(f.history).toHaveLength(2);
}finally{f.dispose()}
});
test('same pending replay never refreshes time; earlier completion cannot reopen',()=>{
const f=fixture();try{
f.write(f.event('PostToolUse','toolu_older'));
f.write(f.event());const before=fs.readFileSync(f.recorder.file,'utf8');f.write(f.event());expect(fs.readFileSync(f.recorder.file,'utf8')).toBe(before);
f.write(f.event('PostToolUse','toolu_older'));expect(f.read()?.toolUseId).toBe('toolu_pending');
f.write(f.event('PreToolUse','toolu_older'));expect(f.read()?.toolUseId).toBe('toolu_pending');
f.write(f.event('PostToolUse'));f.write(f.event('PreToolUse','toolu_older'));expect(f.read()).toBeUndefined();
}finally{f.dispose()}
});
test.each(['cwd','session','subagent'])('foreign %s cannot clear, replace or poison an owned parent',kind=>{
const f=fixture();try{
f.write(f.event());const before=fs.readFileSync(f.recorder.file,'utf8');
for(const hook of ['PreToolUse','PostToolUse','PostToolUseFailure']){
const e:Record<string,unknown>=f.event(hook,'toolu_foreign');
if(kind==='cwd')e.cwd=path.join(f.root,'foreign');
if(kind==='session'){e.session_id='foreign-session';e.transcript_path=path.join(f.project,'foreign-session.jsonl');fs.writeFileSync(e.transcript_path as string,'')}
if(kind==='subagent')e.agent_id='child-agent';
f.write(e);expect(f.status().status).toBe('pending');expect(fs.readFileSync(f.recorder.file,'utf8')).toBe(before);
}
}finally{f.dispose()}
});
test('only allowlisted Edit can become pending; foreign concurrent mutation invalidates ambiguity',()=>{
for(const kind of ['Write','unowned','different-id','different-path']){
const f=fixture();try{
const e=f.event('PreToolUse',kind==='different-id'?'toolu_other':'toolu_pending');
if(kind==='Write')e.tool_name='Write';
if(kind==='unowned')e.tool_input.file_path=path.join(f.stateRoot,'config');
if(kind==='different-path'){e.tool_input.file_path=path.join(path.dirname(f.artifact),'2026-09-09-other.md');fs.writeFileSync(e.tool_input.file_path,'old')}
if(kind==='Write'||kind==='unowned'){f.write(e);expect(f.status().status).toBe('idle')}
f.write(f.event());f.write(e);expect(f.status().status).toBe('invalid');expect(f.read()).toBeUndefined();
f.write(f.event('PostToolUse'));f.write(f.event('PreToolUse','toolu_later'));expect(f.status().status).toBe('invalid');
}finally{f.dispose()}
}
});
test.each(['same-file','foreign-path'])('unseen mutation completion on %s invalidates a pending request',target=>{
const f=fixture();try{
f.write(f.event());const e=f.event('PostToolUse','toolu_unseen');
if(target==='foreign-path')e.tool_input.file_path=path.join(f.stateRoot,'config');
f.write(e);expect(f.status().status).toBe('invalid');expect(f.read()).toBeUndefined();
}finally{f.dispose()}
});
test('malformed owned requests fail closed without persisting their contents',()=>{
for(const kind of ['bad-json','missing-cwd','outside-transcript','wrong-tool','empty-old','same-body','replace-all','oversized']){
const f=fixture();try{
const e:Record<string,any>=f.event();
if(kind==='missing-cwd')delete e.cwd;
if(kind==='outside-transcript'){e.transcript_path=path.join(f.root,f.session+'.jsonl');fs.writeFileSync(e.transcript_path,'')}
if(kind==='wrong-tool')e.tool_name='Read';
if(kind==='empty-old')e.tool_input.old_string='';
if(kind==='same-body')e.tool_input.new_string=e.tool_input.old_string;
if(kind==='replace-all')e.tool_input.replace_all=true;
if(kind==='oversized')e.tool_input.new_string='x'.repeat(4*1024*1024);
f.write(kind==='bad-json'?'{not-json':e);expect(f.status().status).toBe('invalid');expect(f.read()).toBeUndefined();
expect(fs.readFileSync(f.recorder.file,'utf8')).not.toContain('BODY_NEW_SENTINEL');
}finally{f.dispose()}
}
});
test('read binds actual session, interval and published identity precedence',()=>{
const f=fixture();try{
f.write(f.event());expect(f.read()).toBeDefined();
expect(f.read([],f.startedAt)).toBeUndefined();
expect(f.read(f.history.map(e=>({...e,sessionId:'other'})))).toBeUndefined();
expect(f.read([...f.history,{...f.history[0]!,sessionId:'other'}])).toBeUndefined();
expect(f.read(f.history,Date.now()+1000)).toBeUndefined();expect(f.read(f.history,f.startedAt,f.startedAt)).toBeUndefined();
for(const kind of ['use','result'] as const)expect(f.read([...f.history,{kind,sessionId:f.session,toolUseId:'toolu_pending',timestamp:new Date().toISOString()}])).toBeUndefined();
expect(readPendingAutoplanArtifact(undefined,f.cwd,f.config,f.stateRoot,f.startedAt,f.history)).toBeUndefined();
expect(readPendingAutoplanArtifact(f.recorder.file,f.cwd,null,f.stateRoot,f.startedAt,f.history)).toBeUndefined();
expect(readPendingAutoplanArtifact(f.recorder.file,f.cwd,f.config,path.join(f.root,'other'),f.startedAt,f.history)).toBeUndefined();
}finally{f.dispose()}
});
test.each([NaN, Infinity])('a non-finite reader clock %s cannot supply pending authority',now=>{
const f=fixture();try{f.write(f.event());expect(f.read(f.history,f.startedAt,now)).toBeUndefined()}finally{f.dispose()}
});
test('busy, oversized, symlinked and exhausted state cannot supply authority',()=>{
for(const kind of ['busy','oversized','symlinked','exhausted']){
const f=fixture();try{
f.write(f.event());
if(kind==='busy')fs.writeFileSync(f.recorder.file+'.lock','');
if(kind==='oversized')fs.writeFileSync(f.recorder.file,' '.repeat(64*1024+1));
if(kind==='symlinked'){const backup=f.recorder.file+'.backup';fs.renameSync(f.recorder.file,backup);fs.symlinkSync(backup,f.recorder.file)}
if(kind==='exhausted'){f.write(f.event('PostToolUse'));const s=JSON.parse(fs.readFileSync(f.recorder.file,'utf8'));s.seenIds=Array.from({length:128},(_,i)=>'toolu_'+i);fs.writeFileSync(f.recorder.file,JSON.stringify(s));f.write(f.event('PreToolUse','toolu_overflow'))}
expect(f.read()).toBeUndefined();expect(f.status().status).toBe(kind==='busy'?'busy':'invalid');
}finally{f.dispose()}
}
});
test('all generated hook commands are bounded, correctly quoted and silent',()=>{
const f=fixture();try{
expect(Object.keys(f.recorder.hooks).sort()).toEqual(['PostToolUse','PostToolUseFailure','PreToolUse']);
for(const kind of ['PreToolUse','PostToolUse','PostToolUseFailure'] as const){
const entries=f.recorder.hooks[kind];expect(entries).toHaveLength(1);expect(entries[0]!.matcher).toBe('^(Write|Edit)$');
const hook=entries[0]!.hooks[0]!;expect(hook.timeout).toBe(5);
for(const input of [JSON.stringify(f.event(kind)),'{broken']){
const child=spawnSync('bash',['-c',hook.command],{cwd:f.cwd,input,encoding:'utf8',timeout:6000});
expect(child.error).toBeUndefined();expect(child.status).toBe(0);expect(child.stdout).toBe('');expect(child.stderr).toBe('');
}
}
}finally{f.dispose()}
});
test('a hook without input EOF exits silently within its internal timeout',async()=>{
const f=fixture(),hook=f.recorder.hooks.PreToolUse[0]!.hooks[0]!;
const child=Bun.spawn(['bash','-c','exec '+hook.command],{cwd:f.cwd,stdin:'pipe',stdout:'pipe',stderr:'pipe'});
let forced=false;const timer=setTimeout(()=>{forced=true;child.kill('SIGKILL')},6000);
try{
const [code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
expect(forced).toBe(false);expect(code).toBe(0);expect(out).toBe('');expect(err).toBe('');
expect(f.status()).toEqual({status:'invalid',reason:'stdin_timeout'});
}finally{clearTimeout(timer);child.stdin.end();if(child.exitCode===null){child.kill('SIGKILL');await child.exited}f.dispose()}
},7000);
test('recorder disposal removes all owned state and the new inputs select only Autoplan',()=>{
const f=fixture();f.write(f.event());f.dispose();expect(fs.existsSync(f.recorder.file)).toBe(false);
for(const file of ['test/helpers/autoplan-artifact-recorder.ts','test/autoplan-artifact-recorder.test.ts','test/autoplan-pending-artifact.test.ts','test/fixtures/autoplan-pending-artifact-ae.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
});
// Real hook subprocesses and the public JSONL reader; no synthetic phase success
// and no reconstruction of the unpublished paid request.
function approvalFixture(approve=true, activate=true) {
const f=fixture(approve), startedAt=Date.now();
if (approve && activate) f.recorder.startEditApproval!(startedAt);
const event=f.event();event.tool_input.old_string='Prior retained behavior.';event.tool_input.new_string='Reviewed retained behavior.';
const records:any[]=[
{message:{role:'assistant',id:'msg_prior',content:[{type:'tool_use',id:'toolu_prior',name:'Write',input:{file_path:f.artifact}}]}},
{message:{role:'user',content:[{type:'tool_result',tool_use_id:'toolu_prior',is_error:false,content:'Written'}]}},
].map(record=>({cwd:f.cwd,sessionId:f.session,isSidechain:false,timestamp:new Date(startedAt).toISOString(),...record}));
const publish=()=>fs.writeFileSync(f.transcript,records.map(record=>JSON.stringify(record)+'\n').join(''));
const current=()=>({cwd:f.cwd,sessionId:f.session,isSidechain:false,timestamp:new Date(startedAt).toISOString(),requestId:'req_current',
message:{role:'assistant',id:'msg_current',content:[{type:'tool_use',id:event.tool_use_id,name:'Edit',input:{...event.tool_input}}]}});
const invoke=(input:unknown=event)=>{
const child=spawnSync('bash',['-c',f.recorder.hooks.PreToolUse[0]!.hooks[0]!.command],
{cwd:f.cwd,input:JSON.stringify(input),encoding:'utf8',timeout:6000});
expect(child.error).toBeUndefined();expect(child.status).toBe(0);expect(child.stderr).toBe('');
return child.stdout ? JSON.parse(child.stdout) : undefined;
};
publish();return {...f,startedAt,event,records,publish,current,invoke};
}
describe('explicit native approval for owned Autoplan artifact Edits',()=>{
test.each(['unpublished','published','eng-artifact'])('%s approves once without changing bytes, JSONL or phase evidence',kind=>{
const f=approvalFixture();try{
if(kind==='eng-artifact'){
f.event.tool_input.file_path=path.join(path.dirname(path.dirname(f.artifact)),'branch-test-plan-20260911-055000.md');
fs.writeFileSync(f.event.tool_input.file_path,fs.readFileSync(f.artifact));
f.records[0].message.content[0].input.file_path=f.event.tool_input.file_path;
}
if(kind==='published')f.records.push(f.current());
f.publish();const journal=fs.readFileSync(f.transcript,'utf8'),before=fs.readFileSync(f.event.tool_input.file_path,'utf8');
expect(f.invoke()).toEqual({hookSpecificOutput:{hookEventName:'PreToolUse',permissionDecision:'allow'}});
const state=fs.readFileSync(f.recorder.file,'utf8');
expect(JSON.parse(state).pending.editDigest).toBeDefined();expect(JSON.parse(state).seenIds).toEqual([f.event.tool_use_id]);
for(const body of [f.event.tool_input.old_string,f.event.tool_input.new_string,'tool_response'])expect(state).not.toContain(body);
expect(f.invoke()).toBeUndefined();expect(fs.readFileSync(f.recorder.file,'utf8')).toBe(state);
expect(f.invoke({...f.event,hook_event_name:'PostToolUse',tool_response:'Not native history'})).toBeUndefined();
expect(f.status().status).toBe('idle');expect(f.invoke()).toBeUndefined();
expect(fs.readFileSync(f.event.tool_input.file_path,'utf8')).toBe(before);expect(fs.readFileSync(f.transcript,'utf8')).toBe(journal);
}finally{f.dispose()}
});
test.each(['passive','not-started'])('%s remains silent even for a valid owned Edit',kind=>{
const f=approvalFixture(kind!=='passive',false);try{expect(f.invoke()).toBeUndefined()}finally{f.dispose()}
});
test.each(['no-history','old-history','failed-write','unresolved-write','concurrent-edit','foreign-mutation','duplicate-use',
'completed-current','conflicting-current','foreign-session','subagent-history','agent-id-history','other-journal-history','future-history','wrong-cwd','subagent',
'wrong-session','foreign-transcript','missing-old','duplicate-old','future-mtime','replace-all','Write','Bash',
'snapshot','config','native-plan','foreign-file','symlink','locked'])('%s cannot grant a native decision',kind=>{
const f=approvalFixture();try{
const e=f.event as Record<string,any>;
if(kind==='no-history')f.records.length=0;
if(kind==='old-history')f.records.forEach(r=>r.timestamp=new Date(f.startedAt-1).toISOString());
if(kind==='failed-write')f.records[1].message.content[0].is_error=true;
if(kind==='unresolved-write')f.records.pop();
if(kind==='concurrent-edit'||kind==='foreign-mutation'){
const other=f.current();other.message.content[0].id='toolu_conflict';
if(kind==='foreign-mutation')other.message.content[0].input.file_path=path.join(f.stateRoot,'config');
f.records.push(other);
}
if(kind==='duplicate-use')f.records.splice(1,0,f.records[0]);
if(kind==='completed-current')f.records.push(f.current(),{...f.records[1],message:{role:'user',content:[{type:'tool_result',tool_use_id:e.tool_use_id,is_error:false}]}});
if(kind==='conflicting-current'){const current=f.current();current.message.content[0].input.new_string='Different request';f.records.push(current)}
if(kind==='foreign-session')f.records.forEach(r=>r.sessionId='another-parent');
if(kind==='subagent-history')f.records.forEach(r=>r.isSidechain=true);
if(kind==='agent-id-history')f.records.forEach(r=>r.agentId='child-agent');
if(kind==='other-journal-history'){const other=path.join(f.config,'projects','other',f.session+'.jsonl');fs.mkdirSync(path.dirname(other));fs.writeFileSync(other,fs.readFileSync(f.transcript));f.records.length=0;}
if(kind==='future-history')f.records.forEach(r=>r.timestamp=new Date(Date.now()+10000).toISOString());
if(kind==='wrong-cwd')e.cwd=path.join(f.root,'other');
if(kind==='subagent')e.agent_id='child';
if(kind==='wrong-session'){e.session_id='another-parent';e.transcript_path=path.join(f.project,e.session_id+'.jsonl');fs.writeFileSync(e.transcript_path,'')}
if(kind==='foreign-transcript'){e.transcript_path=path.join(f.root,f.session+'.jsonl');fs.writeFileSync(e.transcript_path,'')}
if(kind==='missing-old')e.tool_input.old_string='Absent text';
if(kind==='duplicate-old')fs.writeFileSync(f.artifact,'Prior retained behavior. Prior retained behavior.');
if(kind==='future-mtime')fs.utimesSync(f.artifact,new Date(),new Date(Date.now()+10000));
if(kind==='replace-all')e.tool_input.replace_all=true;
if(kind==='Write'||kind==='Bash')e.tool_name=kind;
if(['snapshot','config','native-plan','foreign-file'].includes(kind)){
e.tool_input.file_path=kind==='snapshot'?path.join(path.dirname(f.artifact),'snapshot.md'):
kind==='config'?path.join(f.stateRoot,'config'):kind==='native-plan'?path.join(f.config,'plans','owned-plan.md'):path.join(f.root,'foreign.md');
fs.mkdirSync(path.dirname(e.tool_input.file_path),{recursive:true});fs.writeFileSync(e.tool_input.file_path,'Prior retained behavior.');
f.records[0].message.content[0].input.file_path=e.tool_input.file_path;
}
if(kind==='symlink'){fs.renameSync(f.artifact,f.artifact+'.original');fs.symlinkSync(f.artifact+'.original',f.artifact)}
if(kind==='locked')fs.writeFileSync(f.recorder.file+'.lock','');
f.publish();expect(f.invoke()).toBeUndefined();
}finally{f.dispose()}
});
test('rejected owned Edit stops before UI fallback; initial Write leaves unrelated navigation available',()=>{
const f=approvalFixture();try{
expect(f.invoke({...f.event,tool_name:'Write'})).toBeUndefined();
expect(autoplanArtifactApprovalBoundary(f.status())).toBe('clear');
f.records[1].message.content[0].is_error=true;f.publish();
expect(f.invoke()).toBeUndefined();
expect(f.status()).toEqual({status:'invalid',reason:'approval_withheld'});
expect(autoplanArtifactApprovalBoundary(f.status())).toBe('failed');
// The failed route is sticky across repeated hooks and cannot send UI input.
expect(f.invoke()).toBeUndefined();expect(autoplanArtifactApprovalBoundary(f.status())).toBe('failed');
const retained=JSON.parse(fs.readFileSync(f.recorder.file,'utf8'));
expect(retained.pending.toolUseId).toBe(f.event.tool_use_id);
}finally{f.dispose()}
const g=approvalFixture();try{
expect(g.invoke()?.hookSpecificOutput.permissionDecision).toBe('allow');
expect(autoplanArtifactApprovalBoundary(g.status())).toBe('pending');
expect(g.invoke({...g.event,hook_event_name:'PostToolUse'})).toBeUndefined();
expect(autoplanArtifactApprovalBoundary(g.status())).toBe('clear');
}finally{g.dispose()}
});
test('changed replay cannot refresh approval; completion hooks cannot fabricate required history',()=>{
const f=approvalFixture();try{
expect(f.invoke()?.hookSpecificOutput.permissionDecision).toBe('allow');
fs.writeFileSync(f.artifact,'Changed bytes.');expect(f.invoke()).toBeUndefined();expect(f.status().status).toBe('invalid');
}finally{f.dispose()}
const g=approvalFixture();try{
g.records.length=0;g.publish();expect(g.invoke()).toBeUndefined();
expect(g.invoke({...g.event,hook_event_name:'PostToolUse',tool_response:{success:true}})).toBeUndefined();
g.event.tool_use_id='toolu_next';expect(g.invoke()).toBeUndefined();
}finally{g.dispose()}
});
test.each(['repeat','future','before-recorder'])('invalid %s activation fails closed',kind=>{
const f=approvalFixture(true,kind==='repeat');try{
const started=kind==='future'?Date.now()+10000:kind==='before-recorder'?f.startedAt-10000:f.startedAt;
expect(()=>f.recorder.startEditApproval!(started)).toThrow();expect(f.invoke()).toBeUndefined();
}finally{f.dispose()}
});
});
+144
View File
@@ -0,0 +1,144 @@
import { capturedPathRebaser } from './helpers/captured-paths';
import {expect,test} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
import fixture from './fixtures/autoplan-artifact-stall-as.json';
import * as permission from './helpers/autoplan-artifact-permission';
import {readPendingAutoplanArtifact,autoplanArtifactRecorderStatus} from './helpers/autoplan-artifact-recorder';
import {readPlanCountTranscript,type NativePublicToolEvent} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
test('captured path rebasing preserves JSON strings and emits canonical native file paths',()=>{
const destination=String.raw`C:\a\repo`,source={file:'/captured/plans/plan.md',content:'First\n/captured/notes\nLast'};
const rebase=capturedPathRebaser([['/captured',destination]]);
const display=destination.split(path.sep).join('/');
expect(rebase.json(source)).toEqual({file:path.normalize(display+'/plans/plan.md'),content:'First\n'+display+'/notes\nLast'});
expect(source.file).toBe('/captured/plans/plan.md');
});
test('captured path rebasing preserves malformed and foreign ownership inputs',()=>{
const destination=path.join(path.parse(process.cwd()).root,'replayed');
const rebase=capturedPathRebaser([['/captured',destination]]);
for(const suffix of ['../foreign.md','plans/../plan.md','plans//plan.md','plans/./plan.md']){
expect(rebase.json({file:'/captured/'+suffix}).file).toBe(destination+path.sep+suffix.split('/').join(path.sep));
}
expect(rebase.json({file:'../foreign.md'}).file).toBe('..'+path.sep+'foreign.md');
expect(rebase.json({file:'/foreign/plans/../plan.md'}).file).toBe(path.sep+'foreign'+path.sep+'plans'+path.sep+'..'+path.sep+'plan.md');
});
function replay() {
const root=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-ap-stall-'));
const runtimeBefore=path.dirname(path.dirname(fixture.stateRoot));
const runtime=path.join(root,path.basename(runtimeBefore)),cwd=path.join(root,path.basename(fixture.cwd));
const rebase=capturedPathRebaser([[runtimeBefore,runtime],[fixture.cwd,cwd]]);
const hook=rebase.json(fixture.hook),stateRoot=rebase.file(fixture.stateRoot),config=rebase.file(fixture.config);
const events=rebase.json(fixture.publicTools) as NativePublicToolEvent[];
const now=Date.parse(fixture.viewportCapturedAt),startedAt=Date.parse(fixture.commandStartedAt);
const file=hook.pending.file,nativePlan=events.filter(e=>e.kind==='use'&&e.name==='Edit').at(-1)!.input!.file_path as string;
for(const [target,content] of [[file,fixture.before],[nativePlan,fixture.nativePlanBefore]]) {
fs.mkdirSync(path.dirname(target),{recursive:true});fs.writeFileSync(target,content);
const at=new Date(Date.parse(hook.pending.timestamp)-1000);fs.utimesSync(target,at,at);
}
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(hook.pending.transcriptPath),{recursive:true});
const records=events.map(e=>({sessionId:e.sessionId,cwd,isSidechain:false,timestamp:e.timestamp,requestId:e.requestId,
message:{id:e.messageId,role:e.kind==='use'?'assistant':'user',content:e.kind==='use'?[{type:'tool_use',id:e.toolUseId,name:e.name,input:e.input}]:[{type:'tool_result',tool_use_id:e.toolUseId,content:e.content??'',is_error:e.isError}]}}));
fs.writeFileSync(hook.pending.transcriptPath,records.map(r=>JSON.stringify(r)).join('\n')+'\n');
const hookFile=path.join(root,'hook.json');fs.writeFileSync(hookFile,JSON.stringify(hook)+'\n');
const publicTools:NativePublicToolEvent[]=[];const transcript=readPlanCountTranscript(config,cwd,e=>publicTools.push(e));
const pending=readPendingAutoplanArtifact(hookFile,cwd,config,stateRoot,startedAt,publicTools,now,true);
const context={cwd,ownedStateRoot:stateRoot,ownedNativePlansRoot:path.join(config,'plans'),commandStartedAt:startedAt,
now,viewportCapturedAt:now,transcriptStatus:transcript.status,publicTools,pending};
const viewport=rebase.text(fixture.viewport);
const invoke=(screen=viewport,ctx=context,seen=new Set<string>())=>permission.publishedAutoplanArtifactPermissionInput(screen,ctx,seen);
return {root,hook,hookFile,config,file,nativePlan,context,viewport,invoke,dispose:()=>fs.rmSync(root,{recursive:true,force:true})};
}
type Replay=ReturnType<typeof replay>;
const current=(r:Replay)=>r.context.publicTools.find(e=>e.toolUseId===r.hook.pending.toolUseId&&e.kind==='use')!;
const queued=(r:Replay)=>r.context.publicTools.filter(e=>e.kind==='use'&&e.name==='Edit'&&Date.parse(e.timestamp)>Date.parse(r.hook.pending.timestamp));
function reject(cases:Array<[string,(r:Replay)=>void]>) {
for(const [name,change] of cases){const r=replay();try{change(r);expect(r.invoke(),name).toBeNull()}finally{r.dispose()}}
}
test('the retained pending CEO edit remains distinct from later published native-plan edits',()=>{
const r=replay();try{
expect(autoplanArtifactRecorderStatus(r.hookFile,r.context.cwd,r.config,r.context.ownedStateRoot)).toEqual({status:'pending'});
expect(r.context.pending?.toolUseId).toBe(fixture.hook.pending.toolUseId);
expect(queued(r)).toHaveLength(2);
expect(permission.autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(permission.pendingAutoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(r.invoke()).toEqual({input:'1\r',signature:fixture.hook.sessionId+':'+fixture.hook.pending.toolUseId,file:r.file});
expect(fixture.provenance.retrospectivePass).toBe(false);
expect(r.invoke(r.viewport,r.context,new Set([r.hook.sessionId+':'+r.hook.pending.toolUseId]))).toBeNull();
expect(r.invoke(r.viewport,r.context,new Set([permission.autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
}finally{r.dispose()}
});
test('a bare current panel and its bound redraw labels represent the same one-time permission',()=>{
const r=replay();try{
const title=r.viewport.indexOf('● Update('),panel=r.viewport.indexOf('────────────────');
expect(r.invoke(r.viewport.slice(title))?.input).toBe('1\r');
expect(r.invoke(r.viewport.slice(panel))?.input).toBe('1\r');
}finally{r.dispose()}
});
test('only unstarted same-batch publications to the launcher-owned native plans root may wait behind it',()=>{
reject([
['no launcher root',r=>{delete (r.context as any).ownedNativePlansRoot}],
['foreign launcher root',r=>{r.context.ownedNativePlansRoot=path.join(r.root,'foreign')}],
['foreign message',r=>{queued(r)[0]!.messageId='msg_other'}],
['foreign request',r=>{queued(r)[0]!.requestId='req_other'}],
['foreign session',r=>{queued(r)[0]!.sessionId='other'}],
['foreign target',r=>{queued(r)[0]!.input!.file_path=r.file+'.other'}],
['queued Write',r=>{queued(r)[0]!.name='Write'}],
['replace-all successor',r=>{queued(r)[0]!.input!.replace_all=true}],
['already started successor',r=>{r.context.pending!.hookSeenIds!.push(queued(r)[0]!.toolUseId)}],
['successor completion',r=>{const q=queued(r)[0]!;r.context.publicTools.push({kind:'result',sessionId:q.sessionId,toolUseId:q.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:false})}],
['successor failure',r=>{const q=queued(r)[0]!;r.context.publicTools.push({kind:'result',sessionId:q.sessionId,toolUseId:q.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:true})}],
['successor published after viewport',r=>{queued(r)[0]!.timestamp=new Date(r.context.viewportCapturedAt+1).toISOString()}],
['missing native plan',r=>{fs.unlinkSync(r.nativePlan)}],
['native plan changed after current hook',r=>{fs.utimesSync(r.nativePlan,new Date(r.context.now),new Date(r.context.now))}],
['symlink native plan',r=>{const other=path.join(r.root,'other.md');fs.renameSync(r.nativePlan,other);fs.symlinkSync(other,r.nativePlan)}],
['successful Read cannot replace native-plan mutation history',r=>{for(const e of r.context.publicTools)if(e.kind==='use'&&e.input?.file_path===r.nativePlan&&Date.parse(e.timestamp)<Date.parse(r.hook.pending.timestamp))e.name='Read'}],
['no successful native-plan history',r=>{const ids=new Set(r.context.publicTools.filter(e=>e.input?.file_path===r.nativePlan).map(e=>e.toolUseId));for(const e of r.context.publicTools)if(e.kind==='result'&&ids.has(e.toolUseId))e.isError=true}],
]);
});
test('the active hook, current digest, successful owned history and time remain mandatory',()=>{
reject([
['no current hook',r=>{r.context.pending=undefined}],['foreign hook',r=>{r.context.pending!.sessionId='other'}],
['wrong current ID',r=>{r.context.pending!.toolUseId=queued(r)[0]!.toolUseId}],
['no digest',r=>{delete r.context.pending!.editDigest}],
['changed digest',r=>{r.context.pending!.editDigest!.requestSHA256='0'.repeat(64)}],
['changed replacement',r=>{current(r).input!.new_string+=' changed'}],
['changed current file',r=>{fs.appendFileSync(r.file,'changed');fs.utimesSync(r.file,new Date(0),new Date(0))}],
['completed current',r=>{const q=current(r);r.context.publicTools.push({kind:'result',sessionId:q.sessionId,toolUseId:q.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:false})}],
['current file newer than hook',r=>{fs.utimesSync(r.file,new Date(r.context.now),new Date(r.context.now))}],
['pending after viewport',r=>{r.context.pending!.timestamp=new Date(r.context.now+1).toISOString()}],
['stale hook',r=>{r.context.pending!.timestamp=new Date(r.context.commandStartedAt-1).toISOString()}],
['unavailable transcript',r=>{r.context.transcriptStatus='missing'}],
]);
const r=replay();try{
fs.writeFileSync(r.hookFile+'.invalid','{"reason":"concurrent_pending"}');
expect(readPendingAutoplanArtifact(r.hookFile,r.context.cwd,r.config,r.context.ownedStateRoot,r.context.commandStartedAt,r.context.publicTools,r.context.now,true)).toBeUndefined();
}finally{r.dispose()}
});
test('completed output and redraw labels cannot hide a foreign, quoted or persistent-permission panel',()=>{
const changes:Array<[string,(s:string)=>string]>=[
['example prefix',s=>'Example:\n'+s],['quoted whole pane',s=>s.split('\n').map(r=>'> '+r).join('\n')],
['arbitrary output',s=>s.replace('"changed": true','"changed": false')],
['foreign completed command',s=>s.replace('with-skills/.clau','foreign/.clau')],
['missing one redraw',s=>s.replace('● Updated plan','')],['extra redraw',s=>s.replace('● Updated plan','● Updated plan\n● Updated plan')],
['arbitrary redraw prose',s=>s.replace('● Updated plan','● Example plan')],
['foreign current title',s=>s.replace('Update(~/.gstack/','Update(/foreign/')],
['foreign displayed project',s=>s.replace('…-207152-jk89F3/skill-home-bOPSw5/.gstack/projects/gstack-autoplan-chain-kVh2Sb','…projects/foreign')],
['different requested addition',s=>s.replace('## Reviewer Concerns','## An unrelated edit')],
['wrong menu file',s=>s.replace('user-dashboard.md?','other.md?')],
['persistent session approval',s=>s.replace(' 1. Yes',' 2. Yes')],['trailing prose',s=>s+'\nAnother prompt'],
];
for(const [name,edit] of changes){const r=replay();try{expect(r.invoke(edit(r.viewport)),name).toBeNull()}finally{r.dispose()}}
});
test('only Autoplan discovers the permission regression and its captured fixture',()=>{
for(const file of ['test/autoplan-artifact-stall-as.test.ts','test/fixtures/autoplan-artifact-stall-as.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
+81
View File
@@ -0,0 +1,81 @@
import { expect, test } from 'bun:test';
import { readFileSync, existsSync, readdirSync } from 'node:fs';
import { spawnSync } from 'node:child_process';
import { createNativeReviewState } from './helpers/plan-count-fixture';
import { getHermeticDirs } from './helpers/hermetic-env';
import { resolve } from 'node:path';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const root = resolve(import.meta.dir, '..');
const read = (file: string) => readFileSync(resolve(root, file), 'utf8');
const fixture = 'test/fixtures/plans/autoplan-dashboard.md';
test('the chain fixture retains the complete original UI/API scope', () => {
const original = read('test/fixtures/plans/ui-heavy-feature.md');
const complete = read(fixture);
expect(complete.startsWith(original + '\n')).toBe(true);
// This supplements dependency facts; it does not supply a completed review,
// prescribe its decisions, or pre-build the feature exercised by the chain.
expect(complete).not.toMatch(/Phase \d|GSTACK REVIEW REPORT|AUTO-DECIDE|all findings resolved/i);
expect(complete).toContain('there are no dashboard-specific tests yet');
expect(complete).toContain('not completed work');
});
test('the new fixture is isolated to the chain and its selection dependencies', () => {
expect(read('test/skill-e2e-autoplan-chain.test.ts')).toContain("'plans', 'autoplan-dashboard.md'");
expect(read('test/skill-e2e-plan-design-with-ui.test.ts')).toContain("'plans', 'ui-heavy-feature.md'");
for (const file of [fixture, 'test/autoplan-chain-fixture.test.ts']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
}
expect(selectTests(['test/fixtures/plans/ui-heavy-feature.md'], E2E_TOUCHFILES).selected)
.toEqual(['plan-design-with-ui-scope']);
});
test('native sequencing config reaches the real CLI reader without changing shared state', () => {
const shared = getHermeticDirs().gstackHome;
const before = readFileSync(resolve(shared, 'config.yaml'), 'utf8');
const first = createNativeReviewState();
const second = createNativeReviewState();
try {
expect(first.env.GSTACK_HOME).not.toBe(shared);
expect(first.env.GSTACK_HOME).not.toBe(second.env.GSTACK_HOME);
expect(first.env.GSTACK_STATE_ROOT).toBe(first.env.GSTACK_HOME);
const result = spawnSync('bash', [resolve(root, 'bin/gstack-config'), 'get', 'codex_reviews'], {
cwd: root, env: { ...process.env, ...first.env }, encoding: 'utf8', timeout: 5000,
});
expect(result.status, result.stderr).toBe(0);
expect(result.stdout.trim()).toBe('disabled');
for (const marker of readdirSync(shared).filter(name => name === '.activated' ||
/^\..*(?:-seen|-prompted|-shown)$/.test(name) || name.startsWith('.feature-prompted-'))) {
expect(readFileSync(resolve(first.env.GSTACK_HOME!, marker), 'utf8'))
.toBe(readFileSync(resolve(shared, marker), 'utf8'));
}
first.cleanup();
first.cleanup();
expect(existsSync(first.env.GSTACK_HOME!)).toBe(false);
expect(existsSync(second.env.GSTACK_HOME!)).toBe(true);
expect(readFileSync(resolve(shared, 'config.yaml'), 'utf8')).toBe(before);
} finally {
first.cleanup();
second.cleanup();
}
expect(existsSync(second.env.GSTACK_HOME!)).toBe(false);
});
test('the UI/API chain requires all four native phases and registers its config dependency', () => {
const source = read('test/skill-e2e-autoplan-chain.test.ts');
const plan = read(fixture);
expect(plan).toContain('## UI Scope');
expect(plan).toContain('New REST endpoint `GET /api/dashboard`');
expect(source).toContain('env: nativeState.env');
expect(source).toContain('if (!ceo || !design || !dx || !eng)');
expect(source).toContain('expect(ceo.ts).toBeLessThan(design.ts)');
expect(source).toContain('expect(design.ts).toBeLessThan(dx.ts)');
expect(source).toContain('expect(dx.ts).toBeLessThan(eng.ts)');
expect(source).toContain('nativeState?.cleanup()');
for (const file of ['test/helpers/plan-count-fixture.ts', 'test/plan-count-fixture.test.ts', 'bin/gstack-config']) {
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain(file);
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('autoplan-chain-pty');
}
});
+97
View File
@@ -0,0 +1,97 @@
import {test,expect,afterEach} from 'bun:test';import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
import fixture from './fixtures/autoplan-clipped-suffix-aq.json';
import {createAutoplanEditDigest,validAutoplanEditDigest,matchesAutoplanDigestRows} from './helpers/autoplan-artifact-digest';
import {createAutoplanArtifactRecorder,recordAutoplanArtifact,readPendingAutoplanArtifact,autoplanArtifactRecorderStatus} from './helpers/autoplan-artifact-recorder';
import {pendingAutoplanArtifactPermissionInput,autoplanArtifactMenuKey} from './helpers/autoplan-artifact-permission';
import {E2E_TOUCHFILES} from './helpers/touchfiles-data';
const cleanup:Array<()=>void>=[];afterEach(()=>{for(const f of cleanup.splice(0))f()});
function replay(before=fixture.before,removed=fixture.request.old_string,added=fixture.request.new_string){
const root=fs.mkdtempSync(path.join(os.tmpdir(),'ap-suffix-')),cwd=path.join(root,path.basename(fixture.cwd)),config=path.join(root,'config'),stateRoot=path.join(root,'home/.gstack');
const file=path.normalize(fixture.hook.pending.file.replace(fixture.stateRoot,stateRoot)),native=path.join(config,'projects/owned',fixture.hook.sessionId+'.jsonl');
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(file),{recursive:true});fs.mkdirSync(path.dirname(native),{recursive:true});fs.writeFileSync(native,'');fs.writeFileSync(file,before);fs.utimesSync(file,new Date(0),new Date(0));
const recorder=createAutoplanArtifactRecorder(cwd,config,stateRoot);cleanup.push(()=>{recorder.dispose();fs.rmSync(root,{recursive:true,force:true})});
const event={hook_event_name:'PreToolUse',tool_name:'Edit',session_id:fixture.hook.sessionId,tool_use_id:fixture.hook.pending.toolUseId,cwd,transcript_path:native,tool_input:{file_path:file,old_string:removed,new_string:added}};
recordAutoplanArtifact(JSON.stringify(event),recorder.file,cwd,config,stateRoot);
const publicTools=structuredClone(fixture.publicTools) as any[];for(const e of publicTools)if(e.input)e.input.file_path=file;
const commandStartedAt=Date.parse(publicTools[0].timestamp)-1;
const pending=readPendingAutoplanArtifact(recorder.file,cwd,config,stateRoot,commandStartedAt,publicTools);
const context={cwd,ownedStateRoot:stateRoot,commandStartedAt,transcriptStatus:'ready',publicTools,pending,now:Date.now()+1000,viewportCapturedAt:Date.now()};
const invoke=(viewport=fixture.viewport,seen=new Set<string>())=>pendingAutoplanArtifactPermissionInput(viewport,context,seen);
return {root,cwd,config,stateRoot,file,recorder,event,context,invoke};
}
const menu=fixture.viewport.slice(fixture.viewport.indexOf('╌'));
const panel=(rows:string[])=>rows.join('\n')+'\n'+menu;
test('exact current clipped pane requires new recorded suffix commitments and preserves original request bytes',()=>{
const r=replay(),digest=r.context.pending!.editDigest!;
expect(digest.beforeSHA256).toBe(fixture.provenance.beforeSHA256);expect(digest.requestSHA256).toBe(fixture.provenance.requestSHA256);
expect(digest.oldLineHashes).toEqual(fixture.hook.pending.editDigest.oldLineHashes);expect(digest.newLineHashes).toEqual(fixture.hook.pending.editDigest.newLineHashes);
expect(digest.clippedAdditions?.status).toBe('complete');expect(r.invoke()?.input).toBe('1\r');
delete digest.clippedAdditions;expect(r.invoke()).toBeNull();expect(r.invoke(fixture.viewport.split('\n').slice(1).join('\n'))?.input).toBe('1\r');
expect(fixture.provenance.actualCoverage).toContain('no phase credit');
});
test('first, middle and last changed lines support full120-column crops and following context',()=>{
const lines=Array.from({length:32},(_,i)=>'Line '+i+' '+String.fromCharCode(65+i%26).repeat(180));
const r=replay('Heading\nAnchor\nAfter one\nAfter two\n','Anchor',lines.join('\n'));
expect(r.context.pending!.editDigest!.clippedAdditions?.status).toBe('complete');
for(const i of [0,15,31]){
const row=i+2,tail=lines[i]!.slice(-114),next=i+1<lines.length?`${row+1} +${lines[i+1]}`:`${row+1} After one`,second=i+2<lines.length?`${row+2} +${lines[i+2]}`:`${row+2} ${i+1<lines.length?'After one':'After two'}`;
const column=String(row+1).length+2;
const viewport=panel([' '.repeat(column)+'+'+tail,' '+next,' '+second]);
expect(r.invoke(viewport)?.input).toBe('1\r');expect(r.invoke(viewport.replace(tail,'foreign'+tail))).toBeNull();
}
expect(fs.statSync(r.recorder.file).size).toBeLessThan(1024*1024);
});
test('exact suffix, corresponding line, next line and complete crop content are all mandatory',()=>{
for(const change of [
(s:string)=>s.replace(/^ \+t\./,' +x.'), (s:string)=>s.replace(/^ \+t\./,' +t!'),
(s:string)=>s.replace(/^ \+t\./,' +t.'),(s:string)=>s.replace(/^ \+t\./,' +t.'),
(s:string)=>s.replace(/^ \+t\./,' -t.'),(s:string)=>s.replace(/^ \+t\./,' Source: t.'),
// A forged deletion marker cannot make rejected digest rows use legacy authority.
(s:string)=>s.replace(/^ \+t\./,' -t.').replace(/^ 139 /m,' 140 '),
(s:string)=>s.replace(/^ \+t\./,' -t.').replace('Snapshot consistency','Foreign consistency'),
(s:string)=>s.replace(/^ 139 /m,' 140 '),(s:string)=>s.replace('Snapshot consistency','Foreign consistency'),
(s:string)=>s.replace('authoritative gate','unrequested gate'),(s:string)=>'> source\n'+s,
(s:string)=>s.replace('3. No','3. Maybe'),(s:string)=>s.replace(' 1. Yes',' 2. Yes'),
(s:string)=>s+'\nUnrelated menu',
]){const r=replay();expect(r.invoke(change(fixture.viewport))).toBeNull()}
});
test('wrong digest, file, current native history and previously seen menu remain denied',()=>{
for(const edit of [
(r:any)=>{r.context.pending.sessionId='foreign';},(r:any)=>{r.context.pending.editDigest.beforeSHA256='0'.repeat(64);},
(r:any)=>{r.context.pending.editDigest.clippedAdditions.lines[0].lineHash='0'.repeat(64);},
(r:any)=>{r.context.publicTools[1].isError=true;},(r:any)=>{r.context.publicTools=[];},
(r:any)=>{r.context.viewportCapturedAt=Date.parse(r.context.pending.timestamp)-1;},
(r:any)=>{fs.appendFileSync(r.file,'changed');fs.utimesSync(r.file,new Date(0),new Date(0));},
(r:any)=>{r.context.pending.file=r.file.replace('user-dashboard','foreign-dashboard');},
(r:any)=>{r.context.publicTools.push({kind:'use',name:'Edit',sessionId:r.context.pending.sessionId,toolUseId:'queued',timestamp:new Date().toISOString(),input:{file_path:r.file}});},
]){const r=replay();edit(r);expect(r.invoke()).toBeNull()}
const r=replay();expect(r.invoke(fixture.viewport,new Set([autoplanArtifactMenuKey(fixture.viewport)]))).toBeNull();expect(r.invoke(fixture.viewport,new Set([r.context.pending!.sessionId+':'+r.context.pending!.toolUseId]))).toBeNull();
});
test('suffix commitments are not body persistence and current replay cannot retain stale hashes',()=>{
const r=replay(),raw=fs.readFileSync(r.recorder.file,'utf8');for(const text of ['old_string','new_string','Preconditions heading','Snapshot consistency'])expect(raw).not.toContain(text);
recordAutoplanArtifact(JSON.stringify(r.event),r.recorder.file,r.cwd,r.config,r.stateRoot);expect(fs.readFileSync(r.recorder.file,'utf8')).toBe(raw);
r.event.tool_input.new_string+='changed';recordAutoplanArtifact(JSON.stringify(r.event),r.recorder.file,r.cwd,r.config,r.stateRoot);
expect(autoplanArtifactRecorderStatus(r.recorder.file,r.cwd,r.config,r.stateRoot)).toEqual({status:'invalid',reason:'conflicting_replay'});
});
test('legacy digest replay is harmless and partial-edge requests do not manufacture suffix authority',()=>{
const r=replay(),state=JSON.parse(fs.readFileSync(r.recorder.file,'utf8'));delete state.pending.editDigest.clippedAdditions;
fs.writeFileSync(r.recorder.file,JSON.stringify(state)+'\n');const raw=fs.readFileSync(r.recorder.file,'utf8');recordAutoplanArtifact(JSON.stringify(r.event),r.recorder.file,r.cwd,r.config,r.stateRoot);expect(fs.readFileSync(r.recorder.file,'utf8')).toBe(raw);
const q=replay('Prefix Anchor suffix\nAfter one\nAfter two\n','Anchor','New');expect(q.context.pending!.editDigest!.clippedAdditions).toBeUndefined();
});
test('malformed, sparse, tampered and excessive suffix records fail closed',()=>{
for(const edit of [
(c:any)=>{c.version=2;},(c:any)=>{c.extra=true;},(c:any)=>{c.startLine=0;},(c:any)=>{c.lines=Array(2);},
(c:any)=>{c.lines[0].suffixHashes=Array(2);},(c:any)=>{c.lines[0].suffixHashes=Array(257).fill('0'.repeat(64));},
(c:any)=>{c.lines[0].nextLineHash='0'.repeat(64);},(c:any)=>{c.lines[0].line++;},
]){const r=replay(),d=r.context.pending!.editDigest!;edit(d.clippedAdditions);expect(validAutoplanEditDigest(d)).toBe(false);expect(r.invoke()).toBeNull()}
const r=replay(),c=r.context.pending!.editDigest!.clippedAdditions;if(c?.status!=='complete')throw Error('missing');const target=c.lines.find(x=>x.line===138)!;target.suffixHashes[1]='0'.repeat(64);expect(r.invoke()).toBeNull();
});
test('overflow is explicit for every crop while complete-row legacy authority remains intact',()=>{
const lines=Array.from({length:40},(_,i)=>'Line '+i+' '+String.fromCharCode(65+i%26).repeat(300));const r=replay('Anchor\nAfter one\nAfter two\n','Anchor',lines.join('\n'));const d=r.context.pending!.editDigest!;
expect(d.clippedAdditions).toEqual({version:1,status:'overflow'});expect(validAutoplanEditDigest(d)).toBe(true);
for(const i of [0,20,39]){const n=i+1,next=i+1<lines.length?lines[i+1]:'After one',last=i+2<lines.length?lines[i+2]:'After two';expect(matchesAutoplanDigestRows([' '.repeat(String(n+1).length+2)+'+'+lines[i]!.slice(-114),` ${n+1} +${next}`,` ${n+2} +${last}`],Buffer.from('Anchor\nAfter one\nAfter two\n'),d)).toBe(false)}
expect(matchesAutoplanDigestRows([' 1 +'+lines[0],' 2 +'+lines[1]],Buffer.from('Anchor\nAfter one\nAfter two\n'),d)).toBe(true);
});
test('new regression files register only the actual Autoplan owner',()=>{
for(const p of ['test/autoplan-clipped-suffix-aq.test.ts','test/fixtures/autoplan-clipped-suffix-aq.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes(p)).map(([owner])=>owner)).toEqual(['autoplan-chain-pty']);
});
+206
View File
@@ -0,0 +1,206 @@
import { capturedPathRebaser } from './helpers/captured-paths';
import { expect, test } from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { createHash } from 'node:crypto';
import fixture from './fixtures/autoplan-command-prefix-au.json';
import * as permission from './helpers/autoplan-artifact-permission';
import { readPendingAutoplanArtifact } from './helpers/autoplan-artifact-recorder';
import { readPlanCountTranscript, type NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
function replay() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-ap-command-'));
const old = path.dirname(path.dirname(fixture.stateRoot));
const runtime = path.join(root, path.basename(old)), cwd = path.join(root, path.basename(fixture.cwd));
const rebase = capturedPathRebaser([[old,runtime],[fixture.cwd,cwd]]);
const hook = rebase.json(fixture.hook);
const stateRoot = rebase.file(fixture.stateRoot), config = rebase.file(fixture.config), file = hook.pending.file;
const events = rebase.json(fixture.publicTools) as NativePublicToolEvent[];
fs.mkdirSync(cwd, { recursive: true }); fs.mkdirSync(path.dirname(file), { recursive: true });
fs.writeFileSync(file, fixture.before, { mode: fixture.targetStat.mode });
const mtime = Number(BigInt(fixture.targetStat.mtimeNs)) / 1e9;
fs.utimesSync(file, mtime, mtime);
fs.mkdirSync(path.dirname(hook.pending.transcriptPath), { recursive: true });
const records = events.map(e => ({ sessionId: e.sessionId, cwd, isSidechain: false, timestamp: e.timestamp,
requestId: e.requestId, message: { id: e.messageId, role: e.kind === 'use' ? 'assistant' : 'user',
content: e.kind === 'use' ? [{ type: 'tool_use', id: e.toolUseId, name: e.name, input: e.input }]
: [{ type: 'tool_result', tool_use_id: e.toolUseId, content: e.content, is_error: e.isError }] } }));
fs.writeFileSync(hook.pending.transcriptPath, records.map(r => JSON.stringify(r)).join('\n') + '\n');
const hookFile = path.join(root, 'hook.json'); fs.writeFileSync(hookFile, JSON.stringify(hook));
const publicTools: NativePublicToolEvent[] = [];
const transcript = readPlanCountTranscript(config, cwd, e => publicTools.push(e));
const now = Date.parse(fixture.viewportCapturedAt), commandStartedAt = Date.parse(fixture.commandTimestamp);
const pending = readPendingAutoplanArtifact(hookFile, cwd, config, stateRoot, commandStartedAt, publicTools, now, true);
const context = { cwd, ownedStateRoot: stateRoot, ownedNativePlansRoot: path.join(config, 'plans'),
commandStartedAt, now, viewportCapturedAt: now, transcriptStatus: transcript.status, publicTools, pending };
return { root, file, context, viewport: rebase.text(fixture.viewport), dispose: () => fs.rmSync(root, { recursive: true, force: true }) };
}
type Replay = ReturnType<typeof replay>;
const pick = (r: Replay, seen = new Set<string>()) => permission.pendingAutoplanArtifactPermissionInput(r.viewport, r.context, seen);
const panel = (viewport: string) => viewport.slice(viewport.search(/^[─╌]{8,}\n {0,3}Edit file/m));
// Exact AY public native prefix; only its owned archive path is relocated onto
// this existing digest fixture. The unpublished Bash body is not reconstructed.
function nativeCards(r: Replay): string {
const relative = path.relative(r.context.ownedStateRoot, r.file).split(path.sep).join('/');
return [
`● Update(~/.gstack/${relative})`, ' ', '● Updated plan', ' ', '● Updated plan', ' ',
'● Bash(mkdir -p ~/.gstack/analytics',
` echo '{"skill":"plan-ceo-review","via":"autoplan","ts":"'$(date -u`,
` +%Y-%m-%dT%H:%M:%SZ)'","iterations":3,"issues_found":56,"issues_…)`,
' ⎿  Waiting…', '', '', '',
].join('\n') + panel(r.viewport);
}
test('native plan redraws and a queued command preserve only the digest-bound pending Edit', () => {
const r = replay(); try {
r.viewport = nativeCards(r);
const granted = pick(r);
expect(granted).toEqual({ input: '1\r', signature: `${r.context.pending!.sessionId}:${r.context.pending!.toolUseId}`, file: r.file });
expect(permission.autoplanArtifactPermissionInput(r.viewport, r.context, new Set())).toBeNull();
expect(permission.publishedAutoplanArtifactPermissionInput(r.viewport, r.context, new Set())).toBeNull();
expect(pick(r, new Set([granted!.signature]))).toBeNull();
expect(pick(r, new Set([permission.autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
} finally { r.dispose(); }
});
const nativeScreens: Array<[string, (s: string) => string]> = [
['foreign Update title', s => s.replace('Update(~/.gstack/', 'Update(/foreign/')],
['unbound redraw', s => s.replace('● Updated plan', '● Updated another file')],
['second Update', s => s.replace('● Updated plan', '● Update(/foreign/plan.md)')],
['second Bash', s => s.replace('● Updated plan', '● Bash(echo another…)')],
['no native redraw', s => s.replaceAll('● Updated plan', '')],
['completed command', s => s.replace('Waiting…', 'Done')],
['missing Waiting marker', s => s.replace(' ⎿  Waiting…', '')],
['unclosed command card', s => s.replace('"issues_…)', '"issues_…')],
['unindented command continuation', s => s.replace(' echo ', 'echo ')],
['competing permission', s => s.replace(' echo ', ' Do you want to proceed? ')],
['indented native action', s => s.replace(' echo ', ' ● Read ')],
['indented question', s => s.replace(' echo ', ' 1. ')],
['source prefix', s => 'Source:\n' + s],
['quoted pane', s => s.split('\n').map(row => '> ' + row).join('\n')],
['code pane', s => '```text\n' + s + '\n```'],
['second edit panel', s => s + '\n' + panel(s)],
['foreign active panel', s => s.replace('projects/gstack-autoplan-chain-zmFsqo/', 'projects/foreign/')],
['foreign menu', s => s.replace('edit to 2026-09-10-user-dashboard.md?', 'edit to other.md?')],
['persistent approval', s => s.replace(' 1. Yes', ' 2. Yes')],
['changed digest-bound addition', s => s.replace('zero before advancing', 'ten before advancing')],
];
for (const [name, change] of nativeScreens) test(`native batch cards cannot hide another authority: ${name}`, () => {
const r = replay(); try { r.viewport = change(nativeCards(r)); expect(pick(r)).toBeNull(); } finally { r.dispose(); }
});
test('the exact public command display preserves only the current unpublished Edit approval', () => {
const r = replay(); try {
expect(r.context.transcriptStatus).toBe('ready'); expect(r.context.publicTools).toHaveLength(2);
expect(r.context.pending?.toolUseId).toBe(fixture.hook.pending.toolUseId);
expect(r.context.publicTools.some(e => e.toolUseId === r.context.pending?.toolUseId)).toBe(false);
expect(createHash('sha256').update(fs.readFileSync(r.file)).digest('hex')).toBe(fixture.provenance.beforeSHA256);
expect(Math.floor(fs.statSync(r.file).mtimeMs)).toBe(Number(BigInt(fixture.targetStat.mtimeNs) / 1_000_000n));
expect(permission.autoplanArtifactPermissionInput(r.viewport, r.context, new Set())).toBeNull();
expect(permission.publishedAutoplanArtifactPermissionInput(r.viewport, r.context, new Set())).toBeNull();
const expected = { input: '1\r', signature: `${fixture.hook.sessionId}:${fixture.hook.pending.toolUseId}`, file: r.file };
expect(pick(r)).toEqual(expected);
expect(pick(r, new Set([expected.signature]))).toBeNull();
expect(pick(r, new Set([permission.autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
r.viewport = panel(r.viewport); expect(pick(r)).toEqual(expected);
expect(fixture.provenance.paidOutcomesReclassified).toBe(false);
} finally { r.dispose(); }
});
test('a plain native command description and wrapped display supply no command authority', () => {
for (const prefix of ['● Recording review metrics\n ⎿ $ echo recorded\n\n',
'⏺ Running local diagnostics\n ⎿ $ bun test\n echo finished\n\n']) {
const r = replay(); try { r.viewport = prefix + panel(r.viewport); expect(pick(r)?.input).toBe('1\r'); }
finally { r.dispose(); }
}
});
const screens: Array<[string, (s: string) => string]> = [
['source introduction', s => 'Source:\n' + s], ['example introduction', s => 'Example:\n' + s],
['whole quotation', s => s.split('\n').map(line => '> ' + line).join('\n')],
['whole code block', s => '```text\n' + s + '\n```'],
['quoted title', s => s.replace('● Appending spec-review metrics', '● "Appending spec-review metrics"')],
['source title', s => s.replace('● Appending spec-review metrics', '● Source: an example command')],
['second native action', s => s.replace(' echo logged', '● Another tool\n ⎿ $ echo other')],
['indented second action', s => s.replace(' echo logged', ' ● Another tool')],
['Bash confirmation', s => s.replace(' echo logged', ' Do you want to proceed?')],
['Bash permission', s => s.replace(' echo logged', ' Bash command requires permission')],
['second question', s => s.replace(' echo logged', ' 1. Approve this command')],
['missing command marker', s => s.replace('⎿ $', '⎿ ')],
['unbound command prose', s => s.replace(' echo logged', 'Unrelated current prose')],
['second edit panel', s => s + '\n' + panel(s)],
['foreign panel path', s => s.replace('projects/gstack-autoplan-chain-zmFsqo/', 'projects/another-project/')],
['basename-only panel', s => s.replace(/^ …[^\n]+$/m, ' 2026-09-10-user-dashboard.md')],
['foreign menu target', s => s.replace('edit to 2026-09-10-user-dashboard.md?', 'edit to another.md?')],
['session approval cursor', s => s.replace(' 1. Yes', ' 2. Yes')],
['extra current prompt', s => s + '\nChoose another action'],
['changed added rows', s => s.replace('zero before advancing', 'ten before advancing')],
['removed-line gap', s => s.replace(' 98 -', ' 100 -')],
['added-line gap', s => s.replace(' 98 +', ' 100 +')],
['different reset start', s => s.replace(' 97 +', ' 96 +')],
['duplicate removed row', s => s.replace(/(^ 98 -[^\n]*\n)/m, '$1$1')],
['duplicate added row', s => s.replace(/(^ 98 \+[^\n]*\n)/m, '$1$1')],
['multiple resets', s => s.replace(' 108 5.', s.slice(s.indexOf(' 97 -'), s.indexOf(' 108 5.')) + ' 108 5.')],
['truncated removed block', s => s.replace(/^ 99 -[^\n]*\n/m, '')],
['truncated added block', s => s.replace(/^ 107 \+[^\n]*\n/m, '')],
['missing panel separator', s => s.replace(/^[─]{8,}\n/m, '')],
];
for (const [name, change] of screens) test(`current panel remains unambiguous: ${name}`, () => {
const r = replay(); try { r.viewport = change(r.viewport); expect(pick(r)).toBeNull(); } finally { r.dispose(); }
});
const bindings: Array<[string, (r: Replay) => void]> = [
['missing hook', r => { r.context.pending = undefined; }],
['wrong hook tool', r => { r.context.pending!.tool = 'Write' as 'Edit'; }],
['foreign hook session', r => { r.context.pending!.sessionId = 'foreign'; }],
['foreign hook path', r => { r.context.pending!.file += '.other'; }],
['missing digest', r => { delete r.context.pending!.editDigest; }],
['invalid digest', r => { r.context.pending!.editDigest!.beforeSHA256 = 'invalid'; }],
['wrong before digest', r => { r.context.pending!.editDigest!.beforeSHA256 = '0'.repeat(64); }],
['wrong addition commitments', r => { r.context.pending!.editDigest!.newLineHashes.fill('0'.repeat(64)); }],
['current file changed', r => { fs.appendFileSync(r.file, '\nchanged'); fs.utimesSync(r.file, new Date(0), new Date(0)); }],
['file newer than pending', r => { fs.utimesSync(r.file, new Date(r.context.now), new Date(r.context.now)); }],
['history is Read', r => { r.context.publicTools[0]!.name = 'Read'; }],
['foreign history file', r => { r.context.publicTools[0]!.input!.file_path = r.file + '.other'; }],
['failed history', r => { r.context.publicTools[1]!.isError = true; }],
['unresolved mutation', r => { r.context.publicTools.pop(); }],
['published pending request', r => { r.context.publicTools.push({ kind: 'use', name: 'Edit', sessionId: r.context.pending!.sessionId,
toolUseId: r.context.pending!.toolUseId, timestamp: r.context.pending!.timestamp, input: { file_path: r.file } }); }],
['missing transcript', r => { r.context.transcriptStatus = 'missing'; }],
['future hook', r => { r.context.pending!.timestamp = new Date(r.context.now + 1).toISOString(); }],
['viewport predates hook', r => { r.context.viewportCapturedAt = Date.parse(r.context.pending!.timestamp) - 1; }],
['command after hook', r => { r.context.commandStartedAt = Date.parse(r.context.pending!.timestamp) + 1; }],
];
for (const [name, change] of bindings) test(`pending authorization is retained: ${name}`, () => {
const r = replay(); try { change(r); expect(pick(r)).toBeNull(); } finally { r.dispose(); }
});
for (const [name, change] of bindings) test(`native cards retain pending authorization: ${name}`, () => {
const r = replay(); try { r.viewport = nativeCards(r); change(r); expect(pick(r)).toBeNull(); } finally { r.dispose(); }
});
test('only the Autoplan workflow selects this fixture and behavioral regression', () => {
for (const file of ['test/autoplan-command-prefix-au.test.ts', 'test/fixtures/autoplan-command-prefix-au.json'])
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['autoplan-chain-pty']);
});
test('removed row order is bound to both the current file and pending digest', () => {
const r = replay(); try {
const rows = r.viewport.split('\n'), a = rows.findIndex(row => /^ 97 -/.test(row)), b = rows.findIndex(row => /^ 98 -/.test(row));
expect(a).toBeGreaterThan(0); expect(b).toBe(a + 1);
const first = rows[a]!.slice(6), second = rows[b]!.slice(6);
rows[a] = rows[a]!.slice(0, 6) + second; rows[b] = rows[b]!.slice(0, 6) + first;
r.viewport = rows.join('\n'); expect(pick(r)).toBeNull();
} finally { r.dispose(); }
});
test('added row order is bound to the complete pending replacement digest', () => {
const r = replay(); try {
const rows = r.viewport.split('\n'), a = rows.findIndex(row => /^ 97 \+/.test(row)), b = rows.findIndex(row => /^ 98 \+/.test(row));
expect(a).toBeGreaterThan(0); expect(b).toBe(a + 1);
const first = rows[a]!.slice(6), second = rows[b]!.slice(6);
rows[a] = rows[a]!.slice(0, 6) + second; rows[b] = rows[b]!.slice(0, 6) + first;
r.viewport = rows.join('\n'); expect(pick(r)).toBeNull();
} finally { r.dispose(); }
});
+116
View File
@@ -0,0 +1,116 @@
import {expect,test} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';import {createHash} from 'node:crypto';
import fixture from './fixtures/autoplan-cropped-command-av.json';
import * as permission from './helpers/autoplan-artifact-permission';
import {E2E_TOUCHFILES,LLM_JUDGE_TOUCHFILES,selectTests} from './helpers/touchfiles';
type Context=Parameters<typeof permission.publishedAutoplanArtifactPermissionInput>[1];
function replay(){
const root=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-cropped-command-'));
const replace=(s:string)=>s.replaceAll(path.dirname(fixture.context.cwd),root);
const context=JSON.parse(replace(JSON.stringify(fixture.context))) as Context;
const nativePlan=replace(fixture.nativePlan.path),file=context.pending!.file;
for(const [name,body,mtime] of [[file,fixture.before,fixture.beforeMtimeMs],[nativePlan,fixture.nativePlan.text,fixture.nativePlan.mtimeMs]] as const){
fs.mkdirSync(path.dirname(name),{recursive:true});fs.writeFileSync(name,body);fs.utimesSync(name,mtime/1000,mtime/1000);
}
fs.mkdirSync(context.cwd,{recursive:true});
return {root,file,nativePlan,context,viewport:replace(fixture.viewport),dispose:()=>fs.rmSync(root,{recursive:true,force:true})};
}
type Replay=ReturnType<typeof replay>;
const invoke=(r:Replay,seen=new Set<string>())=>permission.publishedAutoplanArtifactPermissionInput(r.viewport,r.context,seen);
const current=(r:Replay)=>r.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId===r.context.pending!.toolUseId)!;
const queued=(r:Replay)=>r.context.publicTools.find(e=>e.kind==='use'&&e.name==='Edit'&&e.input?.file_path===r.nativePlan&&
!r.context.publicTools.some(result=>result.kind==='result'&&result.toolUseId===e.toolUseId))!;
const bash=(r:Replay)=>r.context.publicTools.find(e=>e.kind==='use'&&e.name==='Bash')!;
const complete=(r:Replay,e:ReturnType<typeof bash>,isError=false)=>r.context.publicTools.push({kind:'result',sessionId:e.sessionId,
toolUseId:e.toolUseId,timestamp:new Date(r.context.now!).toISOString(),isError,content:'completed'});
const panel=(r:Replay)=>r.viewport.slice(r.viewport.search(/^[─╌]{8,}\n {0,3}Edit file/m));
const show=(r:Replay,command:string,rows=[command])=>{bash(r).input!.command=command;r.viewport=' ⎿ $ '+rows.join('\n ')+'\n\n'+panel(r)};
test('the retained captionless queued command grants only the current digest-bound Edit',()=>{const r=replay();try{
expect(r.context.publicTools).toHaveLength(7);
expect(createHash('sha256').update(fs.readFileSync(r.file)).digest('hex')).toBe(fixture.beforeSha256);
const expected={input:'1\r',signature:fixture.context.pending.sessionId+':'+fixture.context.pending.toolUseId,file:r.file};
expect(permission.autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(permission.pendingAutoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(invoke(r)).toEqual(expected);
expect(invoke(r,new Set([expected.signature]))).toBeNull();
expect(invoke(r,new Set([permission.autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
r.viewport=panel(r);expect(invoke(r)).toEqual(expected);
expect(fixture.provenance.paidOutcomesReclassified).toBe(false);
}finally{r.dispose()}});
for(const [name,change] of [
['single row',(r:Replay)=>show(r,bash(r).input!.command)],
['different soft wrap',(r:Replay)=>{const command=bash(r).input!.command as string;const at=command.indexOf(' && ');show(r,command,[command.slice(0,at),command.slice(at+1)])}],
['CRLF renderer',(r:Replay)=>{r.viewport=r.viewport.replaceAll('\n','\r\n')}],
['nonbreaking native gutter',(r:Replay)=>{r.viewport=r.viewport.replace('⎿ $','⎿\u00a0 $')}],
['quoted argument with literal spaces',(r:Replay)=>show(r,"printf '%s' 'two words'")],
['soft wrap inside a quoted argument',(r:Replay)=>show(r,"printf '%s' 'two words'",["printf '%s' 'two","words'"])],
] as const)test(`complete public command binding accepts ${name}`,()=>{const r=replay();try{change(r);expect(invoke(r)?.signature).toBe(`${r.context.pending!.sessionId}:${r.context.pending!.toolUseId}`)}finally{r.dispose()}});
const identityCases:Array<[string,(r:Replay)=>void]>=[
['missing Bash publication',r=>{r.context.publicTools=r.context.publicTools.filter(e=>e!==bash(r))}],
['foreign Bash message',r=>{bash(r).messageId='msg_foreign'}],['foreign Bash request',r=>{bash(r).requestId='req_foreign'}],
['foreign Bash session',r=>{bash(r).sessionId='foreign'}],['Bash with no identity',r=>{bash(r).toolUseId=''}],
['different command',r=>{bash(r).input!.command+=' && echo other'}],['missing command',r=>{delete bash(r).input!.command}],
['multiline command',r=>{bash(r).input!.command+='\n'}],['control byte in command',r=>{bash(r).input!.command+='\x1b'}],
['another tool name',r=>{bash(r).name='Read'}],['started command',r=>{r.context.pending!.hookSeenIds!.push(bash(r).toolUseId)}],
['completed command',r=>complete(r,bash(r))],['failed command',r=>complete(r,bash(r),true)],
['ambiguous queued commands',r=>{r.context.publicTools.push({...structuredClone(bash(r)),toolUseId:'toolu_duplicate'})}],
['second unmatched queued command',r=>{r.context.publicTools.push({...structuredClone(bash(r)),toolUseId:'toolu_other',input:{command:'echo other'}})}],
['command after viewport',r=>{r.context.viewportCapturedAt=Date.parse(bash(r).timestamp)-1}],
['command before queued Edit',r=>{const e=bash(r),q=queued(r),at=r.context.publicTools.indexOf(q);e.timestamp=current(r).timestamp;r.context.publicTools.pop();r.context.publicTools.splice(at,0,e)}],
['no queued mutation',r=>{const q=queued(r);r.context.publicTools=r.context.publicTools.filter(e=>e!==q)}],
['foreign queued mutation path',r=>{queued(r).input!.file_path='/tmp/foreign.md'}],
['foreign queued message',r=>{queued(r).messageId='msg_foreign'}],['foreign queued request',r=>{queued(r).requestId='req_foreign'}],
['queued Write',r=>{queued(r).name='Write'}],['queued replace all',r=>{queued(r).input!.replace_all=true}],
['started queued Edit',r=>{r.context.pending!.hookSeenIds!.push(queued(r).toolUseId)}],
['completed queued Edit',r=>complete(r,queued(r))],
['failed native-plan history',r=>{const previous=r.context.publicTools.find(e=>e.kind==='use'&&e.input?.file_path===r.nativePlan&&e!==queued(r))!;r.context.publicTools.find(e=>e.kind==='result'&&e.toolUseId===previous.toolUseId)!.isError=true}],
['Read is not native-plan mutation history',r=>{r.context.publicTools.find(e=>e.kind==='use'&&e.input?.file_path===r.nativePlan&&e!==queued(r))!.name='Read'}],
['native plan modified after hook',r=>{fs.utimesSync(r.nativePlan,new Date(r.context.now!),new Date(r.context.now!))}],
['foreign native-plan root',r=>{r.context.ownedNativePlansRoot=path.join(r.root,'foreign')}],
['missing hook',r=>{r.context.pending=undefined}],['missing current publication',r=>{const c=current(r);r.context.publicTools=r.context.publicTools.filter(e=>e!==c)}],
['foreign current message',r=>{current(r).messageId='msg_foreign'}],['foreign current request',r=>{current(r).requestId='req_foreign'}],
['foreign current session',r=>{current(r).sessionId='foreign'}],['completed current Edit',r=>complete(r,current(r))],
['current request changed',r=>{current(r).input!.new_string+=' changed'}],
['missing digest',r=>{delete r.context.pending!.editDigest}],['wrong request digest',r=>{r.context.pending!.editDigest!.requestSHA256='0'.repeat(64)}],
['wrong before digest',r=>{r.context.pending!.editDigest!.beforeSHA256='0'.repeat(64)}],
['file changed',r=>{fs.appendFileSync(r.file,'changed');fs.utimesSync(r.file,0,0)}],
['file modified after hook',r=>{fs.utimesSync(r.file,new Date(r.context.now!),new Date(r.context.now!))}],
['missing native transcript',r=>{r.context.transcriptStatus='missing'}],['wrong pending identity',r=>{r.context.pending!.toolUseId='toolu_other'}],
['failed archive history',r=>{r.context.publicTools.find(e=>e.kind==='result')!.isError=true}],
['command before launched review',r=>{r.context.commandStartedAt=r.context.now!+1}],
];
for(const[name,change]of identityCases)test(`caption crop retains native authority: ${name}`,()=>{const r=replay();try{change(r);expect(invoke(r)).toBeNull()}finally{r.dispose()}});
const displayCases:Array<[string,(r:Replay)=>void]>=[
['example introduction',r=>{r.viewport='Example:\n'+r.viewport}],['historical introduction',r=>{r.viewport='Historical screen:\n'+r.viewport}],
['quoted display',r=>{r.viewport=r.viewport.split('\n').map(line=>'> '+line).join('\n')}],
['fenced display',r=>{r.viewport='```text\n'+r.viewport+'\n```'}],
['caption instead of native cropped prefix',r=>{r.viewport='● Approve everything\n'+r.viewport}],
['missing dollar marker',r=>{r.viewport=r.viewport.replace('⎿ $','⎿ ')}],
['different command prefix',r=>{r.viewport=r.viewport.replace('mkdir -p','mkdir -m 777 -p')}],
['truncated command',r=>{r.viewport=r.viewport.replace('&& echo logged','…')}],
['extra command suffix',r=>{r.viewport=r.viewport.replace('&& echo logged','&& echo logged; echo other')}],
['missing wrapped row',r=>{r.viewport=r.viewport.split('\n').filter((_,i)=>i!==1).join('\n')}],
['blank row in command',r=>{r.viewport=r.viewport.replace('\n +%','\n\n +%')}],
['extra non-command row',r=>{r.viewport=r.viewport.replace('\n \n','\n completed successfully\n')}],
['second dollar command',r=>{r.viewport=r.viewport.replace('\n \n','\n ⎿ $ echo other\n')}],
['Bash approval menu',r=>{r.viewport='Bash command permission\nDo you want to run this command?\n'+r.viewport}],
['duplicate Edit panel',r=>{r.viewport+=panel(r)}],
['foreign Edit target',r=>{r.viewport=r.viewport.replace('gstack-autoplan-chain-ZdZS9F','gstack-autoplan-chain-foreign')}],
['altered added diff row',r=>{r.viewport=r.viewport.replace('server clock','attacker clock')}],
['persistent permission selected',r=>{r.viewport=r.viewport.replace(' 1. Yes',' 2. Yes')}],
['trailing unrelated prose',r=>{r.viewport+='\nAnother current request'}],
['within-row quoted whitespace contradiction',r=>{show(r,"printf '%s' 'two words'");r.viewport=r.viewport.replace('two words','two words')}],
['within-row unquoted whitespace contradiction',r=>{r.viewport=r.viewport.replace('mkdir -p','mkdir -p')}],
];
for(const[name,change]of displayCases)test(`caption crop rejects unrelated display: ${name}`,()=>{const r=replay();try{change(r);expect(invoke(r)).toBeNull()}finally{r.dispose()}});
test('only Autoplan selects the public fixture and focused regression',()=>{
for(const file of ['test/autoplan-cropped-command-av.test.ts','test/fixtures/autoplan-cropped-command-av.json']){
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
expect(selectTests([file],LLM_JUDGE_TOUCHFILES,[]).selected).toEqual([]);
}
});
+95
View File
@@ -0,0 +1,95 @@
import {expect, test} from 'bun:test';
import {readFileSync} from 'node:fs';
import {autoplanBlockingQuestionBoundary, autoplanSetupDecision} from './helpers/autoplan-setup-question';
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
import {E2E_TOUCHFILES, GLOBAL_TOUCHFILES} from './helpers/touchfiles';
import capture from './fixtures/autoplan-cropped-gate-av.json';
const fixture = (): {screen:string; context:Parameters<typeof autoplanBlockingQuestionBoundary>[1]} => ({screen:capture.screen,
context:{commandStartedAt:capture.commandStartedAt,viewportCapturedAt:capture.viewportCapturedAt,
transcript:{status:'ready',calls:[structuredClone(capture.call)],assistantMessages:[]},publicTools:[structuredClone(capture.publicUse)]}});
const call=(f:ReturnType<typeof fixture>)=>f.context.transcript.calls[0]!;
const detect=(f=fixture())=>autoplanBlockingQuestionBoundary(f.screen,f.context);
const expected={sessionId:capture.call.sessionId,toolUseId:capture.call.toolUseId,source:'native'};
const rebind=(f:ReturnType<typeof fixture>)=>{f.context.publicTools[0]!.input!.questions=structuredClone(call(f).questions);};
type Change=(f:ReturnType<typeof fixture>)=>void;
test('exact AV crop proves a human wait without answer, phase credit or evidence mutation',()=>{
const f=fixture(),before=JSON.stringify(f);expect(detect(f)).toEqual(expected);
expect(autoplanSetupDecision(f.screen,new Set(),call(f))).toEqual({kind:'unrelated'});
expect(autoplanPhaseCompletions(f.context.transcript,f.context.commandStartedAt)).toEqual([]);
expect(call(f).answered).toBe(false);expect(call(f).failed).toBe(false);expect(JSON.stringify(f)).toBe(before);
});
test('wrapping and crop position may vary while the owned excerpt and choices remain exact',()=>{
const controls:Change[]=[
f=>{f.screen=f.screen.replace(/\n/g,'\r\n');},f=>{f.screen=f.screen.replace(/^│ /gm,'┃ ');},
f=>{f.screen=f.screen.replace('wall…','wall-clock time');},
f=>{f.screen=f.screen.replace('│ Pros / cons:\n','│ Pros /\n│ cons:\n');},
f=>{f.screen=f.screen.slice(f.screen.indexOf('│ Stakes if'));},
f=>{call(f).questions[0]!.header='Final approval gate';call(f).questions[0]!.question=call(f).questions[0]!.question.replace('D1 — Final Approval Gate: approve the reviewed plan?','D8 — Final Approval: approve the amended plan?');rebind(f);},
// A native human wait stays real even if the question body retracts approval.
f=>{call(f).questions[0]!.question+='\nThis final approval gate is withdrawn.';rebind(f);},
];for(const [i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toEqual(expected);}
});
test('native identity, public use, no acknowledgment, current session and time remain mandatory',()=>{
const controls:Change[]=[
f=>{f.context.transcript.status='missing';},f=>{f.context.transcript.status='error';},f=>{f.context.transcript.calls=[];},f=>{f.context.publicTools=[];},
f=>{call(f).answered=true;},f=>{call(f).failed=true;},f=>{call(f).toolUseId='foreign';},f=>{call(f).sessionId='foreign';},
f=>{f.context.publicTools[0]!.toolUseId='foreign';},f=>{f.context.publicTools[0]!.sessionId='foreign';},f=>{f.context.publicTools[0]!.name='Read';},
f=>{f.context.publicTools[0]!.timestamp='bad';},f=>{f.context.publicTools[0]!.timestamp=new Date(f.context.viewportCapturedAt+1).toISOString();},
f=>{f.context.commandStartedAt=Date.parse(capture.publicUse.timestamp)+1;},f=>{f.context.commandStartedAt=NaN;},f=>{f.context.viewportCapturedAt=Infinity;},
f=>{f.context.publicTools[0]!.input!.questions=[];},f=>{f.context.publicTools[0]!.input!.questions=[{header:'Foreign',question:'Other?'}];},
f=>{f.context.publicTools.push(structuredClone(f.context.publicTools[0]!));},
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:false} as any);},
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:true} as any);},
f=>{f.context.transcript.calls.push({...structuredClone(call(f)),toolUseId:'another'});},
f=>{f.context.transcript.assistantMessages.push({sessionId:'foreign',timestamp:capture.publicUse.timestamp,text:'Unrelated'});},
f=>{call(f).questions[0]!.multiSelect=true;rebind(f);},f=>{call(f).questions.push(structuredClone(call(f).questions[0]!));rebind(f);},
f=>{const pending={...structuredClone(call(f)),source:'pre_tool_use' as const};f.context.transcript.calls=[];f.context.transcript.assistantMessages=[{sessionId:pending.sessionId,timestamp:capture.publicUse.timestamp,text:'Preparing'}];f.context.publicTools=[];f.context.pending=pending;},
];for(const[i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
});
test('copied, ambiguous, partial and mismatched crop displays cannot identify a current gate',()=>{
const controls:Change[]=[
f=>{f.screen='Source panel:\n'+f.screen;},f=>{f.screen='│ Source panel:\n'+f.screen;},f=>{f.screen='Example:\n'+f.screen;},
f=>{f.screen='Historical example:\n'+f.screen;},f=>{f.screen='```text\n'+f.screen;},f=>{f.screen='│ ```text\n'+f.screen;},
f=>{f.screen='> '+f.screen.replace(/\n/g,'\n> ');},f=>{f.screen=' '+f.screen.replace(/\n/g,'\n ');},
f=>{f.screen=f.screen.replace(/^│ /gm,'');},f=>{f.screen=f.screen.slice(f.screen.indexOf(' 1.'));},
f=>{f.screen=f.screen.replace('the confirmation modal','the unrelated confirmation');},f=>{f.screen=f.screen.replace('│ Pros / cons:\n','');},
f=>{f.screen=f.screen.replace('│ Pros / cons:\n','│ Different question?\n');},f=>{f.screen=f.screen.replace(' 1.',' 1.');},
f=>{f.screen=f.screen.replace(' 2.',' 2.');},f=>{f.screen=f.screen.replace(' 2.',' 7.');},
f=>{f.screen=f.screen.replace('1. Approve as-is (recommended)','1. Ship immediately');},
f=>{f.screen=f.screen.replace('Accept all 117 auto-decisions','Reject all 117 auto-decisions');},
f=>{f.screen=f.screen.replace(' Accept all 117 auto-decisions and the 4 taste recommendations; write review logs; suggest /ship.\n','');},
f=>{f.screen=f.screen.replace(' 5. Type something.',' 5. Submit answers');},f=>{f.screen=f.screen.replace(' 6. Chat about this','');},
f=>{f.screen=f.screen.replace(' 6. Chat about this',' 6. Chat about this\n 7. Another option');},
f=>{f.screen=f.screen.replace('Esc to cancel','Esc to');},f=>{f.screen+='Another current panel\n';},
f=>{f.screen=f.screen.replace('│ Pros / cons:','│ ☐ Other gate\n│ Pros / cons:');},
f=>{f.screen=f.screen.replace(' 5. Type something.',' 5. Type something.\nOther confirmation');},
f=>{f.screen=f.screen.replace(' 6. Chat about this',' 6. Chat about this\nOther confirmation');},
f=>{call(f).questions[0]!.header='Setup';rebind(f);},f=>{call(f).questions[0]!.question='Example: '+call(f).questions[0]!.question;rebind(f);},
f=>{call(f).questions[0]!.question='"'+call(f).questions[0]!.question+'"';rebind(f);},
];for(const[i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
});
test('unchanged production loop fails as blocked and sends no input while preserving missing phases',async()=>{
const source=readFileSync(new URL('./skill-e2e-autoplan-chain.test.ts',import.meta.url),'utf8');
const begin=source.indexOf(' // This new repository offers routing'),end=source.indexOf('\n }\n } finally',begin);
expect(begin).toBeGreaterThan(0);expect(end).toBeGreaterThan(begin);
const AsyncFunction=Object.getPrototypeOf(async()=>{}).constructor;
const loop=new AsyncFunction('autoplanBlockingQuestionBoundary','autoplanSetupDecision','ctx',new Bun.Transpiler({loader:'ts'}).transformSync(`
async function run(){const {commandStartedAt,viewportCapturedAt,transcript,publicTools}=ctx;
const hits=[],methodologyAudit=['ceo','design','dx','eng'].map(phase=>({phase,passed:true})),pendingSetupQuestion=undefined;
let outcome='timeout',evidence='',blockedQuestion=null,unsupportedSetup=null;
const inputs=[],seenSetupQuestions=new Set(),session={send:(s)=>inputs.push(s)},Bun={sleep:async()=>{}};
const selectPtyNumberedOption=async(_session,n)=>session.send(String(n)+'\\r'),isPlanReadyVisible=()=>false;
for(const visible of [ctx.screen,ctx.screen]){const viewport=visible;${source.slice(begin,end)}}
return {outcome,blockedQuestion,hits,inputs};}`)+'return run();');
const f=fixture(),result=await loop(autoplanBlockingQuestionBoundary,autoplanSetupDecision,{...f.context,screen:f.screen});
expect(result).toEqual({outcome:'blocked_on_question',blockedQuestion:expected,hits:[],inputs:[]});
const errorStart=source.indexOf(" if (outcome === 'blocked_on_question')"),errorEnd=source.indexOf(" if (outcome === 'exited'",errorStart);
const raise=new Function('outcome','hits','blockedQuestion','transcript','artifacts','evidence',new Bun.Transpiler({loader:'ts'}).transformSync(source.slice(errorStart,errorEnd)));
expect(()=>raise(result.outcome,result.hits,result.blockedQuestion,f.context.transcript,{},f.screen)).toThrow('missing phase markers=[1,2,2.5,3]');
});
test('new fixture and test select only the Autoplan owner',()=>{
for(const p of ['test/autoplan-cropped-gate-av.test.ts','test/fixtures/autoplan-cropped-gate-av.json']){
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(p)).map(([name])=>name)).toEqual(['autoplan-chain-pty']);expect(GLOBAL_TOUCHFILES).not.toContain(p);
}
});
+153
View File
@@ -0,0 +1,153 @@
import {test,expect,afterEach} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';import {spawnSync} from 'node:child_process';
import fixture from './fixtures/autoplan-edit-digests-al.json';
import {createAutoplanArtifactRecorder,recordAutoplanArtifact,readPendingAutoplanArtifact,autoplanArtifactRecorderStatus} from './helpers/autoplan-artifact-recorder';
import {pendingAutoplanArtifactPermissionInput,autoplanArtifactMenuKey} from './helpers/autoplan-artifact-permission';
import {createAutoplanEditDigest,validAutoplanEditDigest,autoplanEditLineHash} from './helpers/autoplan-artifact-digest';
import type {NativePublicToolEvent} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
const cleanups:Array<()=>void>=[];afterEach(()=>{for(const cleanup of cleanups.splice(0))cleanup()});
function replay(record=true) {
const root=fs.mkdtempSync(path.join(os.tmpdir(),'ap-digest-')),cwd=path.join(root,path.basename(fixture.cwd)),ownedStateRoot=path.join(root,'home','.gstack'),config=path.join(root,'config');
const file=fixture.pending.file.replace(fixture.ownedStateRoot,ownedStateRoot),transcript=path.join(config,'projects','owned',fixture.pending.sessionId+'.jsonl');
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(file),{recursive:true});fs.mkdirSync(path.dirname(transcript),{recursive:true});fs.writeFileSync(file,fixture.before);fs.writeFileSync(transcript,'');
const old=new Date(Date.parse(fixture.pending.timestamp)-1000);fs.utimesSync(file,old,old);
const recorder=createAutoplanArtifactRecorder(cwd,config,ownedStateRoot);cleanups.push(()=>{recorder.dispose();fs.rmSync(root,{recursive:true,force:true})});
const event={hook_event_name:'PreToolUse',tool_name:'Edit',session_id:fixture.pending.sessionId,tool_use_id:fixture.pending.toolUseId,cwd,transcript_path:transcript,tool_input:{file_path:file,...fixture.reconstructedRequest}};
const history=structuredClone(fixture.events) as NativePublicToolEvent[];for(const e of history)if(e.input?.file_path===fixture.pending.file)e.input.file_path=file;
const startedAt=Date.parse(history[0]!.timestamp)-1;
if(record)recordAutoplanArtifact(JSON.stringify(event),recorder.file,cwd,config,ownedStateRoot);
const pending=readPendingAutoplanArtifact(recorder.file,cwd,config,ownedStateRoot,startedAt,history);
const context={cwd,ownedStateRoot,commandStartedAt:startedAt,transcriptStatus:'ready',publicTools:history,pending,viewportCapturedAt:Date.now(),now:Date.now()+1000};
return {root,file,config,recorder,event,context,viewport:fixture.viewport};
}
const pick=(r:ReturnType<typeof replay>,seen=new Set<string>())=>pendingAutoplanArtifactPermissionInput(r.viewport,r.context,seen);
test('actual added-only three-digit pane rejects without request digests, accepts a separately recorded reconstructed insertion',()=>{
const r=replay();expect(r.context.pending?.editDigest).toBeDefined();expect(pick(r)?.input).toBe('1\r');
delete r.context.pending!.editDigest;expect(pick(r)).toBeNull();
expect(fixture.provenance.reconstruction).toContain('not the original');
});
test('hook persists bounded digests from its input, never request or result text',()=>{
const r=replay(false),hook=r.recorder.hooks.PreToolUse[0]!.hooks[0]!;
const child=spawnSync('bash',['-c',hook.command],{input:JSON.stringify({...r.event,tool_response:'PRIVATE_RESULT_SENTINEL'}),encoding:'utf8',timeout:6000});
expect(child.status).toBe(0);expect(child.stdout).toBe('');expect(child.stderr).toBe('');
const raw=fs.readFileSync(r.recorder.file,'utf8'),state=JSON.parse(raw);expect(validAutoplanEditDigest(state.pending.editDigest)).toBe(true);
for(const secret of ['old_string','new_string','PRIVATE_RESULT_SENTINEL','Owner: the user.','Toast stacking ahead'])expect(raw).not.toContain(secret);
expect(state.pending.editDigest).toEqual(createAutoplanEditDigest(r.file,r.event.tool_input.old_string,r.event.tool_input.new_string));
expect(fs.statSync(r.recorder.file).size).toBeLessThan(1024*1024);expect(fs.statSync(r.recorder.file).mode&0o777).toBe(0o600);
const legacy=structuredClone(state);delete legacy.pending.editDigest.clippedAdditions;
expect(Buffer.byteLength(JSON.stringify(legacy))).toBeLessThan(64*1024);
});
test('digests of a different request cannot authorize the displayed additions',()=>{
const r=replay();r.context.pending!.editDigest=createAutoplanEditDigest(r.file,'Owner: the user.\n','Owner: the user.\nDifferent requested insertion.\n')!;expect(pick(r)).toBeNull();
r.context.pending!.editDigest=createAutoplanEditDigest(r.file,r.event.tool_input.old_string,r.event.tool_input.new_string)!;
r.context.pending!.editDigest.newLineHashes=r.context.pending!.editDigest.oldLineHashes;expect(pick(r)).toBeNull();
});
test('current file hash, native identity, predecessor and single use stay required',()=>{
const mutations:Array<(r:ReturnType<typeof replay>)=>void>=[
r=>{r.context.pending!.sessionId='foreign';},r=>{r.context.pending!.file=path.join(r.root,'foreign.md');},
r=>{r.context.pending!.editDigest!.beforeSHA256='0'.repeat(64);},
r=>{fs.writeFileSync(r.file,fixture.before+'Changed concurrently.');const old=new Date(0);fs.utimesSync(r.file,old,old);},
r=>{r.context.pending!.timestamp=new Date(r.context.now+1000).toISOString();},r=>{r.context.viewportCapturedAt=Date.parse(r.context.pending!.timestamp)-1;},
r=>{r.context.commandStartedAt=r.context.now+1;},r=>{r.context.publicTools[1]!.isError=true;},r=>{r.context.publicTools=[];},
r=>{r.context.publicTools.push({kind:'result',sessionId:r.context.pending!.sessionId,toolUseId:r.context.pending!.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:false});},
r=>{r.context.publicTools.push({kind:'use',name:'Write',sessionId:r.context.pending!.sessionId,toolUseId:'successor',timestamp:new Date(r.context.now).toISOString(),input:{file_path:r.file}});},
];for(const change of mutations){const r=replay();change(r);expect(pick(r)).toBeNull()}
const r=replay();expect(pick(r,new Set([r.context.pending!.sessionId+':'+r.context.pending!.toolUseId]))).toBeNull();expect(pick(r,new Set([autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
});
test('exact three-digit marker column rejects wrong gutters, arbitrary source rows and malformed numbering',()=>{
for(const change of [
(s:string)=>s.replace(/^ \+/m,' +'),(s:string)=>s.replace(/^ \+/m,' +'),
(s:string)=>s.replace(/^ \+/m,' Source: '),(s:string)=>'> quoted example\n'+s,
(s:string)=>s.replace(/^ 140 /m,' 0 '),(s:string)=>s.replace(/^ 141 /m,' 139 '),
(s:string)=>s.replace(/^ 140 /m,' 999999999999999999999 '),(s:string)=>s.replace(/^ 140 \+/m,' 140 -'),
(s:string)=>s.replace('3. No','3. Maybe'),(s:string)=>s.replace(' 1. Yes',' 2. Yes'),
(s:string)=>s.replace('2026-09-10-user-dashboard.md?','foreign.md?'),(s:string)=>s+'\nUnrelated prompt',
]){const r=replay();r.viewport=change(r.viewport);expect(pick(r)).toBeNull()}
});
test('four-space continuation is accepted only with the matching two-digit numbered gutter',()=>{
const r=replay();r.viewport=r.viewport.replace(/^ (1[4][0-9]) /gm,(_,n)=>' '+(Number(n)-130)+' ').replace(/^ ([+ -])/gm,' $1');expect(pick(r)?.input).toBe('1\r');
});
test('original or context rows cannot supply insertion authority',()=>{
const r=replay(),menu=r.viewport.slice(r.viewport.indexOf('╌'));
r.context.pending!.editDigest=createAutoplanEditDigest(r.file,'Owner: the user.\n','Owner: the user.\nNew actual request.\n')!;
r.viewport=' 140 +Owner: the user.\n 141 +Owner: the user.\n'+menu;expect(pick(r)).toBeNull();
r.viewport=' 140 Owner: the user.\n 141 Owner: the user.\n'+menu;expect(pick(r)).toBeNull();
});
test('malformed, sparse and high-volume persisted digest records fail closed',()=>{
for(const change of [(d:any)=>{d.version=2},(d:any)=>{d.extra='text'},(d:any)=>{d.beforeSHA256='bad'},(d:any)=>{d.newLineHashes=[]},(d:any)=>{d.newLineHashes=Array(513).fill('a'.repeat(64))},(d:any)=>{d.oldLineHashes[0]=null}]){
const r=replay(),s=JSON.parse(fs.readFileSync(r.recorder.file,'utf8'));change(s.pending.editDigest);fs.writeFileSync(r.recorder.file,JSON.stringify(s));expect(autoplanArtifactRecorderStatus(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot).status).toBe('invalid');
r.context.pending!.editDigest=s.pending.editDigest;expect(pick(r)).toBeNull();
}
const r=replay(),sparse={...r.context.pending!.editDigest!,newLineHashes:Array(2)};expect(validAutoplanEditDigest(sparse)).toBe(false);
});
test('unavailable or oversized before/request data yields no new digest authority',()=>{
const r=replay();expect(createAutoplanEditDigest(r.file,'missing original','new')).toBeUndefined();expect(createAutoplanEditDigest(r.file,'Owner: the user.\n','x\n'.repeat(513))).toBeUndefined();
const link=path.join(r.root,'linked');fs.symlinkSync(r.file,link);expect(createAutoplanEditDigest(link,r.event.tool_input.old_string,r.event.tool_input.new_string)).toBeUndefined();
fs.writeFileSync(r.file,'x'.repeat(1024*1024+1));expect(createAutoplanEditDigest(r.file,'x','new')).toBeUndefined();fs.unlinkSync(r.file);expect(createAutoplanEditDigest(r.file,'old','new')).toBeUndefined();
});
test('normalization joins display wrapping but keeps changed nonwhitespace bytes distinct',()=>{
expect(autoplanEditLineHash('same body\t')).toBe(autoplanEditLineHash('samebody'));expect(autoplanEditLineHash('same body')).not.toBe(autoplanEditLineHash('different body'));
const r=replay();r.viewport=r.viewport.replace('Toast stacking','Toast stacKING');expect(pick(r)).toBeNull();
});
test('only Autoplan owns the new digest helper and regression evidence',()=>{
const owner=E2E_TOUCHFILES['autoplan-chain-pty']!;for(let i=0;i<owner.length;i++){expect(Object.hasOwn(owner,i)).toBe(true);expect(typeof owner[i]).toBe('string');}
for(const file of ['test/helpers/autoplan-artifact-digest.ts','test/autoplan-edit-digests-al.test.ts','test/fixtures/autoplan-edit-digests-al.json'])expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
test('identical pending hook replay cannot refresh digest or timestamp',()=>{
const r=replay(),before=fs.readFileSync(r.recorder.file,'utf8');recordAutoplanArtifact(JSON.stringify(r.event),r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot);expect(fs.readFileSync(r.recorder.file,'utf8')).toBe(before);
});
test.each(['changed-new','changed-old','whitespace-only','missing-input','over-limit','replace-all','before-file'])('same pending identity with %s invalidates prior digest authority',kind=>{
const r=replay(),e=structuredClone(r.event) as any;
if(kind==='changed-new')e.tool_input.new_string+='A different final action.\n';
if(kind==='changed-old')e.tool_input.old_string='Owner: the user.';
if(kind==='whitespace-only')e.tool_input.new_string=e.tool_input.new_string.replace('Owner: the user.','Owner: the user.');
if(kind==='missing-input')delete e.tool_input.new_string;
if(kind==='over-limit')e.tool_input.new_string='new\n'.repeat(513);
if(kind==='replace-all')e.tool_input.replace_all=true;
if(kind==='before-file')fs.writeFileSync(r.file,fixture.before+'Unobserved change.');
recordAutoplanArtifact(JSON.stringify(e),r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot);
expect(autoplanArtifactRecorderStatus(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot)).toEqual({status:'invalid',reason:'conflicting_replay'});
expect(readPendingAutoplanArtifact(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot,r.context.commandStartedAt,r.context.publicTools,r.context.now)).toBeUndefined();
});
// Synthetic legacy crops use the actual generated PreToolUse subprocess. They
// preserve the frozen deletion/context policy, not new insertion-only authority.
test.each([
{name:'leading partial deletion',rows:[' -full line',' 11 -Old second',' 12 +New replacement'],removed:'First original full line\nOld second',added:'New replacement'},
{name:'leading partial context',rows:[' full line',' 11 -Old second',' 12 +New replacement'],removed:'Old second',added:'New replacement'},
{name:'deletion-only rows',rows:[' 10 -First original full line',' 11 -Old second',' 12 Context'],removed:'First original full line\nOld second\n',added:''},
{name:'old/new line numbering reset',rows:[' 10 -First original full line',' 11 -Old second',' 10 +New first',' 11 +New second',' 12 Context'],removed:'First original full line\nOld second',added:'New first\nNew second'},
])('recording a digest preserves an owned legacy $name crop',c=>{
const r=replay(false),before='First original full line\nOld second\nContext\n';
fs.writeFileSync(r.file,before);const old=new Date(Date.parse(fixture.pending.timestamp)-1000);fs.utimesSync(r.file,old,old);
for(const e of r.context.publicTools)if(e.name==='Write'&&e.input?.file_path===r.file)e.input.content=before;
r.event.tool_input.old_string=c.removed;r.event.tool_input.new_string=c.added;
const child=spawnSync('bash',['-c',r.recorder.hooks.PreToolUse[0]!.hooks[0]!.command],{input:JSON.stringify(r.event),encoding:'utf8',timeout:6000});
expect(child.status).toBe(0);expect(child.stdout).toBe('');expect(child.stderr).toBe('');
r.context.pending=readPendingAutoplanArtifact(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot,r.context.commandStartedAt,r.context.publicTools);
r.context.viewportCapturedAt=Date.now();r.context.now=Date.now()+1000;
const menu=r.viewport.slice(r.viewport.indexOf('Do you want to make this edit'));
r.viewport=c.rows.join('\n')+'\n'+'╌'.repeat(20)+'\n'+menu;
expect(validAutoplanEditDigest(r.context.pending?.editDigest)).toBe(true);
const digest=structuredClone(r.context.pending!.editDigest!);
expect(pick(r)?.input).toBe('1\r');
delete r.context.pending!.editDigest;expect(pick(r)?.input).toBe('1\r');
r.context.pending!.editDigest={...digest,beforeSHA256:'0'.repeat(64)};expect(pick(r)).toBeNull();
r.context.pending!.editDigest={...digest,beforeSHA256:'malformed'};expect(pick(r)).toBeNull();
r.context.pending!.editDigest=digest;
const viewport=r.viewport;r.viewport=r.viewport.replace(/^((?: {0,3}\d+ | {4})-).*$/gm,'$1Foreign unowned deletion');expect(pick(r)).toBeNull();r.viewport=viewport;
// The digest's request ownership remains binding through the legacy crop path.
r.context.pending!.editDigest={...digest,oldLineHashes:[autoplanEditLineHash('Context')]};expect(pick(r)).toBeNull();
r.context.pending!.editDigest=digest;
if(c.rows.some(row=>/^[ ]*\d+ \+/.test(row))){
r.viewport=viewport.replace(/^([ ]*\d+ \+).*$/gm,'$1Context');expect(pick(r)).toBeNull();r.viewport=viewport;
}
if(c.name==='leading partial deletion'){
r.viewport=viewport.replace(' -full line',' -Context');expect(pick(r)).toBeNull();r.viewport=viewport;
}
if(c.name==='leading partial context'){
r.viewport=viewport.replace(' full line',' +full line');expect(pick(r)).toBeNull();r.viewport=viewport;
}
fs.unlinkSync(r.file);expect(pick(r)).toBeNull();
});
+121
View File
@@ -0,0 +1,121 @@
import { capturedPathRebaser } from './helpers/captured-paths';
import {expect,test} from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import fixture from './fixtures/autoplan-edit-edges-an.json';
import * as permission from './helpers/autoplan-artifact-permission';
import {readPendingAutoplanArtifact} from './helpers/autoplan-artifact-recorder';
import {createAutoplanEditDigest} from './helpers/autoplan-artifact-digest';
import {readPlanCountTranscript,type NativePublicToolEvent} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
function setup(changeRecords?:(records:any[])=>void){
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-ap-edges-'));
const cwd=path.join(dir,path.basename(fixture.cwd)),config=path.join(dir,'config');
const stateRoot=path.join(dir,'gstack-hermetic-2546450-gfwm4G/skill-home-k7zGB1/.gstack');
const rebase=capturedPathRebaser([[fixture.stateRoot,stateRoot],[fixture.cwd,cwd],[fixture.config,config]]);
const hook=rebase.json(fixture.hook);
const file=hook.pending.file,nativeFile=path.join(config,'projects','owned',hook.sessionId+'.jsonl');hook.pending.transcriptPath=nativeFile;
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(file),{recursive:true});fs.mkdirSync(path.dirname(nativeFile),{recursive:true});
fs.writeFileSync(file,fixture.before);fs.utimesSync(file,new Date(fixture.now-1000000),new Date(Date.parse(hook.pending.timestamp)-1000));
const events=rebase.json(fixture.publicTools) as (NativePublicToolEvent & {messageId?:string;requestId?:string})[];
const records=events.map(e=>({sessionId:e.sessionId,cwd,isSidechain:false,timestamp:e.timestamp,requestId:e.requestId,message:{id:e.messageId,role:e.kind==='use'?'assistant':'user',content:e.kind==='use'?[{type:'tool_use',id:e.toolUseId,name:e.name,input:e.input}]:[{type:'tool_result',tool_use_id:e.toolUseId,content:'',is_error:e.isError}]}}));
changeRecords?.(records);
fs.writeFileSync(nativeFile,records.map(r=>JSON.stringify(r)).join('\n')+'\n');
const hookFile=path.join(dir,'hook.json');fs.writeFileSync(hookFile,JSON.stringify(hook)+'\n');
const publicTools:NativePublicToolEvent[]=[];const native=readPlanCountTranscript(config,cwd,e=>publicTools.push(e));
const pending=(readPendingAutoplanArtifact as any)(hookFile,cwd,config,stateRoot,fixture.commandStartedAt,publicTools,fixture.now,true);
const context={cwd,ownedStateRoot:stateRoot,commandStartedAt:fixture.commandStartedAt,now:fixture.now,viewportCapturedAt:fixture.now,transcriptStatus:native.status,publicTools,pending};
const invoke=(screen=fixture.viewport,ctx:any=context,seen=new Set<string>())=>(permission as any).publishedAutoplanArtifactPermissionInput?.(screen,ctx,seen)??null;
return {dir,cwd,config,stateRoot,hook,hookFile,file,nativeFile,publicTools,context,invoke,dispose:()=>fs.rmSync(dir,{recursive:true,force:true})};
}
test('exact published Edit keeps unchanged suffixes in complete native preview rows',()=>{
const s=setup();try{
expect(s.context.pending?.toolUseId).toBe(fixture.hook.pending.toolUseId);
expect(permission.autoplanArtifactPermissionInput(fixture.viewport,s.context,new Set())).toBeNull();
expect(permission.pendingAutoplanArtifactPermissionInput(fixture.viewport,s.context,new Set())).toBeNull();
expect(s.invoke()).toEqual({input:'1\r',signature:s.hook.sessionId+':'+s.hook.pending.toolUseId,file:s.file});
}finally{s.dispose()}
});
type Replay=ReturnType<typeof setup>;
const current=(s:Replay)=>s.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId===fixture.hook.pending.toolUseId)!;
const queued=(s:Replay)=>s.context.publicTools.filter(e=>e.kind==='use'&&e.name==='Edit'&&e.toolUseId!==fixture.hook.pending.toolUseId).at(-1)!;
function panel(s:Replay,rows:string[]){const bar='─'.repeat(120);return `${bar}\n Edit file\n ${s.file}\n${bar}\n${rows.join('\n')}\n${bar}\n Do you want to make this edit to ${path.basename(s.file)}?\n 1. Yes\n 2. Yes, and switch to accept edits (auto-approve file edits and common file commands) for this session (shift+tab)\n 3. No\n\n Esc to cancel · Tab to amend\n`;}
function request(s:Replay,before:string,old:string,replacement:string){
fs.writeFileSync(s.file,before);fs.utimesSync(s.file,new Date(0),new Date(Date.parse(s.hook.pending.timestamp)-1000));
const input=current(s).input!;input.old_string=old;input.new_string=replacement;
s.context.pending!.editDigest=createAutoplanEditDigest(s.file,old,replacement)!;
}
test('unique request edges reconstruct exact prefix, suffix, newline and file boundaries',()=>{
const cases:Array<[string,string,string,string,string[]]>=[
['both edges','prefix OLD suffix\n','OLD','NEW',[' 1 -prefix OLD suffix',' 1 +prefix NEW suffix']],
['file start','OLD suffix\n','OLD','NEW',[' 1 -OLD suffix',' 1 +NEW suffix']],
['file end','prefix OLD','OLD','NEW',[' 1 -prefix OLD',' 1 +prefix NEW']],
['line start','head\nOLD suffix\n','OLD','NEW',[' 2 -OLD suffix',' 2 +NEW suffix']],
['multiline edges','prefix first\nsecond suffix\n','first\nsecond','one\ntwo',[' 1 -prefix first',' 2 -second suffix',' 1 +prefix one',' 2 +two suffix']],
['trailing newline','prefix OLD\nnext\n','OLD\n','NEW\n',[' 1 -prefix OLD',' 1 +prefix NEW',' 2 next']],
['leading newline','head\nOLD suffix\n','\nOLD','\nNEW',[' 1 head',' 2 -OLD suffix',' 2 +NEW suffix']],
['insert newline','prefix OLD suffix\n','OLD','NEW\nNEXT',[' 1 -prefix OLD suffix',' 1 +prefix NEW',' 2 +NEXT suffix']],
['remove middle text','keep token tail\n','token ','',[' 1 -keep token tail',' 1 +keep tail']],
];
for(const [name,before,old,replacement,rows] of cases){const s=setup();try{request(s,before,old,replacement);expect(s.invoke(panel(s,rows))?.input,name).toBe('1\r');}finally{s.dispose()}}
});
test('viewport edges must be exact unchanged file bytes and cannot come from queued edits',()=>{
const s=setup();try{
expect(s.invoke(fixture.viewport.replaceAll('the envelope becomes the response','the envelope leaks a secret'))).toBeNull();
request(s,'prefix OLD suffix\n','OLD','NEW');
for(const rows of [
[' 1 -foreign OLD suffix',' 1 +foreign NEW suffix'],
[' 1 -prefix OLD forged',' 1 +prefix NEW forged'],
[' 1 -prefix OLD suffix',' 1 +prefix UNREQUESTED suffix'],
[' 1 -prefix OLD suffix',' 1 +prefix NEW suffix',' 2 +queued sibling change'],
[' 1 prefix OLD suffix',' 1 +prefix OLD suffix'],
])expect(s.invoke(panel(s,rows))).toBeNull();
// A repeated old snippet must not select an arbitrary copy even when the pane matches one.
request(s,'prefix OLD suffix\nanother OLD line\n','OLD','NEW');
expect(s.context.pending!.editDigest).toBeUndefined();
expect(s.invoke(panel(s,[' 1 -prefix OLD suffix',' 1 +prefix NEW suffix']))).toBeNull();
const direct={...s.context,publicTools:s.context.publicTools.filter(e=>e.toolUseId===current(s).toolUseId||e.kind==='result'||e.toolUseId===fixture.publicTools[0]!.toolUseId)};
expect(permission.autoplanArtifactPermissionInput(panel(s,[' 1 -prefix OLD suffix',' 1 +prefix NEW suffix']),direct,new Set())).toBeNull();
}finally{s.dispose()}
});
test('exact digest, current ownership and batch authority stay mandatory for the actual partial-line pane',()=>{
const cases:Array<[string,(s:Replay)=>void]>=[
['before digest',s=>{s.context.pending!.editDigest.beforeSHA256='0'.repeat(64)}],
['request digest',s=>{s.context.pending!.editDigest.requestSHA256='0'.repeat(64)}],
['changed file',s=>{fs.appendFileSync(s.file,'\nChanged');fs.utimesSync(s.file,new Date(0),new Date(0))}],
['changed request',s=>{current(s).input!.new_string+=' '}],
['stale hook',s=>{s.context.pending!.timestamp=new Date(fixture.commandStartedAt-1).toISOString()}],
['stale viewport',s=>{s.context.viewportCapturedAt=Date.parse(s.hook.pending.timestamp)-1}],
['foreign session',s=>{s.context.pending!.sessionId='foreign'}],
['foreign file',s=>{current(s).input!.file_path=s.file+'.other'}],
['foreign queued batch',s=>{queued(s).requestId='req_foreign'}],
['hooked queued sibling',s=>{s.context.pending!.hookSeenIds!.push(queued(s).toolUseId)}],
['no successful prior write',s=>{for(const e of s.context.publicTools)if(e.kind==='result')e.isError=true}],
['completed current request',s=>{s.context.publicTools.push({kind:'result',sessionId:s.hook.sessionId,toolUseId:current(s).toolUseId,timestamp:s.hook.pending.timestamp,isError:false})}],
];
for(const [name,change] of cases){const s=setup();try{change(s);expect(s.invoke(),name).toBeNull()}finally{s.dispose()}}
const s=setup();try{
expect(s.invoke(fixture.viewport,s.context,new Set([s.hook.sessionId+':'+s.hook.pending.toolUseId]))).toBeNull();
expect(s.invoke(fixture.viewport,s.context,new Set([permission.autoplanArtifactMenuKey(fixture.viewport)]))).toBeNull();
expect(s.invoke('Source excerpt:\n'+fixture.viewport)).toBeNull();
expect(s.invoke(fixture.viewport.split('\n').map(row=>'> '+row).join('\n'))).toBeNull();
expect(s.invoke(fixture.viewport.replace(' 1. Yes',' 2. Yes'))).toBeNull();
expect(s.invoke(fixture.viewport.replace('3. No','3. Maybe'))).toBeNull();
}finally{s.dispose()}
});
test('the partial-line fixture and tests register only the Autoplan owner densely',()=>{
const owner=E2E_TOUCHFILES['autoplan-chain-pty']!;
expect(Object.keys(owner)).toHaveLength(owner.length);
expect(Array.from(owner).every(x=>typeof x==='string')).toBe(true);
for(const file of ['test/autoplan-edit-edges-an.test.ts','test/fixtures/autoplan-edit-edges-an.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
+109
View File
@@ -0,0 +1,109 @@
import { afterEach, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { autoplanArtifactPermissionInput, pendingAutoplanArtifactPermissionInput, autoplanArtifactMenuKey } from './helpers/autoplan-artifact-permission';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/autoplan-edit-header-ag.json';
const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, {recursive:true,force:true}); });
function replay() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-edit-header-')); roots.push(root);
const cwd = path.join(root,path.basename(captured.cwd));
const ownedStateRoot = path.join(root,'home','.gstack');
const file = path.normalize(captured.pending.file.replace(captured.ownedStateRoot,ownedStateRoot));
fs.mkdirSync(cwd,{recursive:true}); fs.mkdirSync(path.dirname(file),{recursive:true});
fs.writeFileSync(file,captured.before);
const beforeTime = new Date(Date.parse(captured.pending.timestamp)-1000);
fs.utimesSync(file,beforeTime,beforeTime);
const publicTools = structuredClone(captured.events) as NativePublicToolEvent[];
for (const event of publicTools) if (event.input?.file_path === captured.pending.file) event.input.file_path = file;
const pending = {...captured.pending,file,source:'pre_tool_use' as const,tool:'Edit' as const};
const context = {cwd,ownedStateRoot,commandStartedAt:Date.parse(publicTools[0]!.timestamp)-1,
now:Date.parse(captured.viewportCapturedAt),viewportCapturedAt:Date.parse(captured.viewportCapturedAt),
transcriptStatus:'ready',publicTools,pending};
const viewport = captured.viewport.replace(/^ (…[^\n]+)$/m,' …'+file.slice(root.length+1));
return {root,file,context,viewport};
}
const pick = (r:ReturnType<typeof replay>, seen = new Set<string>()) =>
pendingAutoplanArtifactPermissionInput(r.viewport,r.context,seen);
test('the captured native edit header preserves the current owned hook and diff', () => {
const r = replay();
expect(r.context.publicTools).toHaveLength(88);
expect(autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(pick(r)).toEqual({input:'1\r',signature:captured.pending.sessionId+':'+captured.pending.toolUseId,file:r.file});
});
test('exact absolute paths, full owned suffixes and launcher-owned aliases bind the same one-time request', () => {
for (const absoluteTitle of [false,true]) for (const displayedPath of ['cropped','absolute','relative']) {
const r = replay();
if (absoluteTitle) r.viewport = r.viewport.replace(/^([●⏺] Update\()[^\n]+(?=\)$)/m,'$1'+r.file);
if (displayedPath === 'absolute') r.viewport = r.viewport.replace(/^ …[^\n]+$/m,' '+r.file);
if (displayedPath === 'relative') r.viewport = r.viewport.replace(/^ …[^\n]+$/m,' …'+path.relative(r.context.ownedStateRoot,r.file));
const result = pick(r); expect(result?.input).toBe('1\r');
expect(pick(r,new Set([result!.signature]))).toBeNull();
expect(pick(r,new Set([autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
}
});
test('a retained header does not permit unrelated, ambiguous or quoted prefix rows', () => {
const changes = [
(s:string) => s.replace('● Update(', '● Write('),
(s:string) => s.replace(/^● Update\([^\n]+\)/, '● Update(/tmp/foreign.md)'),
(s:string) => s.replace('~/.gstack/projects/', '~/.gstack/../projects/'),
(s:string) => s.replace(/(^ …[^\n]+)dashboard.md/m, '$1other.md'),
(s:string) => s.replace(/^ …[^\n]+$/m, ' …2026-09-10-user-dashboard.md'),
(s:string) => s.replace(/^ …[^\n]+$/m, ' …projects/sibling/ceo-plans/2026-09-10-user-dashboard.md'),
(s:string) => s.replace(' Edit file', ' Read file'),
(s:string) => s.replace(' Edit file', ' Run this first\n Edit file'),
(s:string) => s.replace(' Edit file', ' Edit file\n Edit file'),
(s:string) => 'Example:\n'+s,
(s:string) => '> '+s.replaceAll('\n','\n> '),
(s:string) => '```text\n'+s+'\n```',
(s:string) => s+'\nRun another action.',
(s:string) => s.replace(' 1. Yes',' 1. Yes, always allow'),
(s:string) => s.replace('to 2026-09-10-user-dashboard.md?','to sibling.md?'),
];
for (const change of changes) { const r=replay(); r.viewport=change(r.viewport); expect(pick(r),change.toString()).toBeNull(); }
});
test('framed edits retain stale, wrong-tool, foreign-path and success-history gates', () => {
const changes: Array<(r:ReturnType<typeof replay>)=>void> = [
r=>{r.context.pending.tool='Write' as 'Edit';},
r=>{r.context.pending.sessionId='foreign';},
r=>{r.context.pending.file=r.file+'.sibling';},
r=>{r.context.viewportCapturedAt=Date.parse(r.context.pending.timestamp)-1;},
r=>{r.context.pending.timestamp=new Date(r.context.now+1000).toISOString();},
r=>{r.context.publicTools.push({kind:'result',sessionId:r.context.pending.sessionId,toolUseId:r.context.pending.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:false});},
r=>{r.context.publicTools.push({kind:'use',sessionId:r.context.pending.sessionId,toolUseId:'unresolved-other',name:'Write',timestamp:new Date(r.context.now).toISOString(),input:{file_path:r.file}});},
r=>{for(const event of r.context.publicTools) if(event.kind==='result') event.isError=true;},
r=>{fs.writeFileSync(r.file,'Unrelated replacement content');},
r=>{fs.utimesSync(r.file,new Date(r.context.now+1000),new Date(r.context.now+1000));},
];
for(const change of changes) { const r=replay();change(r);expect(pick(r),change.toString()).toBeNull(); }
});
test('the same header works for fully published synthetic Edit inputs without replacing their comparison', () => {
const r = replay();
const oldString = captured.before.split('\n')[0]!;
const newString = oldString+' (revised)';
const lines = r.viewport.split('\n');
const menu = r.viewport.slice(r.viewport.indexOf(' Do you want'));
r.viewport = lines.slice(0,6).join('\n')+'\n 1 -'+oldString+'\n 1 +'+newString+'\n────────\n'+menu;
r.context.publicTools.push({kind:'use',sessionId:r.context.pending.sessionId,toolUseId:r.context.pending.toolUseId,
name:'Edit',timestamp:r.context.pending.timestamp,input:{file_path:r.file,old_string:oldString,new_string:newString}});
expect(pendingAutoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
expect(autoplanArtifactPermissionInput(r.viewport,r.context,new Set())?.input).toBe('1\r');
r.context.publicTools.at(-1)!.input!.new_string='Different unpublished replacement';
expect(autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
});
test('the new native header evidence selects only the existing Autoplan paid case', () => {
for(const file of ['test/autoplan-edit-header-ag.test.ts','test/fixtures/autoplan-edit-header-ag.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes(file)).map(([owner])=>owner)).toEqual(['autoplan-chain-pty']);
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
}
});
+100
View File
@@ -0,0 +1,100 @@
import { afterEach, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import captured from './fixtures/autoplan-edit-panel-aj.json';
import published from './fixtures/autoplan-edit-prefix-ai.json';
import { autoplanArtifactPermissionInput, pendingAutoplanArtifactPermissionInput, autoplanArtifactMenuKey } from './helpers/autoplan-artifact-permission';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, { recursive: true, force: true }); });
function replay() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-edit-panel-')); roots.push(root);
const cwd = path.join(root, path.basename(captured.cwd)), ownedStateRoot = path.join(root, 'home', '.gstack');
const file = path.normalize(captured.pending.file.replace(captured.ownedStateRoot, ownedStateRoot));
fs.mkdirSync(cwd, { recursive: true }); fs.mkdirSync(path.dirname(file), { recursive: true }); fs.writeFileSync(file, captured.before);
const time = new Date(Date.parse(captured.pending.timestamp) - 1000); fs.utimesSync(file, time, time);
const events = structuredClone(captured.events) as NativePublicToolEvent[];
for (const event of events) if (event.input?.file_path === captured.pending.file) event.input.file_path = file;
const context = { cwd, ownedStateRoot, commandStartedAt: Date.parse(events[0]!.timestamp) - 1,
now: captured.viewportCapturedAt, viewportCapturedAt: captured.viewportCapturedAt,
pending: { ...captured.pending, source: 'pre_tool_use' as const, tool: 'Edit' as const, file }, transcriptStatus: 'ready', publicTools: events };
const viewport = captured.viewport.replace(/^ …[^\n]+$/m, ' …' + path.relative(ownedStateRoot, file));
return { root, file, context, viewport };
}
const pick = (r: ReturnType<typeof replay>, seen = new Set<string>()) => pendingAutoplanArtifactPermissionInput(r.viewport, r.context, seen);
test('the exact standalone native Edit panel binds the owned current unpublished request', () => {
const r = replay();
expect(pick(r)).toEqual({ input: '1\r', signature: r.context.pending.sessionId + ':' + r.context.pending.toolUseId, file: r.file });
expect(autoplanArtifactPermissionInput(r.viewport, r.context, new Set())).toBeNull();
});
test('complete absolute, home alias and full relative suffix paths retain ownership', () => {
for (const displayed of ['absolute', 'alias', 'suffix'] as const) {
const r = replay(), relative = path.relative(r.context.ownedStateRoot, r.file).split(path.sep).join('/');
const value = displayed === 'absolute' ? r.file : displayed === 'alias' ? '~/.gstack/' + relative : '…' + relative;
r.viewport = r.viewport.replace(/^ …[^\n]+$/m, ' ' + value); expect(pick(r)?.input).toBe('1\r');
}
const crop = replay(); crop.viewport = crop.viewport.split('\n').slice(4).join('\n'); expect(pick(crop)?.input).toBe('1\r');
});
test('missing, foreign, quoted and ambiguous headers do not authorize the current file', () => {
for (const change of [
(s: string) => s.replace(/^ …[^\n]+$/m, ' /tmp/foreign.md'),
(s: string) => s.replace(/^ …[^\n]+$/m, ' …' + path.basename(captured.pending.file)),
(s: string) => s.replace(/^ …[^\n]+$/m, ' …projects/sibling/ceo-plans/' + path.basename(captured.pending.file)),
(s: string) => s.replace(' Edit file\n', ''),
(s: string) => s.replace(' Edit file', ' Read file'),
(s: string) => s.split('\n').slice(1).join('\n'),
(s: string) => s.replace(/^─+\n/, '--------\n'),
(s: string) => '> quoted panel\n' + s,
(s: string) => '```text\n' + s + '\n```',
(s: string) => s.split('\n').slice(0, 4).join('\n') + '\n' + s,
(s: string) => '● Update(/tmp/foreign.md)\n\n' + s,
(s: string) => s + '\n' + s,
]) { const r = replay(); r.viewport = change(r.viewport); expect(pick(r)).toBeNull(); }
});
test('current hook, observed time, same-file history and one-time menu remain required', () => {
const once = replay(), granted = pick(once)!;
expect(pick(once, new Set([granted.signature]))).toBeNull();
expect(pick(once, new Set([autoplanArtifactMenuKey(once.viewport)]))).toBeNull();
for (const change of [
(r: ReturnType<typeof replay>) => { r.context.pending.sessionId = 'foreign'; },
(r: ReturnType<typeof replay>) => { r.context.pending.file = r.file + '.foreign'; },
(r: ReturnType<typeof replay>) => { r.context.viewportCapturedAt = Date.parse(r.context.pending.timestamp) - 1; },
(r: ReturnType<typeof replay>) => { r.context.publicTools[1]!.isError = true; },
(r: ReturnType<typeof replay>) => { r.context.publicTools.push({ kind: 'result', sessionId: r.context.pending.sessionId, toolUseId: r.context.pending.toolUseId, timestamp: new Date(r.context.now).toISOString(), isError: false }); },
(r: ReturnType<typeof replay>) => { r.context.publicTools.push({ kind: 'use', sessionId: r.context.pending.sessionId, toolUseId: 'newer', timestamp: new Date(r.context.now).toISOString(), name: 'Write', input: { file_path: r.file } }); },
(r: ReturnType<typeof replay>) => { fs.writeFileSync(r.file, 'Foreign content'); },
(r: ReturnType<typeof replay>) => { fs.renameSync(r.file, r.file + '.target'); fs.symlinkSync(r.file + '.target', r.file); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace(' 1. Yes', ' 2. Yes'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace('3. No', '3. Maybe'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace(' 10 ', ' 0 '); },
]) { const r = replay(); change(r); expect(pick(r)).toBeNull(); }
});
test('published edits retain exact old/new content guards with the standalone presentation', () => {
const r = replay(), events = structuredClone(published.events) as NativePublicToolEvent[];
const edit = events.find(e => e.kind === 'use' && e.toolUseId === published.pending.toolUseId)!;
const oldFile = edit.input!.file_path;
const file = path.normalize((oldFile as string).replace(published.ownedStateRoot, r.context.ownedStateRoot));
const cwd = path.join(r.root, path.basename(published.cwd)); fs.mkdirSync(cwd, { recursive: true });
fs.mkdirSync(path.dirname(file), { recursive: true }); fs.writeFileSync(file, published.before);
for (const event of events) if (event.input?.file_path === oldFile) event.input.file_path = file;
const header = published.viewport.lastIndexOf('\n● Update(') + 1;
const viewport = published.viewport.slice(header).split('\n').slice(2).join('\n').replace(/^ …[^\n]+$/m, ' …' + path.relative(r.context.ownedStateRoot, file));
const context = { cwd, ownedStateRoot: r.context.ownedStateRoot, commandStartedAt: Date.parse(events[0]!.timestamp) - 1, now: Date.parse(published.viewportCapturedAt), transcriptStatus: 'ready', publicTools: events };
expect(autoplanArtifactPermissionInput(viewport, context, new Set())?.input).toBe('1\r');
const original = edit.input!.new_string; edit.input!.new_string = 'Unrelated replacement';
expect(autoplanArtifactPermissionInput(viewport, context, new Set())).toBeNull();
edit.input!.new_string = original; edit.input!.old_string = 'Unrelated original';
expect(autoplanArtifactPermissionInput(viewport, context, new Set())).toBeNull();
});
test('only Autoplan owns the standalone panel regression inputs', () => {
for (const file of ['test/autoplan-edit-panel-aj.test.ts', 'test/fixtures/autoplan-edit-panel-aj.json'])
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['autoplan-chain-pty']);
});
+121
View File
@@ -0,0 +1,121 @@
import { afterEach, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import captured from './fixtures/autoplan-edit-prefix-ai.json';
import { autoplanArtifactPermissionInput, pendingAutoplanArtifactPermissionInput, autoplanArtifactMenuKey } from './helpers/autoplan-artifact-permission';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, { recursive: true, force: true }); });
function replay() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-edit-prefix-')); roots.push(root);
const cwd = path.join(root, path.basename(captured.cwd)), ownedStateRoot = path.join(root, 'home', '.gstack');
const events = structuredClone(captured.events) as NativePublicToolEvent[];
const latest = events.find(e => e.kind === 'use' && e.toolUseId === captured.pending.toolUseId)!
const original = latest.input!.file_path as string, file = path.normalize(original.replace(captured.ownedStateRoot, ownedStateRoot));
fs.mkdirSync(cwd, { recursive: true }); fs.mkdirSync(path.dirname(file), { recursive: true }); fs.writeFileSync(file, captured.before);
const time = new Date(Date.parse(latest.timestamp) - 1000); fs.utimesSync(file, time, time);
for (const event of events) if (event.input?.file_path === original) event.input.file_path = file;
const context = { cwd, ownedStateRoot, commandStartedAt: Date.parse(events[0]!.timestamp) - 1,
now: Date.parse(captured.viewportCapturedAt), viewportCapturedAt: Date.parse(captured.viewportCapturedAt), transcriptStatus: 'ready', publicTools: events };
const viewport = captured.viewport.replace(/^ …[^\n]+$/m, ' …' + path.relative(ownedStateRoot, file));
const header = viewport.lastIndexOf('\n● Update(') + 1;
return { root, file, current: latest, context, viewport, prefix: viewport.slice(0, header), panel: viewport.slice(header) };
}
const pick = (r: ReturnType<typeof replay>, seen = new Set<string>()) => autoplanArtifactPermissionInput(r.viewport, r.context, seen);
test('the exact retained prior diff output does not hide the current published owned edit', () => {
const r = replay();
expect(r.prefix.split('\n')).toHaveLength(17);
expect(pick(r)).toEqual({ input: '1\r', signature: r.current.sessionId + ':' + r.current.toolUseId, file: r.file });
r.viewport = r.panel;
expect(pick(r)?.input).toBe('1\r');
});
test('completed diff rows are ignored only before one complete current native panel', () => {
for (const prefix of [' 1 +Previous completed output\n\n', ' +cropped prior row\n 12 +next prior row\n +wrapped row\n\n', ' 1 -Old value\n 1 +New value\n\n']) {
const r = replay(); r.viewport = prefix + r.panel; expect(pick(r)?.input).toBe('1\r');
}
});
test('competing headers, previous panels, misleading prose and quotes remain rejected', () => {
for (const prefix of [
'● Update(/tmp/foreign.md)\n\n',
'● Update(~/.gstack/projects/gstack-autoplan-chain-9599im/ceo-plans/2026-09-10-user-dashboard.md)\n ⎿ Added 1 line\n\n',
' Edit file\n /tmp/foreign.md\n────────\n',
'Example:\n', '> quoted output\n', '```diff\n 1 +quoted\n```\n',
]) { const r = replay(); r.viewport = prefix + r.viewport; expect(pick(r)).toBeNull(); }
const priorPanel = replay(); priorPanel.viewport = priorPanel.panel + '\n' + priorPanel.panel; expect(pick(priorPanel)).toBeNull();
});
test('malformed completed-output gutters cannot become a panel delimiter', () => {
for (const prefix of [' 1 +wrong indent\n', ' 0 +zero line\n', ' 9007199254740992 +unsafe line\n', ' 11 +row\n +short wrap\n', ' 11 +row\n -wrong kind\n', ' +only a cropped fragment\n']) {
const r = replay(); r.viewport = prefix + r.panel; expect(pick(r)).toBeNull();
}
});
for (const [numbered, continuation] of [
[' 7 ', ' '], [' 17 ', ' '],
[' 116 ', ' '], [' 1024 ', ' '],
] as const) test(`completed prefix ${numbered.trim()} infers one column before checking cropped and wrapped rows`, () => {
const r = replay();
const prefix = `${continuation}+leading cropped fragment\n${numbered}+Previous completed\n${continuation}+ output\n\n`;
r.viewport = prefix + r.panel;
expect(pick(r)?.input).toBe('1\r');
for (const invalid of [
prefix.replaceAll(continuation + '+', continuation.slice(1) + '+'),
prefix.replaceAll(continuation + '+', ' ' + continuation + '+'),
prefix.replace(continuation + '+ output', continuation + '- output'),
prefix + numbered.replace(/(\d+) /, '$10 ') + '+mixed column\n',
prefix.replace(numbered + '+', ' ' + numbered.trim() + ' +'),
prefix.replace(numbered + '+Previous completed\n', ''),
'Example:\n' + prefix,
]) { r.viewport = invalid + r.panel; expect(pick(r), invalid).toBeNull(); }
});
test('the complete current header, exact target, menu and requested replacement remain binding', () => {
for (const change of [
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace('● Update(~/.gstack/', '● Update(/foreign/'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace(/^ …[^\n]+$/m, ' …projects/sibling/ceo-plans/2026-09-10-user-dashboard.md'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace(' Edit file', ' Read file'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace(' 1. Yes', ' 2. Yes'); },
(r: ReturnType<typeof replay>) => { r.viewport = r.viewport.replace('3. No', '3. Maybe'); },
(r: ReturnType<typeof replay>) => { r.viewport += '\nDo another action.'; },
(r: ReturnType<typeof replay>) => { r.current.input!.new_string = 'Unrelated replacement'; },
(r: ReturnType<typeof replay>) => { fs.writeFileSync(r.file, 'Unrelated current file'); },
]) { const r = replay(); change(r); expect(pick(r)).toBeNull(); }
});
test('seen, completed, foreign or superseded native requests cannot borrow the valid panel', () => {
const once = replay(), granted = pick(once)!;
expect(pick(once, new Set([granted.signature]))).toBeNull();
for (const change of [
(r: ReturnType<typeof replay>) => { const e = r.current; r.context.publicTools.push({ kind: 'result', sessionId: e.sessionId, toolUseId: e.toolUseId, timestamp: new Date(r.context.now).toISOString(), isError: false }); },
(r: ReturnType<typeof replay>) => { r.current.sessionId = 'foreign'; },
(r: ReturnType<typeof replay>) => { r.current.name = 'Write'; },
(r: ReturnType<typeof replay>) => { r.current.input!.file_path = r.file + '.foreign'; },
(r: ReturnType<typeof replay>) => { r.context.publicTools.find(e => e.kind === 'result')!.isError = true; },
(r: ReturnType<typeof replay>) => { const e = structuredClone(r.current); e.toolUseId = 'newer-edit'; r.context.publicTools.push(e); },
]) { const r = replay(); change(r); expect(pick(r)).toBeNull(); }
});
test('metadata fallback uses the same panel boundary while published inputs stay authoritative', () => {
const r = replay(), current = r.current;
const pending = { source: 'pre_tool_use' as const, tool: 'Edit' as const, sessionId: current.sessionId, toolUseId: current.toolUseId, timestamp: captured.pending.timestamp, file: r.file };
expect(pendingAutoplanArtifactPermissionInput(r.viewport, { ...r.context, pending }, new Set())).toBeNull();
// Synthetic missing-publication projection; actual AI request was published.
r.context.publicTools = r.context.publicTools.filter(e => e.toolUseId !== current.toolUseId);
const context = { ...r.context, pending };
expect(pendingAutoplanArtifactPermissionInput(r.viewport, context, new Set())?.input).toBe('1\r');
expect(pendingAutoplanArtifactPermissionInput(r.viewport, context, new Set([autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
expect(pendingAutoplanArtifactPermissionInput(r.viewport, { ...context, viewportCapturedAt: Date.parse(pending.timestamp) - 1 }, new Set())).toBeNull();
r.viewport = 'Example:\n' + r.viewport;
expect(pendingAutoplanArtifactPermissionInput(r.viewport, context, new Set())).toBeNull();
});
test('the exact prefix fixture and controls select only Autoplan', () => {
for (const file of ['test/autoplan-edit-prefix-ai.test.ts', 'test/fixtures/autoplan-edit-prefix-ai.json'])
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['autoplan-chain-pty']);
});
+189
View File
@@ -0,0 +1,189 @@
import { capturedPathRebaser } from './helpers/captured-paths';
import {expect,test} from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import fixture from './fixtures/autoplan-edit-queue-am.json';
import * as permission from './helpers/autoplan-artifact-permission';
import {readPendingAutoplanArtifact} from './helpers/autoplan-artifact-recorder';
import {readPlanCountTranscript,type NativePublicToolEvent} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
function setup(changeRecords?:(records:any[])=>void){
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-ap-queue-'));
const cwd=path.join(dir,path.basename(fixture.cwd)),config=path.join(dir,'config');
const stateRoot=path.join(dir,'gstack-hermetic-2101964-HvDZyN/skill-home-zgCNxG/.gstack');
const rebase=capturedPathRebaser([[fixture.stateRoot,stateRoot],[fixture.cwd,cwd],[fixture.config,config]]);
const hook=rebase.json(fixture.hook);
const file=hook.pending.file,nativeFile=path.join(config,'projects','owned',hook.sessionId+'.jsonl');hook.pending.transcriptPath=nativeFile;
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(file),{recursive:true});fs.mkdirSync(path.dirname(nativeFile),{recursive:true});
fs.writeFileSync(file,fixture.before);fs.utimesSync(file,new Date(fixture.now-1000000),new Date(Date.parse(hook.pending.timestamp)-1000));
const events=rebase.json(fixture.publicTools) as (NativePublicToolEvent & {messageId?:string;requestId?:string})[];
const records=events.map(e=>({sessionId:e.sessionId,cwd,isSidechain:false,timestamp:e.timestamp,requestId:e.requestId,message:{id:e.messageId,role:e.kind==='use'?'assistant':'user',content:e.kind==='use'?[{type:'tool_use',id:e.toolUseId,name:e.name,input:e.input}]:[{type:'tool_result',tool_use_id:e.toolUseId,content:'',is_error:e.isError}]}}));
changeRecords?.(records);
fs.writeFileSync(nativeFile,records.map(r=>JSON.stringify(r)).join('\n')+'\n');
const hookFile=path.join(dir,'hook.json');fs.writeFileSync(hookFile,JSON.stringify(hook)+'\n');
const publicTools:NativePublicToolEvent[]=[];const native=readPlanCountTranscript(config,cwd,e=>publicTools.push(e));
const pending=(readPendingAutoplanArtifact as any)(hookFile,cwd,config,stateRoot,fixture.commandStartedAt,publicTools,fixture.now,true);
const context={cwd,ownedStateRoot:stateRoot,commandStartedAt:fixture.commandStartedAt,now:fixture.now,viewportCapturedAt:fixture.now,transcriptStatus:native.status,publicTools,pending};
const invoke=(screen=fixture.viewport,ctx:any=context,seen=new Set<string>())=>(permission as any).publishedAutoplanArtifactPermissionInput?.(screen,ctx,seen)??null;
return {dir,cwd,config,stateRoot,hook,hookFile,file,nativeFile,publicTools,context,invoke,dispose:()=>fs.rmSync(dir,{recursive:true,force:true})};
}
test('the actual active hook binds its published request amid later queued edits and both native prefix forms',()=>{
const s=setup();try{
expect(permission.autoplanArtifactPermissionInput(fixture.viewport,s.context,new Set())).toBeNull();
expect(permission.pendingAutoplanArtifactPermissionInput(fixture.viewport,s.context,new Set())).toBeNull();
expect(s.context.pending?.toolUseId).toBe(fixture.hook.pending.toolUseId);
expect(s.invoke()?.signature).toBe(`${fixture.hook.sessionId}:${fixture.hook.pending.toolUseId}`);
expect(s.invoke()?.file).toBe(s.file);
expect(s.invoke()?.input).toBe('1\r');
}finally{s.dispose()}
});
test('the default metadata-only reader continues excluding a published request',()=>{
const s=setup();try{expect(readPendingAutoplanArtifact(s.hookFile,s.cwd,s.config,s.stateRoot,fixture.commandStartedAt,s.publicTools,fixture.now)).toBeUndefined();}finally{s.dispose()}
});
type Replay=ReturnType<typeof setup>;
const current=(s:Replay)=>s.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId===fixture.hook.pending.toolUseId)!;
const queued=(s:Replay)=>s.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId==='toolu_01SYiANcdq3hLqGxEhDQVNJf')!;
function rejects(cases:Array<[string,(s:Replay)=>void]>){
for(const [name,change] of cases){const s=setup();try{change(s);expect(s.invoke(),name).toBeNull()}finally{s.dispose()}}
}
test('only exact native message and request identifiers establish queued membership',()=>{
const s=setup();try{
expect(current(s).messageId).toBe('msg_011CeuYDnRH9L1Qoom8gBVdc');
expect(current(s).requestId).toBe('req_011CeuYDk9cAd6Yozh8QnH62');
expect(queued(s).messageId).toBe(current(s).messageId);
}finally{s.dispose()}
for(const change of [
(r:any)=>{delete r.message.id},(r:any)=>{delete r.requestId},
(r:any)=>{r.message.id='quoted msg_example'},(r:any)=>{r.requestId='req_'+ 'a'.repeat(161)},
]){const s=setup(records=>{for(const r of records)if(r.message.content[0]?.id===fixture.hook.pending.toolUseId)change(r)});try{
expect(current(s).messageId).toBeUndefined();expect(current(s).requestId).toBeUndefined();expect(s.invoke()).toBeNull();
}finally{s.dispose()}}
});
test.each(['one native record','equal timestamps'])('ordered later blocks in %s remain queued behind the current hook',shape=>{
const ids=['toolu_01SYiANcdq3hLqGxEhDQVNJf','toolu_01LgaibBToDfuxGNFBKew9PS','toolu_01VqJFXfD5cfdjiar1gAjpkV'];
const s=setup(records=>{
const active=records.find(r=>r.message.content[0]?.id===fixture.hook.pending.toolUseId)!;
for(let i=records.length-1;i>=0;i--){const r=records[i];if(!ids.includes(r.message.content[0]?.id))continue;
if(shape==='one native record'){active.message.content.splice(1,0,r.message.content[0]);records.splice(i,1)}
else r.timestamp=active.timestamp;
}
});try{
const active=current(s),remaining=s.publicTools.filter(e=>e.kind==='use'&&ids.includes(e.toolUseId));
expect(remaining.map(e=>e.toolUseId)).toEqual(ids);
expect(remaining.every(e=>e.timestamp===active.timestamp&&e.messageId===active.messageId&&e.requestId===active.requestId)).toBe(true);
expect(s.invoke()?.signature).toBe(s.hook.sessionId+':'+s.hook.pending.toolUseId);
// Moving a same-time unresolved block ahead of the current request is not a queued successor.
const earlier=remaining[0]!,events=s.context.publicTools;events.splice(events.indexOf(earlier),1);events.splice(events.indexOf(active),0,earlier);
expect(s.invoke()).toBeNull();
}finally{s.dispose()}
});
test('another batch, session, path, tool, malformed edit or already hooked successor cannot be ignored',()=>{
rejects([
['foreign message',s=>{queued(s).messageId='msg_other'}],
['foreign request',s=>{queued(s).requestId='req_other'}],
['missing message',s=>{delete queued(s).messageId}],
['foreign session',s=>{queued(s).sessionId='foreign'}],
['foreign file',s=>{queued(s).input!.file_path=s.file+'.other'}],
['queued Write',s=>{queued(s).name='Write'}],
['empty old request',s=>{queued(s).input!.old_string=''}],
['missing replacement',s=>{delete queued(s).input!.new_string}],
['replace all',s=>{queued(s).input!.replace_all=true}],
['already hooked',s=>{s.context.pending!.hookSeenIds!.push(queued(s).toolUseId)}],
['older unresolved',s=>{s.context.publicTools=s.context.publicTools.filter(e=>!(e.kind==='result'&&e.toolUseId==='toolu_01BbKwZ7JFFdm2FLFdcNQXPq'))}],
]);
});
test('current hook identity, completed or failed requests and ordering cannot be overridden',()=>{
rejects([
['foreign pending',s=>{s.context.pending!.sessionId='foreign'}],
['wrong current hook',s=>{s.context.pending!.toolUseId=queued(s).toolUseId}],
['missing hook',s=>{s.context.pending=undefined}],
['missing tombstones',s=>{delete s.context.pending!.hookSeenIds}],
['duplicate tombstone',s=>{s.context.pending!.hookSeenIds!.push(fixture.hook.pending.toolUseId)}],
['unseen current',s=>{s.context.pending!.hookSeenIds=[]}],
['duplicate current',s=>{const at=s.context.publicTools.indexOf(current(s));s.context.publicTools.splice(at,0,structuredClone(current(s)))}],
['completion',s=>{s.context.publicTools.push({kind:'result',sessionId:s.hook.sessionId,toolUseId:current(s).toolUseId,timestamp:s.hook.pending.timestamp,isError:false})}],
['failure',s=>{s.context.publicTools.push({kind:'result',sessionId:s.hook.sessionId,toolUseId:current(s).toolUseId,timestamp:s.hook.pending.timestamp,isError:true})}],
['completed queued',s=>{s.context.publicTools.push({kind:'result',sessionId:s.hook.sessionId,toolUseId:queued(s).toolUseId,timestamp:s.hook.pending.timestamp,isError:false})}],
['failed queued',s=>{s.context.publicTools.push({kind:'result',sessionId:s.hook.sessionId,toolUseId:queued(s).toolUseId,timestamp:s.hook.pending.timestamp,isError:true})}],
['late predecessor completion',s=>{s.context.publicTools.at(-1)!.timestamp=new Date(Date.parse(s.hook.pending.timestamp)+1).toISOString()}],
['no successful predecessor',s=>{for(const e of s.context.publicTools)if(e.kind==='result')e.isError=true}],
['out of order',s=>{s.context.publicTools.reverse()}],
['future publication',s=>{queued(s).timestamp=new Date(fixture.now+1).toISOString()}],
]);
});
test('exact digest and current before file are required independently of the visible subset',()=>{
rejects([
['missing digest',s=>{delete s.context.pending!.editDigest}],
['malformed digest',s=>{s.context.pending!.editDigest.version=2}],
['different request hash',s=>{s.context.pending!.editDigest.requestSHA256='0'.repeat(64)}],
['different before hash',s=>{s.context.pending!.editDigest.beforeSHA256='0'.repeat(64)}],
['different old lines',s=>{s.context.pending!.editDigest.oldLineHashes=['0'.repeat(64)]}],
['different new lines',s=>{s.context.pending!.editDigest.newLineHashes=['0'.repeat(64)]}],
['changed old request',s=>{current(s).input!.old_string+=' '}],
['changed replacement',s=>{current(s).input!.new_string+=' '}],
['missing current file',s=>{fs.unlinkSync(s.file)}],
['changed current file with old mtime',s=>{fs.writeFileSync(s.file,fixture.before+'\nChanged.');fs.utimesSync(s.file,new Date(0),new Date(0))}],
['file updated after hook',s=>{fs.utimesSync(s.file,new Date(fixture.now),new Date(fixture.now))}],
['stale viewport',s=>{s.context.viewportCapturedAt=Date.parse(s.hook.pending.timestamp)-1}],
['stale hook',s=>{s.context.pending!.timestamp=new Date(fixture.commandStartedAt-1).toISOString()}],
['future viewport',s=>{s.context.viewportCapturedAt=fixture.now+1}],
['unavailable native',s=>{s.context.transcriptStatus='missing'}],
]);
const s=setup();try{
expect(s.invoke(fixture.viewport,s.context,new Set([s.hook.sessionId+':'+s.hook.pending.toolUseId]))).toBeNull();
expect(s.invoke(fixture.viewport,s.context,new Set([permission.autoplanArtifactMenuKey(fixture.viewport)]))).toBeNull();
}finally{s.dispose()}
});
test('invalid, busy, foreign or ambiguous persisted hook state supplies no current authority',()=>{
for(const change of [
(s:Replay)=>{fs.writeFileSync(s.hookFile+'.invalid','{"reason":"conflicting_replay"}')},
(s:Replay)=>{fs.writeFileSync(s.hookFile+'.lock','')},
(s:Replay)=>{s.hook.pending.transcriptPath=path.join(s.dir,'foreign.jsonl');fs.writeFileSync(s.hookFile,JSON.stringify(s.hook))},
(s:Replay)=>{s.hook.pending.hookSeenIds=[];fs.writeFileSync(s.hookFile,JSON.stringify(s.hook))},
]){const s=setup();try{change(s);expect(readPendingAutoplanArtifact(s.hookFile,s.cwd,s.config,s.stateRoot,fixture.commandStartedAt,s.publicTools,fixture.now,true)).toBeUndefined()}finally{s.dispose()}}
});
test('existing prefix forms compose but cannot hide a competing title, source or malformed current panel',()=>{
const s=setup();try{
const first=fixture.viewport.indexOf('● Update('),screen=fixture.viewport.slice(first);
const titles=screen.match(/^● Update\([^\n]+\)\n/gm)!;
expect(titles).toHaveLength(4);
expect(s.invoke(screen)?.input).toBe('1\r');
expect(s.invoke('\n\n'+screen)?.input).toBe('1\r');
expect(s.invoke(fixture.viewport.replaceAll(titles[0]!,''))).toBeNull(); // A completed prefix still needs its current tool boundary.
expect(s.invoke(screen.slice(screen.indexOf('────────────────')))?.input).toBe('1\r');
let one=screen;for(let n=0;n<3;n++)one=one.replace(titles[0]!,'');
expect(s.invoke(one.trimStart())?.input).toBe('1\r');
for(const [name,changed] of [
['foreign first title',fixture.viewport.replace(titles[0]!,titles[0]!.replace('user-dashboard.md','foreign.md'))],
['quoted whole pane',fixture.viewport.split('\n').map(row=>'> '+row).join('\n')],
['source prefix','Example:\n'+fixture.viewport],
['arbitrary indented prose',' This is an example.\n'+fixture.viewport],
['competing completed panel','● Update(/tmp/foreign.md)\n'+fixture.viewport],
['broken wrap kind',fixture.viewport.replace(/^ \+/m,' -')],
['foreign displayed path',fixture.viewport.replace('…2101964-HvDZyN','…foreign')],
['wrong menu target',fixture.viewport.replace('user-dashboard.md?','foreign.md?')],
['persistent edit mode',fixture.viewport.replace(' 1. Yes',' 2. Yes')],
['malformed no',fixture.viewport.replace('3. No','3. Maybe')],
['changed addition',fixture.viewport.replace(/^( {0,3}\d+ \+).*/m,'$1A different current edit')],
])expect(s.invoke(changed),name).toBeNull();
}finally{s.dispose()}
});
test('the new queue regression files select only the Autoplan owner with dense registration',()=>{
const owner=E2E_TOUCHFILES['autoplan-chain-pty']!;
for(let i=0;i<owner.length;i++){expect(Object.hasOwn(owner,i)).toBe(true);expect(typeof owner[i]).toBe('string')}
for(const file of ['test/autoplan-edit-queue-am.test.ts','test/fixtures/autoplan-edit-queue-am.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
+137
View File
@@ -0,0 +1,137 @@
import { expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import {
buildPaidShardArgs, buildRunManifest, parseCliOptions, parseRunManifest,
planPaidShards, resolvePaidShardBudget, retriesForFiles, runPaidShard,
verifySliceResults, type PaidRunManifest, type SliceResult,
} from '../scripts/test-paid-shards';
import { AUTOPLAN_CHAIN_BUDGET as budget, assertPaidTestBudget, ALL_TIERS, PTY_LONG_MS } from './helpers/eval-budgets';
test('the one specified exception fits nested supervision and both unchanged retries', () => {
for (const ms of [budget.workMs, budget.sessionMs, budget.testMs, budget.shardMs]) {
expect(Number.isSafeInteger(ms) && ms > 0).toBe(true);
}
expect(budget.workMs).toBe(4 * PTY_LONG_MS);
expect(budget.workMs).toBeLessThan(budget.sessionMs);
expect(budget.sessionMs).toBeLessThan(budget.testMs);
expect(budget.testMs * (retriesForFiles([budget.file]) + 1) + budget.shardReserveMs).toBe(budget.shardMs);
expect(budget.shardMs + budget.ciReserveMs).toBe(budget.ciJobMs);
expect(Math.max(...Object.values(ALL_TIERS))).toBe(PTY_LONG_MS);
expect(() => assertPaidTestBudget(budget.file, budget.testMs)).not.toThrow();
for (const [file, ms] of [[budget.file, budget.testMs + 1], ['test/other.test.ts', budget.testMs],
[budget.file, Infinity], [budget.file, NaN], [budget.file, -1]] as const) {
expect(() => assertPaidTestBudget(file, ms)).toThrow('Unregistered');
}
});
test('only Autoplan receives the default exception and it cannot inflate a packed neighbor', () => {
expect(resolvePaidShardBudget([budget.file])).toEqual({ timeoutMs: budget.shardMs, source: 'registered', policyId: budget.id });
expect(resolvePaidShardBudget(['test/other.test.ts'])).toEqual({ timeoutMs: 1_800_000, source: 'default', policyId: null });
expect(() => resolvePaidShardBudget([budget.file, 'test/other.test.ts'])).toThrow('own shard');
const shards = planPaidShards(['test/a.test.ts', budget.file, 'test/z.test.ts'], { maxFilesPerShard: 3 });
expect(shards.find(files => files.includes(budget.file))).toEqual([budget.file]);
expect(shards.flat().sort()).toEqual(['test/a.test.ts', budget.file, 'test/z.test.ts'].sort());
for (const value of [NaN, Infinity, -1, 0, 1.5, 2_147_483_648]) {
expect(() => resolvePaidShardBudget([budget.file], value)).toThrow('timer-safe');
}
});
test('CLI and environment distinguish user limits from the ordinary default', () => {
const implicit = parseCliOptions([], {});
expect(implicit.timeoutMs).toBe(1_800_000);
expect(implicit.timeoutExplicit).toBe(false);
for (const explicit of [parseCliOptions(['--timeout', '12'], {}), parseCliOptions([], { EVALS_SHARD_TIMEOUT_MS: '12000' })]) {
expect(explicit.timeoutExplicit).toBe(true);
expect(resolvePaidShardBudget([budget.file], explicit.timeoutMs).timeoutMs).toBe(12_000);
}
expect(buildPaidShardArgs([budget.file], budget.shardMs, 2, retriesForFiles([budget.file])))
.toContain('--timeout=' + budget.shardMs);
expect(retriesForFiles([budget.file])).toBe(1);
expect(() => parseCliOptions(['--autoplan-slice'], {})).toThrow('--emit-plan');
});
function planned(): PaidRunManifest {
return buildRunManifest({ tier: 'periodic', sliceCount: 7, dedicatedAutoplanSlice: true,
evalsAll: true, env: { EVALS_ALL: '1' } });
}
function results(manifest: PaidRunManifest): SliceResult[] {
return Array.from({ length: manifest.sliceCount }, (_, index) => ({ version: 1, tier: manifest.tier,
sliceIndex: index + 1, sliceCount: manifest.sliceCount,
outcomes: manifest.entries.filter(e => e.status === 'planned' && e.slice === index + 1).map(e => ({
files: [e.file], status: 'passed', exitCode: 0, elapsedMs: 1, executedTests: 1, skippedTests: 0,
...(e.budget ? { budget: e.budget } : {}),
})),
}));
}
test('the seventh periodic slice isolates Autoplan and retains the full ordinary census', () => {
const manifest = planned();
const ordinary = buildRunManifest({ tier: 'periodic', sliceCount: 6, evalsAll: true, env: { EVALS_ALL: '1' } });
expect(manifest.entries.map(e => e.file)).toEqual(ordinary.entries.map(e => e.file));
expect(manifest.entries.filter(e => e.slice === 7).map(e => e.file)).toEqual([budget.file]);
expect(manifest.entries.filter(e => e.file !== budget.file && e.status === 'planned').every(e => e.slice <= 6)).toBe(true);
expect(parseRunManifest(JSON.stringify(manifest))).toEqual(manifest);
expect(verifySliceResults(manifest, results(manifest))).toEqual({ ok: true, problems: [] });
for (const mutate of [
(m: PaidRunManifest) => { m.entries = m.entries.filter(e => e.file !== budget.file); },
(m: PaidRunManifest) => { m.entries.push(m.entries.find(e => e.file === budget.file)!); },
(m: PaidRunManifest) => { m.entries.find(e => e.file === budget.file)!.slice = 1; },
(m: PaidRunManifest) => { delete m.entries.find(e => e.file === budget.file)!.budget; },
(m: PaidRunManifest) => { m.entries.find(e => e.file === budget.file)!.budget!.timeoutMs = 999; },
]) {
const invalid = structuredClone(manifest); mutate(invalid);
expect(() => parseRunManifest(JSON.stringify(invalid))).toThrow();
expect(verifySliceResults(invalid, results(manifest)).ok).toBe(false);
}
expect(verifySliceResults(manifest, results(manifest).slice(0, 6)).ok).toBe(false);
const duplicate = results(manifest); duplicate[0]!.outcomes.push(duplicate[6]!.outcomes[0]!);
expect(verifySliceResults(manifest, duplicate).ok).toBe(false);
for (const change of [
(o: SliceResult['outcomes'][number]) => { o.executedTests = 0; },
(o: SliceResult['outcomes'][number]) => { o.skippedTests = 1; },
(o: SliceResult['outcomes'][number]) => { o.exitCode = 1; },
(o: SliceResult['outcomes'][number]) => { o.files = ['test/other.test.ts', budget.file]; },
]) { const bad = results(manifest); change(bad[6]!.outcomes[0]!); expect(verifySliceResults(manifest, bad).ok).toBe(false); }
const reordered = structuredClone(manifest);
const entry = reordered.entries.find(e => e.file === budget.file)!;
entry.budget = { policyId: budget.id, source: 'registered', timeoutMs: budget.shardMs };
expect(() => parseRunManifest(JSON.stringify(reordered))).not.toThrow();
const wrongWall = results(manifest); delete wrongWall[6]!.outcomes[0]!.budget;
expect(verifySliceResults(manifest, wrongWall).ok).toBe(false);
const lower = results(manifest); lower[6]!.timeoutOverrideMs = 12000;
lower[6]!.outcomes[0]!.budget = resolvePaidShardBudget([budget.file], 12000);
expect(verifySliceResults(manifest, lower).ok).toBe(true);
});
test('a real fake subprocess records the chosen wall and obeys an explicit shorter deadline', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-wall-'));
try {
const common = { jobs: 1, log: () => {}, logDir: dir,
env: { ...process.env, GSTACK_CLAUDE_CLI_VERSION: 'fixture-no-cli' } };
const pass = await runPaidShard([budget.file], 1, 1, { ...common,
commandFor: () => ({ command: process.execPath, args: ['-e', 'console.log(" 1 pass\\n 0 fail\\nRan 1 tests across 1 files. [1ms]")'] }) });
expect(pass.status).toBe('passed');
expect(pass.budget).toEqual(resolvePaidShardBudget([budget.file]));
const start = Date.now();
const stopped = await runPaidShard([budget.file], 1, 1, { ...common, timeoutMs: 150,
commandFor: () => ({ command: process.execPath, args: ['-e', 'setInterval(()=>{},1000)'] }) });
expect(stopped.status).toBe('timed-out');
expect(stopped.budget).toEqual(resolvePaidShardBudget([budget.file], 150));
expect(Date.now() - start).toBeLessThan(5000);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
}, 10_000);
// Wiring is execution policy: a planner-only seventh slice would silently leave
// the long case unexecuted, or a smaller job cap would preempt both attempts.
test('periodic CI allocates and executes the dedicated seventh slice inside its existing cap', () => {
const yaml = fs.readFileSync(path.resolve(import.meta.dir, '../.github/workflows/evals-periodic.yml'), 'utf8');
expect(yaml).toMatch(/--emit-plan[^\n]+--slices 7 --autoplan-slice/);
const slices = yaml.split(' eval-slices:')[1]!.split('\n report:')[0]!;
expect(slices).toContain('slice: [1, 2, 3, 4, 5, 6, 7]');
expect(slices).toContain('timeout-minutes: 200');
expect(slices).toContain('EVALS_JOBS: "2"');
expect(slices).toContain('--plan /tmp/paid-plan/manifest.json --slice ${{ matrix.slice }}');
});
+176
View File
@@ -0,0 +1,176 @@
import {expect, test} from 'bun:test';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
import {autoplanBlockingQuestionBoundary, autoplanSetupDecision} from './helpers/autoplan-setup-question';
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
import {readPendingQuestion, createPendingQuestionRecorder, recordPendingQuestion} from './helpers/plan-count-pending-question';
import {readPlanCountTranscript} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES} from './helpers/touchfiles';
import capture from './fixtures/autoplan-final-gate-ao.json';
const fixture = (): {screen:string;context:Parameters<typeof autoplanBlockingQuestionBoundary>[1]} => ({screen:capture.screen, context:{commandStartedAt:capture.commandStartedAt,viewportCapturedAt:capture.observedAt,
transcript:structuredClone(capture.transcript),publicTools:[structuredClone(capture.gateUse)]}});
const detect = (f=fixture()) => autoplanBlockingQuestionBoundary(f.screen,f.context);
const gateCall = (f:ReturnType<typeof fixture>) => f.context.transcript.calls.find(c => c.toolUseId===capture.call.toolUseId)!;
function rebind(f:ReturnType<typeof fixture>) { f.context.publicTools[0]!.input!.questions=structuredClone(gateCall(f).questions); }
test('exact AO unanswered gate stops observation but supplies no missing phase or approval', () => {
const f=fixture();const before=JSON.stringify(f);
expect(detect(f)).toEqual({sessionId:capture.call.sessionId,toolUseId:capture.call.toolUseId,source:'native'});
expect(autoplanSetupDecision(f.screen,new Set(),gateCall(f)).kind).toBe('unrelated');
// The separate dash repair recognizes DX; recorded original hits stay historical.
expect(autoplanPhaseCompletions(f.context.transcript,capture.commandStartedAt)).toEqual([
...capture.hits,{phase:2.5,ts:1789042284933},
]);
expect(capture.hits.map(h=>h.phase)).toEqual([1,2]);
expect(JSON.stringify(f)).toBe(before);
expect(gateCall(f).answered).toBe(false);
});
test('native question identity, status, chronology and current project are mandatory', () => {
const controls: Array<(f:ReturnType<typeof fixture>)=>void> = [
f=>{f.context.transcript.status='missing';}, f=>{f.context.transcript.status='error';},
f=>{f.context.publicTools=[];}, f=>{f.context.publicTools[0]!.timestamp='invalid';},
f=>{f.context.commandStartedAt=Date.parse(capture.gateUse.timestamp)+1;},
f=>{f.context.viewportCapturedAt=Date.parse(capture.gateUse.timestamp)-1;},
f=>{f.context.publicTools[0]!.sessionId='foreign';}, f=>{f.context.publicTools[0]!.toolUseId='foreign';},
f=>{f.context.publicTools[0]!.name='Read';},
f=>{f.context.publicTools[0]!.input!.questions=[null];},
f=>{f.context.publicTools[0]!.input!.questions=[{header:'Approval',question:'Partial'}];}, f=>{f.context.publicTools[0]!.input!.questions=[];},
f=>{f.context.publicTools.push(structuredClone(f.context.publicTools[0]!));},
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:false} as any);},
f=>{gateCall(f).answered=true;}, f=>{gateCall(f).failed=true;},
f=>{gateCall(f).sessionId='foreign';}, f=>{gateCall(f).questions[0]!.multiSelect=true;},
f=>{gateCall(f).questions.push(structuredClone(gateCall(f).questions[0]!));},
f=>{f.context.transcript.calls.push({...structuredClone(gateCall(f)),toolUseId:'other'});},
f=>{f.context.commandStartedAt=NaN;},
];
for(const [i,change] of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
});
function render(f:ReturnType<typeof fixture>) {
const q=gateCall(f).questions[0]!;rebind(f);
f.screen=`${q.header}\n\n${q.question.split('\n').map(s=>'│ '+s).join('\n')}\n\n`+
q.options.map((o,i)=>`${i===0?' ': ' '}${i+1}. ${o.label}\n${o.description?.split('\n').map(s=>' '+s).join('\n')??''}`).join('\n')+
`\n ${q.options.length+1}. Type something.\n ${q.options.length+2}. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
}
test('copied, stale and incomplete displays do not prove a current blocking question', () => {
for(const change of [
(f:ReturnType<typeof fixture>)=>{f.screen='Source panel:\n'+f.screen;},
f=>{f.screen='Example:\n'+f.screen;}, f=>{f.screen='```text\n'+f.screen+'\n```';},
f=>{f.screen=f.screen.split('\n').map(row=>'> '+row).join('\n');},
f=>{f.screen=f.screen.split('\n').map(row=>' '+row).join('\n');},
f=>{f.screen+='\nContinuing the review.';}, f=>{f.screen=f.screen.replace('Esc to cancel','Esc to');},
f=>{f.screen=f.screen.replace(' 6. Chat about this','');},
f=>{f.screen=f.screen.replace('4. Revise the plan or reject','4. Unmatched current choice');},
f=>{f.screen=f.screen.replace('D2 — Final Approval','D3 — Final Approval');},
f=>{f.screen=f.screen.replace(' 1.',' 1.');},
]){const f=fixture();change(f);expect(detect(f)).toBeNull();}
});
test('an actual current human wait remains blocking regardless of source or withdrawn body semantics', () => {
for(const change of [
(q:any)=>{q.question=q.question.replace('ELI10: ','ELI10: Source excerpt, not a current assessment: ');},
(q:any)=>{q.question=q.question.replace('ELI10: ','ELI10: If approved, ');},
(q:any)=>{q.question=q.question.replace('\nELI10:','\nSource excerpt:\nELI10:');},
(q:any)=>{q.question+='\nThis final approval gate is cancelled.';},
(q:any)=>{q.question+=' This approval gate is withdrawn.';},
(q:any)=>{q.question+='\nCorrection: this final gate is not current.';},
(q:any)=>{q.question+='\n> Historical note: the old gate was cancelled.';},
(q:any)=>{q.question='Choose one of these approaches?';q.header='Approach';},
(q:any)=>{q.question=q.question.replace(/^D2 /,'D9 ');},
(q:any)=>{q.options[0].label='Start implementation';},
]){const f=fixture();change(gateCall(f).questions[0]);render(f);expect(detect(f)?.source).toBe('native');}
});
test('validated owned pending-hook fallback retains stale/foreign/completed rejection', () => {
const root=fs.mkdtempSync(path.join(os.tmpdir(),'autoplan-final-gate-'));
const cwd=path.join(root,path.basename(capture.cwd)),config=path.join(root,'config');
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.join(config,'projects','owned'),{recursive:true});
const transcriptPath=path.join(config,'projects','owned',capture.call.sessionId+'.jsonl');fs.writeFileSync(transcriptPath,'');
const recorder=createPendingQuestionRecorder(cwd,config),startedAt=Date.now()-10;
const transcript:any={status:'ready',calls:[],assistantMessages:[{sessionId:capture.call.sessionId,timestamp:new Date(startedAt).toISOString(),text:'Finishing this review.'}]};
const event={hook_event_name:'PreToolUse',cwd,session_id:capture.call.sessionId,tool_name:'AskUserQuestion',tool_use_id:capture.call.toolUseId,transcript_path:transcriptPath,tool_input:{questions:capture.call.questions}};
try{
recordPendingQuestion(JSON.stringify(event),recorder.file,cwd,config);
const get=(t=transcript,cwdArg=cwd,start=startedAt)=>readPendingQuestion(recorder.file,cwdArg,config,start,t);
const check=(pending=get(),t=transcript)=>autoplanBlockingQuestionBoundary(capture.screen,{commandStartedAt:startedAt,viewportCapturedAt:Date.now(),transcript:t,publicTools:[],pending});
expect(check()?.source).toBe('pre_tool_use');
expect(get(transcript,cwd+'-foreign')).toBeUndefined();
expect(get(transcript,cwd,Date.now()+1000)).toBeUndefined();
expect(get({...transcript,assistantMessages:[{...transcript.assistantMessages[0],sessionId:'foreign'}]})).toBeUndefined();
for(const failed of [false,true]){
const completed={...transcript,calls:[{...capture.call,answered:!failed,failed}]};
expect(get(completed)).toBeUndefined();expect(check(undefined,completed)).toBeNull();
}
recordPendingQuestion(JSON.stringify({...event,hook_event_name:'PostToolUse'}),recorder.file,cwd,config);
expect(get()).toBeUndefined();expect(check()).toBeNull();
// The native route consumes the same cwd-scoped public reader as production.
// Only this local test envelope is synthetic; question bytes stay exact.
const record={cwd,sessionId:capture.call.sessionId,isSidechain:false,timestamp:new Date().toISOString(),
message:{role:'assistant',content:[{type:'tool_use',id:capture.call.toolUseId,name:'AskUserQuestion',input:{questions:capture.call.questions}}]}};
const native=(owner=cwd)=>{
const events:any[]=[];const transcript=readPlanCountTranscript(config,owner,e=>events.push(e));
return autoplanBlockingQuestionBoundary(capture.screen,{commandStartedAt:startedAt,viewportCapturedAt:Date.now(),transcript,publicTools:events});
};
fs.writeFileSync(transcriptPath,JSON.stringify(record)+'\n');
expect(native()?.source).toBe('native');expect(native(cwd+'-foreign')).toBeNull();
fs.writeFileSync(transcriptPath,JSON.stringify({...record,isSidechain:true})+'\n');expect(native()).toBeNull();
}finally{recorder.dispose();fs.rmSync(root,{recursive:true,force:true});}
});
test('production loop fails without answering; allowed and repeated setup keep their old behavior', async () => {
const source=fs.readFileSync(path.join(import.meta.dir,'skill-e2e-autoplan-chain.test.ts'),'utf8');
const begin=source.indexOf(' // This new repository offers routing');
const end=source.indexOf('\n }\n } finally',begin);
const block=source.slice(begin,end);expect(begin).toBeGreaterThan(0);expect(end).toBeGreaterThan(begin);
const AsyncFunction=Object.getPrototypeOf(async()=>{}).constructor;
const loop=new AsyncFunction('autoplanBlockingQuestionBoundary','autoplanSetupDecision','ctx',
new Bun.Transpiler({loader:'ts'}).transformSync(`async function observeBoundedLoop(){
const {methodologyAudit,hits,commandStartedAt,viewportCapturedAt,transcript,publicTools,pendingSetupQuestion,panes}=ctx;
let outcome='timeout',evidence='',blockedQuestion=null,unsupportedSetup=null;
const inputs=[],seenSetupQuestions=new Set(),session={send:(s)=>inputs.push(s)},Bun={sleep:async()=>{}};
const selectPtyNumberedOption=async(_session,n)=>session.send(String(n)+'\\r'),isPlanReadyVisible=()=>false;
for(const visible of panes){const viewport=visible;${block}}
return {outcome,blockedQuestion,hits,inputs};}`)+'return observeBoundedLoop();');
const f=fixture(),ctx={...f.context,panes:[f.screen,f.screen],hits:structuredClone(capture.hits),methodologyAudit:['ceo','design','dx','eng'].map(phase=>({phase,passed:true}))};
const run=(x=ctx)=>loop(autoplanBlockingQuestionBoundary,autoplanSetupDecision,x);
const result=await run();expect(result).toMatchObject({outcome:'blocked_on_question',hits:capture.hits,inputs:[]});
expect(await run({...ctx,methodologyAudit:[{phase:'eng',passed:false}]})).toMatchObject({outcome:'incomplete_methodology',inputs:[]});
expect(await run({...ctx,publicTools:[]})).toMatchObject({outcome:'timeout',inputs:[]});
const partial=f.screen.replace('Esc to cancel','Esc to');
expect(await run({...ctx,panes:[partial,partial]})).toMatchObject({outcome:'timeout',inputs:[]});
expect(await run({...ctx,panes:[partial,f.screen]})).toMatchObject({outcome:'blocked_on_question',inputs:[]});
const setup=fixture(),q=gateCall(setup).questions[0]!;
q.header='Routing';q.question='Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>';
q.options=[{label:'Add routing rules (Recommended)',description:'Add project routing.'},{label:'Skip, invoke manually',description:'Keep manual invocation.'}];render(setup);
expect(autoplanSetupDecision(setup.screen,new Set(),gateCall(setup))).toMatchObject({kind:'input',input:'1'});
const repeated=await run({...ctx,...setup.context,panes:[setup.screen,setup.screen]});
expect(repeated).toMatchObject({outcome:'timeout',blockedQuestion:null,inputs:['1']});
const complete=[1,2,2.5,3].map((phase,index)=>({phase,ts:capture.commandStartedAt+index+1}));
expect(await run({...ctx,hits:complete})).toMatchObject({outcome:'chain_complete',inputs:[]});
const errorStart=source.indexOf(" if (outcome === 'blocked_on_question')");
const errorEnd=source.indexOf(" if (outcome === 'exited'",errorStart);
const throwBlocked=new Function('outcome','hits','blockedQuestion','transcript','artifacts','evidence',
new Bun.Transpiler({loader:'ts'}).transformSync(source.slice(errorStart,errorEnd)));
expect(()=>throwBlocked(result.outcome,result.hits,result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('missing phase markers=[2.5,3]');
expect(()=>throwBlocked('blocked_on_question',[...ctx.hits,{phase:2.5,ts:capture.observedAt-1}],result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('missing phase markers=[3]');
// Even an impossible caller state with all markers cannot turn this disposition into success.
expect(()=>throwBlocked('blocked_on_question',complete,result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('outcome=blocked_on_question');
const validation=source.slice(source.indexOf(' // Phase 3 (Eng) MUST have been seen.'),source.indexOf(' } finally {\n try { fs.rmSync(tempDir',source.indexOf(' // Phase 3 (Eng) MUST have been seen.')));
const validate=new Function('hits','methodologyAudit','expect','transcript','artifacts','evidence',new Bun.Transpiler({loader:'ts'}).transformSync(validation));
const check=(hits:any[],audit=ctx.methodologyAudit)=>validate(hits,audit,expect,f.context.transcript,{},'Retained final gate');
expect(()=>check(ctx.hits)).toThrow('Required phase markers missing');expect(()=>check(complete)).not.toThrow();
expect(()=>check(complete,[])).toThrow();
expect(()=>check(complete.map(h=>h.phase===2.5?{...h,ts:capture.commandStartedAt+10}:h))).toThrow();
});
test('only the Autoplan owner adds the exact fixtures and every indexed entry stays dense', () => {
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/autoplan-final-gate-ao.test.ts');
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-final-gate-ao.json');
for(const paths of Object.values(E2E_TOUCHFILES))for(let i=0;i<paths.length;i++){
expect(Object.hasOwn(paths,i)).toBe(true);expect(typeof paths[i]).toBe('string');
}
});
+243
View File
@@ -0,0 +1,243 @@
import { afterEach, expect, test } from 'bun:test';
import { createHash } from 'node:crypto';
import { existsSync, linkSync, mkdirSync, mkdtempSync, readFileSync, readdirSync, rmSync, statSync, symlinkSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import { spawnSync } from 'node:child_process';
import { prepareMethodology } from '../bin/gstack-autoplan-snapshot';
const ROOT = resolve(import.meta.dir, '..');
const TOOL = join(ROOT, 'bin/gstack-autoplan-snapshot.ts');
const original = readFileSync(join(ROOT, 'test/fixtures/plans/autoplan-dashboard.md'));
const hash = (value: Buffer | string) => createHash('sha256').update(value).digest('hex');
const owned: string[] = [];
function fixture() {
const dir = mkdtempSync(join(tmpdir(), 'gstack-autoplan-init-')); owned.push(dir);
const source = join(dir, 'source plan.md');
const active = join(dir, 'assigned plan.md');
const restore = join(dir, 'restore point.md');
writeFileSync(source, original);
return { dir, source, active, restore };
}
function cli(...args: string[]) {
if (args[0] === 'create' && args.length === 4) args.push(prepareMethodology(args[1]!, join(ROOT, `plan-${args[1] === 'dx' ? 'devex' : args[1]}-review`, 'SKILL.md'), args[3]!).methodologyPath);
return spawnSync(process.execPath, [TOOL, ...args], {
cwd: ROOT, encoding: 'utf8', timeout: 10_000, maxBuffer: 1024 * 1024,
});
}
function invoke(...args: string[]) {
const result = cli(...args);
expect(result.error).toBeUndefined();
expect(result.status, result.stderr).toBe(0);
return JSON.parse(result.stdout);
}
afterEach(() => { for (const dir of owned.splice(0)) rmSync(dir, { recursive: true, force: true }); });
test('actual raw R input initializes before scope and reaches a complete CEO dispatch payload', () => {
const f = fixture();
expect(original.length).toBe(4607);
expect(hash(original)).toBe('2fdf0ece590925869fe25ae941301894f8f4505da6674302df25d5c4546fddbc');
const missing = cli('scope', f.source);
expect(missing.status).toBe(1);
expect(missing.stderr).toContain('Expected one Implementation plan section');
const initialized = invoke('init', f.source, f.source, f.restore);
expect(initialized.activePlan).toBe(f.source);
expect(initialized.originalBytes).toBe(4607);
expect(initialized.originalSha256).toBe(hash(original));
expect(readFileSync(f.restore)).toEqual(original);
expect(initialized.scope.sha256).toBe(hash(original));
expect(initialized.scope.matchCount).toBe(21);
expect(initialized.scope.dxRequired).toBe(true);
expect(invoke('scope', f.source)).toEqual(initialized.scope);
const ceo = invoke('create', 'ceo', f.source, f.restore);
expect(readFileSync(ceo.snapshotPath)).toEqual(original);
expect(ceo.nativePrompt.endsWith(original.toString())).toBe(true);
expect(ceo.nativePrompt).toContain('Mutations already require CSRF tokens');
expect(ceo.sha256).toBe(initialized.scope.sha256);
writeFileSync(f.source, readFileSync(f.source, 'utf8') + '<!-- autoplan-accepted:ceo -->\nNone: Initialization only; no review decisions yet.\n<!-- /autoplan-accepted:ceo -->\n');
expect(invoke('check', 'ceo', f.source, ceo.snapshotPath, 'unchanged').changed).toBe(false);
});
test('assigned active path preserves the separate original source and exact restore', () => {
for (const emptyAssigned of [false, true]) {
const f = fixture();
if (emptyAssigned) writeFileSync(f.active, '');
const sourceMtime = statSync(f.source).mtimeMs;
const initialized = invoke('init', f.source, f.active, f.restore);
expect(initialized.activePlan).toBe(f.active);
expect(initialized.restorePath).toBe(f.restore);
expect(readFileSync(f.source)).toEqual(original);
expect(statSync(f.source).mtimeMs).toBe(sourceMtime);
expect(readFileSync(f.restore)).toEqual(original);
expect(readFileSync(f.active, 'utf8')).toContain('## Implementation plan\n' + original.toString());
expect(readdirSync(f.dir).sort()).toEqual(['assigned plan.md', 'restore point.md', 'source plan.md']);
}
});
test('the observed missing harness plans directory is initialized without a separate mkdir step', () => {
const f = fixture();
const active = join(f.dir, 'harness', 'plans', 'assigned.md');
const restore = join(f.dir, 'state', 'project', 'restore.md');
const initialized = invoke('init', f.source, active, restore);
expect(initialized.activePlan).toBe(active);
expect(initialized.restorePath).toBe(restore);
expect(initialized.scope.dxRequired).toBe(true);
expect(readFileSync(f.source)).toEqual(original);
expect(readFileSync(restore)).toEqual(original);
expect(invoke('init', f.source, active, restore).reused).toBe(true);
});
test('same initialization is idempotent but later amendments never get reset', () => {
for (const separate of [false, true]) {
const f = fixture(); const active = separate ? f.active : f.source;
invoke('init', f.source, active, f.restore);
const bytes = readFileSync(active); const backup = readFileSync(f.restore);
const stamp = statSync(active).mtimeMs; const backupStamp = statSync(f.restore).mtimeMs;
expect(invoke('init', f.source, active, f.restore).reused).toBe(true);
expect(readFileSync(active)).toEqual(bytes);
expect(statSync(active).mtimeMs).toBe(stamp);
expect(statSync(f.restore).mtimeMs).toBe(backupStamp);
writeFileSync(active, bytes.toString() + 'Accepted review decision.\n');
const changed = readFileSync(active);
const retry = cli('init', f.source, active, f.restore);
expect(retry.status).toBe(1);
expect(retry.stderr).toContain('Existing restore does not match');
expect(readFileSync(active)).toEqual(changed);
expect(readFileSync(f.restore)).toEqual(backup);
}
});
test('already structured input keeps the review record separate from blind inputs', () => {
const f = fixture();
const body = 'API and endpoint.\n> ## Review record\n```text\n## Review record\n```\n';
const source = `# Earlier plan\n## Implementation plan\n${body}## Review record\nPrivate earlier review\n`;
writeFileSync(f.source, source);
const initialized = invoke('init', f.source, f.active, f.restore);
expect(readFileSync(f.source, 'utf8')).toBe(source);
expect(readFileSync(f.restore, 'utf8')).toBe(source);
expect(readFileSync(f.active, 'utf8')).toContain('## Review record\nPrivate earlier review');
expect(initialized.scope.sha256).toBe(hash(body));
const ceo = invoke('create', 'ceo', f.active, f.restore);
expect(readFileSync(ceo.snapshotPath, 'utf8')).toBe(body);
expect(ceo.nativePrompt).not.toContain('Private earlier review');
});
test('invalid or ambiguous input fails without publishing any backup or active plan', () => {
for (const source of [Buffer.from(''), Buffer.from([0xff, 0xfe]),
Buffer.from('## Implementation plan\nMissing review boundary\n'),
Buffer.from('## Review record\nPrivate audit\n'),
Buffer.from('```text\nUnclosed fence\n'),
Buffer.from('## Implementation plan\nAPI\n## Review record\nAudit\n## Review record\nAgain\n')]) {
const f = fixture(); writeFileSync(f.source, source);
const result = cli('init', f.source, f.active, f.restore);
expect(result.status).toBe(1);
expect(result.stdout).toBe('');
expect(readFileSync(f.source)).toEqual(source);
expect(readdirSync(f.dir)).toEqual(['source plan.md']);
}
});
test('existing foreign destinations, symlinks and path aliases cannot be overwritten', () => {
for (const kind of ['active-content', 'restore-content', 'active-link', 'restore-link', 'hardlink', 'restore-is-source', 'same-destinations', 'relative']) {
const f = fixture(); let active = f.active; let restore = f.restore;
if (kind === 'active-content') writeFileSync(active, 'Other assigned work');
if (kind === 'restore-content') writeFileSync(restore, 'Existing history');
if (kind === 'active-link') symlinkSync('missing', active);
if (kind === 'restore-link') symlinkSync('missing', restore);
if (kind === 'hardlink') linkSync(f.source, active);
if (kind === 'restore-is-source') restore = f.source;
if (kind === 'same-destinations') restore = active;
if (kind === 'relative') active = 'relative.md';
const entries = readdirSync(f.dir).sort();
const result = cli('init', f.source, active, restore);
expect(result.status, kind).toBe(1);
expect(result.stdout).toBe('');
expect(readdirSync(f.dir).sort()).toEqual(entries);
expect(readFileSync(f.source)).toEqual(original);
if (kind === 'active-content') expect(readFileSync(active, 'utf8')).toBe('Other assigned work');
if (kind === 'restore-content') expect(readFileSync(restore, 'utf8')).toBe('Existing history');
}
});
test('line endings and Unicode survive normalization; only a missing final separator LF is added', () => {
for (const text of ['最後の API 要件 🧪\r\nREST must remain.\r\n', 'API and endpoint without final newline']) {
const f = fixture(); const restore = join(f.dir, process.platform === 'win32' ? "restore -- quoted '名前'.md" : 'restore -- quoted "名前".md');
writeFileSync(f.source, text);
invoke('init', f.source, f.active, restore);
expect(readFileSync(restore, 'utf8')).toBe(text);
const ceo = invoke('create', 'ceo', f.active, restore);
expect(readFileSync(ceo.snapshotPath, 'utf8')).toBe(text + (text.endsWith('\n') ? '' : '\n'));
expect(invoke('init', f.source, f.active, restore).reused).toBe(true);
}
});
test('large file identities remain distinct on reuse while real hardlink aliases are rejected', () => {
const f = fixture();
const worker = join(f.dir, 'large-file-ids.ts');
writeFileSync(worker, `import { mock } from 'bun:test';
const real = { ...await import('node:fs') };
const ids = new Map();
function observed(kind, file, options) {
const exact = real[kind](file, { ...options, bigint: true });
if (!exact) return exact;
const key = exact.dev + ':' + exact.ino;
if (!ids.has(key)) ids.set(key, 2n ** 60n + BigInt(ids.size));
const ino = ids.get(key);
const state = options?.bigint ? exact : real[kind](file, options);
return new Proxy(state, { get(target, key, receiver) {
return key === 'ino' ? (options?.bigint ? ino : Number(ino)) : Reflect.get(target, key, receiver);
} });
}
mock.module('node:fs', () => ({ ...real,
statSync: (file, options) => observed('statSync', file, options),
lstatSync: (file, options) => observed('lstatSync', file, options),
}));
const { initializePlan } = await import(${JSON.stringify(TOOL)});
const [source, active, restore, alias] = process.argv.slice(2);
const initial = initializePlan(source, active, restore);
const reused = initializePlan(source, active, restore);
real.linkSync(source, alias);
let rejected = false;
try { initializePlan(source, alias, restore + '.other'); }
catch (error) { rejected = error.message.includes('ambiguous alias'); }
console.log(JSON.stringify({ initial: initial.reused, reused: reused.reused, rejected,
roundedIds: new Set([...ids.values()].map(Number)).size, exactIds: ids.size }));
`);
const result = spawnSync(process.execPath, [worker, f.source, f.active, f.restore, join(f.dir, 'hardlink.md')], {
encoding: 'utf8', timeout: 10_000,
});
expect(result.status, result.stderr).toBe(0);
const report = JSON.parse(result.stdout);
expect(report).toMatchObject({ initial: false, reused: true, rejected: true, roundedIds: 1 });
expect(report.exactIds).toBeGreaterThan(2);
expect(readFileSync(f.source)).toEqual(original);
expect(readFileSync(f.restore)).toEqual(original);
expect(existsSync(f.restore + '.other')).toBe(false);
});
test('staging failure cleans owned temporary files without changing source or active bytes', () => {
const f = fixture();
const active = join(f.dir, 'harness', 'plans', 'assigned.md');
const restore = join(f.dir, 'state', 'project', 'restore.md');
const worker = join(f.dir, 'fail-stage.ts');
writeFileSync(worker, `import { mock } from 'bun:test';
const real = { ...await import('node:fs') };
mock.module('node:fs', () => ({ ...real, mkdtempSync(prefix, options) {
if (String(prefix).includes('.gstack-autoplan-restore-')) throw new Error('Injected restore staging failure');
return real.mkdtempSync(prefix, options);
} }));
const { initializePlan } = await import(${JSON.stringify(TOOL)});
try { initializePlan(...process.argv.slice(2)); process.exitCode = 5; }
catch (error) { console.error(error.message); process.exitCode = 1; }
`);
const result = spawnSync(process.execPath, [worker, f.source, active, restore], {
encoding: 'utf8', timeout: 10_000,
});
expect(result.status).toBe(1);
expect(result.stdout).toBe('');
expect(result.stderr).toContain('Injected restore staging failure');
expect(readFileSync(f.source)).toEqual(original);
expect(readdirSync(f.dir).sort()).toEqual(['fail-stage.ts', 'source plan.md']);
expect(existsSync(active)).toBe(false);
expect(existsSync(restore)).toBe(false);
});
+167
View File
@@ -0,0 +1,167 @@
import { afterEach, describe, expect, test } from 'bun:test';
import { chmodSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import { spawnSync } from 'node:child_process';
import { pathToFileURL } from 'node:url';
import { prepareMethodology, createSnapshot } from '../bin/gstack-autoplan-snapshot';
import { auditAutoplanMethodReads, loadAutoplanMethodologyBinding } from './helpers/autoplan-method-read-audit';
import { readPlanCountTranscript, type NativePublicToolEvent } from './helpers/plan-count-transcript';
import recorded from './fixtures/autoplan-method-read-aa-events.json';
const ROOT = resolve(import.meta.dir, '..');
const clone = <T>(value: T): T => JSON.parse(JSON.stringify(value));
const events = () => clone(recorded.events) as NativePublicToolEvent[];
const binding = () => clone(recorded.binding);
function complete(): NativePublicToolEvent[] {
const list = events();
const last = list.pop()!;
const common = { sessionId: last.sessionId, toolUseId: 'synthetic-tail-repair-before-dispatch' };
list.push({ ...common, kind: 'use', timestamp: '2026-09-09T12:58:00.000Z', name: 'Read',
input: { file_path: recorded.binding.path, offset: 2200, limit: 60 } });
list.push({ ...common, kind: 'result', timestamp: '2026-09-09T12:58:00.020Z', isError: false,
file: { filePath: recorded.binding.path, startLine: 2200, numLines: 60, totalLines: 2259,
content: recorded.binding.content.split('\n').slice(2199).join('\n') } });
list.push(last); return list;
}
const audit = (list = events()) => auditAutoplanMethodReads(list, binding)[0]!;
const owned: string[] = [];
afterEach(() => { for (const dir of owned.splice(0)) rmSync(dir, { recursive: true, force: true }); });
describe('actual AA parent methodology delivery boundary', () => {
test('actual six full-content ranges miss the tail at the actual dispatch', () => {
const actual = audit();
expect(actual.passed).toBe(false);
expect(actual.ranges).toHaveLength(6);
expect(actual.missing).toEqual([{ startLine: 2200, endLine: 2259 }]);
expect(actual.at).toBe('2026-09-09T12:59:34.639Z');
});
test('explicitly synthetic final range repair before dispatch completes the same content', () => {
const fixed = audit(complete());
expect(fixed.passed).toBe(true);
expect(fixed.missing).toEqual([]);
expect(fixed.ranges.at(-1)).toMatchObject({ startLine: 2200, endLine: 2259 });
});
test('a later result, foreign session, wrong path, error, or unpaired result supplies no tail', () => {
for (const change of ['later', 'foreign-session', 'wrong-path', 'error', 'unpaired']) {
const list = complete(); const result = list.at(-2)!;
if (change === 'later') { list.splice(list.length - 2, 1); list.push(result); }
if (change === 'foreign-session') result.sessionId = 'different-parent';
if (change === 'wrong-path') (result.file as any).filePath += '.other';
if (change === 'error') result.isError = true;
if (change === 'unpaired') result.toolUseId += '-orphan';
expect(audit(list).passed, change).toBe(false);
}
});
test('matching ranges/hash claims cannot replace exact delivered bytes or EOF accounting', () => {
for (const change of ['content', 'missing-eof', 'total', 'start', 'limit']) {
const list = complete(); const result = list.at(-2)!; const file = result.file as any;
if (change === 'content') file.content = file.content.replace('MODE COMPARISON', 'METHOD COMPLETE');
if (change === 'missing-eof') { file.content = file.content.slice(0, -1); file.numLines--; }
if (change === 'total') file.totalLines--;
if (change === 'start') file.startLine--;
if (change === 'limit') list.at(-3)!.input!.limit = 59;
expect(audit(list).passed, change).toBe(false);
}
});
test('malformed dispatch identities/timestamps and empty Read IDs cannot supply coverage', () => {
for (const change of ['timestamp', 'session', 'dispatch-id', 'read-id']) {
const list = complete();
if (change === 'timestamp') list.at(-1)!.timestamp = 'not-a-timestamp';
if (change === 'session') for (const event of list) event.sessionId = '';
if (change === 'dispatch-id') list.at(-1)!.toolUseId = ' ';
if (change === 'read-id') list.at(-3)!.toolUseId = list.at(-2)!.toolUseId = '';
expect(audit(list).passed, change).toBe(false);
}
});
test('conflicting same-ID results fail; identical repeated records add no extra credit', () => {
const list = complete(); list.splice(list.length - 1, 0, clone(list.at(-2)!));
expect(audit(list).passed).toBe(true);
(list.at(-2)!.file as any).content += 'altered';
expect(audit(list).passed).toBe(false);
expect(audit(list).error).toContain('Conflicting');
});
test('self-report, child reads and backward request/result chronology do not fill the gap', () => {
const list = complete(); list.at(-2)!.timestamp = '2026-09-09T12:57:00.000Z';
expect(audit(list).passed).toBe(false);
const child = complete(); child.at(-3)!.sessionId = child.at(-2)!.sessionId = 'child';
expect(audit(child).passed).toBe(false);
const claim = events(); claim.splice(claim.length - 1, 0, { sessionId: recorded.events[0]!.sessionId,
timestamp: '2026-09-09T12:59:00.000Z', toolUseId: 'claim', kind: 'result',
content: 'Methodology read completely (lines 1-2259); sha256 ' + recorded.binding.sha256 });
expect(audit(claim).passed).toBe(false);
});
test('existing owned native transcript filter projects public tool content only', () => {
const dir = mkdtempSync(join(tmpdir(), 'gstack-method-events-')); owned.push(dir);
const cwd = '/fixture/cwd'; const project = join(dir, 'projects', 'fixture'); mkdirSync(project, { recursive: true });
const records = complete().map(event => ({ cwd, isSidechain: false, sessionId: event.sessionId,
timestamp: event.timestamp, message: { role: event.kind === 'use' ? 'assistant' : 'user', content: [event.kind === 'use'
? { type: 'tool_use', id: event.toolUseId, name: event.name, input: event.input }
: { type: 'tool_result', tool_use_id: event.toolUseId, is_error: event.isError, content: event.content ?? '' }] },
toolUseResult: event.kind === 'result' ? { file: event.file } : undefined }));
const text = records.map(row => JSON.stringify(row)).join('\n') + '\n';
const native = join(project, `${recorded.events[0]!.sessionId}.jsonl`); writeFileSync(native, text);
const projection: NativePublicToolEvent[] = []; const transcript = readPlanCountTranscript(dir, cwd, e => projection.push(e));
expect(transcript.status).toBe('ready'); expect(audit(projection).passed).toBe(true);
writeFileSync(native, records.map(row => JSON.stringify({ ...row, isSidechain: true })).join('\n') + '\n');
const foreign: NativePublicToolEvent[] = []; readPlanCountTranscript(dir, cwd, e => foreign.push(e));
expect(foreign).toEqual([]);
writeFileSync(native, text.slice(0, text.lastIndexOf('\n', text.length - 2) + 1) + '{"incomplete":');
const partial: NativePublicToolEvent[] = []; readPlanCountTranscript(dir, cwd, e => partial.push(e));
expect(auditAutoplanMethodReads(partial, binding)).toEqual([]);
});
});
describe('actual immutable snapshot methodology binding', () => {
function fixture() {
const dir = mkdtempSync(join(tmpdir(), 'gstack-method-binding-')); owned.push(dir);
const restore = join(dir, 'restore.md'); const active = join(dir, 'active.md');
writeFileSync(restore, '# Original\n'); writeFileSync(active, '## Implementation plan\n# Original\n\n## Review record\n');
const method = prepareMethodology('ceo', join(ROOT, 'plan-ceo-review/SKILL.md'), restore);
const snapshot = createSnapshot('ceo', active, restore, method.methodologyPath);
return { dir, method, snapshot };
}
test('dispatch binds actual immutable snapshot/method bytes independent of filtered helper stdout', () => {
const f = fixture(); const result = loadAutoplanMethodologyBinding(f.snapshot.nativeDispatchPrompt, [f.dir]);
expect(result.sha256).toBe(f.method.sha256); expect(result.content).toBe(readFileSync(f.method.methodologyPath, 'utf8'));
expect(result.lines).toBe(f.method.lines);
expect(() => loadAutoplanMethodologyBinding(f.snapshot.nativeDispatchPrompt.replace('CEO', 'DESIGN'), [f.dir])).toThrow();
const other = fixture(); expect(() => loadAutoplanMethodologyBinding(f.snapshot.nativeDispatchPrompt, [other.dir])).toThrow();
});
test('mutable or altered snapshot/method data cannot supply valid dispatch coverage', () => {
for (const kind of ['mutable', 'native', 'method', 'manifest']) {
const f = fixture(); const target = kind === 'native' ? f.snapshot.nativePromptPath : kind === 'manifest'
? join(f.method.methodologyPath, '..', 'methodology.json') : f.method.methodologyPath;
chmodSync(target, 0o600);
if (kind !== 'mutable') { writeFileSync(target, readFileSync(target, 'utf8') + 'tamper'); chmodSync(target, 0o444); }
if (kind === 'mutable') {
if (process.platform !== 'win32') expect(() => loadAutoplanMethodologyBinding(f.snapshot.nativeDispatchPrompt, [f.dir]), kind).toThrow();
// Windows does not use POSIX permission bits. Exercise that policy in
// an isolated process with observed modes, keeping actual artifact
// paths/bytes and both the immutable control and writable rejection.
const worker = join(f.dir, 'observed-mode.ts');
writeFileSync(worker, `import { mock } from 'bun:test';
const real = { ...await import('node:fs') };
await import('node:path');
const input = JSON.parse(await Bun.stdin.text());
Object.defineProperty(process, 'platform', { value: 'linux' });
mock.module('node:fs', () => ({ ...real, lstatSync(file) {
const stat = real.lstatSync(file);
stat.mode = (stat.mode & ~0o777) | (file === input.target && input.writable ? 0o600 : 0o444);
return stat;
} }));
const { loadAutoplanMethodologyBinding } = await import(${JSON.stringify(pathToFileURL(join(ROOT, 'test/helpers/autoplan-method-read-audit.ts')).href)});
loadAutoplanMethodologyBinding(input.prompt, input.roots);
`);
for (const writable of [false, true]) {
const result = spawnSync(process.execPath, [worker], { encoding: 'utf8', timeout: 10_000,
input: JSON.stringify({ prompt: f.snapshot.nativeDispatchPrompt, roots: [f.dir], target, writable }) });
expect(result.error).toBeUndefined();
expect(result.status, result.stderr).toBe(writable ? 1 : 0);
if (writable) expect(result.stderr).toContain('Artifact is not immutable bounded regular data');
}
} else {
expect(() => loadAutoplanMethodologyBinding(f.snapshot.nativeDispatchPrompt, [f.dir]), kind).toThrow();
}
}
});
});
+479
View File
@@ -0,0 +1,479 @@
import { afterEach, expect, test } from 'bun:test';
import { createHash } from 'node:crypto';
import { chmodSync, mkdtempSync, readFileSync, rmSync, statSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { spawnSync } from 'node:child_process';
import { createSnapshot, prepareMethodology, extractImplementationPlan } from '../bin/gstack-autoplan-snapshot';
const TOOL = join(import.meta.dir, '../bin/gstack-autoplan-snapshot.ts');
const captured = JSON.parse(readFileSync(join(import.meta.dir, 'fixtures/autoplan/t-ceo-omitted-obligations.json'), 'utf8'));
const lost = JSON.parse(readFileSync(join(import.meta.dir, 'fixtures/autoplan/u-ceo-original-loss.json'), 'utf8'));
const dangling = JSON.parse(readFileSync(join(import.meta.dir, 'fixtures/autoplan/v-ceo-dangling-references.json'), 'utf8'));
function methodology(phase: string, restore: string) {
return prepareMethodology(phase, join(import.meta.dir, '..', `plan-${phase === 'dx' ? 'devex' : phase}-review`, 'SKILL.md'), restore).methodologyPath;
}
const owned: string[] = [];
afterEach(() => { for (const dir of owned.splice(0)) rmSync(dir, { recursive: true, force: true }); });
function invoke(...args: string[]) {
const result = spawnSync(process.execPath, [TOOL, ...args], {
encoding: 'utf8', timeout: 10_000, maxBuffer: 2 * 1024 * 1024,
});
if (result.error) throw result.error;
return result;
}
function setup(body = 'Build the dashboard.\n') {
const dir = mkdtempSync(join(tmpdir(), 'gstack-obligations-')); owned.push(dir);
const active = join(dir, 'plan.md'); const restore = join(dir, 'restore.md');
writeFileSync(active, `## Implementation plan\n${body}## Review record\n`);
writeFileSync(restore, 'Original restore bytes\n');
const snapshot = createSnapshot('ceo', active, restore, methodology('ceo', restore));
return { dir, active, restore, snapshot };
}
const block = (phase: string, body: string) => `<!-- autoplan-accepted:${phase} -->\n${body}\n<!-- /autoplan-accepted:${phase} -->\n`;
const appendRecord = (active: string, value: string) => writeFileSync(active, readFileSync(active, 'utf8') + value);
for (const command of ['amend', 'check']) {
test(`actual V dangling local requirements reject ${command} without changing preserved baseline`, () => {
const f = setup(dangling.initialImplementation);
writeFileSync(f.active, dangling.activeAfterAmend);
const result = invoke(command, 'ceo', f.active, f.snapshot.snapshotPath, ...(command === 'check' ? ['changed'] : []));
expect(result.status).toBe(1);
expect(result.stderr).toContain('Review-record-only Section 6');
expect(readFileSync(f.active, 'utf8')).toBe(dangling.activeAfterAmend);
});
}
test('actual V dangling references cannot create the next blind input even if close was skipped', () => {
const f = setup(dangling.initialImplementation);
writeFileSync(f.active, dangling.activeAfterAmend);
expect(() => createSnapshot('design', f.active, f.restore, methodology('design', f.restore))).toThrow('Review-record-only Section 6');
expect(readFileSync(f.active, 'utf8')).toBe(dangling.activeAfterAmend);
expect(readFileSync(f.restore, 'utf8')).toBe('Original restore bytes\n');
});
test('local requirement references reject before first publication and ignore fake implementation headings', () => {
for (const fake of ['', '```md\n### Section 8: Metrics\n```\n', '> ### Section 8: Metrics\n']) {
const f = setup('Build the dashboard.\n' + fake);
appendRecord(f.active, '### Section 8: Metrics\nCount errors.\n' + block('ceo', '- Instrumentation as specified in Section 8.'));
const before = readFileSync(f.active, 'utf8');
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status).toBe(1);
expect(result.stderr).toContain('Review-record-only Section 8');
expect(readFileSync(f.active, 'utf8')).toBe(before);
}
});
test('local requirement references cannot hide behind an unrelated filename or URL', () => {
for (const body of ['- Add all tests in Section 6; update README.md.',
'- Add all tests in Section 6; see https://example.test/other.',
'- Read README.md and add all tests in Section 6.',
'- Follow https://example.test/other and add all tests in Section 6.']) {
const f = setup();
appendRecord(f.active, '### Section 6: Tests\nRun coverage.\n' + block('ceo', body));
const before = readFileSync(f.active, 'utf8');
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status).toBe(1);
expect(result.stderr).toContain('Review-record-only Section 6');
expect(readFileSync(f.active, 'utf8')).toBe(before);
}
});
test('inline adopted tests and instrumentation instead of transporting V review-only references', () => {
const f = setup(dangling.initialImplementation);
const full = dangling.activeAfterAmend as string;
const tests = full.match(/FLOW \/ CODEPATH[\s\S]*?No → T-S6-12/)![0];
const metrics = full.match(/Metric\/log[\s\S]*?api\.dashboard\.partial_failure_rate[^\n]*/)![0];
const logs = full.match(/- Endpoint entry:[\s\S]*?- Mutation:[^\n]*/)![0];
const indented = (value: string) => value.split('\n').map(line => ' ' + line).join('\n');
const revised = full.slice(full.indexOf('## Review record\n') + '## Review record\n'.length)
.replace('- All 12 test scenarios in Section 6 required before rollout.', '- Required test scenarios before rollout:\n' + indented(tests))
.replace('- Dashboard instrumentation: metrics and structured logs as specified in Section 8.', '- Required instrumentation:\n' + indented(metrics) + '\n' + indented(logs));
appendRecord(f.active, revised);
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, result.stderr).toBe(0);
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(0);
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
const input = readFileSync(next.snapshotPath, 'utf8');
expect(input.startsWith(dangling.initialImplementation)).toBe(true);
for (const term of ['T-S6-1', 'T-S6-12', 'dashboard.page.loaded', 'api.dashboard.partial_failure_rate', 'dashboard_fetch_start', 'snapshot_time']) {
expect(input).toContain(term);
}
expect(input).not.toContain('as specified in Section 8');
expect(input).not.toContain('All 12 test scenarios in Section 6');
expect(input).not.toContain('CEO DUAL VOICES');
expect(input).not.toContain('autoplan-accepted:');
});
test('reference checks preserve external, unresolved, ambiguous, quoted and satisfied local references', () => {
const examples = [
{ base: '### Section 6: Tests\nRun regression coverage.\n', body: '- Run tests in Section 6 before rollout.' },
{ body: '- Run tests in Section 6 of docs/testing.md.' },
{ body: '- Run tests in Section 6 (https://example.test/spec).' },
{ body: '- Follow https://example.test/spec as specified in Section 6.' },
{ body: '- Run tests in Section 60 before rollout.' },
{ body: '- Run tests in Section 6.1 before rollout.' },
{ body: '- Display the literal "as specified in Section 6".' },
{ body: '- Display `as specified in Section 6` as example text.' },
{ body: '- Document an example:\n ```md\n tests as specified in Section 6.\n ```' },
{ body: '- Document an example:\n > tests as specified in Section 6.' },
{ review: '### Section 6: First\nTests.\n### Section 6: Second\nOther tests.\n', body: '- Run tests in Section 6 before rollout.' },
{ review: '```md\n### Section 6: Tests\n```\n', body: '- Run tests in Section 6 before rollout.' },
];
for (const example of examples) {
const f = setup(example.base || 'Build the dashboard.\n');
appendRecord(f.active, (example.review || '### Section 6: Tests\nRun coverage.\n') + block('ceo', example.body));
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, example.body + result.stderr).toBe(0);
expect(() => createSnapshot('design', f.active, f.restore, methodology('design', f.restore))).not.toThrow();
}
});
test('new reference checks do not bind a later phase to an earlier phase review heading or baseline prose', () => {
const f = setup('Existing external contract uses tests in Section 6.\n');
appendRecord(f.active, '### Section 6: CEO tests\nOriginal review.\n' + block('ceo', '- Keep all authorization checks.'));
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const design = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
appendRecord(f.active, block('design', '- Run tests in Section 6 before rollout.'));
expect(invoke('amend', 'design', f.active, design.snapshotPath).status).toBe(0);
expect(() => createSnapshot('dx', f.active, f.restore, methodology('dx', f.restore))).not.toThrow();
});
test('actual T changed-only plan cannot close with accepted obligations only in its review', () => {
const f = setup(captured.initialImplementation);
writeFileSync(f.active, captured.activeAtBoundary);
const result = invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed');
expect(result.status).toBe(1);
expect(result.stderr).toContain('Missing accepted-obligations record for ceo');
expect(readFileSync(f.active, 'utf8')).toBe(captured.activeAtBoundary);
});
test('whole recorded T obligations retain omitted guards and every nested verification in the next input', () => {
const f = setup(captured.initialImplementation);
const accepted = block('ceo', captured.acceptedObligations.trimEnd());
writeFileSync(f.active, captured.activeAtBoundary + accepted);
const reviewBefore = readFileSync(f.active, 'utf8').split('## Review record\n')[1];
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(1);
const rewritten = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(rewritten);
writeFileSync(f.active, '## Implementation plan\n' + captured.initialImplementation + '## Review record\n' + reviewBefore);
const amended = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(amended.status, amended.stderr).toBe(0);
const checked = invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed');
expect(checked.status, checked.stderr).toBe(0);
const result = JSON.parse(checked.stdout);
expect(result.limitation).toContain('enumeration and semantic correctness still require review');
expect(result.implementation).toContain(accepted);
for (const detail of ['only one request fired', 'Failed to mark as read. Try again.', 'Panel-level retry',
'Screen reader: live region', 'Reduced-motion', 'session expiry mid-page-load', 'RTL test with mixed panel results']) {
expect(result.implementation).toContain(detail);
}
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
expect(readFileSync(next.sourceSnapshotPath, 'utf8')).toContain(accepted);
expect(readFileSync(next.snapshotPath, 'utf8')).toContain(captured.acceptedObligations.trimEnd());
expect(readFileSync(next.snapshotPath, 'utf8')).not.toContain('autoplan-accepted:');
expect(next.nativePrompt).not.toContain('autoplan-accepted:');
expect(next.nativeDispatchPrompt).not.toContain(next.sourceSnapshotPath);
expect(readFileSync(next.snapshotPath, 'utf8')).not.toContain('CEO DUAL VOICES');
expect(readFileSync(f.active, 'utf8').split('## Review record\n')[1]).toBe(reviewBefore);
expect(readFileSync(f.restore, 'utf8')).toBe('Original restore bytes\n');
});
const editRecord = (phase: string, sourceSha256: string, replacements: Array<{ oldText: string; newText: string }>) =>
`<!-- autoplan-baseline-edits:${phase} ${JSON.stringify({ sourceSha256, replacements })} -->\n`;
test('actual U canonical block cannot conceal a rewritten original baseline at amend or check', () => {
const f = setup(lost.initialImplementation);
writeFileSync(f.active, lost.activeAfterAmend);
for (const command of ['amend', 'check']) {
const result = invoke(command, 'ceo', f.active, f.snapshot.snapshotPath, ...(command === 'check' ? ['changed'] : []));
expect(result.status).toBe(1);
expect(result.stderr).toContain('Unrecorded Implementation rewrite');
expect(readFileSync(f.active, 'utf8')).toBe(lost.activeAfterAmend);
}
});
test('unchanged U source retains every original byte when accepted requirements are appended', () => {
const f = setup(lost.initialImplementation);
appendRecord(f.active, block('ceo', '- Preserve each contract and add a loading state.\n Verify: reject cross-workspace requests.'));
const amended = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(amended.status, amended.stderr).toBe(0);
expect(JSON.parse(amended.stdout).implementation.startsWith(lost.initialImplementation)).toBe(true);
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
expect(readFileSync(next.snapshotPath, 'utf8').startsWith(lost.initialImplementation)).toBe(true);
expect(readFileSync(f.snapshot.sourceSnapshotPath, 'utf8')).toBe(lost.initialImplementation);
});
test('exact replacements and deletion preserve untouched CRLF/Unicode bytes and produce only effective blind input', () => {
const baseline = 'Keep café ✓.\r\nUse a blue button.\r\nObsolete behavior.\r\nKeep 日本語.\r\n';
for (const alreadyEdited of [false, true]) {
const f = setup(baseline);
const replacements = [{ oldText: 'blue', newText: 'green' }, { oldText: 'Obsolete behavior.\r\n', newText: '' }];
const edited = baseline.replace('blue', 'green').replace('Obsolete behavior.\r\n', '');
if (alreadyEdited) writeFileSync(f.active, `## Implementation plan\n${edited}## Review record\n`);
appendRecord(f.active, editRecord('ceo', f.snapshot.sourceSha256, replacements) +
block('ceo', '- Replace blue with green; remove obsolete behavior.\n Verify: green renders; obsolete behavior is absent.'));
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, result.stderr).toBe(0);
expect(JSON.parse(result.stdout).implementation.startsWith(edited)).toBe(true);
const first = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
expect(readFileSync(f.active, 'utf8')).toBe(first);
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(0);
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
expect(next.nativePrompt).not.toContain('autoplan-baseline-edits');
expect(next.nativePrompt).not.toContain('Use a blue button.');
expect(next.nativePrompt.endsWith(readFileSync(next.snapshotPath, 'utf8'))).toBe(true);
expect(readFileSync(next.snapshotPath, 'utf8').startsWith(edited)).toBe(true);
}
});
test('baseline edit record rejects stale source, ambiguous/overlapping anchors and malformed or quoted edits without writes', () => {
const f = setup('Keep owner permission.\nRepeat repeat.\n');
const original = readFileSync(f.active, 'utf8');
const accepted = block('ceo', '- Preserve scope.\n Verify: permission is checked.');
const hash = f.snapshot.sourceSha256;
const good = editRecord('ceo', hash, [{ oldText: 'owner', newText: 'member' }]);
const bad = [
editRecord('ceo', '0'.repeat(64), [{ oldText: 'owner', newText: 'member' }]),
editRecord('ceo', hash, [{ oldText: '', newText: 'inserted' }]),
editRecord('ceo', hash, [{ oldText: 'missing', newText: 'present' }]),
editRecord('ceo', hash, [{ oldText: 'e', newText: 'E' }]),
editRecord('ceo', hash, [{ oldText: 'owner', newText: 'member' }, { oldText: 'owner', newText: 'admin' }]),
editRecord('ceo', hash, [{ oldText: 'owner permission', newText: 'member' }, { oldText: 'permission', newText: 'scope' }]),
editRecord('ceo', hash, [{ oldText: 'owner', newText: '\ud800' }]),
good + good, good.replace('"replacements":', '"unknown":'), good.replace('ceo ', 'invalid '),
good.replace(' -->', ''), good.replace('"oldText":"owner"', '"oldText":"owner","extra":true'),
good.replace('"oldText":"owner"', '"oldText":"other","oldText":"owner"'),
editRecord('ceo', hash, [{ oldText: 'owner', newText: '\n' + good }]),
editRecord('ceo', hash, [{ oldText: 'owner', newText: 'owner\n## Review record\n' }]),
];
for (const record of bad) {
const plan = original + accepted + record; writeFileSync(f.active, plan);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status, record).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(plan);
}
for (const record of ['```html\n' + good + '```\n', '> ' + good, ' ' + good]) {
const plan = original.replace('owner', 'member') + accepted + record; writeFileSync(f.active, plan);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(plan);
}
});
test('later exact baseline revisions preserve earlier accepted blocks and reject edits into them', () => {
const f = setup('Use blue.\n');
const ceo = block('ceo', '- Preserve owner authorization.\n Verify: reject cross-user access.');
appendRecord(f.active, ceo);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const design = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
const original = readFileSync(f.active, 'utf8');
const accepted = block('design', '- Replace the blue baseline with green.\n Verify: green keeps the authorized action.');
for (const oldText of ['owner authorization', ceo, 'Use blue.\n\n' + ceo]) {
const plan = original + accepted + editRecord('design', design.sourceSha256, [{ oldText, newText: 'replacement' }]);
writeFileSync(f.active, plan);
expect(invoke('amend', 'design', f.active, design.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(plan);
}
writeFileSync(f.active, original + accepted + editRecord('design', design.sourceSha256, [{ oldText: 'blue', newText: 'green' }]));
const result = invoke('amend', 'design', f.active, design.snapshotPath);
expect(result.status, result.stderr).toBe(0);
expect(JSON.parse(result.stdout).implementation).toContain(ceo);
expect(JSON.parse(result.stdout).implementation.startsWith('Use green.\n')).toBe(true);
const next = createSnapshot('dx', f.active, f.restore, methodology('dx', f.restore));
expect(readFileSync(next.snapshotPath, 'utf8')).toContain('Preserve owner authorization.');
expect(next.nativePrompt).not.toContain('sourceSha256');
});
test('empty exact-edit list supports honest unchanged closure; declared edits cannot hide behind None', () => {
const f = setup(); const original = readFileSync(f.active, 'utf8');
appendRecord(f.active, block('ceo', 'None: Existing baseline suffices.') + editRecord('ceo', f.snapshot.sourceSha256, []));
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, result.stderr).toBe(0);
expect(JSON.parse(result.stdout).changed).toBe(false);
writeFileSync(f.active, original + block('ceo', 'None: Existing baseline suffices.') +
editRecord('ceo', f.snapshot.sourceSha256, [{ oldText: 'dashboard', newText: 'inbox' }]));
const before = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(before);
});
test('create returns an exact current-source edit record without adding reviewer metadata to immutable payloads', () => {
const f = setup();
expect(f.snapshot.baselineEdits.record).toBe(editRecord('ceo', f.snapshot.sourceSha256, []).trimEnd());
expect(f.snapshot.baselineEdits.instructions).toContain('not approval or completeness');
const manifest = readFileSync(join(f.snapshot.snapshotPath, '..', 'snapshot.json'), 'utf8');
expect(manifest).not.toContain('baselineEdits');
expect(readFileSync(f.snapshot.nativePromptPath, 'utf8')).toBe(f.snapshot.nativePrompt);
expect(f.snapshot.nativePrompt).not.toContain('autoplan-baseline-edits');
expect(statSync(f.snapshot.sourceSnapshotPath).mode & 0o777).toBe(0o444);
});
test('amend is idempotent and check rejects a dropped condition or verification line', () => {
const f = setup(); const accepted = block('ceo', '- Disable while pending.\n Verify: two clicks fire one request.');
appendRecord(f.active, accepted);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const first = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
expect(readFileSync(f.active, 'utf8')).toBe(first);
writeFileSync(f.active, first.replace(' Verify: two clicks fire one request.\n', ''));
const checked = invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed');
expect(checked.status).toBe(1);
expect(checked.stderr).toContain('not retained exactly');
});
test('the current phase can grow its accepted block without permitting an unrelated baseline rewrite', () => {
const f = setup('Use blue.\nKeep ownership checks.\n');
const first = block('ceo', '- Disable the action while pending.\n Verify: one request.');
const second = block('ceo', '- Disable the action while pending.\n Verify: one request.\n- Replace blue with green.\n Verify: green renders.');
appendRecord(f.active, first);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const current = readFileSync(f.active, 'utf8');
const boundary = current.indexOf('## Review record\n');
const revised = current.slice(0, boundary) + current.slice(boundary).replace(first, second) +
editRecord('ceo', f.snapshot.sourceSha256, [{ oldText: 'blue', newText: 'green' }]);
writeFileSync(f.active, revised.replace('Keep ownership checks.\n', ''));
const invalid = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(invalid);
writeFileSync(f.active, revised);
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, result.stderr).toBe(0);
expect(JSON.parse(result.stdout).implementation).toBe('Use green.\nKeep ownership checks.\n\n' + second);
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(0);
});
test('none requires a reason and unchanged implementation, without creating a fake amendment', () => {
const f = setup(); appendRecord(f.active, block('ceo', 'None: All current requirements were retained.'));
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'unchanged').status).toBe(0);
expect(extractImplementationPlan(readFileSync(f.active, 'utf8'))).toBe('Build the dashboard.\n');
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(1);
});
test('quoted, fenced, duplicate, malformed and mixed None records cannot authorize amendment', () => {
const f = setup(); const original = readFileSync(f.active, 'utf8');
const good = block('ceo', '- Add error handling.\n Verify: request failure shows retry.');
const bad = [
'```markdown\n' + good + '```\n', good.split('\n').map(l => '> ' + l).join('\n'),
good + good, good.replace('/autoplan-accepted:ceo', '/autoplan-accepted:design'),
good.replace('<!-- autoplan-accepted:ceo -->', '<!-- autoplan-accepted:unknown -->'),
block('ceo', '- Severity: critical'), block('ceo', '- **Severity:** critical'), block('ceo', '- Add a guard.\n Consensus: CONFIRMED'),
block('ceo', '- Add guard.\n## CEO Review'), block('ceo', ''), block('ceo', 'None:'), block('ceo', 'None: No changes.\n- Also add a new feature.'),
];
for (const record of bad) {
const plan = original + record; writeFileSync(f.active, plan);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(plan);
}
});
test('phase/path/snapshot identity still rejects before any amendment', () => {
const f = setup(); appendRecord(f.active, block('ceo', '- Add error handling.'));
const before = readFileSync(f.active, 'utf8');
expect(invoke('amend', 'design', f.active, f.snapshot.snapshotPath).status).toBe(1);
const another = join(f.dir, 'another.md'); writeFileSync(another, before);
expect(invoke('amend', 'ceo', another, f.snapshot.snapshotPath).status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(before);
expect(readFileSync(another, 'utf8')).toBe(before);
});
test('later phases retain prior registered obligations and cannot erase them with None', () => {
const f = setup(); const ceo = block('ceo', '- Handle network failure.\n Verify: offer retry.');
appendRecord(f.active, ceo); expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const design = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
const next = block('design', '- Show a named error control.\n Verify: keyboard reaches retry.');
appendRecord(f.active, next);
expect(invoke('amend', 'design', f.active, design.snapshotPath).status).toBe(0);
const whole = readFileSync(f.active, 'utf8');
expect(extractImplementationPlan(whole)).toContain(ceo);
expect(extractImplementationPlan(whole)).toContain(next);
writeFileSync(f.active, whole.replaceAll(ceo, ''));
expect(invoke('check', 'design', f.active, design.snapshotPath, 'changed').status).toBe(1);
writeFileSync(f.active, whole.replaceAll(ceo, '').replace('## Review record\n', '## Review record\n' + block('ceo', 'None: no changes')));
expect(invoke('amend', 'design', f.active, design.snapshotPath).status).toBe(1);
});
test('UTF-8 and CRLF requirements survive exact copying and a repeated no-change review', () => {
const f = setup('Keep café and 日本語.\r\n');
const accepted = block('ceo', '- Show ✓ for success.\n Verify: naïve input stays intact.').replaceAll('\n', '\r\n');
appendRecord(f.active, accepted);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
expect(extractImplementationPlan(readFileSync(f.active, 'utf8'))).toContain(accepted);
const again = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore));
expect(invoke('amend', 'ceo', f.active, again.snapshotPath).status).toBe(0);
expect(invoke('check', 'ceo', f.active, again.snapshotPath, 'unchanged').status).toBe(0);
});
test('closing marker at EOF cannot swallow the Review-record boundary or lose obligation bytes', () => {
const f = setup();
const accepted = block('ceo', '- Keep the final requirement ✓.\n Verify: the final assertion stays.').trimEnd();
appendRecord(f.active, accepted);
const result = invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath);
expect(result.status, result.stderr).toBe(0);
const plan = readFileSync(f.active, 'utf8');
expect(extractImplementationPlan(plan)).toContain(accepted + '\n');
expect(plan.endsWith(accepted)).toBe(true);
expect(invoke('check', 'ceo', f.active, f.snapshot.snapshotPath, 'changed').status).toBe(0);
});
test('changing both earlier copies cannot erase the authorization obligation from the immutable input', () => {
const f = setup();
const original = block('ceo', '- Preserve owner authorization.\n Verify: reject cross-user access.');
appendRecord(f.active, original);
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const design = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
const changed = readFileSync(f.active, 'utf8').replaceAll(original,
block('ceo', '- Permit cross-user access.\n Verify: cross-user access succeeds.')) +
block('design', '- Label the owner control.\n Verify: accessible name.');
writeFileSync(f.active, changed);
const result = invoke('amend', 'design', f.active, design.snapshotPath);
expect(result.status).toBe(1);
expect(result.stderr).toContain('Prior accepted obligations changed: ceo');
expect(invoke('check', 'design', f.active, design.snapshotPath, 'changed').status).toBe(1);
expect(readFileSync(f.active, 'utf8')).toBe(changed);
});
test('blind projection preserves UTF-8/CRLF bodies and fenced examples, while binding the full source', () => {
const example = '```html\n<!-- autoplan-accepted:design -->\nLiteral documentation example\n<!-- /autoplan-accepted:design -->\n```\n';
const f = setup(example);
const body = '- Preserve café ✓ and 日本語.\r\n Verify: the last condition survives.\r\n';
appendRecord(f.active, '<!-- autoplan-accepted:ceo -->\r\n' + body + '<!-- /autoplan-accepted:ceo -->\r\n');
expect(invoke('amend', 'ceo', f.active, f.snapshot.snapshotPath).status).toBe(0);
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
const transport = readFileSync(next.snapshotPath, 'utf8');
expect(transport).toContain(example);
expect(transport).toContain(body);
expect(transport).not.toContain('autoplan-accepted:ceo');
expect(next.nativePrompt.endsWith(transport)).toBe(true);
appendRecord(f.active, block('design', 'None: Existing requirements suffice.'));
expect(invoke('check', 'design', f.active, next.snapshotPath, 'unchanged').status).toBe(0);
const manifestPath = join(next.sourceSnapshotPath, '..', 'snapshot.json');
// Deliberate corruption owns these files; production snapshots remain read-only.
for (const file of [next.sourceSnapshotPath, next.snapshotPath, manifestPath]) {
expect(statSync(file).mode & 0o222).toBe(0);
chmodSync(file, 0o600);
}
const original = readFileSync(next.sourceSnapshotPath, 'utf8');
writeFileSync(next.sourceSnapshotPath, original.replace('last condition', 'different condition'));
expect(invoke('check', 'design', f.active, next.snapshotPath, 'unchanged').status).toBe(1);
writeFileSync(next.sourceSnapshotPath, original);
const manifest = JSON.parse(readFileSync(manifestPath, 'utf8'));
writeFileSync(manifestPath, JSON.stringify({ ...manifest, sourceSnapshotPath: f.active }));
expect(invoke('amend', 'design', f.active, next.snapshotPath).status).toBe(1);
writeFileSync(manifestPath, JSON.stringify({ ...manifest, schemaVersion: 1 }));
expect(invoke('check', 'design', f.active, next.snapshotPath, 'unchanged').status).toBe(1);
const { sourceSnapshotPath, sourceSha256, sourceBytes, ...downgraded } = manifest;
writeFileSync(manifestPath, JSON.stringify({ ...downgraded, schemaVersion: 1 }));
expect(invoke('check', 'design', f.active, next.snapshotPath, 'unchanged').status).toBe(1);
const altered = transport.replace('last condition', 'different condition');
writeFileSync(next.snapshotPath, altered);
writeFileSync(manifestPath, JSON.stringify({ ...manifest, sha256: createHash('sha256').update(altered).digest('hex') }));
const mismatch = invoke('check', 'design', f.active, next.snapshotPath, 'unchanged');
expect(mismatch.status).toBe(1);
expect(mismatch.stderr).toContain('blind review projection does not match');
});
@@ -0,0 +1,70 @@
import {expect,test} from 'bun:test';
import fs from 'node:fs';
import {autoplanPermissionProgressKey} from './helpers/autoplan-artifact-permission';
import type {NativePublicToolEvent} from './helpers/plan-count-transcript';
import capture from './fixtures/autoplan-overwrite-progress-ax.json';
const before=()=>structuredClone(capture.beforeEvents) as NativePublicToolEvent[];
const after=()=>structuredClone(capture.afterEvents) as NativePublicToolEvent[];
test('the acknowledged 92-line Write distinguishes the next identical overwrite footer',()=>{
expect(capture.before.slice(-500)).toBe(capture.after.slice(-500));
const oldKey=autoplanPermissionProgressKey(capture.before,before());
const newKey=autoplanPermissionProgressKey(capture.after,after());
expect(oldKey).toEndWith(':toolu_01RBorP8UERrbVheRiXSN1v4');
expect(newKey).toEndWith(':toolu_01RPGbV4z5AAMcnzwD4qcx9p');
expect(newKey).not.toBe(oldKey);
});
test('the same still-pending dialog has no new progress epoch',()=>{
const events=before(),key=autoplanPermissionProgressKey(capture.before,events);
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
events.push(after()[2]!); // Published use alone has not completed.
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
events.push({...after()[3]!,isError:true});
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
});
test('unrelated results and same-basename files in other directories do not advance the epoch',()=>{
const key=autoplanPermissionProgressKey(capture.before,before());
for(const mutate of [
(events:NativePublicToolEvent[])=>{events[2]!.name='Read';},
events=>{events[2]!.name='Bash';},
events=>{events[2]!.input!.file_path=String(events[2]!.input!.file_path).replace('/ceo-plans/','/other-plans/');},
events=>{events[2]!.input!.file_path=String(events[2]!.input!.file_path).replace('/ceo-plans/','/ceo-plans-sibling/');},
events=>{events[3]!.isError=undefined;},
events=>{events[3]!.toolUseId='unrelated-result';},
events=>{events[3]!.timestamp='invalid';},
events=>{events[3]!.timestamp='2026-09-11T02:00:00Z';},
]){const events=after();mutate(events);expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);}
});
test('missing path authority, mixed sessions and duplicate uses supply no matching progress',()=>{
expect(autoplanPermissionProgressKey(capture.after,[])).toBeUndefined();
expect(autoplanPermissionProgressKey(capture.after.replace('overwrite 2026-09-11-user-dashboard.md','overwrite other.md'),after())).toBeUndefined();
expect(autoplanPermissionProgressKey(capture.after.replace('always allow access to','access to'),after())).toBeUndefined();
const mixed=after();mixed[3]!.sessionId='other';expect(autoplanPermissionProgressKey(capture.after,mixed)).toBeUndefined();
const duplicate=after();duplicate.splice(3,0,structuredClone(duplicate[2]!));
expect(autoplanPermissionProgressKey(capture.after,duplicate)).toBe(autoplanPermissionProgressKey(capture.before,before()));
});
test('the actual generic permission branch preserves classification and waits for selection',async()=>{
const source=fs.readFileSync(new URL('./skill-e2e-autoplan-chain.test.ts',import.meta.url),'utf8');
const block=source.slice(source.indexOf(' const recentTail = visible.slice(-1500);'),source.indexOf(' // This new repository offers routing'));
expect(block.match(/continue;/g)).toHaveLength(1);
const sends:string[]=[];let release:(()=>void)|undefined;
const select=async()=>{sends.push('selected');await new Promise<void>(r=>{release=r;});sends.push('confirmed');};
const make=new Function('autoplanPermissionProgressKey','selectPtyNumberedOption','Bun',`
let lastPermSig='',lastPermissionProgress='';
return async(visible,publicTools,allowed=true)=>{
const transcript={status:'ready'},session={};
const isNumberedOptionListVisible=()=>allowed,isPermissionDialogVisible=()=>allowed;
${block.replace('continue;','return;')}
};
`);
const step=make(autoplanPermissionProgressKey,select,{sleep:async()=>{}});
const first=step(capture.before,before());await Promise.resolve();expect(sends).toEqual(['selected']);release!();await first;
await step(capture.after,before());expect(sends).toEqual(['selected','confirmed']);
await step(capture.after,after(),false);expect(sends).toHaveLength(2); // Existing AUQ/permission classification still decides.
const next=step(capture.after,after());await Promise.resolve();expect(sends).toHaveLength(3);release!();await next;
await step(capture.after,after());expect(sends).toEqual(['selected','confirmed','selected','confirmed']);
});
+187
View File
@@ -0,0 +1,187 @@
import { afterEach, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
import fixture from './fixtures/autoplan-pending-artifact-ae.json';
import { autoplanArtifactPermissionInput, pendingAutoplanArtifactPermissionInput, autoplanArtifactMenuKey } from './helpers/autoplan-artifact-permission';
import { createAutoplanArtifactRecorder, recordAutoplanArtifact, readPendingAutoplanArtifact } from './helpers/autoplan-artifact-recorder';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
const roots:string[]=[];
afterEach(()=>{for(const root of roots.splice(0))fs.rmSync(root,{recursive:true,force:true});});
function replay(relative='ceo-plans/2026-09-09-user-dashboard.md') {
const root=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-pending-artifact-test-'));roots.push(root);
const cwd=path.join(root,path.basename(fixture.cwd)),ownedStateRoot=path.join(root,'home','.gstack'),config=path.join(root,'config');
fs.mkdirSync(cwd);const file=path.join(ownedStateRoot,'projects',path.basename(cwd),relative);
fs.mkdirSync(path.dirname(file),{recursive:true});fs.writeFileSync(file,fixture.before);
const native=path.join(config,'projects','fixture',fixture.sessionId+'.jsonl');fs.mkdirSync(path.dirname(native),{recursive:true});fs.writeFileSync(native,'');
const publicTools=structuredClone(fixture.events) as NativePublicToolEvent[];
for(const e of publicTools)if(e.input)e.input.file_path=file;
const recorder=createAutoplanArtifactRecorder(cwd,config,ownedStateRoot);
// Synthetic hook: only its identity/path are retained. Neither this input
// nor the displayed additions are claimed to reproduce the unpublished body.
const event={hook_event_name:'PreToolUse',tool_name:'Edit',session_id:fixture.sessionId,tool_use_id:'synthetic-current-edit',
cwd,transcript_path:native,tool_input:{file_path:file,old_string:'Synthetic old content',new_string:'Synthetic new content',replace_all:false}};
const record=(change:Record<string,unknown>={})=>recordAutoplanArtifact(JSON.stringify({...event,...change}),recorder.file,cwd,config,ownedStateRoot);
record();
const context={cwd,ownedStateRoot,commandStartedAt:fixture.commandStartedAt,now:Date.now(),viewportCapturedAt:Date.now(),
transcriptStatus:'ready',publicTools,pending:readPendingAutoplanArtifact(recorder.file,cwd,config,ownedStateRoot,fixture.commandStartedAt,publicTools)};
const screen=fixture.viewport.replaceAll(path.basename(fixture.file),path.basename(file));
roots.push(path.dirname(recorder.file));
return {root,file,native,config,recorder,event,record,context,screen};
}
const pick=(r:ReturnType<typeof replay>,seen=new Set<string>())=>pendingAutoplanArtifactPermissionInput(r.screen,r.context,seen);
test('actual public pane stays blocked without hook identity; synthetic owned metadata enables only one option',()=>{
const r=replay();
expect(autoplanArtifactPermissionInput(r.screen,r.context,new Set())).toBeNull();
expect(pick({...r,context:{...r.context,pending:undefined}})).toBeNull();
expect(pick(r)).toEqual({input:'1\r',signature:fixture.sessionId+':synthetic-current-edit',file:r.file});
expect(pick(r,new Set([pick(r)!.signature]))).toBeNull();
expect(JSON.stringify(r.context.pending)).not.toContain('Synthetic old content');
expect(r.context.publicTools).toHaveLength(fixture.events.length);
});
test('all130 actual published tool events preserve the same metadata-only fallback boundary',()=>{
const r=replay();r.context.publicTools=structuredClone(fixture.allPublicTools) as NativePublicToolEvent[];
for(const e of r.context.publicTools)if(e.input?.file_path===fixture.file)e.input.file_path=r.file;
expect(r.context.publicTools).toHaveLength(130);
expect(r.context.publicTools.filter(e=>e.kind==='use' && ['Write','Edit'].includes(e.name??''))).toHaveLength(39);
expect(pick(r)?.input).toBe('1\r');
});
test('completed or published requests and newer identities on an old granted viewport remain closed',()=>{
const r=replay(),first=pick(r)!;
const seen=new Set([first.signature,autoplanArtifactMenuKey(r.screen)]);
r.record({hook_event_name:'PostToolUse'});
expect(readPendingAutoplanArtifact(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot,r.context.commandStartedAt,r.context.publicTools)).toBeUndefined();
r.record({tool_use_id:'newer-request'});
r.context.now=Date.now();r.context.viewportCapturedAt=r.context.now;
r.context.pending=readPendingAutoplanArtifact(r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot,r.context.commandStartedAt,r.context.publicTools);
expect(pick(r,seen)).toBeNull();
r.context.publicTools.push({kind:'result',sessionId:fixture.sessionId,toolUseId:'newer-request',timestamp:new Date().toISOString(),isError:false});
expect(pick(r)).toBeNull();
});
test('hook after viewport, invalid clocks, future/stale/foreign IDs and missing success cannot authorize input',()=>{
const changes:Array<(r:ReturnType<typeof replay>)=>void>=[
r=>{r.context.viewportCapturedAt=Date.parse(r.context.pending!.timestamp)-1;},
r=>{r.context.now=NaN;},r=>{r.context.now=Infinity;},r=>{r.context.viewportCapturedAt=NaN;},
r=>{r.context.pending!.timestamp=new Date(r.context.now+10000).toISOString();},
r=>{r.context.pending!.timestamp=new Date(r.context.commandStartedAt-1).toISOString();},
r=>{r.context.pending!.sessionId='foreign';},r=>{r.context.pending!.toolUseId='';},
r=>{r.context.pending!.toolUseId='invalid:id';},r=>{r.context.pending!.file=42 as any;},r=>{r.context.publicTools=[];},
r=>{r.context.transcriptStatus='error';},
r=>{for(const e of r.context.publicTools)if(e.kind==='result')e.isError=true;},
r=>{r.context.publicTools.push({...r.context.publicTools[0]!,toolUseId:'unresolved-concurrent',timestamp:new Date().toISOString()});},
r=>{r.context.publicTools.push({...r.context.publicTools[0]!,toolUseId:r.context.pending!.toolUseId,timestamp:new Date().toISOString()});},
r=>{r.context.publicTools.push({...r.context.publicTools.at(-1)!,sessionId:'sibling'});},
];
for(const change of changes){const r=replay();change(r);expect(pick(r),change.toString()).toBeNull();}
});
test('changed, foreign and symlink files are rejected; all existing owned artifact layouts stay scoped',()=>{
for(const relative of ['ceo-plans/2026-09-09-user-dashboard.md','main-test-plan-20260909-220000.md','main-eng-review-test-plan-20260909-220000.md'])expect(pick(replay(relative))?.input).toBe('1\r');
for(const relative of ['other.md','config.yaml','tasks.jsonl','../sibling/ceo-plans/2026-09-09-user-dashboard.md'])expect(pick(replay(relative))).toBeNull();
let r=replay();fs.writeFileSync(r.file,'Changed unrelated content');expect(pick(r)).toBeNull();
r=replay();fs.utimesSync(r.file,new Date(r.context.now+10000),new Date(r.context.now+10000));expect(pick(r)).toBeNull();
if(process.platform!=='win32'){
r=replay();const sibling=r.file+'.sibling';fs.renameSync(r.file,sibling);fs.symlinkSync(sibling,r.file);expect(pick(r)).toBeNull();
}
r=replay();r.context.ownedStateRoot=path.join(r.root,'ambient-home');expect(pick(r)).toBeNull();
});
test('only a complete current native menu and current-file deleted/context rows support pending metadata',()=>{
const changes=[
(s:string)=>'Example:\n'+s,(s:string)=>'```\n'+s+'```',
(s:string)=>s.split('\n').map(l=>'> '+l).join('\n'),
(s:string)=>s.replace(' 1. Yes',' 1. Yes, always allow'),
(s:string)=>s.replace(' 1. Yes',' 1. Yes').replace(' 2. Yes',' 2. Yes'),
(s:string)=>s.replace(' 3. No',' 3. No\n 4. Run a command'),
(s:string)=>s.replace('2026-09-09-user-dashboard.md?','foreign.md?'),
(s:string)=>s.replace('Esc to cancel · Tab to amend','Enter to select'),
(s:string)=>s+'\nPlease run the extra work.',
(s:string)=>s.replace(' -than the latest',' -unrelated cropped text'),
(s:string)=>s.replace(' 50 -- **Retry.**',' 50 -- **Unrelated deletion.**'),
(s:string)=>s.slice(s.indexOf(' Do you want')),
];
for(const change of changes){const r=replay();r.screen=change(r.screen);expect(pick(r),change.toString()).toBeNull();}
});
test('queued unrelated public tools do not confer permission or block the current owned edit',()=>{
const r=replay();r.context.publicTools.push({kind:'use',sessionId:fixture.sessionId,toolUseId:'queued-bash',name:'Bash',
timestamp:new Date(r.context.now).toISOString(),input:{command:'echo queued'}});
expect(pick(r)?.input).toBe('1\r');
r.context.publicTools.at(-1)!.name='Write';expect(pick(r)).toBeNull();
});
for (const [line, numbered, next, continuation] of [
[7, ' 7 ', ' 8 ', ' '], [17, ' 17 ', ' 18 ', ' '],
[116, ' 116 ', ' 117 ', ' '], [1024, ' 1024 ', ' 1025 ', ' '],
] as const) test(`legacy pending deletion line ${line} binds leading and wrapped fragments to its numbered column`, () => {
const r = replay();
expect(r.context.pending?.editDigest).toBeUndefined();
const before = Array.from({ length: line - 2 }, (_, n) => `Context ${n}`)
.concat('Head before crop tail', 'Old complete row', 'Context').join('\n');
fs.writeFileSync(r.file, before);
const at = new Date(Date.parse(r.context.pending!.timestamp) - 1); fs.utimesSync(r.file, at, at);
const menu = r.screen.slice(r.screen.indexOf(' Do you want'));
const rows = `${continuation}-tail\n${numbered}-Old complete\n${continuation}- row\n` +
`${numbered}+New complete\n${continuation}+ row\n${next} Context\n`;
const pane = rows + '╌'.repeat(20) + '\n' + menu;
r.screen = pane;
expect(pick(r)?.input).toBe('1\r');
expect(pick(r, new Set([pick(r)!.signature]))).toBeNull();
for (const invalid of [
pane.replaceAll(continuation + '-', continuation.slice(1) + '-'),
pane.replaceAll(continuation + '-', ' ' + continuation + '-'),
pane.replace(continuation + '- row', continuation + '+ row'),
pane.replace(next + ' Context', ' ' + next + ' Context'),
pane.replace('Old complete', 'Unrelated deleted'),
pane.replace(continuation + '-tail', continuation + '-foreign suffix'),
pane.replaceAll(numbered, ' 0 '),
]) { r.screen = invalid; expect(pick(r), invalid).toBeNull(); }
});
test.skipIf(process.platform==='win32')('real launcher installs only opt-in owned hooks and removes records on close or early exit',async()=>{
const root=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-artifact-launch-'));roots.push(root);
const fake=path.join(root,'fake-claude');fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import * as fs from 'node:fs';
fs.writeFileSync(process.env.ARTIFACT_RECORD,JSON.stringify({pid:process.pid,args:process.argv.slice(2)}));
if(process.env.ARTIFACT_FAIL==='1')process.exit(19);
process.stdout.write('ARTIFACT_READY\n');process.stdin.resume();
`,{mode:0o755});
const runner=pathToFileURL(path.join(import.meta.dir,'helpers/claude-pty-runner.ts')).href;
for(const variant of ['enabled','approval','disabled','explicit-home','explicit-config','early-exit']){
const cwd=path.join(root,variant);fs.mkdirSync(cwd);const result=path.join(cwd,'result.json');
const extra=variant==='explicit-home'?{HOME:cwd}:variant==='explicit-config'?{CLAUDE_CONFIG_DIR:cwd}:{};
const worker=path.join(cwd,'worker.ts');fs.writeFileSync(worker,`
import * as fs from 'node:fs';
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(runner)};
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('fake binding');
const session=await launchClaudePty({cwd:${JSON.stringify(cwd)},seedSkills:true,observeAutoplanArtifacts:${variant!=='disabled'},approveAutoplanArtifactEdits:${variant==='approval'||variant.startsWith('explicit-')},timeoutMs:8000,
env:${JSON.stringify({...extra,ARTIFACT_RECORD:result,ARTIFACT_FAIL:variant==='early-exit'?'1':'0'})}});
try {try{await session.waitFor('ARTIFACT_READY',{timeoutMs:2000,pollMs:20});}catch(e){if(${variant!=='early-exit'})throw e;}
const r=JSON.parse(fs.readFileSync(${JSON.stringify(result)},'utf8'));
r.file=session.pendingAutoplanArtifactFile??null;r.stateRoot=session.hermeticSkillStateRoot??null;
r.canStart=typeof session.startAutoplanArtifactEditApproval==='function';
if(${variant==='approval'}){const start=Date.now();session.startAutoplanArtifactEditApproval(start);r.started=JSON.parse(fs.readFileSync(r.file,'utf8')).approvalStartedAt===start;}
r.exists=r.file?fs.existsSync(r.file):false;fs.writeFileSync(${JSON.stringify(result)},JSON.stringify(r));
} finally {await session.close();}
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,EVALS_HERMETIC:'1'},stdout:'pipe',stderr:'pipe'});
const timer=setTimeout(()=>child.kill('SIGKILL'),15000);
try{const [code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);expect(code,out+err).toBe(0);}finally{clearTimeout(timer);}
const resultData=JSON.parse(fs.readFileSync(result,'utf8')),enabled=['enabled','approval','early-exit'].includes(variant);
expect(resultData.canStart).toBe(variant==='approval');
if(variant==='approval')expect(resultData.started).toBe(true);
expect(Boolean(resultData.file)).toBe(enabled);expect(resultData.exists).toBe(enabled);
if(enabled){const settings=JSON.parse(resultData.args[resultData.args.indexOf('--settings')+1]);
expect(Object.keys(settings.hooks).sort()).toEqual(['PostToolUse','PostToolUseFailure','PreToolUse']);
for(const entries of Object.values(settings.hooks) as any[]){expect(entries).toHaveLength(1);expect(entries[0].matcher).toBe('^(Write|Edit)$');expect(entries[0].hooks[0].timeout).toBe(5);expect(entries[0].hooks[0].command).toContain(resultData.stateRoot);expect(entries[0].hooks[0].command.includes('--approve-edits')).toBe(variant==='approval');}
expect(fs.existsSync(resultData.file)).toBe(false);
}else expect(resultData.args).not.toContain('--settings');
expect(()=>process.kill(resultData.pid,0)).toThrow();
}
},90000);
+241
View File
@@ -0,0 +1,241 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { spawnSync } from 'node:child_process';
import { createPendingQuestionRecorder, readPendingQuestion, recordPendingQuestion } from './helpers/plan-count-pending-question';
import { readPlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/autoplan-routing-manual-skills-ac.json';
// Exact retained public question; hook envelopes and owned temp paths are
// synthetic controls. AC retry 2's unpublished native payload is unknown.
function fixture() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), "pending-auq-' quote $()-"));
const cwd = path.join(root, 'repo with spaces');
const config = path.join(root, 'config');
const project = path.join(config, 'projects', 'fixture');
fs.mkdirSync(cwd, { recursive: true });
fs.mkdirSync(project, { recursive: true });
const session = captured.call.sessionId;
const transcriptPath = path.join(project, `${session}.jsonl`);
const startedAt = Date.now();
fs.writeFileSync(transcriptPath, JSON.stringify({ cwd, sessionId: session, isSidechain: false,
timestamp: new Date().toISOString(), message: { role: 'assistant', content: [{ type: 'text', text: 'Review setup.' }] },
}) + '\n');
const recorder = createPendingQuestionRecorder(cwd, config);
const event = (kind = 'PreToolUse', id = captured.call.toolUseId) => ({ hook_event_name: kind,
tool_name: 'AskUserQuestion', session_id: session, tool_use_id: id, cwd,
transcript_path: transcriptPath, tool_input: { questions: structuredClone(captured.call.questions) },
});
const write = (value: unknown) => recordPendingQuestion(typeof value === 'string' ? value : JSON.stringify(value), recorder.file, cwd, config);
const transcript = () => readPlanCountTranscript(config, cwd);
const read = (native = transcript(), after = startedAt) => readPendingQuestion(recorder.file, cwd, config, after, native);
const dispose = () => { recorder.dispose(); fs.rmSync(root, { recursive: true, force: true }); };
return { root, cwd, config, recorder, startedAt, session, transcriptPath, event, write, read, transcript, dispose };
}
describe('opt-in pending native AskUserQuestion capture', () => {
test('a scoped request supplies pending identity, never an answer or coverage', () => {
const f = fixture();
try {
expect(f.read()).toBeUndefined();
f.write(f.event());
expect(f.read()).toMatchObject({ sessionId: f.session, toolUseId: captured.call.toolUseId,
source: 'pre_tool_use', answered: false, questions: captured.call.questions });
expect(f.read()?.answers).toBeUndefined();
expect(f.read()?.answeredAt).toBeUndefined();
} finally { f.dispose(); }
});
test.each(['PostToolUse', 'PostToolUseFailure'])('%s closes a request and a replay cannot reopen it', kind => {
const f = fixture();
try {
f.write(f.event());
expect(f.read()).toBeDefined();
f.write(f.event(kind));
expect(f.read()).toBeUndefined();
f.write(f.event());
expect(f.read()).toBeUndefined();
f.write(f.event('PreToolUse', 'toolu_new_owned_question'));
expect(f.read()?.toolUseId).toBe('toolu_new_owned_question');
} finally { f.dispose(); }
});
test('a duplicate pending request does not refresh its evidence timestamp', async () => {
const f = fixture();
try {
f.write(f.event());
expect(f.read()).toBeDefined();
const before = fs.readFileSync(f.recorder.file, 'utf8');
await Bun.sleep(5);
f.write(f.event());
expect(fs.readFileSync(f.recorder.file, 'utf8')).toBe(before);
} finally { f.dispose(); }
});
test('a completion observed before its request cannot reopen', () => {
const f = fixture();
try {
f.write(f.event('PostToolUse'));
f.write(f.event());
expect(f.read()).toBeUndefined();
} finally { f.dispose(); }
});
test('concurrent calls and conflicting same-call payloads poison the ambiguous pending capture', () => {
for (const sameId of [false, true]) {
const f = fixture();
try {
f.write(f.event());
const next = f.event('PreToolUse', sameId ? captured.call.toolUseId : 'toolu_other_pending');
if (sameId) next.tool_input.questions[0]!.question += ' Changed.';
f.write(next);
expect(f.read()).toBeUndefined();
f.write(f.event('PostToolUse'));
f.write(f.event('PreToolUse', 'toolu_later'));
expect(f.read()).toBeUndefined();
} finally { f.dispose(); }
}
});
test.each(['cwd', 'session', 'subagent'])('foreign %s cannot replace or clear the owned request', field => {
const f = fixture();
try {
f.write(f.event());
const prior = f.read();
for (const kind of ['PreToolUse', 'PostToolUse', 'PostToolUseFailure']) {
const event: Record<string, unknown> = f.event(kind);
if (field === 'cwd') event.cwd = path.join(f.root, 'foreign');
if (field === 'session') {
event.session_id = '3f7e6255-331b-4c3f-b5b6-cc9481be0548';
event.transcript_path = path.join(path.dirname(f.transcriptPath), `${event.session_id}.jsonl`);
fs.writeFileSync(event.transcript_path as string, '');
}
if (field === 'subagent') event.agent_id = 'foreign-subagent';
f.write(event);
expect(f.read()).toEqual(prior);
}
} finally { f.dispose(); }
});
test.each(['empty questions', 'too many questions', 'too few options', 'too many options', 'invalid option',
'oversize', 'outside transcript path', 'wrong tool name', 'missing cwd'])('malformed scoped %s fails closed', kind => {
const f = fixture();
try {
f.write(f.event());
const event = f.event();
const question = event.tool_input.questions[0]!;
if (kind === 'empty questions') event.tool_input.questions = [];
if (kind === 'too many questions') event.tool_input.questions = Array.from({ length: 5 }, () => structuredClone(question));
if (kind === 'too few options') question.options.splice(1);
if (kind === 'too many options') question.options = Array.from({ length: 5 }, (_, i) => ({ label: `Option ${i}`, description: '' }));
if (kind === 'invalid option') question.options[0]!.label = '';
if (kind === 'oversize') question.question = 'x'.repeat(128 * 1024);
if (kind === 'outside transcript path') event.transcript_path = path.join(f.root, `${f.session}.jsonl`);
if (kind === 'wrong tool name') event.tool_name = 'Write';
if (kind === 'missing cwd') delete (event as Partial<typeof event>).cwd;
f.write(event);
expect(f.read()).toBeUndefined();
f.write(f.event());
expect(f.read()).toBeUndefined();
} finally { f.dispose(); }
});
test('a late result for another call cannot clear the current pending request', () => {
const f = fixture();
try {
f.write(f.event());
const current = f.read();
f.write(f.event('PostToolUse', 'toolu_earlier_call'));
expect(f.read()).toEqual(current);
f.write(f.event('PreToolUse', 'toolu_earlier_call'));
expect(f.read()).toEqual(current);
} finally { f.dispose(); }
});
test('read requires the same ready native session, isolated directory and current epoch', () => {
const f = fixture();
try {
f.write(f.event());
const native = f.transcript();
expect(f.read(native)).toBeDefined();
expect(f.read(native, Date.now() + 1_000)).toBeUndefined();
for (const status of ['missing', 'error'] as const) {
expect(f.read({ ...native, status })).toBeUndefined();
}
const foreign = structuredClone(native);
foreign.assistantMessages[0]!.sessionId = '3f7e6255-331b-4c3f-b5b6-cc9481be0548';
expect(f.read(foreign)).toBeUndefined();
const mixed = structuredClone(native);
mixed.assistantMessages.push({ ...foreign.assistantMessages[0]! });
expect(f.read(mixed)).toBeUndefined();
expect(readPendingQuestion(undefined, f.cwd, f.config, f.startedAt, native)).toBeUndefined();
expect(readPendingQuestion(f.recorder.file, f.cwd, null, f.startedAt, native)).toBeUndefined();
expect(readPendingQuestion(f.recorder.file, path.join(f.root, 'other'), f.config, f.startedAt, native)).toBeUndefined();
} finally { f.dispose(); }
});
test('published native identity takes precedence even while unanswered or failed', () => {
const f = fixture();
try {
f.write(f.event());
for (const state of [{ answered: false }, { answered: true }, { answered: false, failed: true }]) {
const native = f.transcript();
native.calls.push({ ...structuredClone(captured.call), sessionId: f.session, ...state });
expect(f.read(native)).toBeUndefined();
}
} finally { f.dispose(); }
});
test('generated shell hooks are exact, bounded, correctly quoted and silent on valid or broken input', () => {
const f = fixture();
try {
expect(Object.keys(f.recorder.hooks).sort()).toEqual(['PostToolUse', 'PostToolUseFailure', 'PreToolUse']);
for (const kind of ['PreToolUse', 'PostToolUse', 'PostToolUseFailure'] as const) {
const matchers = f.recorder.hooks[kind];
expect(matchers).toHaveLength(1);
expect(matchers[0]!.matcher).toBe('^AskUserQuestion$');
expect(matchers[0]!.hooks).toHaveLength(1);
const hook = matchers[0]!.hooks[0]!;
expect(hook.timeout).toBe(5);
for (const input of [JSON.stringify(f.event(kind)), '{invalid json']) {
const result = spawnSync('bash', ['-c', hook.command], { cwd: f.cwd, input, encoding: 'utf8', timeout: 6_000 });
expect(result.error).toBeUndefined();
expect(result.status).toBe(0);
expect(result.stdout).toBe('');
}
}
} finally { f.dispose(); }
});
test('the helper and new free test select only the two opted-in workflows', () => {
for (const file of ['test/helpers/plan-count-pending-question.ts', 'test/autoplan-pending-question.test.ts']) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected.sort()).toEqual(['autoplan-chain-pty', 'plan-ceo-mode-routing']);
}
});
test('a hook whose input never ends closes within its own bound and remains silent', async () => {
const f = fixture();
const hook = f.recorder.hooks.PreToolUse[0]!.hooks[0]!;
const child = Bun.spawn(['bash', '-c', ['exec', hook.command].join(' ')], { cwd: f.cwd,
stdin: 'pipe', stdout: 'pipe', stderr: 'pipe' });
// Use an actual shell command argument; no fixture string is interpolated.
let forced = false;
const timer = setTimeout(() => { forced = true; child.kill('SIGKILL'); }, 6_000);
try {
const [code, stdout] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
expect(forced).toBe(false);
expect(code).toBe(0);
expect(stdout).toBe('');
expect(f.read()).toBeUndefined();
expect(fs.existsSync(f.recorder.file + '.invalid')).toBe(true);
} finally {
clearTimeout(timer);
child.stdin.end();
child.kill('SIGKILL');
await child.exited;
f.dispose();
}
}, 7_000);
});
+78
View File
@@ -0,0 +1,78 @@
import { expect, test } from 'bun:test';
import fixture from './fixtures/autoplan-phase-dash-ao.json';
import { autoplanPhaseCompletions } from './helpers/autoplan-phase-observer';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
const at = Date.parse(fixture.message.timestamp);
const transcript = (text = fixture.message.text): PlanCountTranscript => ({
status: 'ready', calls: [], assistantMessages: [{ ...fixture.message, text }],
});
const hits = (text: string) => autoplanPhaseCompletions(transcript(text), at - 1);
test('exact owned DX dash declaration adds only DX at its native timestamp', () => {
expect(hits(fixture.message.text)).toEqual([{ phase: 2.5, ts: at }]);
const all = autoplanPhaseCompletions({ status: 'ready', calls: [],
assistantMessages: fixture.orderedMessages }, fixture.commandLowerBound);
expect(all).toEqual([...fixture.actualHits, { phase: 2.5, ts: at }]);
expect(all.map(hit => hit.phase)).toEqual([1, 2, 2.5]);
});
test('em and en dash spacing share the existing completed declaration forms', () => {
for (const dash of ['—', '']) for (const before of ['', ' ']) for (const after of ['', ' ']) {
expect(hits(fixture.message.text.replace('complete—', `complete${before}${dash}${after}`)))
.toEqual([{ phase: 2.5, ts: at }]);
for (const phase of [1, 2, 2.5, 3]) for (const state of ['complete', 'completed', 'done', 'finished', 'wrapped up']) {
expect(hits(`Phase ${phase} is ${state}${before}${dash}${after}Work retained.`))
.toEqual([{ phase, ts: at }]);
}
}
expect(hits('**Phase 2.5 complete** — Work retained.')).toEqual([{ phase: 2.5, ts: at }]);
});
test('dash continuations cannot turn a conditional, quotation, question or denial into completion', () => {
for (const dash of ['—', '']) for (const tail of [
'', 'if approved.', 'unless the checks fail.', 'when review finishes.',
'once the reviewer signs off.', 'pending final checks.', 'maybe tomorrow.',
'perhaps it is complete.', 'would be complete after review.',
'not complete yet.', 'the phase is not complete.', 'this completion is withdrawn.',
'actually never finished.', 'this completion is superseded.',
'provided the remaining checks pass.', 'this completion is rejected.',
'the completion announcement is retracted.', 'actually incomplete.',
'the review remains pending.', 'Work retained?', 'is this complete?',
'Source excerpt: Work retained.', 'Earlier review: Work retained.',
'the historical example says work is retained.', '"Work retained."',
]) expect(hits(`Phase 2.5 complete ${dash} ${tail}`), tail).toEqual([]);
for (const text of [
'If approved, Phase 2.5 complete—Work retained.',
'Phase 2.5 is not complete—Work retained.',
'Phase 2.5 complete?—Work retained.',
'> Phase 2.5 complete—Work retained.',
'"Phase 2.5 complete—Work retained."',
'Source excerpt:\nPhase 2.5 complete—Work retained.',
'Example:\nPhase 2.5 complete—Work retained.\nPhase 3 complete—Work retained.',
'```text\nPhase 2.5 complete—Work retained.\n```',
' Phase 2.5 complete—Work retained.',
'# Phase 2.5 complete—Work retained.',
'Phase 2.5 (Eng review) complete—Work retained.',
]) expect(hits(text), text).toEqual([]);
});
test('dash support keeps ready/current native evidence and first-hit ordering', () => {
for (const status of ['missing', 'error'] as const) {
expect(autoplanPhaseCompletions({ ...transcript(), status }, at - 1)).toEqual([]);
}
expect(autoplanPhaseCompletions(transcript(), at + 1)).toEqual([]);
expect(autoplanPhaseCompletions({ ...transcript(), assistantMessages: [
{ ...fixture.message, timestamp: 'invalid' },
] }, at - 1)).toEqual([]);
const later = { ...fixture.message, timestamp: new Date(at + 1).toISOString() };
expect(autoplanPhaseCompletions({ ...transcript(), assistantMessages: [later, fixture.message] }, at - 1))
.toEqual([{ phase: 2.5, ts: at }]);
});
test('dash fixture and regression select only the existing AP owner', () => {
for (const file of ['test/autoplan-phase-dash-ao.test.ts', 'test/fixtures/autoplan-phase-dash-ao.json']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
}
});
+300
View File
@@ -0,0 +1,300 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
import { autoplanPhaseCompletions } from './helpers/autoplan-phase-observer';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const START = Date.parse('2026-09-08T16:00:00.000Z');
const transcript = (...messages: Array<[number, string]>): PlanCountTranscript => ({
status: 'ready', calls: [],
assistantMessages: messages.map(([ms, text]) => ({
sessionId: 'autoplan-fixture', timestamp: new Date(START + ms).toISOString(), text,
})),
});
describe('native autoplan phase observation', () => {
test('AF public wrapped-up announcement retains its actual Phase 1 timestamp', () => {
// Exact reader-projected public parent narration from the AF run; no raw
// signature or private block is stored. The next-phase mention adds no hit.
const announcement = {
sessionId: '6ce27a04-7422-4687-9de9-138e68f308d8',
text: 'Phase 1 wrapped up: 11 findings from the Claude subagent, 30 of 34 spec issues fixed after 3 review rounds, and 21 obligations carried forward with 3 disagreements flagged as taste items. Moving on to Phase 2 (design review) now that UI scope was detected.\n\n',
timestamp: '2026-09-10T00:00:03.893Z',
};
const at = Date.parse(announcement.timestamp);
expect(autoplanPhaseCompletions({ status: 'ready', calls: [],
assistantMessages: [announcement] }, at - 1)).toEqual([{ phase: 1, ts: at }]);
});
test('AF bare done announcement retains Phase 2 without crediting its DX transition', () => {
const announcement2 = {
sessionId: "6ce27a04-7422-4687-9de9-138e68f308d8",
text: "Phase 2 done: Claude subagent found 15 issues (1 critical, 8 high, 6 medium), with 14 fully accepted and 1 partially accepted; design score rose from 6/10 to 8.7/10, and 10 accepted items are now carried into the Implementation plan. Moving on to Phase 2.5 (DX Review) since developer-facing scope was detected.",
timestamp: "2026-09-10T00:08:31.892Z",
};
const at = Date.parse(announcement2.timestamp);
expect(autoplanPhaseCompletions({ status: 'ready', calls: [],
assistantMessages: [announcement2] }, at - 1)).toEqual([{ phase: 2, ts: at }]);
});
test('AF named DX completion retains Phase 2.5 without crediting its Eng transition', () => {
const announcement3 = {
sessionId: "6ce27a04-7422-4687-9de9-138e68f308d8",
text: "Phase 2.5 (DX review) is done: DX score rose from 5.1 to 8.0/10, all 19 findings reviewed with 12 items accepted into the plan and one taste item flagged for gating. All pre-checks pass, so I'm moving on to Phase 3, the final Engineering Review of the amended plan.\n\n",
timestamp: "2026-09-10T00:16:29.888Z",
};
const at = Date.parse(announcement3.timestamp);
expect(autoplanPhaseCompletions({ status: 'ready', calls: [],
assistantMessages: [announcement3] }, at - 1)).toEqual([{ phase: 2.5, ts: at }]);
});
test('completed-state declarations retain supported phases, punctuation and first native time', () => {
for (const state of ['wrapped up', 'done']) for (const phase of [1, 2, 2.5, 3]) for (const tail of ['', '.', ': Work retained.', '. Work retained.']) {
expect(autoplanPhaseCompletions(transcript([1, `Phase ${phase} ${state}${tail}`]), START))
.toEqual([{ phase, ts: START + 1 }]);
}
expect(autoplanPhaseCompletions(transcript([1, '**Phase 1 wrapped up.**']), START))
.toEqual([{ phase: 1, ts: START + 1 }]);
expect(autoplanPhaseCompletions(transcript([4, 'Phase 1 wrapped up.'],
[2, 'Phase 3 wrapped up.'], [3, 'Phase 1 wrapped up: Moving to Phase 2.']), START))
.toEqual([{ phase: 3, ts: START + 2 }, { phase: 1, ts: START + 3 }]);
});
test('affirmative completion words share punctuation and optional is without future tense', () => {
for (const state of ['complete', 'completed', 'done', 'finished', 'wrapped up']) {
for (const copula of ['', 'is ']) for (const tail of ['', '.', ': Work retained.']) {
expect(autoplanPhaseCompletions(transcript([1, `Phase 3 ${copula}${state}${tail}`]), START))
.toEqual([{ phase: 3, ts: START + 1 }]);
}
for (const text of [`Phase 3 ${state}?`, `Phase 3 will be ${state}.`,
`Phase 3 is not ${state}.`, `Phase 3 ${state} if the reviewer finishes.`,
`Phase 3 ${state} when the work ends.`, `**Phase 3 ${state}** if approved.`,
`Example:\nPhase 3 ${state}.`, `> Phase 3 ${state}.`,
`Phase 3 (Eng review) is not ${state}.`, `Phase 3 (Eng review) ${state} if approved.`]) {
expect(autoplanPhaseCompletions(transcript([1, text]), START), text).toEqual([]);
}
}
});
test('known phase names must agree with their number and never supply completion alone', () => {
const names = [[1, 'CEO'], [2, 'Design'], [2.5, 'DX'], [3, 'Eng'], [3, 'Engineering']] as const;
for (const [phase, name] of names) for (const suffix of ['', ' review']) {
const declaration = `Phase ${phase} (${name}${suffix}) is finished.`;
expect(autoplanPhaseCompletions(transcript([1, declaration]), START)).toEqual([{ phase, ts: START + 1 }]);
expect(autoplanPhaseCompletions(transcript([1, `**${declaration}**`]), START)).toEqual([{ phase, ts: START + 1 }]);
for (const other of [1, 2, 2.5, 3].filter(n => n !== phase)) {
expect(autoplanPhaseCompletions(transcript([1, declaration.replace(`Phase ${phase}`, `Phase ${other}`)]), START))
.toEqual([]);
}
}
for (const text of ['Phase 3 (Eng review).', 'Phase 2.5 (future DX review) is done.',
'Phase 2.5 (DX review if approved) is done.', 'Phase 1 (source) complete.',
'Phase 2.5 ((DX review)) is done.', 'Phase 2.5 (DX review) finished soon.',
'# Phase 2.5 (DX review) is done.', 'Example:\nPhase 2.5 (DX review) is done.']) {
expect(autoplanPhaseCompletions(transcript([1, text]), START), text).toEqual([]);
}
});
test('future, conditional, negative and quoted wrap-up claims do not complete a phase', () => {
for (const text of [
'Phase 1 will wrap up.', 'Phase 1 has not wrapped up.', 'Phase 1 is not wrapped up.',
'Phase 1 wrapped up if the reviewer finishes.', 'Phase 1 wrapped up when the review ends.',
'Phase 1 wrapped up but is not complete.', 'Phase 1 wrapped up?',
'Once Phase 1 wrapped up, we would start Phase 2.', 'I will announce Phase 1 wrapped up.',
'**Phase 1 wrapped up** if the tests pass.', 'Phase 4 wrapped up.', 'Phase 2.1 wrapped up.',
'# Phase 1 wrapped up.', '> Phase 1 wrapped up.', '"Phase 1 wrapped up."',
'- Phase 1 wrapped up.', '| Phase 1 wrapped up. |', ' Phase 1 wrapped up.',
'```text\nPhase 1 wrapped up.\n```', '~~~text\nPhase 1 wrapped up.\n~~~',
'Example:\nPhase 1 wrapped up.\nPhase 2 wrapped up.',
'The template says:\n\nPhase 1 wrapped up.',
'**Phase 1 wrapped up.** Emit phase-transition summary:',
]) for (const declaration of [text, text.replace(/wrapped up/g, 'done')]) {
expect(autoplanPhaseCompletions(transcript([1, declaration]), START), declaration).toEqual([]);
}
});
test('wrapped-up declarations retain ready transcript and native timestamp requirements', () => {
const current = transcript([1, 'Phase 1 wrapped up.']);
for (const status of ['missing', 'error'] as const) {
expect(autoplanPhaseCompletions({ ...current, status }, START)).toEqual([]);
}
expect(autoplanPhaseCompletions(current, START + 2)).toEqual([]);
expect(autoplanPhaseCompletions({ ...current, assistantMessages: current.assistantMessages.map(
message => ({ ...message, timestamp: 'invalid' })) }, START)).toEqual([]);
});
test('retains actual completion timestamps when several phases arrive between polls', () => {
expect(autoplanPhaseCompletions(transcript(
[1, '**Phase 1 complete.** Codex: 2 concerns. Native: 3 issues.'],
[2, 'Phase 2 complete. Design outputs are in the plan.'],
[3, '**Phase 2.5 complete.** DX overall: 8/10.'],
[4, 'Phase 3 complete. Both engineering reviews finished.'],
), START)).toEqual([1, 2, 2.5, 3].map((phase, index) => ({ phase, ts: START + index + 1 })));
});
test('duplicate announcements retain their first timestamp without reordering phases', () => {
expect(autoplanPhaseCompletions(transcript(
[3, 'Phase 1 complete.'], [2, 'Phase 2 complete.'],
[1, '**Phase 1 complete.**'], [4, '**Phase 3 complete.**'],
), START)).toEqual([{ phase: 1, ts: START + 1 }, { phase: 2, ts: START + 2 }, { phase: 3, ts: START + 4 }]);
});
test('headings, quoted skill text, examples, tables and planned checklists provide no completion', () => {
for (const text of [
'## Phase 3 complete.',
'**PHASE 3 COMPLETE.** Emit phase-transition summary:',
'> **Phase 3 complete.** Codex: [N concerns].',
'The section says "**Phase 3 complete.**".',
'Example: **Phase 3 complete.**',
'Example announcement:\n**Phase 3 complete.**',
'Example announcement:\nPhase 2 complete.\nPhase 3 complete.',
'The template requires this completion marker:\n\n**Phase 3 complete.**',
' **Phase 3 complete.**',
'\tPhase 3 complete.',
'```markdown\n**Phase 3 complete.**\n```',
'```markdown\n```still-code\nPhase 3 complete.\n```',
'````markdown\n```\nPhase 3 complete.\n````',
'~~~markdown\nPhase 3 complete.\n~~~',
'```markdown\nPhase 3 complete.',
'| **Phase 3 complete.** | pending |',
'- [ ] **Phase 3 complete.**',
'1. Phase 3 complete.',
'I will announce **Phase 3 complete.** after the review.',
'Phase 3 complete when the engineering review ends.',
'Phase 3 complete?',
'Phase 4 complete.',
]) expect(autoplanPhaseCompletions(transcript([1, text]), START), text).toEqual([]);
});
test('a real announcement after a closed example fence still establishes completion', () => {
expect(autoplanPhaseCompletions(transcript([1, 'Example:\n```markdown\nPhase 3 complete.\n```\n\n**Phase 1 complete.**']), START))
.toEqual([{ phase: 1, ts: START + 1 }]);
});
test('accepts a template-compliant quoted transition with actual consensus while rejecting quoted examples', () => {
const actual = '> **Phase 1 complete.** Codex: 2 concerns. Claude subagent: 3 issues.\n' +
'> Consensus: 4/6 confirmed, 2 disagreements → surfaced at gate.\n> Passing to Phase 2.';
expect(autoplanPhaseCompletions(transcript([1, actual]), START)).toEqual([{ phase: 1, ts: START + 1 }]);
expect(autoplanPhaseCompletions(transcript([1, actual.replace('Codex: 2 concerns.', 'Codex: unavailable.')]), START))
.toEqual([{ phase: 1, ts: START + 1 }]);
for (const quoted of [
actual.replace('4/6', 'X/6'),
actual.replace('2 concerns', '[N concerns]'),
'The template contains this example:\n' + actual,
'Example announcement:\n' + actual + '\n' + actual.replace('Phase 1', 'Phase 3'),
'> **Phase 3 complete.**',
]) expect(autoplanPhaseCompletions(transcript([1, quoted]), START), quoted).toEqual([]);
});
test('missing/error transcripts and pre-command declarations cannot establish coverage', () => {
const prior = transcript([-1, '**Phase 3 complete.**']);
expect(autoplanPhaseCompletions(prior, START)).toEqual([]);
for (const status of ['missing', 'error'] as const) {
expect(autoplanPhaseCompletions({ ...transcript([1, '**Phase 3 complete.**']), status }, START)).toEqual([]);
}
});
test('does not repair an out-of-order chain or synthesize omitted phases', () => {
expect(autoplanPhaseCompletions(transcript([1, 'Phase 3 complete.'], [2, 'Phase 1 complete.']), START))
.toEqual([{ phase: 3, ts: START + 1 }, { phase: 1, ts: START + 2 }]);
expect(autoplanPhaseCompletions(transcript([1, 'Phase 2 skipped — no UI scope.']), START)).toEqual([]);
});
test('phase observer changes select the autoplan eval', () => {
for (const file of ['test/helpers/autoplan-phase-observer.ts', 'test/autoplan-phase-observer.test.ts']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
}
});
test.skipIf(process.platform === 'win32')('ANSI-rendered completions use native evidence while displayed Read/source markers do not', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-phase-replay-'));
const fake = path.join(dir, 'fake-claude');
const worker = path.join(dir, 'worker.ts');
const recordFile = path.join(dir, 'events.jsonl');
const resultFile = path.join(dir, 'result.json');
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
import * as fs from 'node:fs';
import * as path from 'node:path';
fs.writeFileSync(process.env.PHASE_RECORD, JSON.stringify({pid:process.pid}) + '\n');
const sessionId = 'fake-autoplan-session';
const dir = path.join(process.env.CLAUDE_CONFIG_DIR, 'projects', 'fixture');
fs.mkdirSync(dir, {recursive:true});
const base = Date.now();
const write = (role, content, offset, extra = {}) => fs.appendFileSync(path.join(dir, sessionId + '.jsonl'), JSON.stringify({
sessionId, cwd:process.cwd(), isSidechain:false, timestamp:new Date(base + offset).toISOString(), message:{role,content}, ...extra,
}) + '\n');
write('user', [{type:'tool_result', tool_use_id:'read', content:'**Phase 3 complete.**'}], 0);
write('assistant', [{type:'text', text:'> **Phase 3 complete.** is the quoted source marker.'}], 0);
write('assistant', [{type:'text', text:'**Phase 3 wrapped up.**'}], 0, {isSidechain:true});
write('assistant', [{type:'text', text:'**Phase 3 wrapped up.**'}], 0, {cwd:path.join(process.cwd(), 'foreign')});
process.stdin.setRawMode?.(true);
let sent = false;
process.stdin.on('data', data => {
if (sent || !data.toString().includes('\r')) return;
sent = true;
process.stdout.write('\x1b[2J\x1b[H');
[1, 2, 2.5, 3].forEach((phase, index) => {
const message = '**Phase ' + phase + (phase === 1 ? ' wrapped up.**' : phase === 2 ? ' done.**' : phase === 2.5 ? ' (DX review) is finished.**' : ' complete.**');
write('assistant', [{type:'text', text:message}], index + 1);
process.stdout.write('● \x1b[1mPhase ' + phase + ' complete.\x1b[22m\n');
});
process.stdout.write('NATIVE_PHASES_READY\n');
});
process.stdout.write('Read: autoplan/sections/eng-phase.md\n> **Phase 3 complete.**\nSOURCE_READY\n');
process.on('SIGINT', () => process.exit(0));
process.stdin.resume();
`);
fs.chmodSync(fake, 0o755);
const moduleUrl = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
fs.writeFileSync(worker, `
import { launchClaudePty } from ${JSON.stringify(moduleUrl('claude-pty-runner.ts'))};
import { readPlanCountTranscript } from ${JSON.stringify(moduleUrl('plan-count-transcript.ts'))};
import { autoplanPhaseCompletions } from ${JSON.stringify(moduleUrl('autoplan-phase-observer.ts'))};
const start = Date.now();
const session = await launchClaudePty({cwd:${JSON.stringify(dir)}, timeoutMs:8000, env:{PHASE_RECORD:${JSON.stringify(recordFile)}}});
try {
await session.waitFor('SOURCE_READY', {timeoutMs:4000, pollMs:20});
const read = () => readPlanCountTranscript(session.hermeticConfigDir, ${JSON.stringify(dir)});
const sourceOnly = autoplanPhaseCompletions(read(), start);
const oldPattern = /\\*\\*Phase\\s+(\\d+(?:\\.\\d+)?)\\s+complete\\.?\\*\\*/g;
const sourceFalsePositive = [...session.visibleText().matchAll(oldPattern)].length;
const since = session.mark();
session.send('\\r');
await session.waitFor('NATIVE_PHASES_READY', {timeoutMs:4000, pollMs:20});
const rendered = session.visibleSince(since);
const oldPatternMissed = [...rendered.matchAll(oldPattern)].length;
const hits = autoplanPhaseCompletions(read(), start);
await Bun.write(${JSON.stringify(resultFile)}, JSON.stringify({sourceOnly, sourceFalsePositive, oldPatternMissed, hits, rendered}));
} finally { await session.close(); }
`);
const child = Bun.spawn([process.execPath, worker], {
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' }, stdout: 'pipe', stderr: 'pipe',
});
const timer = setTimeout(() => child.kill('SIGKILL'), 12_000);
try {
const [code, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
expect(code, stdout + stderr).toBe(0);
const result = JSON.parse(fs.readFileSync(resultFile, 'utf8'));
expect(result.sourceOnly).toEqual([]);
expect(result.sourceFalsePositive).toBe(1);
expect(result.oldPatternMissed).toBe(0);
expect(result.rendered).toContain('Phase 3 complete.');
expect(result.hits.map((hit: { phase: number }) => hit.phase)).toEqual([1, 2, 2.5, 3]);
for (let index = 1; index < result.hits.length; index++) {
expect(result.hits[index].ts).toBeGreaterThan(result.hits[index - 1].ts);
}
const pid = JSON.parse(fs.readFileSync(recordFile, 'utf8').split('\n')[0]!).pid;
expect(() => process.kill(pid, 0)).toThrow();
} finally {
clearTimeout(timer); child.kill('SIGKILL');
if (fs.existsSync(recordFile)) {
const pid = JSON.parse(fs.readFileSync(recordFile, 'utf8').split('\n')[0]!).pid;
try { process.kill(pid, 'SIGKILL'); } catch { /* already reaped */ }
}
fs.rmSync(dir, {recursive:true, force:true});
}
}, 15_000);
});
+169
View File
@@ -68,3 +68,172 @@ describe('autoplan phase order (Eng always last)', () => {
expect(ceo).toContain('Final');
});
});
describe('autoplan phase execution checkpoints', () => {
const tmpl = read('autoplan/SKILL.md.tmpl');
const phases = ['ceo', 'design', 'dx', 'eng'];
test('loads full review skills at phase entry instead of prefetching future phases', () => {
const intake = tmpl.split('### Step 3:')[1]?.split('## Phase 0.5:')[0] ?? '';
expect(intake).toContain("Resolve this phase's source to absolute `<REVIEW_SKILL>`; load via its checkpoint");
expect(intake).toContain('Do not prefetch future phase sections or review skills');
for (const phase of phases) {
const section = read(`autoplan/sections/${phase}-phase.md.tmpl`);
expect(section).toMatch(/^Before dispatch, Read \{\{AUTOPLAN_REVIEW_FILE:plan-[a-z-]+:with-sections\}\}/);
const load = section.split('**Override rules:**')[0]!;
expect(load).toContain('per `readRanges`');
expect(load).toContain('log successful ranges/total');
expect(load).toContain('to EOF');
expect(load).toContain('Skip-listed: load only');
expect(section.indexOf(':with-sections}}')).toBeLessThan(section.indexOf('create ' + phase));
expect(section).toContain(`create ${phase} "<ACTIVE_PLAN>" "<RESTORE_PATH>" "<methodologyPath>"`);
}
});
for (const phase of phases) {
test(`${phase} places schema-aware dispatch and the actual completion wait before outside review`, () => {
const section = read(`autoplan/sections/${phase}-phase.md.tmpl`);
const native = section.indexOf(`**{{NATIVE_LABEL}} ${phase === 'design' ? 'design' : phase === 'dx' ? 'DX' : phase === 'ceo' ? 'CEO' : 'eng'} subagent**`);
const outside = section.indexOf('{{OUTSIDE_INVOCATION:autoplan}}');
expect(native).toBeGreaterThan(-1);
expect(native).toBeLessThan(outside);
const dispatch = section.slice(native, outside);
expect(dispatch).toContain('run_in_background: false');
expect(dispatch).toContain("ONLY/FINAL tool call");
expect(dispatch).toContain('Keep native Reads enabled');
expect(dispatch).toContain('Child first Reads `nativePromptPath` to EOF');
expect(dispatch).toContain('all criteria + plan');
expect(dispatch).toContain('if its schema exposes it');
expect(dispatch).toContain('isAsync: true');
expect(dispatch).toContain('Claude Code: end response immediately');
expect(dispatch).toContain('No further tool calls/review until');
expect(dispatch).toContain('Completed-native INPUT must match snapshot phase/hash');
expect(dispatch).toContain('Retry invalid input once; then failure policy if still invalid');
expect(dispatch).toContain('Other hosts await that ID');
expect(dispatch).toContain("Then outside → this phase's review ONLY");
expect(dispatch).toContain('No inline substitute; apply failure policy');
// Provider preflight, timeout and native fallback remain at every call.
expect(section).toContain('Outer tool timeout: 720000ms');
expect(section).toContain('disabled → skip outside. Both retain the native pass.');
expect(section).toContain(`{{OUTSIDE_PROVENANCE:${phase}}}`);
expect(section).toContain(phase === 'ceo' ? 'Outside disabled/unavailable' : 'Missing/disabled');
expect(section).toContain('N/A');
expect(section).toContain('primary cannot replace');
expect(section).toContain(phase === 'design' ? 'not CONFIRMED' : 'never CONFIRMED');
});
}
test('the parent completes only the current phase and cannot waive native work for context pressure', () => {
const contract = tmpl.split('## Sequential Execution')[1]?.split('---')[0] ?? '';
expect(contract).toContain('Keep ONE phase active');
expect(contract).toContain('Never draft future-phase reviews or outputs');
expect(contract).toContain('After compaction, reload current phase instructions/skill/sections; reconcile disk progress before resuming');
expect(contract).toContain('Load its phase instructions and full skill/sections');
expect(contract).toContain('Create the fresh snapshot and dispatch its nativeDispatchPrompt unchanged');
expect(contract).toContain('Consume native completion, then enabled outside results; only then do the full primary review');
expect(contract).toContain("Persist outputs/amendments and run the phase's implementation check/readback");
expect(contract).toContain('Send the phase completion summary as a standalone user-facing message');
expect(contract).toContain("Only then make the next phase's tool calls");
expect(contract).toContain('for Eng, send it before final synthesis and the approval question');
expect(contract).toContain('A missing gate means the current phase remains open');
expect(contract).toContain('Read requests/self-reports and INPUT hashes do not prove uptake or review quality');
expect(contract).toContain('Pending is not unavailable');
expect(contract).toContain('Time/context pressure or your own review never permits\nskipping native passes or required sections');
expect(contract).toContain('Never read raw agent transcripts');
expect(tmpl).toContain('LOG each decision; record ALL accepted obligations below and run `amend` before continuing');
});
test('each completed phase announces only after persisted full outputs and settled reviewers', () => {
for (const [phase, number] of [['ceo', '1'], ['design', '2'], ['dx', '2.5'], ['eng', '3']]) {
const section = read(`autoplan/sections/${phase}-phase.md.tmpl`);
const barrier = section.indexOf('**Close this phase:**');
const announcement = section.indexOf(`\n**Phase ${number} complete.**\n`);
expect(barrier).toBeGreaterThan(-1);
expect(barrier).toBeLessThan(announcement);
const checkpoint = section.slice(barrier, announcement);
expect(checkpoint).toContain('Require full skill/section ranges');
expect(checkpoint).toContain('successful writes');
expect(checkpoint).toContain('terminal reviewers');
expect(checkpoint).toContain('matched completed-native INPUT');
expect(checkpoint).toContain('(unavailable/disabled allowed)');
expect(checkpoint).toContain('successful writes/check');
expect(checkpoint).toContain(phase === 'eng'
? 'After sending it, proceed to final synthesis/approval'
: 'After sending it, load/create/dispatch the next phase');
expect(checkpoint).toContain('EVERY accepted requirement/condition/test');
expect(checkpoint).toContain('in its block');
expect(checkpoint).toContain('Reconcile full review');
expect(checkpoint).toContain('Read back fully');
expect(checkpoint).toContain('retention ≠ approval/completeness/correctness');
expect(checkpoint).toContain('Taste provisional');
expect(checkpoint).toContain('User Challenges keep original');
expect(checkpoint).toContain(`amend ${phase} "<ACTIVE_PLAN>" "<${phase.toUpperCase()}_INPUT>"`);
expect(checkpoint).toContain('None: reason checks unchanged');
expect(checkpoint).toContain('Only then send this completion summary as a standalone user-facing message');
}
});
test('Design hands off to conditional DX and DX never requests a future Eng result', () => {
const design = read('autoplan/sections/design-phase.md.tmpl');
const dx = read('autoplan/sections/dx-phase.md.tmpl');
expect(design).toContain('Passing to Phase 2.5 (DX Review) if DX scope was detected; otherwise Phase 3');
expect(design).not.toContain('> Passing to Phase 3.');
expect(dx).toContain("Design: <insert Design consensus summary, or 'skipped, no UI scope'>");
expect(dx).not.toContain('Eng: <insert Eng consensus summary>');
});
});
describe('autoplan current implementation-plan identity', () => {
test('pins the assigned active plan and keeps accepted amendments separate from review analyses', () => {
const intake = read('autoplan/SKILL.md.tmpl').split('## Phase 0: Intake')[1]?.split('### Step 2:')[0] ?? '';
expect(intake).toContain('ACTIVE_PLAN (harness-assigned plan, else SOURCE_PLAN)');
expect(intake).toContain('Write all amendments/outputs to ACTIVE_PLAN');
expect(intake).toContain("init backs up SOURCE_PLAN exactly");
expect(intake).toContain('without losing requirements');
expect(intake).toContain('init "<SOURCE_PLAN>" "<ACTIVE_PLAN>" "<RESTORE_PATH>"');
expect(intake).toContain('Use returned paths/`scope`');
expect(intake).toContain('On helper errors, stop');
expect(intake).toContain('analysis stays in `## Review record`');
// Binding belongs to the lazy execution site, not a stale intake variable.
expect(intake).not.toContain('Bind `<review_plan_path>`');
});
test('DX scope consumes the deterministic full-input result and permits only enabling overrides', () => {
const intake = read('autoplan/SKILL.md.tmpl').split('### Step 2: Read context')[1]?.split('### Step 3:')[0] ?? '';
expect(intake).toContain('scope "<ACTIVE_PLAN>"');
expect(intake).toContain('Use returned `dxRequired`');
expect(intake).toContain('threshold is 2+ term matches');
expect(intake).toContain('`--developer-tool` or `--agent-primary`');
expect(intake).toContain('no context label can negate a positive result');
expect(intake).toContain('false and neither semantic trigger applies');
});
test('every native and outside call site binds the fresh snapshot, retaining requested outside consensus', () => {
for (const phase of ['ceo', 'design', 'dx', 'eng']) {
const section = read(`autoplan/sections/${phase}-phase.md.tmpl`);
const bind = section.indexOf("**Bind phase input:**");
const native = section.indexOf('subagent**');
const outside = section.indexOf('{{OUTSIDE_INVOCATION:autoplan}}');
expect(bind).toBeGreaterThan(-1);
expect(bind).toBeLessThan(native);
expect(native).toBeLessThan(outside);
const preparation = section.slice(bind, native);
expect(preparation).toContain(`create ${phase} "<ACTIVE_PLAN>" "<RESTORE_PATH>"`);
expect(preparation).toContain('`snapshotPath` as `<' + phase.toUpperCase() + '_INPUT>` for both voices');
expect(preparation).toContain('excludes `Review record`');
expect(section).toContain('Send `nativeDispatchPrompt` verbatim');
expect(section).toContain('Reads `nativePromptPath` to EOF');
expect(section).toContain(`Outside prompt: inline the full contents of <${phase.toUpperCase()}_INPUT>`);
expect(section).toContain(`amend ${phase} "<ACTIVE_PLAN>" "<${phase.toUpperCase()}_INPUT>"`);
expect(section).toContain('None: reason checks unchanged');
expect(read('autoplan/SKILL.md.tmpl')).toContain('checks exact retention');
expect(section).not.toContain('<review_plan_path>');
expect(section).not.toContain('<plan_path>');
expect(section).toContain('no summaries or prior reviews');
}
const eng = read('autoplan/sections/eng-phase.md.tmpl');
expect(eng).toContain('no summaries or prior reviews');
expect(eng).toContain('DX: <insert DX consensus table summary');
});
});
@@ -0,0 +1,129 @@
import { expect, test } from 'bun:test';
import { execFileSync } from 'node:child_process';
import { existsSync, mkdirSync, mkdtempSync, readFileSync, readdirSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import { seedAutoplanOnboarding } from './helpers/autoplan-preconfigured-fixture';
import { DESIGN_DOC_DISCOVERY_BLOCK } from '../scripts/resolvers/design-doc-discovery';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const root = resolve(import.meta.dir, '..');
const read = (file: string) => readFileSync(join(root, file), 'utf8');
const original = read('test/fixtures/plans/autoplan-dashboard.md');
function fixture(plan = original) {
const temp = mkdtempSync(join(tmpdir(), 'gstack-chain-onboarding-'));
const cwd = join(temp, 'project');
const home = join(temp, 'home');
const state = join(temp, 'state');
for (const dir of [join(cwd, '.claude/plans'), home, state]) mkdirSync(dir, { recursive: true });
const planFile = join(cwd, '.claude/plans/ui-heavy-feature.md');
writeFileSync(planFile, plan);
execFileSync('git', ['init', '-q', '-b', 'main'], { cwd });
// This free test controls unrelated startup side effects, not the seeded repo.
writeFileSync(join(state, 'config.yaml'), 'update_check: false\nartifacts_sync: off\ntelemetry: off\n');
writeFileSync(join(state, '.proactive-prompted'), '');
const env = { PATH: process.env.PATH!, HOME: home, GSTACK_HOME: state, GSTACK_STATE_ROOT: state };
const start = () => execFileSync(join(root, 'bin/gstack-skill-start'), ['--skill', 'autoplan'],
{ cwd, env, encoding: 'utf8', timeout: 30_000 });
const discover = () => execFileSync('bash', ['-c', DESIGN_DOC_DISCOVERY_BLOCK],
{ cwd, env: { ...env, SLUG: 'chain-fixture', BRANCH: 'main' }, encoding: 'utf8', timeout: 5000 }).trim();
return { cwd, home, state, planFile, start, discover, cleanup: () => rmSync(temp, { recursive: true, force: true }) };
}
test('real skill-start and canonical discovery see the configured chain prerequisites', () => {
const f = fixture();
try {
const before = f.start();
expect(before).toContain('SESSION_KIND: interactive');
expect(before).toContain('HAS_ROUTING: no');
expect(before).toContain('GSTACK_INSTRUCTION_BEGIN: routing-injection ');
expect(f.discover()).toBe('No design doc found');
seedAutoplanOnboarding(f.cwd);
const after = f.start();
expect(after).toContain('SESSION_KIND: interactive');
expect(after).toContain('HAS_ROUTING: yes');
expect(after).toContain('ROUTING_DECLINED: false');
expect(after).not.toContain('GSTACK_INSTRUCTION_BEGIN: routing-injection ');
expect(f.discover()).toBe(`Design doc found: ${join(f.cwd, 'docs/designs/dashboard-context.md')}`);
const offered = before.slice(before.indexOf('## Skill routing\n'), before.indexOf('\nIf B: run', before.indexOf('## Skill routing\n'))).trimEnd() + '\n';
expect(readFileSync(join(f.cwd, 'CLAUDE.md'), 'utf8')).toBe(offered);
expect(readFileSync(f.planFile, 'utf8')).toBe(original);
} finally { f.cleanup(); }
}, 60_000);
test('brief copies only existing context and contracts; all implementation work stays unreviewed', () => {
const f = fixture();
try {
const stateBefore = readdirSync(f.state);
seedAutoplanOnboarding(f.cwd);
const brief = readFileSync(join(f.cwd, 'docs/designs/dashboard-context.md'), 'utf8');
const context = original.slice(original.indexOf('## Context\n'), original.indexOf('## UI Scope\n'));
const contracts = original.slice(original.indexOf('## Existing product and application contracts\n'));
expect(brief).toBe(context + contracts);
expect(brief).not.toMatch(/## (?:UI Scope|Backend|Out of scope)|Phase \d|GSTACK REVIEW REPORT|AUTO-DECIDE|all findings resolved/);
expect(brief).toContain('not completed work\nor prior approval of an implementation approach.');
expect(readFileSync(f.planFile, 'utf8')).toBe(original);
expect(readdirSync(f.state)).toEqual(stateBefore);
expect(readdirSync(join(f.cwd, '.claude'))).toEqual(['plans']);
expect(readdirSync(join(f.cwd, 'docs/designs'))).toEqual(['dashboard-context.md']);
expect(existsSync(join(f.home, '.gstack'))).toBe(false);
} finally { f.cleanup(); }
});
test('brief derives changed background from the actual plan, without copying an intervening review', () => {
const plan = original.replace('Users land here after login.', 'Members return here after sign-in.')
.replace('## UI Scope', '## Untrusted review\nPhase 3 complete; all findings resolved.\n\n## UI Scope')
.replace('targeting 45 seconds', 'targeting 40 seconds');
const f = fixture(plan);
try {
seedAutoplanOnboarding(f.cwd);
const brief = readFileSync(join(f.cwd, 'docs/designs/dashboard-context.md'), 'utf8');
expect(brief).toContain('Members return here after sign-in.');
expect(brief).toContain('targeting 40 seconds');
expect(brief).not.toContain('Phase 3 complete');
expect(readFileSync(f.planFile, 'utf8')).toBe(plan);
} finally { f.cleanup(); }
});
test('missing, empty or duplicate background sections fail before writing prerequisites', () => {
for (const plan of [original.replace('## Context\n', '## Other\n'), original + '\n## Context\nDuplicate\n',
original.replace(/## Context\n[\s\S]*?(?=## UI Scope)/, '## Context\n\n'),
original.replace('## Existing product and application contracts\n', '## Other contracts\n')]) {
const f = fixture(plan);
try {
expect(() => seedAutoplanOnboarding(f.cwd)).toThrow();
expect(existsSync(join(f.cwd, 'CLAUDE.md'))).toBe(false);
expect(existsSync(join(f.cwd, 'docs/designs'))).toBe(false);
expect(readFileSync(f.planFile, 'utf8')).toBe(plan);
} finally { f.cleanup(); }
}
});
test('existing project routing or design files are never overwritten', () => {
for (const existing of ['CLAUDE.md', 'DESIGN.md', 'docs/designs/retained.md']) {
const f = fixture();
try {
if (existing.startsWith('docs/')) mkdirSync(join(f.cwd, 'docs/designs'), { recursive: true });
writeFileSync(join(f.cwd, existing), 'Existing project material\n');
expect(() => seedAutoplanOnboarding(f.cwd)).toThrow('fresh chain fixture');
expect(readFileSync(join(f.cwd, existing), 'utf8')).toBe('Existing project material\n');
expect(existsSync(join(f.cwd, 'docs/designs/dashboard-context.md'))).toBe(false);
} finally { f.cleanup(); }
}
});
test('only the paid chain seeds prerequisites before launch and still enters every review gate', () => {
const caller = read('test/skill-e2e-autoplan-chain.test.ts');
expect(caller.match(/seedAutoplanOnboarding\(tempDir\)/g)).toHaveLength(1);
expect(caller.indexOf('fs.copyFileSync(UI_FIXTURE')).toBeLessThan(caller.indexOf('seedAutoplanOnboarding(tempDir)'));
expect(caller.indexOf('seedAutoplanOnboarding(tempDir)')).toBeLessThan(caller.indexOf("gitRun(['add', '.'])"));
expect(caller.indexOf('seedAutoplanOnboarding(tempDir)')).toBeLessThan(caller.indexOf('launchClaudePty({'));
expect(caller).toContain("session.send('/autoplan\\r')");
expect(caller).toContain('if (!ceo || !design || !dx || !eng)');
expect(caller).toContain("for (const phase of ['ceo', 'design', 'dx', 'eng'])");
expect(caller).toContain('expect(methodologyAudit.some(audit => audit.phase === phase && audit.passed)).toBe(true)');
expect(read('test/helpers/plan-count-fixture.ts')).not.toContain('seedAutoplanOnboarding');
for (const file of ['test/helpers/autoplan-preconfigured-fixture.ts', 'test/autoplan-preconfigured-onboarding-ar.test.ts']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
}
});
+134
View File
@@ -0,0 +1,134 @@
import {expect,test} from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import {readPlanCountTranscript,type NativePublicToolEvent} from './helpers/plan-count-transcript';
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
import fixture from './fixtures/autoplan-public-narration-ad.json';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
const at=Date.parse(fixture.provenance.timestamp);
function read(blocks: unknown[]= [fixture.block],delta: any={},complete=true) {
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'public-narration-')),cwd=path.join(dir,'repo');
const project=path.join(dir,'projects','owned');fs.mkdirSync(project,{recursive:true});
const record={cwd,isSidechain:false,sessionId:fixture.provenance.sessionId,
timestamp:fixture.provenance.timestamp,uuid:fixture.provenance.uuid,
message:{role:'assistant',content:blocks},...delta};
const file=path.join(project,fixture.provenance.sessionId+'.jsonl');
fs.writeFileSync(file,JSON.stringify(record)+(complete?'\n':''));
const tools:NativePublicToolEvent[]=[];
try{return {transcript:readPlanCountTranscript(dir,cwd,event=>tools.push(event)),tools};}
finally{fs.rmSync(dir,{recursive:true,force:true});}
}
test('actual public server narration is projected at its original session and timestamp',()=>{
const {transcript,tools}=read();
expect(transcript.assistantMessages).toEqual([{sessionId:fixture.provenance.sessionId,
timestamp:fixture.provenance.timestamp,text:fixture.block.thinking}]);
expect(transcript.calls).toEqual([]);expect(tools).toEqual([]);
expect(autoplanPhaseCompletions(transcript,at-1)).toEqual([{phase:1,ts:at}]);
});
test('actual Phase 1 is done declaration independently matches the phase marker',()=>{
expect(autoplanPhaseCompletions({status:'ready',calls:[],assistantMessages:[{
sessionId:fixture.provenance.sessionId,timestamp:fixture.provenance.timestamp,
text:fixture.block.thinking}]},at-1)).toEqual([{phase:1,ts:at}]);
});
// Minimal synthetic protobuf envelopes exercise public-tag classification only.
// No opaque native signature or untagged model text is stored in this fixture.
const vi=(n:number):number[]=>{const bytes:number[]=[];do{const b=n%128;n=Math.floor(n/128);bytes.push(b+(n?128:0));}while(n);return bytes;};
const field=(n:number,body:Uint8Array)=>Buffer.from([...vi(n*8+2),...vi(body.length),...body]);
const tagged=(kind='narration')=>field(2,field(1,field(8,Buffer.from(kind))));
const summary=(signature=fixture.block.signature,thinking=fixture.block.thinking)=>({type:'thinking',thinking,signature});
const encode=(bytes:Uint8Array)=>Buffer.from(bytes).toString('base64');
test('unknown or malformed signature envelopes never become public narration',()=>{
const invalid:unknown[]=[undefined,null,'',4,'narration','%%%',' '+fixture.block.signature,
fixture.block.signature+'=',encode(tagged('reasoning')),encode(tagged('Narration')),
encode(field(1,field(1,field(8,Buffer.from('narration'))))),
encode(field(2,field(2,field(8,Buffer.from('narration'))))),
encode(field(2,field(1,field(7,Buffer.from('narration'))))),
encode(tagged().subarray(0,-1)),encode(Buffer.from([0x12,0x80])),
encode(Buffer.from([0x12,0xff,0xff,0xff,0xff,0xff,0xff,0xff,0xff,0x7f])),
encode(Buffer.concat([tagged(),Buffer.from([0])])),
encode(Buffer.concat([tagged(),Buffer.from([0x0f])])),
encode(Buffer.concat([tagged(),tagged('reasoning')])),
encode(field(2,Buffer.concat([field(1,field(8,Buffer.from('narration'))),field(1,field(8,Buffer.from('reasoning')))]))),
encode(field(2,field(1,Buffer.concat([field(8,Buffer.from('narration')),field(8,Buffer.from('reasoning'))])))),
encode(field(2,field(1,Buffer.concat([field(8,Buffer.from('narration')),field(8,Buffer.from('narration'))])))),
encode(Buffer.concat([tagged(),field(3,Buffer.alloc(64*1024))])),
];
for(const signature of invalid){
const {transcript,tools}=read([{...summary(),signature}]);
expect(transcript.assistantMessages,JSON.stringify(signature)?.slice(0,80)).toEqual([]);
expect(transcript.calls).toEqual([]);expect(tools).toEqual([]);
}
});
test('supported unrelated envelope fields preserve an exact public tag',()=>{
const bytes=Buffer.concat([Buffer.from([0x08,0x01]),tagged(),field(3,Buffer.from('opaque'))]);
expect(read([summary(encode(bytes))]).transcript.assistantMessages[0]?.text).toBe(fixture.block.thinking);
});
test('private, empty and wrong-kind blocks remain outside the public projection',()=>{
for(const block of [{type:'thinking',thinking:'SYNTHETIC_PRIVATE_TEXT'},
{...summary(),signature:encode(tagged('reasoning')),thinking:'SYNTHETIC_PRIVATE_TEXT'},
{...summary(),thinking:''},{...summary(),thinking:' '},{...summary(),thinking:7},
{...summary(),type:'redacted_thinking'},{...summary(),type:'tool_result'},
{type:'thinking',thinking:'SYNTHETIC_PRIVATE_TEXT',block_kind:'narration'}]){
expect(read([block]).transcript.assistantMessages).toEqual([]);
}
});
test('parent role, cwd, native filename and complete valid timestamp remain required',()=>{
for(const delta of [{cwd:'/foreign'},{sessionId:'foreign'},
{isSidechain:true},{isSidechain:undefined},{timestamp:'invalid'},
{timestamp:null},{message:{role:'user',content:[fixture.block]}},
{message:{role:'system',content:[fixture.block]}}]){
expect(read([fixture.block],delta).transcript.assistantMessages).toEqual([]);
}
expect(read([fixture.block],{},false).transcript.assistantMessages).toEqual([]);
});
test('public narration does not manufacture questions, replies or plan approval',()=>{
const {transcript,tools}=read([summary(),{type:'text',text:'Ordinary assistant prose.'},
{type:'thinking',thinking:'SYNTHETIC_PRIVATE_TEXT'},
{type:'tool_use',id:'actual-read',name:'Read',input:{file_path:'/owned/PLAN.md'}}]);
expect(transcript.assistantMessages.map(m=>m.text)).toEqual([fixture.block.thinking,'Ordinary assistant prose.']);
expect(transcript.calls).toEqual([]);expect(transcript.planReadyRequests).toBeUndefined();
expect(tools.map(e=>[e.kind,e.toolUseId,e.name])).toEqual([['use','actual-read','Read']]);
});
function hits(text:string,start=at-1,timestamp=fixture.provenance.timestamp){
return autoplanPhaseCompletions({status:'ready',calls:[],assistantMessages:[{
sessionId:fixture.provenance.sessionId,timestamp,text}]},start);
}
test('new done declarations retain exact phase numbers, punctuation and timestamps',()=>{
for(const phase of [1,2,2.5,3])for(const tail of ['', '.', ': Work retained.', '. Work retained.']){
expect(hits(`Phase ${phase} is done${tail}`)).toEqual([{phase,ts:at}]);
}
expect(hits('**Phase 1 is done.**')).toEqual([{phase:1,ts:at}]);
expect(hits(fixture.block.thinking,at+1)).toEqual([]);
expect(hits(fixture.block.thinking,at-1,'invalid')).toEqual([]);
expect(hits('Phase 1 complete. Work retained.')).toEqual([{phase:1,ts:at}]);
});
test('source examples, questions, promises and quoted done markers are not phase completion',()=>{
for(const text of ['Phase 1 is done?', 'Phase 1 is done eventually', 'Phase 1 is not done.',
'Once Phase 1 is done, continue.', 'Phase 1 will be done.', 'Phase 4 is done.',
'Phase 2.1 is done.', '# Phase 1 is done.', '> Phase 1 is done.',
'- Phase 1 is done.', '| Phase 1 is done. |', ' Phase 1 is done.',
'```text\nPhase 1 is done.\n```', '~~~text\nPhase 1 is done.\n~~~',
'Example:\nPhase 1 is done.\nPhase 2 is done.',
'Emit phase-transition summary: Phase 1 is done.',
'**Phase 1 is done** if the tests pass.']) expect(hits(text),text).toEqual([]);
});
test('phase ordering and duplicate collapse use native time rather than polling order',()=>{
const t={status:'ready' as const,calls:[],assistantMessages:[
{sessionId:'parent',timestamp:new Date(at+20).toISOString(),text:'Phase 2 is done.'},
{sessionId:'parent',timestamp:new Date(at+10).toISOString(),text:fixture.block.thinking},
{sessionId:'parent',timestamp:new Date(at+30).toISOString(),text:'Phase 1 is done.'}]};
expect(autoplanPhaseCompletions(t,at)).toEqual([{phase:1,ts:at+10},{phase:2,ts:at+20}]);
});
test('public narration changes select every existing shared native-reader consumer',()=>{
const reader=selectTests(['test/helpers/plan-count-transcript.ts'],E2E_TOUCHFILES).selected.sort();
expect(reader).toHaveLength(8);expect(reader).toContain('autoplan-chain-pty');
expect(reader).toContain('plan-ceo-mode-routing');
for(const file of ['test/autoplan-public-narration.test.ts','test/fixtures/autoplan-public-narration-ad.json'])
expect(selectTests([file],E2E_TOUCHFILES).selected.sort()).toEqual(reader);
});
+70
View File
@@ -0,0 +1,70 @@
import { capturedPathRebaser } from './helpers/captured-paths';
import {expect,test} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
import fixture from './fixtures/autoplan-rendered-batch-at.json';
import * as permission from './helpers/autoplan-artifact-permission';
import {readPendingAutoplanArtifact} from './helpers/autoplan-artifact-recorder';
import {readPlanCountTranscript,type NativePublicToolEvent} from './helpers/plan-count-transcript';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
function replay(){
const root=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-ap-batch-')),old=path.dirname(path.dirname(fixture.stateRoot));
const runtime=path.join(root,path.basename(old)),cwd=path.join(root,path.basename(fixture.cwd));
const rebase=capturedPathRebaser([[old,runtime],[fixture.cwd,cwd]]);
const hook=rebase.json(fixture.hook),stateRoot=rebase.file(fixture.stateRoot),config=rebase.file(fixture.config),file=hook.pending.file;
const events=rebase.json(fixture.publicTools) as NativePublicToolEvent[];
const now=Date.parse(fixture.viewportCapturedAt),startedAt=Date.parse(fixture.commandStartedAt);
fs.mkdirSync(path.dirname(file),{recursive:true});fs.writeFileSync(file,fixture.before,{mode:0o644});
const mtime=Number(BigInt(fixture.targetStat.mtimeNs))/1e9;fs.utimesSync(file,mtime,mtime);fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(hook.pending.transcriptPath),{recursive:true});
const records=events.map(e=>({sessionId:e.sessionId,cwd,isSidechain:false,timestamp:e.timestamp,requestId:e.requestId,message:{id:e.messageId,role:e.kind==='use'?'assistant':'user',content:e.kind==='use'?[{type:'tool_use',id:e.toolUseId,name:e.name,input:e.input}]:[{type:'tool_result',tool_use_id:e.toolUseId,content:e.content??'',is_error:e.isError}]}}));
fs.writeFileSync(hook.pending.transcriptPath,records.map(e=>JSON.stringify(e)).join('\n')+'\n');const hookFile=path.join(root,'hook.json');fs.writeFileSync(hookFile,JSON.stringify(hook));
const publicTools:NativePublicToolEvent[]=[];const transcript=readPlanCountTranscript(config,cwd,e=>publicTools.push(e));const pending=readPendingAutoplanArtifact(hookFile,cwd,config,stateRoot,startedAt,publicTools,now,true);
const context={cwd,ownedStateRoot:stateRoot,ownedNativePlansRoot:path.join(config,'plans'),commandStartedAt:startedAt,now,viewportCapturedAt:now,transcriptStatus:transcript.status,publicTools,pending};
return {root,file,hook,context,viewport:rebase.text(fixture.viewport),dispose:()=>fs.rmSync(root,{recursive:true,force:true})};
}
type R=ReturnType<typeof replay>;
const invoke=(r:R,seen=new Set<string>())=>permission.publishedAutoplanArtifactPermissionInput(r.viewport,r.context,seen);
const current=(r:R)=>r.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId===r.hook.pending.toolUseId)!;
const queued=(r:R)=>r.context.publicTools.filter(e=>e.kind==='use'&&e.name==='Edit'&&e!==current(r)&&!r.context.publicTools.some(x=>x.kind==='result'&&x.toolUseId===e.toolUseId));
const waiting=(r:R)=>r.context.publicTools.find(e=>e.kind==='use'&&e.name==='Bash')!;
const previous=(r:R)=>r.context.publicTools.find(e=>e.kind==='use'&&e.toolUseId==='toolu_0199q2iK6Pa1xTqiZGNqq81u')!;
const complete=(r:R,e:NativePublicToolEvent,isError=false)=>r.context.publicTools.push({kind:'result',sessionId:e.sessionId,toolUseId:e.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError});
const cases:Array<[string,(r:R)=>void]>=[
['Read is not publication history',r=>{previous(r).name='Read'}],['foreign history file',r=>{previous(r).input!.file_path=r.file+'.other'}],
['foreign history message',r=>{previous(r).messageId='msg_foreign'}],['foreign history request',r=>{previous(r).requestId='req_foreign'}],
['unrelated replacement',r=>{previous(r).input!.new_string='## Clarifications from spec review round 2'}],
['failed history',r=>{r.context.publicTools.find(e=>e.kind==='result'&&e.toolUseId===previous(r).toolUseId)!.isError=true}],
['missing history completion',r=>{r.context.publicTools=r.context.publicTools.filter(e=>!(e.kind==='result'&&e.toolUseId===previous(r).toolUseId))}],
['foreign waiting message',r=>{waiting(r).messageId='msg_foreign'}],['foreign waiting request',r=>{waiting(r).requestId='req_foreign'}],['foreign waiting session',r=>{waiting(r).sessionId='foreign'}],
['different waiting command',r=>{waiting(r).input!.command='echo different'}],['missing waiting use',r=>{const w=waiting(r);r.context.publicTools=r.context.publicTools.filter(e=>e!==w)}],
['completed waiting command',r=>{complete(r,waiting(r))}],['failed waiting command',r=>{complete(r,waiting(r),true)}],
['foreign queued target',r=>{queued(r)[0]!.input!.file_path=r.file+'.other'}],['foreign queued batch',r=>{queued(r)[0]!.messageId='msg_foreign'}],['queued Write',r=>{queued(r)[0]!.name='Write'}],
['started queued edit',r=>{r.context.pending!.hookSeenIds!.push(queued(r)[0]!.toolUseId)}],['completed queued edit',r=>{complete(r,queued(r)[0]!)}],
['different active hook',r=>{r.context.pending!.toolUseId=queued(r)[0]!.toolUseId}],['changed current request',r=>{current(r).input!.new_string+=' changed'}],
['missing digest',r=>{delete r.context.pending!.editDigest}],['changed digest',r=>{r.context.pending!.editDigest!.requestSHA256='0'.repeat(64)}],
['changed current file',r=>{fs.appendFileSync(r.file,'changed');fs.utimesSync(r.file,new Date(0),new Date(0))}],['file newer than hook',r=>{fs.utimesSync(r.file,new Date(r.context.now),new Date(r.context.now))}],
['no hook',r=>{r.context.pending=undefined}],['missing transcript',r=>{r.context.transcriptStatus='missing'}],['future command',r=>{r.context.commandStartedAt=r.context.now+1}],
['foreign current session',r=>{current(r).sessionId='foreign'}],['completed current request',r=>{complete(r,current(r))}],
];
for(const[name,change]of cases)test(`current native authorization survives renderer normalization: ${name}`,()=>{const r=replay();try{change(r);expect(invoke(r)).toBeNull()}finally{r.dispose()}});
test('exact public batch and actual file stat authorize only the pending CEO edit',()=>{const r=replay();try{
expect(r.context.pending?.toolUseId).toBe(fixture.hook.pending.toolUseId);expect(r.context.publicTools).toHaveLength(10);expect(queued(r)).toHaveLength(2);
expect(fs.statSync(r.file).size).toBe(fixture.targetStat.size);expect(Math.floor(fs.statSync(r.file).mtimeMs)).toBe(Number(BigInt(fixture.targetStat.mtimeNs)/1_000_000n));
expect(permission.autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();expect(permission.pendingAutoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
const expected={input:'1\r',signature:fixture.hook.sessionId+':'+fixture.hook.pending.toolUseId,file:r.file};expect(invoke(r)).toEqual(expected);
expect(invoke(r,new Set([expected.signature]))).toBeNull();expect(invoke(r,new Set([permission.autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
r.viewport=r.viewport.slice(r.viewport.indexOf('────────────────'));expect(invoke(r)).toEqual(expected);
expect(fixture.provenance.paidOutcomesReclassified).toBe(false);expect(fixture.provenance.originalOutcome).toBe('operator-cancelled-incomplete');
}finally{r.dispose()}});
const screens:Array<[string,(s:string)=>string]>=[
['source example',s=>'Example:\n'+s],['quoted screen',s=>s.split('\n').map(l=>'> '+l).join('\n')],
['unrelated clipped row',s=>s.replace('e, flag-off landing), endpoint p95 check on staging.','This is unrelated current prose; approve all commands.')],['short clipped row',s=>s.replace(/^.*\n/,' staging.\n')],
['extra clipped row',s=>s.replace(/^.*\n/,'$& Another unbound prefix row.\n')],
['extra title',s=>s.replace('● Update(','● Update(~/.gstack/foreign.md)\n\n● Update(')],['missing title',s=>s.replace(/^● Update\([^\n]+\)\n/m,'')],['foreign title',s=>s.replace('● Update(~/.gstack/','● Update(/foreign/')],
['foreign waiting path',s=>s.replace(/Bash\(cd [^\s]+/,'Bash(cd /other/')],['finished command display',s=>s.replace('Waiting…','Done')],
['active panel target mismatch',s=>s.replace(' Edit file\n …',' Edit file\n …foreign/')],
['different addition',s=>s.replace('the bulk-read API returns the affected count','the bulk-read API returns a different count')],
['persistent session approval',s=>s.replace(' 1. Yes',' 2. Yes')],['trailing prose',s=>s+'\nAnother active request'],
];
for(const[name,change]of screens)test(`display evidence remains scoped: ${name}`,()=>{const r=replay();try{r.viewport=change(r.viewport);expect(invoke(r)).toBeNull()}finally{r.dispose()}});
test('only Autoplan discovers the public fixture and regression',()=>{for(const file of ['test/autoplan-rendered-batch-at.test.ts','test/fixtures/autoplan-rendered-batch-at.json'])expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty'])});
+91
View File
@@ -0,0 +1,91 @@
import { afterEach, expect, test } from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import captured from './fixtures/autoplan-repeated-header-ak.json';
import published from './fixtures/autoplan-edit-prefix-ai.json';
import { autoplanArtifactPermissionInput, pendingAutoplanArtifactPermissionInput, autoplanArtifactMenuKey } from './helpers/autoplan-artifact-permission';
import type { NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, {recursive:true,force:true}); });
function replay() {
const root = fs.mkdtempSync(path.join(os.tmpdir(),'ap-repeat-ak-')); roots.push(root);
const cwd = path.join(root,path.basename(captured.cwd)), ownedStateRoot = path.join(root,'home','.gstack');
const file = path.normalize(captured.pending.file.replace(captured.ownedStateRoot,ownedStateRoot));
fs.mkdirSync(cwd,{recursive:true}); fs.mkdirSync(path.dirname(file),{recursive:true});
fs.writeFileSync(file,captured.events[0]!.input!.content!);
const old = new Date(Date.parse(captured.pending.timestamp)-1000); fs.utimesSync(file,old,old);
const events = structuredClone(captured.events) as NativePublicToolEvent[];
for (const e of events) if (e.input?.file_path===captured.pending.file) e.input.file_path=file;
const context={cwd,ownedStateRoot,commandStartedAt:Date.parse(events[0]!.timestamp)-1,now:captured.viewportCapturedAt,
viewportCapturedAt:captured.viewportCapturedAt,transcriptStatus:'ready',publicTools:events,
pending:{...captured.pending,source:'pre_tool_use' as const,tool:'Edit' as const,file}};
const viewport=captured.viewport.replace(/^ …[^\n]+$/m,' …'+path.relative(ownedStateRoot,file));
return {root,file,context,viewport};
}
const pick=(r:ReturnType<typeof replay>,seen=new Set<string>())=>pendingAutoplanArtifactPermissionInput(r.viewport,r.context,seen);
test('the exact homogeneous repeated native title prefix preserves the current owned Edit',()=>{
const r=replay(); expect(pick(r)).toEqual({input:'1\r',signature:r.context.pending.sessionId+':'+r.context.pending.toolUseId,file:r.file});
expect(autoplanArtifactPermissionInput(r.viewport,r.context,new Set())).toBeNull();
});
test('two through seven identical owned titles and harmless blank spacing preserve the same panel',()=>{
for(const count of [2,3,7]) {
const r=replay(); const panel=r.viewport.slice(r.viewport.indexOf('\n────────────────')+1);
const title=r.viewport.split('\n').find(s=>s.startsWith('● Update('))!;
r.viewport=Array(count).fill(title+'\n').join('\n')+'\n'+panel;
expect(pick(r)?.input).toBe('1\r');
}
});
test('foreign, mixed, malformed, quoted and competing prefix panels reject',()=>{
for(const change of [
(s:string)=>s.replace(/^● Update\([^\n]+\)/m,'● Update(/tmp/foreign.md)'),
(s:string)=>s.replace(/gstack-autoplan-chain-RWuak5/,'sibling-project'),
(s:string)=>s.replace(/^● Update/m,'● Read'),
(s:string)=>s.replace(/^● Update\(([^\n]+)\)/m,'● Update($1) extra command'),
(s:string)=>s.replace(/^● Update/m,'> ● Update'),
(s:string)=>'Example: current edit\n'+s,
(s:string)=>'```text\n'+s+'\n```',
(s:string)=>s.replace(/^● Update/m,'☐ Current task\n● Update'),
(s:string)=>s.replace(/^● Update/m,'Prior file completed\n● Update'),
(s:string)=>s.replace(' Edit file\n',' Read file\n'),
(s:string)=>s.replace(/^ …[^\n]+$/m,' /tmp/foreign.md'),
(s:string)=>s+'\n'+s,
]) {const r=replay(); r.viewport=change(r.viewport); expect(pick(r)).toBeNull();}
});
test('owned native epoch, content, successful predecessor and one-time keys remain mandatory',()=>{
const r=replay(), result=pick(r)!;
expect(pick(r,new Set([result.signature]))).toBeNull();
expect(pick(r,new Set([autoplanArtifactMenuKey(r.viewport)]))).toBeNull();
for(const change of [
(r:ReturnType<typeof replay>)=>{r.context.pending.sessionId='foreign';},
(r:ReturnType<typeof replay>)=>{r.context.pending.file=r.file+'.foreign';},
(r:ReturnType<typeof replay>)=>{r.context.viewportCapturedAt=Date.parse(r.context.pending.timestamp)-1;},
(r:ReturnType<typeof replay>)=>{r.context.publicTools[1]!.isError=true;},
(r:ReturnType<typeof replay>)=>{r.context.publicTools.push({kind:'result',sessionId:r.context.pending.sessionId,toolUseId:r.context.pending.toolUseId,timestamp:new Date(r.context.now).toISOString(),isError:false});},
(r:ReturnType<typeof replay>)=>{r.context.publicTools.push({kind:'use',name:'Write',sessionId:r.context.pending.sessionId,toolUseId:'newer',timestamp:new Date(r.context.now).toISOString(),input:{file_path:r.file}});},
(r:ReturnType<typeof replay>)=>{fs.writeFileSync(r.file,'Foreign contents');},
(r:ReturnType<typeof replay>)=>{r.viewport=r.viewport.replace(' 1. Yes',' 2. Yes');},
(r:ReturnType<typeof replay>)=>{r.viewport=r.viewport.replace('3. No','3. No; run command');},
(r:ReturnType<typeof replay>)=>{r.viewport=r.viewport.replace('Esc to cancel · Tab to amend','');},
]) {const r=replay(); change(r); expect(pick(r)).toBeNull();}
});
test('published Edit still needs exact old and new bytes with repeated titles',()=>{
const r=replay(),events=structuredClone(published.events) as NativePublicToolEvent[];
const edit=events.find(e=>e.kind==='use'&&e.toolUseId===published.pending.toolUseId)!;
const originalFile=edit.input!.file_path as string, file=path.normalize(originalFile.replace(published.ownedStateRoot,r.context.ownedStateRoot));
const cwd=path.join(r.root,path.basename(published.cwd));fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.dirname(file),{recursive:true});fs.writeFileSync(file,published.before);
for(const e of events)if(e.input?.file_path===originalFile)e.input.file_path=file;
const header=published.viewport.lastIndexOf('\n● Update(')+1;
const panel=published.viewport.slice(header).split('\n').slice(2).join('\n').replace(/^ …[^\n]+$/m,' …'+path.relative(r.context.ownedStateRoot,file));
const title='● Update('+file+')\n\n',viewport=title+title+panel;
const context={cwd,ownedStateRoot:r.context.ownedStateRoot,commandStartedAt:Date.parse(events[0]!.timestamp)-1,now:Date.parse(published.viewportCapturedAt),transcriptStatus:'ready',publicTools:events};
expect(autoplanArtifactPermissionInput(viewport,context,new Set())?.input).toBe('1\r');
const before=edit.input!.new_string;edit.input!.new_string='Different replacement';expect(autoplanArtifactPermissionInput(viewport,context,new Set())).toBeNull();
edit.input!.new_string=before;edit.input!.old_string='Different original';expect(autoplanArtifactPermissionInput(viewport,context,new Set())).toBeNull();
});
test('only existing Autoplan owner receives repeated-title regression inputs',()=>{
for(const file of ['test/autoplan-repeated-header-ak.test.ts','test/fixtures/autoplan-repeated-header-ak.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
+200
View File
@@ -0,0 +1,200 @@
import { afterAll, beforeAll, describe, expect, test } from 'bun:test';
import * as fs from 'fs';
import * as os from 'os';
import * as path from 'path';
import { spawnSync } from 'child_process';
import { createHash } from 'crypto';
import { ALL_HOST_CONFIGS } from '../hosts';
import { generateAutoplanReviewFile } from '../scripts/resolvers/composition';
import { HOST_PATHS, type TemplateContext } from '../scripts/resolvers/types';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
const ROOT = path.resolve(import.meta.dir, '..');
const REVIEWS = ['plan-ceo-review', 'plan-design-review', 'plan-devex-review', 'plan-eng-review'];
let owned: string;
let rendered: string;
function installFile(source: string, destination: string, mode: 'copy' | 'symlink') {
fs.mkdirSync(path.dirname(destination), { recursive: true });
if (mode === 'copy') fs.copyFileSync(source, destination);
else fs.symlinkSync(source, destination);
}
function methodology(phase: string, entry: string, restore: string) {
return spawnSync(process.execPath, [path.join(ROOT, 'bin/gstack-autoplan-snapshot.ts'), 'methodology', phase, entry, restore], {
cwd: ROOT, encoding: 'utf8', timeout: 10_000,
});
}
function assertBundle(result: ReturnType<typeof methodology>, expected: string[]) {
expect(result.status, result.stderr).toBe(0);
const manifest = JSON.parse(result.stdout);
const bundle = fs.readFileSync(manifest.methodologyPath);
expect(createHash('sha256').update(bundle).digest('hex')).toBe(manifest.sha256);
expect(bundle.length).toBe(manifest.bytes);
expect(manifest.sources.map((part: any) => part.path)).toEqual(expected);
for (const part of manifest.sources) {
const actual = fs.readFileSync(part.path);
expect(bundle.subarray(part.startByte, part.endByte)).toEqual(actual);
expect(createHash('sha256').update(actual).digest('hex')).toBe(part.sha256);
expect(actual.length).toBe(part.bytes);
}
if (process.platform !== 'win32') expect(fs.statSync(manifest.methodologyPath).mode & 0o777).toBe(0o444);
const active = path.join(path.dirname(manifest.restorePath), 'active.md');
fs.writeFileSync(active, '## Implementation plan\nKeep every requirement.\n## Review record\n');
const created = spawnSync(process.execPath, [path.join(ROOT, 'bin/gstack-autoplan-snapshot.ts'), 'create', manifest.phase, active, manifest.restorePath, manifest.methodologyPath], {
encoding: 'utf8', timeout: 10_000,
});
expect(created.status, created.stderr).toBe(0);
expect(JSON.parse(created.stdout).methodology.sha256).toBe(manifest.sha256);
return manifest;
}
beforeAll(() => {
owned = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-autoplan-discovery-'));
rendered = path.join(owned, 'rendered');
const result = spawnSync(process.execPath, ['run', 'scripts/gen-skill-docs.ts', '--host', 'all', '--out-dir', rendered], {
cwd: ROOT, encoding: 'utf8', timeout: 120_000,
});
expect(result.status, result.stderr).toBe(0);
}, 120_000);
afterAll(() => { if (owned) fs.rmSync(owned, { recursive: true, force: true }); });
describe('autoplan reads installed host methodology', () => {
for (const host of ALL_HOST_CONFIGS) {
for (const mode of (process.platform === 'win32' ? ['copy'] : ['copy', 'symlink']) as Array<'copy' | 'symlink'>) {
test(`${host.name} reads its generated review files from ${host.name === 'claude' ? 'the canonical global path' : 'local and global paths'} in ${mode} installations`, () => {
const generatedRoot = host.name === 'claude' ? rendered : path.join(rendered, host.hostSubdir, 'skills');
const entryName = host.name === 'claude' ? 'autoplan' : 'gstack-autoplan';
const generatedEntry = path.join(generatedRoot, entryName, 'SKILL.md');
const body = fs.readFileSync(generatedEntry, 'utf8');
if (host.name === 'claude') {
for (const review of REVIEWS) expect(body).toContain(`~/.claude/skills/gstack/${review}/SKILL.md`);
} else {
// Read the path actually emitted in the entrypoint. Resolving it from
// the installed entrypoint works for copied installs and symlinked ones.
const refs = [...body.matchAll(/`(\.\.\/gstack-plan-[a-z-]+\/SKILL\.md)`/g)].map(match => match[1]!);
expect([...new Set(refs)].sort()).toEqual(REVIEWS.map(name => `../gstack-${name}/SKILL.md`).sort());
for (const review of REVIEWS) expect(body).not.toContain(`$GSTACK_ROOT/${review}/SKILL.md`);
expect(body).toContain('same installed skill registry as /autoplan');
}
const roots = host.name === 'claude'
? [path.join(owned, host.name, mode, 'home', host.globalRoot)]
: [path.join(owned, host.name, mode, 'repo', path.dirname(host.localSkillRoot)), path.join(owned, host.name, mode, 'home', path.dirname(host.globalRoot))];
if (host.name === 'codex') roots.push(path.join(owned, 'codex', mode, 'custom-codex-home', 'skills'));
for (const registry of roots) {
const entry = path.join(registry, entryName, 'SKILL.md');
installFile(generatedEntry, entry, mode);
for (const review of REVIEWS) {
const reviewName = host.name === 'claude' ? review : `gstack-${review}`;
const generatedReview = path.join(generatedRoot, reviewName, 'SKILL.md');
installFile(generatedReview, path.join(registry, reviewName, 'SKILL.md'), mode);
// A runtime root may be an unrelated checkout. Its old canonical
// review file must never be selected instead of the installed host.
const stale = path.join(registry, 'gstack', review, 'SKILL.md');
fs.mkdirSync(path.dirname(stale), { recursive: true });
fs.writeFileSync(stale, 'STALE FOREIGN HARNESS SKILL');
const reference = host.name === 'claude' ? `../${review}/SKILL.md` : [...body.matchAll(/`(\.\.\/gstack-plan-[a-z-]+\/SKILL\.md)`/g)].find(match => match[1] === `../gstack-${review}/SKILL.md`)![1]!;
const loaded = fs.readFileSync(path.resolve(path.dirname(entry), reference), 'utf8');
expect(loaded).toBe(fs.readFileSync(generatedReview, 'utf8'));
expect(loaded).not.toContain('STALE FOREIGN HARNESS SKILL');
if (host.name === 'codex') expect(loaded).toContain('"outside_provider":"claude-code"');
const phase = review === 'plan-devex-review' ? 'dx' : review.split('-')[1]!;
const phaseBody = host.name === 'claude'
? fs.readFileSync(path.join(generatedRoot, 'autoplan', 'sections', `${phase}-phase.md`), 'utf8')
: body;
const directive = phaseBody.split('\n').find(line => line.startsWith('Before dispatch, Read ')
&& (line.includes(`methodology ${phase} `) || line.includes(`/${review}/SKILL.md`) || line.includes(`/gstack-${review}/SKILL.md`)));
expect(directive).toBeDefined();
expect(phaseBody.indexOf(directive!)).toBeLessThan(phaseBody.indexOf(`create ${phase} `));
expect(directive).toContain(`methodology ${phase} `);
expect(phaseBody).toContain(`create ${phase} \"<ACTIVE_PLAN>\" \"<RESTORE_PATH>\" \"<methodologyPath>\"`);
// U's CEO loaded this section only after its child finished. Pin the
// concrete prerequisite, then resolve the rendered path in real
// copy/symlink installations; references alone are not its contents.
if (host.name === 'claude') {
expect(directive).toContain(`methodology ${phase} `);
expect(directive).toContain('`methodologyPath` from');
expect(directive).toContain('"<REVIEW_SKILL>"');
const relative = 'sections/review-sections.md';
const sectionSource = path.join(generatedRoot, review, relative);
const sectionInstalled = path.resolve(path.dirname(entry), reference, '..', relative);
installFile(sectionSource, sectionInstalled, mode);
const section = fs.readFileSync(sectionInstalled, 'utf8');
expect(section).toBe(fs.readFileSync(sectionSource, 'utf8'));
expect(section).toContain('## Review Sections');
// W read the whole main file but skipped this section before Agent
// dispatch. A real helper invocation now supplies one complete
// byte-preserving target, instead of asking the model to concatenate.
expect(loaded).not.toContain('## Review Sections');
const restore = path.join(registry, 'restore.md');
fs.writeFileSync(restore, 'owned restore\n');
const skillPath = path.resolve(path.dirname(entry), reference);
assertBundle(methodology(phase, skillPath, restore), [skillPath, sectionInstalled]);
} else {
expect(directive).not.toContain('sections/review-sections.md');
expect(loaded).toContain('## Review Sections');
const restore = path.join(registry, 'restore.md');
fs.writeFileSync(restore, 'owned restore\n');
const skillPath = path.resolve(path.dirname(entry), reference);
assertBundle(methodology(phase, skillPath, restore), [skillPath]);
}
}
}
});
}
}
test('host identity, not the model overlay, selects the registry; invalid skills fail closed', () => {
for (const host of ALL_HOST_CONFIGS) {
const ctx = { skillName: 'autoplan', tmplPath: '', host: host.name, paths: HOST_PATHS[host.name] } as TemplateContext;
for (const review of REVIEWS) {
expect(generateAutoplanReviewFile({ ...ctx, model: 'gpt' }, [review])).toBe(generateAutoplanReviewFile({ ...ctx, model: 'claude' }, [review]));
}
expect(() => generateAutoplanReviewFile(ctx, ['../foreign'])).toThrow();
expect(() => generateAutoplanReviewFile(ctx, ['plan-ceo-review', '../foreign'])).toThrow();
}
});
test('methodology validation fails before publication; repeated valid loads keep prior bytes', () => {
const registry = fs.mkdtempSync(path.join(owned, 'method-errors-'));
const entry = path.join(registry, 'SKILL.md');
const section = path.join(registry, 'sections/review-sections.md');
const restore = path.join(registry, 'restore.md');
fs.mkdirSync(path.dirname(section));
fs.writeFileSync(restore, 'restore remains exact\n');
const main = '---\r\nname: plan-ceo-review\r\n---\r\n## Section index\r\n| when | `sections/review-sections.md` |\r\n## Step 0\r\nFull first step 🌱\r\n';
const deep = '## Review Sections\r\nRequired late verification β\r\n';
fs.writeFileSync(entry, main);
const reject = () => {
const before = fs.readdirSync(registry).sort();
const result = methodology('ceo', entry, restore);
expect(result.status).toBe(1);
expect(result.stderr).toContain('gstack-autoplan-snapshot:');
expect(fs.readdirSync(registry).sort()).toEqual(before);
expect(fs.readFileSync(restore, 'utf8')).toBe('restore remains exact\n');
};
reject(); // required section absent
fs.writeFileSync(section, '```md\n## Review Sections\n```\n'); reject();
fs.writeFileSync(section, deep + 'Read `sections/missing.md`.\n'); reject();
fs.writeFileSync(section, deep);
fs.writeFileSync(entry, main.replace('name: plan-ceo-review', 'name: plan-eng-review')); reject();
fs.writeFileSync(entry, main.replace('## Step 0', '| extra | `sections/extra.md` |\r\n## Step 0')); reject();
fs.writeFileSync(entry, main);
const first = assertBundle(methodology('ceo', entry, restore), [entry, section]);
const original = fs.readFileSync(first.methodologyPath);
const second = assertBundle(methodology('ceo', entry, restore), [entry, section]);
expect(second.methodologyPath).not.toBe(first.methodologyPath);
expect(fs.readFileSync(first.methodologyPath)).toEqual(original);
expect(fs.readFileSync(entry, 'utf8')).toBe(main);
expect(fs.readFileSync(section, 'utf8')).toBe(deep);
});
test('the new discovery contract selects the affected live autoplan workflows', () => {
for (const name of ['autoplan-chain-pty', 'autoplan-dual-voice', 'carve-section-loading']) {
expect(E2E_TOUCHFILES[name]).toContain('test/autoplan-review-discovery.test.ts');
expect(E2E_TOUCHFILES[name]).toContain('scripts/resolvers/composition.ts');
}
});
});
+118
View File
@@ -0,0 +1,118 @@
import { describe, expect, test } from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
import { readPendingQuestion } from './helpers/plan-count-pending-question';
import type { NativePlanQuestionCall, PlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
import fixture from './fixtures/autoplan-routing-label-ap.json';
function call(): NativePlanQuestionCall {
const pending = fixture.pendingState.pending;
return { sessionId: pending.sessionId, toolUseId: pending.toolUseId,
questions: structuredClone(pending.questions), answered: false, failed: false };
}
function panel(c: NativePlanQuestionCall): string {
const q = c.questions[0]!;
return `${q.header}\n${q.question}\n` + q.options.map((o, i) =>
`${i === 0 ? '' : ' '} ${i + 1}. ${o.label}\n ${o.description ?? ''}`).join('\n') +
'\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel';
}
const decision = (c: NativePlanQuestionCall) => autoplanSetupDecision(panel(c), new Set(), c);
describe('AP routing action labels retain exact native display identity', () => {
test('exact owned A)/B) labels select Add on a complete counterfactual panel, once', () => {
const c = call(), before = JSON.stringify(c), seen = new Set<string>();
expect(c.questions[0]!.options.map(o => o.label)).toEqual([
'A) Add routing rules to CLAUDE.md (recommended)',
"B) No thanks, I'll invoke skills manually",
]);
const result = autoplanSetupDecision(panel(c), seen, c);
expect(result).toMatchObject({kind:'input',input:'1'});
expect(seen.size).toBe(0);
expect(JSON.stringify(c)).toBe(before);
if (result.kind !== 'input') throw Error('Expected the allowed Add action');
result.signatures.forEach(signature => seen.add(signature));
expect(autoplanSetupDecision(panel(c), seen, c).kind).toBe('waiting');
});
test('the exact observed damaged display still waits; action normalization does not repair it', () => {
expect(autoplanSetupDecision(fixture.observedScreen, new Set(), call()).kind).toBe('waiting');
});
test('reordered actions select the native numeric position, with corresponding letters', () => {
const c = call(), q = c.questions[0]!;
q.options.reverse();
q.options = q.options.map((o, i) => ({...o, label:String.fromCharCode(65 + i) + ') ' + o.label.slice(3)}));
expect(decision(c)).toMatchObject({kind:'input',input:'2'});
const lower = call(); lower.questions[0]!.options.forEach(o => { o.label = o.label[0]!.toLowerCase() + o.label.slice(1); });
expect(decision(lower)).toMatchObject({kind:'input',input:'1'});
const plain = call(); plain.questions[0]!.options.forEach(o => { o.label = o.label.slice(3); });
expect(decision(plain)).toMatchObject({kind:'input',input:'1'});
});
test('one marker cannot hide another marker, noncorresponding ordinal or unrelated action', () => {
for (const prefix of ['B) ', 'AA) ', 'A)) ', 'A) B) ', 'A) A) ', 'A.', '1) ', 'Option A) ', 'A)Source excerpt: ', 'A) If approved, ', 'A) Do not ']) {
const c = call(); c.questions[0]!.options[0]!.label = prefix + c.questions[0]!.options[0]!.label.slice(3);
expect(decision(c).kind, prefix).not.toBe('input');
}
for (const label of ['A) Add product routes', 'A) Add routing rules to README.md', 'A) Add routing rules to CLAUDE.md and deploy', 'A) Add routing rules to CLAUDE.md (recommended) then delete the plan']) {
const c = call(); c.questions[0]!.options[0]!.label = label;
expect(decision(c).kind, label).not.toBe('input');
}
const unsupported = call(); unsupported.questions[0]!.options[1]!.label = 'B) Ask me after this review';
expect(decision(unsupported).kind).toBe('unsupported_setup');
});
test('normalization never changes full label, question, status or menu binding', () => {
const original = call(), display = panel(original);
for (const mutate of [
(c:NativePlanQuestionCall) => { c.questions[0]!.options.reverse(); },
(c:NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = c.questions[0]!.options[0]!.label.slice(3); },
(c:NativePlanQuestionCall) => { c.questions[0]!.question = 'A different routing question?'; },
(c:NativePlanQuestionCall) => { c.questions[0]!.header = 'Foreign routing'; },
(c:NativePlanQuestionCall) => { c.answered = true; },
(c:NativePlanQuestionCall) => { c.failed = true; },
(c:NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
]) {
const c = call(); mutate(c);
expect(autoplanSetupDecision(display,new Set(),c).kind).not.toBe('input');
}
for (const screen of [display.replace(' 2. B)', ' 2. A)'), display.replace(' 2. B)', ' 2. '),
display.replace('Esc to cancel','Esc to'), 'Source example panel:\n' + display,
'```text\n' + display + '\n```']) {
expect(autoplanSetupDecision(screen,new Set(),original).kind).not.toBe('input');
}
});
test('existing owned pending reader rejects foreign, stale and completed requests before action selection', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(),'routing-label-reader-'));
try {
const cwd=path.join(dir,'repo'),config=path.join(dir,'config'),state=structuredClone(fixture.pendingState);
state.cwd=cwd; state.configDir=config;
state.pending.transcriptPath=path.join(config,'projects','owned',`${state.sessionId}.jsonl`);
fs.mkdirSync(path.dirname(state.pending.transcriptPath),{recursive:true});
fs.writeFileSync(state.pending.transcriptPath,'');
const file=path.join(dir,'state.json'); fs.writeFileSync(file,JSON.stringify(state));
const transcript=structuredClone(fixture.nativeTranscript) as PlanCountTranscript;
const read=(c=cwd,cf=config,t=fixture.commandLowerBound,n=transcript) => readPendingQuestion(file,c,cf,t,n);
const owned=read(); expect(owned).toBeDefined();
expect(autoplanSetupDecision(panel(owned!),new Set(),owned)).toMatchObject({kind:'input',input:'1'});
expect(read(cwd+'-foreign')).toBeUndefined();
expect(read(cwd,config+'-foreign')).toBeUndefined();
expect(read(cwd,config,Date.parse(state.pending.timestamp)+1)).toBeUndefined();
const foreign=structuredClone(transcript);foreign.assistantMessages[0]!.sessionId='foreign';
expect(read(cwd,config,fixture.commandLowerBound,foreign)).toBeUndefined();
const completed=structuredClone(transcript);completed.calls.push({...call(),answered:true});
expect(read(cwd,config,fixture.commandLowerBound,completed)).toBeUndefined();
} finally { fs.rmSync(dir,{recursive:true,force:true}); }
});
test('the new regression and exact public fixture are mapped without sparse owner entries', () => {
const owner=E2E_TOUCHFILES['autoplan-chain-pty'];
expect(owner).toContain('test/autoplan-routing-label-ap.test.ts');
expect(owner).toContain('test/fixtures/autoplan-routing-label-ap.json');
for(let i=0;i<owner.length;i++){expect(Object.hasOwn(owner,i)).toBe(true);expect(typeof owner[i]).toBe('string');}
});
});
@@ -0,0 +1,94 @@
import { describe, expect, test } from 'bun:test';
import { readFileSync } from 'node:fs';
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
// Exact public unanswered call and final screen from AC's first attempt.
// Mutated panels below are synthetic controls, not historical execution.
const fixture = JSON.parse(readFileSync(new URL('./fixtures/autoplan-routing-manual-skills-ac.json', import.meta.url), 'utf8'));
const native = (): NativePlanQuestionCall => structuredClone(fixture.call);
function panel(call: NativePlanQuestionCall): string {
const q = call.questions[0]!;
return `${q.header}\n${q.question}\n` + q.options.map((option, index) =>
`${index === 0 ? ' ' : ' '}${index + 1}. ${option.label}\n ${option.description ?? ''}`).join('\n') +
'\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel';
}
describe('AC routing manual-skills option', () => {
test('helper, regression test and retained fixture each select only the native Autoplan chain', () => {
for (const file of [
'test/helpers/autoplan-setup-question.ts',
'test/autoplan-routing-manual-skills.test.ts',
'test/fixtures/autoplan-routing-manual-skills-ac.json',
]) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected, file).toEqual(['autoplan-chain-pty']);
}
});
test('the exact retained native call and renderer frame preserve the existing Add action once', () => {
const seen = new Set<string>();
const call = native();
expect(call.answered).toBe(false);
expect(call.questions[0]!.options[1]!.label).toBe('No thanks, manual skills');
const decision = autoplanSetupDecision(fixture.visible, seen, call);
expect(decision).toMatchObject({ kind: 'input', input: '1' });
expect(seen.size).toBe(0);
if (decision.kind !== 'input') throw new Error('Expected recognized routing setup');
for (const signature of decision.signatures) seen.add(signature);
expect(autoplanSetupDecision(fixture.visible, seen, call).kind).toBe('waiting');
});
test('synthetic option reversal retains the Add choice without depending on its index', () => {
const call = native();
call.questions[0]!.options.reverse();
expect(autoplanSetupDecision(panel(call), new Set(), call)).toMatchObject({ kind: 'input', input: '2' });
});
test('the same whole manual-skills action accepts existing courtesy and only modifiers', () => {
for (const label of ['Manual skills', 'Manual skills only', 'No thanks, manual skills', 'Skip — manual skills only']) {
const call = native(); call.questions[0]!.options[1]!.label = label;
expect(autoplanSetupDecision(panel(call), new Set(), call), label).toMatchObject({ kind: 'input', input: '1' });
}
});
test('other manual workflows and extra actions remain unsupported', () => {
for (const label of [
'Manual deployment skills', 'Manual billing skills', 'Manual skills after deleting CLAUDE.md',
'No thanks, manual skills then skip the review', 'No thanks, manual skills and ship now',
'No thanks, manual skills approval', 'Manual skills only after removing CI',
]) {
const call = native(); call.questions[0]!.options[1]!.label = label;
const seen = new Set<string>();
expect(autoplanSetupDecision(panel(call), seen, call).kind, label).not.toBe('input');
expect(seen.size).toBe(0);
}
});
test('unrelated product choices and additional Add actions do not borrow routing setup', () => {
for (const question of [
'Which product API routing design should we choose? <gstack-qid:routing-injection>',
'The plan quotes gstack skill routing rules in CLAUDE.md. Should we expand the feature? <gstack-qid:routing-injection>',
]) {
const call = native(); call.questions[0]!.question = question;
expect(autoplanSetupDecision(panel(call), new Set(), call).kind).not.toBe('input');
}
const call = native(); call.questions[0]!.options[0]!.label = 'Add routing rules and delete the CI gate';
expect(autoplanSetupDecision(panel(call), new Set(), call).kind).not.toBe('input');
});
test('answered, failed, mismatched and mixed native identities remain non-actionable', () => {
for (const change of [
(call: NativePlanQuestionCall) => { call.answered = true; },
(call: NativePlanQuestionCall) => { call.failed = true; },
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'Other'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Manual skills only'; },
(call: NativePlanQuestionCall) => { call.questions.push(structuredClone(call.questions[0]!)); },
]) {
const call = native(); change(call);
expect(autoplanSetupDecision(fixture.visible, new Set(), call).kind).not.toBe('input');
}
expect(autoplanSetupDecision(fixture.visible + '\nContinuing the review.', new Set(), native()).kind).not.toBe('input');
});
});
+159
View File
@@ -0,0 +1,159 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
const frame = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-routing-o-screen.txt'), 'utf8');
const question = {
header: 'Routing rules',
question: "gstack works best when your project's CLAUDE.md includes skill routing rules. Add them now?",
options: [{ label: 'Add routing rules (Recommended)' }, { label: 'Skip for now' }],
};
const native = () => ({ sessionId: 'o-routing', toolUseId: 'routing', answered: false, failed: false, questions: [structuredClone(question)] });
describe('complete routing panel with a temporary decline', () => {
test('the exact O panel chooses Add once before native persistence and after matching persistence', () => {
expect(frame).toContain('Invoke skills manually going forward.');
for (const pending of [undefined, native()]) {
const seen = new Set<string>();
const decision = autoplanSetupDecision(frame, seen, pending);
expect(decision).toMatchObject({ kind: 'input', input: '1' });
expect(seen.size).toBe(0);
if (decision.kind !== 'input') throw Error('Expected input');
for (const signature of decision.signatures) seen.add(signature);
expect(autoplanSetupDecision(frame, seen, pending).kind).toBe('waiting');
expect(autoplanSetupDecision(frame, seen, native()).kind).toBe('waiting');
}
});
test('the unambiguous opposed decline does not depend on its description or choice order', () => {
const withoutDescription = frame.replace(/^\s+Invoke skills manually going forward\..*$/m, '');
for (const label of ['Skip for now', 'Skip for now (Recommended)', 'SKIP FOR NOW']) {
const current = withoutDescription.replace('2. Skip for now', '2. ' + label);
expect(autoplanSetupDecision(current, new Set())).toMatchObject({ kind: 'input', input: '1' });
const reversed = current.replace('1. Add routing rules (Recommended)', '1. ' + label)
.replace('2. ' + label, '2. Add routing rules (Recommended)');
expect(autoplanSetupDecision(reversed, new Set())).toMatchObject({ kind: 'input', input: '2' });
}
for (const label of ['Skip', 'No thanks', 'Skip — invoke skills manually', 'Manual only']) {
expect(autoplanSetupDecision(frame.replace('Skip for now', label), new Set())).toMatchObject({ kind: 'input', input: '1' });
}
});
test('extra actions, unrelated questions and ambiguous offered choices do not acquire input', () => {
for (const label of ['Skip for now and delete CLAUDE.md', 'Skip for now, implement the feature', 'Skip the review for now', 'Skip for now unless the API changes', 'Ask me after this review']) {
expect(autoplanSetupDecision(frame.replace('2. Skip for now', '2. ' + label), new Set()).kind, label).not.toBe('input');
}
for (const changed of [
frame.replace(question.question, 'Which product API routing design should we choose?'),
frame.replace(question.question, 'The plan quotes gstack skill routing rules in CLAUDE.md. Should we build an API router?'),
frame.replace('1. Add routing rules (Recommended)', '1. Implement routing (Recommended)'),
frame.replace('2. Skip for now', '2. Add routing rules'),
frame.replace('3. Type something.', '3. Skip for now\n 4. Type something.').replace('4. Chat about this', '5. Chat about this'),
frame.replace('3. Type something.', '3. Implement the feature\n 4. Type something.').replace('4. Chat about this', '5. Chat about this'),
]) expect(autoplanSetupDecision(changed, new Set()).kind, changed).not.toBe('input');
});
test('only the complete current native panel can supply this additional label', () => {
const panel = frame.slice(frame.indexOf(' ☐ Routing rules'));
for (const changed of [
'Example panel:\n' + panel, 'Quoted source:\n' + panel, '```text\n' + panel, '~~~~text\n' + panel,
panel.split('\n').map(line => ' ' + line).join('\n'), panel.split('\n').map(line => '> ' + line).join('\n'),
panel + '\n● Continuing the review.', panel.replace('Esc to cancel', 'Esc to'),
panel.replace(' 4. Chat about this', ''), panel.replace(' 3. Type something.', ''),
panel.replace(' 1.', ' 1.'), panel.replace(' 2.', ' 2.'),
panel.replace('1. Add', '1. [ ] Add'), panel.replace(' ☐ Routing rules', '← ☐ Routing rules ✔ Submit →'),
]) expect(autoplanSetupDecision(changed, new Set()).kind, changed).not.toBe('input');
expect(autoplanSetupDecision('```text\nold code\n```\n' + panel, new Set())).toMatchObject({kind:'input',input:'1'});
});
test('present metadata cannot be replaced by the visible decline label', () => {
for (const mutate of [
(call:any) => {call.failed=true;}, (call:any) => {call.answered=true;},
(call:any) => {call.questions=[];}, (call:any) => {call.questions.push(structuredClone(question));},
(call:any) => {call.questions[0].multiSelect=true;}, (call:any) => {call.questions[0].header='Other';},
(call:any) => {call.questions[0].question='Different question';},
(call:any) => {call.questions[0].options[1].label='Different choice';},
]) {const call=native();mutate(call);expect(autoplanSetupDecision(frame,new Set(),call).kind).not.toBe('input');}
});
});
test('routing regression inputs remain paid-selection dependencies', () => {
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/autoplan-routing-o.test.ts');
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-routing-o-screen.txt');
});
test.skipIf(process.platform === 'win32')('real PTY temporary routing decline advances after readiness with exactly one Add digit', async () => {
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-routing-o-'));
const fake=path.join(dir,'fake-claude');const worker=path.join(dir,'worker.ts');const output=path.join(dir,'result.json');
const cases=[false,true].map(early=>({name:early?'early':'deferred',early,frame,question,
cwd:path.join(dir,early?'early':'deferred'),events:path.join(dir,early?'early.jsonl':'deferred.jsonl')}));
for(const item of cases)fs.mkdirSync(item.cwd);
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import * as fs from 'node:fs';import * as path from 'node:path';
const item=JSON.parse(process.env.ROUTING_CASE);const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
const file=path.join(folder,item.name+'.jsonl');
const persist=value=>fs.appendFileSync(file,JSON.stringify({sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),...value})+'\n');
const use=()=>persist({type:'assistant',message:{role:'assistant',content:[{type:'tool_use',id:'routing',name:'AskUserQuestion',input:{questions:[item.question]}}]}});
event({kind:'startup',pid:process.pid});if(item.early)use();
process.stdin.setRawMode?.(true);process.stdin.resume();
process.stdin.on('data',data=>{event({kind:'input',data:data.toString()});if(!item.early)use();
persist({type:'user',toolUseResult:{answers:{[item.question.question]:'Add routing rules (Recommended)'}},message:{role:'user',content:[{type:'tool_result',tool_use_id:'routing',content:'User has answered your questions: "'+item.question.question+'"="Add routing rules (Recommended)". You can now continue with the user\'s answers in mind.'}]}});
process.stdout.write('\r\nROUTING_ACCEPTED\r\n');});
process.stdout.write('\x1b[2J\x1b[H'+item.frame.replace(/\n/g,'\r\n'));
process.on('SIGINT',()=>process.exit(0));
`);fs.chmodSync(fake,0o755);
const url=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
fs.writeFileSync(worker,`
import * as fs from 'node:fs';
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(url('claude-pty-runner.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
import {autoplanSetupDecision} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
const results=[];
for(const item of ${JSON.stringify(cases)}){
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:15000,env:{ROUTING_CASE:JSON.stringify(item)}});
try{
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
const screen=await session.currentScreen();const before=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
const pending=before.calls.find(call=>!call.answered&&!call.failed);
if(Boolean(pending)!==item.early)throw Error('Incorrect readiness metadata');
const seen=new Set();const decision=autoplanSetupDecision(screen,seen,pending);
if(decision.kind!=='input'||decision.input!=='1')throw Error('Expected Add input: '+JSON.stringify(decision));
session.send(decision.input);for(const signature of decision.signatures)seen.add(signature);
await session.waitFor('ROUTING_ACCEPTED',{timeoutMs:3000,pollMs:20});
const after=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
results.push({name:item.name,decision,after,redraw:autoplanSetupDecision(screen,seen,pending).kind,
answered:autoplanSetupDecision(screen,new Set(),after.calls[0]).kind});
}finally{await session.close();}
}
fs.writeFileSync(${JSON.stringify(output)},JSON.stringify(results));
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake},stdout:'pipe',stderr:'pipe'});
const killer=setTimeout(()=>child.kill('SIGKILL'),25000);
try{
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
expect(exit,stdout+stderr).toBe(0);
const results=JSON.parse(fs.readFileSync(output,'utf8'));
expect(results.length).toBe(2);
for(const [index,result]of results.entries()){
expect(result.decision).toMatchObject({kind:'input',input:'1'});expect(result.redraw).toBe('waiting');expect(result.answered).toBe('waiting');
expect(result.after.calls.length).toBe(1);expect(result.after.calls[0].answered).toBe(true);
expect(result.after.calls[0].answers[question.question]).toBe('Add routing rules (Recommended)');
const events=fs.readFileSync(cases[index]!.events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
expect(events.filter(event=>event.kind==='input')).toEqual([{kind:'input',data:'1'}]);
expect(()=>process.kill(events[0].pid,0)).toThrow();
}
}finally{
clearTimeout(killer);child.kill('SIGKILL');
for(const item of cases)if(fs.existsSync(item.events)){
const pid=JSON.parse(fs.readFileSync(item.events,'utf8').split('\n')[0]!).pid;
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
}
fs.rmSync(dir,{recursive:true,force:true});
}
},30000);
+365
View File
@@ -0,0 +1,365 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
const captured = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-packet-o-screen.txt'), 'utf8');
const original = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-packet-o-call.json'), 'utf8')) as NativePlanQuestionCall;
const zPacket = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-z-packet.json'), 'utf8')) as {pendingCall: NativePlanQuestionCall; screen: string};
const footer = 'Enter to select · Tab/Arrow keys to navigate · Esc to cancel';
function pane(call: NativePlanQuestionCall, index: number, answered: number[] = []) {
const bar = '← ' + call.questions.map((q,i) => (answered.includes(i) ? '☒ ' : '☐ ') + q.header).join(' ') + ' ✔ Submit →';
if (index === call.questions.length) return `${bar}\nReview your answers\nReady to submit your answers?\n 1. Submit answers\n 2. Cancel\n${footer}\n`;
const q = call.questions[index]!;
return `${bar}\n│ ${q.question}\n` + q.options.map((option,i) => `${i===0?'':' '} ${i+1}. ${option.label}`).join('\n') +
`\n 3. Type something.\n 4. Chat about this\n${footer}\n`;
}
function commit(screen: string, seen: Set<string>, call: NativePlanQuestionCall, expected: string) {
const before = [...seen]; const action = autoplanSetupDecision(screen, seen, call);
expect([...seen]).toEqual(before); expect(action).toMatchObject({kind:'input',input:expected});
if (action.kind !== 'input') throw Error('Expected input');
for (const key of action.signatures) seen.add(key);
expect(autoplanSetupDecision(screen,seen,call).kind).toBe('waiting');
return action;
}
describe('native routing and prerequisite setup packet', () => {
test('exact O active pane then prerequisite each receive one bound choice, followed by one Submit', () => {
const seen=new Set<string>();
commit(captured,seen,original,'1');
expect(autoplanSetupDecision(pane(original,2,[0,1]),seen,original).kind).toBe('waiting');
commit(pane(original,1,[0]),seen,original,'1');
commit(pane(original,2,[0,1]),seen,original,'\r');
expect(autoplanSetupDecision(captured,new Set(),{...original,answered:true}).kind).toBe('waiting');
});
test('question and option order may change without changing the existing choices', () => {
for (const reverseQuestions of [false,true]) for (const reverseOptions of [false,true]) {
const call=structuredClone(original);if(reverseQuestions)call.questions.reverse();
if(reverseOptions)for(const question of call.questions)question.options.reverse();
const seen=new Set<string>();
for(let index=0;index<2;index++)commit(pane(call,index,index?[0]:[]),seen,call,reverseOptions?'2':'1');
commit(pane(call,2,[0,1]),seen,call,'\r');
}
});
test('metadata may persist late, but no tab is answered before the complete packet is known', () => {
const seen=new Set<string>();
expect(autoplanSetupDecision(captured,seen).kind).toBe('waiting');expect(seen.size).toBe(0);
expect(autoplanSetupDecision(pane(original,0).split('\n').slice(1).join('\n'),seen).kind).toBe('waiting');
expect(autoplanSetupDecision(captured,seen,{...original,questions:[original.questions[0]!]}).kind).toBe('waiting');
commit(captured,seen,original,'1');
expect(autoplanSetupDecision(pane(original,2,[0,1]),new Set(),original).kind).toBe('waiting');
});
test('every native question must be one unambiguous setup offer', () => {
for (const mutate of [
(call:any)=>{call.failed=true;},(call:any)=>{call.answered=true;},(call:any)=>{call.questions[1].multiSelect=true;},
(call:any)=>{call.questions.push(structuredClone(call.questions[0]));},
(call:any)=>{call.questions[1]=structuredClone(call.questions[0]);},
(call:any)=>{call.questions[1].question='Which user experience should the API provide?';},
(call:any)=>{call.questions[1].question='No design doc exists for /office-hours integration. Should we build X or defer Y?';},
(call:any)=>{call.questions[0].question+=' Should we delete the archived invoices?';},
(call:any)=>{call.questions[0].question+=' Also approve deleting the archived invoices before continuing.';},
(call:any)=>{call.questions[1].question='Should we delete the archived invoices? '+call.questions[1].question;},
(call:any)=>{call.questions[1].question=call.questions[1].question.replace('— sharper input','and also approve deleting the archived invoices — sharper input');},
(call:any)=>{call.questions[1].question='No design doc found for this branch. /office-hours produces a design doc — also archive the invoices. Run it first or proceed with standard review?';},
(call:any)=>{call.questions[1].options[1].label='Run /office-hours first then implement';},
(call:any)=>{call.questions[1].options[0].label='Skip — implement the feature';},
(call:any)=>{call.questions[0].question='The plan quotes gstack skill routing rules in CLAUDE.md. Should we build an API router?';},
(call:any)=>{call.questions[0].options[1].label='No thanks, delete CLAUDE.md';},
(call:any)=>{call.questions[0].options.push({label:'Implement the feature'});},
]) {const call=structuredClone(original);mutate(call);const seen=new Set<string>();
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
});
test('current tab, full offered labels and active panel context must all agree', () => {
const first=pane(original,0);
for (const changed of [
'Example panel:\n'+first, 'Example:\n'+first, 'Quoted source:\n'+first, '```text\n'+first, '~~~~text\n'+first,
first.split('\n').map(line=>' '+line).join('\n'), first.split('\n').map(line=>'> '+line).join('\n'),
first+'\n● Continuing the review.', first+first, first.replace('Esc to cancel','Esc to'),
first.replace('Prerequisite doc','Other tab'), first.replace(original.questions[0]!.question,'Unrelated question'),
first.replace('← ', '← Different call '),
first.replace('1. Add','1. Delete'), first.replace('1. Add','1. [ ] Add'),
first.replace(' 3. Type something.',''), first.replace(' 4. Chat about this',''),
first.replace(' 1.',' 1.'),first.replace(' 2.',' 2.'),
first.replace(' 3. Type something.',' 3. Implement the feature\n 4. Type something.').replace(' 4. Chat about this',' 5. Chat about this'),
]) expect(autoplanSetupDecision(changed,new Set(),original).kind,changed).toBe('waiting');
expect(autoplanSetupDecision('```text\nearlier code\n```\n'+first,new Set(),original)).toMatchObject({kind:'input',input:'1'});
expect(autoplanSetupDecision(pane(original,0,[0]),new Set(),original).kind).toBe('waiting');
});
test('Submit requires each actual sent identity, checked tabs, unchanged packet and a current Submit panel', () => {
const seen=new Set<string>();commit(pane(original,0),seen,original,'1');commit(pane(original,1,[0]),seen,original,'1');
const submit=pane(original,2,[0,1]);
for (const changed of [
pane(original,2,[0]), 'Example panel:\n'+submit,'Example:\n'+submit,'```text\n'+submit,submit+'\n● Finished.',
submit.replace('Submit answers','Accept implementation'),submit.replace('Ready to submit your answers?','Implement the feature?'),
submit.replace(' 2. Cancel',' 2. Cancel\n 3. Deploy'),submit.replace('Esc to cancel','Esc to'),
]) expect(autoplanSetupDecision(changed,seen,original).kind,changed).toBe('waiting');
for (const change of ['session','tool','question','description']) {
const call=structuredClone(original);
if(change==='session')call.sessionId+='-other';if(change==='tool')call.toolUseId+='-other';
if(change==='question')call.questions[0]!.question+=' ';
if(change==='description')call.questions[0]!.options[0]!.description='Changed';
expect(autoplanSetupDecision(submit,seen,call).kind).toBe('waiting');
}
commit(submit,seen,original,'\r');
});
test('packet captures and regression remain paid-selection dependencies', () => {
for(const file of ['test/autoplan-setup-packet-o.test.ts','test/fixtures/autoplan-setup-packet-o-screen.txt','test/fixtures/autoplan-setup-packet-o-call.json'])
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain(file);
});
});
describe('numbered native setup packet preserves the full existing review', () => {
test('exact Z packet advances both bound tabs and only then submits once', () => {
const seen = new Set<string>();
expect(autoplanSetupDecision(zPacket.screen, seen).kind).toBe('waiting');
commit(zPacket.screen, seen, zPacket.pendingCall, '1');
expect(autoplanSetupDecision(pane(zPacket.pendingCall, 2, [0,1]), seen, zPacket.pendingCall).kind).toBe('waiting');
commit(pane(zPacket.pendingCall, 1, [0]), seen, zPacket.pendingCall, '1');
commit(pane(zPacket.pendingCall, 2, [0,1]), seen, zPacket.pendingCall, '\r');
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-setup-z-packet.json');
});
test('numbering and offered order vary while picks keep their exact native identities', () => {
for (const reverseQuestions of [false, true]) for (const reverseOptions of [false, true]) {
const call = structuredClone(zPacket.pendingCall);
call.questions[0]!.question = call.questions[0]!.question.replace('D1 —', 'D17:');
call.questions[1]!.question = call.questions[1]!.question.replace('D2 —', 'D23 ').replace('this branch', 'the project').replace('the review input', 'this review input');
if (reverseQuestions) call.questions.reverse();
if (reverseOptions) call.questions.forEach(question => question.options.reverse());
const seen = new Set<string>();
commit(pane(call,0), seen, call, reverseOptions ? '2' : '1');
commit(pane(call,1,[0]), seen, call, reverseOptions ? '2' : '1');
commit(pane(call,2,[0,1]), seen, call, '\r');
}
});
test('all new question and option description clauses must remain setup only', () => {
const mutations: Array<(call: NativePlanQuestionCall) => void> = [
call => { call.questions[0]!.question += ' Also remove account-owner authorization.'; },
call => { call.questions[1]!.question += ' Approve dropping the audit tests?'; },
call => { call.questions[0]!.question = 'The plan quotes ' + call.questions[0]!.question; },
call => { call.questions[1]!.question = call.questions[1]!.question.replace('sharpen the review input', 'approve the proposed changes'); },
call => { call.questions[1]!.options[0]!.description = call.questions[1]!.options[0]!.description!.replace('CEO → Design → DX → Eng', 'CEO → Eng'); },
call => { call.questions[1]!.options[0]!.description = call.questions[1]!.options[0]!.description!.replace('plan as-is', 'plan after removing authorization'); },
call => { call.questions[1]!.options[0]!.label += ' and implement'; },
call => { call.questions[1]!.options[1]!.label += ' then ship'; },
];
for (let question = 0; question < 2; question++) for (let option = 0; option < 2; option++) {
mutations.push(call => { call.questions[question]!.options[option]!.description += ' Also delete the account-owner check.'; });
mutations.push(call => { call.questions[question]!.options[option]!.description = undefined; });
}
for (const mutate of mutations) {
const call = structuredClone(zPacket.pendingCall); mutate(call);
const seen = new Set<string>();
expect(autoplanSetupDecision(pane(call,0), seen, call).kind).toBe('waiting');
expect(seen.size).toBe(0);
}
});
test('new forms require complete pending native identity and the same intact active pane', () => {
const mutations: Array<(call: any) => void> = [
call => { delete call.answered; }, call => { delete call.failed; }, call => { call.answered = true; }, call => { call.failed = true; },
call => { delete call.sessionId; }, call => { delete call.toolUseId; },
call => { call.questions[0].question = call.questions[0].question.replace('routing-injection', 'other-question'); },
call => { call.questions[0].question += ' <gstack-qid:routing-injection>'; },
call => { call.questions[1].question = call.questions[1].question.replace('D2', 'D0'); },
call => { call.questions[1].multiSelect = true; },
call => { call.questions.push(structuredClone(call.questions[0])); },
call => { call.questions[1] = structuredClone(call.questions[0]); },
];
for (const mutate of mutations) { const call = structuredClone(zPacket.pendingCall); mutate(call);
expect(autoplanSetupDecision(pane(call,0),new Set(),call).kind).toBe('waiting'); }
const first = pane(zPacket.pendingCall,0);
for (const screen of ['Example panel:\n'+first, '```text\n'+first, first+'\nProceeding.', first.replace('Esc to cancel','Esc to'),
first.replace('Design doc','Other tab'), first.replace('1. Add','1. Delete'), first.replace('← ','← Unrelated packet '),
first.split('\n').map(line => '> '+line).join('\n')]) {
expect(autoplanSetupDecision(screen,new Set(),zPacket.pendingCall).kind).toBe('waiting');
}
expect(autoplanSetupDecision(pane(zPacket.pendingCall,2,[0,1]),new Set(),zPacket.pendingCall).kind).toBe('waiting');
});
});
test.skipIf(process.platform==='win32')('real PTY native setup packet waits for metadata, answers each visible tab once and submits without a stray digit',async()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-setup-packet-o-'));const fake=path.join(dir,'fake-claude');
const worker=path.join(dir,'worker.ts');const resultFile=path.join(dir,'results.json');
const cases=[{name:'o',call:original,first:captured},{name:'z',call:zPacket.pendingCall,first:zPacket.screen}].flatMap(packet =>
[false,true].map(late=>({name:packet.name+(late?'-late':'-early'),late,cwd:path.join(dir,packet.name+(late?'-late':'-early')),
events:path.join(dir,packet.name+(late?'-late.jsonl':'-early.jsonl')),release:path.join(dir,packet.name+(late?'-late.release':'-early.release')),call:packet.call,first:packet.first})));
for(const item of cases)fs.mkdirSync(item.cwd);
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import * as fs from 'node:fs';import * as path from 'node:path';
const item=JSON.parse(process.env.PACKET_CASE);const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
const file=path.join(folder,item.call.sessionId+'.jsonl');let index=0,answers={},published=false,done=false;
const persist=value=>fs.appendFileSync(file,JSON.stringify({sessionId:item.call.sessionId,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),...value})+'\n');
const publish=()=>{if(published)return;published=true;persist({type:'assistant',message:{role:'assistant',content:[{type:'tool_use',id:item.call.toolUseId,name:'AskUserQuestion',input:{questions:item.call.questions}}]}});event({kind:'metadata'});process.stdout.write('\r\nMETADATA_READY\r\n');render();};
function render(){const q=item.call.questions;let screen=item.first;
if(index>0){const bar='← '+q.map(question=>(answers[question.question]?'☒ ':'☐ ')+question.header).join(' ')+' ✔ Submit →';
screen=index<q.length?bar+'\n│ '+q[index].question+'\n'+q[index].options.map((o,i)=>(i===0?'':' ')+' '+(i+1)+'. '+o.label).join('\n')+'\n 3. Type something.\n 4. Chat about this':bar+'\nReview your answers\nReady to submit your answers?\n 1. Submit answers\n 2. Cancel';
screen+='\nEnter to select · Tab/Arrow keys to navigate · Esc to cancel\n';}
process.stdout.write('\x1b[2J\x1b[H'+screen.replace(/\n/g,'\r\n'));}
event({kind:'startup',pid:process.pid});process.stdin.setRawMode?.(true);process.stdin.resume();
process.stdin.on('data',data=>{const input=data.toString();event({kind:'input',input,index,published});if(!published||done)throw Error('Unexpected input lifecycle');
if(index<2){if(!/^[12]$/.test(input))throw Error('One native digit required');answers[item.call.questions[index].question]=item.call.questions[index].options[Number(input)-1].label;index++;render();}
else{if(input!=='\r')throw Error('Raw Submit required');done=true;persist({type:'user',toolUseResult:{answers},message:{role:'user',content:[{type:'tool_result',tool_use_id:item.call.toolUseId,content:'Answered.'}]}});event({kind:'submitted',answers});process.stdout.write('\x1b[2J\x1b[HNATIVE_PACKET_COMPLETE\r\n');}});
render();if(!item.late)publish();const timer=setInterval(()=>{if(item.late&&fs.existsSync(item.release))publish();},10);
process.on('SIGINT',()=>{clearInterval(timer);process.exit(0);});
`);fs.chmodSync(fake,0o755);
const url=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
fs.writeFileSync(worker,`
import * as fs from 'node:fs';
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(url('claude-pty-runner.ts'))};
import {autoplanSetupDecision} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
const results=[];
for(const item of ${JSON.stringify(cases)}){
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:15000,env:{PACKET_CASE:JSON.stringify(item)}});
try{
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});const seen=new Set();
if(item.late){const pending=readPlanCountTranscript(session.hermeticConfigDir,item.cwd).calls[0];if(pending)throw Error('Expected missing native packet');
if(autoplanSetupDecision(await session.currentScreen(),seen,pending).kind!=='waiting'||seen.size)throw Error('Guessed before native identity');
fs.writeFileSync(item.release,'release');}
await session.waitFor('METADATA_READY',{timeoutMs:3000,pollMs:20});
const inputs=[];
for(let step=0;step<3;step++){
const current=await session.currentScreen();const call=readPlanCountTranscript(session.hermeticConfigDir,item.cwd).calls[0];
const action=autoplanSetupDecision(current,seen,call);
if(action.kind!=='input')throw Error('Expected input at '+step+': '+JSON.stringify({action,current,call}));
session.send(action.input);inputs.push(action.input);for(const signature of action.signatures)seen.add(signature);
if(autoplanSetupDecision(current,seen,call).kind!=='waiting')throw Error('Repeated input on unchanged pane');
await session.waitFor(step===0?'☒ '+item.call.questions[0].header:step===1?'Ready to submit your answers?':'NATIVE_PACKET_COMPLETE',{timeoutMs:3000,pollMs:20});
}
const transcript=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);results.push({name:item.name,inputs,transcript});
}finally{await session.close();}}
fs.writeFileSync(${JSON.stringify(resultFile)},JSON.stringify(results));
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake},stdout:'pipe',stderr:'pipe'});
const killer=setTimeout(()=>child.kill('SIGKILL'),26000);
try{
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);expect(exit,stdout+stderr).toBe(0);
const results=JSON.parse(fs.readFileSync(resultFile,'utf8'));expect(results.length).toBe(4);
for(const [index,result]of results.entries()){
expect(result.inputs).toEqual(['1','1','\r']);expect(result.transcript.calls.length).toBe(1);expect(result.transcript.calls[0].answered).toBe(true);
expect(result.transcript.calls[0].answers).toEqual(Object.fromEntries(cases[index]!.call.questions.map(q=>[q.question,q.options[0]!.label])));
const events=fs.readFileSync(cases[index]!.events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
expect(events.filter(e=>e.kind==='input').map(e=>({input:e.input,index:e.index,published:e.published}))).toEqual([
{input:'1',index:0,published:true},{input:'1',index:1,published:true},{input:'\r',index:2,published:true}]);
expect(events.filter(e=>e.kind==='submitted').length).toBe(1);expect(()=>process.kill(events[0].pid,0)).toThrow();
}
}finally{
clearTimeout(killer);child.kill('SIGKILL');for(const item of cases)if(fs.existsSync(item.events)){
const pid=JSON.parse(fs.readFileSync(item.events,'utf8').split('\n')[0]!).pid;
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
}fs.rmSync(dir,{recursive:true,force:true});
}
},30000);
const adV2Packet = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-ad-v2-packet.json'), 'utf8')) as {pendingCall: NativePlanQuestionCall; screen: string};
test('AD v2 actual setup packet chooses routing and standard review with the existing native identity',()=>{
const seen=new Set<string>(),call=adV2Packet.pendingCall;
expect(autoplanSetupDecision(adV2Packet.screen,seen).kind).toBe('waiting');
commit(adV2Packet.screen,seen,call,'1');
// Only the first pane was retained live. Later panes are explicit native-question projections.
commit(pane(call,1,[0]),seen,call,'2');
commit(pane(call,2,[0,1]),seen,call,'\r');
});
test('AD v2 setup policy uses the task and actions across presentation and option order',()=>{
for(const variant of ['numbered','unprefixed','different explanation'])for(const reverseQuestions of [false,true])for(const reverseOptions of [false,true]){
const call=structuredClone(adV2Packet.pendingCall);
call.questions.forEach((q,index)=>{
q.question=q.question.replace(/^D\d+\s*[—–:-]\s*/,variant==='unprefixed'?'':`D${31+index}: `);
if(variant==='different explanation')q.question=q.question.split('\n')[0]+'\nProject/branch/task: disposable review fixture, another branch and release.\nELI10: This setup changes how later sessions find workflow context.\nStakes if we pick wrong: an extra setup step.\nRecommendation: Keep the offered actions explicit.\nNet: setup now versus a direct review.';
});
if(reverseQuestions)call.questions.reverse();if(reverseOptions)call.questions.forEach(q=>q.options.reverse());
const seen=new Set<string>();
for(let i=0;i<2;i++){
const ordinary=call.questions[i]!.header==='Routing'?1:2;
commit(pane(call,i,i===1?[0]:[]),seen,call,String(reverseOptions?3-ordinary:ordinary));
}
commit(pane(call,2,[0,1]),seen,call,'\r');
}
});
test('AD v2 setup cannot borrow a header, subject or adjacent question for a different decision',()=>{
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
c=>{c.questions[0]!.header='Product router';},
c=>{c.questions[1]!.header='Deployment';},
c=>{[c.questions[0]!.header,c.questions[1]!.header]=[c.questions[1]!.header,c.questions[0]!.header];},
c=>{c.questions[0]!.question=c.questions[0]!.question.replace(/^.*\n/,'D1 — Should the application route requests through a proxy?\n');},
c=>{c.questions[1]!.question=c.questions[1]!.question.replace(/^.*\n/,'D2 — Should we add an office-hours page to the product?\n');},
c=>{c.questions[0]!.question='The plan quotes: '+c.questions[0]!.question;},
c=>{c.questions[1]!.question='```text\n'+c.questions[1]!.question+'\n```';},
c=>{c.questions[1]!.question=c.questions[1]!.question.split('\n').map(l=>'> '+l).join('\n');},
c=>{c.questions[1]!.question+=' Should we remove the authorization check?';},
c=>{c.questions[1]={...structuredClone(c.questions[1]!),question:'Approve deployment to production?',header:'Approval'};},
];
for(const change of changes){const call=structuredClone(adV2Packet.pendingCall);change(call);const seen=new Set<string>();
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
});
test('AD v2 setup rejects conditional, contradictory and ambiguous actions in either tab',()=>{
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
c=>{c.questions[0]!.options[0]!.description='Do not add routing rules to CLAUDE.md.';},
c=>{c.questions[0]!.options[1]!.description='Add routing rules to CLAUDE.md after declining.';},
c=>{c.questions[1]!.options[0]!.description='Skip the design doc and begin the review now.';},
c=>{c.questions[1]!.options[1]!.description='Run /office-hours first, then proceed with standard review.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review after completing /office-hours.';},
c=>{c.questions[1]!.options[1]!.description='No review will run.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review?';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review but do not run it.';},
c=>{c.questions[1]!.options[1]!.description='Review starts now, but not yet.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review when /office-hours completes.';},
c=>{c.questions[1]!.options[1]!.description='Review starts immediately after completing /office-hours.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review once the design doc is complete.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review if the tests pass.';},
c=>{c.questions[1]!.options[1]!.description='Skip the CEO review and proceed directly to engineering.';},
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review only if the tests pass.';},
c=>{c.questions[1]!.question+=' You must run /office-hours before the review.';},
c=>{c.questions[1]!.question+=' Standard review is forbidden until /office-hours completes.';},
c=>{c.questions[0]!.options[0]!.label+=' and implement the feature';},
c=>{c.questions[1]!.options[1]!.label+=' if the tests pass';},
c=>{c.questions[0]!.options[0]!.description+=' Also delete the authorization check.';},
c=>{c.questions[1]!.options[1]!.description+=' Also deploy to production.';},
c=>{c.questions[0]!.options[1]=structuredClone(c.questions[0]!.options[0]!);},
c=>{c.questions[1]!.options.push({label:'Skip the remaining review phases'});},
];
for(const change of changes){const call=structuredClone(adV2Packet.pendingCall);change(call);const seen=new Set<string>();
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
});
test('AD v2 setup retains complete native identity and current-pane requirements',()=>{
const first=adV2Packet.screen,call=adV2Packet.pendingCall;
// The example label must introduce the panel, not precede unrelated earlier transcript rows.
for(const screen of ['Example panel:\n'+pane(call,0),'```text\n'+first,first+'\nContinuing.',
first.replace('Design doc','Different tab'),first.replace('Esc to cancel','Esc to'),
first.replace('Add routing rules to CLAUDE.md (recommended)','Add routing rules to OTHER.md (recommended)')]){
expect(screen).not.toBe(first);expect(autoplanSetupDecision(screen,new Set(),call).kind).toBe('waiting');
}
for(const delta of [{answered:true},{failed:true},{sessionId:''},{toolUseId:''}])
expect(autoplanSetupDecision(first,new Set(),{...call,...delta}).kind).toBe('waiting');
});
test('AD v2 setup fixture selects the existing Autoplan paid case only',()=>{
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-setup-ad-v2-packet.json');
const owners=Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes('test/fixtures/autoplan-setup-ad-v2-packet.json')).map(([name])=>name);
expect(owners).toEqual(['autoplan-chain-pty']);
});
test('AD v2 selected review action allows short affirmative descriptions with dynamic tradeoffs',()=>{
for(const description of ['Proceed with standard review. The plan already states its goals.', 'Review begins now using the existing plan. No separate design artifact is created.', 'Start the standard review immediately with the supplied context.']){
const call=structuredClone(adV2Packet.pendingCall);call.questions[1]!.options[1]!.description=description;
expect(autoplanSetupDecision(pane(call,0),new Set(),call).kind).toBe('input');
}
});
+916
View File
@@ -0,0 +1,916 @@
import { describe, expect, test } from 'bun:test';
import { autoplanRoutingSetupInput, autoplanSetupDecision } from './helpers/autoplan-setup-question';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
const CLIPPED_ROUTING_N = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-routing-n-screen.txt'), 'utf8');
describe('current routing title survives a scrolled native header before metadata flushes', () => {
test('exact N frame selects the offered Add action once with a native digit only', () => {
expect(CLIPPED_ROUTING_N).not.toMatch(/[☐□]/);
const seen = new Set<string>();
const decision = autoplanSetupDecision(CLIPPED_ROUTING_N, seen);
expect(decision.kind).toBe('input');
if (decision.kind !== 'input') throw Error('Expected native setup input');
expect(decision.input).toBe('1');
expect(seen.size).toBe(0);
for (const signature of decision.signatures) seen.add(signature);
expect(autoplanSetupDecision(CLIPPED_ROUTING_N, seen).kind).toBe('waiting');
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-routing-n-screen.txt');
});
test('equivalent direct title and reordered opposed choices retain picker binding', () => {
const frame = CLIPPED_ROUTING_N.replace('D1 — Add skill', 'D9 — Add gstack skill');
const swapped = frame.replace('1. Add routing rules (Recommended)', '1. Skip, invoke manually')
.replace('2. Skip, invoke manually', '2. Add routing rules (Recommended)');
expect(autoplanSetupDecision(frame, new Set())).toMatchObject({kind:'input',input:'1'});
expect(autoplanSetupDecision(swapped, new Set())).toMatchObject({kind:'input',input:'2'});
});
test('copied, stale, incomplete, ambiguous and substantive panels cannot borrow the top routing identity', () => {
for (const frame of [
'Example panel:\n' + CLIPPED_ROUTING_N,
'Quoted source:\n' + CLIPPED_ROUTING_N,
'```text\n' + CLIPPED_ROUTING_N + '\n```',
'~~~~text\n' + CLIPPED_ROUTING_N,
CLIPPED_ROUTING_N.split('\n').map(line => ' ' + line).join('\n'),
CLIPPED_ROUTING_N.split('\n').map(line => '> ' + line).join('\n'),
CLIPPED_ROUTING_N + '\n⏺ Continuing the review.',
CLIPPED_ROUTING_N.replace('Esc to cancel', 'Esc to'),
CLIPPED_ROUTING_N.replace(' 1.', ' 1.'),
CLIPPED_ROUTING_N.replace(' 1.', ' 1.').replace(' 2.', ' 2.'),
CLIPPED_ROUTING_N.replace(' 2.', ' 2.'),
CLIPPED_ROUTING_N.replace('1. Add', '1. [ ] Add'),
CLIPPED_ROUTING_N.replace('│\n│ Project', '│ ← ☐ Routing ✔ Submit →\n│ Project'),
CLIPPED_ROUTING_N.replace(' 4. Chat about this', ''),
CLIPPED_ROUTING_N.replace('2. Skip, invoke manually', '2. Add routing rules (Recommended)'),
CLIPPED_ROUTING_N.replace('2. Skip, invoke manually', '2. Delete routing and migrate the product'),
CLIPPED_ROUTING_N.replace('routing-injection>', 'product-routing>'),
CLIPPED_ROUTING_N.replace('routing-injection>', 'routing-injection'),
CLIPPED_ROUTING_N.replace('│ Project/branch:', '│ <gstack-qid:routing-injection>\n│ Project/branch:'),
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'Choose the product API router for CLAUDE.md?'),
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'The spec quotes Add skill routing rules to CLAUDE.md?'),
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'Add skill routing rules to README.md?'),
CLIPPED_ROUTING_N.replace(' <gstack-qid:routing-injection>', '').replace('│ Net:', '│ <gstack-qid:routing-injection> Net:'),
CLIPPED_ROUTING_N.replace('│ ELI10:', '│ ```text\n│ ELI10:'),
CLIPPED_ROUTING_N.replace('│ ELI10:', '│ > Quoted source:\n│ ELI10:'),
]) expect(autoplanSetupDecision(frame, new Set()).kind, frame).not.toBe('input');
});
test('present native metadata keeps its full existing identity binding', () => {
const before = CLIPPED_ROUTING_N.split(' 1.')[0]!.replace(/^[│┃] ?/gm, '').trim();
const call: any = {toolUseId:'n-routing',sessionId:'n',timestamp:'2026-09-09T01:10:05Z',answered:false,failed:false,
questions:[{header:'Routing',question:before,options:[{label:'Add routing rules (Recommended)'},{label:'Skip, invoke manually'}]}]};
expect(autoplanSetupDecision(CLIPPED_ROUTING_N,new Set(),call)).toMatchObject({kind:'input',input:'1'});
for (const mutate of [
(q:any) => {q.failed=true;}, (q:any) => {q.answered=true;}, (q:any) => {q.questions=[];},
(q:any) => {q.questions.push(structuredClone(q.questions[0]));},
(q:any) => {q.questions[0].multiSelect=true;},
(q:any) => {q.questions[0].question='Unrelated finding <gstack-qid:routing-injection>';},
(q:any) => {q.questions[0].options[1].label='Another choice';},
]) {const changed=structuredClone(call);mutate(changed);expect(autoplanSetupDecision(CLIPPED_ROUTING_N,new Set(),changed).kind).not.toBe('input');}
const seen=new Set<string>();
const early=autoplanSetupDecision(CLIPPED_ROUTING_N,seen);
if(early.kind!=='input')throw Error('Expected initial input');
for(const signature of early.signatures)seen.add(signature);
expect(autoplanSetupDecision(CLIPPED_ROUTING_N,seen,call).kind).toBe('waiting');
});
});
test.skipIf(process.platform === 'win32')('real PTY clipped routing advances from the exact current panel with one digit and no Enter', async () => {
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-clipped-routing-'));
const fake=path.join(dir,'fake-claude');const events=path.join(dir,'events.jsonl');
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import * as fs from 'node:fs';
const emit=value=>fs.appendFileSync(process.env.ROUTING_EVENTS,JSON.stringify(value)+'\n');
emit({kind:'started',pid:process.pid});
process.stdin.setRawMode?.(true);process.stdin.resume();
process.stdin.on('data',data=>{emit({kind:'input',data:data.toString()});process.stdout.write('\r\nNATIVE_SETUP_ACCEPTED\r\n');});
process.stdout.write(fs.readFileSync(process.env.ROUTING_SCREEN,'utf8').replace(/\n/g,'\r\n'));
process.on('SIGINT',()=>process.exit(0));
`);fs.chmodSync(fake,0o755);
const worker=path.join(dir,'worker.ts');const resultFile=path.join(dir,'result.json');
const helper=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
fs.writeFileSync(worker,`
import * as fs from 'node:fs';
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(helper('claude-pty-runner.ts'))};
import {autoplanSetupDecision} from ${JSON.stringify(helper('autoplan-setup-question.ts'))};
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
const session=await launchClaudePty({cwd:${JSON.stringify(dir)},observeScreen:true,timeoutMs:15000,
env:{ROUTING_EVENTS:process.env.ROUTING_EVENTS,ROUTING_SCREEN:process.env.ROUTING_SCREEN}});
try{
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
const screen=await session.currentScreen();
const decision=autoplanSetupDecision(screen,new Set());
if(decision.kind!=='input'||decision.input!=='1')throw Error('Expected current setup: '+JSON.stringify(decision));
session.send(decision.input);
await session.waitFor('NATIVE_SETUP_ACCEPTED',{timeoutMs:3000,pollMs:20});
fs.writeFileSync(${JSON.stringify(resultFile)},JSON.stringify({screen,decision}));
}finally{await session.close();}
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,
ROUTING_EVENTS:events,ROUTING_SCREEN:path.join(import.meta.dir,'fixtures/autoplan-routing-n-screen.txt')},stdout:'pipe',stderr:'pipe'});
const killer=setTimeout(()=>child.kill('SIGKILL'),17000);
try {
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
expect(exit,stdout+stderr).toBe(0);
const result=JSON.parse(fs.readFileSync(resultFile,'utf8'));
expect(result.screen).not.toMatch(/[☐□]/);
expect(result.decision).toMatchObject({kind:'input',input:'1'});
const recorded=fs.readFileSync(events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
expect(recorded.filter(e=>e.kind==='input')).toEqual([{kind:'input',data:'1'}]);
expect(()=>process.kill(recorded[0].pid,0)).toThrow();
} finally {
clearTimeout(killer);child.kill('SIGKILL');
if(fs.existsSync(events)){
const pid=JSON.parse(fs.readFileSync(events,'utf8').split('\n')[0]!).pid;
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
}
fs.rmSync(dir,{recursive:true,force:true});
}
},20000);
// Sanitized terminal frame from the 2026-09-08 autoplan timeout. The qid is
// visibly incomplete; the prompt body and explicit choices remain intact.
const CAPTURE = [
'─'.repeat(120),
'Planning: /tmp/hermetic/.claude/plans/modular-bouncing-swing.md',
'─'.repeat(120),
' ☐ Routing rules',
"│ gstack works best when your project's CLAUDE.md includes skill routing rules. Add them now?",
'│<gstck-qid:routing-injectin>',
'1.Addroutingrules(Recommended)',
'CreatesCLAUDE.mdwithskillroutingrulessogstackknowswhentoinvoke/office-hours,/autoplan,/ship,/qa,',
"etc.automatically.We'lldothisafterthereview.",
'2.Nothanks',
"Skip—I'llinvokeskillsmanually.Youcanenablethislaterbyrunninggstack-configsetrouting_declinedfalse.",
'3.Typesomething.',
'4.Chataboutthis',
'Entertoselect·↑/↓tonavigate·Esctocancel',
].join('\r\r');
// Targeted-a stalled on this complete menu for the full test budget. Parsing
// retained its identity and choices; the setup helper rejected their wording.
const CURRENT_CAPTURE = [
' ☐ Routing rules',
'',
'Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>',
'',
'1.AddtoCLAUDE.md(recommended)',
'',
'Appendsa##SkillroutingsectiontoCLAUDE.mdandcommitsit.Futuresessionswillauto-invoketherightskill',
'(/investigateforbugs,/shipforPRs,/qafortesting,etc.)withoutmanualinvocation.',
'',
'2.Skip—invokemanually',
'',
"Nofilechanges.You'llcontinuecallingskillsbyname.Canaddroutingruleslater.",
'',
'3.Typesomething.',
'─'.repeat(120),
'4.Chataboutthis',
'Entertoselect·↑/↓tonavigate·Esctocancel',
].join('\n');
// Targeted-b's first attempt stayed on this complete setup menu until its
// 15-minute deadline. The parser retained the prompt and both labels, but
// the setup selector rejected "No thanks, invoke manually".
const B_CAPTURE = [
' ☐ CLAUDE.md',
'',
'│ D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>',
'│',
'│ELI10:ThisprojecthasnoCLAUDE.md.Thatfileiswheregstacklooksforroutingrules—instructionstellingClaude',
'│Codewhichskilltoauto-invokeforwhichrequest(e.g."ship→/ship","bugs→/investigate").Withoutityoutype',
'│theskillnameeverytime.Withit,gstackcanrecognizeyourintentandrouteautomatically.',
'│',
'│Stakesifweskip:Noauto-routing;youinvokeskillsmanuallyeachsession.',
'│',
'│Recommendation:A—one-timesetup,saveskeystrokesoneveryfuturesession.',
'│Completeness:A=9/10,B=5/10',
'',
'1.AddroutingrulestoCLAUDE.md(Recommended)',
'AppendsthestandardgstackroutingblocktoanewCLAUDE.mdandcommitsit.Doneonce,activeforever.',
'2.Nothanks,invokemanually',
'SkipCLAUDE.mdsetup.Youcontinuecalling/autoplan,/ship,/qa,etc.bynameeachtime.',
'3.Typesomething.',
'─'.repeat(120),
'4.Chataboutthis',
'Entertoselect·↑/↓tonavigate·Esctocancel',
].join('\r\r');
// Fresh broad retry: the complete setup menu uses a noun for the manual
// alternative. This is the same opposed setup action as "invoke manually".
const FRESH_RETRY_CAPTURE = [
'☐Routingsetup',
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
'1.Addroutingrules(Recommended)',
'AppendskillroutingrulestoCLAUDE.mdsoClaudeautomaticallyinvokestherightskillforproduct,engineering,',
'design,andshipworkflows.Willbedoneafterplanapproval(planmodeisactivenow).',
'2.Nothanks,manualinvocation',
"Skip—I'llinvokeskillsmanually.Thispromptwon'tappearagain.",
'3.Typesomething.',
'─'.repeat(120),
'4.Chataboutthis',
'Entertoselect·↑/↓tonavigate·Esctocancel',
].join('\n');
describe('autoplan routing setup handling', () => {
// Source-F retry's first complete frame preceded damaged terminal redraws.
const F_SETUP_CAPTURE = [
'Planning: /tmp/hermetic/.claude/plans/deep-coalescing-valiant.md',
'☐Skillrouting',
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
'1.AddroutingrulestoCLAUDE.md',
'AppendsskillroutingrulestoCLAUDE.mdsogstackauto-invokestherightskillforcommonrequests(review,ship,',
'investigate,etc.).Willbecommittedtotherepo.(recommended)',
'2.Nothanks,skip',
"I'llinvokeskillsmanually.Youcanaddroutinglater.",
'3.Typesomething.',
'4.Chataboutthis',
'Entertoselect·↑/↓tonavigate·Esctocancel',
].join('\r\r');
test('answers the captured combined decline action once, regardless of option order', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE, seen)).toBeNull();
const reordered = F_SETUP_CAPTURE.replace('1.AddroutingrulestoCLAUDE.md', '1.Nothanks,skip')
.replace('2.Nothanks,skip', '2.AddroutingrulestoCLAUDE.md');
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip—invoke skills manually'), new Set())).toBe('1');
});
test('does not infer a routing answer from damaged, ambiguous, or unrelated setup choices', () => {
for (const frame of [
F_SETUP_CAPTURE.replace('Addroutingrules', 'Addrutingrules'),
F_SETUP_CAPTURE.replace('Nothanks,skip', 'Nothank,skip'),
F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip the review'),
F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip then delete CLAUDE.md'),
F_SETUP_CAPTURE.replace('3.Typesomething.', '3.Skip'),
F_SETUP_CAPTURE.replace("gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?", 'Which routing design should the application use?'),
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
});
test('answers the fresh retry manual-invocation setup once in either option order', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE, seen)).toBeNull();
const reordered = FRESH_RETRY_CAPTURE.replace('1.Addroutingrules(Recommended)', '1.Nothanks,manualinvocation')
.replace('2.Nothanks,manualinvocation', '2.Addroutingrules(Recommended)');
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
});
test('requires opposed manual setup actions and rejects ambiguous or unrelated choices', () => {
for (const decline of [
'No thanks, delete the file manually',
'No thanks, manual data migration',
'No thanks, invoke the deploy manually',
'Manual deployment invocation',
'Accept recommendation',
'No thanks, manual invocation then delete CLAUDE.md',
]) {
const frame = FRESH_RETRY_CAPTURE.replace('Nothanks,manualinvocation', decline);
expect(autoplanRoutingSetupInput(frame, new Set()), decline).toBeNull();
}
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE.replace('3.Typesomething.', '3.Add routing rules'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE.replace('3.Typesomething.', '3.Skip—invoke manually'), new Set())).toBeNull();
const review = FRESH_RETRY_CAPTURE.replace(
"gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
'Which product routing design should we ship? <gstack-qid:routing-injection>',
);
expect(autoplanRoutingSetupInput(review, new Set())).toBeNull();
});
test('answers the captured setup once, using the full question identity', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(CAPTURE, seen)).toBeNull();
expect(autoplanRoutingSetupInput(CAPTURE.replace('works best', 'works best'), seen)).toBeNull();
});
test('chooses Add routing rules by label when option order changes', () => {
const reordered = CAPTURE.replace('1.Addroutingrules(Recommended)', '1.Nothanks')
.replace('2.Nothanks', '2.Add routing rules (Recommended)');
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
});
test('accepts the full option labels captured from the subsequent live setup prompt', () => {
const fullLabels = CAPTURE.replace('Addroutingrules(Recommended)', 'Add routing rules to CLAUDE.md (Recommended)')
.replace('2.Nothanks', "2.No thanks, I'll invoke skills manually");
expect(autoplanRoutingSetupInput(fullLabels, new Set())).toBe('1');
expect(autoplanRoutingSetupInput(fullLabels.replace('CLAUDE.md (Recommended)', 'product routes (Recommended)'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(fullLabels.replace("I'll invoke skills manually", 'delete the existing rules'), new Set())).toBeNull();
});
test('answers the current captured CLAUDE.md setup, including reordered choices, once', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE, seen)).toBeNull();
const reordered = CURRENT_CAPTURE.replace('1.AddtoCLAUDE.md(recommended)', '1.Skip—invokemanually')
.replace('2.Skip—invokemanually', '2.AddtoCLAUDE.md(recommended)');
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE.replace('to CLAUDE.md?', "to this project's CLAUDE.md?"), new Set())).toBe('1');
});
test('the current wording still requires both explicit setup choices and the CLAUDE.md target', () => {
for (const frame of [
CURRENT_CAPTURE.replace('to CLAUDE.md?', 'to the application API?'),
CURRENT_CAPTURE.replace('AddtoCLAUDE.md(recommended)', 'Acceptrecommendation'),
CURRENT_CAPTURE.replace('Skip—invokemanually', 'Deferthisfinding'),
CURRENT_CAPTURE.replace('AddtoCLAUDE.md(recommended)', 'Deletetheexistingroutingrules'),
CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Should we expand the current feature?'),
]) expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
});
test('recognizes the native A retry packet with its abbreviated manual-decline label', () => {
const retry = CAPTURE.replace('Addroutingrules(Recommended)', 'Add to CLAUDE.md (Recommended)')
.replace('2.Nothanks', '2.No thanks, manual');
expect(autoplanRoutingSetupInput(retry, new Set())).toBe('1');
expect(autoplanRoutingSetupInput(retry.replace('No thanks, manual', 'No thanks, delete it'), new Set())).toBeNull();
});
test('answers the exact B timeout menu by its routing label, in either order', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(B_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(B_CAPTURE, seen)).toBeNull();
const reordered = B_CAPTURE.replace('1.AddroutingrulestoCLAUDE.md(Recommended)', '1.Nothanks,invokemanually')
.replace('2.Nothanks,invokemanually', '2.AddroutingrulestoCLAUDE.md(Recommended)');
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
expect(autoplanRoutingSetupInput(B_CAPTURE.replace('Nothanks,invokemanually', 'Nothanks,deletethefilemanually'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(B_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Which routing design should the application use?'), new Set())).toBeNull();
});
test('recognizes the setup premise without depending on its closing sentence', () => {
const openings = [
"gstack works best when your project's CLAUDE.md includes skill routing rules. Would you like to add them?",
"gstack works best when your project's CLAUDE.md includes skill routing rules. Enable them for this repository?",
'Should we configure skill routing rules for gstack in CLAUDE.md?',
'Set up gstack skill routing rules in CLAUDE.md.',
];
for (const opening of openings) {
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>', opening);
expect(autoplanRoutingSetupInput(frame, new Set()), opening).toBe('1');
}
});
test('recognizes an intact setup qid with an explicit CLAUDE.md action and opposed manual decline', () => {
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Configure this projects CLAUDE.md?');
expect(autoplanRoutingSetupInput(frame, new Set())).toBe('1');
expect(autoplanRoutingSetupInput(frame.replace('gstack-qid:routing-injection', 'gstack-qid:product-routing'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(frame.replace('AddtoCLAUDE.md(recommended)', 'Acceptrecommendation'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(frame.replace('Skip—invokemanually', 'Deferthisfinding'), new Set())).toBeNull();
});
test('keeps generic review, quoted premises and different routing targets out of setup handling', () => {
for (const question of [
'Which dashboard layout should we ship?',
'Add routing rules to the application API? <gstack-qid:product-routing>',
'The plan quotes gstack CLAUDE.md skill routing rules. Which API design should we use?',
'The document references gstack skill routing rules in CLAUDE.md. Should we expand the feature?',
]) {
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>', question);
expect(autoplanRoutingSetupInput(frame, new Set()), question).toBeNull();
}
});
test('waits for complete recognized choices rather than guessing a default', () => {
expect(autoplanRoutingSetupInput(CAPTURE.replace('2.Nothanks', '2.Ask me later'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput(CAPTURE.replace('Addroutingrules(Recommended)', 'Accept recommendation'), new Set())).toBeNull();
expect(autoplanRoutingSetupInput('1.Addroutingrules(Recommended)\r2.Nothanks', new Set())).toBeNull();
});
test('never answers review or taste questions, even with a routing qid or the same choices', () => {
const prompts = [
'Which visual direction should this settings page use?',
'Should the payment handler bypass the existing dispatcher?',
'Add routing rules to the product API now? <gstack-qid:routing-injection>',
'The plan quotes CLAUDE.md skill routing rules. Should we change this feature?',
];
for (const prompt of prompts) {
const frame = `☐ Review decision\r${prompt}\r1.Addroutingrules(Recommended)\r2.Nothanks`;
expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
}
});
test('setup helper and captured-frame changes select the autoplan eval only', () => {
for (const file of ['test/helpers/autoplan-setup-question.ts', 'test/autoplan-setup-question.test.ts', 'test/fixtures/autoplan-routing-n-screen.txt']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
}
});
});
// Source-G's retry remained at this actual captured menu until shard timeout.
// The action is intact; cumulative ANSI stripping loses the courtesy's 'o'.
// A real xterm replay retains it in the prior screen cell.
const G_ROUTING_CAPTURE = [
'☐Routingrules',
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Wouldyouliketoaddthem?",
'1.AddroutingrulestoCLAUDE.md',
'AppendsstandardskillroutingrulestoCLAUDE.md(creatingitifabsent)andcommits.Meansgstackskillslike',
'/autoplan,/ship,/qaetc.getinvokedautomaticallywhenthetaskmatches.(recommended)',
"2. N thanks, I'll invokeskillsmanually",
'Skiprouting setup. You can re-enable later by removing the routing_declined flag.',
'3.Typesomething.',
'4.Chataboutthis',
'Enter toselect · ↑/↓ to navigate · Esc to cancel',
].join('\r');
describe('autoplan routing action survives courtesy repaint', () => {
test('selects the explicit Add action once in the captured G menu, in both orders', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE, seen)).toBeNull();
const reversed = G_ROUTING_CAPTURE.replace('1.AddroutingrulestoCLAUDE.md', "1.N thanks, I'll invokeskillsmanually")
.replace("2. N thanks, I'll invokeskillsmanually", '2.AddroutingrulestoCLAUDE.md');
expect(autoplanRoutingSetupInput(reversed, new Set())).toBe('2');
});
test('the actual manual-invocation action needs no courtesy formula', () => {
for (const action of ['Manual invocation', 'Invoke skills manually', "I'll invoke skills manually", 'Thanks, invoke manually']) {
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE.replace("N thanks, I'll invokeskillsmanually", action), new Set()), action).toBe('1');
}
});
test('still requires exact opposed setup actions and a genuine routing premise', () => {
for (const label of [
'N thanks', 'Invoke the deployment manually', 'N thanks, manual data migration',
'Delete CLAUDE.md, invoke skills manually', 'No thanks, invoke skills manually then delete CLAUDE.md',
'Skip the review, invoke skills manually', 'Skip the review thanks, invoke skills manually',
]) expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE.replace("N thanks, I'll invokeskillsmanually", label), new Set()), label).toBeNull();
for (const frame of [
G_ROUTING_CAPTURE.replace("gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Wouldyouliketoaddthem?", 'Which application router should we implement?'),
G_ROUTING_CAPTURE.replace('AddroutingrulestoCLAUDE.md', 'AddrutingrulestoCLAUDE.md'),
G_ROUTING_CAPTURE.replace('3.Typesomething.', '3.Invoke skills manually'),
G_ROUTING_CAPTURE.replace('3.Typesomething.', '3.Add routing rules'),
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
});
});
const PREREQUISITE_CAPTURE = " ☐ Design doc\n\n│ No design doc found for this branch. /office-hours produces a structured problem statement, premise challenge, and\n│ explored alternatives — it gives this review much sharper input to work with. Takes about 10 minutes. The design doc\n│ is per-feature, not per-product — it captures the thinking behind this specific change. Run /office-hours first?\n\n 1. Run /office-hours now\n Runs /office-hours to produce a design doc first, then picks up the full autoplan review right after. (~10 min)\n 2. Skip — proceed with standard review\n Skips /office-hours and runs the autoplan review pipeline now using the existing plan file as input.\n 3. Type something.\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n 4. Chat about this\n\nEnter to select · ↑/↓ to navigate · Esc to cancel\n";
const prerequisiteQuestion = {
header: 'Design doc',
question: "No design doc found for this branch. /office-hours produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input to work with. Takes about 10 minutes. The design doc is per-feature, not per-product — it captures the thinking behind this specific change. Run /office-hours first?",
options: [{ label: 'Run /office-hours now' }, { label: 'Skip — proceed with standard review' }],
};
const prerequisiteCall = () => ({
sessionId: 'prerequisite-session', toolUseId: 'prerequisite-call',
answered: false, failed: false, questions: [structuredClone(prerequisiteQuestion)],
});
function prerequisiteMenu(reverse = false) {
if (!reverse) return PREREQUISITE_CAPTURE;
return PREREQUISITE_CAPTURE
.replace('1. Run /office-hours now', '1. Skip — proceed with standard review')
.replace('2. Skip — proceed with standard review', '2. Run /office-hours now');
}
describe('autoplan optional design-doc prerequisite', () => {
test('the exact K native screen declines the optional prerequisite by label', () => {
for (const reverse of [false, true]) {
const frame = prerequisiteMenu(reverse);
expect(autoplanRoutingSetupInput(frame, new Set())).toBe(reverse ? '1' : '2');
const native = prerequisiteCall(); if (reverse) native.questions[0]!.options.reverse();
expect(autoplanRoutingSetupInput(frame, new Set(), native)).toBe(reverse ? '1' : '2');
}
});
test('quoted panels and menus followed by new output are not active input', () => {
for (const frame of [
'Example panel:\n```text\n' + PREREQUISITE_CAPTURE + '\n```\n',
'Example panel:\n~~~text\n' + PREREQUISITE_CAPTURE,
'Example panel:\n' + PREREQUISITE_CAPTURE,
PREREQUISITE_CAPTURE.split('\n').map(line => ' ' + line).join('\n'),
'The document quotes this panel:\n────────────────────\n' + PREREQUISITE_CAPTURE,
PREREQUISITE_CAPTURE + '\n⏺ Continuing the review without office hours.\n',
PREREQUISITE_CAPTURE + '\n 1. A new menu\n 2. Another choice\n',
]) for (const native of [undefined, prerequisiteCall()]) {
expect(autoplanRoutingSetupInput(frame, new Set(), native)).toBeNull();
}
expect(autoplanRoutingSetupInput('```text\nearlier real code\n```\n────────────────────\n' + PREREQUISITE_CAPTURE, new Set())).toBe('2');
});
test('late native identity does not re-answer the retained menu', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen)).toBe('2');
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, prerequisiteCall())).toBeNull();
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen)).toBeNull();
});
test('unrelated, failed, mixed and checkbox native calls do not borrow the setup menu', () => {
for (const mutate of [
(call: ReturnType<typeof prerequisiteCall>) => { call.questions[0]!.question = 'Should we change the dashboard design?'; },
(call: ReturnType<typeof prerequisiteCall>) => { call.failed = true; },
(call: ReturnType<typeof prerequisiteCall>) => { call.answered = true; },
(call: ReturnType<typeof prerequisiteCall>) => { call.questions.push({ header:'Finding', question:'Fix missing auth?', options:[{label:'Fix it'},{label:'Defer'}] }); },
(call: ReturnType<typeof prerequisiteCall>) => { Object.assign(call.questions[0]!, {multiSelect:true}); },
]) {
const native = prerequisiteCall(); mutate(native);
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, native)).toBeNull();
// Waiting for correct metadata must not mark an unanswered UI as sent.
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, prerequisiteCall())).toBe('2');
}
});
test('arbitrary skip, outside offers, mixed actions and prose examples remain unanswered', () => {
for (const frame of [
PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Skip this security check'),
PREREQUISITE_CAPTURE.replaceAll('/office-hours', '/codex'),
PREREQUISITE_CAPTURE.replace('3. Type something.', '3. Fix the missing authorization check'),
PREREQUISITE_CAPTURE.replace('No design doc found for this branch.', 'A dashboard design issue was found.'),
PREREQUISITE_CAPTURE.replace(' ☐ Design doc', 'Example choices:').replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
]) expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
});
});
test.skipIf(process.platform === 'win32')('real PTY prerequisite answer survives early and deferred native records without a second key', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-autoplan-prereq-'));
const fake = path.join(dir, 'fake-claude');
const worker = path.join(dir, 'worker.ts');
const resultFile = path.join(dir, 'result.json');
const cases = [false, true].flatMap(early => [false, true].map(reverse => {
const name = `${early ? 'early' : 'deferred'}-${reverse ? 'reversed' : 'original'}`;
const q = structuredClone(prerequisiteQuestion); if (reverse) q.options.reverse();
return { name, early, cwd: path.join(dir, name), record: path.join(dir, name + '.jsonl'),
question: q, frame: prerequisiteMenu(reverse), expected: reverse ? '1' : '2' };
}));
for (const item of cases) fs.mkdirSync(item.cwd);
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
import * as fs from 'node:fs';
import * as path from 'node:path';
const item = JSON.parse(process.env.PREREQUISITE_REPLAY);
const record = event => fs.appendFileSync(item.record, JSON.stringify(event) + '\n');
record({type:'startup',pid:process.pid});
const folder = path.join(process.env.CLAUDE_CONFIG_DIR, 'projects', 'fixture');
fs.mkdirSync(folder, {recursive:true});
const transcript = path.join(folder, item.name + '.jsonl');
let logged = false;
function writeCall() {
if (logged) return; logged = true;
fs.appendFileSync(transcript, JSON.stringify({type:'assistant',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),
message:{role:'assistant',content:[{type:'tool_use',id:'prerequisite',name:'AskUserQuestion',input:{questions:[item.question]}}]}})+'\n');
}
if (item.early) writeCall();
process.stdin.setRawMode?.(true);
let answered = false;
process.stdin.on('data', data => {
record({type:'input',data:data.toString()});
for (const key of data.toString()) if (/^[12]$/.test(key) && !answered) {
answered = true; writeCall();
const label = item.question.options[Number(key)-1].label;
fs.appendFileSync(transcript, JSON.stringify({type:'user',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),
toolUseResult:{answers:{[item.question.question]:label}},
message:{role:'user',content:[{type:'tool_result',tool_use_id:'prerequisite',content:'answered'}]}})+'\n');
process.stdout.write('\x1b[2J\x1b[H'+item.frame+'\nSETUP_ANSWERED\n');
}
});
process.stdout.write('\x1b[2J\x1b[H'+item.frame);
process.on('SIGINT', () => process.exit(0));
process.stdin.resume();
`);
fs.chmodSync(fake, 0o755);
const moduleUrl = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
fs.writeFileSync(worker, `
import {launchClaudePty} from ${JSON.stringify(moduleUrl('claude-pty-runner.ts'))};
import {autoplanRoutingSetupInput} from ${JSON.stringify(moduleUrl('autoplan-setup-question.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(moduleUrl('plan-count-transcript.ts'))};
const results = await Promise.all(${JSON.stringify(cases)}.map(async item => {
const session = await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:20000,env:{PREREQUISITE_REPLAY:JSON.stringify(item)}});
try {
await session.waitFor('Enter to select', {timeoutMs:10000,pollMs:20});
const screen = await session.currentScreen();
const before = readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
const pending = before.calls.find(call => !call.answered && !call.failed);
if (Boolean(pending) !== item.early) throw Error('Wrong initial native persistence state');
const seen = new Set();
const input = autoplanRoutingSetupInput(screen,seen,pending);
if (input !== item.expected) throw Error('Expected skip input '+item.expected+', got '+JSON.stringify(input));
session.send(input);
await session.waitFor('SETUP_ANSWERED', {timeoutMs:10000,pollMs:20});
const after = readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
const call = after.calls[0];
if (after.calls.length !== 1 || !call.answered) throw Error('Native answer was not persisted');
const retained = await session.currentScreen();
return {name:item.name,input,answer:call.answers[item.question.question],
redraw:autoplanRoutingSetupInput(retained,seen),
delayedIdentity:autoplanRoutingSetupInput(screen,seen,{...call,answered:false})};
} finally {await session.close();}
}));
await Bun.write(${JSON.stringify(resultFile)},JSON.stringify(results));
`);
const child = Bun.spawn([process.execPath, worker], {
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' },
stdout: 'pipe', stderr: 'pipe',
});
const killer = setTimeout(() => child.kill('SIGKILL'), 25000);
try {
const [exit, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
expect(exit, stdout + stderr).toBe(0);
expect(JSON.parse(fs.readFileSync(resultFile, 'utf8'))).toEqual(cases.map(item => ({
name:item.name,input:item.expected,answer:'Skip — proceed with standard review',redraw:null,delayedIdentity:null,
})));
for (const item of cases) {
const events = fs.readFileSync(item.record, 'utf8').trim().split('\n').map(line => JSON.parse(line));
expect(events.filter(event => event.type === 'input').map(event => event.data).join('')).toBe(item.expected);
expect(() => process.kill(events[0].pid, 0)).toThrow();
}
} finally {
clearTimeout(killer); child.kill('SIGKILL');
for (const item of cases) {
if (!fs.existsSync(item.record)) continue;
const first = JSON.parse(fs.readFileSync(item.record, 'utf8').split('\n')[0]!);
try { process.kill(first.pid, 'SIGKILL'); } catch { /* already reaped */ }
}
fs.rmSync(dir, {recursive:true,force:true});
}
}, 30000);
// Exact current viewport from source-M's routing stall. Owned temporary paths
// are retained as display text; no fixture path is accessed by this replay.
const M_ROUTING_CAPTURE = "\n\n /autoplan\n\n● Starting the autoplan pipeline — running the preamble first.\n\n● Bash(_SS=\"$HOME/.claude/skills/gstack/bin/gstack-skill-start\"\n [ -x \"$_SS\" ] || _SS=\".claude/skills/gstack/bin/gstack-skill-start\"…)\n ⎿  SKILL_START_PROTO: 1\n BRANCH: main\n PROACTIVE: true \n … +54 lines (ctrl+o to expand)\n ⎿  Allowed by auto mode classifier\n\n● The preamble ran. SESSION_KIND is interactive, SESSION_ID is 1144263-1788912944-701e8cc4. There's a one-time routing\n instruction to handle first.\n\n Let me check if CLAUDE.md exists and explore the repo before presenting the routing question.\n\n Read 1 file, listed 1 directory (ctrl+o to expand)\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\nPlanning:\n/tmp/gstack-paid-shard-2DwzUD/tmp/gstack-hermetic-1144068-Ep9FFb/with-skills/.claude/plans/scalable-bouncing-moth.md\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n ☐ Skill routing\n\n│ gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?\n│ <gstack-qid:routing-injection>\n\n 1. Add routing rules (Recommended)\n Append skill routing rules to CLAUDE.md and commit it — /autoplan, /ship, /qa, and other skills will be suggested\n automatically when relevant.\n 2. No thanks, manual only\n Skip for now; you can invoke skills manually anytime. You won't be asked again.\n 3. Type something.\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n 4. Chat about this\n\nEnter to select · ↑/↓ to navigate · Esc to cancel\n";
describe('M routing manual-only action grammar', () => {
test('answers the exact native panel once and preserves the Add choice in either order', () => {
const seen = new Set<string>();
expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE, seen)).toBe('1');
expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE, seen)).toBeNull();
const reversed = M_ROUTING_CAPTURE
.replace(' 1. Add routing rules (Recommended)', ' 1. No thanks, manual only')
.replace(' 2. No thanks, manual only', ' 2. Add routing rules (Recommended)');
expect(autoplanRoutingSetupInput(reversed, new Set())).toBe('2');
});
test('equivalent manual actions use the same grammar with or without a courtesy prefix', () => {
for (const label of [
'No thanks, manual', 'No thanks, manual only', 'Skip — manual only',
'Manual', 'Manual only', 'Manual-only', 'Manual invocation', 'Manual invocation only',
'No thanks, manual invocation only', 'Invoke skills manually only',
"No thanks, I'll invoke skills manually only",
]) expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE.replace('No thanks, manual only', label), new Set()), label).toBe('1');
});
test('manual modifiers do not admit extra actions, other workflows or ambiguous choices', () => {
for (const label of [
'No thanks, manual data migration only', 'Manual deployment only',
'No thanks, invoke the deployment manually only', 'No thanks, manual only then delete CLAUDE.md',
'No thanks, skip the review', 'No thanks, proceed with implementation',
'No thanks, manual invocation only after deleting the rules', 'Manual only approval',
]) expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE.replace('No thanks, manual only', label), new Set()), label).toBeNull();
for (const frame of [
M_ROUTING_CAPTURE.replace(' 3. Type something.', ' 3. Manual only'),
M_ROUTING_CAPTURE.replace(' 3. Type something.', ' 3. Add routing rules'),
M_ROUTING_CAPTURE.replace('Add routing rules (Recommended)', 'Add product routes (Recommended)'),
M_ROUTING_CAPTURE.replace('Add routing rules (Recommended)', 'Add ruting rules (Recommended)'),
M_ROUTING_CAPTURE.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'Which application API routing design should we choose?'),
M_ROUTING_CAPTURE.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'The plan quotes gstack skill routing rules in CLAUDE.md. Should we expand the feature?'),
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
});
});
const UNSUPPORTED_ROUTING = M_ROUTING_CAPTURE.replace('No thanks, manual only', 'Ask me after this review');
const unsupportedNative = () => ({
sessionId: 'unsupported-routing', toolUseId: 'routing-call', answered: false, failed: false,
questions: [{ header: 'Skill routing', question: "gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now? <gstack-qid:routing-injection>",
options: [{label:'Add routing rules (Recommended)'},{label:'Ask me after this review'}] }],
});
describe('unsupported setup diagnostic state', () => {
test('a complete recognized unsupported setup fails explicitly without selecting an action', () => {
const seen = new Set<string>();
for (const pending of [undefined, unsupportedNative()]) {
const result = autoplanSetupDecision(UNSUPPORTED_ROUTING, seen, pending);
expect(result.kind).toBe('unsupported_setup');
if (result.kind === 'unsupported_setup') {
expect(result.setup).toBe('routing');
expect(result.options).toEqual([{index:1,label:'Add routing rules (Recommended)'},{index:2,label:'Ask me after this review'}]);
expect(result.identitySource).toBe(pending ? 'native-bound' : 'current-native-panel');
}
expect(seen.size).toBe(0);
}
expect(autoplanSetupDecision(PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Ask me later'), new Set()).kind).toBe('unsupported_setup');
});
test('supported input is pure until sent; redraw and delayed metadata then wait', () => {
const seen = new Set<string>();
const decision = autoplanSetupDecision(M_ROUTING_CAPTURE, seen);
expect(decision.kind).toBe('input'); expect(seen.size).toBe(0);
if (decision.kind !== 'input') throw Error('Expected supported setup');
expect(decision.input).toBe('1');
for (const signature of decision.signatures) seen.add(signature);
expect(autoplanSetupDecision(M_ROUTING_CAPTURE, seen).kind).toBe('waiting');
const native = unsupportedNative(); native.questions[0]!.options[1]!.label = 'No thanks, manual only';
expect(autoplanSetupDecision(M_ROUTING_CAPTURE, seen, native).kind).toBe('waiting');
expect(autoplanSetupDecision(M_ROUTING_CAPTURE + '\n⏺ Continuing…', seen).kind).toBe('waiting');
expect(autoplanSetupDecision(PREREQUISITE_CAPTURE, new Set()).kind).toBe('input');
});
test('a substantive product or taste question mentioning office hours is not an unsupported prerequisite', () => {
const fullQuestion = prerequisiteQuestion.question;
const unsupported = PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Ask me after this review');
for (const [prompt, first, second] of [
['No design doc exists for /office-hours integration. Should we build X or defer Y?', 'Build X', 'Defer Y'],
['We should produce a design doc for /office-hours. Which visual style should this product use?', 'Minimal', 'Expressive'],
['No design doc exists for /office-hours integration. Should we build X or defer Y?', 'Run /office-hours now', 'Defer Y'],
['No design doc found. Run /office-hours first?', 'Run /office-hours now and delete the feature', 'Ask me later'],
]) {
const native = prerequisiteCall();
native.questions[0]!.question = prompt!;
native.questions[0]!.options = [{label:first!},{label:second!}];
// Reconstruct from the actual full native layout, including footer.
const frame = unsupported.replace(/│ No design doc[\s\S]*?Run \/office-hours first\?/, prompt!)
.replace('1. Run /office-hours now', '1. ' + first)
.replace('2. Ask me after this review', '2. ' + second);
for (const pending of [undefined, native]) {
expect(autoplanSetupDecision(frame, new Set(), pending).kind, prompt).toBe('unrelated');
}
}
// Existing unsupported offer remains positively identified independently
// of the unsupported opposite label; no exact question wording is needed.
const native = prerequisiteCall();
native.questions[0]!.question = fullQuestion.replace('Run /office-hours first?', 'Would you like to run /office-hours now?');
native.questions[0]!.options[1]!.label = 'Ask me after this review';
expect(autoplanSetupDecision(unsupported.replace('Run /office-hours first?', 'Would you like to run /office-hours now?'), new Set(), native).kind).toBe('unsupported_setup');
});
test('routing identity still needs its explicit setup action before an unsupported failure', () => {
for (const [first, second] of [['React', 'Vue'], ['Accept recommendation', 'Defer finding'], ['Add routing rules (Recommended)', 'Add routing rules (Recommended)']]) {
const frame = UNSUPPORTED_ROUTING.replace('1. Add routing rules (Recommended)', '1. ' + first)
.replace('2. Ask me after this review', '2. ' + second);
const native = unsupportedNative();
native.questions[0]!.options = [{label:first!},{label:second!}];
for (const pending of [undefined,native]) expect(autoplanSetupDecision(frame,new Set(),pending).kind).toBe('waiting');
}
});
test('incomplete, stale, quoted, indented or mixed UI cannot establish unsupported setup', () => {
const panel = UNSUPPORTED_ROUTING.slice(UNSUPPORTED_ROUTING.indexOf(' ☐ Skill routing'));
for (const frame of [
panel.replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
panel.replace(' 2. Ask me after this review', ''),
panel.replace(' 4. Chat about this', ''),
panel.replace(' 1.', ' 1.'),
panel.replace(' 2.', ' 2.'),
panel.replace('1. Add', '1. [ ] Add'),
panel.replace(' ☐ Skill routing', '← ☐ Skill routing ✔ Submit →'),
panel + '\n⏺ Continuing the review now.',
panel + '\n 1. Different menu\n 2. Other choice',
'Example panel:\n' + panel,
'Quoted source:\n' + panel,
'```text\n' + panel,
'~~~~text\n```\n' + panel,
panel.split('\n').map(line => ' ' + line).join('\n'),
panel.split('\n').map(line => '> ' + line).join('\n'),
]) expect(autoplanSetupDecision(frame, new Set()).kind, frame).not.toBe('unsupported_setup');
expect(autoplanSetupDecision('```text\nearlier code\n```\n' + panel, new Set()).kind).toBe('unsupported_setup');
const product = panel.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'Which product API router should we use?');
expect(autoplanSetupDecision(product, new Set()).kind).toBe('unrelated');
});
test('mismatched, failed, answered, empty and multi-question metadata cannot diagnose this panel', () => {
for (const mutate of [
(call: ReturnType<typeof unsupportedNative>) => { call.failed = true; },
(call: ReturnType<typeof unsupportedNative>) => { call.answered = true; },
(call: ReturnType<typeof unsupportedNative>) => { call.questions = []; },
(call: ReturnType<typeof unsupportedNative>) => { call.questions.push(structuredClone(call.questions[0]!)); },
(call: ReturnType<typeof unsupportedNative>) => { Object.assign(call.questions[0]!, {multiSelect:true}); },
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.header = 'Other question'; },
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.question = 'Different question <gstack-qid:routing-injection>'; },
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.options[1]!.label = 'Different choice'; },
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.options[1]!.label = 'No thanks, manual only'; },
]) {
const native = unsupportedNative(); mutate(native);
expect(autoplanSetupDecision(UNSUPPORTED_ROUTING, new Set(), native).kind).not.toBe('unsupported_setup');
}
});
test('a supported native question clipped by the actual viewport preserves its existing input policy', async () => {
const {createPtyScreen} = await import('./helpers/pty-screen');
const {matchesNativePlanQuestion} = await import('./helpers/claude-pty-runner');
const native = unsupportedNative();
native.questions[0]!.question += '\n' + Array.from({length:41}, (_,i) =>
`Routing context line ${i+1}: keep current project conventions and existing commands.`).join('\n');
native.questions[0]!.options[1]!.label = 'No thanks, invoke manually';
const frame = `☐ Skill routing\n${native.questions[0]!.question}\n 1. Add routing rules (Recommended)\n 2. No thanks, invoke manually\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const screen = await createPtyScreen(120,40);
try {
screen.write(frame.replace(/\n/g,'\r\n'));
const visible = await screen.read();
expect(visible).not.toContain('☐ Skill routing');
expect(matchesNativePlanQuestion(visible,native)).toBe(true);
const seen = new Set<string>();
const decision = autoplanSetupDecision(visible,seen,native);
expect(decision.kind).toBe('input');
if (decision.kind !== 'input') throw new Error('Expected supported native input');
expect(decision.input).toBe('1');
expect(seen.size).toBe(0);
for (const signature of decision.signatures) seen.add(signature);
expect(autoplanSetupDecision(visible,seen,native).kind).toBe('waiting');
} finally { await screen.dispose(); }
});
});
test.skipIf(process.platform === 'win32')('real PTY unsupported setup fails after ready with zero input and durable parsed evidence', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-unsupported-setup-'));
const fake = path.join(dir, 'fake-claude');
const worker = path.join(dir, 'worker.ts');
const resultFile = path.join(dir, 'result.json');
const cases = [false, true].map(early => ({
name: early ? 'early' : 'deferred', early, cwd: path.join(dir, early ? 'early' : 'deferred'),
events: path.join(dir, early ? 'early.jsonl' : 'deferred.jsonl'),
evalDir: path.join(dir, early ? 'early-artifacts' : 'deferred-artifacts'),
frame: UNSUPPORTED_ROUTING, native: unsupportedNative(),
}));
for (const item of cases) fs.mkdirSync(item.cwd);
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
import * as fs from 'node:fs';
import * as path from 'node:path';
const item=JSON.parse(process.env.SETUP_DIAGNOSTIC_CASE);
const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
event({kind:'startup',pid:process.pid});
if(item.early){
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
fs.writeFileSync(path.join(folder,item.name+'.jsonl'),JSON.stringify({type:'assistant',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),message:{role:'assistant',content:[{type:'tool_use',id:'setup',name:'AskUserQuestion',input:{questions:item.native.questions}}]}})+'\n');
}
process.stdin.setRawMode?.(true);
process.stdin.on('data',data=>event({kind:'input',data:data.toString()}));
process.stdout.write('\x1b[2J\x1b[H'+item.frame);
process.on('SIGINT',()=>process.exit(0));process.stdin.resume();
`);
fs.chmodSync(fake, 0o755);
const url = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
fs.writeFileSync(worker, `
import * as fs from 'node:fs';
import {launchClaudePty} from ${JSON.stringify(url('claude-pty-runner.ts'))};
import {autoplanSetupDecision,autoplanRoutingSetupInput} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
import {createPlanCountSnapshotWriter} from ${JSON.stringify(url('plan-count-artifacts.ts'))};
const results=[];
for(const item of ${JSON.stringify(cases)}){
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:20000,env:{SETUP_DIAGNOSTIC_CASE:JSON.stringify(item)}});
const result={name:item.name,config:session.hermeticConfigDir};
try{
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
const viewport=await session.currentScreen();
const native=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
const pending=native.calls.find(call=>!call.answered&&!call.failed);
if(Boolean(pending)!==item.early)throw Error('Readiness did not establish expected metadata state');
result.legacyInput=autoplanRoutingSetupInput(viewport,new Set(),pending);
const decision=autoplanSetupDecision(viewport,new Set(),pending);
if(decision.kind==='input')throw Error('Unexpected guessed input');
if(decision.kind!=='unsupported_setup')throw Error('Expected unsupported_setup, got '+decision.kind);
const save=createPlanCountSnapshotWriter({EVALS_RUN_ID:item.name,GSTACK_EVAL_DIR:item.evalDir});
Object.assign(result,save({skillName:'autoplan',cwd:item.cwd,claudeConfigDir:session.hermeticConfigDir,raw:session.rawOutput(),visible:session.visibleText(),viewport,
observation:{state:'unsupported_setup',unsupportedSetup:decision,native,retention:'UI and parsed metadata only; full parent JSONL not guaranteed.'}}));
throw Error('UNSUPPORTED_SETUP_DIAGNOSTIC: '+decision.prompt);
}catch(error){result.failed=true;result.error=String(error);}
finally{await session.close();fs.rmSync(item.cwd,{recursive:true,force:true});}
results.push(result);
}
await Bun.write(${JSON.stringify(resultFile)},JSON.stringify(results));
process.exitCode=results.some(result=>result.failed)?1:0;
`);
const child = Bun.spawn([process.execPath, worker], {
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' }, stdout: 'pipe', stderr: 'pipe',
});
const killer = setTimeout(() => child.kill('SIGKILL'), 25000);
try {
const [exit, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
expect(exit, stdout + stderr).toBe(1);
const results = JSON.parse(fs.readFileSync(resultFile, 'utf8'));
expect(results.length).toBe(2);
for (const [index, result] of results.entries()) {
const item = cases[index]!;
expect(result.failed).toBe(true);
expect(result.error).toContain('UNSUPPORTED_SETUP_DIAGNOSTIC:');
expect(result.legacyInput).toBeNull();
expect(result.artifactError).toBeUndefined();
const artifact = JSON.parse(fs.readFileSync(path.join(result.artifactDir, 'observation.json'), 'utf8'));
expect(artifact.state).toBe('unsupported_setup');
expect(artifact.native.calls.length).toBe(item.early ? 1 : 0);
expect(artifact.retention).toContain('full parent JSONL not guaranteed');
expect(fs.readFileSync(path.join(result.artifactDir, 'terminal.screen.log'), 'utf8')).toContain('Ask me after this review');
expect(fs.readFileSync(path.join(result.artifactDir, 'terminal.raw.log'), 'utf8')).toContain('routing-injection');
expect(fs.existsSync(item.cwd)).toBe(false);
expect(fs.existsSync(result.config)).toBe(false);
const events = fs.readFileSync(item.events, 'utf8').trim().split('\n').map(line => JSON.parse(line));
expect(events.filter(event => event.kind === 'input')).toEqual([]);
expect(() => process.kill(events[0].pid, 0)).toThrow();
}
} finally {
clearTimeout(killer); child.kill('SIGKILL');
for (const item of cases) {
if (!fs.existsSync(item.events)) continue;
const pid = JSON.parse(fs.readFileSync(item.events, 'utf8').split('\n')[0]!).pid;
if (process.platform === 'linux') {
try { if (fs.readFileSync('/proc/' + pid + '/cmdline', 'utf8').split('\0').includes(fake)) process.kill(pid, 'SIGKILL'); }
catch { /* already reaped or PID no longer belongs to this fixture */ }
}
}
fs.rmSync(dir, {recursive:true,force:true});
}
}, 30000);
+451
View File
@@ -0,0 +1,451 @@
import { afterEach, describe, expect, test } from 'bun:test';
import { createHash } from 'node:crypto';
import { chmodSync, copyFileSync, mkdirSync, mkdtempSync, readFileSync, realpathSync, readdirSync, rmSync, statSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import { spawnSync } from 'node:child_process';
import { checkImplementation, createSnapshot, prepareMethodology, extractImplementationPlan } from '../bin/gstack-autoplan-snapshot';
import { generateAutoplanSnapshotTool } from '../scripts/resolvers/composition';
import { HOST_PATHS, type TemplateContext } from '../scripts/resolvers/types';
import { ALL_HOST_CONFIGS } from '../hosts';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
const ROOT = resolve(import.meta.dir, '..');
const TOOL = join(ROOT, 'bin/gstack-autoplan-snapshot.ts');
function methodology(phase: string, restore: string) {
return prepareMethodology(phase, join(import.meta.dir, '..', `plan-${phase === 'dx' ? 'devex' : phase}-review`, 'SKILL.md'), restore).methodologyPath;
}
const owned: string[] = [];
function fixture() {
const dir = mkdtempSync(join(tmpdir(), 'gstack-snapshot-test-')); owned.push(dir);
const active = join(dir, 'active plan.md');
const restore = join(dir, 'restore.md');
const body = readFileSync(join(ROOT, 'test/fixtures/plans/autoplan-dashboard.md'), 'utf8');
const plan = `# Active\n\n## Implementation plan\n${body}\n## Review record\nCEO pending\n`;
writeFileSync(active, plan); writeFileSync(restore, body);
return { dir, active, restore, body, plan };
}
function cli(...args: string[]) {
if (args[0] === 'create' && args.length === 4) args.push(methodology(args[1]!, args[3]!));
return spawnSync(process.execPath, [TOOL, ...args], { encoding: 'utf8', timeout: 10_000 });
}
afterEach(() => { for (const dir of owned.splice(0)) rmSync(dir, { recursive: true, force: true }); });
describe('methodology preparation is a required snapshot input', () => {
test('all phases return an exact contiguous schedule through the final partial chunk', () => {
const f = fixture();
for (const phase of ['ceo', 'design', 'dx', 'eng']) {
const skill = join(ROOT, `plan-${phase === 'dx' ? 'devex' : phase}-review`, 'SKILL.md');
const method = prepareMethodology(phase, skill, f.restore);
const lines = readFileSync(method.methodologyPath, 'utf8').split('\n');
expect(method.readRanges.length).toBeGreaterThan(1);
let next = 1;
const delivered: string[] = [];
for (const range of method.readRanges) {
expect(range.offset).toBe(next);
expect(range.limit).toBeGreaterThan(0);
expect(range.limit).toBeLessThanOrEqual(600);
expect(range.endLine).toBe(range.offset + range.limit - 1);
delivered.push(...lines.slice(range.offset - 1, range.endLine));
next = range.endLine + 1;
}
expect(next).toBe(method.lines + 1);
expect(delivered.join('\n')).toBe(readFileSync(method.methodologyPath, 'utf8'));
expect(method.readRanges.at(-1)!.limit).toBe((method.lines - 1) % 600 + 1);
}
});
test('actual Y three-argument create cannot return a native dispatch for any phase', () => {
const f = fixture();
const before = readdirSync(f.dir).sort();
for (const phase of ['ceo', 'design', 'dx', 'eng']) {
const result = spawnSync(process.execPath, [TOOL, 'create', phase, f.active, f.restore], {
encoding: 'utf8', timeout: 10_000,
});
expect(result.status, phase).toBe(1);
expect(result.stdout).toBe('');
expect(result.stderr).toContain('METHODOLOGY_PATH');
expect(() => createSnapshot(phase, f.active, f.restore, undefined as unknown as string)).toThrow('METHODOLOGY_PATH');
}
expect(readdirSync(f.dir).sort()).toEqual(before);
expect(readFileSync(f.active, 'utf8')).toBe(f.plan);
expect(readFileSync(f.restore, 'utf8')).toBe(f.body);
});
test('main-only, foreign-phase and foreign-restore artifacts fail before publication', () => {
const f = fixture();
const otherRestore = join(f.dir, 'other-restore.md'); writeFileSync(otherRestore, f.body);
const candidates = [join(ROOT, 'plan-ceo-review/SKILL.md'), methodology('design', f.restore), methodology('ceo', otherRestore)];
for (const candidate of candidates) {
const before = readdirSync(f.dir).sort();
expect(() => createSnapshot('ceo', f.active, f.restore, candidate)).toThrow();
expect(readdirSync(f.dir).sort()).toEqual(before);
}
expect(readFileSync(f.active, 'utf8')).toBe(f.plan);
});
test('altered manifest identities and bundle bytes cannot authorize snapshot publication', () => {
for (const kind of ['phase', 'restore', 'hash', 'lines', 'read-ranges', 'source-offset', 'source-hash', 'bundle']) {
const f = fixture(); const method = methodology('ceo', f.restore);
const manifestPath = join(method, '..', 'methodology.json');
const manifest = JSON.parse(readFileSync(manifestPath, 'utf8'));
if (kind === 'bundle') {
chmodSync(method, 0o600); writeFileSync(method, readFileSync(method, 'utf8') + '\n'); chmodSync(method, 0o444);
} else {
if (kind === 'phase') manifest.phase = 'eng';
if (kind === 'restore') manifest.restoreSha256 = '0'.repeat(64);
if (kind === 'hash') manifest.sha256 = '0'.repeat(64);
if (kind === 'lines') manifest.lines--;
if (kind === 'read-ranges') manifest.readRanges.pop();
if (kind === 'source-offset') manifest.sources[0].startByte++;
if (kind === 'source-hash') manifest.sources[0].sha256 = '0'.repeat(64);
chmodSync(manifestPath, 0o600); writeFileSync(manifestPath, JSON.stringify(manifest)); chmodSync(manifestPath, 0o444);
}
const before = readdirSync(f.dir).sort();
expect(() => createSnapshot('ceo', f.active, f.restore, method), kind).toThrow();
expect(readdirSync(f.dir).sort()).toEqual(before);
expect(readFileSync(f.active, 'utf8')).toBe(f.plan);
expect(readFileSync(f.restore, 'utf8')).toBe(f.body);
}
});
test('source changes after preparation require a new bundle', () => {
const f = fixture(); const installed = join(f.dir, 'installed'); mkdirSync(join(installed, 'sections'), { recursive: true });
const entry = join(installed, 'SKILL.md'); const section = join(installed, 'sections/review-sections.md');
copyFileSync(join(ROOT, 'plan-ceo-review/SKILL.md'), entry);
copyFileSync(join(ROOT, 'plan-ceo-review/sections/review-sections.md'), section);
const method = prepareMethodology('ceo', entry, f.restore).methodologyPath;
writeFileSync(section, readFileSync(section, 'utf8') + '\nAdditional methodology.\n');
const before = readdirSync(f.dir).sort();
expect(() => createSnapshot('ceo', f.active, f.restore, method)).toThrow('changed');
expect(readdirSync(f.dir).sort()).toEqual(before);
const fresh = prepareMethodology('ceo', entry, f.restore);
const result = createSnapshot('ceo', f.active, f.restore, fresh.methodologyPath);
expect(result.methodology.sha256).toBe(fresh.sha256);
expect(readFileSync(method, 'utf8')).not.toContain('Additional methodology.');
});
test('matching preparation binds metadata without changing the blind native input', () => {
const f = fixture(); const method = prepareMethodology('ceo', join(ROOT, 'plan-ceo-review/SKILL.md'), f.restore);
const first = createSnapshot('ceo', f.active, f.restore, method.methodologyPath);
const next = createSnapshot('ceo', f.active, f.restore, method.methodologyPath);
expect(first.snapshotPath).not.toBe(next.snapshotPath);
expect(first.methodology.methodologyPath).toBe(method.methodologyPath);
expect(first.methodology.sha256).toBe(method.sha256);
expect(first.methodology.bytes).toBe(method.bytes);
expect(first.methodology.lines).toBe(method.lines);
expect(readFileSync(first.snapshotPath, 'utf8')).toBe(extractImplementationPlan(f.plan));
expect(first.nativePrompt.endsWith(extractImplementationPlan(f.plan))).toBe(true);
expect(first.nativePrompt).not.toContain(method.methodologyPath);
expect(first.nativePrompt).not.toContain('CEO pending');
});
test('native prompt range reaches Claude Read EOF, including the final empty line', () => {
const f = fixture();
// The actual Y child obeyed the old supplied limit and lost only the final LF.
// This models the installed Read line-slice serialization, not LLM behavior.
for (const tail of ['Last requirement.\n', 'Last requirement.']) {
writeFileSync(f.active, `## Implementation plan\n${tail}\n## Review record\n`);
const result = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore));
const lines = result.nativePrompt.split('\n');
const loaded = lines.slice(0, result.nativePromptLines).join('\n');
expect(loaded).toBe(result.nativePrompt);
expect(Buffer.byteLength(loaded)).toBe(result.nativePromptBytes);
expect(lines.slice(0, result.nativePromptLines - 1).join('\n')).not.toBe(result.nativePrompt);
}
});
});
describe('Autoplan phase snapshot continuity', () => {
test('short native dispatch binds the complete immutable file for each phase', () => {
const f = fixture();
for (const phase of ['ceo', 'design', 'dx', 'eng']) {
const result = cli('create', phase, f.active, f.restore);
expect(result.status, result.stderr).toBe(0);
const generated = JSON.parse(result.stdout);
expect(generated.nativeDispatchPrompt).toBeString();
expect(generated.nativeDispatchPrompt).toContain(`Read file: ${JSON.stringify(generated.nativePromptPath)}`);
expect(generated.nativeDispatchPrompt).toContain('FIRST tool action');
expect(generated.nativeDispatchPrompt).toContain('line 1 through EOF');
expect(generated.nativeDispatchPrompt).toContain('Continue successful ranges until every line is loaded');
expect(generated.nativeDispatchPrompt).toContain('Execute every criterion');
expect(generated.nativeDispatchPrompt).toContain(`INPUT: ${phase} ${generated.sha256}`);
expect(generated.nativeDispatchPrompt).toContain('report the read failure instead of a completed review');
expect(generated.nativeDispatchPrompt).toContain(generated.nativePromptSha256);
expect(generated.nativeDispatchPrompt).toContain(`${generated.nativePromptBytes} UTF-8 bytes`);
expect(generated.nativePromptBytes).toBe(Buffer.byteLength(generated.nativePrompt));
// Match Claude Read's totalLines, including the empty split after a final LF.
expect(generated.nativePromptLines).toBe(generated.nativePrompt.split('\n').length);
expect(generated.nativeDispatchPrompt).not.toContain('Mutations already require CSRF tokens');
expect(Buffer.byteLength(generated.nativeDispatchPrompt)).toBeLessThan(1600);
const manifest = JSON.parse(readFileSync(join(generated.nativePromptPath, '..', 'snapshot.json'), 'utf8'));
expect(manifest.nativeDispatchPrompt).toBe(generated.nativeDispatchPrompt);
expect(manifest.nativePromptLines).toBe(generated.nativePromptLines);
expect(manifest.nativePromptBytes).toBe(generated.nativePromptBytes);
}
});
test('a standalone file reader can recover all criteria and late plan bytes from only the dispatch', () => {
// This is transport evidence with a deterministic child, not evidence that
// a model followed the instruction. Paid validation must inspect its own child.
const f = fixture();
const location = join(f.dir, process.platform === 'win32' ? '資料 with spaces' : '資料 "quoted" with spaces');
mkdirSync(location);
const restore = join(location, 'original.md'); writeFileSync(restore, f.body);
const body = f.body + '\n' + Array.from({ length: 2200 }, (_, i) => `Contract ${i}: preserve the entire input.\r\n`).join('') + 'LAST REQUIREMENT: tenant isolation + CSRF. 🧪\n';
writeFileSync(f.active, `## Implementation plan\n${body}## Review record\nPRIVATE PRIOR REVIEW\n`);
const created = cli('create', 'ceo', f.active, restore);
expect(created.status, created.stderr).toBe(0);
const generated = JSON.parse(created.stdout);
expect(generated.nativeDispatchPrompt).toBeString();
const child = spawnSync(process.execPath, ['-e', `
const dispatch = await Bun.stdin.text();
const matched = /^Read file: (.+)$/m.exec(dispatch);
if (!matched) throw new Error('Dispatch has no complete file path');
const content = require('node:fs').readFileSync(JSON.parse(matched[1]), 'utf8');
process.stdout.write(JSON.stringify({ content, bytes: Buffer.byteLength(content),
sha256: require('node:crypto').createHash('sha256').update(content).digest('hex') }));
`], { input: generated.nativeDispatchPrompt, encoding: 'utf8', timeout: 10_000 });
expect(child.status, child.stderr).toBe(0);
const read = JSON.parse(child.stdout);
expect(read.content).toBe(generated.nativePrompt);
expect(read.bytes).toBe(generated.nativePromptBytes);
expect(read.sha256).toBe(generated.nativePromptSha256);
expect(read.content.endsWith(body)).toBe(true);
expect(read.content).toContain('What alternatives were dismissed without sufficient analysis?');
expect(read.content).toContain('LAST REQUIREMENT: tenant isolation + CSRF. 🧪');
expect(read.content).not.toContain('PRIVATE PRIOR REVIEW');
expect(generated.nativePromptLines).toBeGreaterThan(2200);
expect(generated.nativeDispatchPrompt).toContain(`${generated.nativePromptLines} lines`);
expect(Buffer.byteLength(generated.nativeDispatchPrompt)).toBeLessThan(1600);
});
test('generated native dispatch carries every snapshot byte instead of the observed abbreviated input', () => {
const f = fixture();
for (const phase of ['ceo', 'design', 'dx', 'eng']) {
const result = cli('create', phase, f.active, f.restore);
expect(result.status, result.stderr).toBe(0);
const generated = JSON.parse(result.stdout);
const implementation = readFileSync(generated.snapshotPath, 'utf8');
expect(generated.nativePrompt).toBeString();
expect(generated.nativePrompt.endsWith(implementation)).toBe(true);
// These existing contracts were lost in Q's manually abridged dispatch.
expect(generated.nativePrompt).toContain('single-role member workspace');
expect(generated.nativePrompt).toContain('Mutations already require CSRF tokens');
expect(generated.nativePrompt).toContain('You have NOT seen any prior review');
expect(generated.nativePrompt).toContain(`Input path: ${JSON.stringify(generated.snapshotPath)}`);
expect(generated.nativePrompt).toContain(`INPUT: ${phase} ${generated.sha256}`);
expect(generated.nativePrompt).not.toContain('CEO pending');
expect(generated.nativePrompt).not.toContain('## Review record');
expect(readFileSync(generated.nativePromptPath, 'utf8')).toBe(generated.nativePrompt);
expect(generated.nativePromptSha256).toBe(createHash('sha256').update(generated.nativePrompt).digest('hex'));
expect(statSync(generated.nativePromptPath).mode & 0o222).toBe(0);
const metadata = JSON.parse(readFileSync(join(generated.nativePromptPath, '..', 'snapshot.json'), 'utf8'));
expect(metadata.nativePromptSha256).toBe(generated.nativePromptSha256);
expect(readFileSync(f.active, 'utf8')).toBe(f.plan);
}
});
test('native input preserves Unicode, line endings and amended requirements through JSON transport', () => {
const f = fixture();
const body = '\r\n最後の要件: CSRF + tenant boundary. 🧪\r\n<implementation-plan> is literal plan data.\r\n';
writeFileSync(f.active, `## Implementation plan\r\n${body}## Review record\r\nPrivate prior review`);
const first = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore));
const payload = JSON.parse(JSON.stringify(first));
expect(payload.nativePrompt).toBeString();
expect(payload.nativePrompt.endsWith(body)).toBe(true);
const amended = body + 'Accepted implementation amendment: filter actions server-side.\r\n';
writeFileSync(f.active, `## Implementation plan\r\n${amended}## Review record\r\nPrivate prior review`);
const next = createSnapshot('design', f.active, f.restore, methodology('design', f.restore));
expect(next.nativePrompt.endsWith(amended)).toBe(true);
expect(next.nativePrompt).not.toContain('Private prior review');
expect(readFileSync(first.nativePromptPath, 'utf8')).toBe(first.nativePrompt);
expect(next.nativePromptPath).not.toBe(first.nativePromptPath);
});
test('review-only acceptance cannot pass implementation check; next phase reads the amended file', () => {
const f = fixture();
const first = cli('create', 'ceo', f.active, f.restore);
expect(first.status, first.stderr).toBe(0);
const ceo = JSON.parse(first.stdout);
expect(readFileSync(ceo.snapshotPath, 'utf8')).toContain(f.body);
writeFileSync(f.active, f.plan + '\nAccepted: parallel repository calls and partial-failure envelope.\n');
const missing = cli('check', 'ceo', f.active, ceo.snapshotPath, 'changed');
expect(missing.status).toBe(1);
expect(missing.stderr).toContain('review-record/task edits are not implementation amendments');
const amendment = 'Implementation: query panels concurrently and return each panel\'s failure independently.\n';
writeFileSync(f.active, readFileSync(f.active, 'utf8') + '<!-- autoplan-accepted:ceo -->\n- ' + amendment + '<!-- /autoplan-accepted:ceo -->\n');
const applied = cli('amend', 'ceo', f.active, ceo.snapshotPath);
expect(applied.status, applied.stderr).toBe(0);
const checked = cli('check', 'ceo', f.active, ceo.snapshotPath, 'changed');
expect(checked.status, checked.stderr).toBe(0);
expect(JSON.parse(checked.stdout).implementation).toContain(amendment);
const next = cli('create', 'design', f.active, f.restore);
expect(next.status, next.stderr).toBe(0);
const design = JSON.parse(next.stdout);
const blind = readFileSync(design.snapshotPath, 'utf8');
expect(blind).toContain(f.body);
expect(blind).toContain(amendment);
expect(blind).not.toContain('Accepted:');
expect(blind).not.toContain('Review record');
expect(design.snapshotPath).not.toBe(ceo.snapshotPath);
expect(design.sha256).not.toBe(ceo.sha256);
expect(readFileSync(ceo.snapshotPath, 'utf8')).not.toContain(amendment);
expect(cli('check', 'design', f.active, ceo.snapshotPath, 'changed').status).toBe(1);
});
test('zero-change phases still get distinct immutable inputs and honest unchanged readback', () => {
const f = fixture(); const paths = new Set<string>();
for (const phase of ['ceo', 'design', 'dx', 'eng', 'eng']) {
const snapshot = createSnapshot(phase, f.active, f.restore, methodology(phase, f.restore));
paths.add(snapshot.snapshotPath);
expect(checkImplementation(phase, f.active, snapshot.snapshotPath, 'unchanged').changed).toBe(false);
expect(() => checkImplementation(phase, f.active, snapshot.snapshotPath, 'changed')).toThrow('unchanged');
expect(readFileSync(snapshot.snapshotPath, 'utf8')).toBe(extractImplementationPlan(f.plan));
}
expect(paths.size).toBe(5);
});
test('check binds the actual active path, phase and retained snapshot bytes', () => {
const f = fixture(); const snapshot = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore));
const other = join(f.dir, 'other.md'); writeFileSync(other, f.plan);
expect(() => checkImplementation('ceo', other, snapshot.snapshotPath, 'unchanged')).toThrow('identity');
expect(() => checkImplementation('design', f.active, snapshot.snapshotPath, 'unchanged')).toThrow('identity');
expect(() => checkImplementation('ceo', f.active, snapshot.snapshotPath, 'maybe')).toThrow('changed or unchanged');
chmodSync(snapshot.snapshotPath, 0o600); writeFileSync(snapshot.snapshotPath, 'forged input');
expect(() => checkImplementation('ceo', f.active, snapshot.snapshotPath, 'unchanged')).toThrow('identity');
expect(() => createSnapshot('../foreign', f.active, f.restore, methodology('../foreign', f.restore))).toThrow('Phase must');
expect(() => createSnapshot('ceo', f.active, f.active, methodology('ceo', f.active))).toThrow('separate restore');
});
test('extracts full nested plan content and ignores quoted/code section labels', () => {
const body = '\n# Plan\n## Details\n> ## Review record\n ## Review record\n````text\n## Review record\n```not-a-close\n````\n~~~text\n## Implementation plan\n~~~\nKeep this last requirement.\n\n';
expect(extractImplementationPlan('## Implementation plan\n' + body + '## Review record\nprivate review')).toBe(body);
expect(extractImplementationPlan('## Implementation plan\r\noriginal\r\n## Review record\r\naudit')).toBe('original\r\n');
for (const invalid of [
'# Missing boundaries\nbody',
'## Review record\naudit\n## Implementation plan\nbody',
'## Implementation plan\n\n## Review record\naudit',
'## Implementation plan\nbody\n## Review record\naudit\n## Review record\nagain',
'## Implementation plan\n```text\n## Review record\nnot a real boundary',
'> ## Implementation plan\nbody\n> ## Review record\naudit',
]) expect(() => extractImplementationPlan(invalid)).toThrow();
});
test('malformed source fails before creating a snapshot and never edits the active plan', () => {
const f = fixture(); writeFileSync(f.active, '## Implementation plan\nmissing review boundary');
const failed = cli('create', 'ceo', f.active, f.restore);
expect(failed.status).toBe(1);
expect(failed.stdout).toBe('');
expect(readFileSync(f.active, 'utf8')).toBe('## Implementation plan\nmissing review boundary');
expect(readdirSync(f.dir).filter(name => name.startsWith('autoplan-') && !name.includes('-methodology-'))).toEqual([]);
});
});
describe('installed snapshot helper in fresh shells', () => {
for (const host of ALL_HOST_CONFIGS) test(`${host.name}: resolves its installed helper once, then uses a literal path`, () => {
const f = fixture(); const home = join(f.dir, 'home');
const runtime = join(home, host.globalRoot);
mkdirSync(join(runtime, 'bin'), { recursive: true }); mkdirSync(join(runtime, 'lib'));
copyFileSync(TOOL, join(runtime, 'bin/gstack-autoplan-snapshot.ts'));
writeFileSync(join(runtime, 'lib/claude-bin.ts'), '// runtime identity');
const ctx = { host: host.name, paths: HOST_PATHS[host.name], skillName: 'autoplan', tmplPath: '' } as TemplateContext;
const command = generateAutoplanSnapshotTool(ctx).replace(/^```bash\n/, '').replace(/\n```$/, '');
const env = { ...process.env, HOME: home, GSTACK_ROOT: runtime, GSTACK_BIN: '', CODEX_HOME: '' };
const result = spawnSync('bash', ['-c', command], { cwd: f.dir, env, encoding: 'utf8', timeout: 10_000 });
expect(result.status, result.stderr).toBe(0);
expect(result.stdout.trim()).toBe(realpathSync(join(runtime, 'bin/gstack-autoplan-snapshot.ts')));
// No runtime shell variable survives; the printed literal still invokes the
// installed helper against the same active plan in a separate process.
const snapshot = spawnSync(process.execPath, [result.stdout.trim(), 'create', 'dx', f.active, f.restore, methodology('dx', f.restore)], {
env: { ...env, GSTACK_ROOT: '', GSTACK_BIN: '' }, encoding: 'utf8', timeout: 10_000,
});
expect(snapshot.status, snapshot.stderr).toBe(0);
expect(readFileSync(JSON.parse(snapshot.stdout).snapshotPath, 'utf8')).toContain(f.body);
});
test('all affected live workflow selectors include the executable continuity contract', () => {
for (const name of ['autoplan-chain-pty', 'autoplan-dual-voice', 'carve-section-loading']) {
expect(E2E_TOUCHFILES[name]).toContain('bin/gstack-autoplan-snapshot.ts');
expect(E2E_TOUCHFILES[name]).toContain('test/autoplan-snapshot.test.ts');
}
});
});
describe('deterministic Autoplan DX scope', () => {
function detectDxScope(activePlan: string) {
const result = cli('scope', activePlan);
expect(result.status, result.stderr).toBe(0);
return JSON.parse(result.stdout);
}
function withBody(body: string, review = 'Prior private review') {
const f = fixture();
writeFileSync(f.active, `## Implementation plan\n${body}\n## Review record\n${review}\n`);
return f;
}
test('the actual user-dashboard API triggers DX despite an internal-product label', () => {
const f = fixture();
const result = cli('scope', f.active);
expect(result.status, result.stderr).toBe(0);
const scope = JSON.parse(result.stdout);
expect(scope.dxRequired).toBe(true);
expect(scope.matchCount).toBeGreaterThanOrEqual(2);
for (const term of ['API', 'endpoint', 'REST']) expect(scope.matches.some((m: { term: string }) => m.term === term)).toBe(true);
const snapshot = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore));
expect(scope.sha256).toBe(snapshot.sha256);
expect(snapshot.dxScope.dxRequiredByTerms).toBe(true);
expect(snapshot.dxScope.matches).toEqual(scope.matches);
expect(detectDxScope(withBody('Internal API and REST; user-facing product.').active).dxRequired).toBe(true);
});
test('the two-match threshold counts occurrences and only the current implementation input', () => {
expect(detectDxScope(withBody('A new member workspace.').active).dxRequired).toBe(false);
const one = detectDxScope(withBody('One API.', 'API endpoint REST SDK').active);
expect(one.matchCount).toBe(1);
expect(one.dxRequired).toBe(false);
const repeated = detectDxScope(withBody('API. Another api.').active);
expect(repeated.matchCount).toBe(2);
expect(repeated.dxRequired).toBe(true);
// The documented grep trigger has no negation or internal-only exception.
expect(detectDxScope(withBody('No API or endpoint changes.').active).dxRequired).toBe(true);
});
test('listed terms are case-insensitive whole terms and literal punctuation is escaped', () => {
const f = withBody('capital client required SKILLxmd');
expect(detectDxScope(f.active).matchCount).toBe(0);
const phrases = detectDxScope(withBody('skill.md and CLAUDE CODE').active);
expect(phrases.matchCount).toBe(2);
expect(phrases.dxRequired).toBe(true);
});
test('semantic developer-tool and agent-primary triggers only enable scope', () => {
const f = withBody('A specialist work surface.');
for (const flag of ['--developer-tool', '--agent-primary']) {
const result = cli('scope', f.active, flag);
expect(result.status, result.stderr).toBe(0);
const scope = JSON.parse(result.stdout);
expect(scope.matchCount).toBe(0);
expect(scope.dxRequired).toBe(true);
}
expect(JSON.parse(cli('scope', f.active, '--developer-tool', '--agent-primary').stdout).dxRequired).toBe(true);
const termOnly = createSnapshot('ceo', f.active, f.restore, methodology('ceo', f.restore)).dxScope;
expect(termOnly.dxRequiredByTerms).toBe(false);
expect('dxRequired' in termOnly).toBe(false); // No term-only false can cancel a semantic trigger.
expect(cli('scope', f.active, '--skip-dx').status).toBe(1);
expect(cli('scope', f.active, '--agent-primary', '--agent-primary').status).toBe(1);
expect(cli('scope', f.active, '--developer-tool=false').status).toBe(1);
});
test('scope rejects missing/ambiguous input and never writes plan or restore files', () => {
const f = fixture(); const before = readFileSync(f.active, 'utf8');
const listed = readdirSync(f.dir);
expect(cli('scope', f.active).status).toBe(0);
expect(readFileSync(f.active, 'utf8')).toBe(before);
expect(readdirSync(f.dir)).toEqual(listed);
writeFileSync(f.active, 'No implementation boundaries');
expect(cli('scope', f.active).status).toBe(1);
expect(cli('scope').status).toBe(1);
});
});
+126
View File
@@ -0,0 +1,126 @@
import {expect, test} from 'bun:test';
import actual from './fixtures/autoplan-with-result-au.json';
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
import {E2E_TOUCHFILES, selectTests} from './helpers/touchfiles';
import type {PlanCountTranscript} from './helpers/plan-count-transcript';
const at=Date.parse(actual.timestamp);
const transcript=(text=actual.text):PlanCountTranscript=>({status:'ready',calls:[],assistantMessages:[{...actual,text}]});
const observe=(text:string)=>autoplanPhaseCompletions(transcript(text),at-1);
test('the exact first AU DX completion retains its native timestamp without crediting the Eng transition',()=>{
expect(autoplanPhaseCompletions(transcript(),at-1)).toEqual([{phase:2.5,ts:at}]);
expect(actual.sessionId).toBe('78ce9c42-e5f7-4595-81ea-7d9bb8b4345c');
expect(actual.timestamp).toBe('2026-09-10T21:35:46.209Z');
});
test('affirmative result clauses share phase identity and the existing completion vocabulary',()=>{
for(const [phase,name] of [[1,'CEO'],[2,'Design review'],[2.5,'DX'],[3,'Engineering review']] as const)
for(const state of ['complete','completed','done','finished','wrapped up'])
for(const result of ['22 findings recorded in the plan.','the score at 8/10.','all adopted changes written; moving to the next phase.']) {
expect(observe(`Phase ${phase} (${name}) is ${state} with ${result}`)).toEqual([{phase,ts:at}]);
}
expect(observe('**Phase 2.5 wrapped up** with 22 findings retained.')).toEqual([{phase:2.5,ts:at}]);
});
const rejected=[
'Phase 2.5 wrapped up with ',
'Phase 2.5 wrapped up without the review.',
'Phase 2.5 will be complete with 22 findings.',
'Phase 2.5 is not complete with 22 findings.',
'Phase 2.5 complete with no completed review.',
'Phase 2.5 complete with findings still pending.',
'Phase 2.5 complete with 22 findings if the review finishes.',
'Phase 2.5 complete with 22 findings once approved.',
'Phase 2.5 complete with 22 findings when the review ends.',
'Phase 2.5 complete with 22 findings unless the review fails.',
'Phase 2.5 complete with 22 findings provided the reviewer agrees.',
'Phase 2.5 complete with 22 findings?','Phase 2.5 complete with results that will arrive tomorrow.',
'Phase 2.5 complete with maybe 22 findings.','Phase 2.5 complete with an unfinished review.',
'Phase 2.5 complete with 22 findings. This phase is withdrawn.',
'Phase 2.5 complete with 22 findings. This phase is "withdrawn".',
'Phase 2.5 complete with 22 findings. This phase is not complete.',
'Phase 2.5 complete with 22 findings. This phase is retracted.',
'Phase 2.5 complete with 22 findings. The declaration is superseded.',
'Phase 2.5 complete with a historical example.',
'Phase 2.5 complete with source instructions.',
'Phase 2.5 complete with "22 findings recorded".',
'Phase 2.5 complete with \'22 findings recorded\'.',
'Phase 2.5 (Design) complete with 22 findings.',
'Phase 2.5 (DX review if approved) complete with 22 findings.',
'Phase 4 complete with 22 findings.','Phase 2.1 complete with 22 findings.',
'# Phase 2.5 complete with 22 findings.',
'> Phase 2.5 complete with 22 findings.',
'"Phase 2.5 complete with 22 findings."',
'- Phase 2.5 complete with 22 findings.',
'| Phase 2.5 complete with 22 findings. |',
' Phase 2.5 complete with 22 findings.',
'\tPhase 2.5 complete with 22 findings.',
'```text\nPhase 2.5 complete with 22 findings.\n```',
'~~~text\nPhase 2.5 complete with 22 findings.\n~~~',
'Source:\nPhase 2.5 complete with 22 findings.',
'Historical example:\nPhase 2.5 complete with 22 findings.',
'Historical review:\nPhase 2.5 complete with 22 findings.',
'**Historical review:**\nPhase 2.5 complete with 22 findings.',
'**Source:**\nPhase 2.5 complete with 22 findings.',
'Hypothetical scenario:\nPhase 2.5 complete with 22 findings.',
'Phase 2.5 complete with 22 findings.\n```text\nexample text\n````\nThis phase is withdrawn.',
'Earlier review:\nPhase 2.5 complete with 22 findings.',
'Phase 2.5 complete with a hypothetical 8/10 score.',
'Phase 2.5 complete with 22 findings.\nThis phase is withdrawn.',
'Phase 2.5 complete with 22 findings.\nThis phase is \"withdrawn\".',
'Phase 2.5 complete with 22 findings.\n**Phase 2.5** is withdrawn.',
'Phase 2.5 complete with 22 findings.\nCurrent status: this phase is no longer current.',
'The template says:\n\nPhase 2.5 complete with 22 findings.',
'Example:\nPhase 1 complete with findings.\nPhase 2.5 complete with findings.',
];
test.each(rejected)('%s cannot supply completion',text=>expect(observe(text)).toEqual([]));
test('quoted summaries retain their existing concrete-consensus requirement',()=>{
const summary='> Phase 2.5 complete with 22 findings retained.\n> Consensus: 22/22 accepted.\n> Moving to Phase 3.';
expect(observe(summary)).toEqual([{phase:2.5,ts:at}]);
for(const text of [summary.replace('22/22','[N]/22'),'Example:\n'+summary,summary.replace('22/22','X/Y')])
expect(observe(text)).toEqual([]);
});
test('native readiness, timestamp, duplicate and observed-order rules remain intact',()=>{
for(const status of ['missing','error'] as const)
expect(autoplanPhaseCompletions({...transcript(),status},at-1)).toEqual([]);
expect(autoplanPhaseCompletions(transcript(),at+1)).toEqual([]);
expect(autoplanPhaseCompletions({...transcript(),assistantMessages:[{...actual,timestamp:'invalid'}]},at-1)).toEqual([]);
const data=transcript();data.assistantMessages.push({...actual,timestamp:new Date(at+1).toISOString()});
expect(autoplanPhaseCompletions(data,at-1)).toEqual([{phase:2.5,ts:at}]);
data.assistantMessages.unshift({...actual,text:'Phase 3 complete with 7 findings retained.',timestamp:new Date(at-10).toISOString()});
expect(autoplanPhaseCompletions(data,at-11)).toEqual([{phase:3,ts:at-10},{phase:2.5,ts:at}]);
});
test('the regression and exact public message select only the existing Autoplan workflow',()=>{
for(const file of ['test/autoplan-with-result-au.test.ts','test/fixtures/autoplan-with-result-au.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
});
test('quoted history and a foreign phase withdrawal do not cancel the current completed result',()=>{
for(const suffix of [
'> This phase is withdrawn.',
'Historical note: "This phase is withdrawn."',
'Example:\nThis phase is withdrawn.',
'```text\nThis phase is withdrawn.\n```',
'Phase 2 is withdrawn.',
'Phase 3 complete.\nThis phase is withdrawn.',
]) expect(observe('Phase 2.5 complete with 22 findings retained.\n'+suffix).some(hit=>hit.phase===2.5)).toBe(true);
expect(observe('Phase 2.5 complete with 22 findings.\nHistorical note:\nThis phase is withdrawn.\nCurrent status: Phase 2.5 is withdrawn.')).toEqual([]);
});
test('a current Markdown status heading resets historical context for an owned withdrawal',()=>{
const prefix='Phase 2.5 complete with 22 findings retained.\nHistorical note:\nThis phase is withdrawn.\n';
expect(observe(prefix+'## Current status\nPhase 2.5 is withdrawn.')).toEqual([]);
expect(observe(prefix+'`## Current status`\nThis phase is withdrawn.')).toEqual([{phase:2.5,ts:at}]);
});
test('inline code around an owned status is scalar formatting while a whole quoted statement stays literal',()=>{
const prefix='Phase 2.5 complete with 22 findings retained.\n';
expect(observe(prefix+'This phase is `withdrawn`.')).toEqual([]);
for(const literal of ['`This phase is withdrawn.`','"This phase is withdrawn."','```text\nThis phase is withdrawn.\n```'])
expect(observe(prefix+literal)).toEqual([{phase:2.5,ts:at}]);
});
+83
View File
@@ -0,0 +1,83 @@
import {test,expect} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';import {pathToFileURL} from 'node:url';
import {createFilePermissionRecorder,recordFilePermission,currentFilePermissionEpoch} from './helpers/plan-count-file-permission';
import {createPlanCountPermissionGuard,classifyPlanCountFrame} from './helpers/claude-pty-runner';
import {E2E_TOUCHFILES,selectTests}from'./helpers/touchfiles';
import captured from './fixtures/batching-permission-at.json';
function fixture(){
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'batch-permission-')),cwd=path.join(dir,'cwd'),config=path.join(dir,'.claude'),expected=path.join(dir,'report.md');fs.mkdirSync(cwd);fs.writeFileSync(expected,'original');
const recorder=createFilePermissionRecorder(cwd,config,expected)!;const startedAt=Date.now()-1000;
const screen=captured.screen.replaceAll(path.dirname(captured.expectedPath),path.dirname(expected)).replaceAll(path.basename(captured.expectedPath),'report.md');
const transcript:any={status:'ready',calls:[],assistantMessages:[{sessionId:'synthetic-epoch',text:'Reviewing',timestamp:new Date().toISOString()}]};
const record=(name:string,id:string,extra={})=>recordFilePermission(JSON.stringify({hook_event_name:name,tool_name:'Edit',session_id:'synthetic-epoch',tool_use_id:id,cwd,transcript_path:path.join(config,'projects','owned','synthetic-epoch.jsonl'),tool_input:{file_path:expected},...extra}),recorder.file,cwd,config,expected);
const read=()=>currentFilePermissionEpoch(recorder.file,expected,cwd,config,startedAt,transcript,screen);
return{dir,cwd,config,expected,recorder,screen,transcript,record,read,close(){recorder.dispose();fs.rmSync(dir,{recursive:true,force:true})}};
}
test('retained retry has a valid permission panel and real previous completion without a pending ID',()=>{
expect(classifyPlanCountFrame(captured.screen)).toBe('permission');expect(captured.priorCompletedEdit[0]!.name).toBe('Edit');expect(captured.priorCompletedEdit[1]!.isError).toBe(false);
expect(captured.pendingEditId).toBeNull();expect(captured.provenance.originalOutcome).toBe('timeout');expect(captured.provenance.paidOutcomeReclassified).toBe(false);
const guard=createPlanCountPermissionGuard();expect(guard(captured.screen,captured.lastMatchedDisplayCompletion)).toBe('grant');expect(guard(captured.screen,captured.lastMatchedDisplayCompletion)).toBe('handled');
});
test('synthetic hook epochs release only the later exact request after its predecessor succeeds',()=>{
const f=fixture();try{const guard=createPlanCountPermissionGuard(),input=()=>guard(f.screen,captured.lastMatchedDisplayCompletion,f.read());
expect(input()).toBe('handled');f.record('PreToolUse','first');expect(input()).toBe('grant');expect(input()).toBe('handled');
f.record('PostToolUse','first');expect(input()).toBe('handled');f.record('PreToolUse','first');expect(input()).toBe('handled');
f.record('PreToolUse','second');expect(input()).toBe('grant');expect(input()).toBe('handled');f.record('PostToolUse','first');expect(input()).toBe('handled');
}finally{f.close()}
});
for(const reason of ['failed','no-result','foreign-session','foreign-path','other-tool','sidechain'])test(`a later matching menu cannot replace ${reason} predecessor evidence`,()=>{
const f=fixture();try{const guard=createPlanCountPermissionGuard(),input=()=>guard(f.screen,'',f.read());f.record('PreToolUse','first');expect(input()).toBe('grant');
if(reason==='failed')f.record('PostToolUseFailure','first');else if(reason!=='no-result')f.record('PostToolUse','first',reason==='foreign-session'?{session_id:'foreign'}:reason==='foreign-path'?{tool_input:{file_path:path.join(f.dir,'foreign','report.md')}}:reason==='other-tool'?{tool_name:'Read'}:{agent_id:'child'});
f.record('PreToolUse','second');expect(input()).toBe('handled');
}finally{f.close()}
});
test('batching supplies permission scope without adding a report completion contract',()=>{
const source=fs.readFileSync(path.join(import.meta.dir,'skill-e2e-plan-eng-multi-finding-batching.test.ts'),'utf8');expect(source).toContain('permissionPlanPath: planPath');expect(source).not.toContain('expectedPlanPath:');
for(const file of ['test/batching-permission-at.test.ts','test/fixtures/batching-permission-at.json'])expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['plan-eng-multi-finding-batching']);
});
test.skipIf(process.platform==='win32')('real fake CLI observes two file epochs without imposing terminal report validation',async()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'batch-permission-pty-')),fake=path.join(dir,'fake-claude'),worker=path.join(dir,'worker.ts'),events=path.join(dir,'events.jsonl'),output=path.join(dir,'output.json'),expected=path.join(dir,'report.md');fs.writeFileSync(expected,'original');
const screen=captured.screen.replaceAll(path.dirname(captured.expectedPath),path.dirname(expected)).replaceAll(path.basename(captured.expectedPath),'report.md');
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import * as fs from 'node:fs';import * as path from 'node:path';
const item=JSON.parse(process.env.FILE_EPOCH_CASE);const log=e=>fs.appendFileSync(item.events,JSON.stringify(e)+'\n');
const sid='epoch-main';const nativePath=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','epoch',sid+'.jsonl');fs.mkdirSync(path.dirname(nativePath),{recursive:true});
const native=(role,content,extra={})=>fs.appendFileSync(nativePath,JSON.stringify({cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date().toISOString(),message:{role,content},...extra})+'\n');
native('assistant',[{type:'text',text:'Reviewing fixture.'}]);log({type:'start',pid:process.pid,cwd:process.cwd()});
const settings=JSON.parse(process.argv[process.argv.indexOf('--settings')+1]);
if(settings.hooks.PreToolUse[0].matcher!=='^ExitPlanMode$')throw Error('Exit recorder changed');
const hook=async(name,id)=>{
const entries=(settings.hooks[name]??[]).filter(h=>h.matcher==='^(Write|Edit)$');
if(entries.length!==1)throw Error('Expected exactly one caller-owned file recorder');
for(const entry of entries){
const event={hook_event_name:name,tool_name:'Edit',session_id:sid,tool_use_id:id,cwd:process.cwd(),transcript_path:nativePath,tool_input:{file_path:item.activePlan?path.join(process.cwd(),'PLAN.md'):item.expected,old_string:'old',new_string:'new'}};
const p=Bun.spawn(['bash','-c',entry.hooks[0].command],{stdin:new Blob([JSON.stringify(event)]),stdout:'pipe',stderr:'pipe'});
const [code,out,err]=await Promise.all([p.exited,new Response(p.stdout).text(),new Response(p.stderr).text()]);if(code||out||err)throw Error('hook was not silent');log({type:'hook',name,id});
}
};
let stage='startup';const paint=()=>process.stdout.write('\x1b[2J\x1b[H'+item.screen.replaceAll('__ACTIVE_PLAN_PATH__',path.join(process.cwd(),'PLAN.md')).replaceAll('\n','\r\n'));
process.stdin.setRawMode?.(true);process.stdin.on('data',async data=>{
const input=data.toString();log({type:'input',stage,input});
if(stage==='startup'){stage='first';await hook('PreToolUse','first');paint();return;}
if(stage==='old-pane'||stage==='done'){log({type:'unexpected'});return;}
if(input!=='1\r')throw Error('default permission input changed');
if(stage==='first'){stage='old-pane';await hook('PostToolUse','first');if(item.intervening){await hook('PreToolUse','automatic');await hook('PostToolUse','automatic');}paint();setTimeout(async()=>{await hook('PreToolUse','second');stage='second';paint();},3200);return;}
stage='done';await hook('PostToolUse','second');
const q={header:'Finding',question:'Apply this repair?',options:[{label:'Fix'},{label:'Keep'}]};
native('assistant',[{type:'tool_use',name:'AskUserQuestion',id:'finding',input:{questions:[q]}}]);native('user',[{type:'tool_result',tool_use_id:'finding',content:'Answered'}],{toolUseResult:{answers:{[q.question]:'Fix'}}});
process.stdout.write('\x1b[2J\x1b[HCompletion summary\r\n');
});process.on('SIGINT',()=>process.exit(0));process.stdin.resume();
`);fs.chmodSync(fake,0o755);
fs.writeFileSync(worker,`import {runPlanSkillCounting} from ${JSON.stringify(pathToFileURL(path.join(import.meta.dir,'helpers/claude-pty-runner.ts')).href)};const o=await runPlanSkillCounting({skillName:'plan-eng-review',slashCommand:'/plan-eng-review',followUpPrompt:'Review this disposable batching fixture.',permissionPlanPath:${JSON.stringify(expected)},isLastStep0AUQ:()=>false,isReviewAUQ:()=>true,reviewCountCeiling:2,timeoutMs:28000,env:{FILE_EPOCH_CASE:${JSON.stringify(JSON.stringify({events,expected,screen}))}}});await Bun.write(${JSON.stringify(output)},JSON.stringify(o));`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,EVALS_HERMETIC:'1'},stdout:'pipe',stderr:'pipe'});const killer=setTimeout(()=>child.kill('SIGKILL'),33000);
try{const[code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);expect(code,out+err).toBe(0);
const o=JSON.parse(fs.readFileSync(output,'utf8'));expect(o.outcome,JSON.stringify(o)).toBe('completion_summary');expect(o.reviewCount).toBe(1);expect(fs.readFileSync(expected,'utf8')).toBe('original');
const rows=fs.readFileSync(events,'utf8').trim().split('\n').map(l=>JSON.parse(l));expect(rows.filter(e=>e.type==='input').map(e=>e.input)).toEqual(['/plan-eng-review\r','1\r','1\r']);expect(rows.some(e=>e.type==='unexpected')).toBe(false);
expect(()=>process.kill(rows[0].pid,0)).toThrow();expect(fs.existsSync(rows[0].cwd)).toBe(false);
}finally{clearTimeout(killer);child.kill('SIGKILL');if(fs.existsSync(events)){const first=JSON.parse(fs.readFileSync(events,'utf8').split('\n')[0]!);try{process.kill(first.pid,'SIGKILL');}catch{}}fs.rmSync(dir,{recursive:true,force:true});}
},35000);
+98 -18
View File
@@ -19,7 +19,7 @@
* Reader-side fix folded from community PR #1851 by @harjothkhara.
*/
import { describe, test, expect } from 'bun:test';
import { execSync } from 'child_process';
import { execSync, spawnSync } from 'child_process';
import * as fs from 'fs';
import * as os from 'os';
import * as path from 'path';
@@ -34,15 +34,61 @@ const PATH_ADJACENT = /\/\$\{?_BRANCH|\$\{_BRANCH\}\/|\$_BRANCH\//;
// Raw $_BRANCH as a filename prefix (…-reviews.jsonl and friends).
const FILENAME_PREFIX = /\$\{?_BRANCH\}?[A-Za-z0-9._-]*\.(?:jsonl|json|md|txt|log)/;
function renderedSkillFiles(): string[] {
const out = execSync(
`find "${ROOT}" -name 'SKILL.md' -not -path '*/node_modules/*' -not -path '*/.claude/*' ; find "${ROOT}" -path '*/sections/*.md' -not -path '*/node_modules/*' -not -path '*/.claude/*'`,
{ encoding: 'utf-8', timeout: 30_000 },
);
return out.split('\n').filter(Boolean);
function renderedSkillFiles(root = ROOT): string[] {
// Enumerate managed render trees without buffering a shell's file census.
// .context holds archived/experimental copies, not shipped skill output.
const excluded = new Set(['node_modules', '.claude', '.context', '.git']);
const files: string[] = [];
function visit(dir: string, inSections = false) {
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
const file = path.join(dir, entry.name);
if (entry.isDirectory()) {
if (!excluded.has(entry.name)) visit(file, inSections || entry.name === 'sections');
} else if (entry.name === 'SKILL.md' || (inSections && entry.name.endsWith('.md'))) {
files.push(file);
}
}
}
visit(root);
return files;
}
describe('branch slug hygiene (#2550, #1851)', () => {
test('render discovery excludes scratch copies and retains every managed host without a pipe-size limit', () => {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-render-census-'));
const add = (relative: string) => {
const file = path.join(root, relative);
fs.mkdirSync(path.dirname(file), { recursive: true });
fs.writeFileSync(file, '# Render fixture\n');
return file;
};
try {
const expected = [
add('review/SKILL.md'), add('review/sections/analysis.md'),
add('.agents/skills/gstack-review/SKILL.md'),
add('.kiro/skills/gstack-review/sections/nested/analysis.md'),
add('skill with $quotes/sections/line\nbreak.md'),
];
for (const excluded of ['.context', '.claude', '.git', 'node_modules']) {
add(`${excluded}/old-render/review/SKILL.md`);
add(`${excluded}/old-render/review/sections/analysis.md`);
}
// The old execSync census failed at its 1 MiB stdout default once
// enough isolated host renders existed in a workspace.
for (let i = 0; i < 4500; i++) {
expected.push(add(`host-output/skill-${i}-${'x'.repeat(210)}/SKILL.md`));
}
expect(Buffer.byteLength(expected.join('\n'))).toBeGreaterThan(1024 * 1024);
const actual = renderedSkillFiles(root);
const expectedSet = new Set(expected);
expect(actual).toHaveLength(expected.length);
expect(new Set(actual).size).toBe(expected.length);
expect(actual.every(file => expectedSet.has(file))).toBe(true);
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
});
test('no generated SKILL.md or section interpolates raw $_BRANCH in a path position', () => {
const offenders: string[] = [];
for (const file of renderedSkillFiles()) {
@@ -84,7 +130,7 @@ describe('branch slug hygiene (#2550, #1851)', () => {
);
});
test('live round-trip: gstack-review-log writes, Context Recovery probe finds it (slash branch)', () => {
test('live round-trip: Context Recovery finds slugged reviews and raw timeline branches in a fresh shell', () => {
const home = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-home-'));
const repo = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-repo-'));
try {
@@ -110,22 +156,56 @@ describe('branch slug hygiene (#2550, #1851)', () => {
const proj = path.join(home, 'projects', slug);
expect(fs.existsSync(path.join(proj, 'feat-slug-hygiene-reviews.jsonl'))).toBe(true);
// Reader: execute the rendered probe line with $BRANCH from gstack-slug.
// Reader: execute the complete rendered block without inheriting the
// external skill-start process's private shell variables.
const ctx: TemplateContext = {
skillName: 'test-skill', tmplPath: 'test.tmpl', host: 'claude',
paths: HOST_PATHS.claude, preambleTier: 2,
paths: { ...HOST_PATHS.claude, binDir: '"$TEST_BIN"' }, preambleTier: 2,
};
const probeLine = generateContextRecovery(ctx)
.split('\n')
.find((l) => l.includes('-reviews.jsonl'))!;
const script = `_PROJ="${proj}"\nBRANCH="${branch}"\n${probeLine.trim()}`;
const out = execSync(`bash -c '${script.replace(/'/g, `'\\''`)}'`, {
cwd: repo, encoding: 'utf-8', timeout: 30_000,
});
expect(out).toContain('REVIEWS: 1 entries');
const script = generateContextRecovery(ctx).match(/```bash\n([\s\S]*?)\n```/)![1];
fs.writeFileSync(path.join(proj, 'timeline.jsonl'), [
{ branch: 'feat/slug-hygiene', event: 'completed', skill: 'review' },
{ branch: 'feat/slug-hygiene', event: 'started', skill: 'unfinished' },
{ branch, event: 'completed', skill: 'slugged-decoy' },
{ branch: 'stale/parent-branch', event: 'completed', skill: 'inherited-decoy' },
{ branch: 'unknown', event: 'completed', skill: 'fallback' },
{ branch: 'feat/slug-hygiene', event: 'completed', skill: 'ship' },
].map(entry => JSON.stringify(entry)).join('\n') + '\n');
const recover = (cwd: string, inheritedBranch?: string) => {
const result = spawnSync('bash', ['-c', script], {
cwd, encoding: 'utf8', timeout: 30_000,
env: { ...env, GSTACK_PROJECT_SLUG: slug, TEST_BIN: path.join(ROOT, 'bin'), _BRANCH: inheritedBranch },
});
expect(result.status, result.stderr).toBe(0);
expect(result.stderr).not.toContain('not a git repository');
return result.stdout;
};
for (const inheritedBranch of [undefined, 'stale/parent-branch']) {
const out = recover(repo, inheritedBranch);
expect(out).toContain('REVIEWS: 1 entries');
expect(out.split('\n').filter(line => line.startsWith('LAST_SESSION:'))).toEqual([
'LAST_SESSION: {"branch":"feat/slug-hygiene","event":"completed","skill":"ship"}',
]);
expect(out.split('\n').filter(line => line.startsWith('RECENT_PATTERN:'))).toEqual([
'RECENT_PATTERN: review,ship,',
]);
}
// Negative control: the raw-branch probe (the pre-fix shape) misses.
expect(fs.existsSync(path.join(proj, 'feat/slug-hygiene-reviews.jsonl'))).toBe(false);
// Both an unnamed checkout and a non-repository use the same unknown
// timeline identity as skill-start/end, never an inherited branch.
execSync('git checkout -q --detach', { cwd: repo, timeout: 30_000 });
for (const cwd of [repo, home]) {
const out = recover(cwd, 'stale/parent-branch');
expect(out.split('\n').filter(line => line.startsWith('LAST_SESSION:'))).toEqual([
'LAST_SESSION: {"branch":"unknown","event":"completed","skill":"fallback"}',
]);
expect(out.split('\n').filter(line => line.startsWith('RECENT_PATTERN:'))).toEqual([
'RECENT_PATTERN: fallback,',
]);
}
} finally {
fs.rmSync(home, { recursive: true, force: true });
fs.rmSync(repo, { recursive: true, force: true });
+81
View File
@@ -0,0 +1,81 @@
import { expect, test } from 'bun:test';
// Run the ownership probe in a child: an affected Bun can close arbitrary
// recycled descriptors, including the test runner's own sockets and pipes.
// Playwright uses the same extra-stdio slots for Chromium's CDP transport.
// Upstream ownership fixes: oven-sh/bun#32520 and oven-sh/bun#33828.
const fixture = String.raw`
const { spawn } = require('node:child_process');
const rounds = 4;
const listenersPerRound = 8;
for (let round = 0; round < rounds; round++) {
let child = spawn(process.execPath, ['-e', ''], {
stdio: ['ignore', 'ignore', 'ignore', 'pipe', 'pipe', 'pipe'],
});
await new Promise((resolve, reject) => {
child.once('exit', resolve);
child.once('error', reject);
});
await Promise.all(child.stdio.slice(3).map(socket => {
if (socket.closed) return;
return new Promise(resolve => {
// Subscribe before destroy: descriptor reuse must follow the actual
// close event, not a delay that may expire before close under load.
socket.once('close', resolve);
socket.destroy();
});
}));
// The OS can now reuse the closed extra-stdio descriptors for these
// listeners. Capture their URLs before GC; affected runtimes can also
// invalidate server.port when the underlying listener vanishes.
const listeners = Array.from({ length: listenersPerRound }, () => {
const server = Bun.serve({
hostname: '127.0.0.1', port: 0, fetch: () => new Response('alive'),
});
return { server, url: 'http://127.0.0.1:' + server.port + '/' };
});
child = null;
Bun.gc(true);
await Bun.sleep(0);
Bun.gc(true);
const failures = [];
for (const { url } of listeners) {
try {
const response = await fetch(url, { signal: AbortSignal.timeout(1_000) });
const body = await response.text();
if (response.status !== 200 || body !== 'alive') {
failures.push({ url, status: response.status, body });
}
} catch (error) {
failures.push({ url, error: String(error) });
}
}
if (failures.length) {
console.error(JSON.stringify({ bun: Bun.version, round, failures }));
// Only this isolated process is affected. Avoid asking the broken
// runtime to close descriptors again; process exit releases them.
process.exit(1);
}
for (const { server } of listeners) await server.stop(true);
}
console.log(JSON.stringify({ checkedListeners: rounds * listenersPerRound }));
`;
test.skipIf(process.platform === 'win32')('extra-stdio cleanup preserves unrelated HTTP listeners after GC', async () => {
const child = Bun.spawn([process.execPath, '-e', fixture], {
stdin: 'ignore', stdout: 'pipe', stderr: 'pipe', timeout: 10_000, killSignal: 'SIGKILL',
});
const [stdout, stderr, code] = await Promise.all([
new Response(child.stdout).text(), new Response(child.stderr).text(), child.exited,
]);
const diagnosis = `Bun ${Bun.version} failed the subprocess descriptor-ownership probe. `
+ 'Install the repository\'s pinned Bun version (1.4.0 or newer); '
+ 'older Bun can double-close extra stdio and destroy unrelated browser/server sockets '
+ '(oven-sh/bun#32520, #33828).\n'
+ `exit=${code}\nstdout:\n${stdout}\nstderr:\n${stderr}`;
expect(code, diagnosis).toBe(0);
expect(JSON.parse(stdout), diagnosis).toEqual({ checkedListeners: 32 });
expect(stderr, diagnosis).toBe('');
}, 15_000);
+14
View File
@@ -73,4 +73,18 @@ describe('bun version pins', () => {
expect(versions, `bun version drift across CI surfaces:\n${detail}`).toHaveLength(1);
expect(versions[0]).toMatch(/^\d+\.\d+\.\d+$/);
});
test('every CI surface requires Bun 1.4.0 or newer for safe extra-stdio ownership', () => {
// Matching pins alone would allow every lane to regress together. Older
// Linux Bun releases double-close extra stdio FDs during subprocess GC,
// which can close unrelated listeners after the OS reuses an FD number.
// https://github.com/oven-sh/bun/issues/34785#issuecomment-5020318035
for (const pin of collectPins()) {
expect(pin.version, `${pin.surface} must pin a stable numeric version`).toMatch(/^\d+\.\d+\.\d+$/);
expect(
Bun.semver.satisfies(pin.version, '>=1.4.0'),
`${pin.surface} pins Bun ${pin.version}; Bun >=1.4.0 is required for safe extra-stdio ownership`,
).toBe(true);
}
});
});
+2 -2
View File
@@ -22,7 +22,7 @@
import { test, expect } from 'bun:test';
import { CAPTURE_LONG_MS } from './helpers/eval-budgets';
import { describeE2ETier } from './helpers/e2e-gate';
import { setupSkillDir, skillFromWorktree, captureSectionReads } from './helpers/auq-sdk-capture';
import { setupSkillDir, skillFromWorktree, captureSectionReads, LONG_SECTION_CAPTURE_MS } from './helpers/auq-sdk-capture';
import { CARVE_GUARDS } from './helpers/carve-guards';
const describeE2E = describeE2ETier('periodic');
@@ -78,7 +78,7 @@ describeE2E('carve behavioral section-loading (periodic, SDK capture)', () => {
// their required section reads inside 60s but need 300-450s of
// wall clock to finish the report on slower sandboxes — a timeout
// there reads as a loading failure when the carve invariant held.
timeout: 480_000,
timeout: LONG_SECTION_CAPTURE_MS,
});
const missing = guard.requiredReads.filter((s) => !readSections.has(s));
+218
View File
@@ -0,0 +1,218 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import captured from './fixtures/ceo-annotation-aj.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const call = (index = 3): any => structuredClone(captured.cases.paired.calls[index]);
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
function edit(c: any, change: (s: string) => string) {
const q = c.questions[0], answer = c.answers[q.question];
q.question = change(q.question); c.answers = { [q.question]: answer };
}
test('completed native receipt and retry findings retain identity through section annotations', () => {
for (const index of [3, 4]) expect(ceoFirstReviewAUQ(fp(call(index)))).toBe(true);
});
test('actual setup remains excluded before the two completed assertion findings', () => {
let started = false;
const phases = captured.cases.paired.calls.map(c => {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
return phase.preReview;
});
expect(phases).toEqual([true, true, true, false, false]);
const approach = structuredClone(captured.cases.distinct.calls[2]);
expect(ceoFirstReviewAUQ(fp(approach))).toBe(false);
});
test('the new captured inputs belong only to the existing CEO count owner', () => {
for (const dependency of ['test/ceo-annotation-aj.test.ts', 'test/fixtures/ceo-annotation-aj.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name))
.toEqual(['plan-ceo-finding-count']);
}
});
test('section references do not replace finding or native option identity', () => {
for (const index of [3, 4]) {
const c = call(index);
edit(c, s => s.replace(/\(Sections? [^)]+\)/, '(Sections 3, 5 and 8, Error Handling)'));
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
c.answers[c.questions[0].question] = c.questions[0].options[1].label;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
const renamed = call(index), q = renamed.questions[0], old = String(index - 2);
edit(renamed, s => s.replace(new RegExp('Finding F' + old), 'Finding F9')
.replace(new RegExp('^Recommendation: ' + old, 'm'), 'Recommendation: 9')
.replace(new RegExp('^' + old + '([A-Z][)])', 'gm'), '9$1'));
q.header = q.header.replace(/^F\d+/, 'F9');
q.options.forEach((o: any) => { o.label = o.label.replace(/^\d+/, '9'); });
renamed.answers = { [q.question]: q.options[0].label };
expect(ceoFirstReviewAUQ(fp(renamed))).toBe(true);
}
});
test('source frames and conditional or missing assessments cannot supply a current finding', () => {
for (const index of [3, 4]) for (const change of [
(s: string) => 'Example: ' + s,
(s: string) => '> ' + s,
(s: string) => '```\n' + s + '\n```',
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: `$1`'),
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Suppose '),
(s: string) => s.replace(/^ELI10: .+$/m, ''),
(s: string) => s + '\nELI10: A second contradictory assessment.',
(s: string) => s.replace(/: the (success|repeated)/, ': the hypothetical $1'),
(s: string) => s.replace(/\(Sections? [^)]+\)/, '(Section 6, Quoted Source)'),
(s: string) => s.replace(/\(Sections? [^)]+\)/, '(Section 6, Historical Example)'),
]) { const c = call(index); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
});
test('same-brief withdrawals defeat a finding while attributed historical quotes do not', () => {
for (const index of [3, 4]) {
for (const tail of ['This issue is withdrawn.', 'We have withdrawn this finding.',
'There is no current defect or unresolved issue.', `F${index - 2} is rejected.`]) {
const c = call(index); edit(c, s => s + '\n' + tail); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
const quoted = call(index); edit(quoted, s => s + '\nOld note: "This issue is withdrawn."');
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
}
});
test('administrative options and stale action rows do not amend the current contract', () => {
for (const index of [3, 4]) {
for (const wording of ['Start review', 'Pause', 'Write the completed report']) {
const stale = call(index), native = stale.questions[0];
native.options.forEach((o: any, i: number) => {
o.label = `${index - 2}${String.fromCharCode(65 + i)}: ${wording}`;
o.description = wording;
});
stale.answers = { [native.question]: native.options[0].label };
expect(ceoFirstReviewAUQ(fp(stale))).toBe(false);
}
const c = call(index), q = c.questions[0];
q.options.forEach((o: any, i: number) => {
o.label = `${index - 2}${String.fromCharCode(65 + i)}: Archive the completed report ${i}`;
o.description = 'Save the completed review for reference.';
});
c.answers = { [q.question]: q.options[0].label };
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
const noGap = call(index);
edit(noGap, s => s.replace(/^D\d+[^\n]+/, `D4 — Finding F${index - 2} (Sections 2 and 6): where should the completed report be stored?`)
.replace(/^ELI10: .+$/m, 'ELI10: The review is complete and all assertions already enforce the full contract.'));
expect(ceoFirstReviewAUQ(fp(noGap))).toBe(false);
const negated = call(index);
edit(negated, s => s.replace(/asserts only/, 'does not assert only')
.replace(/^ELI10: .+$/m, 'ELI10: The assertions enforce the complete receipt and retry contracts.'));
expect(ceoFirstReviewAUQ(fp(negated))).toBe(false);
}
});
test('completion, recommendation, identity and actual offered options remain required', () => {
for (const index of [3, 4]) for (const mutate of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.sessionId = ''; },
(c: any) => { c.answers = {}; },
(c: any) => { c.answers[c.questions[0].question] = 'Foreign answer'; },
(c: any) => { c.questions[0].multiSelect = true; },
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
(c: any) => { c.questions[0].header = 'Approach'; },
(c: any) => { c.questions[0].header = 'Finding 99'; },
(c: any) => { c.questions[0].options[1].description = ''; },
(c: any) => { c.questions[0].options[1].label = '99B: Different finding'; },
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, 'Recommendation: 99Z')),
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, '')),
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-eng-review-finding>'),
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
for (const index of [3, 4]) {
const original = fp(call(index));
expect(ceoFirstReviewAUQ({ ...original, signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...original, nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...original, options: original.options.slice(1) })).toBe(false);
}
});
test('completed dotted issue briefs retain their full identity and section option binding', () => {
for (const source of captured.cases.distinct.calls.slice(4)) {
expect(ceoFirstReviewAUQ(fp(source))).toBe(true);
}
let started = false;
const phases = captured.cases.distinct.calls.map(c => {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
return phase.preReview;
});
expect(phases).toEqual([true, true, true, true, false, false, false, false, false]);
});
test('dotted issue syntax never supplies missing current defect or remedy evidence', () => {
for (const source of captured.cases.distinct.calls.slice(4)) {
const identity = /\(Issue ([\d.]+)\)/.exec(source.questions[0]!.question)![1]!;
for (const change of [
(s: string) => 'Example: ' + s,
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Historical example: '),
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: `$1`'),
(s: string) => s.replace(/^ELI10: .+$/m, ''),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The completed review has no current defect or unresolved issue.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The existing behavior satisfies every contract and needs no change.'),
(s: string) => s + '\nThis issue has been resolved.',
(s: string) => s + `\nIssue ${identity} is rejected.`,
(s: string) => s.replace(/^Recommendation: \d+[A-Z]/m, 'Recommendation: 99Z'),
(s: string) => s + '\n<gstack-qid:plan-eng-review-finding>',
]) { const c = structuredClone(source); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
const sourceOnly = structuredClone(source), q = sourceOnly.questions[0]!;
q.options.forEach((o, i) => { o.label = `${identity.split('.')[0]}${String.fromCharCode(65 + i)}: Archive the completed report ${i}`; o.description = 'Save the completed review.'; });
sourceOnly.answers = { [q.question]: q.options[0]!.label };
expect(ceoFirstReviewAUQ(fp(sourceOnly))).toBe(false);
const quoted = structuredClone(source); edit(quoted, s => s + `\nOld note: "Issue ${identity} is rejected."`);
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
}
});
test('dotted identities remain complete while the option prefix names the containing section', () => {
for (const source of captured.cases.distinct.calls.slice(4)) {
const identity = /\(Issue ([\d.]+)\)/.exec(source.questions[0]!.question)![1]!;
for (const mutate of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.answers = {}; },
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
(c: any) => { c.questions[0].header = 'Approach'; },
(c: any) => { c.questions[0].header = `Finding ${identity.split('.')[0]}`; },
(c: any) => { c.questions[0].header = 'Issue 99.1'; },
(c: any) => { c.questions[0].options[1].label = '99B: Borrowed option'; },
(c: any) => { c.questions[0].options[1].description = ''; },
(c: any) => { c.questions[0].multiSelect = true; },
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
(c: any) => edit(c, s => s.replace(`Issue ${identity}`, 'Issue 1.0')),
]) { const c = structuredClone(source); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
const header = structuredClone(source); header.questions[0]!.header = `Finding ${identity}`;
expect(ceoFirstReviewAUQ(fp(header))).toBe(true);
const localQid = structuredClone(source); edit(localQid, s => s + '\n<gstack-qid:plan-ceo-review-handler-correctness>');
expect(ceoFirstReviewAUQ(fp(localQid))).toBe(true);
expect(ceoFirstReviewAUQ({ ...fp(source), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(source), nativeCall: undefined })).toBe(false);
}
});
test('an owning assessment declaration cannot relabel source or hypothetical prose as a current finding', () => {
for (const source of [...captured.cases.paired.calls.slice(3), ...captured.cases.distinct.calls.slice(4)]) {
for (const frame of [
'The following is a quoted source excerpt.',
'The following is a hypothetical example.',
'This assessment is only a historical example.',
]) {
const c = structuredClone(source);
edit(c, s => s.replace(/^ELI10: /m, 'ELI10: ' + frame + ' '));
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
const quoted = structuredClone(source);
edit(quoted, s => s.replace(/^(ELI10: .+)$/m, '$1 Old note: "The following is a hypothetical example."'));
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
}
});
+139
View File
@@ -0,0 +1,139 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import captured from './fixtures/ceo-annotation-header-at.json';
const call = (): any => structuredClone(captured.calls[2]);
const fp = (value: any) => nativePlanCallFingerprint(value, 0, true);
const matches = (value: any) => ceoFirstReviewAUQ(fp(value));
function edit(value: any, change: (text: string) => string) {
const q = value.questions[0], answer = value.answers[q.question];
q.question = change(q.question);
value.answers = { [q.question]: answer };
}
test('the exact completed section-annotated mail rescue finding opens review', () => {
const original = call();
expect(matches(original)).toBe(true);
expect(original).toEqual(captured.calls[2]);
let started = false;
const phases = captured.calls.map(value => {
const phase = planCountQuestionPhase(fp(value), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
return phase.preReview;
});
expect(phases).toEqual([true, true, false, false, false, false, false, false]);
});
test('descriptive headers and matching finding renames preserve the same rich decision', () => {
for (const header of ['Email rescue', 'Mail failure', 'Receipt retry', 'Finding 2', 'Issue 2', 'F2 rescue']) {
const value = call(); value.questions[0].header = header;
expect(matches(value)).toBe(true);
}
for (const label of call().questions[0].options.map((option: any) => option.label)) {
const value = call(); value.answers[value.questions[0].question] = label;
expect(matches(value)).toBe(true);
}
const renamed = call();
edit(renamed, text => text.replace('Finding 2 (Section 2', 'Finding 9 (Section 2')
.replace(/^Recommendation: 2A/m, 'Recommendation: 9A'));
renamed.questions[0].options.forEach((option: any) => { option.label = option.label.replace(/^2/, '9'); });
renamed.answers = { [renamed.questions[0].question]: renamed.questions[0].options[0].label };
expect(matches(renamed)).toBe(true);
const decision = call(); edit(decision, text => text.replace(/^D2/, 'D19'));
expect(matches(decision)).toBe(true);
});
test('section metadata cannot override conflicting or malformed identities', () => {
for (const header of ['Finding 9', 'Issue 9', 'F9 rescue', 'Finding 2.1', 'Finding zero', 'Section 9', 'Section 2']) {
const value = call(); value.questions[0].header = header;
expect(matches(value)).toBe(false);
}
for (const change of [
(text: string) => text.replace('Finding 2 (Section 2, CRITICAL GAP)', 'Finding 0 (Section 2, CRITICAL GAP)'),
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 0, CRITICAL GAP)'),
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, Historical Example)'),
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, Quoted Source)'),
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, CRITICAL GAP) (Section 9)'),
(text: string) => text.replace('Finding 2 (Section 2, CRITICAL GAP)', 'Finding 2 and Finding 9 (Section 2, CRITICAL GAP)'),
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2)'),
(text: string) => text.replace(/^Recommendation: 2A/m, 'Recommendation: 9A'),
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
});
test('source, hypothetical, withdrawn or missing assessments do not open review', () => {
for (const change of [
(text: string) => 'Example: ' + text,
(text: string) => '> ' + text,
(text: string) => '```\n' + text + '\n```',
(text: string) => text.replace('\nProject/branch/task:', '\nSource:\nProject/branch/task:'),
(text: string) => text.replace(/^ELI10: /m, 'ELI10: If approved, '),
(text: string) => text.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
(text: string) => text.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: This handler has no current defect and needs no amendment.'),
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: The plan needs no change.'),
(text: string) => text.replace(/^ELI10: .+$/m, ''),
(text: string) => text + '\nELI10: Another assessment.',
(text: string) => text + '\nThis finding is withdrawn.',
(text: string) => text + '\nThis finding is "withdrawn".',
(text: string) => text + '\nThis finding is hypothetical.',
(text: string) => text + '\nThis finding is not current.',
(text: string) => text + '\nThis finding is no longer current.',
(text: string) => text + '\nThis finding is "no longer current".',
(text: string) => text + '\nThis finding is superseded.',
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
const historical = call(); edit(historical, text => text + '\nOld note: "This finding is withdrawn."');
expect(matches(historical)).toBe(true);
const resolvedHistory = call();
edit(resolvedHistory, text => text.replace(/^(ELI10: .+)$/m, '$1 Old note: "This handler has no current defect and needs no amendment."'));
expect(matches(resolvedHistory)).toBe(true);
});
test('only current offered remedies can supply the amendment', () => {
for (const prefix of ['Source: ', 'If approved, ', 'This remedy is withdrawn. ', 'This remedy is "withdrawn". ']) {
const value = call();
value.questions[0].options.forEach((option: any) => { option.description = prefix + option.description; });
expect(matches(value)).toBe(false);
}
const report = call();
report.questions[0].options.forEach((option: any, index: number) => {
option.label = `2${String.fromCharCode(65 + index)}) Archive the completed report ${index}`;
option.description = 'Save the completed review for reference.';
});
report.answers = { [report.questions[0].question]: report.questions[0].options[0].label };
expect(matches(report)).toBe(false);
});
test('the completed native identity, offered choice and answer remain required', () => {
for (const change of [
(value: any) => { value.answered = false; },
(value: any) => { value.failed = true; },
(value: any) => { value.sessionId = ''; },
(value: any) => { value.toolUseId = ''; },
(value: any) => { value.unansweredQuestionIndices = [0]; },
(value: any) => { value.answeredAt = 'invalid'; },
(value: any) => { value.answers = {}; },
(value: any) => { value.answers[value.questions[0].question] = 'Foreign answer'; },
(value: any) => { value.questions[0].multiSelect = true; },
(value: any) => { value.questions.push(structuredClone(value.questions[0])); },
(value: any) => { value.questions[0].header = 'Approach'; },
(value: any) => { value.questions[0].options[1].description = ''; },
(value: any) => { value.questions[0].options[1].label = '9B) Borrowed amendment'; },
(value: any) => edit(value, text => text.replace(/^Recommendation: .+$/m, '')),
(value: any) => edit(value, text => text + '\n<gstack-qid:plan-eng-review-finding>'),
]) { const value = call(); change(value); expect(matches(value)).toBe(false); }
const original = fp(call());
for (const changed of [
{ ...original, signature: 'foreign:call' },
{ ...original, nativeCall: undefined },
{ ...original, nativeQuestionIndex: 1 },
{ ...original, options: original.options.slice(1) },
]) expect(ceoFirstReviewAUQ(changed)).toBe(false);
});
test('new retained inputs belong only to the CEO finding-count workflow', () => {
for (const file of ['test/ceo-annotation-header-at.test.ts', 'test/fixtures/ceo-annotation-header-at.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name))
.toEqual(['plan-ceo-finding-count']);
}
});
+261
View File
@@ -0,0 +1,261 @@
import { describe, expect, test } from 'bun:test';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { capturePlanCountQuestion, nativePlanCallFingerprint, planCountQuestionInput } from './helpers/claude-pty-runner';
import { pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import { pickCeoCountQuestion, pickCeoRecommendedApproach } from './helpers/ceo-approach-pick';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import recorded from './fixtures/ceo-approach-q-call.json';
import pairedRecorded from './fixtures/ceo-approach-q-paired-call.json';
import handoffs from './fixtures/ceo-completion-handoff-m-call.json';
import recordedY from './fixtures/ceo-approach-y-call.json';
import recordedAA from './fixtures/ceo-approach-aa-call.json';
function pending(source: NativePlanQuestionCall = recorded as NativePlanQuestionCall): NativePlanQuestionCall {
const call = structuredClone(source);
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
return call;
}
const fingerprint = (call: NativePlanQuestionCall, preReview = true) => nativePlanCallFingerprint(call, 0, preReview);
describe('numbered native approach identity', () => {
test('the actual AA question selects its offered recommendation with a projected pending binding', () => {
// Only completed live versions survived capture. Preserve the actual A
// answer; this projection tests routing, not live metadata availability.
const call = pending(recordedAA as NativePlanQuestionCall);
const active = capturePlanCountQuestion(screen(call), new Set(), 0, true, call)!;
expect(active.nativeCall).toBe(call);
expect(pickCeoCountQuestion(fingerprint(call), active)).toBe(2);
expect(planCountQuestionInput(screen(call), active, 2)).toBe('2');
expect(recordedAA.answers[recordedAA.questions[0]!.question]).toBe(recordedAA.questions[0]!.options[0]!.label);
expect(pickCeoCountQuestion(fingerprint(recordedAA as NativePlanQuestionCall))).toBeNull();
});
test('decision numbers and option positions may change together without changing policy', () => {
for (const decision of ['2', '37']) {
const call = pending(recordedAA as NativePlanQuestionCall);
const q = call.questions[0]!;
q.question = q.question.replace(/^D1/, `D${decision}`).replace('approach-d1>', `approach-d${decision}>`);
q.options.reverse();
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
q.options.unshift(q.options.pop()!);
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(3);
}
});
test('numbered identities must agree with the explicit decision and remain a supported approach id', () => {
for (const id of ['plan-ceo-review-approach-d2', 'plan-ceo-review-approach-d0',
'plan-ceo-review-approach-d01', 'plan-ceo-review-approach-d1-extra',
'plan-eng-review-approach-d1', 'plan-ceo-review-mode-d1']) {
const call = pending(recordedAA as NativePlanQuestionCall);
call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-review-approach-d1', id);
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
}
for (const prefix of ['', 'D2 — ', 'Example: D1 — ', '> D1 — ']) {
const call = pending(recordedAA as NativePlanQuestionCall);
call.questions[0]!.question = call.questions[0]!.question.replace(/^D1 — /, prefix);
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
}
});
test('numbered ids retain the native binding, phase, question and sole recommendation guards', () => {
const call = pending(recordedAA as NativePlanQuestionCall);
const fp = fingerprint(call);
const unbound = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
expect(pickCeoCountQuestion(fp, unbound)).toBeNull();
expect(pickCeoCountQuestion({...fp, preReview: false})).toBeNull();
expect(pickCeoCountQuestion({...fp, signature: 'foreign:call'})).toBeNull();
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('should this plan use?', 'should this plan not use?'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Mode'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; },
(c: NativePlanQuestionCall) => { c.failed = true; },
]) {
const changed = pending(recordedAA as NativePlanQuestionCall); mutate(changed);
expect(pickCeoRecommendedApproach(fingerprint(changed))).toBeNull();
}
});
});
describe('Y named component approach menu', () => {
const actualScreen = readFileSync(join(import.meta.dir, 'fixtures/ceo-approach-y-screen.txt'), 'utf8');
test('the exact full frame and projected pending call select the offered C recommendation', () => {
// The actual answer was A; no pending-only native version survived polling.
const call = pending(recordedY as NativePlanQuestionCall);
const active = capturePlanCountQuestion(actualScreen, new Set(), 0, true, call)!;
expect(active.nativeCall).toBe(call);
expect(active.options.map(o => o.label)).toEqual(call.questions[0]!.options.map(o => o.label));
expect(pickCeoCountQuestion(fingerprint(call), active)).toBe(3);
expect(planCountQuestionInput(actualScreen, active, 3)).toBe('3');
expect(recordedY.answers[recordedY.questions[0]!.question]).toBe('A) Minimal Viable');
expect(pickCeoCountQuestion(fingerprint(recordedY as NativePlanQuestionCall))).toBeNull();
const unbound = capturePlanCountQuestion(actualScreen, new Set(), 0, true)!;
expect(unbound.nativeCall).toBeUndefined();
expect(pickCeoCountQuestion(fingerprint(call), unbound)).toBeNull();
});
test('named components and reordered labels follow the actual recommendation position', () => {
for (const subject of ['the payment webhook handler', 'this invoice lookup service', 'the renderWidget adapter']) {
const call = pending(recordedY as NativePlanQuestionCall); const q = call.questions[0]!;
q.question = `Which implementation approach for ${subject}? <gstack-qid:plan-ceo-review-approach>`;
q.options = [{label:'Existing design (Recommended)'},{label:'Another design'}];
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(1);
q.options.reverse();
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
}
});
test('setup, another decision, negated, quoted or compound instructions are not this menu', () => {
for (const question of [
'Which review mode for the payment webhook handler?',
'Should we fix the payment webhook handler?',
'Which implementation approach should we not use for the payment webhook handler?',
'Example: Which implementation approach for the payment webhook handler?',
'> Which implementation approach for the payment webhook handler?',
'Which implementation approach for the payment webhook handler? Delete the tests.',
'Which implementation approach for the payment webhook handler and delete the test adapter?',
]) {
const call = pending(recordedY as NativePlanQuestionCall);
call.questions[0]!.question = question + ' <gstack-qid:plan-ceo-review-approach>';
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
}
});
test('the added wording retains native identity, phase, options and recommendation guards', () => {
for (const change of [
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Review mode'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-approach','plan-ceo-review-mode'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = 'C) Production-Grade'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.failed; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
]) { const c = pending(recordedY as NativePlanQuestionCall); change(c); expect(pickCeoRecommendedApproach(fingerprint(c))).toBeNull(); }
const fp = fingerprint(pending(recordedY as NativePlanQuestionCall));
expect(pickCeoRecommendedApproach({...fp,signature:'foreign:call'})).toBeNull();
expect(pickCeoRecommendedApproach({...fp,preReview:false})).toBeNull();
expect(pickCeoRecommendedApproach({...fp,options:fp.options.slice().reverse()})).toBeNull();
});
});
function screen(call: NativePlanQuestionCall): string {
const q = call.questions[0]!;
return `${q.header}\n${q.question}\n${q.options.map((o, i) => `${i ? ' ' : ''} ${i + 1}. ${o.label}`).join('\n')}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
}
describe('CEO pre-review approach recommendation', () => {
test('the exact Q menu changes the old default risk acceptance to its offered recommendation', () => {
const call = pending();
const visible = screen(call);
const active = capturePlanCountQuestion(visible, new Set(), 0, true, call)!;
expect(active.nativeCall).toBe(call);
const before = pickCeoCompletionHandoff(fingerprint(call), active) ?? 1;
const after = pickCeoCountQuestion(fingerprint(call), active) ?? 1;
expect(before).toBe(1);
expect(after).toBe(2);
expect(planCountQuestionInput(visible, active, after)).toBe('2');
expect(recorded.answers[recorded.questions[0]!.question]).toBe(recorded.questions[0]!.options[0]!.label);
expect(pickCeoCountQuestion(fingerprint(recorded as NativePlanQuestionCall))).toBeNull();
});
test('the paired first native approach uses the same offered recommendation policy', () => {
const call = pending(pairedRecorded as NativePlanQuestionCall);
const visible = screen(call);
const active = capturePlanCountQuestion(visible, new Set(), 0, true, call)!;
expect(active.nativeCall).toBe(call);
expect(pickCeoCompletionHandoff(fingerprint(call), active) ?? 1).toBe(1);
const after = pickCeoCountQuestion(fingerprint(call), active) ?? 1;
expect(after).toBe(2);
expect(planCountQuestionInput(visible, active, after)).toBe('2');
expect(pairedRecorded.answers[pairedRecorded.questions[0]!.question]).toBe(pairedRecorded.questions[0]!.options[0]!.label);
expect(pickCeoCountQuestion(fingerprint(pairedRecorded as NativePlanQuestionCall))).toBeNull();
});
test('paired approach grammar is function-agnostic and follows reordered options', () => {
const call = pending(pairedRecorded as NativePlanQuestionCall);
const q = call.questions[0]!;
q.question = 'D3 — Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-review-impl-approach>';
q.options = [{ label: 'A) Custom renderer' }, { label: 'B) Existing renderer (Recommended)' }];
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
q.options.reverse();
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(1);
});
test.each([
['non-approach question', 'Should the renderWidget() tests be deleted? <gstack-qid:plan-ceo-review-impl-approach>'],
['negated question', 'Which implementation approach should the renderWidget() tests not use? <gstack-qid:plan-ceo-review-impl-approach>'],
['negated test subject', 'Which implementation approach for not testing renderWidget()? <gstack-qid:plan-ceo-review-impl-approach>'],
['wrong approach identity', 'Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-approach>'],
['extra action before question', 'Delete the tests. Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-review-impl-approach>'],
])('does not apply paired approach selection to %s', (_name, question) => {
const call = pending(pairedRecorded as NativePlanQuestionCall);
call.questions[0]!.question = question;
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
});
test('recommendation follows actual option position and arbitrary approach content, never seed words', () => {
for (const order of [[0, 1, 2], [1, 2, 0], [2, 0, 1]]) {
const call = pending();
const q = call.questions[0]!;
const options = [{ label: 'A) Compare two renderers' }, { label: 'B) Existing renderer (Recommended)' }, { label: 'C) Custom renderer' }];
q.options = order.map(index => options[index]!);
q.question = 'D1 — Which implementation approach should this plan use? <gstack-qid:plan-ceo-approach>';
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(order.indexOf(1) + 1);
}
});
test.each([
['no recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Secure Baseline'; }],
['duplicate recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; }],
['duplicate offered label', (c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = c.questions[0]!.options[1]!.label; }],
['negated recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Not (Recommended)'; }],
['conflicting recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Not recommended here (Recommended)'; }],
['description-only recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Secure Baseline'; c.questions[0]!.options[1]!.description = 'Recommended'; }],
['unknown qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-approach', 'plan-ceo-security'); }],
['missing qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('<gstack-qid:plan-ceo-approach>', ''); }],
['malformed extra qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question += '<gstack-qid:broken'; }],
['duplicate qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question += '<gstack-qid:plan-ceo-approach>'; }],
['negated approach question', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('should this plan use?', 'should this plan not use?'); }],
['non-approach question', (c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we accept this security risk? <gstack-qid:plan-ceo-approach>'; }],
['non-approach header', (c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Review Mode'; }],
['multi-select', (c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; }],
['mixed packet', (c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); }],
['failed native call', (c: NativePlanQuestionCall) => { c.failed = true; }],
])('keeps the old caller/default policy for %s', (_name, change) => {
const call = pending();
change(call);
const fp = fingerprint(call);
expect(pickCeoRecommendedApproach(fp)).toBeNull();
expect(pickCeoCountQuestion(fp)).toBe(pickCeoCompletionHandoff(fp));
});
test('requires current native binding and pre-review phase', () => {
const call = pending();
const fp = fingerprint(call);
expect(pickCeoRecommendedApproach({ ...fp, preReview: false })).toBeNull();
expect(pickCeoRecommendedApproach({ ...fp, signature: 'foreign:call' })).toBeNull();
expect(pickCeoRecommendedApproach({ ...fp, nativeQuestionIndex: 1 })).toBeNull();
expect(pickCeoRecommendedApproach({ ...fp, options: fp.options.slice().reverse() })).toBeNull();
const visibleOnly = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
expect(visibleOnly.nativeCall).toBeUndefined();
expect(pickCeoCountQuestion(fp, visibleOnly)).toBeNull();
const foreign = '☐ Finding\nShould we add validation?\n 1. Add fix\n 2. Defer\nEnter to select · ↑/↓ to navigate · Esc to cancel';
const active = capturePlanCountQuestion(foreign, new Set(), 0, true, call)!;
expect(active.nativeCall).toBeUndefined();
expect(pickCeoCountQuestion(fp, active)).toBeNull();
});
test('the existing completed-review manual picker still runs after approach selection declines', () => {
const call = structuredClone(handoffs.calls.at(-1)!) as NativePlanQuestionCall;
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
const fp = fingerprint(call, false);
const expected = pickCeoCompletionHandoff(fp);
expect(expected).not.toBeNull();
expect(pickCeoCountQuestion(fp)).toBe(expected);
});
test('both count callers use the composed picker while leaving first-scope and count predicates intact', () => {
const caller = readFileSync(join(import.meta.dir, 'skill-e2e-plan-ceo-finding-count.test.ts'), 'utf8');
expect(caller.match(/pickAUQ: pickCeoCountQuestion/g)).toHaveLength(2);
expect(caller.match(/isFirstReviewAUQ: ceoFirstReviewAUQ/g)).toHaveLength(2);
expect(caller).toContain('firstAUQPick: pickSkipInterview');
});
});
+89
View File
@@ -0,0 +1,89 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, nativePlanCallFingerprint, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
import fixture from './fixtures/ceo-assertion-header-am-calls.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const calls = fixture.calls as AskUserQuestionFingerprint[];
const findings = calls.slice(2);
function change(fp: AskUserQuestionFingerprint, edit: (q: NonNullable<AskUserQuestionFingerprint['nativeCall']>['questions'][number]) => void) {
const call = structuredClone(fp.nativeCall!);
const selected = call.questions[0]!.options.findIndex(o => o.label === call.answers?.[call.questions[0]!.question]);
edit(call.questions[0]!);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[selected]!.label };
return nativePlanCallFingerprint(call, fp.observedAtMs, fp.preReview);
}
for (const [i, fp] of findings.entries()) {
test(`the actual completed assertion finding ${i + 1} starts review with its descriptive header`, () => {
expect(ceoFirstReviewAUQ(fp)).toBe(true);
});
}
test('routing and implementation layout remain setup', () => {
for (const fp of calls.slice(0, 2)) expect(ceoFirstReviewAUQ(fp)).toBe(false);
});
test('the same current issue is already recognized with an explicit numbered header', () => {
findings.forEach((fp, i) => expect(ceoFirstReviewAUQ(change(fp, q => { q.header = `Issue ${i + 1}`; }))).toBe(true));
});
test('a competing numbered header cannot borrow the title issue', () => {
findings.forEach((fp, i) => expect(ceoFirstReviewAUQ(change(fp, q => { q.header = `Issue ${i + 2}`; }))).toBe(false));
});
test('only the completed owned native decision supplies the finding', () => {
for (const fp of findings) {
for (const mutate of [
(x: AskUserQuestionFingerprint) => { x.nativeCall!.answered = false; },
(x: AskUserQuestionFingerprint) => { x.nativeCall!.failed = true; },
(x: AskUserQuestionFingerprint) => { x.nativeCall!.unansweredQuestionIndices = [0]; },
(x: AskUserQuestionFingerprint) => { x.signature = 'foreign:call'; },
(x: AskUserQuestionFingerprint) => { x.nativeCall!.answers = {}; },
(x: AskUserQuestionFingerprint) => { x.options[0]!.label = 'different menu'; },
]) {
const modified = structuredClone(fp); mutate(modified);
expect(ceoFirstReviewAUQ(modified)).toBe(false);
}
}
});
test('descriptive headers and decision ordinals do not replace the issue identity', () => {
for (const fp of findings) {
for (const titlePrefix of ['D19', 'd4']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/^D\d+/, titlePrefix); }))).toBe(true);
}
expect(ceoFirstReviewAUQ(change(fp, q => { q.header = 'Test contract'; }))).toBe(true);
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/^Recommendation: \d+/m, 'Recommendation: 9'); }))).toBe(false);
}
});
test('source, earlier and conditional framing cannot own the current assessment', () => {
for (const fp of findings) {
for (const prefix of ['Source excerpt:', 'The following assessment is hypothetical.', 'Earlier review assessment:', 'If approved:', 'Source:', 'Example:', 'Historical review:']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', `\n${prefix}\nELI10:`); }))).toBe(false);
}
for (const prefix of ['Source excerpt. ', 'Previously, ', 'If approved, ']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', `ELI10: ${prefix}`); }))).toBe(false);
}
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\nArchived wording: "Source excerpt."\nELI10:'); }))).toBe(true);
}
});
test('the assertion gap and offered remedy must still be current', () => {
for (const fp of findings) {
for (const correction of ['This finding is withdrawn.', 'No current defect remains.']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + correction; }))).toBe(false);
}
expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace(/^\d+[A-Z]\)[\s\S]*?(?=^Net:)/m, '');
const issue = q.options[0]!.label.match(/^\d+/)![0];
q.options.forEach((option, i) => {
option.label = `${issue}${String.fromCharCode(65 + i)}: Keep the current assertion`;
option.description = 'Leave the assertion unchanged.';
});
}))).toBe(false);
}
});
test('regression inputs belong only to the existing CEO finding owner without sparse paths', () => {
for (const input of ['test/ceo-assertion-header-am.test.ts', 'test/fixtures/ceo-assertion-header-am-calls.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(input)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
}
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
for (let i = 0; i < paths.length; i++) {
expect(Object.hasOwn(paths, i)).toBe(true);
expect(typeof paths[i]).toBe('string');
}
});
+127
View File
@@ -0,0 +1,127 @@
import { describe, expect, test } from 'bun:test';
import { hasNativePostAnswerCeoPosture, nextCeoPostureContinuation } from './helpers/ceo-mode-option';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/ceo-barless-submit-ac.json';
const selectedAt = Date.parse(captured.provenance.modeRequestAt);
const questions = captured.projectedQuestions;
const chrome = 'Planning:\n/tmp/owned-plan.md\n─────────────────\n';
function transcript(native = false, count = 3): PlanCountTranscript {
return { status: 'ready', assistantMessages: [], calls: [structuredClone(captured.modeCall),
...(native ? [{ sessionId: captured.modeCall.sessionId, toolUseId: 'projected-pending-call',
answered: false, failed: false, questions: structuredClone(questions.slice(0, count)) }] : [])] };
}
// Complete projected frames exercise the observed barless layout. The raw
// historical tail is retained separately and is never promoted to a full frame.
function screen(index: number, count = 3): string {
const current = questions[index]!;
return chrome + '← ' + questions.slice(0, count).map((q, i) => `${i < index ? '☒' : '☐'} ${q.header}`).join(' ') +
' ✔ Submit →\n' + current.question + '\n' + current.options.map((option, i) =>
`${i === 0 ? '' : ''}${i + 1}. ${option.label}`).join('\n') +
'\nEnter to select · Tab/Arrow keys to navigate · Esc to cancel\n';
}
function summary(count = 3): string {
return chrome + 'Review your answers\n' + questions.slice(0, count).map(question =>
`│ ● ${question.question.replaceAll('\n', '\n│ ')}\n│ → ${question.options[0]!.label}\n`).join('\n') +
'\nReady to submit your answers?\n1. Submit answers\n2. Cancel\n';
}
function observed(count = 3, native = false) {
const t = transcript(native, count), seen = new Set<string>();
for (let i = 0; i < count; i++) {
expect(nextCeoPostureContinuation(screen(i, count), t, 'HOLD SCOPE', selectedAt, seen, i > 0)).toBe('question');
}
const submit = (visible = summary(count), native = t) =>
nextCeoPostureContinuation(visible, native, 'HOLD SCOPE', selectedAt, seen, true);
return { t, seen, submit };
}
describe('bounded CEO barless packet submission', () => {
test('two through four observed tabs submit once with delayed or eager native identity', () => {
for (const count of [2, 3, 4]) for (const native of [false, true]) {
const p = observed(count, native);
expect(p.submit()).toBe('submission');
expect(p.submit()).toBeNull();
expect(p.submit(screen(0, count))).toBeNull();
expect(hasNativePostAnswerCeoPosture(p.t, 'HOLD SCOPE', /hold\s*scope/i, selectedAt)).toBe(false);
}
});
test('an exact late native packet binds all prior tabs before Submit', () => {
const p = observed();
expect(p.submit(summary(), transcript(true))).toBe('submission');
expect(p.submit(summary(), transcript(true))).toBeNull();
});
test.each(['different question', 'different selected answer', 'reordered questions', 'missing question',
'extra question', 'extra work', 'later assistant prose', 'changed plan chrome', 'missing ready prompt',
'cancel cursor', 'wrong submit action', 'quoted whole summary', 'no complete summary'])('%s does not submit', mutation => {
const p = observed();
let visible = summary();
if (mutation === 'different question') visible = visible.replace('persisted filters remain readable', 'project billing change');
if (mutation === 'different selected answer') visible = visible.replace('→ Version and validate (recommended)', '→ Store opaque filters');
if (mutation === 'reordered questions') {
const first = questions[0]!.question, second = questions[1]!.question;
visible = visible.replace(first.replaceAll('\n', '\n│ '), second.replaceAll('\n', '\n│ '));
}
if (mutation === 'missing question') visible = summary(2);
if (mutation === 'extra question') visible = visible.replace('Ready to submit', '● Delete the release branch?\n→ Yes\nReady to submit');
if (mutation === 'extra work') visible = visible.replace('Ready to submit', 'Also deploy everything.\nReady to submit');
if (mutation === 'later assistant prose') visible += '\n● Starting another question.';
if (mutation === 'changed plan chrome') visible = visible.replace('/tmp/owned-plan.md', '/tmp/foreign-plan.md');
if (mutation === 'missing ready prompt') visible = visible.replace('Ready to submit your answers?', '');
if (mutation === 'cancel cursor') visible = visible.replace('1.', '1.').replace('2. Cancel', '2. Cancel');
if (mutation === 'wrong submit action') visible = visible.replace('Submit answers', 'Approve and deploy');
if (mutation === 'quoted whole summary') visible = visible.split('\n').map(line => `> ${line}`).join('\n');
if (mutation === 'no complete summary') visible = captured.actualTruncatedTail;
expect(p.submit(visible)).toBeNull();
});
test('a bare Submit, a skipped tab or another session cannot borrow the observed packet', () => {
const t = transcript(), seen = new Set<string>();
expect(nextCeoPostureContinuation(summary(), t, 'HOLD SCOPE', selectedAt, seen, false)).toBeNull();
expect(nextCeoPostureContinuation(screen(0), t, 'HOLD SCOPE', selectedAt, seen, false)).toBe('question');
expect(nextCeoPostureContinuation(summary(), t, 'HOLD SCOPE', selectedAt, seen, true)).toBeNull();
const p = observed();
const foreign = transcript(); foreign.calls[0]!.sessionId = 'other-session';
expect(p.submit(summary(), foreign)).toBeNull();
});
test.each(['foreign session', 'foreign ID', 'different question', 'different option', 'answered', 'failed', 'new mode answer'])('late native %s refuses Submit', mutation => {
const p = observed(3, true), t = transcript(true);
const pending = t.calls[1]!;
if (mutation === 'foreign session') pending.sessionId = 'other-session';
if (mutation === 'foreign ID') pending.toolUseId = 'other-pending';
if (mutation === 'different question') pending.questions[0]!.question += ' Changed.';
if (mutation === 'different option') pending.questions[0]!.options[0]!.label = 'Deploy everything';
if (mutation === 'answered') pending.answered = true;
if (mutation === 'failed') pending.failed = true;
if (mutation === 'new mode answer') {
t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'SCOPE EXPANSION';
}
expect(p.submit(summary(), t)).toBeNull();
});
test('only later native assistant posture after a successful answer supplies coverage', () => {
const p = observed(3, true);
expect(p.submit()).toBe('submission');
expect(hasNativePostAnswerCeoPosture(p.t, 'HOLD SCOPE', /hold\s*scope/i, selectedAt)).toBe(false);
p.t.calls[1]!.answered = true;
p.t.calls[1]!.answeredAt = '2026-09-09T16:42:10.000Z';
p.t.calls[1]!.answers = Object.fromEntries(p.t.calls[1]!.questions.map(q => [q.question, q.options[0]!.label]));
p.t.calls[1]!.unansweredQuestionIndices = [];
p.t.assistantMessages.push({ sessionId: captured.modeCall.sessionId, timestamp: '2026-09-09T16:42:11.000Z',
text: 'HOLD SCOPE: keep the saved-view feature fixed and make its failure handling rigorous.' });
expect(hasNativePostAnswerCeoPosture(p.t, 'HOLD SCOPE', /hold\s*scope/i, selectedAt)).toBe(true);
});
test('the free test and fixture select only the mode-routing workflow', () => {
for (const file of ['test/ceo-barless-submit.test.ts', 'test/fixtures/ceo-barless-submit-ac.json']) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-ceo-mode-routing']);
}
});
});
+71
View File
@@ -0,0 +1,71 @@
import { describe, expect, test } from 'bun:test';
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-completion-handoff-l-calls.json';
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
describe('CEO closed review with zero unresolved decisions', () => {
test('the actual final handoff leaves all four independent issue and TODO decisions intact', () => {
const input = calls();
const original = structuredClone(input);
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
for (const call of input) {
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 4, review: 4, administrative: 1 });
expect(input.filter(c => /TODO/i.test(c.questions[0]!.header)).every(c =>
!isCeoCompletionHandoff(fingerprint(c)))).toBe(true);
expect(input).toEqual(original);
});
test('the offered manual action binds to the active native menu in either order', () => {
for (const reverse of [false, true]) {
const call = calls().at(-1)!;
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
const q = call.questions[0]!;
if (reverse) q.options.reverse();
const visible = `${q.header}\n${q.question}\n` + q.options.map((option, i) =>
`${i ? ' ' : ''} ${i + 1}. ${option.label}`).join('\n') +
'\nEnter to select · ↑/↓ to navigate · Esc to cancel';
const active = capturePlanCountQuestion(visible, new Set(), 0, false, call)!;
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
expect(pickCeoCompletionHandoff(fingerprint(call), { ...active, signature: 'other' })).toBeNull();
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
});
test('conditional, unresolved, substantive and unconfirmed variants are not handoffs', () => {
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 unresolved', '1 unresolved'); },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 unresolved decisions.', '0 unresolved decisions after fixing receipt assertions.'); },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete', 'is not complete'); },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('ceo-next-step-eng-review', 'ceo-security-finding'); },
c => { c.questions[0]!.header = 'Receipt gap'; },
c => { c.questions[0]!.options.push({ label: 'Add the missing happy-path assertions' }); },
c => { c.questions.push(calls()[4]!.questions[0]!); },
c => { c.failed = true; },
c => { c.answered = false; },
c => { c.unansweredQuestionIndices = [0]; },
];
for (const mutate of mutations) {
const call = calls().at(-1)!;
mutate(call);
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
const call = calls().at(-1)!;
call.answers = { [call.questions[0]!.question]: 'First add another test' };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
});
});
+352
View File
@@ -0,0 +1,352 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal,
nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-completion-handoff-m-call.json';
import nextStepCapture from './fixtures/ceo-handoff-n-calls.json';
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
const handoff = () => calls().at(-1)!;
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
function reanswer(call: NativePlanQuestionCall) {
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
return call;
}
describe('CEO completion described by a native navigation choice', () => {
test('the exact seven-call session preserves three setup and three finding decisions', () => {
let reviewStarted = false;
const counts = { setup: 0, review: 0, administrative: 0 };
const original = calls();
for (const call of original) {
const phase = planCountQuestionPhase(fingerprint(call), reviewStarted, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
reviewStarted = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 3, review: 3, administrative: 1 });
expect(original).toEqual(calls());
expect(original.slice(3, -1).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false, false]);
});
test('the active pending handoff selects the actual manual option in either order', () => {
for (const reverse of [false, true]) {
const call = handoff();
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
const q = call.questions[0]!;
if (reverse) q.options.reverse();
const screen = `${q.header}\n${q.question}\n 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const active = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
expect(active.nativeCall?.toolUseId).toBe(call.toolUseId);
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(active)).toBe(false);
expect(pickCeoCompletionHandoff(capturePlanCountQuestion(screen, new Set(), 0, false)!)).toBeNull();
expect(pickCeoCompletionHandoff(fingerprint(call), { ...active, signature: 'another:call' })).toBeNull();
}
});
test('completion placement is independent of the next-step wording and option order', () => {
const call = handoff();
const q = call.questions[0]!;
q.question = 'D9 — Next steps: The review is done. Where should we go next? <gstack-qid:plan-ceo-review-next-step>';
q.header = 'Next review';
q.options[0]!.description = 'Eng review is the required shipping gate.';
for (const description of [
'CEO review found 3 specification gaps (all resolved). Continue manually.',
'The CEO review identified gaps; all findings are resolved. Continue manually.',
'CEO review is complete with 0 unresolved decisions. Continue manually.',
]) {
q.options[1]!.description = description;
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
}
});
test('conditional, unfinished, quoted and non-CEO recaps cannot supply completion', () => {
for (const description of [
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved after adding tests).',
'Eng review is the required shipping gate. If CEO review found 3 gaps (all resolved), continue.',
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved); one gap remains.',
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). There is an unresolved test issue.',
'Eng review is the required shipping gate. The document says "CEO review found 3 gaps (all resolved)."',
'Eng review is the required shipping gate. Design review found 3 gaps (all resolved).',
'Eng review is the required shipping gate. CEO review found 3 gaps.',
'Eng review is the required shipping gate. CEO review found 3 specification gaps (not all resolved).',
'Eng review is the required shipping gate. CEO review did not find all gaps resolved.',
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). Also add a new test before proceeding.',
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). Please fix the new missing authorization check before proceeding.',
]) {
const call = handoff();
call.questions[0]!.options[0]!.description = description;
expect(isCeoCompletionHandoff(fingerprint(call)), description).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
}
});
test('native identity, completion, required gate and exclusively administrative choices remain necessary', () => {
for (const mutate of [
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Review complete only after fixing tests. What next? <gstack-qid:plan-ceo-review-next-step>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Should we finish reviewing? <gstack-qid:plan-ceo-review-next-step>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid:plan-ceo-security-issue>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = '<gstack-qid broken> ' + call.questions[0]!.question; },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO decision'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add another TODO'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[0]!.label += ' and fix the missing test'; },
(call: NativePlanQuestionCall) => { for (const option of call.questions[0]!.options) option.description = option.description?.replaceAll('required', 'optional'); },
(call: NativePlanQuestionCall) => { call.questions.push(calls()[3]!.questions[0]!); },
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
(call: NativePlanQuestionCall) => { call.failed = true; },
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
(call: NativePlanQuestionCall) => { call.answered = false; },
]) {
const call = handoff();
mutate(call);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
}
const addedWork = handoff();
addedWork.answers = { [addedWork.questions[0]!.question]: 'First add another payment test' };
expect(isCeoCompletionHandoff(fingerprint(addedWork))).toBe(false);
});
test('the actual report and Exit order permits only the administrative freshness exception', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-handoff-'));
const report = path.join(dir, 'plan.md');
try {
fs.writeFileSync(report, captured.report.content);
const reportAt = Date.parse(captured.report.successfulResult.timestamp) / 1000;
fs.utimesSync(report, reportAt, reportAt);
const transcript = { status: 'ready' as const, calls: calls(), assistantMessages: [],
planReadyRequests: structuredClone(captured.planReadyRequests) };
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
.map(call => `${call.sessionId}:${call.toolUseId}`));
const startedAt = Date.parse('2026-09-09T00:15:27Z');
expect(administrative.size).toBe(1);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
transcript.calls.push({ ...structuredClone(transcript.calls[3]!), toolUseId: 'new-test-obligation',
answeredAt: captured.calls.at(-1)!.answeredAt });
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
});
describe('native next-review navigation with a resolved CEO recap', () => {
const retryCalls = () => structuredClone(captured.distinctRetry.calls) as NativePlanQuestionCall[];
const retryHandoff = () => retryCalls().at(-1)!;
test('the captured retry preserves its four actual findings and the unchanged mechanical band', () => {
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
const original = retryCalls();
for (const call of original) {
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 4, review: 4, administrative: 1 });
expect(original).toEqual(retryCalls());
// The transcript contains four individual findings. The unasked dispatcher
// remedy remains a separate workflow-quality limitation, never a fifth call.
expect(original.slice(4, -1).every(call => !isCeoCompletionHandoff(fingerprint(call)))).toBe(true);
});
test('actual offered manual navigation still requires the matching pending native question', () => {
for (const reverse of [false, true]) {
const call = retryHandoff();
call.answered = false;
delete call.answers;
const q = call.questions[0]!;
if (reverse) q.options.reverse();
const screen = `${q.header}\n${q.question}\n 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const active = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
expect(active.nativeCall?.toolUseId).toBe(call.toolUseId);
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(active)).toBe(false);
expect(pickCeoCompletionHandoff(capturePlanCountQuestion(screen, new Set(), 0, false)!)).toBeNull();
}
});
test('partial, conditional, quoted or still-open recaps never establish this navigation boundary', () => {
for (const recap of [
'This CEO review resolved some security bugs.',
'This CEO review resolved most security bugs.',
'This CEO review resolved all but one security bugs.',
'This CEO review resolved two of three security bugs.',
'This CEO review only resolved the security bugs.',
'This CEO review did not resolve the security bugs.',
'If this CEO review resolved the security bugs, continue.',
'The document says "This CEO review resolved the security bugs."',
'This CEO review resolved the security bugs. One issue remains unresolved.',
'This CEO review resolved the security bugs; validation of that remedy is still pending.',
'This CEO review resolved the security bugs. Please add a new test first.',
'This CEO review will resolve the security bugs.',
]) {
const call = retryHandoff();
call.questions[0]!.options[0]!.description = 'Eng review is the required shipping gate. ' + recap;
expect(isCeoCompletionHandoff(fingerprint(call)), recap).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
}
for (const question of [
'Should we add a missing authorization test as the next step after this CEO review?',
'The CEO review did not finish. What is the next step after this CEO review?',
'Can you first fix the missing authorization check as the next step after this CEO review?',
'If the CEO review finishes, what is the next step after this CEO review?',
'Example: What is the next step after this CEO review?',
]) {
const call = retryHandoff();
call.questions[0]!.question = question + ' <gstack-qid:plan-ceo-next-step>';
reanswer(call);
expect(isCeoCompletionHandoff(fingerprint(call)), question).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call)), question).toBeNull();
}
for (const mutate of [
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Choose a fix for the missing test <gstack-qid:plan-ceo-next-step>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-next-step', 'plan-ceo-test-gap'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace(/ <gstack-qid:[^>]+>/, ''); },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add a missing receipt assertion'; },
(call: NativePlanQuestionCall) => { call.questions.push(retryCalls()[4]!.questions[0]!); },
(call: NativePlanQuestionCall) => { call.failed = true; },
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
]) {
const call = retryHandoff();
mutate(call);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
}
});
test('the final native report edit precedes handoff and still covers every real answer', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-retry-handoff-'));
const report = path.join(dir, 'plan.md');
try {
fs.writeFileSync(report, captured.distinctRetry.reportContent);
const reportAt = Date.parse(captured.distinctRetry.reportUpdate.at(-1)!.timestamp) / 1000;
fs.utimesSync(report, reportAt, reportAt);
const transcript = { status: 'ready' as const, calls: retryCalls(), assistantMessages: [],
planReadyRequests: structuredClone(captured.distinctRetry.planReadyRequests) };
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
.map(call => `${call.sessionId}:${call.toolUseId}`));
const startedAt = Date.parse('2026-09-09T00:23:30Z');
expect(administrative.size).toBe(1);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
transcript.calls.push({ ...structuredClone(transcript.calls[4]!), toolUseId: 'new-independent-finding',
answeredAt: transcript.calls.at(-1)!.answeredAt });
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
});
describe('CEO completed next-step identity in native option order', () => {
const input = () => structuredClone(nextStepCapture.calls) as NativePlanQuestionCall[];
const actual = () => input().at(-1)!;
test('the complete native sequence retains two setup and four real issue decisions', () => {
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
const native = input();
const original = structuredClone(native);
for (const call of native) {
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 2, review: 4, administrative: 1 });
expect(native).toEqual(original);
});
test('only the positively bound pending menu selects its offered manual action', () => {
for (const reverse of [false, true]) {
const call = actual();
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
if (reverse) call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'other:call' })).toBeNull();
}
});
test('the observed identity cannot excuse unfinished work, a finding or a malformed native call', () => {
for (const mutate of [
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete only after fixing authorization.'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete. One issue remains unresolved.'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete. Please fix the missing authorization test.'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'Should we add a missing authorization test before the next review?'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'We should fix the missing authorization test before the next review.'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'We should fix the missing authorization test. What next?'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('required shipping gate', 'optional review'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('ceo-plan-next-steps', 'ceo-plan-test-gap'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid:ceo-plan-next-steps>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.description = 'Proceed to fix the missing authorization test before Eng review.'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[0]!.label += ' and add a missing test'; },
(call: NativePlanQuestionCall) => { call.questions.push(input()[2]!.questions[0]!); },
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
]) {
const call = actual();
mutate(call);
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
}
for (const mutate of [
(call: NativePlanQuestionCall) => { call.failed = true; },
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
(call: NativePlanQuestionCall) => { call.answers = {}; },
(call: NativePlanQuestionCall) => { call.answers = { [call.questions[0]!.question]: 'Build another feature' }; },
]) {
const call = actual();
mutate(call);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
});
test('the actual report precedes handoff but the captured absent Exit remains incomplete', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-next-step-'));
const report = path.join(dir, 'plan.md');
try {
fs.writeFileSync(report, nextStepCapture.report.content);
const written = Date.parse(nextStepCapture.report.successfulUpdateAt) / 1000;
fs.utimesSync(report, written, written);
const calls = input();
expect(Date.parse(calls.at(-2)!.answeredAt!)).toBeLessThan(written * 1000);
expect(Date.parse(calls.at(-1)!.answeredAt!)).toBeGreaterThan(written * 1000);
const transcript = { status: 'ready' as const, calls, assistantMessages: [],
planReadyRequests: structuredClone(nextStepCapture.planReadyRequests) };
const admin = new Set([fingerprint(calls.at(-1)!).signature]);
expect(hasNativePlanTerminal(transcript, report, Date.parse('2026-09-09T01:06:22Z'), 'plan_ready', admin)).toBe(false);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
});
+270
View File
@@ -0,0 +1,270 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-completion-handoff-o-call.json';
import capturedQ from './fixtures/ceo-completion-handoff-q-call.json';
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
const handoff = () => calls().at(-1)!;
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
const reanswer = (call: NativePlanQuestionCall) => {
const question = call.questions[0]!;
call.answers = { [question.question]: question.options[0]!.label };
return call;
};
describe('closed CEO navigation with the native review-prefixed identity', () => {
test('the actual six-call sequence preserves setup and both independent findings', () => {
const original = calls();
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
for (const call of original) {
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 3, review: 2, administrative: 1 });
expect(original).toEqual(calls());
expect(original.slice(3, 5).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false]);
});
test('the offered manual action needs the matching pending native call in either order', () => {
for (const reverse of [false, true]) {
const call = handoff();
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
if (reverse) call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'foreign:call' })).toBeNull();
}
expect(pickCeoCompletionHandoff(fingerprint(handoff()))).toBeNull();
});
test('closed navigation semantics are shared across the bounded review identity family', () => {
for (const id of ['ceo-review-next-step', 'ceo-review-next-steps', 'ceo-review-next-review', 'ceo-plan-next-steps']) {
for (const completion of ['done', 'complete', 'cleared']) {
const call = handoff();
call.questions[0]!.question = call.questions[0]!.question
.replace('ceo-review-next-steps', id).replace('CEO review done.', `CEO review ${completion}.`);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
}
}
const sequencing = handoff();
sequencing.questions[0]!.options[1]!.description = 'Once implementation is finished, run /plan-eng-review. After Eng review is complete, proceed to shipping.';
expect(isCeoCompletionHandoff(fingerprint(sequencing))).toBe(true);
});
test('a completed heading cannot conceal unresolved work or a substantive question', () => {
for (const text of [
'CEO review is not done. What\'s next?',
'CEO review done only after fixing the missing authorization test. What\'s next?',
'CEO review done. Should we add a missing authorization test before Eng?',
'CEO review done. We should fix the missing authorization test. What\'s next?',
'CEO review done. Do you want me to fix the missing authorization test? What\'s next?',
'CEO review done. One contrast issue remains. What\'s next?',
'CEO review done. Validation is still pending. What\'s next?',
'CEO review done. Not all findings are resolved. What\'s next?',
'CEO review done. One test issue is still open. What\'s next?',
'CEO review done. There are not 0 unresolved decisions. What\'s next?',
'CEO review done. If the tests pass, what\'s next?',
'CEO review done. What\'s next? Once the tests pass, all decisions are resolved.',
'CEO review done. What\'s next? After the authorization tests pass, the review is complete.',
'CEO review done. What\'s next? The review is complete when authorization tests pass.',
'CEO review done. What\'s next? Once the tests pass, all decisions will be resolved.',
'CEO review done. What\'s next? All findings become resolved after the tests pass.',
'Example: CEO review done. What\'s next?',
]) {
const call = handoff();
call.questions[0]!.question = call.questions[0]!.question.replace("CEO review done. What's next?", text);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call))), text).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call)), text).toBeNull();
}
for (const description of [
'Proceed to fix the missing authorization test before Eng.',
'The contrast gap remains unresolved; handle it manually.',
'Please add a new regression test before implementation.',
'We could add a missing regression test before Eng.',
'Do you want to add a new test before the next review?',
]) {
const call = handoff();
call.questions[0]!.options[1]!.description = description;
expect(isCeoCompletionHandoff(fingerprint(call)), description).toBe(false);
}
});
test('failed, partial, malformed, unrelated or mixed native calls remain substantive', () => {
for (const mutate of [
(call: NativePlanQuestionCall) => { call.failed = true; },
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
(call: NativePlanQuestionCall) => { call.questions.push(calls()[3]!.questions[0]!); },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'Test gap'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('ceo-review-next-steps', 'ceo-review-test-gap'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('required shipping gate', 'optional review'); },
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add a missing receipt assertion'; },
]) {
const call = handoff();
mutate(call);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
}
const freeform = handoff();
freeform.answers![freeform.questions[0]!.question] = 'Please add another test first';
expect(isCeoCompletionHandoff(fingerprint(freeform))).toBe(false);
});
test('actual report edits precede the handoff and retain the strict native Exit and freshness checks', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-closed-navigation-'));
const report = path.join(dir, 'plan.md');
try {
fs.writeFileSync(report, captured.reportContent);
const reportAt = Date.parse(captured.reportUpdate.at(-1)!.timestamp) / 1000;
fs.utimesSync(report, reportAt, reportAt);
const transcript = { status: 'ready' as const, calls: calls(), assistantMessages: [],
planReadyRequests: structuredClone(captured.planReadyRequests) };
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
.map(call => `${call.sessionId}:${call.toolUseId}`));
const startedAt = Date.parse('2026-09-09T01:46:12Z');
expect(administrative.size).toBe(1);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
transcript.planReadyRequests = [];
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
transcript.planReadyRequests = structuredClone(captured.planReadyRequests);
transcript.calls.push({ ...structuredClone(transcript.calls[3]!), toolUseId: 'new-real-finding',
answeredAt: transcript.calls.at(-1)!.answeredAt });
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
});
describe('CEO completion recap after native project metadata', () => {
const qCalls = () => structuredClone(capturedQ.calls) as NativePlanQuestionCall[];
const qHandoff = () => qCalls().at(-1)!;
test('the exact Q sequence keeps all three substantive calls and four setup calls', () => {
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
for (const call of qCalls()) {
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
if (phase.administrative) counts.administrative++;
else if (phase.preReview) counts.setup++;
else counts.review++;
}
expect(counts).toEqual({ setup: 4, review: 3, administrative: 1 });
expect(qCalls().slice(4, 7).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false, false]);
expect(isCeoCompletionHandoff(fingerprint(qHandoff()))).toBe(true);
});
test('only the current offered manual option is selected, including reordered choices', () => {
for (const reverse of [false, true]) {
const call = qHandoff();
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
if (reverse) call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'foreign:call' })).toBeNull();
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
expect(pickCeoCompletionHandoff(fingerprint(qHandoff()))).toBeNull();
});
test('unconditional line recaps support ordinary completion wording and Eng sequencing', () => {
for (const state of ['done and clear', 'done', 'complete', 'cleared']) {
const call = qHandoff();
call.questions[0]!.question = call.questions[0]!.question.replace('done and clear', state);
call.questions[0]!.options[1]!.description = 'Once implementation is finished, run /plan-eng-review. After Eng review is complete, proceed to shipping.';
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
}
});
test('the recap cannot hide contradictory, conditional, quoted or new work in question or choices', () => {
for (const extra of [
'CEO review is not complete.', 'The review remains incomplete.', 'Not all decisions are resolved.',
'One test gap remains.', 'Validation is still pending.', 'There are unresolved findings.',
'Once tests pass, the CEO review will be complete.', 'All decisions resolved after tests pass.',
'We should fix a missing authorization test.', 'We could repair a missing authorization check.',
'Repair the missing authorization test.', 'Recommendation: repair the missing authorization test.',
'We may repair the missing authorization test.', 'We might fix the missing authorization test.',
'Proceed to add a new regression.', 'Do you want to add a missing test?',
'```text\nCEO review is complete.', '> CEO review is complete.', 'Example: CEO review is complete.',
]) {
for (const target of ['question', 'description']) {
const call = qHandoff();
if (target === 'question') call.questions[0]!.question += `\n${extra}`;
else call.questions[0]!.options[1]!.description += ` ${extra}`;
expect(isCeoCompletionHandoff(fingerprint(reanswer(call))), `${target}: ${extra}`).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call)), `${target}: ${extra}`).toBeNull();
}
}
for (const first of [
'Should we add a missing authorization test as the next step after this CEO review?',
'The CEO review did not finish. What is next after this CEO review?',
'Can you first fix authorization? What is next after this CEO review?',
]) {
const call = qHandoff();
call.questions[0]!.question = call.questions[0]!.question.replace("What's next after this CEO review?", first);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
}
});
test('native failures, mixed choices, absent gates and source copies cannot become administrative', () => {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.questions.push(qCalls()[4]!.questions[0]!); },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Missing tests'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-next-step', 'plan-ceo-new-test'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate', 'optional check'); c.questions[0]!.options[0]!.description = 'Optional check.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix the missing assertion'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('ELI10:', ' ELI10:'); },
]) {
const call = qHandoff(); mutate(call);
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
}
const call = qHandoff();
call.answers![call.questions[0]!.question] = 'Please fix another gap first';
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
});
test('actual full report and Exit chronology retain last substantive-answer freshness', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-metadata-navigation-'));
const report = path.join(dir, 'plan.md');
try {
fs.writeFileSync(report, capturedQ.reportContent);
const reportAt = Date.parse(capturedQ.reportAt) / 1000;
fs.utimesSync(report, reportAt, reportAt);
const transcript = { status: 'ready' as const, calls: qCalls(), assistantMessages: [], planReadyRequests: structuredClone(capturedQ.planReadyRequests) };
const administrative = new Set([`${qHandoff().sessionId}:${qHandoff().toolUseId}`]);
const start = Date.parse('2026-09-09T03:25:54Z');
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', administrative)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', administrative)).toBe(false);
transcript.planReadyRequests = structuredClone(capturedQ.planReadyRequests);
fs.utimesSync(report, start / 1000, start / 1000);
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', administrative)).toBe(false);
fs.writeFileSync(report, 'Incomplete plan');
fs.utimesSync(report, reportAt, reportAt);
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', administrative)).toBe(false);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
});
+949
View File
@@ -0,0 +1,949 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import captures from './fixtures/ceo-completion-handoff-calls.json';
import currentHandoffs from './fixtures/ceo-completion-handoff-j-calls.json';
import kHandoffs from './fixtures/ceo-completion-handoff-k-calls.json';
import rCalls from './fixtures/ceo-completion-handoff-r-calls.json';
import tHandoff from './fixtures/ceo-completion-handoff-t-call.json';
import uHandoff from './fixtures/ceo-completion-handoff-u-call.json';
import vHandoff from './fixtures/ceo-completion-handoff-v-call.json';
import wHandoff from './fixtures/ceo-completion-handoff-w-call.json';
type CapturedCall = typeof captures.cases[number]['calls'][number];
function nativeCall(record: CapturedCall, sessionId = 'native-capture'): NativePlanQuestionCall {
return {
sessionId, toolUseId: record.toolUseId, answered: true, failed: false,
questions: [{ header: record.header, question: record.question,
options: record.options.map(label => ({ label })), multiSelect: false }],
answers: { [record.question]: record.answer }, unansweredQuestionIndices: [],
};
}
const handoff = () => nativeCall(captures.cases[0]!.calls.at(-1)!);
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
describe('W unconditional CLEAR recap and required Eng pronoun navigation', () => {
const actual = () => structuredClone(wHandoff.calls.at(-1)!) as NativePlanQuestionCall;
const pending = (call: NativePlanQuestionCall) => {
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
return fingerprint(copy);
};
const answer = (call: NativePlanQuestionCall) => {
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
return call;
};
test('exact seven calls retain two issue decisions and select the offered manual action', () => {
const calls = structuredClone(wHandoff.calls) as NativePlanQuestionCall[];
expect(replay(calls, false, ceoFirstReviewAUQ))
.toMatchObject({ step0Count: 4, reviewCount: 2, administrativeCount: 1 });
expect(isCeoCompletionHandoff(fingerprint(actual()))).toBe(true);
expect(pickCeoCompletionHandoff(pending(actual()))).toBe(2);
expect(pickCeoCompletionHandoff(fingerprint(actual()))).toBeNull();
expect(calls).toEqual(wHandoff.calls);
});
test('case, gap count and pure navigation option order do not change the meaning', () => {
const call = actual(); const q = call.questions[0]!;
q.question = q.question.toLowerCase().replace(' — ', ' - ');
q.options[0]!.description = q.options[0]!.description!.replace('2 assertion gaps', '12 assertion gaps');
q.options.reverse(); answer(call);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(true);
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
});
test('conditional, negated, quoted or additional question text is not a closed handoff', () => {
const source = actual().questions[0]!.question;
for (const question of [
source.replace('is CLEAR.', 'is not CLEAR.'), source.replace('is CLEAR.', 'will be CLEAR.'),
source.replace('is CLEAR.', 'is CLEAR after tests pass.'), 'Once ' + source,
source.replace('required shipping gate', 'optional shipping check'),
source.replace('Eng review', 'Design review'), source.replace('run it next?', 'repair its findings next?'),
source + ' Remove the failing test.', source + ' Should we change the error contract?',
'> ' + source, 'Example: ' + source, '`' + source + '`',
source + ' <gstack-qid:ceo-plan-next-steps>',
]) {
const call = actual(); call.questions[0]!.question = question; answer(call);
expect(isCeoCompletionHandoff(fingerprint(call)), question).toBe(false);
expect(pickCeoCompletionHandoff(pending(call)), question).toBeNull();
}
});
test('every description sentence must be closed navigation, including unknown action verbs', () => {
for (const extra of [
'Delete the authorization test.', 'Grant access to all accounts.', 'Repair the missing assertion.',
'One gap remains unresolved.', 'The CEO review is CLEAR only if we change the contract.',
'The CEO review will be CLEAR after another fix.', 'Should we add another test?',
'Quoted source: CEO review is CLEAR.',
]) {
for (const index of [0, 1]) {
const call = actual(); call.questions[0]!.options[index]!.description += ' ' + extra;
expect(isCeoCompletionHandoff(fingerprint(call)), extra).toBe(false);
expect(pickCeoCompletionHandoff(pending(call)), extra).toBeNull();
}
}
for (const description of ['', 'This CEO review held scope and resolved some assertion gaps — eng review verifies the test structure is sound.',
'This CEO review held scope and resolved 2 assertion gaps after changing the contract — eng review verifies the test structure is sound.']) {
const call = actual(); call.questions[0]!.options[0]!.description = description;
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
}
});
test('native identity, complete answers, Eng/manual choices and a single question remain required', () => {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.failed; },
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix another issue' }; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'New finding'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = 'Run /plan-design-review'; answer(c); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix remaining issues manually'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[1]!)); },
]) { const call = actual(); mutate(call); expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false); }
expect(pickCeoCompletionHandoff({ ...pending(actual()), signature: 'foreign:call' })).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(actual()), nativeCall: undefined })).toBeNull();
});
test('controlled report time excludes the handoff but still rejects a later real issue answer', () => {
expect(wHandoff.provenance.reportMtimeMs).toBeNull(); // No historical filesystem-time claim.
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-w-handoff-'));
try {
const calls = structuredClone(wHandoff.calls) as NativePlanQuestionCall[];
const issueAt = Date.parse(calls.at(-2)!.answeredAt!);
const navigationAt = Date.parse(calls.at(-1)!.answeredAt!);
const syntheticWritten = Math.floor((issueAt + navigationAt) / 2);
const file = path.join(dir, 'report.md'); fs.writeFileSync(file, wHandoff.reportContent);
fs.utimesSync(file, syntheticWritten / 1000, syntheticWritten / 1000);
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: wHandoff.planReadyRequests };
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
const start = Date.parse('2026-09-09T09:28:55Z');
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', new Set())).toBe(false);
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', admin)).toBe(true);
calls.at(-2)!.answeredAt = new Date(syntheticWritten + 1000).toISOString();
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', admin)).toBe(false);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
});
describe('V closed CEO recap with a resolved-gap count', () => {
const actual = () => structuredClone(vHandoff.calls.at(-1)!) as NativePlanQuestionCall;
const pending = (call: NativePlanQuestionCall) => {
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
return fingerprint(call);
};
test('actual navigation stays outside the two issue decisions and selects manual', () => {
expect(replay(structuredClone(vHandoff.calls) as NativePlanQuestionCall[], false, ceoFirstReviewAUQ))
.toMatchObject({ step0Count: 3, reviewCount: 2, administrativeCount: 1 });
expect(isCeoCompletionHandoff(fingerprint(actual()))).toBe(true);
expect(pickCeoCompletionHandoff(pending(actual())) ?? 1).toBe(2);
const reordered = actual(); reordered.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(pending(reordered))).toBe(1);
});
test('the actual report is fresh after issue decisions but before this navigation', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-v-handoff-'));
try {
const report = path.join(dir, 'report.md'); fs.writeFileSync(report, vHandoff.reportContent);
const written = vHandoff.provenance.reportMtimeMs / 1000; fs.utimesSync(report, written, written);
const calls = structuredClone(vHandoff.calls) as NativePlanQuestionCall[];
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: vHandoff.planReadyRequests };
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
const start = Date.parse('2026-09-09T08:42:53Z');
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', admin)).toBe(true);
calls.at(-2)!.answeredAt = new Date(vHandoff.provenance.reportMtimeMs + 1).toISOString();
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', admin)).toBe(false);
} finally { fs.rmSync(dir, {recursive:true,force:true}); }
});
test('the new recap cannot hide incomplete review, another remedy or altered gate', () => {
const edits: Array<(call: NativePlanQuestionCall) => void> = [
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 critical gaps','1 critical gap'); },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('gaps resolved','gaps unresolved'); },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete','is complete only after tests pass'); },
c => { c.questions[0]!.question += ' Repair the missing authorization test.'; },
c => { c.questions[0]!.question += ' Should we remove the owner check?'; },
c => { c.questions[0]!.options[0]!.description += ' Delete the failing test.'; },
c => { c.questions[0]!.options[1]!.description = 'The CEO review is NOT CLEARED until its gaps are resolved.'; },
c => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate','optional review'); },
c => { c.questions[0]!.header = 'New finding'; },
c => { c.questions[0]!.options[1]!.label = 'Implement a new feature'; },
];
for (const edit of edits) {
const c = actual(); edit(c); c.answers = {[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
}
});
test('completed identity, offered answer and single question remain required', () => {
for (const edit of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
(c: NativePlanQuestionCall) => { c.answers = {[c.questions[0]!.question]:'Add a new task'}; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
]) { const c=actual();edit(c);expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
expect(pickCeoCompletionHandoff({...pending(actual()),signature:'foreign:call'})).toBeNull();
expect(pickCeoCompletionHandoff({...pending(actual()),nativeCall:undefined})).toBeNull();
});
});
describe('U completed CEO metadata navigation with scoped review explanations', () => {
const captured = () => structuredClone(uHandoff.calls.at(-1)!) as NativePlanQuestionCall;
const pending = (call: NativePlanQuestionCall) => {
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
return fingerprint(copy);
};
test('the actual six-call stream retains two issues and selects the offered manual stop', () => {
const calls = structuredClone(uHandoff.calls) as NativePlanQuestionCall[];
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 2, administrativeCount: 1 });
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
expect(calls).toEqual(uHandoff.calls);
});
test('native identity, completed answer and real option order remain required', () => {
const call = captured(); call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign:call' })).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.failed; },
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix one more issue first' }; },
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
});
test('unfinished, conditional, quoted and additional-work descriptions remain substantive', () => {
for (const extra of [
'Delete the failing regression test before Eng.', 'Remove the owner check before Eng.',
'Change the guarantee to permit old results.', 'Rewrite the acceptance criteria before shipping.',
'Repair the missing authorization test.', 'We may repair the missing authorization test.',
'All findings become resolved after the tests pass.', 'There is an outstanding authorization gap.',
'Should we add another test before Eng?', 'Stakes if we pick wrong: delete the owner check.',
'No UI scope was detected, so the CEO review is not complete.',
]) {
const c = captured(); c.questions[0]!.options[1]!.description += ' ' + extra;
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
}
for (const [from, to] of [
['The CEO review is done.', 'The CEO review is not done.'],
['The CEO review is done.', 'The CEO review is done if tests pass.'],
['Two assertion spec gaps were caught and resolved.', 'Not all assertion spec gaps were resolved.'],
['Two assertion spec gaps were caught and resolved.', 'Two assertion spec gaps remain unresolved.'],
['No UI scope was detected, so a design review is not needed.', 'The CEO review is not needed.'],
['No UI scope was detected, so a design review is not needed.', 'No UI scope was detected, so a design review is not complete.'],
['Stakes if we pick wrong:', 'The CEO review is complete only if we pick correctly:'],
]) {
const c = captured(); const q = c.questions[0]!; const old = q.question; q.question = old.replace(from!, to!);
c.answers = { [q.question]: c.answers![old]! };
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
}
for (const prefix of ['> ', '```text\n', 'Example: ']) {
const c = captured(); c.questions[0]!.options[1]!.description = prefix + c.questions[0]!.options[1]!.description;
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we fix the missing authorization check?'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Authorization gap'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add a new TODO' }); },
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); expect(pickCeoCompletionHandoff(pending(c))).toBeNull(); }
});
test('metadata headings cannot shelter an extra obligation or conditional completion', () => {
for (const extra of ['Delete the owner check.', 'Remove the failing regression.', 'Change the guarantee.',
'Rewrite the acceptance criteria.', 'All decisions are resolved after the tests pass.',
'Should we approve one more issue?', 'The CEO review is not complete.']) {
const c = captured(); const q = c.questions[0]!; const old = q.question;
q.question += ' ' + extra; c.answers = { [q.question]: c.answers![old]! };
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
}
});
test('the retained pending Exit and report still require fresh substantive decisions', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-u-handoff-')); const file = path.join(dir, 'plan.md');
try {
fs.writeFileSync(file, uHandoff.reportContent);
fs.utimesSync(file, uHandoff.reportAtMs / 1000, uHandoff.reportAtMs / 1000);
const calls = structuredClone(uHandoff.calls) as NativePlanQuestionCall[];
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: [{
sessionId: uHandoff.pendingExit.sessionId, toolUseId: uHandoff.pendingExit.toolUseId,
timestamp: uHandoff.pendingExit.timestamp, failed: false, source: 'pre_tool_use' as const,
}] };
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
transcript.planReadyRequests[0]!.sessionId = 'foreign-session';
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.sessionId = uHandoff.pendingExit.sessionId;
calls[3]!.answeredAt = new Date(uHandoff.reportAtMs + 1000).toISOString();
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
});
describe('T completed CEO next-review navigation', () => {
const captured = () => structuredClone(tHandoff.calls.at(-1)!) as NativePlanQuestionCall;
const pending = (call: NativePlanQuestionCall) => {
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
return fingerprint(copy);
};
test('the actual nine-call stream retains five issues and selects the offered manual stop', () => {
const calls = structuredClone(tHandoff.calls) as NativePlanQuestionCall[];
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 5, administrativeCount: 1 });
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
expect(calls).toEqual(tHandoff.calls);
});
test('native identity, completed answer and real option order remain required', () => {
const call = captured(); call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign:call' })).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix one more issue first' }; },
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
});
test('unfinished, conditional, quoted and additional-work descriptions remain substantive', () => {
for (const extra of [
'Delete the failing regression test before Eng.', 'Remove the owner check before Eng.',
'Change the guarantee to permit old results.', 'Rewrite the acceptance criteria before shipping.',
'Repair the missing authorization test.', 'We may repair the missing authorization test.',
'All findings become resolved after the tests pass.', 'There is an outstanding authorization gap.',
'Should we add another test before Eng?',
]) {
const c = captured(); c.questions[0]!.options[1]!.description += ' ' + extra;
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
}
for (const replacement of ['resolved some findings', 'did not resolve all findings', 'will resolve all findings after tests pass']) {
const c = captured(); c.questions[0]!.options[1]!.description = c.questions[0]!.options[1]!.description!.replace('resolved all findings', replacement);
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
}
for (const prefix of ['> ', '```text\n', 'Example: ']) {
const c = captured(); c.questions[0]!.options[1]!.description = prefix + c.questions[0]!.options[1]!.description;
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we fix the missing authorization check?'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Authorization gap'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add a new TODO' }); },
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); expect(pickCeoCompletionHandoff(pending(c))).toBeNull(); }
});
test('the retained pending Exit and report still require fresh substantive decisions', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-t-handoff-')); const file = path.join(dir, 'plan.md');
try {
fs.writeFileSync(file, tHandoff.reportContent);
fs.utimesSync(file, tHandoff.reportAtMs / 1000, tHandoff.reportAtMs / 1000);
const calls = structuredClone(tHandoff.calls) as NativePlanQuestionCall[];
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: [{
sessionId: tHandoff.pendingExit.sessionId, toolUseId: tHandoff.pendingExit.toolUseId,
timestamp: tHandoff.pendingExit.timestamp, failed: false, source: 'pre_tool_use' as const,
}] };
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
transcript.planReadyRequests[0]!.sessionId = 'foreign-session';
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.sessionId = tHandoff.pendingExit.sessionId;
calls[3]!.answeredAt = new Date(tHandoff.reportAtMs + 1000).toISOString();
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
});
describe('native direct Eng/manual handoff with described CEO closure', () => {
const captured = () => structuredClone(rCalls.at(-1)!) as NativePlanQuestionCall;
const pending = (call: NativePlanQuestionCall) => {
const copy = structuredClone(call); copy.answered = false; delete copy.answers;
delete copy.unansweredQuestionIndices;
return fingerprint(copy);
};
test('actual R calls retain zero findings and choose offered manual instead of starting Eng', () => {
const calls = structuredClone(rCalls) as NativePlanQuestionCall[];
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 0, administrativeCount: 1, reviewStarted: true });
expect(replay(calls, false, ceoFirstReviewAUQ).reviewCount).toBeLessThan(2); // Existing paired floor still fails.
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
expect(calls).toEqual(rCalls);
});
test('manual choice follows real option order and still requires pending native identity', () => {
const call = captured(); call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign-call' })).toBeNull();
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
call.failed = true;
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
});
test('same native menu retains every incomplete, conditional, quoted or substantive obligation', () => {
const changes: Array<(c: NativePlanQuestionCall) => void> = [
c => { c.questions[0]!.question = 'Should we fix the missing authorization test before the next review?'; },
c => { c.questions[0]!.question += ' First repair the missing assertion.'; },
c => { c.questions[0]!.question = 'The review did not finish. ' + c.questions[0]!.question; },
c => { c.questions[0]!.header = 'Authorization gap'; },
c => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
c => { c.questions[0]!.options[1]!.label = 'Skip'; },
c => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
c => { c.questions[0]!.options.push({ ...c.questions[0]!.options[1]! }); },
c => { c.questions[0]!.options.push({ label: 'Run /plan-design-review' }); },
c => { c.questions[0]!.options[1]!.description = 'The CEO review is not clear.'; },
c => { c.questions[0]!.options[1]!.description = 'The CEO review remains incomplete.'; },
c => { c.questions[0]!.options[1]!.description = 'The CEO review is clear once tests pass.'; },
c => { c.questions[0]!.options[1]!.description = 'Once tests pass, the CEO review will be clear.'; },
c => { c.questions[0]!.options[1]!.description += ' All findings become resolved after tests pass.'; },
c => { c.questions[0]!.options[1]!.description += ' The contrast gap remains unresolved.'; },
c => { c.questions[0]!.options[1]!.description += ' Not all decisions are resolved.'; },
c => { c.questions[0]!.options[1]!.description += ' Repair the missing authorization test.'; },
c => { c.questions[0]!.options[1]!.description += ' Recommendation: repair the missing assertion.'; },
c => { c.questions[0]!.options[1]!.description += ' We may repair the missing assertion.'; },
c => { c.questions[0]!.options[1]!.description += ' We must add the authorization test.'; },
c => { c.questions[0]!.options[1]!.description += ' Delete the failing regression test before Eng.'; },
c => { c.questions[0]!.options[1]!.description += ' Remove the owner check before Eng.'; },
c => { c.questions[0]!.options[1]!.description += ' Change the guarantee to permit old results.'; },
c => { c.questions[0]!.options[1]!.description += ' Rewrite the acceptance criteria before shipping.'; },
c => { c.questions[0]!.options[1]!.description += ' Do you want me to fix the missing test?'; },
c => { c.questions[0]!.options[1]!.description = 'Example: The CEO review is clear.'; },
c => { c.questions[0]!.options[1]!.description = '> The CEO review is clear.'; },
c => { c.questions[0]!.options[1]!.description = '```text\nThe CEO review is clear.'; },
];
for (const change of changes) {
const call = captured(); change(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
}
});
test('an unconditional completed recap permits next Eng sequencing but no failed or free-form answer', () => {
const call = captured();
call.questions[0]!.options[1]!.description = 'The CEO review is complete. Run /plan-eng-review after implementation and before shipping.';
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(true);
expect(pickCeoCompletionHandoff(pending(call))).toBe(2);
call.answers = { [call.questions[0]!.question]: 'First fix the missing receipt assertion' };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
call.unansweredQuestionIndices = [0];
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
});
});
function replay(calls: NativePlanQuestionCall[], reviewStarted = true, firstReview = (_fp: ReturnType<typeof fingerprint>) => true) {
const counts = { step0Count: 0, reviewCount: 0, administrativeCount: 0 };
const classifications = [];
for (const call of calls) {
const fp = fingerprint(call);
const phase = planCountQuestionPhase(fp, reviewStarted, ceoStep0Boundary,
// A completion summary can mention defects; even a broad positive
// first-finding predicate must not promote a handoff into coverage.
firstReview, undefined, isCeoCompletionHandoff);
if (phase.administrative) counts.administrativeCount++;
else if (phase.preReview) counts.step0Count++;
else counts.reviewCount++;
reviewStarted = phase.reviewStarted;
classifications.push(phase);
}
return { ...counts, reviewStarted, classifications };
}
describe('CEO completion handoff classification and selection', () => {
test('captured first attempts keep every finding/TODO and exclude only the handoff; substantive retry still fails its band', () => {
for (const scenario of captures.cases) {
const calls = scenario.calls.map(c => nativeCall(c, scenario.sessionId));
const original = structuredClone(calls);
const result = replay(calls);
expect(result.reviewCount).toBe(scenario.expectedReviewCount);
expect(result.administrativeCount).toBe(scenario.name === 'five-retry' ? 0 : 1);
expect(result.step0Count).toBe(0);
expect(calls).toEqual(original); // Classification never discards or rewrites native evidence.
for (const [i, call] of calls.entries()) {
if (/TODO/i.test(call.questions[0]!.header)) expect(result.classifications[i]!.administrative).toBeUndefined();
}
}
expect(replay(captures.cases[2]!.calls.map(c => nativeCall(c))).reviewCount).toBeGreaterThan(7);
});
test('handoff-only replay adds no findings or setup and cannot establish a first finding', () => {
const result = replay([handoff()], false);
expect(result).toMatchObject({ step0Count: 0, reviewCount: 0, administrativeCount: 1, reviewStarted: false });
expect(result.classifications[0]).toEqual({ preReview: false, reviewStarted: false, administrative: 'completion-handoff' });
});
test('manual/done action is selected in either option order only while the matching native question is pending', () => {
for (const reverse of [false, true]) {
const call = handoff(); call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
if (reverse) call.questions[0]!.options.reverse();
const fp = fingerprint(call);
expect(pickCeoCompletionHandoff(fp)).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(fp)).toBe(false);
}
expect(pickCeoCompletionHandoff(fingerprint(handoff()))).toBeNull();
});
test('substantive choices mentioning another review retain the normal choice and finding count', () => {
const call = nativeCall(captures.cases[0]!.calls[0]!);
call.questions[0]!.question += ' Run /plan-eng-review next after deciding how to fix this issue.';
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(replay([call]).reviewCount).toBe(1);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
});
test('mixed packets and unknown action choices are not classified as an administrative handoff', () => {
const mixed = handoff();
const finding = nativeCall(captures.cases[0]!.calls[0]!);
mixed.questions.push(finding.questions[0]!);
mixed.answers = { ...mixed.answers, ...finding.answers };
expect(isCeoCompletionHandoff(fingerprint(mixed))).toBe(false);
expect(replay([mixed]).reviewCount).toBe(1);
mixed.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(mixed))).toBeNull();
const unknown = handoff(); unknown.questions[0]!.options.push({ label: 'Add another payment test before continuing' });
expect(isCeoCompletionHandoff(fingerprint(unknown))).toBe(false);
});
test('unknown identities and generic skip choices remain counted', () => {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:plan-ceo-security-finding>'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Test gap'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Skip'; },
]) {
const call = handoff(); mutate(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(replay([call]).reviewCount).toBe(1);
}
const call = handoff(); call.answered = false;
const mismatched = { ...fingerprint(call), signature: 'another-native-call' };
expect(pickCeoCompletionHandoff(mismatched)).toBeNull();
});
test('pending, failed, partial and free-form answers never create an exclusion', () => {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'First add a refund test' }; },
]) {
const call = handoff(); mutate(call);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
});
test('UI-only and unrelated pending metadata cannot steer the active menu', () => {
const pending = handoff(); pending.answered = false; delete pending.answers;
const q = pending.questions[0]!;
const active = `${q.header}\n${q.question}\n 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const bound = capturePlanCountQuestion(active, new Set(), 0, false, pending)!;
expect(pickCeoCompletionHandoff(fingerprint(pending), bound)).toBe(2);
const uiOnly = capturePlanCountQuestion(active, new Set(), 0, false)!;
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
const issue = '☐ Security finding\nChoose how to parameterize the SQL query.\n 1. Fix query\n 2. Add a TODO\nEnter to select · ↑/↓ to navigate · Esc to cancel';
const unbound = capturePlanCountQuestion(issue, new Set(), 0, false, pending)!;
expect(unbound.nativeCall).toBeUndefined();
expect(pickCeoCompletionHandoff(fingerprint(pending), unbound)).toBeNull();
});
});
describe('completed CEO handoff with native next-step identity', () => {
function capturedHandoff(): NativePlanQuestionCall {
const question = 'D7 — CEO review is complete. Run /plan-eng-review next (the required shipping gate)? <gstack-qid:plan-ceo-review-next-step>';
return {
sessionId: 'e10cf0b4-525b-442d-9c2a-7a48d6b39f50',
toolUseId: 'toolu_01FmkkRpoE3s6Y93KX6zLN1q',
answered: true,
failed: false,
questions: [{
question,
header: 'Next review',
multiSelect: false,
options: [
{ label: 'Run /plan-eng-review next (recommended)' },
{ label: "Skip — I'll handle reviews manually" },
],
}],
answers: { [question]: 'Run /plan-eng-review next (recommended)' },
unansweredQuestionIndices: [],
};
}
test('captured completed-review menu is administrative and retains every independent finding and TODO', () => {
const calls = captures.cases[1]!.calls.slice(0, -1).map(c => nativeCall(c));
const result = replay([...calls, capturedHandoff()]);
expect(result).toMatchObject({ reviewCount: 4, administrativeCount: 1, step0Count: 0 });
expect(result.classifications.slice(0, -1).every(p => !p.administrative)).toBe(true);
});
test('only the positively bound pending handoff selects manual, in either option order', () => {
for (const reverse of [false, true]) {
const call = capturedHandoff();
call.answered = false;
delete call.answers;
if (reverse) call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
});
test('incomplete review, missing gate, findings, mixed choices, and unoffered answers stay substantive', () => {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete', 'has an unresolved test gap'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate', 'optional follow-up'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-next-step', 'plan-ceo-security-finding'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'TODO: email queue'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add missing staging validation to this plan' }); },
]) {
const call = capturedHandoff();
mutate(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(replay([call]).reviewCount).toBe(1);
}
const call = capturedHandoff();
call.answers = { [call.questions[0]!.question]: 'First add the missing retry test' };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
});
});
const CAPTURED_PAIRED_RETRY_CALLS: NativePlanQuestionCall[] = [
{
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
"toolUseId": "toolu_01CG8hh817d7CvFk9kH5ZFW4",
"questions": [
{
"question": "D6 — Section 2 finding: the 502 failure path test's assertion is under-specified. What does 'fails clean' mean as an observable outcome? <gstack-qid:plan-ceo-fails-clean>",
"header": "502 failure mode",
"multiSelect": false,
"options": [
{
"label": "Specify the exception type in the plan (Recommended)",
"description": "Update the plan to name the exception class processPayment() raises after 502 exhaustion (e.g. 'assert raises Stripe::APIConnectionError' or 'assert raises PaymentFailedError'). The test must assert a concrete observable: the exception class, not just 'something goes wrong.' Effort: add 1 line to the plan. Verify: test fails with wrong exception type.",
"preview": "REMEDY:\n Plan change: add to item 2 under ## Tests:\n 'The 502 test must assert the specific exception class\n (or nil return, or error struct) processPayment() raises\n after retry exhaustion. The test factory already exposes\n mock call history; the test should also assert exactly 2\n charge attempts and 1 backoff sleep call.'\n\nWhy: without this, the implementer will write\n expect { processPayment() }.not_to raise_error\nwhich passes on the wrong behavior (swallowed exception)."
},
{
"label": "Accept 'fails clean' as implementation-determined",
"description": "Trust the implementer to look at processPayment() and assert whatever behavior they find. The test is still useful. Risk: if processPayment() silently swallows the error (no raise, no return value), the test will pass even when payment silently fails."
}
]
}
],
"answered": true,
"failed": false,
"answers": {
"D6 — Section 2 finding: the 502 failure path test's assertion is under-specified. What does 'fails clean' mean as an observable outcome? <gstack-qid:plan-ceo-fails-clean>": "Specify the exception type in the plan (Recommended)"
},
"unansweredQuestionIndices": [],
"answeredAt": "2026-09-08T20:57:40.308Z"
},
{
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
"toolUseId": "toolu_01Bc1mwoqXgNQK7NVx8MA21L",
"questions": [
{
"question": "D7 — Section 4 finding: the happy path assertion 'correct receipt is generated' needs to be field-specific to be a correctness test. <gstack-qid:plan-ceo-receipt-fields>",
"header": "Receipt assertion",
"multiSelect": false,
"options": [
{
"label": "Add field-level assertion requirement to the plan (Recommended)",
"description": "Update the plan: the happy path test must assert specific receipt fields (at minimum: amount matches charged amount, stripe_charge_id matches the mock's returned charge ID). Prevents the test from being just a nil-check smoke test. Effort: add 1 line to the plan. Verify: test fails if receipt has wrong charge ID.",
"preview": "REMEDY:\n Plan change: add to item 1 under ## Tests:\n 'The happy path test must assert field-level receipt\n correctness: at minimum, the receipt amount equals the\n charged amount and the receipt stripe_charge_id matches\n the charge ID returned by the Stripe mock.\n assert receipt.amount == expected_amount\n assert receipt.stripe_charge_id == mock_charge.id'\n\nWhy: 'assert receipt is generated' is a smoke test.\n It passes even if receipt contains wrong amount or\n no charge ID, which is the correctness we care about."
},
{
"label": "Accept 'correct receipt generated' as-is",
"description": "Leave it to the implementer. Risk: a future refactor that breaks receipt field accuracy (e.g., amount in wrong currency unit) could pass the test silently."
}
]
}
],
"answered": true,
"failed": false,
"answers": {
"D7 — Section 4 finding: the happy path assertion 'correct receipt is generated' needs to be field-specific to be a correctness test. <gstack-qid:plan-ceo-receipt-fields>": "Add field-level assertion requirement to the plan (Recommended)"
},
"unansweredQuestionIndices": [],
"answeredAt": "2026-09-08T20:58:04.465Z"
},
{
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
"toolUseId": "toolu_014MYMCNbEGQfNYwqFkQQmKm",
"questions": [
{
"question": "D8 — TODO candidate: test the orphaned-payment scenario (receipt-builder fails after Stripe charge succeeds). Add to TODOS.md? <gstack-qid:plan-ceo-todo-orphaned-payment>",
"header": "TODO: orphaned payment",
"multiSelect": false,
"options": [
{
"label": "Add to TODOS.md (Recommended)",
"description": "Stripe charge succeeds, then receipt-builder throws — processPayment() returns an error to the caller. If the caller retries, the card gets charged twice. The existing receipt-builder regression tests cover receipt-builder itself, but nothing tests processPayment()'s behavior in this specific sequence. P2, M effort (human: ~1h / CC: ~8min).",
"preview": "TODO entry:\n What: Test orphaned-payment scenario in processPayment()\n Why: Stripe charge succeeds, receipt-builder throws,\n caller retries → double charge. No test covers this.\n Where: payment test factory already supports this setup.\n Effort: M (human ~1h / CC ~8min)\n Priority: P2\n Depends on: this PR (test infra in place)"
},
{
"label": "Skip — not valuable enough",
"description": "The receipt-builder regression tests provide sufficient coverage. Double-charge scenario is handled by idempotency keys at the Stripe level."
},
{
"label": "Build it now in this PR",
"description": "Add a third test case to this PR: receipt-builder throws after successful charge — assert processPayment() returns the expected error and Stripe mock shows only 1 charge attempt (no retry on receipt failure). Expands scope from HOLD SCOPE decision."
}
]
}
],
"answered": true,
"failed": false,
"answers": {
"D8 — TODO candidate: test the orphaned-payment scenario (receipt-builder fails after Stripe charge succeeds). Add to TODOS.md? <gstack-qid:plan-ceo-todo-orphaned-payment>": "Add to TODOS.md (Recommended)"
},
"unansweredQuestionIndices": [],
"answeredAt": "2026-09-08T20:59:02.919Z"
},
{
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
"toolUseId": "toolu_019ppgizjxzRiJd2QXPV7rYQ",
"questions": [
{
"question": "D9 — CEO review complete. Run /plan-eng-review next? <gstack-qid:plan-ceo-next-review>",
"header": "Next review",
"multiSelect": false,
"options": [
{
"label": "Run /plan-eng-review next (Recommended)",
"description": "Eng review is the required shipping gate. It covers architecture, code quality, and test correctness at the code level — what the CEO review doesn't dig into. The 2 spec gaps found here (exception type, receipt fields) should be verified at the code level too."
},
{
"label": "Skip — handle reviews manually",
"description": "Proceed without running eng review now. You can run it later with /plan-eng-review. Note: eng review is the only gate that blocks shipping by default."
}
]
}
],
"answered": true,
"failed": false,
"answers": {
"D9 — CEO review complete. Run /plan-eng-review next? <gstack-qid:plan-ceo-next-review>": "Run /plan-eng-review next (Recommended)"
},
"unansweredQuestionIndices": [],
"answeredAt": "2026-09-08T21:03:34.802Z"
}
];
describe('completed CEO next-review declaration and final report order', () => {
test('canonical identity alone never replaces actual completion and the next-review header', () => {
for (const mutate of [
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Should we finish reviewing? <gstack-qid:plan-ceo-next-steps>'; },
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'New security issue'; },
]) {
const call = structuredClone(CAPTURED_PAIRED_RETRY_CALLS.at(-1)!);
call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-next-review', 'plan-ceo-next-steps');
call.questions[0]!.options[1]!.label = "Skip — I'll handle reviews manually";
mutate(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
});
test('the captured retry keeps its three findings/TODOs and recognizes only the completed handoff', () => {
expect(replay(structuredClone(CAPTURED_PAIRED_RETRY_CALLS))).toMatchObject({
reviewCount: 3, administrativeCount: 1, step0Count: 0,
});
const call = structuredClone(CAPTURED_PAIRED_RETRY_CALLS.at(-1)!);
call.answered = false;
delete call.answers;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(2);
call.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(1);
});
test('a report written before the administrative handoff can reach the real plan-approval gate', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-handoff-order-'));
const file = path.join(dir, 'plan.md');
try {
fs.writeFileSync(file, '# Plan\n\n## GSTACK REVIEW REPORT\n\n' +
'| Review | Runs | Status | Findings |\n|---|---|---|---|\n| CEO | 1 | COMPLETE | 3 |\n\n' +
'VERDICT: CEO CLEARED\n\nNO UNRESOLVED DECISIONS\n');
// Native Write succeeded at this time, before the final handoff. The
// live inode was cleaned up; this fixture replays that observed order.
const reportAt = Date.parse('2026-09-08T21:01:38.295Z') / 1000;
fs.utimesSync(file, reportAt, reportAt);
const calls = structuredClone(CAPTURED_PAIRED_RETRY_CALLS);
const transcript = {
status: 'ready' as const,
calls,
assistantMessages: [],
planReadyRequests: [{
sessionId: calls[0]!.sessionId,
toolUseId: 'toolu_01XK7amzoCx4VTm1r2bHdtsH',
timestamp: '2026-09-08T21:03:46.725Z',
failed: false,
}],
};
const admin = new Set(calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
.map(call => `${call.sessionId}:${call.toolUseId}`));
const startedAt = Date.parse('2026-09-08T20:51:50Z');
expect(admin.size).toBe(1);
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
// A new substantive answer after the Write remains a freshness boundary.
calls.splice(-1, 0, { ...structuredClone(calls[0]!), toolUseId: 'later-substantive-fix',
answeredAt: '2026-09-08T21:03:00.000Z' });
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
});
describe('captured CEO next-step prefixes and immediate review menus', () => {
test('next-step prefixes and a CLEAN declaration still identify only the completed handoff', () => {
for (const scenario of currentHandoffs.cases) {
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
const before = structuredClone(call);
expect(replay([call])).toMatchObject({ reviewCount: 0, administrativeCount: 1, step0Count: 0 });
expect(call).toEqual(before);
}
});
test('the bound pending menu selects the offered manual action in either order', () => {
for (const scenario of currentHandoffs.cases) for (const reverse of [false, true]) {
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
if (reverse) call.questions[0]!.options.reverse();
const q = call.questions[0]!;
const active = `${q.header}\n${q.question}\n 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const bound = capturePlanCountQuestion(active, new Set(), 0, false, call)!;
expect(bound.nativeCall?.toolUseId).toBe(call.toolUseId);
expect(pickCeoCompletionHandoff(fingerprint(call), bound)).toBe(reverse ? 1 : 2);
expect(isCeoCompletionHandoff(bound)).toBe(false);
const uiOnly = capturePlanCountQuestion(active, new Set(), 0, false)!;
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
}
});
test('conditional completion, substantive actions, and mismatched identities still cannot authorize a handoff', () => {
for (const scenario of currentHandoffs.cases) for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: If the CEO review is complete, should we run the next review? Eng review is the required shipping gate.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: The CEO review is not complete. Eng review is the required shipping gate.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: CEO review is CLEAN only after fixing this security gap. Eng review is the required shipping gate.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Security finding'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' and implement the fixes'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add missing retry coverage to TODOS.md' }); },
]) {
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
mutate(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
expect(replay([call]).reviewCount).toBe(1);
call.answered = false; delete call.answers;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
}
for (const scenario of currentHandoffs.cases) {
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
call.answered = false;
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'other-session:other-call' })).toBeNull();
}
});
});
describe('native CEO completed handoffs with deferred implementation', () => {
test('captured full sessions keep all substantive questions and classify only the final handoff', () => {
for (const scenario of kHandoffs.cases) {
const calls = structuredClone(scenario.calls) as NativePlanQuestionCall[];
const original = structuredClone(calls);
const result = replay(calls, false, ceoFirstReviewAUQ);
expect(result).toMatchObject({ step0Count: scenario.expectedSetupCount,
reviewCount: scenario.expectedReviewCount, administrativeCount: 1 });
expect(result.classifications.slice(0, -1).every(p => !p.administrative)).toBe(true);
expect(calls).toEqual(original);
}
});
test('active native handoffs choose manual in either order, never implementation or another review', () => {
for (const scenario of kHandoffs.cases) for (const reverse of [false, true]) {
const call = structuredClone(scenario.calls.at(-1)!) as NativePlanQuestionCall;
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
const q = call.questions[0]!;
if (reverse) q.options.reverse();
const options = q.options.map((option, i) => `${i === 0 ? '' : ' '} ${i + 1}. ${option.label}`).join('\n');
const screen = `${q.header}\n${q.question}\n${options}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
const bound = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
expect(bound.nativeCall?.toolUseId).toBe(call.toolUseId);
expect(pickCeoCompletionHandoff(fingerprint(call), bound)).toBe(q.options.findIndex(o => /handle.*manually/i.test(o.label)) + 1);
expect(isCeoCompletionHandoff(bound)).toBe(false);
const uiOnly = capturePlanCountQuestion(screen, new Set(), 0, false)!;
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
}
});
test('conditional declarations and new implementation obligations remain substantive', () => {
for (const scenario of kHandoffs.cases) for (const question of [
'ELI10: If the CEO review is done and the plan is cleared, choose the next step.',
'ELI10: The CEO review is done only after resolving the test gap.',
'ELI10: The CEO review is done and the plan is cleared after you add retry tests.',
'ELI10: The CEO review is not done and the plan is not cleared.',
]) {
const call = structuredClone(scenario.calls.at(-1)!) as NativePlanQuestionCall;
call.questions[0]!.question = question + ' The required shipping gate is an Eng Review.';
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
}
for (const option of [
{ label: 'Implement now, eng review later', description: 'Add the missing receipt test, then implement.' },
{ label: 'Implement now, eng review later', description: 'Implement the approved tasks and add a new receipt assertion before the next review.' },
{ label: 'Implement now, eng review later', description: 'The plan has no approved tasks; decide the missing error contract during implementation.' },
{ label: 'Implement new retry behavior now, eng review later', description: 'The plan already has approved tasks.' },
{ label: 'Add another TODO before implementing', description: 'Use the approved plan.' },
]) {
const call = structuredClone(kHandoffs.cases[0]!.calls.at(-1)!) as NativePlanQuestionCall;
call.questions[0]!.options[1] = option;
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
call.answered = false;
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
}
});
test('a real native approval after the completed report still requires all substantive answers in that report', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-k-handoff-'));
const file = path.join(dir, 'plan.md');
try {
for (const scenario of kHandoffs.cases) {
fs.writeFileSync(file, '# Plan\n\n## GSTACK REVIEW REPORT\n\n' +
'| Review | Runs | Status | Findings |\n|---|---|---|---|\n| CEO | 1 | COMPLETE | 4 |\n\n' +
'VERDICT: CEO CLEARED\n\nNO UNRESOLVED DECISIONS\n');
fs.utimesSync(file, scenario.reportAtMs / 1000, scenario.reportAtMs / 1000);
const calls = structuredClone(scenario.calls) as NativePlanQuestionCall[];
const transcript = { status: 'ready' as const, calls, assistantMessages: [],
planReadyRequests: structuredClone(scenario.planReadyRequests) };
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
const startedAt = Date.parse('2026-09-08T22:17:54Z');
expect(admin.size).toBe(1);
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready')).toBe(false);
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(true);
transcript.planReadyRequests[0]!.failed = true;
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
transcript.planReadyRequests[0]!.failed = false;
calls.splice(-1, 0, { ...structuredClone(calls[2]!), toolUseId: 'new-substantive-answer',
answeredAt: new Date(scenario.reportAtMs + 1000).toISOString() });
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
}
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
});
+181
View File
@@ -0,0 +1,181 @@
import { expect, test } from 'bun:test';
import captured from './fixtures/ceo-contract-assertions-ag.json';
import retry from './fixtures/ceo-contract-assertions-ag-retry.json';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
function reanswer(call: NativePlanQuestionCall) {
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
return call;
}
test('actual declarative assertion defects start review after routing and approach', () => {
let started = false;
const counts = { setup: 0, review: 0 };
for (const call of calls()) {
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
counts[phase.preReview ? 'setup' : 'review']++;
}
expect(counts).toEqual({ setup: 2, review: 2 });
for (const call of calls().slice(2)) expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
// Correct classification cannot retroactively complete the original paid run.
expect(captured.observedOutcome).toBe('no_review_questions');
expect(captured.observedReviewCount).toBe(0);
});
test('assertion briefs still require completed native identity and their actual remedy', () => {
for (const original of calls().slice(2)) {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Issue 99'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/Recommendation: \d[A-Z]/, 'Recommendation: 99Z'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.forEach((o, i) => { o.label = `${i + 1}A) Keep`; o.description = 'Keep the saved report.'; }); },
]) {
const call = structuredClone(original); mutate(call);
if (call.answers && Object.keys(call.answers).length) reanswer(call);
expect(ceoFirstReviewAUQ(fp(call))).toBe(false);
}
expect(ceoFirstReviewAUQ({ ...fp(original), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(original), options: [] })).toBe(false);
}
});
test('historical, hypothetical, quoted and withdrawn assertion problems are not current findings', () => {
for (const original of calls().slice(2)) {
for (const prefix of ['If ', 'Example: ', 'Whether ', 'Unless ']) {
const call = structuredClone(original);
call.questions[0]!.question = call.questions[0]!.question.replace(/(Issue \d+: )/, `$1${prefix}`);
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
for (const replacement of [
'test 2 can detect all retry or backoff regressions',
'test 2 previously could not detect retry or backoff regressions',
'test 1 does not accept any truthy value as a correct receipt',
'"test 2 cannot detect retry or backoff regressions"',
]) {
const call = structuredClone(original);
call.questions[0]!.question = call.questions[0]!.question.replace(/(Issue \d+: )[^\n]+/, `$1${replacement}`);
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
const withdrawn = structuredClone(original);
withdrawn.questions[0]!.question = withdrawn.questions[0]!.question.replace(/(ELI10:[^\n]+)/, '$1 No current defect exists.');
expect(ceoFirstReviewAUQ(fp(reanswer(withdrawn)))).toBe(false);
}
});
test('the captured assertion regression selects the existing CEO count eval', () => {
for (const file of ['test/ceo-contract-assertions-ag.test.ts', 'test/fixtures/ceo-contract-assertions-ag.json']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
}
});
test('actual retry contract wording recognizes its first repair and counts three review decisions', () => {
let started = false;
const counts = { setup: 0, review: 0 };
for (const call of structuredClone(retry.calls) as NativePlanQuestionCall[]) {
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
counts[phase.preReview ? 'setup' : 'review']++;
}
expect(counts).toEqual({ setup: 2, review: 3 });
for (const original of retry.calls.slice(2, 4)) {
const call = structuredClone(original) as NativePlanQuestionCall;
expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options.at(-1)!.label };
expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
}
expect(retry.observedOutcome).toBe('no_review_questions');
expect(retry.observedReviewCount).toBe(0);
expect(selectTests(['test/fixtures/ceo-contract-assertions-ag-retry.json'], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
});
test('already complete assertions and layout-only choices do not invent a defect', () => {
const cases = [
[2, 'D2 — Issue 1: test 2 cannot detect retry regressions (historical assessment)', 'The assertion gap was fixed yesterday. The current test pins the retry count and delay; this choice only arranges the already complete tests.'],
[3, 'D3 — Issue 2: test 1 accepts any truthy value as specified by its success contract', 'The contract intentionally accepts every truthy success marker. The current assertion covers the contract completely; this choice only arranges the existing test.'],
] as const;
for (const [index, title, explanation] of cases) {
const call = calls()[index]!;
const q = call.questions[0]!;
q.question = `${title}\nELI10: ${explanation}\nRecommendation: A`;
q.options = [{ label: 'A) Use a table-driven layout', description: 'Use a table-driven layout for the existing assertions.' }, { label: 'B) Keep the existing layout', description: 'Keep the existing assertions in place.' }];
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
});
test('retry assertion brief keeps native identity, exact contract and repair requirements', () => {
for (const original of retry.calls.slice(2, 4)) {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding 99'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Example: ' + c.questions[0]!.question; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('but the contract is', 'but there is no contract for'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^ELI10:.*$/m, 'ELI10: The current assertion covers the contract completely.'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('Fix the assertion?', 'Save the report?'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.forEach(o => { o.label = o.label.replace(/\).*/, ') Use the existing layout'); o.description = 'Use the existing layout.'; }); },
]) {
const call = structuredClone(original) as NativePlanQuestionCall;
mutate(call);
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
}
});
test('one offered option must repair the assertion rather than borrow report and layout actions', () => {
for (const original of [...captured.calls.slice(2), ...retry.calls.slice(2, 4)]) {
for (const administrative of ['Verify the saved report', 'Assert the full report', 'Pin the exact saved plan', 'Verify the expected layout']) {
const call = structuredClone(original) as NativePlanQuestionCall;
const q = call.questions[0]!;
const prefix = /^([1-9]\d*)?[A-Z]/.exec(q.options[0]!.label)![1] ?? '';
q.options = [
{ label: `${prefix}A) ${administrative}`, description: administrative + '.' },
{ label: `${prefix}B) Use a table-driven layout`, description: 'Use a table-driven layout for the existing assertions.' },
];
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
}
});
test('administrative report qualifiers cannot strengthen the unchanged assertion clause', () => {
for (const suffix of [' and include a full report.', '; write an exact report.', '. Save the complete plan.']) {
const call = calls()[2]!;
const q = call.questions[0]!;
q.options = [
{ label: '1A) Assert the error class only', description: 'Assert the error class only' + suffix },
{ label: '1B) Keep the current test', description: 'Leave the current rejection-only assertion unchanged.' },
];
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
});
test('each assertion clause owns its strong qualifier and actual assertion target', () => {
for (const suffix of [' and verify the full report.', ' and check the full report.', ' with a full report.', ' with a complete saved plan.']) {
const call = calls()[2]!;
call.questions[0]!.options = [
{ label: '1A) Assert the error class only', description: 'Assert the error class only' + suffix },
{ label: '1B) Keep the current test', description: 'Leave the current rejection-only assertion unchanged.' },
];
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
}
for (const description of ['Assert the rejection class and exactly two Stripe attempts.', 'Assert the error class only and assert exactly two Stripe attempts.']) {
const call = calls()[2]!;
call.questions[0]!.options[0]!.label = '1A) Strengthen the assertions';
call.questions[0]!.options[0]!.description = description;
call.questions[0]!.options = [call.questions[0]!.options[0]!, { label: '1B) Keep the current test', description: 'Leave the rejection-only assertion unchanged.' }];
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(true);
}
});
+245
View File
@@ -0,0 +1,245 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
import fixture from './fixtures/ceo-contract-question-an.json';
import sectionFixture from './fixtures/ceo-section-finding-an.json';
import contractFixture from './fixtures/ceo-current-contract-an.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
const findings = calls.slice(2);
const sectionCalls = sectionFixture.fingerprints as AskUserQuestionFingerprint[];
const sectionFindings = sectionCalls.slice(4, 6);
function change(fp: AskUserQuestionFingerprint, edit: (q: any, call: any, fp: any) => void) {
const copy = structuredClone(fp), call = copy.nativeCall!, q = call.questions[0]!;
const answerIndex = q.options.findIndex(o => o.label === call.answers?.[q.question]);
edit(q, call, copy);
call.answers = { [q.question]: q.options[answerIndex]?.label ?? '' };
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
return copy;
}
test('both actual completed contract questions start review; routing and test layout remain setup', () => {
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, true, true]);
});
test('the decision ordinal, punctuation and form of the remedy question do not carry the finding', () => {
for (const fp of findings) for (const title of [
'd19 — Test 1 checks only truthiness; what should the exact assertion verify?',
'D4 - Test 1 asserts only truthiness. How should the test check the full contract?',
'D7 — Test 1 checks only truthiness: assert the contract or keep this check?',
]) {
expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace(q.question.split('\n')[0], title);
q.header = 'Test contract';
}))).toBe(true);
}
});
test('a competing test header or explicit foreign issue cannot borrow a test assertion', () => {
for (const header of ['Test 99 assert', 'Finding 3', 'Issue 1'])
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.header = header; }))).toBe(false);
});
test('a title alone or an administrative response does not establish a review finding', () => {
for (const fp of findings) {
expect(ceoFirstReviewAUQ({ ...fp, nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
q.options = [
{ label: 'A) Keep the current assertion (recommended)', description: 'Leave the test unchanged.' },
{ label: 'B) Archive the review', description: 'Save the existing report without changing tests.' },
];
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { q.header = 'Approach'; }))).toBe(false);
}
});
test('the full native identity, selected answer and completed result remain required', () => {
for (const fp of findings) {
for (const edit of [
(_q: any, c: any) => { c.answered = false; },
(_q: any, c: any) => { c.failed = true; },
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; },
(_q: any, _c: any, f: any) => { f.signature = 'foreign:tool'; },
(q: any) => { q.multiSelect = true; },
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
const answer = change(fp, () => {}); answer.nativeCall!.answers = {};
expect(ceoFirstReviewAUQ(answer)).toBe(false);
const menu = change(fp, () => {}); menu.options[0]!.label = 'Foreign selection';
expect(ceoFirstReviewAUQ(menu)).toBe(false);
}
});
test('source and conditional frames cannot own the current assertion assessment', () => {
for (const intro of ['Source:', 'Example:', 'Earlier review assessment:', 'The following assessment is hypothetical.'])
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('\nELI10:', '\n' + intro + '\nELI10:'); }))).toBe(false);
for (const intro of ['Source excerpt: ', 'Previously, ', 'If approved, ', 'The following is a hypothetical example. '])
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + intro); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: If approved, '); }))).toBe(false);
});
test('literal titles and withdrawn current findings supply no first-review credit', () => {
for (const fp of findings) {
expect(ceoFirstReviewAUQ(change(fp, q => { const lines = q.question.split('\n'); lines[0] = '`' + lines[0] + '`'; q.question = lines.join('\n'); }))).toBe(false);
for (const statement of [
'Correction: this finding is withdrawn.',
'Correction: this finding is "withdrawn".',
'Correction: this explanation is not current.',
'There is no current gap.',
]) expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + statement; }))).toBe(false);
}
});
test('quoted historical notes cannot withdraw the current finding', () => {
expect(ceoFirstReviewAUQ(change(findings[0]!, q => {
q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:');
}))).toBe(true);
});
test('uniform recommendation and option identities remain required', () => {
for (const edit of [
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: Z'); },
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', '9B)'); },
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', 'A)'); },
]) expect(ceoFirstReviewAUQ(change(findings[0]!, edit))).toBe(false);
});
test('the new regression inputs belong only to the dense CEO finding owner', () => {
for (const name of ['test/ceo-contract-question-an.test.ts', 'test/fixtures/ceo-contract-question-an.json', 'test/fixtures/ceo-section-finding-an.json', 'test/fixtures/ceo-current-contract-an.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(name)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
for (let i = 0; i < paths.length; i++) {
expect(Object.hasOwn(paths, i)).toBe(true);
expect(typeof paths[i]).toBe('string');
}
});
test('owned Section finding briefs establish review through their current defect and remedy', () => {
expect(sectionCalls.map(ceoFirstReviewAUQ)).toEqual([false, false, false, false, true, true, false]);
for (const fp of sectionFindings) for (const separator of [':', '—', '-']) {
expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace(/^D\d+ — Section 2 finding (\d):/, `d19 — Section 7 finding $1 ${separator}`);
q.header = 'Section 7';
}))).toBe(true);
}
expect(ceoFirstReviewAUQ(change(sectionFindings[0]!, q => {
q.question = q.question.replace('the lookup reads request.params.userId into a raw SQL fragment', 'the query reads payload.accountId into a raw SQL string');
}))).toBe(true);
for (const term of ['“no error handling”', "'no error handling'", 'no error handling'])
expect(ceoFirstReviewAUQ(change(sectionFindings[1]!, q => {
q.question = q.question.replace('"no error handling"', term);
}))).toBe(true);
});
test('Section dispatch requires an exact completed native question and consistent finding identity', () => {
for (const fp of sectionFindings) for (const edit of [
(_q: any, c: any) => { delete c.answeredAt; },
(_q: any, c: any) => { c.answeredAt = 'not-a-time'; },
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
(_q: any, c: any) => { c.answered = false; },
(_q: any, c: any) => { c.failed = true; },
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; },
(q: any) => { q.header = 'Section 8'; },
(q: any) => { q.header = 'Finding 99'; },
(q: any) => { q.header = 'Section 2 finding 99'; },
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: 99A'); },
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
});
test('Section declarations cannot borrow source, historical, conditional or negated defects', () => {
for (const fp of sectionFindings) for (const prefix of ['Source: ', 'Previously, ', 'If approved, ', 'The hypothetical example: ', 'Earlier review assessment: ', 'For historical context, ']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/(Section 2 finding \d: )/, '$1' + prefix); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
}
for (const fp of sectionFindings) for (const prefix of ['Source:', 'Earlier review assessment:', 'The following is a hypothetical example.'])
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(sectionFindings[0]!, q => {
q.question = q.question.replace('the lookup reads', 'the lookup no longer reads');
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(sectionFindings[1]!, q => {
q.question = q.question.replace('the receipt email has "no error handling"', 'the receipt email no longer has "no error handling"');
}))).toBe(false);
});
test('Section review requires a current offered amendment and an unwithdrawn assessment', () => {
for (const fp of sectionFindings) {
for (const status of ['This finding is withdrawn.', 'Correction: this finding is "withdrawn".', 'This explanation is not current.', 'There is no current gap.'])
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + status; }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
q.options = [
{ label: 'A) Keep the existing implementation', description: 'Leave all behavior unchanged.' },
{ label: 'B) Archive the report', description: 'Export the report.' },
];
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
for (const option of q.options) option.description = 'Source excerpt: ' + option.description;
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
for (const option of q.options) option.description += '\nThis amendment is withdrawn.';
}))).toBe(false);
for (const status of [' This amendment is withdrawn.', ' This remedy is a historical example, not the current option.'])
expect(ceoFirstReviewAUQ(change(fp, q => {
for (const option of q.options) option.description += status;
}))).toBe(false);
for (const prefix of ['Source excerpt: ', 'If approved later: '])
expect(ceoFirstReviewAUQ(change(fp, q => {
for (const option of q.options) option.label = option.label.replace(/^([A-C]\)) /, '$1 ' + prefix);
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:');
}))).toBe(true);
}
});
test('the current plan contract can establish the gap in a later ELI10 sentence', () => {
const fp = contractFixture.fingerprints[2] as AskUserQuestionFingerprint;
expect(ceoFirstReviewAUQ(fp)).toBe(true);
for (const clause of [
"The current plan states 'no error handling on the email leg'.",
'The plan specifies “no error handling on the email leg”.',
'This plan requires "no error handling on the email leg".',
'The plan says no error handling on the email leg.',
]) expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace("The plan says 'no error handling on the email leg'.", clause);
}))).toBe(true);
});
test('later contract declarations retain source, currentness and remedy ownership', () => {
const fp = contractFixture.fingerprints[2] as AskUserQuestionFingerprint;
for (const clause of [
"The old plan said 'no error handling on the email leg'.",
"If approved, the plan says 'no error handling on the email leg'.",
"Source excerpt: the plan says 'no error handling on the email leg'.",
'"The plan says no error handling on the email leg."',
"The plan no longer says 'no error handling on the email leg'.",
"The plan says 'no error handling on the email leg' only in a historical example.",
"The plan says 'no error handling on the email leg”.",
]) expect(ceoFirstReviewAUQ(change(fp, q => {
q.question = q.question.replace("The plan says 'no error handling on the email leg'.", clause);
}))).toBe(false);
for (const edit of [
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: Earlier review assessment: '); },
(q: any) => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: Source excerpt: '); },
(q: any) => { q.question += '\nThis finding is "withdrawn".'; },
(q: any) => { for (const o of q.options) o.description += ' This amendment is withdrawn.'; },
(q: any) => { for (const o of q.options) o.description = 'Source excerpt: ' + o.description; },
(_q: any, c: any) => { delete c.answeredAt; },
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
(q: any) => { q.header = 'Finding 99'; },
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "Source excerpt follows. The plan says 'no error handling"); },
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "Earlier review assessment follows. The plan says 'no error handling"); },
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "If approved later. The plan says 'no error handling"); },
(q: any) => { q.question = q.question.replace("'no error handling on the email leg'.", "'no error handling on the email leg'. This no-error-handling contract is withdrawn."); },
(q: any) => { q.question = q.question.replace("'no error handling on the email leg'.", "'no error handling on the email leg'. This contract is a historical example, not the current plan."); },
(q: any) => { q.question += '\nThis finding is "resolved".'; },
(q: any) => { for (const o of q.options) o.description += '\nThis amendment is "closed".'; },
(q: any) => { for (const o of q.options) o.description += ' This amendment is "closed".'; },
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => {
q.question += '\nArchive note: "This finding is withdrawn."';
}))).toBe(true);
});
test('the assertion assessment and strengthening action retain their own current authority', () => {
for (const edit of [
(_q: any, c: any) => { delete c.answeredAt; },
(_q: any, c: any) => { c.answeredAt = 'invalid'; },
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
(q: any) => { q.question = q.question.replace('ELI10: The plan states', 'ELI10: The historical plan stated'); },
(q: any) => { q.question = q.question.replace('But the planned test only checks', 'But the planned test no longer only checks'); },
(q: any) => { q.options[0].label = q.options[0].label.replace('Assert deep equality with', 'Assert truthiness for'); },
(q: any) => { q.options[0].description = 'Source excerpt:\n' + q.options[0].description; },
(q: any) => { q.options[0].description = 'Earlier review assessment:\n' + q.options[0].description; },
(q: any) => { q.options[0].description += '\nThis amendment is withdrawn.'; },
(q: any) => { q.options[0].description += '\nThis amendment is "withdrawn".'; },
(q: any) => { q.options[0].description += '\nThis amendment is “withdrawn”.'; },
(q: any) => { q.options[0].description += '\nThis remedy is a historical example, not the current option.'; },
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: Earlier review assessment: '); },
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: For historical context, '); },
]) expect(ceoFirstReviewAUQ(change(findings[0]!, edit))).toBe(false);
expect(ceoFirstReviewAUQ(change(findings[0]!, q => {
q.options[1] = { label: 'B) Export documentation', description: 'Export the report.' };
}))).toBe(true);
});
+423
View File
@@ -0,0 +1,423 @@
import { expect, test } from 'bun:test';
import captured from './fixtures/ceo-count-ac-calls.json';
import later from './fixtures/ceo-count-ac-later-calls.json';
import alias from './fixtures/ceo-finding-alias-af.json';
import numberedBrief from './fixtures/ceo-numbered-brief-af.json';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
const finding = () => calls()[2]!;
const handoff = () => calls()[3]!;
function reanswer(c: NativePlanQuestionCall) {
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
return c;
}
function pending(c = handoff()) {
c.answered = false; delete c.answers; delete c.answeredAt;
c.unansweredQuestionIndices = [0]; return c;
}
test('the actual paired attempt has one finding and remains below its two-finding floor', () => {
let started = false;
const counts = { setup: 0, review: 0, administrative: 0 };
for (const c of calls()) {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ,
undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
counts[phase.administrative ? 'administrative' : phase.preReview ? 'setup' : 'review']++;
}
expect(counts).toEqual({ setup: 2, review: 1, administrative: 1 });
expect(counts.review).toBeLessThan(2);
expect(calls()[1]!.answers).toEqual(captured.calls[1]!.answers);
});
test('qidless explicit Findings need a completed matching native decision', () => {
expect(ceoFirstReviewAUQ(fp(finding()))).toBe(true);
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered answer' }; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
]) {
const c = finding(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
expect(ceoFirstReviewAUQ({ ...fp(finding()), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(finding()), nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(finding()), options: [] })).toBe(false);
});
test('setup recaps, quoted titles and foreign qids cannot start a review', () => {
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
const c = finding(); c.questions[0]!.question = prefix + c.questions[0]!.question;
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
for (const header of ['Approach', 'Mode', 'Next review', 'Setup']) {
const c = finding(); c.questions[0]!.header = header;
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
for (const id of ['plan-eng-review-finding', 'plan-ceo-review-mode', 'broken']) {
const c = finding(); c.questions[0]!.question += ` <gstack-qid:${id}>`;
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
expect(ceoFirstReviewAUQ(fp(calls()[1]!))).toBe(false);
expect(ceoFirstReviewAUQ(fp(handoff()))).toBe(false);
});
test('the exact administrative menu chooses manual without awarding completion coverage', () => {
expect(isCeoCompletionHandoff(fp(handoff()))).toBe(true);
expect(pickCeoCompletionHandoff(fp(pending()))).toBe(2);
const c = pending(); c.questions[0]!.options.reverse();
expect(pickCeoCompletionHandoff(fp(c))).toBe(1);
expect(isCeoCompletionHandoff(fp(c))).toBe(false);
expect(pickCeoCompletionHandoff(fp(handoff()))).toBeNull();
});
test('appended obligations and altered navigation context remain substantive', () => {
for (const extra of [' Also add another test.', ' Fix the missing auth check.',
' Once the outstanding gap is resolved.', ' Decide whether to add retry support?',
' The CEO review is not complete.']) {
for (const target of ['question', 'run', 'manual']) {
const c = handoff(), q = c.questions[0]!;
if (target === 'question') q.question += extra;
else q.options[target === 'run' ? 0 : 1]!.description += extra;
expect(isCeoCompletionHandoff(fp(reanswer(c)))).toBe(false);
expect(pickCeoCompletionHandoff(fp(pending(c)))).toBeNull();
}
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('complete and clean', 'not complete'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Add another test'; },
]) { const c = handoff(); mutate(c); expect(isCeoCompletionHandoff(fp(reanswer(c)))).toBe(false); }
expect(pickCeoCompletionHandoff({ ...fp(pending()), signature: 'foreign:call' })).toBeNull();
expect(pickCeoCompletionHandoff({ ...fp(pending()), options: [] })).toBeNull();
});
test('the paired transcript regression remains selected from both new files', () => {
for (const file of ['test/ceo-count-ac.test.ts', 'test/fixtures/ceo-count-ac-calls.json']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
}
});
test('later actual calls count explicit Issue and sectioned Finding titles without crediting a terminal', () => {
for (const [key, expected] of [['distinct', { setup: 4, review: 5 }], ['pairedRetry', { setup: 4, review: 4 }]] as const) {
let started = false;
const count = { setup: 0, review: 0 };
for (const c of structuredClone(later[key].nativeCalls) as NativePlanQuestionCall[]) {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
count[phase.preReview ? 'setup' : 'review']++;
}
expect(count).toEqual(expected);
}
// These attempts were stalled on file permission; count correction supplies
// no terminal, written report or complete methodology evidence.
expect(later.distinct.observedOutcome).toBe('timeout');
expect(later.pairedRetry.observedOutcome).toBe('running');
});
test('numbered Issue/sectioned Finding titles must agree with their native header', () => {
for (const source of [later.distinct.nativeCalls[4]!, later.pairedRetry.nativeCalls[4]!]) {
const c = structuredClone(source) as NativePlanQuestionCall;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
for (const header of ['Issue 7.2', 'Finding 9', 'Mode', 'Next review']) {
c.questions[0]!.header = header;
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
}
expect(selectTests(['test/fixtures/ceo-count-ac-later-calls.json'], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
});
function remedyCall(header: string, title: string, qid?: string) {
const c = finding();
c.questions[0]!.header = header;
c.questions[0]!.question = title + (qid ? `\n<gstack-qid:${qid}>` : '');
c.questions[0]!.options = [{ label: 'Repair the plan' }, { label: 'Keep the plan' }];
return reanswer(c);
}
function assertionCall(qid?: string) {
const c = remedyCall('Receipt shape', 'D2 — Test 1 asserts only that the receipt is truthy, but the plan states the exact receipt contract. Pin the full receipt?', qid);
c.questions[0]!.options = [
{ label: 'A) Assert the exact receipt', description: 'Deep equality against the complete stated receipt.' },
{ label: 'B) Keep truthy-only assertion', description: 'Leave the weaker planned assertion unchanged.' },
];
return reanswer(c);
}
test('an explicit exact-contract assertion gap does not depend on a Finding header or question tuning', () => {
for (const qid of [undefined, 'plan-ceo-review-receipt-contract']) {
const c = assertionCall(qid);
for (const option of c.questions[0]!.options) {
c.answers = { [c.questions[0]!.question]: option.label };
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
}
});
test('assertion-gap evidence needs a direct contract mismatch and opposed assertion choices', () => {
for (const change of [
(s: string) => 'Example: ' + s,
(s: string) => '> ' + s,
(s: string) => s.replace('Test 1 asserts', 'If Test 1 asserts'),
(s: string) => s.replace('the exact receipt contract', 'no required receipt shape'),
(s: string) => s.replace('the exact receipt contract', 'the exact receipt contract is already covered'),
]) {
const c = assertionCall(); c.questions[0]!.question = change(c.questions[0]!.question);
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Skip this review'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = ''; },
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
]) { const c = assertionCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
});
test('completed native remedy headers and numbered Issue titles start CEO review', () => {
for (const c of [
remedyCall('F1 remedy', 'D2 — Test 1: assert the full receipt, or keep the truthy-only assertion?'),
remedyCall('F2 remedy', 'D3 — Test 2: assert attempt count and backoff, or only the rejection?'),
remedyCall('Email leg', 'D4 — Issue 1: where does the notification run relative to commit?', 'plan-ceo-review-email-leg'),
]) {
for (const option of c.questions[0]!.options) {
c.answers = { [c.questions[0]!.question]: option.label };
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
.toEqual({ preReview: false, reviewStarted: true });
}
}
});
test('a remedy header requires a matching completed decision and consistent finding identity', () => {
const source = remedyCall('F1 remedy', 'D2 — Assert the complete receipt?');
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
]) {
const c = structuredClone(source); mutate(c);
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
expect(ceoFirstReviewAUQ({ ...fp(source), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(source), nativeCall: undefined })).toBe(false);
for (const title of ['D2 — Issue 2: Assert the receipt?', 'D2 — Issue 0: Assert the receipt?',
'D2 — Issue 1.0: Assert the receipt?', 'Example: D2 — Assert the receipt?',
'> D2 — Assert the receipt?', '```\nD2 — Assert the receipt?']) {
expect(ceoFirstReviewAUQ(fp(remedyCall('F1 remedy', title)))).toBe(false);
}
for (const header of ['Approach', 'F1', 'Remedy', 'F0 remedy', 'Next review']) {
expect(ceoFirstReviewAUQ(fp(remedyCall(header, 'D2 — Assert the receipt?')))).toBe(false);
}
});
test('numbered Issue titles cannot bypass setup, provider or native-answer checks', () => {
const title = 'D4 — Issue 1: where does the notification run relative to commit?';
for (const qid of ['plan-ceo-review-scope', 'plan-ceo-review-next-steps', 'plan-eng-review-email', 'foreign']) {
expect(ceoFirstReviewAUQ(fp(remedyCall('Email leg', title, qid)))).toBe(false);
}
for (const header of ['Setup', 'Approach', 'Mode', 'Next steps', 'Issue 2']) {
expect(ceoFirstReviewAUQ(fp(remedyCall(header, title, 'plan-ceo-review-email')))).toBe(false);
}
for (const suffix of ['<gstack-qid:plan-ceo-review-email', '<gstack-qid:plan-ceo-review-email:foreign>',
'<gstack-qid:plan-ceo-review-email> <gstack-qid:plan-eng-review-email>']) {
const c = remedyCall('Email leg', title + '\n' + suffix);
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ ...c.questions[0]!.options[0]! }); },
]) {
const c = remedyCall('Email leg', title, 'plan-ceo-review-email'); mutate(c);
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
});
const aliasCalls = () => alias.rows.map(row => structuredClone(row.call) as NativePlanQuestionCall);
test('AF exact native Finding headers and same-number Issue titles start review', () => {
for (const c of aliasCalls()) {
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
.toEqual({ preReview: false, reviewStarted: true });
}
expect(alias.provenance.partial).toBe(true);
expect(alias.provenance.paidCoverageCredit).toBe(false);
});
test('AF Issue and Finding aliases compare the complete native number, not the decision counter', () => {
for (const titleKind of ['Issue', 'Finding']) for (const headerKind of ['Issue', 'Finding']) {
for (const number of ['1', '2.1', '27.3']) {
const c = aliasCalls()[0]!, q = c.questions[0]!;
q.question = q.question.replace(/ <gstack-qid:[^>]+>/, '').replace(/^D4 — Issue 1:/, `D87 — ${titleKind} ${number}:`);
q.header = `${headerKind} ${number}`;
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
for (const wrong of ['9', `${number}.2`]) {
q.header = `${headerKind} ${wrong}`;
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
}
}
});
test('AF aliases preserve section and parenthesized issue header requirements', () => {
for (const title of ['D87 — Issue 2.1 (Section 4): Which assertion should be used?',
'D87 (issue 2.1) — Which assertion should be used?']) {
const c = aliasCalls()[0]!; c.questions[0]!.question = title;
for (const header of ['Issue 2.1', 'Finding 2.1']) {
c.questions[0]!.header = header;
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
}
for (const header of ['Receipt assertion', 'Issue 2', 'Finding 2.2']) {
c.questions[0]!.header = header;
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
}
});
test('AF aliases retain native completion, answer, qid and setup boundaries', () => {
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
c => { c.answered = false; }, c => { c.failed = true; },
c => { c.answers = {}; }, c => { c.unansweredQuestionIndices = [0]; },
c => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
c => { c.questions[0]!.multiSelect = true; },
c => { c.questions.push(structuredClone(c.questions[0]!)); },
c => { c.questions[0]!.options[1]!.label = c.questions[0]!.options[0]!.label; },
];
for (const mutate of mutations) for (const c of aliasCalls()) {
mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
for (const c of aliasCalls()) {
expect(ceoFirstReviewAUQ({ ...fp(c), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(c), nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(c), options: [] })).toBe(false);
for (const header of ['Issue 9', 'Finding 9', 'Setup', 'Approach', 'Mode', 'Next steps']) {
const changed = structuredClone(c); changed.questions[0]!.header = header;
expect(ceoFirstReviewAUQ(fp(changed))).toBe(false);
}
for (const qid of ['plan-ceo-review-setup', 'plan-eng-review-finding']) {
const changed = structuredClone(c); changed.questions[0]!.question = changed.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, `<gstack-qid:${qid}>`);
expect(ceoFirstReviewAUQ(fp(reanswer(changed)))).toBe(false);
}
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
const changed = structuredClone(c); changed.questions[0]!.question = prefix + changed.questions[0]!.question;
expect(ceoFirstReviewAUQ(fp(reanswer(changed)))).toBe(false);
}
}
});
test('AF alias evidence remains registered only to the CEO count workflow', () => {
const file = 'test/fixtures/ceo-finding-alias-af.json';
const owners = Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name);
expect(owners).toEqual(['plan-ceo-finding-count']);
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
});
test('AF complete numbered native briefs identify the three remaining first decisions', () => {
for (const row of numberedBrief.rows) {
expect(ceoFirstReviewAUQ(fp(structuredClone(row.call) as NativePlanQuestionCall))).toBe(true);
}
});
test('AF complete finding identities permit F notation but never contradict the native header', () => {
for (const title of ['D7 — Finding F2.1: Which implementation should be used?', 'D7 — Issue 2.1: Which implementation should be used?']) {
const c = structuredClone(numberedBrief.rows[1]!.call) as NativePlanQuestionCall;
c.questions[0]!.question = title;
for (const header of ['F2.1 remedy', 'Issue F2.1', 'Finding 2.1']) {
c.questions[0]!.header = header; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
}
for (const header of ['F2 remedy', 'Finding 2.1.1', 'Issue F2.1.0', 'F2.1 and F3']) {
c.questions[0]!.header = header; expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
}
const c = aliasCalls()[0]!; c.questions[0]!.question = c.questions[0]!.question.replace('Issue 1:', 'Finding F1:');
c.questions[0]!.header = 'Finding 2'; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
});
test('AF a declarative numbered brief needs current problem evidence and an actual amendment choice', () => {
for (const body of [
'ELI10: The handler has error handling.\nRecommendation: A because it is ready.',
'ELI10: The handler has no current defect.\nRecommendation: A because it is ready.',
'ELI10: If the handler has no error handling, we would repair it.\nRecommendation: A because this is a hypothetical.',
'ELI10: Example: the handler has no error handling.\nRecommendation: A because this is an example.',
'ELI10: "The handler has no error handling."\nRecommendation: A because this quotes the old plan.',
'ELI10: The error contract is not missing.\nRecommendation: A because it is ready.',
'ELI10: The email failure is no longer unhandled.\nRecommendation: A because it is ready.',
]) {
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.question = 'D9 — 1.1 Email leg: transaction boundary and failure handling\n' + body;
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
for (const labels of [['Start review', 'Pause'], ['Write the completed report', 'Save the reviewed plan']]) {
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.options = labels.map((label,i) => ({label:`${i ? 'B' : 'A'}: ${label}`, description:label}));
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
});
test('AF new brief form preserves setup, native answer, quotation and subject binding', () => {
for (const row of numberedBrief.rows) for (const mutation of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Next steps'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
]) {
const c = structuredClone(row.call) as NativePlanQuestionCall; mutation(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.question = prefix + c.questions[0]!.question; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
}
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.header = 'SQL lookup'; expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes('test/fixtures/ceo-numbered-brief-af.json')).map(([name])=>name))
.toEqual(['plan-ceo-finding-count']);
});
test('AF a resolved historical gap and completed-review log check cannot start current review', () => {
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.header = 'Finding 1';
c.questions[0]!.question = 'D4 — Issue 1: Validation was missing in the prior review.\nELI10: The old gap is already resolved. Current validation is complete; this choice only checks the completed review log.\nRecommendation: A because it checks the record.';
c.questions[0]!.options = [
{label:'A) Check the prior review log',description:'Check the prior review log.'},
{label:'B) Keep current report',description:'Keep the current completed report.'},
];
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
});
test('AF currentness uses the whole explanation and saved-log actions remain administrative', () => {
for (const [subject, explanation, action, expected] of [
['Historical missing validation', 'Validation is complete. This task only verifies the stored review log; there is no current defect.', 'Validate the saved review log', false],
['Required validation is missing', 'The required validation is missing.', 'Validate the saved review log', false],
['Historical missing validation', 'The prior review omitted a note; there is no current defect.', 'Validate the incoming request', false],
['Required validation is missing', 'A previous log says "there is no current defect." The current plan still lacks validation.', 'Validate the incoming request', true],
] as const) {
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
c.questions[0]!.header = 'Issue 1';
c.questions[0]!.question = `D1 — Issue 1: ${subject}\nELI10: ${explanation}\nRecommendation: A because it addresses this decision.`;
c.questions[0]!.options = [
{label:`A) ${action}`, description:`${action}.`},
{label:'B) Keep the current report', description:'Leave the stored report unchanged.'},
];
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(expected);
}
});
+125
View File
@@ -0,0 +1,125 @@
import {expect,test} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
import fixture from './fixtures/ceo-count-ad-v2.json';
import {readPlanCountTranscript,type NativePlanQuestionCall} from './helpers/plan-count-transcript';
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
import {isCeoCompletionHandoff,pickCeoCompletionHandoff} from './helpers/ceo-completion-handoff';
import {selectTests, E2E_TOUCHFILES} from './helpers/touchfiles';
const fp=(c:NativePlanQuestionCall)=>nativePlanCallFingerprint(c,0,true);
const get=(which:'distinct'|'paired'|'pairedRetry',index:number)=>structuredClone(fixture.cases[which].calls[index]) as NativePlanQuestionCall;
const realFindings=()=>[get('distinct',4),get('paired',4),get('paired',5),get('pairedRetry',2),get('pairedRetry',3)];
const answer=(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};return c;};
function count(calls:NativePlanQuestionCall[]){let started=false;const n={setup:0,review:0,administrative:0};for(const c of calls){const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=p.reviewStarted;n[p.administrative?'administrative':p.preReview?'setup':'review']++;}return n;}
test('exact public native requests and successful replies reconstruct the captured calls once',()=>{
for(const which of ['distinct','paired','pairedRetry'] as const){const c=fixture.cases[which],dir=fs.mkdtempSync(path.join(os.tmpdir(),'ceo-count-public-'));try{const project=path.join(dir,'projects','owned');fs.mkdirSync(project,{recursive:true});const records=c.nativeRecords.map(r=>JSON.stringify(r)).join('\n')+'\n';fs.writeFileSync(path.join(project,c.calls[0]!.sessionId+'.jsonl'),records+records);expect(readPlanCountTranscript(dir,c.observation.capture.cwd).calls).toEqual(c.calls);for(const a of c.timeAnchors){expect(Date.parse(a.requestAt)).toBeLessThanOrEqual(Date.parse(a.replyAt));expect(Date.parse(a.replyAt)).toBeLessThanOrEqual(Date.parse(c.observation.capture.at));}}finally{fs.rmSync(dir,{recursive:true,force:true});}}
});
for(const [which,index] of [['distinct',4],['paired',4],['paired',5]] as const)test(`actual ${which} issue ${index} starts review from a completed native decision`,()=>expect(ceoFirstReviewAUQ(fp(get(which,index)))).toBe(true));
test('exact snapshots keep real issue counts and separate the administrative handoff',()=>{
expect(count(fixture.cases.distinct.calls as NativePlanQuestionCall[])).toEqual({setup:4,review:1,administrative:0});
expect(count(fixture.cases.paired.calls as NativePlanQuestionCall[])).toEqual({setup:4,review:2,administrative:1});
expect(fixture.cases.distinct.observation.state).toBe('in_progress');expect(fixture.cases.paired.observation.state).toBe('in_progress');
expect(count(fixture.cases.distinct.calls as NativePlanQuestionCall[]).review).toBeLessThan(4);
});
test('finding numbering and matching header identity are presentation, not extra findings',()=>{
for(const c of realFindings()){
const q=c.questions[0]!,oldTitle=q.question.split('\n')[0]!;q.question=q.question.replace(/^D\d+\s*[—–-]\s*/,'');answer(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
q.question=q.question.replace(/^(Finding|Issue)\s+[\d.]+:/,'$1 27.3:');if(/^(Finding|Issue)\s+[\d.]+$/i.test(q.header))q.header=q.header.replace(/[\d.]+/,'27.3');answer(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
expect(oldTitle).toContain('?');
}
});
test('native completion, identity, offered answer and unambiguous single issue remain mandatory',()=>{
const mutations:Array<(c:NativePlanQuestionCall)=>void>=[c=>{c.answered=false;},c=>{c.failed=true;},c=>{c.answers={};},c=>{c.unansweredQuestionIndices=[0];},c=>{c.questions[0]!.multiSelect=true;},c=>{c.answers={[c.questions[0]!.question]:'unoffered'};},c=>{c.questions.push(structuredClone(c.questions[0]!));},c=>{c.questions[0]!.options[1]!.label=c.questions[0]!.options[0]!.label;answer(c);}];
for(const mutate of mutations)for(const c of realFindings()){mutate(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(false);}
for(const c of realFindings()){expect(ceoFirstReviewAUQ({...fp(c),signature:'foreign:call'})).toBe(false);expect(ceoFirstReviewAUQ({...fp(c),nativeCall:undefined})).toBe(false);expect(ceoFirstReviewAUQ({...fp(c),options:[]})).toBe(false);}
});
test('setup, quoted examples, foreign qids and contradictory numbered headers cannot start review',()=>{
for(const prefix of ['Example: ','> ','"','```\n'])for(const c of realFindings()){c.questions[0]!.question=prefix+c.questions[0]!.question;expect(ceoFirstReviewAUQ(fp(answer(c)))).toBe(false);}
for(const header of ['Approach','Mode','Next review','Setup','Finding 88','Issue 88'])for(const c of realFindings()){c.questions[0]!.header=header;expect(ceoFirstReviewAUQ(fp(c))).toBe(false);}
for(const c of realFindings()){c.questions[0]!.question+=' <gstack-qid:plan-eng-review-finding>';expect(ceoFirstReviewAUQ(fp(answer(c)))).toBe(false);}
for(const which of ['distinct','paired'] as const)for(const c of fixture.cases[which].calls.slice(0,4))expect(ceoFirstReviewAUQ(fp(c as NativePlanQuestionCall))).toBe(false);
});
test('a completed pure next-review menu is administrative without granting pending input permission',()=>{
const c=get('paired',6);expect(isCeoCompletionHandoff(fp(c))).toBe(true);expect(ceoFirstReviewAUQ(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();c.answered=false;delete c.answers;delete c.answeredAt;c.unansweredQuestionIndices=[0];expect(isCeoCompletionHandoff(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();
});
test('new work or uncertain closure in the next-review choice stays substantive',()=>{
for(const suffix of ['\nFix the missing authentication check.','\nDelete the CI gate.','\nShip the new endpoint now.','\nWhich new endpoint should we add?']){const c=get('paired',6);c.questions[0]!.question+=suffix;expect(isCeoCompletionHandoff(fp(answer(c)))).toBe(false);}
for(const change of ['CEO review is not complete.','CEO review will be complete.','Example: CEO review complete.']){const c=get('paired',6);c.questions[0]!.question=c.questions[0]!.question.replace('CEO review complete.',change);expect(isCeoCompletionHandoff(fp(answer(c)))).toBe(false);}
for(const mutation of [c=>{c.failed=true;},c=>{c.questions[0].multiSelect=true;},c=>{c.answers={[c.questions[0].question]:'Fix the bug first'};},c=>{c.questions[0].options.push({label:'Fix the security issue',description:'Add a new check.'});}] as Array<(c:NativePlanQuestionCall)=>void>){const c=get('paired',6);mutation(c);expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
});
// The next-gate explanation must never turn conditional CEO closure into a
// completed review. Its narrow normalization is for counting only.
test('next Eng gate timing cannot supply conditional CEO completion', () => {
for (const replacement of [
'The CEO review is complete until someone runs it later.',
'The CEO review is complete if someone runs it later.',
'The CEO review will be complete after someone runs it later.',
'The CEO review still has unresolved findings.',
]) {
const call = get('paired', 6);
call.questions[0]!.question = call.questions[0]!.question.replace(
'The CEO review cleared scope and strengthened both test assertions.', replacement);
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
}
});
test('actual evidence and its regression select the paid CEO counting test', () => {
for (const file of ['test/ceo-count-ad-v2.test.ts', 'test/fixtures/ceo-count-ad-v2.json']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
}
});
test('the actual completed retry keeps two findings and its body-closure handoff administrative', () => {
const calls = fixture.cases.pairedRetry.calls as NativePlanQuestionCall[];
expect(count(calls)).toEqual({setup: 2, review: 2, administrative: 1});
expect(fixture.cases.pairedRetry.observation.outcome).toBe('no_review_questions');
expect(fixture.cases.pairedRetry.observation.completionCredit).toBe(false);
expect(isCeoCompletionHandoff(fp(get('pairedRetry', 4)))).toBe(true);
expect(pickCeoCompletionHandoff(fp(get('pairedRetry', 4)))).toBeNull();
});
test('body closure and echoed choices cannot hide new work or uncertain CEO closure', () => {
for (const text of ['Fix the missing authentication check.', 'Delete the CI gate.', 'Ship the new endpoint now.', 'Which endpoint should we add?']) {
const call = get('pairedRetry', 4);
call.questions[0]!.question += '\n' + text;
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
}
for (const text of ['The CEO review is not done', 'The CEO review will be done', 'The CEO review is done if the fixes land', 'Example: The CEO review is done']) {
const call = get('pairedRetry', 4);
call.questions[0]!.question = call.questions[0]!.question.replace('The CEO review is done', text);
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
}
const pending = get('pairedRetry', 4); pending.answered = false; delete pending.answers; delete pending.answeredAt; pending.unansweredQuestionIndices = [0];
expect(isCeoCompletionHandoff(fp(pending))).toBe(false);
expect(pickCeoCompletionHandoff(fp(pending))).toBeNull();
expect(isCeoCompletionHandoff({...fp(get('pairedRetry', 4)), signature: 'foreign:call'})).toBe(false);
});
test('every offered navigation clause rejects a new repair rather than hiding it under a valid recap', () => {
for (const [which, index] of [['paired', 6], ['pairedRetry', 4]] as const) {
for (const extra of ['Delete the CI gate.', 'Repair the retry assertion.', 'Disable authentication.', 'Please rewrite the endpoint.']) {
for (const optionIndex of [0, 1]) {
const call = get(which, index);
call.questions[0]!.options[optionIndex]!.description += ' ' + extra;
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
}
for (const where of ['before-net', 'inside-eli10'] as const) {
const call = get(which, index);
call.questions[0]!.question = where === 'before-net'
? call.questions[0]!.question.replace('\nNet:', '\n' + extra + '\nNet:')
: call.questions[0]!.question.replace('\nStakes if', ' ' + extra + '\nStakes if');
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
}
}
}
});
test('timing annotations cannot conceal substantive instructions', () => {
for (const [which, index] of [['paired', 6], ['pairedRetry', 4]] as const) {
for (const text of [' (human: Delete the CI gate)', ' (human: ~2 min / CC: Disable authentication)', ' (human: ~2 min / CC: ~1 min; repair the retry assertion)']) {
const c = get(which, index); c.questions[0]!.options[0]!.description += text;
expect(isCeoCompletionHandoff(fp(c))).toBe(false);
}
}
});
+88
View File
@@ -0,0 +1,88 @@
import { describe, expect, test } from 'bun:test';
import { capturePlanCountQuestion, nativePlanCallFingerprint, planCountQuestionInput } from './helpers/claude-pty-runner';
import { pickCeoCountQuestion } from './helpers/ceo-approach-pick';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import recorded from './fixtures/ceo-count-mode-ab-call.json';
function pending(): NativePlanQuestionCall {
const call = structuredClone(recorded) as NativePlanQuestionCall;
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
delete call.answeredAt;
return call;
}
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
function screen(call: NativePlanQuestionCall): string {
const q = call.questions[0]!;
return `${q.header}\n${q.question}\n${q.options.map((o, i) => `${i ? ' ' : ''} ${i + 1}. ${o.label}`).join('\n')}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
}
describe('fixed-scope CEO finding-count mode', () => {
test('the AB native mode menu selects HOLD SCOPE rather than its first expansion option', () => {
// The retained live record already contains the expansion answer. The
// pending state and frame are projections, not proof of live availability.
const call = pending();
const active = capturePlanCountQuestion(screen(call), new Set(), 0, true, call)!;
expect(active.nativeCall).toBe(call);
const selected = pickCeoCountQuestion(fp(call), active) ?? 1;
expect(selected).toBe(3);
expect(planCountQuestionInput(screen(call), active, selected)).toBe('3');
const q = recorded.questions[0]!;
expect(recorded.answers[q.question]).toBe(q.options[0]!.label);
expect(pickCeoCountQuestion(fp(recorded as NativePlanQuestionCall))).toBeNull();
});
test('every offered position chooses the same fixed scope, independent of the recommendation', () => {
for (let shift = 0; shift < 4; shift++) {
const call = pending();
const q = call.questions[0]!;
q.options = [...q.options.slice(shift), ...q.options.slice(0, shift)];
q.options.forEach(o => { o.label = o.label.replace(/ \(Recommended\)$/, ''); });
q.options.find(o => o.label.startsWith('SCOPE EXPANSION'))!.label += ' (Recommended)';
expect(pickCeoCountQuestion(fp(call))).toBe(q.options.findIndex(o => o.label.startsWith('HOLD SCOPE')) + 1);
}
});
test('requires a complete currently bound native pre-review question', () => {
const call = pending();
const fingerprint = fp(call);
const visibleOnly = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
expect(pickCeoCountQuestion(fingerprint, visibleOnly)).toBeNull();
expect(pickCeoCountQuestion({ ...fingerprint, preReview: false })).toBeNull();
expect(pickCeoCountQuestion({ ...fingerprint, signature: 'foreign:call' })).toBeNull();
expect(pickCeoCountQuestion({ ...fingerprint, nativeQuestionIndex: 1 })).toBeNull();
expect(pickCeoCountQuestion({ ...fingerprint, options: fingerprint.options.slice().reverse() })).toBeNull();
for (const mutate of [
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.failed; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options.pop(); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0] = structuredClone(c.questions[0]!.options[1]!); },
]) { const c = pending(); mutate(c); expect(pickCeoCountQuestion(fp(c))).toBeNull(); }
});
test('does not authorize negated, quoted, compound, foreign or finding questions', () => {
for (const question of [
'Which review mode should I not use?',
'Example: Which review mode should I use?',
'> Which review mode should I use?',
'Which review mode should I use? Delete the tests.',
'Should we approve this expansion?',
]) {
const call = pending(); call.questions[0]!.question = question + ' <gstack-qid:ceo-mode-selection>';
expect(pickCeoCountQuestion(fp(call))).toBeNull();
}
for (const id of ['plan-eng-mode', 'ceo-exp-e5-property-based', 'ceo-mode-selection-extra']) {
const call = pending(); call.questions[0]!.question = call.questions[0]!.question.replace('ceo-mode-selection', id);
expect(pickCeoCountQuestion(fp(call))).toBeNull();
}
for (const suffix of [' <gstack-qid:ceo-mode-selection>', ' <gstack-qid:broken']) {
const call = pending(); call.questions[0]!.question += suffix;
expect(pickCeoCountQuestion(fp(call))).toBeNull();
}
const call = pending(); call.questions[0]!.header = 'Finding';
expect(pickCeoCountQuestion(fp(call))).toBeNull();
});
});
+100
View File
@@ -0,0 +1,100 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, isQuestionlessNativePlanExit, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import type { NativePlanQuestionCall, PlanCountTranscript } from './helpers/plan-count-transcript';
import distinct from './fixtures/ceo-count-s-distinct.json';
import paired from './fixtures/ceo-count-s-paired.json';
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
function replay(calls: NativePlanQuestionCall[]) {
let started = false;
const setup = new Set<string>(); const administrative = new Set<string>(); let review = 0;
for (const call of calls) {
const fingerprint = fp(call);
const phase = planCountQuestionPhase(fingerprint, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
if (phase.administrative) administrative.add(fingerprint.signature);
else if (phase.preReview) setup.add(fingerprint.signature);
else review++;
started = phase.reviewStarted;
}
return { setup, administrative, review };
}
function transcript(capture: typeof distinct | typeof paired): PlanCountTranscript {
return { status: 'ready', calls: structuredClone(capture.calls) as NativePlanQuestionCall[],
assistantMessages: [], planReadyRequests: structuredClone(capture.planReadyRequests) };
}
function withReport(capture: typeof distinct | typeof paired, run: (file: string, start: number) => void) {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-s-terminal-'));
const file = path.join(dir, 'report.md');
fs.writeFileSync(file, capture.report);
fs.utimesSync(file, capture.reportMtimeMs / 1000, capture.reportMtimeMs / 1000);
try { run(file, capture.reportMtimeMs - 1000); }
finally { fs.rmSync(dir, { recursive: true, force: true }); }
}
describe('captured S native CEO completion gates', () => {
test('setup-only exit fails promptly without relaxing report freshness for positive coverage', () => {
const t = transcript(distinct); const result = replay(t.calls);
expect(result.setup.size).toBe(4); expect(result.review).toBe(0);
withReport(distinct, (file, start) => {
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen, result.setup)).toBe(true);
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen)).toBe(false);
expect(hasNativePlanTerminal(t, file, start, 'plan_ready')).toBe(false);
for (const mutate of [
(v: PlanCountTranscript) => { v.calls[0]!.answered = false; },
(v: PlanCountTranscript) => { v.calls[0]!.failed = true; },
(v: PlanCountTranscript) => { v.calls[0]!.sessionId = 'foreign'; },
(v: PlanCountTranscript) => { v.calls[0]!.answeredAt = 'invalid'; },
(v: PlanCountTranscript) => { v.calls[0]!.answeredAt = v.planReadyRequests![0]!.timestamp; },
(v: PlanCountTranscript) => { v.calls[0]!.answers = {}; },
(v: PlanCountTranscript) => { v.calls[0]!.unansweredQuestionIndices = [0]; },
]) {
const changed = structuredClone(t); mutate(changed);
expect(isQuestionlessNativePlanExit(changed, file, start, distinct.screen, result.setup)).toBe(false);
}
const incomplete = new Set(result.setup); incomplete.delete(fp(t.calls[0]!).signature);
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen, incomplete)).toBe(false);
});
});
test('paired review retains two issue approvals and excludes only the completed Eng menu', () => {
const t = transcript(paired); const before = structuredClone(t); const result = replay(t.calls);
expect(result.setup.size).toBe(2); expect(result.review).toBe(2); expect(result.administrative.size).toBe(1);
const pending = structuredClone(t.calls.at(-1)!); pending.answered = false; delete pending.answers;
expect(pickCeoCompletionHandoff(fp(pending))).toBe(2);
pending.questions[0]!.options.reverse(); expect(pickCeoCompletionHandoff(fp(pending))).toBe(1);
expect(t).toEqual(before);
withReport(paired, (file, start) => {
expect(hasNativePlanTerminal(t, file, start, 'plan_ready', result.administrative)).toBe(true);
expect(hasNativePlanTerminal(t, file, start, 'plan_ready')).toBe(false);
expect(isQuestionlessNativePlanExit(t, file, start, paired.screen, result.setup)).toBe(false);
});
});
test('the same menu cannot hide a new obligation, ambiguous gate, or unverified answer', () => {
const base = transcript(paired).calls.at(-1)!;
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('It\'s', 'That might become'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' First repair authorization.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Example: ' + c.questions[0]!.question; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('CLEAN', 'CLEAN once tests pass'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Remove the owner check.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Change the guarantee to permit old results.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Tests remain unresolved.'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix the missing assertion'; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.answered = false; },
]) {
const call = structuredClone(base); mutate(call);
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
}
const call = structuredClone(base); call.answers = { [call.questions[0]!.question]: 'First repair the missing test' };
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
const pending = structuredClone(base); pending.answered = false;
expect(pickCeoCompletionHandoff({ ...fp(pending), signature: 'foreign' })).toBeNull();
});
});
+78
View File
@@ -0,0 +1,78 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import fixture from './fixtures/ceo-current-omission-ap.json';
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
const first = calls[3]!;
const originalClause = 'The plan also does not say whether the email runs inside or after the DB transaction.';
function change(edit: (q: any, call: any, fp: any) => void) {
const copy = structuredClone(first), call = copy.nativeCall!, q = call.questions[0]!;
const selected = q.options.findIndex(o => o.label === call.answers?.[q.question]);
edit(q, call, copy);
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
return copy;
}
test('the exact failed retry begins review at its current missing transaction contract', () => {
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, false, true, false, false, false]);
let started = false;
const phases = calls.map(fp => { const p = planCountQuestionPhase(fp, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff); started = p.reviewStarted; return p; });
expect(fixture.actualCounts).toEqual({ setup: 7, review: 0 });
expect(phases.map(p => p.preReview)).toEqual([true, true, true, false, false, false, false]);
expect(phases.filter(p => !p.preReview && !p.administrative)).toHaveLength(4);
});
test('optional also and equivalent present-tense current owners preserve omission meaning', () => {
for (const phrase of ['The plan does not say whether', "This plan also doesn't say whether", 'This plan does not say whether'])
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('The plan also does not say whether', phrase); }))).toBe(true);
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('D2 —', 'D19 —'); }))).toBe(true);
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:'); }))).toBe(true);
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, '"Source: this finding is withdrawn." ' + originalClause); }))).toBe(true);
});
test('source, conditional and historical declarations cannot supply the missing contract', () => {
for (const prefix of ['Source: ', 'If approved, ', 'Previously, ', 'Earlier review assessment: ', 'The following is hypothetical. ']) {
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, prefix + originalClause); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: ' + prefix); }))).toBe(false);
}
for (const prefix of ['Source:', 'Earlier review assessment:', 'If approved:'])
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
for (const wrapped of ['"' + originalClause + '"', '`' + originalClause + '`', '> ' + originalClause, '```' + originalClause + '```'])
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, wrapped); }))).toBe(false);
for (const replacement of ['The previous plan also did not say whether', 'The plan now says whether', 'The example plan also does not say whether'])
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('The plan also does not say whether', replacement); }))).toBe(false);
});
test('the current omission and offered amendment must remain in force', () => {
for (const status of ['This finding is withdrawn.', 'This issue is "rejected".', 'Correction: this contract is not current.', 'There is no current gap.'])
expect(ceoFirstReviewAUQ(change(q => { q.question += '\n' + status; }))).toBe(false);
for (const prefix of ['Source excerpt: ', 'If approved later: ', 'Earlier review assessment: '])
expect(ceoFirstReviewAUQ(change(q => { for (const option of q.options) option.description = prefix + option.description; }))).toBe(false);
for (const status of ['This amendment is withdrawn.', 'This remedy is "cancelled".'])
expect(ceoFirstReviewAUQ(change(q => { for (const option of q.options) option.description += '\n' + status; }))).toBe(false);
expect(ceoFirstReviewAUQ(change(q => { q.options = [{ label: 'A) Keep existing behavior', description: 'Leave the implementation unchanged.' }, { label: 'B) Archive the report', description: 'Save the review text.' }]; }))).toBe(false);
});
test('a current native completion, selected answer and consistent decision are still required', () => {
for (const edit of [
(_q: any, c: any) => { c.answered = false; }, (_q: any, c: any) => { c.failed = true; },
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; }, (_q: any, c: any) => { delete c.answeredAt; },
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; }, (_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
(q: any) => { q.multiSelect = true; }, (q: any) => { q.header = 'Approach'; }, (q: any) => { q.header = 'Issue 99'; },
(q: any) => { q.question = q.question.replace('D2 —', 'D02 —'); },
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: Z'); },
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: 9A'); for (const option of q.options) option.label = '9' + option.label; },
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', '3B)'); },
(q: any) => { q.question = q.question.replace('plan-ceo-review-email-rescue', 'plan-ceo-review-setup'); },
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10 omitted: '); },
(q: any) => { q.question = q.question.replace(/\nProject\/branch\/task:[^\n]+/, ''); },
]) expect(ceoFirstReviewAUQ(change(edit))).toBe(false);
const noAnswer = structuredClone(first); noAnswer.nativeCall!.answers = {}; expect(ceoFirstReviewAUQ(noAnswer)).toBe(false);
const menu = structuredClone(first); menu.options[0]!.label = 'Unowned'; expect(ceoFirstReviewAUQ(menu)).toBe(false);
expect(ceoFirstReviewAUQ({ ...first, nativeCall: undefined })).toBe(false);
});
test('only the existing dense CEO finding owner selects the retry regression', () => {
for (const path of ['test/ceo-current-omission-ap.test.ts', 'test/fixtures/ceo-current-omission-ap.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([, files]) => files.includes(path)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
for (let i = 0; i < paths.length; i++) { expect(Object.hasOwn(paths, i)).toBe(true); expect(typeof paths[i]).toBe('string'); }
});
+63
View File
@@ -0,0 +1,63 @@
import {expect, test} from 'bun:test';
import {ceoFirstReviewAUQ, nativePlanCallFingerprint} from './helpers/claude-pty-runner';
import fixture from './fixtures/ceo-decision-prefix-al.json';
import {E2E_TOUCHFILES} from './helpers/touchfiles-data';
const call=(n=0):any=>structuredClone(fixture.calls[n]);
const accepts=(c:any)=>ceoFirstReviewAUQ(nativePlanCallFingerprint(c,0,true));
function text(c:any,fn:(s:string)=>string){const q=c.questions[0],a=c.answers[q.question];q.question=fn(q.question);c.answers={[q.question]:a};}
function menu(c:any,fn:(o:any,i:number)=>void){const q=c.questions[0],i=q.options.findIndex((o:any)=>o.label===c.answers[q.question]);q.options.forEach(fn);c.answers={[q.question]:q.options[i].label};}
test('the actual completed email finding uses decision-prefixed options and a bare recommendation',()=>expect(accepts(call())).toBe(true));
test('the actual completed SQL finding includes a raw SQL qualifier',()=>expect(accepts(call(1))).toBe(true));
test('decision and finding identifiers remain independent when consistently renamed',()=>{
for(const n of [0,1]){
for(const dotted of [false,true]){const c=call(n);text(c,s=>s.replace(/^D\d+/,'D27').replace(/\(Finding \d+\)/,`(Finding ${dotted?'8.3':'8'})`));menu(c,o=>{o.label=o.label.replace(/^\d+/,'27')});expect(accepts(c)).toBe(true);}
const c=call(n);menu(c,o=>{o.label=o.label.replace(/^\d+/,'')});text(c,s=>s.replace("'no error handling on the email leg'",'no error handling on the email leg').replace('a raw SQL fragment','a SQL fragment'));expect(accepts(c)).toBe(true);
const q=call(n);text(q,s=>s+'\nOld note: "This finding is withdrawn."');expect(accepts(q)).toBe(true);
const lower=call(n);text(lower,s=>s.replace(/^D/,'d'));expect(accepts(lower)).toBe(true);
}
});
test('native ownership, offered answers and unambiguous decision identities are mandatory',()=>{
for(const mutate of [
(c:any)=>{c.answered=false},(c:any)=>{c.failed=true},(c:any)=>{c.unansweredQuestionIndices=[0]},(c:any)=>{c.sessionId=''},
(c:any)=>{c.answers={}},(c:any)=>{c.answers[c.questions[0].question]='A'},(c:any)=>{c.questions[0].multiSelect=true},
(c:any)=>{c.questions[0].header='Finding 9'},(c:any)=>{c.questions[0].header='Approach'},
(c:any)=>text(c,s=>s.replace(/^D4/,'D9')),
(c:any)=>menu(c,o=>{o.label=o.label.replace(/^4/,'9')}),
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4/,'9')}),
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4/,'')}),
(c:any)=>text(c,s=>s.replace(/^Recommendation: A/m,'Recommendation: 9A')),
(c:any)=>text(c,s=>s.replace(/^Recommendation: A/m,'Recommendation: Z')),
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4B/,'4A')}),
(c:any)=>{c.questions[0].options[1].description=''},
]){const c=call();mutate(c);expect(accepts(c)).toBe(false);}
const f=nativePlanCallFingerprint(call(),0,true);expect(ceoFirstReviewAUQ({...f,signature:'foreign:tool'})).toBe(false);
});
test('embedded quoted contract terms cannot supply a hypothetical, historical or withdrawn assessment',()=>{
for(const n of [0,1])for(const fn of [
(s:string)=>'Source excerpt: '+s,(s:string)=>'> '+s,(s:string)=>'```\n'+s+'\n```',
(s:string)=>s.replace(/^ELI10: (.+)$/m,'ELI10: "$1"'),
(s:string)=>s.replace(/^ELI10: /m,'ELI10: If approved, '),
(s:string)=>s.replace(/^ELI10: /m,'ELI10: Source excerpt. '),
(s:string)=>s.replace(/^ELI10: /m,'ELI10: The following is a historical source excerpt. '),
(s:string)=>s.replace(/^ELI10: /m,'ELI10: Previously, '),
(s:string)=>s+'\nThis finding is withdrawn.',
(s:string)=>s+'\nNo current defect remains.',
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan sends the email inline with \'no error handling\' only in a historical example.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan does not send the email inline with \'no error handling\'.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan used to paste the user ID straight into a raw SQL fragment. The current query is parameterized.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: A proposed example pastes the user ID string straight into a raw SQL fragment.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The historical example pastes the user ID string straight into a raw SQL fragment.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: A template pastes the user ID string straight into a raw SQL fragment.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: An unrelated example pastes the user ID string straight into a raw SQL fragment.'),
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan pastes the user ID string straight into a raw SQL fragment only in a hypothetical example.'),
]){const c=call(n);text(c,fn);expect(accepts(c)).toBe(false);}
});
test('a substantive current assessment still needs an offered technical amendment',()=>{
for(const description of ['Archive this report.','If approved: ✅ Rescue named mail exceptions.','Source excerpt: ✅ Rescue named mail exceptions.','❌ Rescue named mail exceptions.','✅ "Rescue named mail exceptions."']){
const c=call();menu(c,(o,i)=>{o.label=`4${String.fromCharCode(65+i)}: Consider candidate ${i}`;o.description=description});expect(accepts(c)).toBe(false);
}
});
test('only the existing CEO count owner selects the captured regression',()=>{
for(const dependency of E2E_TOUCHFILES['plan-ceo-finding-count']) expect(typeof dependency).toBe('string');
for(const d of ['test/ceo-decision-prefix-al.test.ts','test/fixtures/ceo-decision-prefix-al.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,v])=>v.includes(d)).map(([k])=>k)).toEqual(['plan-ceo-finding-count']);
});
+111
View File
@@ -0,0 +1,111 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
import { isCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import fixture from './fixtures/ceo-declarative-premise-ap.json';
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
const first = calls[2]!;
function change(fp: AskUserQuestionFingerprint, edit: (q: any, call: any, fp: any) => void) {
const copy = structuredClone(fp), call = copy.nativeCall!, q = call.questions[0]!;
const selected = q.options.findIndex(o => o.label === call.answers?.[q.question]);
edit(q, call, copy);
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
return copy;
}
test('the exact completed defect premises start review; later calls use unchanged phase continuation', () => {
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, true, true, false, false]);
let started = false;
const phases = calls.map(fp => {
const phase = planCountQuestionPhase(fp, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
started = phase.reviewStarted;
return phase;
});
expect(fixture.actualCounts).toEqual({ setup: 6, review: 0 });
expect(phases.map(p => p.preReview)).toEqual([true, true, false, false, false, false]);
expect(phases.filter(p => !p.preReview && !p.administrative)).toHaveLength(4);
});
test('current metadata, premise and explanation cannot borrow quoted, historical or conditional authority', () => {
for (const fp of calls.slice(2, 4)) {
for (const prefix of ['Source:', 'Earlier review assessment:', 'If approved:', 'Example:'])
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
for (const prefix of ['If approved, ', 'Source excerpt: ', 'Earlier review assessment: ', 'The following is a hypothetical example. ', 'Previously, ', 'Formerly, ']) {
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: ' + prefix); }))).toBe(false);
}
for (const replacement of ['Source: The ', 'If approved, the ', 'The historical ', 'The quoted ', 'The previously '])
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('— The ', '— ' + replacement); }))).toBe(false);
for (const wrap of [(s: string) => `"${s}"`, (s: string) => '`' + s + '`', (s: string) => '> ' + s])
expect(ceoFirstReviewAUQ(change(fp, q => { const title = q.question.split('\n')[0]; q.question = q.question.replace(title, wrap(title)); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/\nProject\/branch\/task:[^\n]+/, ''); }))).toBe(false);
}
});
test('a current finding and its offered amendments cannot be withdrawn', () => {
for (const fp of calls.slice(2, 4)) {
for (const status of ['This finding is withdrawn.', 'This issue is "rejected".', 'Correction: this assessment is not current.', 'There is no current gap.'])
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + status; }))).toBe(false);
for (const prefix of ['Source excerpt: ', 'If approved later: ', 'Previously, ', 'Formerly, '])
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description = prefix + option.description; }))).toBe(false);
for (const status of ['This amendment is withdrawn.', 'This remedy is "cancelled".'])
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description += '\n' + status; }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description = JSON.stringify(option.description); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:'); }))).toBe(true);
}
});
test('statement-only and administrative menus are not review decisions', () => {
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace(' How should the handler treat a mail failure?', ''); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace(' How should the handler treat a mail failure?', ' Record this in the report.'); }))).toBe(false);
expect(ceoFirstReviewAUQ(change(first, q => {
q.options = [
{ label: '2A) Keep the existing implementation', description: 'Leave current behavior unchanged.' },
{ label: '2B) Archive the report', description: 'Save the existing review text.' },
];
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(first, q => { q.header = 'Approach'; }))).toBe(false);
});
test('the native completion, selected answer and issue identities stay bound', () => {
for (const edit of [
(_q: any, c: any) => { c.answered = false; },
(_q: any, c: any) => { c.failed = true; },
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; },
(_q: any, c: any) => { delete c.answeredAt; },
(_q: any, c: any) => { c.answeredAt = 'not-a-time'; },
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; },
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
(q: any) => { q.multiSelect = true; },
(q: any) => { q.header = 'Issue 99'; },
(q: any) => { q.header = 'Issue 0'; },
(q: any) => { q.header = 'Issue 02'; },
(q: any) => { q.question = q.question.replace('(Issue 2)', '(Issue 0)'); },
(q: any) => { q.question = q.question.replace('D4 (', 'D04 ('); },
(q: any) => { q.question = q.question.replace('Recommendation: 2A', 'Recommendation: 9A'); },
(q: any) => { q.options[1].label = q.options[1].label.replace('2B)', '3B)'); },
]) expect(ceoFirstReviewAUQ(change(first, edit))).toBe(false);
const wrongAnswer = structuredClone(first); wrongAnswer.nativeCall!.answers = {};
expect(ceoFirstReviewAUQ(wrongAnswer)).toBe(false);
const wrongMenu = structuredClone(first); wrongMenu.options[0]!.label = 'Foreign';
expect(ceoFirstReviewAUQ(wrongMenu)).toBe(false);
expect(ceoFirstReviewAUQ({ ...first, nativeCall: undefined })).toBe(false);
});
test('equivalent current wording and descriptive or matching issue headers preserve the decision', () => {
for (const header of ['Email contract', 'Issue 2', 'Finding 2'])
expect(ceoFirstReviewAUQ(change(first, q => { q.header = header; }))).toBe(true);
for (const separator of ['—', '', '-'])
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace('D4 (Issue 2) —', `D19 (Issue 2) ${separator}`); }))).toBe(true);
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace('How should the handler treat a mail failure?', 'Which handling should the current implementation use?'); }))).toBe(true);
expect(ceoFirstReviewAUQ(change(calls[3]!, q => { q.question = q.question.replace('request.params.userId', 'payload.accountId'); }))).toBe(true);
});
test('the regression fixture is registered only to the dense CEO finding owner', () => {
for (const path of ['test/ceo-declarative-premise-ap.test.ts', 'test/fixtures/ceo-declarative-premise-ap.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([, files]) => files.includes(path)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
for (let i = 0; i < paths.length; i++) { expect(Object.hasOwn(paths, i)).toBe(true); expect(typeof paths[i]).toBe('string'); }
});
+127
View File
@@ -0,0 +1,127 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { findCeoModeOption, hasNativePostAnswerCeoPosture } from './helpers/ceo-mode-option';
import { parseNumberedOptions } from './helpers/claude-pty-runner';
import { readPlanCountTranscript, type NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/ceo-expansion-auq-ac.json';
const posture = /\b(expansion|10x|delight|dream|cathedral|opt[\s-]?in)\b/i;
const selectedAt = Date.parse(captured.provenance.testSelectionLowerBound.at);
type Records = typeof captured.records;
function replay(change?: (records: Records) => void) {
const records = structuredClone(captured.records);
change?.(records);
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-expansion-auq-'));
const first = records[0]!;
const project = path.join(root, 'projects', 'fixture');
fs.mkdirSync(project, { recursive: true });
fs.writeFileSync(path.join(project, `${first.sessionId}.jsonl`),
records.map(record => JSON.stringify(record)).join('\n') + '\n');
const events: NativePublicToolEvent[] = [];
try {
return { transcript: readPlanCountTranscript(root, first.cwd, event => events.push(event)), events };
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
}
function matches(evidence = replay()) {
return hasNativePostAnswerCeoPosture(evidence.transcript, 'SCOPE EXPANSION', posture, selectedAt, evidence.events);
}
describe('CEO expansion posture in a completed native decision brief', () => {
test('captured option 2 and subsequent answered expansion establish posture without standalone prose', () => {
const evidence = replay();
expect(findCeoModeOption(parseNumberedOptions(captured.visibleAtModeTail), 'SCOPE EXPANSION')).toBe(2);
expect(evidence.transcript.calls.map(call => call.toolUseId)).toEqual([
captured.modeToolUseId, captured.expansionToolUseId,
]);
expect(evidence.transcript.assistantMessages.every(message => Date.parse(message.timestamp) < selectedAt)).toBe(true);
expect(evidence.events.map(event => [event.toolUseId, event.timestamp])).toEqual([
[captured.modeToolUseId, captured.requestReplyTimes[captured.modeToolUseId].request],
[captured.modeToolUseId, captured.requestReplyTimes[captured.modeToolUseId].reply],
[captured.expansionToolUseId, captured.requestReplyTimes[captured.expansionToolUseId].request],
[captured.expansionToolUseId, captured.requestReplyTimes[captured.expansionToolUseId].reply],
]);
expect(matches(evidence)).toBe(true);
// Without the public native request/reply timestamps, no new evidence is inferred.
expect(hasNativePostAnswerCeoPosture(evidence.transcript, 'SCOPE EXPANSION', posture, selectedAt)).toBe(false);
});
test.each([
'wrong selected mode', 'pending selected mode', 'failed selected mode', 'selected answer before selection',
'pending expansion', 'failed expansion', 'unrecognized expansion answer', 'foreign session',
'pre-mode request', 'request after reply', 'reply timestamp mismatch', 'missing request', 'missing reply',
'failed public reply', 'wrong tool name', 'foreign tool identity', 'conflicting request', 'conflicting reply',
'request metadata mismatch', 'batched questions', 'multiselect', 'setup heading', 'quoted brief',
'fenced brief', 'quoted expansion marker', 'wrong expansion choices',
])('%s cannot supply posture coverage', failure => {
const evidence = replay();
const [mode, expansion] = evidence.transcript.calls;
const request = evidence.events[2]!;
const reply = evidence.events[3]!;
const question = expansion!.questions[0]!;
switch (failure) {
case 'wrong selected mode': mode!.answers![mode!.questions[0]!.question] = 'HOLD SCOPE'; break;
case 'pending selected mode': mode!.answered = false; break;
case 'failed selected mode': mode!.failed = true; break;
case 'selected answer before selection': mode!.answeredAt = '2026-09-09T16:40:00.000Z'; break;
case 'pending expansion': expansion!.answered = false; break;
case 'failed expansion': expansion!.failed = true; break;
case 'unrecognized expansion answer': expansion!.answers![question.question] = 'invented approval'; break;
case 'foreign session': expansion!.sessionId = request.sessionId = reply.sessionId = 'unrelated-session'; break;
case 'pre-mode request': request.timestamp = captured.requestReplyTimes[captured.modeToolUseId].request; break;
case 'request after reply': request.timestamp = '2026-09-09T16:41:22.000Z'; break;
case 'reply timestamp mismatch': reply.timestamp = '2026-09-09T16:41:22.000Z'; break;
case 'missing request': evidence.events.splice(2, 1); break;
case 'missing reply': evidence.events.splice(3, 1); break;
case 'failed public reply': reply.isError = true; break;
case 'wrong tool name': request.name = 'Read'; break;
case 'foreign tool identity': request.toolUseId = 'unrelated-call'; break;
case 'conflicting request': evidence.events.push({ ...request, timestamp: '2026-09-09T16:41:20.000Z' }); break;
case 'conflicting reply': evidence.events.push({ ...reply, isError: true }); break;
case 'request metadata mismatch': request.input = { questions: [] }; break;
case 'batched questions': expansion!.questions.push(structuredClone(question)); break;
case 'multiselect': question.multiSelect = true; break;
case 'setup heading': question.question = question.question.replace(/^D5[^\n]+/, 'D5 — Choose the review mode?'); break;
case 'quoted brief': question.question = question.question.split('\n').map(line => `> ${line}`).join('\n'); break;
case 'fenced brief': question.question = '```text\n' + question.question + '\n```'; break;
case 'quoted expansion marker': question.question = 'D5 — Which choice?\n> Expansion 1 of 8: scope opt-in'; break;
case 'wrong expansion choices': question.options[1]!.label = 'Enable telemetry'; break;
}
// Keep parsed call metadata bound to the public request. This makes the
// content negatives exercise semantic guards, not an accidental mismatch.
if (['batched questions', 'multiselect', 'setup heading', 'quoted brief', 'fenced brief',
'quoted expansion marker', 'wrong expansion choices'].includes(failure)) {
request.input = { questions: expansion!.questions };
expansion!.answers = { [question.question]: question.options[0]!.label };
}
expect(matches(evidence)).toBe(false);
});
test('the original caller regex and mode remain required', () => {
const { transcript, events } = replay();
expect(hasNativePostAnswerCeoPosture(transcript, 'SCOPE EXPANSION', /\bcathedral\b/i, selectedAt, events)).toBe(false);
expect(hasNativePostAnswerCeoPosture(transcript, 'HOLD SCOPE', /\bhold\s*scope\b/i, selectedAt, events)).toBe(false);
expect(hasNativePostAnswerCeoPosture(transcript, 'SCOPE EXPANSION', posture, Date.now(), events)).toBe(false);
});
test('reader-level foreign, sidechain, incomplete and failed records cannot create native evidence', () => {
for (const change of [
(records: Records) => { records[5]!.cwd = '/unrelated'; },
(records: Records) => { records[5]!.isSidechain = true; },
(records: Records) => { records.splice(6, 1); },
(records: Records) => { records[6]!.message.content[0]!.is_error = true; },
]) expect(matches(replay(change))).toBe(false);
});
test('the new free test and captured fixture select the paid mode-routing case', () => {
for (const file of ['test/ceo-expansion-auq.test.ts', 'test/fixtures/ceo-expansion-auq-ac.json']) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-ceo-mode-routing']);
}
});
});
+129
View File
@@ -0,0 +1,129 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import captured from './fixtures/ceo-finding-brief-ak.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const call = (index = 4): any => structuredClone(captured.calls[index]);
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
function edit(c: any, change: (s: string) => string) {
const q = c.questions[0], answer = c.answers[q.question];
q.question = change(q.question); c.answers = { [q.question]: answer };
}
function offered(c: any, change: (o: any, i: number) => void) {
const q = c.questions[0], selected = q.options.findIndex((o: any) => o.label === c.answers[q.question]);
q.options.forEach(change); c.answers = { [q.question]: q.options[selected].label };
}
test('the completed parenthesized finding with letter-only choices starts current review', () => {
expect(ceoFirstReviewAUQ(fp(call()))).toBe(true);
});
test('the exact retry phase preserves four setup calls then six substantive choices', () => {
let started = false;
const phases = captured.calls.map(c => {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted; return phase.preReview;
});
expect(phases).toEqual([true, true, true, true, false, false, false, false, false, false]);
});
test('the finding identity is independent of decision number, separator and optional qid', () => {
for (const change of [
(s: string) => s.replace(/^D5/, 'D19'),
(s: string) => s.replace(') — ', ') - '),
(s: string) => s.replace('Finding 1.1', 'Finding 9.4'),
(s: string) => s.replace('Finding 1.1', 'Finding 1'),
(s: string) => s.replace(/\s*<gstack-qid:[^>]+>\s*$/, ''),
(s: string) => s.replace(/\s*<gstack-qid:[^>]+>\s*$/, '') + '\n<gstack-qid:plan-ceo-review-finding-mail>',
(s: string) => s.replace('lets any mail failure', 'allows any mail failure'),
]) { const c = call(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(true); }
for (const label of call().questions[0].options.map((o: any) => o.label)) {
const c = call(); c.answers[c.questions[0].question] = label;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
const numbered = call(); edit(numbered, s => s.replace(/^Recommendation: A/m, 'Recommendation: 1A'));
offered(numbered, o => { o.label = o.label.replace(/^([A-Z])\)/, '1$1)'); });
expect(ceoFirstReviewAUQ(fp(numbered))).toBe(true);
const header = call(); header.questions[0].header = 'Finding 1.1';
expect(ceoFirstReviewAUQ(fp(header))).toBe(true);
});
test('competing finding, section, recommendation and offered choice identities are rejected', () => {
for (const mutate of [
(c: any) => { c.questions[0].header = 'Finding 1'; },
(c: any) => { c.questions[0].header = 'Finding 9.1'; },
(c: any) => { c.questions[0].header = 'Approach'; },
(c: any) => edit(c, s => s.replace('Finding 1.1', 'Finding 1.0')),
(c: any) => edit(c, s => s.replace('Finding 1.1', 'Finding 1.1 and Finding 2.1')),
(c: any) => edit(c, s => s.replace(/^Recommendation: A/m, 'Recommendation: 2A')),
(c: any) => edit(c, s => s.replace(/^Recommendation: A/m, 'Recommendation: Z')),
(c: any) => { c.questions[0].options[0].label = '9A) Foreign issue'; c.answers = { [c.questions[0].question]: c.questions[0].options[0].label }; },
(c: any) => { c.questions[0].options[1].label = 'A) Same choice letter, different action'; },
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[0].label; },
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-eng-review-finding-mail>'),
]) { const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
});
test('a current complete assessment cannot come from source, conditions or a withdrawal', () => {
for (const change of [
(s: string) => 'Example: ' + s,
(s: string) => '> ' + s,
(s: string) => '```\n' + s + '\n```',
(s: string) => s.replace(/^ELI10: .+$/m, ''),
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If approved, '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a quoted source excerpt. '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
(s: string) => s + '\nThis finding is withdrawn.',
(s: string) => s + '\nFinding 1.1 is rejected.',
(s: string) => s + '\nNo current defect remains.',
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan no longer lets mail failures escape the handler. The current named rescue keeps them contained.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan used to let mail failures escape the handler. That was the prior behavior.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan does not let mail failures escape the handler.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Source excerpt: the plan lets mail failures escape the handler.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Previously, the plan lets mail failures escape the handler.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan lets no mail failure escape the handler.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan allows mail failures to escape only in a historical quoted example.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Source excerpt. The plan lets mail failures escape the handler.'),
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan allows mail failures to never escape the handler.'),
]) { const c = call(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
const c = call(); edit(c, s => s + '\nOld note: "Finding 1.1 is rejected."');
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
});
test('current finding prose must offer an actual remedy, not an advisory or hypothetical action', () => {
for (const description of [
'Archive this review for later.',
'Historical source excerpt: ✅ Rescue named mail exceptions.',
'The following is a quoted source excerpt. ✅ Rescue named mail exceptions.',
'If approved: ✅ Rescue named mail exceptions.',
'❌ Rescue named mail exceptions.',
'✅ "Rescue named mail exceptions."',
'✅ If approved, rescue named mail exceptions.',
]) {
const c = call(); offered(c, (o, i) => { o.label = `${String.fromCharCode(65 + i)}) Consider candidate ${i}`; o.description = description; });
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
});
test('the completed native call, exact offered answer and fingerprint remain mandatory', () => {
for (const mutate of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.sessionId = ''; },
(c: any) => { c.toolUseId = ''; },
(c: any) => { c.answers = {}; },
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
(c: any) => { c.questions[0].multiSelect = true; },
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
(c: any) => { c.questions[0].options[1].description = ''; },
]) { const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
const f = fp(call());
expect(ceoFirstReviewAUQ({ ...f, signature: 'foreign:tool' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...f, nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...f, options: f.options.slice(1) })).toBe(false);
});
test('retry fixture and controls select only the existing CEO count owner', () => {
for (const dependency of ['test/ceo-finding-brief-ak.test.ts', 'test/fixtures/ceo-finding-brief-ak.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name)).toEqual(['plan-ceo-finding-count']);
}
});
+91
View File
@@ -0,0 +1,91 @@
import {describe,expect,test} from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import fixture from './fixtures/ceo-handoff-y-call.json';
import zFixture from './fixtures/ceo-handoff-z-call.json';
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
import {ceoFirstReviewAUQ,ceoStep0Boundary,hasNativePlanTerminal,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
import {isCeoCompletionHandoff,pickCeoCompletionHandoff} from './helpers/ceo-completion-handoff';
const actual=()=>structuredClone(fixture.calls.at(-1)!) as NativePlanQuestionCall;
const fp=(c:NativePlanQuestionCall)=>nativePlanCallFingerprint(c,0,false);
const pending=(c:NativePlanQuestionCall)=>{c.answered=false;delete c.answers;delete c.unansweredQuestionIndices;return fp(c);};
function change(c:NativePlanQuestionCall,fn:(s:string)=>string){const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};return c;}
describe('Y bare next-Eng navigation is administrative, not completion evidence',()=>{
test('exact four issues remain while a closed next-workflow menu cannot start review',()=>{
const c=actual();expect(isCeoCompletionHandoff(fp(c))).toBe(true);expect(pickCeoCompletionHandoff(pending(actual()))).toBe(2);
expect(pickCeoCompletionHandoff(fp(c))).toBeNull();expect(c.answers![c.questions[0]!.question]).toBe('A) Run /plan-eng-review next (recommended)');
let started=false;let setup=0,review=0,admin=0;
for(const c of fixture.calls){const p=planCountQuestionPhase(fp(structuredClone(c) as NativePlanQuestionCall),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=p.reviewStarted;if(p.administrative)admin++;else if(p.preReview)setup++;else review++;}
expect({setup,review,admin}).toEqual({setup:2,review:4,admin:1});
expect(planCountQuestionPhase(fp(actual()),false,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff)).toEqual({preReview:false,reviewStarted:false,administrative:'completion-handoff'});
});
test('either offered navigation answer and option order preserve administrative meaning',()=>{
const c=actual();c.questions[0]!.options.reverse();
for(const o of c.questions[0]!.options){c.answers={[c.questions[0]!.question]:o.label};expect(isCeoCompletionHandoff(fp(c))).toBe(true);}
expect(pickCeoCompletionHandoff(pending(c))).toBe(1);
expect(isCeoCompletionHandoff(fp(change(actual(),s=>s.replace('D7 - Next step: run','D17 — Next review: Run').replace('plan-ceo-review-next-step','plan-ceo-review-next-review'))))).toBe(true);
});
test('whole question and description boundaries reject added product work and unfinished choices',()=>{
for(const fn of [(s:string)=>s.replace('run /plan-eng-review?', 'fix the cache before /plan-eng-review?'),(s:string)=>s.replace('run /plan-eng-review?', 'run /plan-eng-review? Also repair the cache.'),(s:string)=>'> '+s,(s:string)=>'Example: '+s,(s:string)=>s.replace('plan-ceo-review-next-step','foreign-next-step'),(s:string)=>s+' <gstack-qid:plan-ceo-review-next-step>'])expect(isCeoCompletionHandoff(fp(change(actual(),fn)))).toBe(false);
for(const i of [0,1])for(const extra of [' Also implement a new cache.',' Resolve the remaining CEO decisions first.',' Should we add another requirement?']){const c=actual();c.questions[0]!.options[i]!.description+=extra;expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
for(const text of ['Resume the unfinished CEO review.','Proceed directly to implementation and add the missing test.','Eng review is optional.']){const c=actual();c.questions[0]!.options[1]!.description=text;expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
});
test('native completion, current offered answer and pending identity remain required',()=>{
for(const mutate of [(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions[0]!.header='Issue';},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'Fix another issue'};}]){const c=actual();mutate(c);expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
expect(isCeoCompletionHandoff({...fp(actual()),signature:'foreign:call'})).toBe(false);expect(isCeoCompletionHandoff({...fp(actual()),options:[]})).toBe(false);
expect(pickCeoCompletionHandoff({...pending(actual()),nativeCall:undefined})).toBeNull();expect(pickCeoCompletionHandoff({...pending(actual()),signature:'foreign:call'})).toBeNull();
});
test('independent fresh report and native Exit still gate completion; menu alone cannot pass',()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-handoff-y-free-'));const report=path.join(dir,'report.md');
try{fs.writeFileSync(report,fixture.report);const calls=structuredClone(fixture.calls) as NativePlanQuestionCall[];const transcript={status:'ready' as const,calls,assistantMessages:[],planReadyRequests:structuredClone(fixture.planReadyRequests)};const handoff=calls.at(-1)!;const admin=new Set([fp(handoff).signature]);const issueAt=Date.parse(calls.at(-2)!.answeredAt!),handoffAt=Date.parse(handoff.answeredAt!);const started=Date.parse(calls[0]!.answeredAt!)-1000;
// Controlled metadata only: original Y report mtime was not captured.
const between=(issueAt+handoffAt)/2;fs.utimesSync(report,between/1000,between/1000);
expect(hasNativePlanTerminal(transcript,report,started,'plan_ready')).toBe(false);expect(hasNativePlanTerminal(transcript,report,started,'plan_ready',admin)).toBe(true);
fs.utimesSync(report,(issueAt-1)/1000,(issueAt-1)/1000);expect(hasNativePlanTerminal(transcript,report,started,'plan_ready',admin)).toBe(false);
fs.utimesSync(report,between/1000,between/1000);expect(hasNativePlanTerminal({...transcript,planReadyRequests:[]},report,started,'plan_ready',admin)).toBe(false);
expect(hasNativePlanTerminal({...transcript,calls:[handoff]},report,started,'plan_ready',admin)).toBe(false);
}finally{fs.rmSync(dir,{recursive:true,force:true});}
});
});
describe('Z completed CEO with an unrun required Eng gate',()=>{
const actualZ=()=>structuredClone(zFixture.calls.at(-1)!) as NativePlanQuestionCall;
const pendingZ=(c=actualZ())=>{c.answered=false;delete c.answers;delete c.answeredAt;delete c.unansweredQuestionIndices;return c;};
const reject=(c:NativePlanQuestionCall)=>{expect(isCeoCompletionHandoff(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();};
test('exact six calls preserve two findings; handoff selects the offered manual route',()=>{
let started=false;const counts={setup:0,review:0,admin:0};
for(const c of zFixture.calls){const phase=planCountQuestionPhase(fp(c as NativePlanQuestionCall),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=phase.reviewStarted;counts[phase.administrative?'admin':phase.preReview?'setup':'review']++;}
expect(counts).toEqual({setup:3,review:2,admin:1});expect(isCeoCompletionHandoff(fp(actualZ()))).toBe(true);expect(pickCeoCompletionHandoff(fp(pendingZ()))).toBe(2);
expect(planCountQuestionPhase(fp(actualZ()),false,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff)).toEqual({preReview:false,reviewStarted:false,administrative:'completion-handoff'});
});
test('number, typography and option order are not semantic requirements',()=>{
const c=change(actualZ(),s=>s.replace('D5 —','D27:').replace("hasn't",'has not').replace("What's",'What is'));c.questions[0]!.options.reverse();
for(const option of c.questions[0]!.options){c.answers={[c.questions[0]!.question]:option.label};expect(isCeoCompletionHandoff(fp(c))).toBe(true);}
expect(pickCeoCompletionHandoff(fp(pendingZ(c)))).toBe(1);
});
test('whole question and role-specific descriptions cannot hide new or conditional work',()=>{
for(const fn of [(s:string)=>s.replace('CEO Review is CLEAR','If CEO Review is CLEAR'),(s:string)=>s.replace('CEO Review is CLEAR','CEO Review is not CLEAR'),(s:string)=>s.replace('required shipping gate','optional shipping gate'),(s:string)=>s.replace("What's next?","What's next? Also add retries."),(s:string)=>'> '+s,(s:string)=>'Example: '+s,(s:string)=>'```\n'+s+'\n```',(s:string)=>s.replace('plan-ceo-next-review','foreign-next-review'),(s:string)=>s+' <gstack-qid:plan-ceo-next-review>'])reject(change(actualZ(),fn));
for(const i of [0,1])for(const extra of [' Also implement the missing checks.',' Rotate credentials.',' Should we add a new requirement?',' Once remaining findings are fixed.']){const c=actualZ();c.questions[0]!.options[i]!.description+=extra;reject(c);}
const swapped=actualZ();[swapped.questions[0]!.options[0]!.description,swapped.questions[0]!.options[1]!.description]=[swapped.questions[0]!.options[1]!.description,swapped.questions[0]!.options[0]!.description];reject(swapped);
const optional=actualZ();optional.questions[0]!.options[1]!.description=optional.questions[0]!.options[1]!.description!.replace('required before shipping','optional before shipping');reject(optional);
});
test('new arm requires explicit native completion and exact producer pending state',()=>{
const mutations=[(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.questions[0]!.header='Issue';},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.options.push({label:'Add a repair',description:'Add a new requirement.'});}];
for(const mutate of mutations){const c=actualZ();mutate(c);reject(c);const p=pendingZ();mutate(p);reject(p);}
for(const mutate of [(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'Repair first'};}]){const c=actualZ();mutate(c);reject(c);}
for(const mutate of [(c:NativePlanQuestionCall)=>{delete (c as any).answered;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[];},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[1];},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answeredAt='2026-09-09T12:00:00Z';}]){const c=pendingZ();mutate(c);reject(c);}
const projected=pendingZ();projected.unansweredQuestionIndices=[0];expect(pickCeoCompletionHandoff(fp(projected))).toBe(2);
for(const variant of [{...fp(pendingZ()),signature:'foreign:call'},{...fp(pendingZ()),options:[]},{...fp(pendingZ()),nativeQuestionIndex:1}])expect(pickCeoCompletionHandoff(variant)).toBeNull();
});
test('retained original mtime passes only with the administrative handoff and real Exit',()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-handoff-z-free-'));const report=path.join(dir,'report.md');
try{fs.writeFileSync(report,zFixture.report);const calls=structuredClone(zFixture.calls) as NativePlanQuestionCall[];const transcript={status:'ready' as const,calls,assistantMessages:[],planReadyRequests:structuredClone(zFixture.planReadyRequests)};const admin=new Set(calls.filter(c=>isCeoCompletionHandoff(fp(c))).map(c=>fp(c).signature));const mtime=Number(BigInt(zFixture.reportOriginalMtimeNs))/1e6;fs.utimesSync(report,mtime/1000,mtime/1000);
expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready')).toBe(false);expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready',admin)).toBe(true);
expect(hasNativePlanTerminal({...transcript,planReadyRequests:[]},report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
expect(hasNativePlanTerminal({...transcript,calls:[calls.at(-1)!]},report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
const stale=Date.parse(calls.at(-2)!.answeredAt!)-1;fs.utimesSync(report,stale/1000,stale/1000);expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
}finally{fs.rmSync(dir,{recursive:true,force:true});}
});
});
+105
View File
@@ -0,0 +1,105 @@
import { expect, test } from 'bun:test';
import { hasNativePostAnswerCeoPosture, nativeCeoModeAnswer } from './helpers/ceo-mode-option';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/ceo-hold-commitment-ar.json';
const posture = /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i;
const original = captured.transcript.assistantMessages[0]!.text;
const replay = () => structuredClone(captured.transcript) as PlanCountTranscript;
const matches = (transcript = replay()) => hasNativePostAnswerCeoPosture(
transcript, 'HOLD SCOPE', posture, captured.selectionStartedAt,
);
const withText = (text: string) => { const t = replay(); t.assistantMessages[0]!.text = text; return matches(t); };
test('the actual failed attempt adopted HOLD through scope, hardening and exclusion', () => {
expect(captured.provenance.actualState).toBe('failed');
expect(nativeCeoModeAnswer(replay(), 'HOLD SCOPE', captured.selectionStartedAt)?.toolUseId)
.toBe('toolu_01E1HnYjRCz79826bo7nNnoK');
expect(posture.test(original)).toBe(false);
expect(matches()).toBe(true);
// This is prospective posture recognition, not evidence of completed work.
for (const prefix of ["I'm keeping", 'I am keeping', 'I will keep', "We'll keep", 'We will keep', 'We are keeping']) {
expect(withText(original.replace("I'll keep", prefix)), prefix).toBe(true);
}
expect(withText(original.replace("I'll", 'Ill').replace("PLAN.md's", 'PLAN.mds'))).toBe(true);
});
test('explicitly future, conditional and quoted statements are not adopted current posture', () => {
for (const text of [
original.replace("I'll keep", 'I will later keep'),
original.replace("I'll keep", 'I will eventually keep'),
original.replace("I'll keep", 'I would keep'),
original.replace("I'll keep", 'I may keep'),
original.replace("I'll keep", "I'll not keep"),
original.replace('scope fixed', 'scope tomorrow fixed'),
original.replace('production visibility', 'production visibility next week'),
original.replace('production visibility', 'production visibility tomorrow'),
...['after approval', 'once approved', 'when approved', 'after launch', 'pending approval', 'subject to approval'].map(when =>
original.replace('production visibility', 'production visibility ' + when)),
'Later, ' + original, 'If you approve, ' + original,
'Hypothetical scenario. ' + original, 'Example only: ' + original,
'"' + original + '"', '> ' + original,
'```text\n' + original + '\n```', '~~~text\n' + original + '\n~~~',
'Read(file)\n' + original, 'The user said: ' + original,
]) expect(withText(text), text).toBe(false);
});
test('all three obligations remain concrete and bound to the selected plan', () => {
for (const [from, to] of [
['PLAN.md', 'OTHER.md'], ['PLAN.md', 'archive/PLAN.md'],
["PLAN.md's four bullets plus the approved schema", 'the future expanded plan'],
['plus the approved schema', 'plus a new unapproved schema'],
[', pressure-testing every stated behavior for failure modes, errors, tests, and production visibility', ''],
['errors, tests, and production visibility', 'word choice and formatting'],
['while deferring anything extra rather than adding it silently', 'while adding anything extra'],
['while deferring', 'while not deferring'], ['pressure-testing', 'not pressure-testing'],
]) expect(withText(original.replace(from!, to!)), from).toBe(false);
for (const contextChange of [
(text: string) => text.replace('PLAN.md', 'PLAN.md and OTHER.md'),
(text: string) => text.replace('schema) approved', 'schema) not approved'),
(text: string) => text.replace('schema) approved', 'schema) discussed'),
...['approved if the user agrees', 'approved once migration finishes', 'approved pending migration', 'approved subject to migration'].map(status =>
(text: string) => text.replace('schema) approved', 'schema) ' + status)),
]) {
const t = replay(); const q = t.calls[0]!.questions[0]!; const prior = q.question;
q.question = contextChange(q.question); t.calls[0]!.answers = { [q.question]: t.calls[0]!.answers![prior]! };
expect(matches(t)).toBe(false);
}
});
test('current corrections withdraw a commitment; quoted corrections do not', () => {
for (const correction of [
'Correction: I will expand scope to include defaults.',
'Correction: I will not keep scope fixed to these requirements.',
'Correction: I am no longer keeping scope to those requirements.',
'The formerly excluded additions are in scope.',
]) {
expect(withText(original + '\n\n' + correction), correction).toBe(false);
for (const quote of ['> ' + correction, '```text\n' + correction + '\n```', '~~~text\n' + correction + '\n~~~', 'A quotation: "' + correction + '"']) {
expect(withText(original + '\n\n' + quote), quote).toBe(true);
}
}
});
test('native selection, session and timestamp evidence remain required', () => {
for (const change of [
(t: PlanCountTranscript) => { t.status = 'missing'; },
(t: PlanCountTranscript) => { t.calls[0]!.answered = false; },
(t: PlanCountTranscript) => { t.calls[0]!.failed = true; },
(t: PlanCountTranscript) => { t.calls[0]!.answeredAt = new Date(captured.selectionStartedAt - 1).toISOString(); },
(t: PlanCountTranscript) => { t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'Scope expansion'; },
(t: PlanCountTranscript) => { t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'Unknown'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.sessionId = 'foreign'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = t.calls[0]!.answeredAt!; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = 'invalid'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = new Date(Date.now() + 60_000).toISOString(); },
(t: PlanCountTranscript) => { t.assistantMessages = []; },
]) { const t = replay(); change(t); expect(matches(t)).toBe(false); }
});
test('new posture evidence selects the existing mode owner', () => {
for (const file of ['test/ceo-hold-commitment-ar.test.ts', 'test/fixtures/ceo-hold-commitment-ar.json']) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-ceo-mode-routing']);
}
});
+212
View File
@@ -0,0 +1,212 @@
import { expect, test } from 'bun:test';
import { hasNativePostAnswerCeoPosture, nativeCeoModeAnswer } from './helpers/ceo-mode-option';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-hold-posture-ag.json';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const posture = /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i;
const original = captured.transcript.assistantMessages[0]!.text;
const replay = () => structuredClone(captured.transcript) as PlanCountTranscript;
const matches = (transcript = replay()) => hasNativePostAnswerCeoPosture(
transcript, 'HOLD SCOPE', posture, captured.selectionStartedAt,
);
test('the captured selected HOLD scope lock and hardening establish posture without a keyword', () => {
const transcript = replay();
expect(captured.provenance.actualState).toBe('failed');
expect(nativeCeoModeAnswer(transcript, 'HOLD SCOPE', captured.selectionStartedAt)?.toolUseId)
.toBe('toolu_011bt3yabPDSEsPNm97EhqV4');
expect(posture.test(original)).toBe(false);
expect(matches(transcript)).toBe(true);
});
test('ordinary current scope declarations preserve the same three obligations', () => {
for (const text of [
original.replace("I'm locking", 'I will lock'),
original.replace("I'm locking", "I'll lock"),
original.replace("I'm locking", 'We are keeping').replace('the four PLAN.md bullets from approach B', 'the agreed plan')
.replace('flagging anything beyond', 'treating everything outside').replace('hunting', 'checking'),
original.replace("I'm locking", 'I am holding').replace('four PLAN.md bullets from approach B', 'PLAN.md requirements')
.replace('flagging', 'marking').replace('hunting', 'looking'),
original.replace("I'm", 'Im').replace('PLAN.md', '**PLAN.md**'),
]) {
const transcript = replay(); transcript.assistantMessages[0]!.text = text;
expect(matches(transcript)).toBe(true);
}
});
test('deferred commitments, conditions and quotation cannot establish the current posture', () => {
for (const text of [
original.replace("I'm locking", 'I would lock'),
original.replace("I'm locking", 'I will later lock'),
'If you approve, ' + original,
'Later, ' + original,
'Example only: ' + original,
'An unproven hypothesis: ' + original,
'Example only. ' + original,
'"' + original + '"',
'> ' + original,
'```text\n' + original + '\n```',
'~~~~\n' + original + '\n~~~~',
'Read(file)\n' + original,
'The user said: ' + original,
original.replace('and hunting', 'and not hunting'),
]) {
const transcript = replay(); transcript.assistantMessages[0]!.text = text;
expect(matches(transcript), text).toBe(false);
}
});
test('all three obligations refer to the selected current scope', () => {
for (const text of [
original.replace('PLAN.md', 'OTHER.md'),
original.replace('PLAN.md', 'archive/PLAN.md'),
original.replace('the four PLAN.md bullets from approach B', 'the future expanded plan'),
original.replace('the four PLAN.md bullets from approach B', 'the two imagined requirements'),
original.replace('out of scope', 'in scope'),
original.replace('as out of scope', 'as not out of scope'),
original.replace('flagging anything beyond that (defaults, sharing, deep links) as out of scope, and ', ''),
original.replace(/, and hunting[^.]+\./, '.'),
original.replace('constraints, error handling, UI edge cases, access-rule leaks', 'word choice and formatting'),
original + ' I am expanding scope to include a new feature.',
original + ' I am adding extra features to scope.',
]) {
const transcript = replay(); transcript.assistantMessages[0]!.text = text;
expect(matches(transcript), text).toBe(false);
}
const ambiguous = replay();
const question = ambiguous.calls[0]!.questions[0]!;
const oldQuestion = question.question;
question.question = question.question.replace('reviewing PLAN.md', 'reviewing PLAN.md and OTHER.md');
ambiguous.calls[0]!.answers = { [question.question]: ambiguous.calls[0]!.answers![oldQuestion]! };
expect(matches(ambiguous)).toBe(false);
});
test('explicit later corrections withdraw scope locking, while quoted examples do not', () => {
const corrections = [
'Correction: the previously excluded defaults, sharing, and deep links are now in scope.',
'Correction: I am no longer locking scope to those requirements.',
'I am not keeping scope to those requirements.',
'The formerly excluded additions are in scope.',
];
for (const correction of corrections) {
const transcript = replay();
transcript.assistantMessages[0]!.text = original + '\n\n' + correction;
expect(matches(transcript), correction).toBe(false);
for (const quote of ['> ' + correction, '```text\n' + correction + '\n```',
'~~~text\n' + correction + '\n~~~', 'An example of withdrawn wording is: "' + correction + '"']) {
transcript.assistantMessages[0]!.text = original + '\n\n' + quote;
expect(matches(transcript), quote).toBe(true);
}
}
});
test('only a real selected HOLD answer followed by its own public statement supplies evidence', () => {
for (const change of [
(t: PlanCountTranscript) => { t.status = 'missing'; },
(t: PlanCountTranscript) => { t.calls[0]!.answered = false; },
(t: PlanCountTranscript) => { t.calls[0]!.failed = true; },
(t: PlanCountTranscript) => { t.calls[0]!.answeredAt = new Date(captured.selectionStartedAt - 1).toISOString(); },
(t: PlanCountTranscript) => { t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'Scope Expansion'; },
(t: PlanCountTranscript) => { t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'Unknown'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.sessionId = 'foreign'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = t.calls[0]!.answeredAt!; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = 'invalid'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = new Date(Date.now() + 60_000).toISOString(); },
(t: PlanCountTranscript) => { t.assistantMessages = []; },
]) {
const transcript = replay(); change(transcript); expect(matches(transcript)).toBe(false);
}
const expansion = replay();
expansion.calls[0]!.answers![expansion.calls[0]!.questions[0]!.question] = 'Scope Expansion';
expect(hasNativePostAnswerCeoPosture(expansion, 'SCOPE EXPANSION', posture, captured.selectionStartedAt)).toBe(false);
});
test('new evidence controls select only the existing mode paid owner', () => {
for (const file of ['test/ceo-hold-posture-ag.test.ts', 'test/fixtures/ceo-hold-posture-ag.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, files]) => files.includes(file)).map(([owner]) => owner))
.toEqual(['plan-ceo-mode-routing']);
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-ceo-mode-routing']);
}
});
// Exact public AY parent narration after the answered HOLD SCOPE mode AUQ.
// Its native ownership controls use the existing PLAN.md / approved-approach-B fixture.
const ambiguityNarration = "I'm holding strictly to the plan's approved scope (Approach B, private-only views) and flagging any ambiguities the sketch leaves undecided as targeted questions rather than expanding scope. First up: what happens when a saved view's filters reference something that's been deleted.\n\n";
const ambiguityReplay = () => {
const transcript = replay();
transcript.assistantMessages[0]!.text = ambiguityNarration;
return transcript;
};
const ambiguityMatches = (text = ambiguityNarration) => {
const transcript = ambiguityReplay(); transcript.assistantMessages[0]!.text = text;
return matches(transcript);
};
test('approved scope plus targeted ambiguity questions applies HOLD without naming the mode', () => {
expect(posture.test(ambiguityNarration)).toBe(false);
expect(ambiguityMatches()).toBe(true);
for (const text of [
ambiguityNarration.replace("I'm holding", 'We are keeping'),
ambiguityNarration.replace("I'm holding", 'I will hold'),
ambiguityNarration.replace('the sketch leaves undecided', 'in the plan').replace('flagging', 'surfacing'),
ambiguityNarration.replace("plan's", "PLAN.md's"),
ambiguityNarration.replace("I'm", 'Im').replace("plan's", 'plans'),
]) expect(ambiguityMatches(text), text).toBe(true);
});
test('ambiguity wording must adopt every obligation without quoting, negating or deferring it', () => {
for (const text of [
'> ' + ambiguityNarration, '"' + ambiguityNarration.trim() + '"',
'```text\n' + ambiguityNarration + '```', '~~~text\n' + ambiguityNarration + '~~~',
'Example only: ' + ambiguityNarration, 'The user said: ' + ambiguityNarration,
'Read(file)\n' + ambiguityNarration, 'If approved, ' + ambiguityNarration,
ambiguityNarration.replace("I'm holding", 'I would hold'),
ambiguityNarration.replace("I'm holding", 'I will later hold'),
ambiguityNarration.replace("I'm holding", "I'm not holding"),
ambiguityNarration.replace('and flagging', 'and not flagging'),
ambiguityNarration.replace('approved scope', 'proposed scope'),
ambiguityNarration.replace("plan's", "OTHER.md's"),
ambiguityNarration.replace('private-only views', 'OTHER.md views'),
ambiguityNarration.replace('Approach B', 'Approach C'),
ambiguityNarration.replace('as targeted questions rather than expanding scope', 'as optional improvements'),
ambiguityNarration.replace('rather than expanding scope', 'while expanding scope'),
ambiguityNarration.replace('ambiguities the sketch leaves undecided', 'word choice and formatting'),
]) expect(ambiguityMatches(text), text).toBe(false);
for (const correction of [
'I am expanding scope to include sharing.',
'Correction: I will add defaults to scope.',
'The previously excluded sharing feature is now in scope.',
'Correction: I am no longer holding scope to this plan.',
"Correction: I am not holding strictly to the plan's approved scope.",
'Correction: I am no longer flagging ambiguities as targeted questions.',
'Correction: this posture is withdrawn.',
'This posture is no longer current.',
]) {
expect(ambiguityMatches(ambiguityNarration + correction), correction).toBe(false);
expect(ambiguityMatches(ambiguityNarration + '> ' + correction), correction).toBe(true);
}
});
test('ambiguity posture stays bound to the approved plan and its actual native answer', () => {
for (const change of [
(t: PlanCountTranscript) => { t.calls[0]!.answered = false; },
(t: PlanCountTranscript) => { t.calls[0]!.failed = true; },
(t: PlanCountTranscript) => { t.calls[0]!.answers![t.calls[0]!.questions[0]!.question] = 'Scope Expansion'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.sessionId = 'foreign'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = t.calls[0]!.answeredAt!; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = 'invalid'; },
(t: PlanCountTranscript) => { t.assistantMessages[0]!.timestamp = new Date(Date.now() + 60_000).toISOString(); },
]) { const transcript = ambiguityReplay(); change(transcript); expect(matches(transcript)).toBe(false); }
for (const [from, to] of [
['PLAN.md', 'PLAN.md and OTHER.md'],
['approved.', 'not approved.'],
['approved.', 'approved if accepted.'],
['approved.', 'discussed.'],
]) {
const transcript = ambiguityReplay(); const q = transcript.calls[0]!.questions[0]!;
const before = q.question; q.question = before.replace(from!, to!);
transcript.calls[0]!.answers = { [q.question]: transcript.calls[0]!.answers![before]! };
expect(matches(transcript), to).toBe(false);
}
});
+108
View File
@@ -0,0 +1,108 @@
import { describe, expect, test } from 'bun:test';
import { findCeoModeOption, nativeCeoModeAnswer, nextCeoModeNavigation } from './helpers/ceo-mode-option';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import { selectTests } from './helpers/touchfiles';
import captured from './fixtures/ceo-mode-colon-at.json';
function transcript(): PlanCountTranscript {
return { status: 'ready', calls: [structuredClone(captured)], assistantMessages: [] };
}
describe('CEO colon-prefixed native mode choices', () => {
test('the exact public menu resolves each named mode by display position', () => {
const options = captured.questions[0]!.options.map((option, i) => ({ index: i + 1, label: option.label }));
expect(findCeoModeOption(options, 'SELECTIVE EXPANSION')).toBe(1);
expect(findCeoModeOption(options, 'SCOPE EXPANSION')).toBe(2);
expect(findCeoModeOption(options, 'HOLD SCOPE')).toBe(3);
expect(findCeoModeOption(options, 'SCOPE REDUCTION')).toBe(4);
});
test('navigation selects expansion in either display order without changing native input', () => {
for (const reverse of [false, true]) {
const call = transcript().calls[0]!;
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
const question = call.questions[0]!;
if (reverse) question.options.reverse();
const original = structuredClone(call);
const visible = `${question.header}\n${question.question}\n` + question.options.map((option, i) =>
`${i ? ' ' : ''} ${i + 1}. ${option.label}`).join('\n') +
'\nEnter to select · ↑/↓ to navigate · Esc to cancel';
const action = nextCeoModeNavigation(visible, 'SCOPE EXPANSION', new Set(), call);
expect(action.kind).toBe('mode');
expect(action.kind === 'mode' && action.index).toBe(reverse ? 3 : 2);
expect(call).toEqual(original);
}
});
test('the recorded wrong selection remains selective expansion, never expansion coverage', () => {
const actual = transcript();
expect(nativeCeoModeAnswer(actual, 'SELECTIVE EXPANSION', 0)?.toolUseId)
.toBe('toolu_01XY3qPeSuJZa3H2uCfatJ8b');
expect(nativeCeoModeAnswer(actual, 'SCOPE EXPANSION', 0)).toBeNull();
expect(actual.calls[0]).toEqual(captured);
});
test('pending, failed, stale and ambiguous native answers cannot prove selection', () => {
for (const change of [
(value: PlanCountTranscript) => { value.calls[0]!.answered = false; },
(value: PlanCountTranscript) => { value.calls[0]!.failed = true; },
(value: PlanCountTranscript) => { delete value.calls[0]!.answers; },
(value: PlanCountTranscript) => { value.calls[0]!.answeredAt = 'invalid'; },
(value: PlanCountTranscript) => {
value.calls[0]!.questions[0]!.options.push({ label: 'E: SELECTIVE EXPANSION' });
},
]) {
const value = transcript();
change(value);
expect(nativeCeoModeAnswer(value, 'SELECTIVE EXPANSION', 0)).toBeNull();
}
expect(nativeCeoModeAnswer(transcript(), 'SELECTIVE EXPANSION', Date.parse(captured.answeredAt) + 1)).toBeNull();
const laterAmbiguous = transcript();
const later = structuredClone(laterAmbiguous.calls[0]!);
later.toolUseId = 'later-ambiguous-mode';
later.answeredAt = new Date(Date.parse(captured.answeredAt) + 1000).toISOString();
later.questions[0]!.options.push({ label: 'E: SELECTIVE EXPANSION' });
laterAmbiguous.calls.push(later);
expect(nativeCeoModeAnswer(laterAmbiguous, 'SELECTIVE EXPANSION', 0)).toBeNull();
});
test('action titles, lookalikes and preview descriptions do not become modes', () => {
for (const label of [
'A: Use HOLD SCOPE for the next review',
'B: Explain SCOPE EXPANSION',
'AA: HOLD SCOPE',
'1: HOLD SCOPE',
'A:: HOLD SCOPE',
'A: HOLD SCOPES',
'A: Fix contrast │ HOLD SCOPE',
'A: Fix contrast ┌ SCOPE EXPANSION',
'A: "HOLD SCOPE"',
'Prior: HOLD SCOPE',
]) expect(findCeoModeOption([{ index: 1, label }], 'HOLD SCOPE')).toBeNull();
expect(() => findCeoModeOption([
{ index: 1, label: 'A: SELECTIVE EXPANSION │ SCOPE EXPANSION' },
{ index: 2, label: 'B: HOLD SCOPE' },
], 'SCOPE EXPANSION')).toThrow('not in option labels');
});
test('duplicate and missing mode titles fail before selection; legacy prefixes still work', () => {
expect(() => findCeoModeOption([
{ index: 1, label: 'A: HOLD SCOPE' },
{ index: 2, label: 'HOLD SCOPE (recommended)' },
], 'HOLD SCOPE')).toThrow('duplicate');
expect(() => findCeoModeOption([{ index: 1, label: 'A: SCOPE REDUCTION' }], 'HOLD SCOPE'))
.toThrow('not in option labels');
for (const label of ['A) HOLD SCOPE', 'A. HOLD SCOPE', 'a: hold scope', 'A: HOLD SCOPE']) {
expect(findCeoModeOption([{ index: 3, label }], 'HOLD SCOPE')).toBe(3);
}
});
test('the exact public fixture and regression select the mode-routing workflow', () => {
for (const file of ['test/ceo-mode-colon-at.test.ts', 'test/fixtures/ceo-mode-colon-at.json']) {
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-ceo-mode-routing']);
}
});
});
+119
View File
@@ -0,0 +1,119 @@
import {describe,expect,test} from 'bun:test';
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
import {hasNativePostAnswerCeoPosture,nextCeoModeNavigation} from './helpers/ceo-mode-option';
import {capturePlanCountQuestion,nativePlanCallFingerprint,planCountPrerequisitePick,planCountQuestionInput} from './helpers/claude-pty-runner';
import {readPlanCountTranscript,type NativePublicToolEvent,type NativePlanQuestionCall} from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-mode-full-ad.json';
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
const pattern=/\b(expansion|10x|delight|dream|cathedral|opt[\s-]?in)\b/i;
function replay(i:number){
const item=captured.cases[i]!,root=fs.mkdtempSync(path.join(os.tmpdir(),'ceo-full-ad-'));
fs.mkdirSync(path.join(root,'projects','owned'),{recursive:true});
fs.writeFileSync(path.join(root,'projects','owned',item.process.sessionId+'.jsonl'),item.records.map(r=>JSON.stringify(r)).join('\n')+'\n');
const events:NativePublicToolEvent[]=[];
try{return {item,transcript:readPlanCountTranscript(root,item.process.cwd,e=>events.push(e)),events};}
finally{fs.rmSync(root,{recursive:true,force:true});}
}
function pending(){const c=structuredClone(replay(0).transcript.calls[0]!);c.answered=false;delete c.answers;delete c.answeredAt;delete c.unansweredQuestionIndices;return c;}
// Full panes projected from exact native questions, not retained historical viewports.
function pane(call:NativePlanQuestionCall,index:number){const q=call.questions[index]!;return [
call.questions.length>1?'← '+call.questions.map((v,i)=>`${i<index?'☒':'☐'} ${v.header}`).join(' ')+' ✔ Submit →':'☐ '+q.header,
q.question,...q.options.map((v,i)=>`${i?' ':''} ${i+1}. ${v.label}`),
`Enter to select · ${call.questions.length>1?'Tab/Arrow keys':'↑/↓'} to navigate · Esc to cancel`].join('\n');}
function frame(c:NativePlanQuestionCall,index:number){const visible=pane(c,index);return {visible,active:capturePlanCountQuestion(visible,new Set(),0,true,c)!,routing:nativePlanCallFingerprint(c,0,true)};}
function match(e= replay(1)){return hasNativePostAnswerCeoPosture(e.transcript,'SCOPE EXPANSION',pattern,e.item.selectedAt!,e.events);}
function rebind(e:ReturnType<typeof replay>){const d=e.transcript.calls[1]!,q=d.questions[0]!;e.events[2]!.input={questions:d.questions};d.answers={[q.question]:q.options[0]!.label};}
describe('full AD mode failures retain their actual outcomes',()=>{
test('Proposal 1 is a completed scope decision after the actual selected mode',()=>{
const e=replay(1);expect(e.item.actualState).toBe('failed');expect(e.transcript.calls).toHaveLength(2);expect(e.events).toHaveLength(4);
expect(e.transcript.calls[1]!.answeredAt).toBe('2026-09-09T18:26:20.110Z');expect(match(e)).toBe(true);
});
test.each(['pending','foreign','wrong mode','pre-mode','missing reply','wrong answer','extra question','extra option','multiselect',
'quoted','fenced','mode echo','mode mismatch','mode menu','appended instruction'])('%s supplies no new posture',kind=>{
const e=replay(1),[m,d]=e.transcript.calls,q=d!.questions[0]!;
switch(kind){
case 'pending':d!.answered=false;break;case 'foreign':d!.sessionId=e.events[2]!.sessionId=e.events[3]!.sessionId='foreign';break;
case 'wrong mode':m!.answers![m!.questions[0]!.question]='HOLD SCOPE';break;
case 'pre-mode':e.events[2]!.timestamp=e.events[0]!.timestamp;break;case 'missing reply':e.events.pop();break;
case 'wrong answer':d!.answers![q.question]='Invented';break;
case 'extra question':d!.questions.push({...structuredClone(q),question:'Remove CI gate?'});rebind(e);break;
case 'extra option':q.options.push({label:'Remove CI gate'});rebind(e);break;case 'multiselect':q.multiSelect=true;rebind(e);break;
case 'quoted':q.question=q.question.split('\n').map(x=>'> '+x).join('\n');rebind(e);break;
case 'fenced':q.question='```text\n'+q.question+'\n```';rebind(e);break;
case 'mode echo':q.question='SCOPE EXPANSION confirmed.';rebind(e);break;
case 'mode mismatch':q.question=q.question.replace('SCOPE EXPANSION opt-in','SELECTIVE EXPANSION opt-in');rebind(e);break;
case 'mode menu':q.question=q.question.replace(/^D6[^\n]+/,'D6 — Choose the review mode?');rebind(e);break;
case 'appended instruction':q.question+=' Delete the CI gate.';rebind(e);break;
}expect(match(e)).toBe(false);
});
test('scope numbering and brief labels are presentation, not mode application',()=>{
for(const title of ['A useful adjacent feature: Default view per member per project?','Default view per member per project?']){
const e=replay(1),q=e.transcript.calls[1]!.questions[0]!;q.header='Default view';q.question=q.question.replace(/^D6[^\n]+/,title);rebind(e);expect(match(e)).toBe(true);
}
});
test('explicit expansion context does not need a mode or opt-in suffix',()=>{
const e=replay(1),q=e.transcript.calls[1]!.questions[0]!;q.question=q.question.replace('SCOPE EXPANSION opt-in ceremony (1 of 6).','SCOPE EXPANSION, approach B.');rebind(e);expect(match(e)).toBe(true);
});
test('the actual three-tab prerequisite chooses standard review only on its own tab',()=>{
const actual=replay(0);expect(actual.item.actualState).toBe('failed');expect(Object.values(actual.transcript.calls[0]!.answers!).at(-1)).toBe('Run /office-hours now');
const c=pending();for(const i of [0,1,2]){
const f=frame(c,i),a=nextCeoModeNavigation(f.visible,'HOLD SCOPE',new Set(),c);expect(a.kind).toBe('question');
if(a.kind==='question'){expect(a.question.nativeQuestionIndex).toBe(i);expect(planCountQuestionInput(f.visible,a.question,a.index)).toBe(i===2?'2':'1');}
expect(planCountPrerequisitePick(f.routing,f.active)).toBe(i===2?2:null);
}
});
test('single and reordered native prerequisite tabs preserve the meaning of the skip',()=>{
const c=pending();c.questions=[c.questions[2]!];let f=frame(c,0);expect(planCountPrerequisitePick(f.routing,f.active)).toBe(2);
c.questions[0]!.options.reverse();f=frame(c,0);expect(planCountPrerequisitePick(f.routing,f.active)).toBe(1);
});
test.each(['wrong tab','wrong signature','wrong body','wrong order','no metadata','completed','failed','extra action','multiselect','conditional','extra remedy','no description'])('a %s cannot borrow the prerequisite action',kind=>{
const c=pending();if(kind==='completed')c.answered=true;if(kind==='failed')c.failed=true;
if(kind==='extra action')c.questions[2]!.options.push({label:'Accept risk'});
if(kind==='multiselect')c.questions[2]!.multiSelect=true;
if(kind==='conditional')c.questions[2]!.options[1]!.description+=' if all tests pass.';
if(kind==='extra remedy')c.questions[2]!.options[1]!.description+=' Remove the CI gate.';
if(kind==='no description')c.questions[2]!.options[1]!.description='';
const f=frame(c,2);let a=f.active;
if(kind==='wrong tab')a={...a,nativeQuestionIndex:0};if(kind==='wrong signature')a={...a,signature:'foreign:tool:question:2'};
if(kind==='wrong body')a={...a,promptSnippet:'Choose a product direction.'};if(kind==='wrong order')a={...a,options:[...a.options].reverse()};
if(kind==='no metadata')a={...a,nativeCall:undefined};
expect(planCountPrerequisitePick(f.routing,a)).toBeNull();
});
});
describe('full AD HOLD retry completed sequencing rationale',()=>{
function hold(){const e=replay(2);return {e,decision:e.transcript.calls[2]!,q:e.transcript.calls[2]!.questions[0]!};}
function matches(e:ReturnType<typeof replay>){return hasNativePostAnswerCeoPosture(e.transcript,'HOLD SCOPE',/\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i,e.item.selectedAt!,e.events);}
function bind(e:ReturnType<typeof replay>){const d=e.transcript.calls[2]!,q=d.questions[0]!;e.events[4]!.input={questions:d.questions};d.answers={[q.question]:q.options[0]!.label};}
test('the actual completed rationale applies HOLD to work in the previously approved approach',()=>{
const {e,decision,q}=hold();expect(e.item.actualState).toBe('failed');expect(e.transcript.calls).toHaveLength(3);
const approach=e.transcript.calls[0]!;expect(Object.values(approach.answers!)).toEqual(['B: ViewState schema (recommended)']);
expect(approach.questions[0]!.options[0]!.description).toContain('URL params');
expect(decision.answeredAt).toBe('2026-09-09T18:35:05.273Z');expect(q.question).toContain('not new scope either way');expect(matches(e)).toBe(true);
});
test('three and four alternatives still express one completed review decision',()=>{
for(const count of [3,4]){const {e,q}=hold();q.options.push({label:'Gate URL sync for the pilot'});if(count===4)q.options.push({label:'Run a limited URL sync pilot'});bind(e);expect(matches(e)).toBe(true);}
});
test.each(['pending','foreign','before mode','missing reply','failed reply','wrong answer','metadata only','bare echo','other mode',
'quoted rationale','fenced rationale','duplicate options','extra question','extra instruction','multiselect'])('%s is not completed HOLD rationale',kind=>{
const {e,decision,q}=hold();
switch(kind){
case 'pending':decision.answered=false;break;case 'foreign':decision.sessionId=e.events[4]!.sessionId=e.events[5]!.sessionId='foreign';break;
case 'before mode':e.events[4]!.timestamp=e.events[0]!.timestamp;break;case 'missing reply':e.events.pop();break;case 'failed reply':e.events[5]!.isError=true;break;
case 'wrong answer':decision.answers![q.question]='Invented';break;
case 'metadata only':q.question=q.question.replace(/ELI10:[\s\S]*?\nStakes/,'ELI10: We will implement the URL codec.\nStakes');bind(e);break;
case 'bare echo':q.question=q.question.replace(/ELI10:[\s\S]*?\nStakes/,'ELI10: HOLD SCOPE confirmed.\nStakes');bind(e);break;
case 'other mode':q.question=q.question.replace(/HOLD SCOPE/g,'SCOPE EXPANSION');bind(e);break;
case 'quoted rationale':q.question=q.question.replace('ELI10: Approach','ELI10:\n> Approach');bind(e);break;
case 'fenced rationale':q.question=q.question.replace('ELI10: Approach','ELI10: ```Approach');bind(e);break;
case 'duplicate options':q.options[1]!.label=q.options[0]!.label;bind(e);break;
case 'extra question':decision.questions.push({...structuredClone(q),question:'Remove CI?'});bind(e);break;
case 'extra instruction':q.question+=' Disable authentication.';bind(e);break;
case 'multiselect':q.multiSelect=true;bind(e);break;
}expect(matches(e)).toBe(false);
});
});
test('the exact full AD regressions select their periodic caller',()=>{
for(const file of ['test/ceo-mode-full-ad.test.ts','test/fixtures/ceo-mode-full-ad.json']) expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['plan-ceo-mode-routing']);
});
+84
View File
@@ -0,0 +1,84 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { findCeoModeOption, nativeCeoModeAnswer, nextCeoModeNavigation } from './helpers/ceo-mode-option';
import { readPlanCountTranscript } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-mode-labels-l.json';
function transcript() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-mode-'));
const first = captured.records[0]!;
const project = path.join(root, 'projects', 'fixture');
fs.mkdirSync(project, { recursive: true });
fs.writeFileSync(path.join(project, `${first.sessionId}.jsonl`),
captured.records.map(record => JSON.stringify(record)).join('\n') + '\n');
try {
return readPlanCountTranscript(root, first.cwd);
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
}
describe('Native CEO letter-prefixed mode choices', () => {
test('the captured menu selects the requested mode in either order without changing labels', () => {
for (const reverse of [false, true]) {
const call = transcript().calls[0]!;
call.answered = false;
delete call.answers;
delete call.unansweredQuestionIndices;
const q = call.questions[0]!;
if (reverse) q.options.reverse();
const original = structuredClone(q.options);
const visible = `${q.header}\n${q.question}\n` + q.options.map((option, i) =>
`${i ? ' ' : ''} ${i + 1}. ${option.label}`).join('\n') +
'\nEnter to select · ↑/↓ to navigate · Esc to cancel';
const action = nextCeoModeNavigation(visible, 'HOLD SCOPE', new Set(), call);
expect(action.kind).toBe('mode');
expect(action.kind === 'mode' && action.index).toBe(reverse ? 2 : 3);
const numbered = q.options.map((option, i) => ({ index: i + 1, label: option.label }));
expect(findCeoModeOption(numbered, 'SCOPE EXPANSION')).toBe(reverse ? 3 : 2);
expect(q.options).toEqual(original);
}
});
test('the historical wrong selection remains SELECTIVE EXPANSION, never evidence of HOLD', () => {
const actual = transcript();
expect(actual.status).toBe('ready');
expect(nativeCeoModeAnswer(actual, 'SELECTIVE EXPANSION', 0)?.toolUseId)
.toBe('toolu_01JCKZEDVZ5DXqRazc6Y7L2Z');
expect(nativeCeoModeAnswer(actual, 'HOLD SCOPE', 0)).toBeNull();
});
test('ordinary action text, lookalikes and preview text do not become mode titles', () => {
for (const label of [
'A) Use HOLD SCOPE for the next review',
'B) Explain SCOPE EXPANSION',
'AA) HOLD SCOPE',
'1) HOLD SCOPE',
'A) HOLD SCOPES',
'A) Fix contrast │ HOLD SCOPE',
'A) Fix contrast ┌ SCOPE EXPANSION',
]) expect(findCeoModeOption([{ index: 1, label }], 'HOLD SCOPE')).toBeNull();
});
test('duplicate or absent target modes fail before an input can be selected', () => {
const duplicate = [
{ index: 1, label: 'A) HOLD SCOPE' },
{ index: 2, label: 'B) HOLD SCOPE (recommended)' },
{ index: 3, label: 'C) SCOPE EXPANSION' },
];
expect(() => findCeoModeOption(duplicate, 'HOLD SCOPE')).toThrow('duplicate');
expect(() => findCeoModeOption([{ index: 1, label: 'A) SCOPE REDUCTION' }], 'HOLD SCOPE'))
.toThrow('not in option labels');
const ambiguous = transcript();
ambiguous.calls[0]!.questions[0]!.options.push({ label: 'E) SELECTIVE EXPANSION' });
expect(nativeCeoModeAnswer(ambiguous, 'SELECTIVE EXPANSION', 0)).toBeNull();
const previous = transcript();
const later = structuredClone(ambiguous.calls[0]!);
later.toolUseId = 'later-ambiguous-mode';
later.answeredAt = new Date(Date.parse(later.answeredAt!) + 1000).toISOString();
previous.calls.push(later);
expect(nativeCeoModeAnswer(previous, 'SELECTIVE EXPANSION', 0)).toBeNull();
});
});
+484
View File
@@ -0,0 +1,484 @@
import { describe, expect, test } from 'bun:test';
import { findCeoModeOption, hasPostAnswerCeoPosture, hasNativePostAnswerCeoPosture, nativeCeoModeAnswer, nextCeoModeNavigation, nextCeoPostureContinuation } from './helpers/ceo-mode-option';
import { parseNumberedOptions, stripAnsi, planCountQuestionInput, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
describe('CEO mode option matching', () => {
test('selects option 4 from the failed Claude Code 2.1.257 menu capture', () => {
// Labels and side-pane residue from the 2026-09-08 paid failure. The
// option existed; literal includes("SCOPE EXPANSION") could not see it.
// The failure log preserves parsed labels, not the original raw frame.
const options = [
{ index: 1, label: 'SELECTIVEEXPANSION┌────────────────────────────────────────────────────────────────────────────────────┐\r (ecommnded) │SELECTIVEEXPANSION│' },
{ index: 2, label: 'HOLD SCOPE │ Hld scope: eview rigorusly fr failure modes, edg ass, observability.│' },
{ index: 3, label: 'SCOPE REDUCTION │ Then surface: cherry-pikableadditions you ca Accept/Defer/Skip.│' },
{ index: 4, label: 'SCOPEEXPANSION│Neutralposture:presentopportunities,stateeffort,youdecide.│\r │ Good for: substantialfeaturewithsolidfoundation,shippedbeforescopelock.│\r└────────────────────────────────────────────────────────────────────────────────────┘' },
];
expect(findCeoModeOption(options, 'SCOPE EXPANSION')).toBe(4);
expect(findCeoModeOption(options, 'HOLD SCOPE')).toBe(2);
expect(findCeoModeOption(options, 'SELECTIVE EXPANSION')).toBe(1);
});
test('recognizes the mode question when every label loses its inter-word spaces', () => {
const frame = stripAnsi([
' 1. HOLD\x1b[1CSCOPE (recommended)',
' 2. SELECTIVE\x1b[1CEXPANSION',
' 3. SCOPE\x1b[1CEXPANSION',
' 4. SCOPE\x1b[1CREDUCTION',
].join('\n'));
const options = parseNumberedOptions(frame);
expect(findCeoModeOption(options, 'SCOPE EXPANSION')).toBe(3);
expect(findCeoModeOption(options, 'SCOPE REDUCTION')).toBe(4);
});
test('retains spaced, mixed-case labels and recommendation suffixes', () => {
expect(findCeoModeOption([
{ index: 1, label: 'Scope Expansion (recommended)' },
{ index: 2, label: 'HOLD SCOPE' },
], 'SCOPE EXPANSION')).toBe(1);
});
test('still fails on the earlier three-option capture with no expansion target', () => {
const options = [
{ index: 1, label: 'HOLD SCOPE (recommended) ┌─────────────────────────────────────────────────────────────┐' },
{ index: 2, label: 'SELECTIVE EXPANSION │HOLD SCOPE │' },
{ index: 3, label: 'SCOPE REDUCTION │ Codeiswritten.Makeitbulletproof.│' },
];
expect(() => findCeoModeOption(options, 'SCOPE EXPANSION'))
.toThrow('target "SCOPE EXPANSION" not in option labels');
});
test('does not select another mode because the side pane mentions the target', () => {
expect(() => findCeoModeOption([
{ index: 1, label: 'HOLD SCOPE │ SCOPE EXPANSION is another option' },
{ index: 2, label: 'SELECTIVEEXPANSION┌ SCOPEEXPANSION' },
], 'SCOPE EXPANSION')).toThrow('target "SCOPE EXPANSION" not in option labels');
});
test('leaves unrelated navigation questions to the existing driver', () => {
expect(findCeoModeOption([
{ index: 1, label: 'Review HOLD SCOPE examples' },
{ index: 2, label: 'Choose a plan │ SCOPE EXPANSION' },
], 'HOLD SCOPE')).toBeNull();
});
test('the shared parser selects both callers while mode-specific regressions stay scoped', () => {
expect(selectTests(['test/helpers/ceo-mode-option.ts'], E2E_TOUCHFILES).selected)
.toEqual(['plan-ceo-mode-routing', 'plan-ceo-finding-count']);
for (const file of ['test/ceo-mode-option.test.ts', 'test/pty-option-selection.test.ts']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['plan-ceo-mode-routing']);
}
});
});
describe('CEO mode navigation replay', () => {
test('advances different setup questions with the same choices and ignores redraws', () => {
const seen = new Set<string>();
const first = '☐Routing\rEnable skill routing?\r1.Enable\r2.Skip';
const next = '☐Learnings\rEnable cross-project learnings?\r1.Enable\r2.Skip';
expect(nextCeoModeNavigation(first, 'HOLD SCOPE', seen).kind).toBe('question');
expect(nextCeoModeNavigation(first, 'HOLD SCOPE', seen).kind).toBe('wait');
expect(nextCeoModeNavigation(next, 'HOLD SCOPE', seen).kind).toBe('question');
expect(seen.size).toBe(2);
});
test('navigates the captured unanswered setup tab before submitting, without counting Submit', () => {
const seen = new Set<string>();
const partial = [
'← ☒ Skill routing ☐ Learnings scope ✔ Submit →',
'Review your answers',
'⚠You have not answered all questions',
'1.Submit aswers',
'2Cancel',
].join('\r\r');
expect(nextCeoModeNavigation(partial, 'HOLD SCOPE', seen)).toEqual({ kind: 'submission', input: '\x1b[Z' });
expect(seen.size).toBe(0);
const question = '← ☒ Skill routing ☐ Learnings scope ✔ Submit →\rEnable cross-project learnings?\r1.Enable\r2.Skip';
expect(nextCeoModeNavigation(`${partial}\r${question}`, 'HOLD SCOPE', seen).kind).toBe('question');
const answered = partial.replace('☐ Learnings scope', '☒ Learnings scope').replace('⚠You have not answered all questions', '');
expect(nextCeoModeNavigation(answered, 'HOLD SCOPE', seen)).toEqual({ kind: 'submission', input: '\r' });
expect(seen.size).toBe(1);
});
test('handles native permission controls before question parsing and dedup', () => {
const seen = new Set<string>();
const permission = 'DoyouwanttooverwriteCLAUDE.md?\r1.Yes\r2.No\rEsctocancel·Tabtoamend';
expect(nextCeoModeNavigation(permission, 'HOLD SCOPE', seen)).toEqual({ kind: 'permission', input: '1\r' });
expect(seen.size).toBe(0);
const actual = '☐ Approaches\rWhich storage strategy?\r1.Server\r2.Local';
expect(nextCeoModeNavigation(`${permission}\r${actual}`, 'HOLD SCOPE', seen).kind).toBe('question');
});
test('file permission lifecycle is shared by navigation and posture without becoming an AUQ', () => {
const seen = new Set<string>();
const permission = 'Do you want to overwrite CLAUDE.md?\n1.Yes\n2.No\nEsc to cancel · Tab to amend';
const transcript: PlanCountTranscript = { status: 'missing', calls: [], assistantMessages: [] };
expect(nextCeoModeNavigation(permission, 'HOLD SCOPE', seen).kind).toBe('permission');
expect(nextCeoModeNavigation(permission, 'HOLD SCOPE', seen).kind).toBe('wait');
expect(nextCeoPostureContinuation(permission, transcript, 'HOLD SCOPE', 0, seen, true)).toBeNull();
const completed = permission + '\n⎿ Added1line\n';
expect(nextCeoPostureContinuation(completed, transcript, 'HOLD SCOPE', 0, seen, true)).toBeNull();
expect(nextCeoModeNavigation(completed, 'HOLD SCOPE', seen).kind).toBe('wait');
const again = completed + permission;
expect(nextCeoPostureContinuation(again, transcript, 'HOLD SCOPE', 0, seen, true)).toBe('permission');
expect(nextCeoModeNavigation(again, 'HOLD SCOPE', seen).kind).toBe('wait');
expect(seen.size).toBe(0);
const question = '☐ Approaches\nWhich storage strategy?\n1.Server\n2.Local';
expect(nextCeoModeNavigation(again + '\n' + question, 'HOLD SCOPE', seen).kind).toBe('question');
expect(seen.size).toBe(1);
// Different sessions retain independent permission state.
expect(nextCeoModeNavigation(permission, 'HOLD SCOPE', new Set()).kind).toBe('permission');
});
test('selects the intended mode from the observed menu, preserving its index', () => {
const frame = '☐ReviewMode\rWhat review posture should I use?\r1.SELECTIVEEXPANSION(recommended)\r2.HOLDSCOPE\r3.SCOPEEXPANSION\r4.SCOPEREDUCTION';
for (const [mode, index] of [['HOLD SCOPE', 2], ['SCOPE EXPANSION', 3]] as const) {
const action = nextCeoModeNavigation(frame, mode, new Set());
expect(action.kind).toBe('mode');
if (action.kind === 'mode') expect(action.index).toBe(index);
}
});
});
describe('CEO posture evidence after mode selection', () => {
const posture = /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i;
const menu = [
'☐ Review mode',
'1.SELECTIVEEXPANSION(recommended)',
'2.HOLD SCOPE │ Code is written. Make it bulletproof.',
'3.SCOPE EXPANSION',
'4.SCOPE REDUCTION',
'Enter to select · ↑/↓ to navigate',
].join('\r');
test('menu redraw and native selected-option echo cannot satisfy the posture gate', () => {
expect(posture.test(menu)).toBe(true); // The old unscoped check passed here.
expect(hasPostAnswerCeoPosture(menu, posture)).toBe(false);
const answer = "⏺ User answered Claude's questions:\r⎿ · Review mode? → HOLD SCOPE";
expect(hasPostAnswerCeoPosture(`${menu}\r${answer}`, posture)).toBe(false);
expect(hasPostAnswerCeoPosture(' HOLD SCOPE\r● HOLD SCOPE', posture)).toBe(false);
expect(hasPostAnswerCeoPosture('● Selected option: HOLD SCOPE', posture)).toBe(false);
expect(hasPostAnswerCeoPosture('● HOLD SCOPE\r✶ Honking… (5s · ↓ 300 tokens)', posture)).toBe(false);
});
test('assistant output must itself contain the existing posture evidence', () => {
expect(hasPostAnswerCeoPosture(`${menu}\r● I will inspect the plan now.`, posture)).toBe(false);
expect(hasPostAnswerCeoPosture('⏺ Read(plan-ceo-review/SKILL.md)\r Review with maximum rigor.', posture)).toBe(false);
expect(hasPostAnswerCeoPosture('● high · /effort\rHOLD SCOPE', posture)).toBe(false);
expect(hasPostAnswerCeoPosture(`● I will inspect the plan now.\r${menu}`, posture)).toBe(false);
});
test('accepts new assistant posture after the answered-question echo, including wrapped prose', () => {
const answer = "⏺UseransweredClaude'squestions:\r⎿Reviewmode?→HOLDSCOPE";
expect(hasPostAnswerCeoPosture(`${answer}\r● HOLD SCOPE. I will review the existing scope for failure modes.`, posture)).toBe(true);
expect(hasPostAnswerCeoPosture(`${answer}\r⏺\rI will apply maximum rigor\rto the agreed scope.`, posture)).toBe(true);
const expansion = /\b(expansion|10x|delight|dream|cathedral|opt[\s-]?in)\b/i;
expect(hasPostAnswerCeoPosture(`${answer}\r● I will explore expansion opportunities that improve the saved-view workflow.`, expansion)).toBe(true);
});
});
describe('native CEO mode posture evidence', () => {
const selectedAt = Date.parse('2026-09-08T15:43:28.000Z');
const posture = /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i;
function transcript(text: string, answer = 'HOLD SCOPE'): PlanCountTranscript {
// Actual question/options/answer shape from targeted-a's false-negative
// HOLD SCOPE run. The native answer selected option3 correctly.
const question = 'Which review mode should I use for this plan? <gstack-qid:plan-ceo-review-mode-selection>';
return { status: 'ready', calls: [{
sessionId: 'mode-session', toolUseId: 'mode-call', answered: true,
answeredAt: '2026-09-08T15:43:30.405Z', answers: { [question]: answer },
questions: [{ header: 'Review mode', question, options: [
{ label: 'SELECTIVE EXPANSION (Recommended)' }, { label: 'SCOPE EXPANSION' },
{ label: 'HOLD SCOPE' }, { label: 'SCOPE REDUCTION' },
] }],
}], assistantMessages: [{ sessionId: 'mode-session', timestamp: '2026-09-08T15:44:01.150Z', text }] };
}
test('recognizes the retained native answer followed by actual HOLD SCOPE analysis', () => {
const captured = 'HOLD SCOPE mode confirmed. Running Step 0D analysis, then reading the review sections file.\n\n**0D — HOLD SCOPE Analysis**\n\n**Complexity check:**\nThe plan introduces: 1 DB migration, 1 SavedView model, 1 CRUD API module (~4 endpoints), 1 view picker UI component, and integration into the existing filter UI.';
expect(hasNativePostAnswerCeoPosture(transcript(captured), 'HOLD SCOPE', posture, selectedAt)).toBe(true);
});
test('wrong, missing, failed, or earlier mode answers cannot establish target routing', () => {
const text = 'I will apply maximum rigor to the existing plan.';
expect(hasNativePostAnswerCeoPosture(transcript(text, 'SCOPE EXPANSION'), 'HOLD SCOPE', posture, selectedAt)).toBe(false);
for (const change of [{ answered: false }, { failed: true }, { answers: {} }, { answeredAt: undefined }]) {
const t = transcript(text); Object.assign(t.calls[0]!, change);
expect(hasNativePostAnswerCeoPosture(t, 'HOLD SCOPE', posture, selectedAt)).toBe(false);
}
expect(hasNativePostAnswerCeoPosture(transcript(text), 'HOLD SCOPE', posture, selectedAt + 10000)).toBe(false);
});
test('prior or foreign assistant prose, a menu, source quotation, and bare confirmation remain insufficient', () => {
for (const text of [
'', 'HOLD SCOPE', '**HOLD SCOPE mode confirmed.**',
'Which mode?\n1. SELECTIVE EXPANSION\n2. HOLD SCOPE\n3. SCOPE EXPANSION',
'```markdown\nReview with maximum rigor.\n```',
'> Review with maximum rigor.',
'Read(SKILL.md)\nReview with maximum rigor.',
]) expect(hasNativePostAnswerCeoPosture(transcript(text), 'HOLD SCOPE', posture, selectedAt)).toBe(false);
for (const change of [{ timestamp: '2026-09-08T15:43:29.000Z' }, { sessionId: 'other-session' }]) {
const t = transcript('I will apply maximum rigor.'); Object.assign(t.assistantMessages[0]!, change);
expect(hasNativePostAnswerCeoPosture(t, 'HOLD SCOPE', posture, selectedAt)).toBe(false);
}
expect(hasNativePostAnswerCeoPosture({ status: 'missing', calls: [], assistantMessages: [] }, 'HOLD SCOPE', posture, selectedAt)).toBe(false);
});
test('continuation requires the confirmed target and permits at most one fresh downstream question', () => {
const t = transcript('');
const fresh = '☐ Architecture\nD4 — Guard the member-scoped lookup?\n1.Add the guard\n2.Defer';
const mode = '☐ Review mode\nWhich review mode?\n1.HOLD SCOPE\n2.SCOPE EXPANSION';
const seen = new Set<string>();
expect(nextCeoPostureContinuation(mode, t, 'HOLD SCOPE', selectedAt, seen, false)).toBeNull();
expect(nextCeoPostureContinuation(fresh, transcript('', 'SCOPE EXPANSION'), 'HOLD SCOPE', selectedAt, seen, false)).toBeNull();
expect(nextCeoPostureContinuation(fresh, t, 'HOLD SCOPE', selectedAt, seen, false)).toBe('question');
expect(nextCeoPostureContinuation(fresh, t, 'HOLD SCOPE', selectedAt, seen, false)).toBeNull();
expect(nextCeoPostureContinuation(fresh.replace('member-scoped', 'project-scoped'), t, 'HOLD SCOPE', selectedAt, seen, true)).toBeNull();
expect(nativeCeoModeAnswer(t, 'HOLD SCOPE', selectedAt)?.toolUseId).toBe('mode-call');
const permission = 'DoyouwanttooverwriteCLAUDE.md?\n1.Yes\n2.No\nEsctocancel·Tabtoamend';
expect(nextCeoPostureContinuation(permission, t, 'HOLD SCOPE', selectedAt, seen, false)).toBe('permission');
expect(nextCeoPostureContinuation(permission, t, 'HOLD SCOPE', selectedAt, seen, false)).toBeNull();
});
test.skipIf(process.platform === 'win32')('one downstream answer releases delayed native prose without passing on the streamed menu', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-posture-flush-'));
const fake = path.join(dir, 'fake-claude');
const worker = path.join(dir, 'worker.ts');
const recordFile = path.join(dir, 'events.jsonl');
const resultFile = path.join(dir, 'result.json');
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
import * as fs from 'node:fs';
import * as path from 'node:path';
const record = event => fs.appendFileSync(process.env.POSTURE_RECORD, JSON.stringify(event) + '\n');
record({type:'startup', pid:process.pid});
const sessionId = 'fake-mode-session';
const dir = path.join(process.env.CLAUDE_CONFIG_DIR, 'projects', 'fixture');
fs.mkdirSync(dir, {recursive:true});
const write = (role, content, extra = {}) => fs.appendFileSync(path.join(dir, sessionId + '.jsonl'), JSON.stringify({
sessionId, cwd:process.cwd(), isSidechain:false, timestamp:new Date().toISOString(), message:{role,content}, ...extra,
}) + '\n');
const question = 'Which mode?';
write('assistant', [{type:'tool_use', id:'mode', name:'AskUserQuestion', input:{questions:[{header:'Mode', question,
options:[{label:'HOLD SCOPE'},{label:'SCOPE EXPANSION'}]}]}}]);
write('user', [{type:'tool_result', tool_use_id:'mode', content:'Answered.'}], {toolUseResult:{answers:{[question]:'SCOPE EXPANSION'}}});
process.stdin.setRawMode?.(true);
let answered = false;
process.stdin.on('data', data => {
record({type:'input', data:data.toString()});
if (data.toString().includes('\r') && !answered) {
answered = true;
write('assistant', [{type:'text', text:'I will explore expansion opportunities that improve saved project views.'}]);
process.stdout.write('\nPOSTURE_FLUSHED\n');
}
});
process.stdout.write('POSTURE_READY\n● I will explore expansion opportunities.\n☐ Expansion 1\nD4 — Add shared project views?\n1.Add to scope\n2.Defer\n');
process.on('SIGINT', () => process.exit(0));
process.stdin.resume();
`);
fs.chmodSync(fake, 0o755);
const moduleUrl = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
fs.writeFileSync(worker, `
import { launchClaudePty, selectPtyNumberedOption } from ${JSON.stringify(moduleUrl('claude-pty-runner.ts'))};
import { readPlanCountTranscript } from ${JSON.stringify(moduleUrl('plan-count-transcript.ts'))};
import { hasNativePostAnswerCeoPosture, nextCeoPostureContinuation } from ${JSON.stringify(moduleUrl('ceo-mode-option.ts'))};
const started = Date.now();
const session = await launchClaudePty({cwd:${JSON.stringify(dir)}, timeoutMs:5000, env:{POSTURE_RECORD:${JSON.stringify(recordFile)}}});
try {
await session.waitFor('POSTURE_READY', {timeoutMs:2000, pollMs:20});
const read = () => readPlanCountTranscript(session.hermeticConfigDir, ${JSON.stringify(dir)});
const posture = /\\b(expansion|10x|delight|dream|cathedral|opt[\\s-]?in)\\b/i;
const before = hasNativePostAnswerCeoPosture(read(), 'SCOPE EXPANSION', posture, started);
const action = nextCeoPostureContinuation(session.visibleText(), read(), 'SCOPE EXPANSION', started, new Set(), false);
if (action === 'question') await selectPtyNumberedOption(session, 1);
await session.waitFor('POSTURE_FLUSHED', {timeoutMs:2000, pollMs:20});
const after = hasNativePostAnswerCeoPosture(read(), 'SCOPE EXPANSION', posture, started);
await Bun.write(${JSON.stringify(resultFile)}, JSON.stringify({before, action, after}));
} finally { await session.close(); }
`);
const child = Bun.spawn([process.execPath, worker], {
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' }, stdout: 'pipe', stderr: 'pipe',
});
const timer = setTimeout(() => child.kill('SIGKILL'), 8000);
try {
const [code, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
expect(code, stdout + stderr).toBe(0);
expect(JSON.parse(fs.readFileSync(resultFile, 'utf8'))).toEqual({before:false, action:'question', after:true});
const events = fs.readFileSync(recordFile, 'utf8').trim().split('\n').map(line => JSON.parse(line));
expect(events.filter(e => e.type === 'input').map(e => e.data).join('')).toBe('1\r');
expect(() => process.kill(events[0].pid, 0)).toThrow();
} finally {
clearTimeout(timer); child.kill('SIGKILL');
if (fs.existsSync(recordFile)) {
const first = JSON.parse(fs.readFileSync(recordFile, 'utf8').split('\n')[0]!);
try { process.kill(first.pid, 'SIGKILL'); } catch { /* already reaped */ }
}
fs.rmSync(dir, {recursive:true, force:true});
}
}, 10_000);
});
const sidebarModeScreen = fs.readFileSync(path.join(import.meta.dir, 'fixtures/ceo-mode-preview-aa-screen.txt'), 'utf8');
// The AA first SCOPE pane was retained in the terminal failure, but its native
// mode call never flushed. This pending call is synthetic identity coverage.
function sidebarPendingMode() {
return {sessionId:'sidebar-fixture', toolUseId:'sidebar-mode', answered:false, failed:false,
questions:[{header:'Review mode', question:'Which review mode should I use for this plan?', multiSelect:false,
options:[
{label:'SELECTIVE EXPANSION — baseline + cherry-pick (Recommended)'},
{label:'HOLD SCOPE — maximum rigor, no expansions'},
{label:'SCOPE EXPANSION — dream big'},
{label:'SCOPE REDUCTION — strip to essentials'},
]}]};
}
describe('AA mode sidebar preview protocol', () => {
test('submits the independently retained CEO count pane when Notes sits on a wrapped option line', () => {
const screen=fs.readFileSync(path.join(import.meta.dir,'fixtures/ceo-count-mode-preview-aa-screen.txt'),'utf8');
const action=nextCeoModeNavigation(screen,'HOLD SCOPE',new Set());
expect(action.kind).toBe('mode');
if(action.kind!=='mode')throw new Error('Expected mode');
expect(action.index).toBe(1);
expect(planCountQuestionInput(screen,action.question,1)).toBe('1\r');
const native=nativePlanCallFingerprint(sidebarPendingMode(),0,true);
for(const altered of [screen.replace(' bigger',' bigger'),screen.replace(' 4. SCOPE EXPANSION — dream',' Unrelated unnumbered message'),screen.replace(' bigger',' bigger│')]) {
expect(planCountQuestionInput(altered,native,1)).toBe('1');
}
});
test('submits all four offered modes from the exact pane, with or without pending metadata', () => {
for (const [mode,index] of [['SELECTIVE EXPANSION',1],['HOLD SCOPE',2],['SCOPE EXPANSION',3],['SCOPE REDUCTION',4]] as const) {
for (const pending of [undefined,sidebarPendingMode()]) {
const action=nextCeoModeNavigation(sidebarModeScreen,mode,new Set(),pending);
expect(action.kind).toBe('mode');
if(action.kind!=='mode')throw new Error('Expected mode');
expect(action.index).toBe(index);
expect(action.question.nativeCall).toBe(pending);
expect(planCountQuestionInput(sidebarModeScreen,action.question,index)).toBe(`${index}\r`);
}
}
});
test('requires a complete aligned preview and exact notes protocol', () => {
const fp=nativePlanCallFingerprint(sidebarPendingMode(),0,true);
for(const frame of [
sidebarModeScreen.replace('Notes: press n to add notes','Notes: press n to run a command'),
sidebarModeScreen.replace(' Notes:',' Notes:'),
sidebarModeScreen.replace(/┌─+┐/,'no preview box'),
sidebarModeScreen.replace(/└─+┘/,'no preview bottom'),
sidebarModeScreen.replace('└──','└─'),
sidebarModeScreen+'\n☐ Next question\nWhat now?\n 1. Continue\n 2. Stop\nEnter to select · ↑/↓ to navigate · n to add notes · Esc to cancel',
sidebarModeScreen.replace(' · n to add notes',''),
sidebarModeScreen.replace(' · Esc to cancel',''),
sidebarModeScreen.replace('☐ Review mode','quoted Review mode'),
sidebarModeScreen.replace('n to add notes · ','n to add notes · n to add notes · '),
sidebarModeScreen+'\nUnrelated active prompt',
sidebarModeScreen.split('\n').map(line=>'> '+line).join('\n'),
])expect(planCountQuestionInput(frame,fp,3)).toBe('3');
const checkbox=sidebarPendingMode();checkbox.questions[0]!.multiSelect=true;
expect(planCountQuestionInput(sidebarModeScreen,nativePlanCallFingerprint(checkbox,0,true),3)).toBe('3');
});
test('keeps the actual retry missing-target failure and independent answer/posture gates', () => {
const omitted=[
{index:1,label:'HOLD SCOPE — make the client-side plan bulletproof (Recommended)'},
{index:2,label:'SELECTIVE EXPANSION — hold core scope but surface cherry-pick options'},
{index:3,label:'SCOPE REDUCTION — cut to absolute minimum'},
{index:4,label:'Type something.'},{index:5,label:'Chat about this'},
];
expect(()=>findCeoModeOption(omitted,'SCOPE EXPANSION')).toThrow('target "SCOPE EXPANSION" not in option labels');
const t:PlanCountTranscript={status:'ready',calls:[sidebarPendingMode()],assistantMessages:[
{sessionId:'sidebar-fixture',timestamp:new Date().toISOString(),text:'I will explore expansion opportunities.'},
]};
expect(hasNativePostAnswerCeoPosture(t,'SCOPE EXPANSION',/expansion/i,0)).toBe(false);
const foreign=sidebarPendingMode();foreign.questions[0]!.header='Other';
const action=nextCeoModeNavigation(sidebarModeScreen,'SCOPE EXPANSION',new Set(),foreign);
expect(action.kind).toBe('mode');
if(action.kind==='mode')expect(action.question.nativeCall).toBeUndefined();
});
});
test.skipIf(process.platform==='win32')('AA sidebar fake CLI requires submission before native mode posture',async()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'mode-sidebar-'));
const fake=path.join(dir,'fake-claude');const worker=path.join(dir,'worker.ts');const output=path.join(dir,'result.json');
const cases=[
{mode:'SELECTIVE EXPANSION',index:1},{mode:'HOLD SCOPE',index:2},
{mode:'SCOPE EXPANSION',index:3},{mode:'SCOPE REDUCTION',index:4},
{mode:'SCOPE EXPANSION',index:3,digitOnly:true},
].map((item,i)=>({...item,cwd:path.join(dir,String(i)),record:path.join(dir,`${i}.jsonl`)}));
for(const item of cases)fs.mkdirSync(item.cwd);
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import fs from 'node:fs';import path from 'node:path';
const item=JSON.parse(process.env.SIDEBAR_REPLAY);const record=row=>fs.appendFileSync(item.record,JSON.stringify(row)+'\n');
const sid='sidebar-'+process.pid;const file=path.join(process.env.CLAUDE_CONFIG_DIR,'projects',sid,sid+'.jsonl');fs.mkdirSync(path.dirname(file),{recursive:true});
const native=(role,content,extra={})=>fs.appendFileSync(file,JSON.stringify({cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date().toISOString(),message:{role,content},...extra})+'\n');
record({type:'start',pid:process.pid});native('assistant',[{type:'text',text:'Review preview: expansion, rigor, and reduction.'}]);
let focused=1;let done=false;process.stdin.setRawMode?.(true);
process.stdin.on('data',data=>{
const input=data.toString();record({type:'input',input});
if(done){record({type:'unexpected',input});return;}
const digit=/[1-4]/.exec(input)?.[0];
if(digit)setTimeout(()=>{focused=Number(digit);record({type:'focus',focused});},25);
if(!input.includes('\r'))return;done=true;
const answer=item.question.options[focused-1].label;
native('assistant',[{type:'tool_use',name:'AskUserQuestion',id:'mode',input:{questions:[item.question]}}]);
native('user',[{type:'tool_result',tool_use_id:'mode',content:'Answered'}],{toolUseResult:{answers:{[item.question.question]:answer}}});
record({type:'answer',answer});
setTimeout(()=>{native('assistant',[{type:'text',text:'I will apply '+answer+' to assess this plan thoroughly.'}]);process.stdout.write('\r\nANSWER_READY\r\n');},25);
});
process.stdout.write('\x1b[2J\x1b[H'+item.screen.replaceAll('\n','\r\n'));process.on('SIGINT',()=>process.exit(0));process.stdin.resume();
`);fs.chmodSync(fake,0o755);
const moduleUrl=(name:string)=>pathToFileURL(path.join(import.meta.dir,'helpers',name)).href;
fs.writeFileSync(worker,`
import {launchClaudePty,selectPtyNumberedOption,planCountQuestionInput} from ${JSON.stringify(moduleUrl('claude-pty-runner.ts'))};
import {nextCeoModeNavigation,hasNativePostAnswerCeoPosture} from ${JSON.stringify(moduleUrl('ceo-mode-option.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(moduleUrl('plan-count-transcript.ts'))};
const cases=${JSON.stringify(cases)};const screen=${JSON.stringify(sidebarModeScreen)};const question=${JSON.stringify(sidebarPendingMode().questions[0])};
const results=await Promise.all(cases.map(async item=>{
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:5000,env:{SIDEBAR_REPLAY:JSON.stringify({...item,screen,question})}});
try{
await session.waitFor('Which review mode',{timeoutMs:2000,pollMs:20});
const visible=await session.currentScreen();const action=nextCeoModeNavigation(visible,item.mode,new Set());
if(action.kind!=='mode')throw new Error('Mode not captured');
const started=Date.now();const before=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
const beforeMatched=hasNativePostAnswerCeoPosture(before,item.mode,new RegExp(item.mode,'i'),started);
const input=item.digitOnly?String(action.index):planCountQuestionInput(visible,action.question,action.index);
if(input.includes('\\r'))await selectPtyNumberedOption(session,action.index);else session.send(input);
if(!item.digitOnly)await session.waitFor('ANSWER_READY',{timeoutMs:1500,pollMs:20});else await Bun.sleep(150);
const transcript=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
return {mode:item.mode,digitOnly:!!item.digitOnly,index:action.index,input,beforeMatched,matched:hasNativePostAnswerCeoPosture(transcript,item.mode,new RegExp(item.mode,'i'),started)};
}finally{await session.close();}
}));await Bun.write(${JSON.stringify(output)},JSON.stringify(results));
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,EVALS_HERMETIC:'1'},stdout:'pipe',stderr:'pipe'});
const timer=setTimeout(()=>child.kill('SIGKILL'),10000);
try{
const [code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
expect(code,out+err).toBe(0);
const results=JSON.parse(fs.readFileSync(output,'utf8'));
for(const [i,item] of cases.entries()){
expect(results[i]).toEqual({mode:item.mode,digitOnly:!!item.digitOnly,index:item.index,input:String(item.index)+(item.digitOnly?'':'\r'),beforeMatched:false,matched:!item.digitOnly});
const rows=fs.readFileSync(item.record,'utf8').trim().split('\n').map(line=>JSON.parse(line));
expect(rows.filter(row=>row.type==='input').map(row=>row.input)).toEqual(item.digitOnly?[String(item.index)]:[String(item.index),'\r']);
expect(rows.filter(row=>row.type==='answer').length).toBe(item.digitOnly?0:1);
expect(rows.some(row=>row.type==='unexpected')).toBe(false);
expect(()=>process.kill(rows[0].pid,0)).toThrow();
}
}finally{
clearTimeout(timer);child.kill('SIGKILL');await child.exited;
for(const item of cases){
if(!fs.existsSync(item.record))continue;const first=JSON.parse(fs.readFileSync(item.record,'utf8').split('\n')[0]!);
try{
const argv=process.platform==='linux'?fs.readFileSync('/proc/'+first.pid+'/cmdline','utf8').split('\0'):Bun.spawnSync(['ps','-p',String(first.pid),'-o','command='],{timeout:1000}).stdout.toString().trim().split(/\s+/);
if(argv.includes(fake))process.kill(first.pid,'SIGKILL');
}catch{/* owned child already closed */}
}
fs.rmSync(dir,{recursive:true,force:true});
}
},12000);
+117
View File
@@ -0,0 +1,117 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { hasNativePostAnswerCeoPosture, nativeCeoModeAnswer } from './helpers/ceo-mode-option';
import { readPlanCountTranscript, type NativePublicToolEvent } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import captured from './fixtures/ceo-mode-posture-ad.json';
const patterns = {
'HOLD SCOPE': /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i,
'SCOPE EXPANSION': /\b(expansion|10x|delight|dream|cathedral|opt[\s-]?in)\b/i,
};
function replay(index: number, change?: (rows: any[]) => void) {
const item = captured.cases[index]!;
const rows = structuredClone(item.records);
change?.(rows);
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-posture-ad-'));
const project = path.join(dir, 'projects', 'owned');
fs.mkdirSync(project, {recursive:true});
fs.writeFileSync(path.join(project, item.process.sessionId+'.jsonl'), rows.map(row=>JSON.stringify(row)).join('\n')+'\n');
const events: NativePublicToolEvent[]=[];
try { return {item, transcript:readPlanCountTranscript(dir,item.process.cwd,event=>events.push(event)),events}; }
finally { fs.rmSync(dir,{recursive:true,force:true}); }
}
function matches(e: ReturnType<typeof replay>) {
const mode=e.item.mode as keyof typeof patterns;
return hasNativePostAnswerCeoPosture(e.transcript,mode,patterns[mode],e.item.selectedAt,e.events);
}
function rebind(e: ReturnType<typeof replay>) {
const decision=e.transcript.calls[1]!;
e.events[2]!.input={questions:decision.questions};
decision.answers={[decision.questions[0]!.question]:decision.questions[0]!.options[0]!.label};
}
for (const index of [0,1]) describe(`${captured.cases[index]!.mode} actual completed mode application`,()=>{
test('the exact mode answer and concrete scope decision supply posture without finalized prose',()=>{
const e=replay(index);
expect(e.item.actualFailure.state).toBe('failed');
expect(e.transcript.calls.map(call=>call.toolUseId)).toEqual([e.item.modeToolUseId,e.item.decisionToolUseId]);
expect(nativeCeoModeAnswer(e.transcript,e.item.mode as keyof typeof patterns,e.item.selectedAt)?.toolUseId).toBe(e.item.modeToolUseId);
expect(e.transcript.assistantMessages.every(message=>Date.parse(message.timestamp)<e.item.selectedAt)).toBe(true);
expect(matches(e)).toBe(true);
expect(hasNativePostAnswerCeoPosture(e.transcript,e.item.mode as keyof typeof patterns,patterns[e.item.mode as keyof typeof patterns],e.item.selectedAt)).toBe(false);
});
test.each(['wrong selected mode','pending mode','failed mode','answer before selection','pending decision','failed decision',
'foreign session','pre-mode request','reply before request','reply before selection','reply timestamp mismatch','wrong tool','missing request','missing reply','failed public reply','duplicate request','duplicate reply',
'request mismatch','unknown answer','extra question','extra option','multiselect','quoted decision','fenced decision',
'mere mode mention','extra obligation','extra imperative','option imperative'])('%s cannot supply posture',failure=>{
const e=replay(index);const [mode,decision]=e.transcript.calls;const q=decision!.questions[0]!;
switch(failure){
case 'wrong selected mode':mode!.answers![mode!.questions[0]!.question]=index===0?'SCOPE EXPANSION':'HOLD SCOPE';break;
case 'pending mode':mode!.answered=false;break;
case 'failed mode':mode!.failed=true;break;
case 'answer before selection':mode!.answeredAt=new Date(e.item.selectedAt-1).toISOString();break;
case 'pending decision':decision!.answered=false;break;
case 'failed decision':decision!.failed=true;break;
case 'foreign session':decision!.sessionId=e.events[2]!.sessionId=e.events[3]!.sessionId='foreign';break;
case 'pre-mode request':e.events[2]!.timestamp=e.events[0]!.timestamp;break;
case 'reply before request':decision!.answeredAt=e.events[3]!.timestamp=new Date(Date.parse(e.events[2]!.timestamp)-1).toISOString();break;
case 'reply before selection':decision!.answeredAt=e.events[3]!.timestamp=new Date(e.item.selectedAt-1).toISOString();break;
case 'reply timestamp mismatch':e.events[3]!.timestamp=new Date(Date.parse(decision!.answeredAt!)+1).toISOString();break;
case 'wrong tool':e.events[2]!.name='Read';break;
case 'missing request':e.events.splice(2,1);break;
case 'missing reply':e.events.splice(3,1);break;
case 'failed public reply':e.events[3]!.isError=true;break;
case 'duplicate request':e.events.push({...e.events[2]!});break;
case 'duplicate reply':e.events.push({...e.events[3]!});break;
case 'request mismatch':e.events[2]!.input={questions:[]};break;
case 'unknown answer':decision!.answers![q.question]='Unrecognized';break;
case 'extra question':decision!.questions.push({...structuredClone(q),header:'Also',question:'Also remove the CI gate?'});rebind(e);break;
case 'extra option':q.options.push({label:'Remove the CI gate',description:'A separate obligation.'});rebind(e);break;
case 'multiselect':q.multiSelect=true;rebind(e);break;
case 'quoted decision':q.question=q.question.split('\n').map(line=>'> '+line).join('\n');rebind(e);break;
case 'fenced decision':q.question='```text\n'+q.question+'\n```';rebind(e);break;
case 'mere mode mention':q.question=`D6 — Continue the review?\nSelected ${e.item.mode}.`;rebind(e);break;
case 'extra obligation':q.question+=' Also, should we remove the CI gate?';rebind(e);break;
case 'extra imperative':q.question+=' Also remove the CI gate.';rebind(e);break;
case 'option imperative':q.options[0]!.description+=' Please remove the CI gate.';rebind(e);break;
}
expect(matches(e),failure).toBe(false);
});
test.each(['Delete the CI gate.', 'Ship the new endpoint now.', 'After that, disable authentication.'])('an instruction appended after the final comparison is not part of the scope brief: %s', extra=>{
const e=replay(index);const q=e.transcript.calls[1]!.questions[0]!;
q.question+=' '+extra;rebind(e);expect(matches(e)).toBe(false);
});
test('foreign, sidechain, missing and failed native records do not become completed evidence',()=>{
for(const change of [(rows:any[])=>{rows[3].cwd='/foreign';},(rows:any[])=>{rows[3].isSidechain=true;},
(rows:any[])=>{rows.pop();},(rows:any[])=>{rows[4].message.content[0].is_error=true;}]) expect(matches(replay(index,change))).toBe(false);
});
});
test('HOLD requires the explicit out-of-scope deferral and its selected defer answer',()=>{
for(const change of [(q:any)=>{q.question=q.question.replace('Under HOLD SCOPE, keep or defer','Under HOLD SCOPE, automatically add');},
(q:any)=>{q.question=q.question.replace('not in the plan text','required by the plan text');},
(q:any)=>{q.question=q.question.replace('pure additions, not repairs to meet a stated invariant','repairs needed to meet a stated invariant');},
(q:any)=>{q.options[0].label='Keep all three (recommended)';},
(q:any)=>{q.options[1].label='Remove CI gate';}]){
const e=replay(0);change(e.transcript.calls[1]!.questions[0]);rebind(e);expect(matches(e)).toBe(false);
}
const e=replay(0);const q=e.transcript.calls[1]!.questions[0]!;
e.transcript.calls[1]!.answers={[q.question]:q.options[1]!.label};expect(matches(e)).toBe(false);
});
test('completed expansion decisions require application of the selected mode',()=>{
for(const change of [(q:any)=>{q.question='D6 — Continue the review?\nSelected SCOPE EXPANSION.';},
(q:any)=>{q.question=q.question.replace('SCOPE EXPANSION mode','SELECTIVE EXPANSION mode');},
(q:any)=>{q.question=q.question.replace('SCOPE EXPANSION mode','HOLD SCOPE mode');},
(q:any)=>{q.options[1].label='Enable telemetry';}]){
const e=replay(1);change(e.transcript.calls[1]!.questions[0]);rebind(e);expect(matches(e)).toBe(false);
}
});
test('the new replay controls and fixture select the actual periodic mode-routing caller',()=>{
for(const file of ['test/ceo-mode-posture-ad.test.ts','test/fixtures/ceo-mode-posture-ad.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['plan-ceo-mode-routing']);
});
+65
View File
@@ -0,0 +1,65 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { hasNativePostAnswerCeoPosture, hasPostAnswerCeoPosture } from './helpers/ceo-mode-option';
import { readPlanCountTranscript } from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-hold-posture-l.json';
const posture = /\b(rigor|bulletproof|hold\s*scope|maximum\s+rigor)\b/i;
function nativeTranscript() {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-posture-'));
const first = captured.records[0]!;
const project = path.join(root, 'projects', 'fixture');
fs.mkdirSync(project, { recursive: true });
fs.writeFileSync(path.join(project, `${first.sessionId}.jsonl`),
captured.records.map(record => JSON.stringify(record)).join('\n') + '\n');
try {
return readPlanCountTranscript(root, first.cwd);
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
}
describe('Native CEO posture with explanatory parentheses', () => {
test('the actual HOLD answer and subsequent scope analysis establish posture', () => {
const transcript = nativeTranscript();
expect(transcript.status).toBe('ready');
expect(transcript.calls).toHaveLength(1);
expect(transcript.assistantMessages).toHaveLength(1);
expect(transcript.calls[0]!.answers).toEqual({ 'Which review mode should I run for this plan?': 'HOLD SCOPE' });
expect(hasNativePostAnswerCeoPosture(transcript, 'HOLD SCOPE', posture,
Date.parse('2026-09-08T23:23:47.560Z'))).toBe(true);
});
test('tool headings, quoted source and a bare mode echo remain insufficient', () => {
for (const visible of [
'● Bash(command)\nHOLD SCOPE confirmed. Approach B (personal DB views) is the baseline.',
'● Read (/skill.md)\nReview with maximum rigor.',
"● User answered Claude's questions:\nHOLD SCOPE",
'● HOLD SCOPE.',
]) expect(hasPostAnswerCeoPosture(visible, posture)).toBe(false);
for (const text of [
'> HOLD SCOPE confirmed. Approach B (personal DB views) is the baseline.',
'```text\nHOLD SCOPE confirmed. Approach B (personal DB views) is the baseline.\n```',
'HOLD SCOPE confirmed.',
]) {
const transcript = nativeTranscript();
transcript.assistantMessages[0]!.text = text;
expect(hasNativePostAnswerCeoPosture(transcript, 'HOLD SCOPE', posture, 0)).toBe(false);
}
});
test('an announced mode before selection, a wrong answer or missing native answer cannot pass', () => {
const earlier = nativeTranscript();
earlier.assistantMessages[0]!.timestamp = '2026-09-08T23:23:41.665Z';
expect(hasNativePostAnswerCeoPosture(earlier, 'HOLD SCOPE', posture, 0)).toBe(false);
const wrong = nativeTranscript();
wrong.calls[0]!.answers = { 'Which review mode should I run for this plan?': 'SCOPE EXPANSION' };
expect(hasNativePostAnswerCeoPosture(wrong, 'HOLD SCOPE', posture, 0)).toBe(false);
const pending = nativeTranscript();
pending.calls[0]!.answered = false;
expect(hasNativePostAnswerCeoPosture(pending, 'HOLD SCOPE', posture, 0)).toBe(false);
});
});
+105
View File
@@ -0,0 +1,105 @@
import {afterAll, beforeAll, expect, test} from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import {spawnSync} from 'node:child_process';
import {getQuestion} from '../scripts/question-registry';
import {E2E_TOUCHFILES} from './helpers/touchfiles-data';
import {CARVE_GUARDS} from './helpers/carve-guards';
const root=path.resolve(import.meta.dir,'..');
const temp=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-ceo-mode-preference-'));
const section=(s:string)=>s.split('### 0F. Mode Selection\n')[1]!.split('\n### 0D-prelude.')[0]!;
const rendered=new Map<string,string>();
const env=(state:string)=>({...process.env,GSTACK_HOME:state,GSTACK_STATE_ROOT:state});
beforeAll(()=>{
for(const host of ['claude','codex']){
const out=path.join(temp,host),state=path.join(temp,'render-state');
const result=spawnSync(process.execPath,['run','scripts/gen-skill-docs.ts','--host',host,'--out-dir',out],{cwd:root,env:env(state),encoding:'utf8',timeout:120_000});
if(result.status!==0)throw new Error(result.stderr||result.stdout);
rendered.set(host,fs.readFileSync(path.join(out,host==='claude'?'plan-ceo-review':'.agents/skills/gstack-plan-ceo-review','SKILL.md'),'utf8'));
}
},120_000);
afterAll(()=>fs.rmSync(temp,{recursive:true,force:true}));
function modeId(document:string){
const ids=[...section(document).matchAll(/`question_id=([^`]+)`/g)].map(m=>m[1]!);
expect(ids).toHaveLength(1);return ids[0]!;
}
function tuning(document:string){
return document.split('## Question Tuning (skip entirely if')[1]!.split('\n## ')[0]!;
}
function renderedCheck(host:string,id:string){
const match=tuning(rendered.get(host)!).match(/`(printf '%s' "<question summary>" \| ([^`]+)\/gstack-question-preference --check "<id>" --summary-stdin)`/)!;
expect(match).not.toBeNull();
expect(match[2]).toBe(host==='claude'?'~/.claude/skills/gstack/bin':'$GSTACK_BIN');
const quote=(value:string)=>"'"+value.replace(/'/g,"'\\''")+"'";
// Run the rendered command, substituting its documented fields and mapping
// the host's installed executable location to this isolated checkout.
return match[1]!.replace('<question summary>','Select the CEO review mode for the current plan.')
.replace('"<id>"',quote(id))
.replace(match[2]!+'/gstack-question-preference',quote(path.join(root,'bin/gstack-question-preference')));
}
function checkWithPreference(host:string,preference?:string,writeId?:string){
const id=modeId(rendered.get(host)!),state=fs.mkdtempSync(path.join(temp,'state-'));
const run=(args:string[],input?:string)=>spawnSync(path.join(root,'bin/gstack-question-preference'),args,{cwd:root,env:env(state),input,encoding:'utf8',timeout:30_000});
if(preference){const written=run(['--write',JSON.stringify({question_id:writeId??id,preference,source:'plan-tune'})]);expect(written.status).toBe(0);}
const result=spawnSync('bash',['-c',renderedCheck(host,id)],{cwd:root,env:env(state),encoding:'utf8',timeout:30_000});
expect(result.status).toBe(0);return {id,result,run};
}
test('source and both isolated host renders bind the shared check, marker and log to the registered mode identity',()=>{
const source=fs.readFileSync(path.join(root,'plan-ceo-review/SKILL.md.tmpl'),'utf8');
for(const document of [source,...rendered.values()]){
const id=modeId(document),s=section(document);
expect(getQuestion(id)).toMatchObject({id:'plan-ceo-review-mode',skill:'plan-ceo-review',category:'routing',door_type:'two-way'});
expect(s).toContain("preamble's Question Tuning check, marker and log");
expect(s).toContain('`auto_decided: true` when automatic');
expect(s).not.toContain('plan-ceo-review-mode-selection');
}
for(const document of rendered.values()){
expect(tuning(document)).toContain('<gstack-qid:{question_id}>');
expect(tuning(document)).toContain('"question_id":"<id>"');
}
});
test('both rendered canonical checks actually honor a stored never-ask preference',()=>{
for(const host of rendered.keys()){
const {result,run}=checkWithPreference(host,'never-ask');
expect(result.stdout.trim()).toBe('AUTO_DECIDE');
const wrong=run(['--check','plan-ceo-review-mode-selection','--summary-stdin'],'Select the CEO review mode.');
expect(wrong.status).toBe(0);expect(wrong.stdout.trim()).toBe('ASK_NORMALLY');
}
});
test('absent, always-ask and foreign preferences do not authorize either host to select a mode',()=>{
for(const host of rendered.keys()){
for(const [preference,id] of [[undefined,undefined],['always-ask',undefined],['never-ask','plan-design-review-mode']] as const){
expect(checkWithPreference(host,preference,id).result.stdout.trim()).toBe('ASK_NORMALLY');
}
}
});
test('only an explicit user selection or enabled successful mode check bypasses asking',()=>{
for(const document of rendered.values()){
const s=section(document),q=tuning(document);
expect(s).toContain('Ask and wait unless the user explicitly selected a mode or tuning is enabled and the actual mode check exits 0 with `AUTO_DECIDE`');
expect(document).toContain('Question Tuning (skip entirely if `QUESTION_TUNING: false`)');
expect(q).toContain('`AUTO_DECIDE` means choose the recommended option');
expect(q).toContain('Auto-decided [summary] → [option] (your preference). Change with /plan-tune.');
expect(q).toContain('`ASK_NORMALLY` means ask.');
expect(s).toContain('This settles only the mode, not approach or scope approval.');
expect(document).toContain('Do NOT proceed to mode selection (0F) without user approval of the chosen approach.');
expect(s).toContain('Every mode requires explicit user approval for scope changes.');
expect(s).toContain('Keep the approved 0C-bis approach; explain and obtain approval for any mode-required change.');
expect(s).toContain('offer all four modes in one AskUserQuestion');
expect(s).toContain('context defaults for RECOMMENDATION');
expect(s).toContain('Do NOT emit `Completeness: N/10` per option');
expect(s).toContain('Note: options differ in kind, not coverage — no completeness score.');
}
});
test('the new render/runtime regression belongs to the existing auto-decide owner',()=>{
expect(Object.entries(E2E_TOUCHFILES).filter(([,v])=>v.includes('test/ceo-mode-preference-al.test.ts')).map(([k])=>k)).toEqual(['auto-decide-preserved']);
});
test('rendered mode contract stays within the existing canonical skeleton cap',()=>{
// --out-dir changes section-link roots only. Undo that output-location
// substitution before measuring the same canonical bytes as parity-suite.
const out=path.join(temp,'claude').replace(/[.*+?^${}()|[\]\\]/g,'\\$&');
const canonical=rendered.get('claude')!.replace(new RegExp(out+'/([^\\s)`"\'*]+/sections/)','g'),(_m,p1)=>'~/.claude/skills/gstack/'+p1);
const cap=Object.values(CARVE_GUARDS).find(g=>g.skill==='plan-ceo-review')!.maxSkeletonBytes;
expect(Buffer.byteLength(canonical)).toBeLessThanOrEqual(cap);
});
+141
View File
@@ -0,0 +1,141 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { pathToFileURL } from 'node:url';
import { execFileSync } from 'node:child_process';
import { nextCeoModeNavigation } from './helpers/ceo-mode-option';
import { planCountQuestionInput } from './helpers/claude-pty-runner';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
import priorCalls from './fixtures/ceo-mode-prerequisite-o-calls.json';
import directProceedCall from './fixtures/ceo-mode-prerequisite-q-call.json';
import fullAd from './fixtures/ceo-mode-full-ad.json';
const fullAdQuestions=fullAd.cases[0]!.records.find(row=>row.message.role==='assistant')!.message.content[0]!.input.questions;
const calls = [...priorCalls, directProceedCall, {call:{sessionId:'full-ad-projected',toolUseId:'full-ad-prerequisite',questions:fullAdQuestions,answered:false,failed:false}}];
function ownedIdentity(pid: number): string | null {
try { return execFileSync('ps', ['-p', String(pid), '-o', 'lstart=', '-o', 'command='], { encoding: 'utf8', timeout: 5000 }).trim(); }
catch { return null; }
}
function cleanupFake(events: string, fake: string): void {
if (!fs.existsSync(events)) return;
const first = JSON.parse(fs.readFileSync(events, 'utf8').split('\n')[0]!);
if (first.type === 'pid' && Number.isInteger(first.pid) && first.identity?.includes(fake) &&
ownedIdentity(first.pid) === first.identity) {
try { process.kill(first.pid, 'SIGKILL'); } catch { /* already exited */ }
}
}
function pending(i: number): NativePlanQuestionCall {
const { answers, answeredAt, ...call } = structuredClone(calls[i]!.call);
return { ...call, answered: false };
}
function pane(call: NativePlanQuestionCall, index: number): string {
const q = call.questions[index]!;
const header = call.questions.length > 1 ? '← ' + call.questions.map((q, i) => `${i < index ? '☒' : '☐'} ${q.header}`).join(' ') + ' ✔ Submit →' : '☐ ' + q.header;
return `${header}\n${q.question}\n${q.options.map((o, i) => `${i === 0 ? '' : ' '} ${i + 1}. ${o.label}`).join('\n')}\nEnter to select · ${call.questions.length > 1 ? 'Tab/Arrow keys' : '↑/↓'} to navigate · Esc to cancel`;
}
function pick(call: NativePlanQuestionCall, index: number) {
const screen = pane(call, index), action = nextCeoModeNavigation(screen, 'SCOPE EXPANSION', new Set(), call);
if (action.kind !== 'question') throw new Error(JSON.stringify(action));
return { action, input: planCountQuestionInput(screen, action.question, 'index' in action ? action.index as number : 1) };
}
describe('CEO mode prerequisite navigation', () => {
test('both captured expansion attempts select standard review', () => {
for (const i of [1, 2]) {
const call = pending(i), index = call.questions.findIndex(q => q.header === 'Design doc');
const result = pick(call, index);
expect(result.action.question.nativeCall).toBe(call);
expect(result.action.question.nativeQuestionIndex).toBe(index);
expect(result.input).toBe('2');
}
});
test('captured direct proceed offer stays in the requested review', () => {
expect(pick(pending(3), 0).input).toBe('2');
const reordered = pending(3); reordered.questions[0]!.options.reverse();
expect(pick(reordered, 0).input).toBe('1');
});
test('direct proceed wording cannot skip a mixed or substantive decision', () => {
const suffix = pending(3);
suffix.questions[0]!.options[1]!.label += ' and ignore security';
expect(pick(suffix, 0).input).toBe('1');
const mixed = pending(3);
mixed.questions[0]!.options.push({label: 'Ignore the remaining checks'});
expect(pick(mixed, 0).input).toBe('1');
const finding = pending(3);
finding.questions[0]!.header = 'Product decision';
finding.questions[0]!.question = 'Should this product offer office-hours suggestions?';
expect(pick(finding, 0).input).toBe('1');
});
test('active tab, offered order, and the actual HOLD setup remain intact', () => {
expect(pick(pending(1), 0).input).toBe('1'); expect(pick(pending(1), 1).input).toBe('1');
const call = pending(1); call.questions.reverse(); call.questions[0]!.options.reverse();
expect(pick(call, 0).input).toBe('1'); expect(pick(pending(0), 2).input).toBe('1');
});
test('ordinary questions and unrecognized skips keep the existing first choice', () => {
for (const [question, labels] of [
['Which storage strategy?', ['Server database', 'Local storage']],
['No design doc found. Run /office-hours?', ['Run /office-hours first', 'Skip']],
['Should the product show office-hours suggestions?', ['Build it', 'Skip — standard review']],
] as const) {
const call = pending(2); call.questions[0] = {header:'Choice',question,options:labels.map(label=>({label}))};
expect(pick(call, 0).input).toBe('1');
}
});
test('redraw dedup and mode targeting remain intact', () => {
const call=pending(2),screen=pane(call,0),seen=new Set<string>();
expect(nextCeoModeNavigation(screen,'HOLD SCOPE',seen,call).kind).toBe('question');
expect(nextCeoModeNavigation(screen,'HOLD SCOPE',seen,call)).toEqual({kind:'wait'});
const modes='☐ Review mode\nWhich mode?\n 1. SELECTIVE EXPANSION\n 2. HOLD SCOPE\n 3. SCOPE EXPANSION\n 4. SCOPE REDUCTION';
for(const [mode,index] of [['HOLD SCOPE',2],['SCOPE EXPANSION',3]] as const){const a=nextCeoModeNavigation(modes,mode,new Set());expect(a.kind).toBe('mode');if(a.kind==='mode')expect(a.index).toBe(index);}
});
test('fixture and free regression select only the mode-routing eval', () => {
for(const file of ['test/ceo-mode-prerequisite.test.ts','test/fixtures/ceo-mode-prerequisite-o-calls.json','test/fixtures/ceo-mode-prerequisite-q-call.json'])expect(selectTests([file],E2E_TOUCHFILES).selected).toEqual(['plan-ceo-mode-routing']);
});
});
for(const fixtureIndex of [1,2,3,4])test.skipIf(process.platform==='win32')(`fake native PTY skips prerequisite ${fixtureIndex} and confirms target posture`,async()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'ceo-mode-prerequisite-')),fake=path.join(dir,'fake-claude'),events=path.join(dir,'events.jsonl'),worker=path.join(dir,'worker.ts');
const root=path.resolve(import.meta.dir,'..'),call=pending(fixtureIndex),screens=call.questions.map((_,i)=>pane(call,i)),expected=call.questions.map(q=>q.header==='Design doc'?'2':'1');
fs.writeFileSync(fake,`#!${process.execPath}\n`+`
import * as fs from 'node:fs';import * as path from 'node:path';import {execFileSync} from 'node:child_process';
const call=${JSON.stringify(call)},screens=${JSON.stringify(screens)},expected=${JSON.stringify(expected)},out=${JSON.stringify(events)};
const dir=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','mode-prerequisite');fs.mkdirSync(dir,{recursive:true});
const native=(role,content,extra={})=>fs.appendFileSync(path.join(dir,call.sessionId+'.jsonl'),JSON.stringify({cwd:process.cwd(),sessionId:call.sessionId,isSidechain:false,timestamp:new Date().toISOString(),message:{role,content},...extra})+'\\n');
const record=e=>fs.appendFileSync(out,JSON.stringify(e)+'\\n');const show=s=>process.stdout.write('\\x1b[2J\\x1b[H'+s.replace(/\\n/g,'\\r\\n'));
const mode={header:'Review mode',question:'Which mode?',options:[{label:'SELECTIVE EXPANSION'},{label:'HOLD SCOPE'},{label:'SCOPE EXPANSION'},{label:'SCOPE REDUCTION'}]};
let at=0,started=false;const answers={};record({type:'pid',pid:process.pid,identity:execFileSync('ps',['-p',String(process.pid),'-o','lstart=','-o','command='],{encoding:'utf8',timeout:5000}).trim()});process.stdin.setRawMode?.(true);process.stdout.write('FIXTURE_READY');
process.stdin.on('data',data=>{const input=data.toString();record({type:'input',input});
if(!started){started=true;native('assistant',[{type:'tool_use',id:call.toolUseId,name:'AskUserQuestion',input:{questions:call.questions}}]);show(screens[0]);return;}
if(at<expected.length){if(input!==expected[at]){show('Office-hours diversion');return;}answers[call.questions[at].question]=call.questions[at].options[Number(input)-1].label;at++;if(at<expected.length){show(screens[at]);return;}
native('user',[{type:'tool_result',tool_use_id:call.toolUseId,content:'Answered.'}],{toolUseResult:{answers}});native('assistant',[{type:'tool_use',id:'mode-choice',name:'AskUserQuestion',input:{questions:[mode]}}]);show('☐ Review mode\\nWhich mode?\\n 1. SELECTIVE EXPANSION\\n 2. HOLD SCOPE\\n 3. SCOPE EXPANSION\\n 4. SCOPE REDUCTION\\nEnter to select · ↑/↓ to navigate · Esc to cancel');return;}
if(input!=='3')throw new Error('wrong mode');native('user',[{type:'tool_result',tool_use_id:'mode-choice',content:'Answered.'}],{toolUseResult:{answers:{[mode.question]:'SCOPE EXPANSION'}}});native('assistant',[{type:'text',text:'I will explore expansion opportunities for the saved-view workflow.'}],{timestamp:new Date(Date.now()+2).toISOString()});show('● I will explore expansion opportunities.');});setInterval(()=>{},1000);
`,{mode:0o755});
const url=(file:string)=>JSON.stringify(pathToFileURL(path.join(root,file)).href);
fs.writeFileSync(worker,`
import {launchClaudePty,planCountQuestionInput} from ${url('test/helpers/claude-pty-runner.ts')};import {nextCeoModeNavigation,hasNativePostAnswerCeoPosture} from ${url('test/helpers/ceo-mode-option.ts')};import {readPlanCountTranscript} from ${url('test/helpers/plan-count-transcript.ts')};
const cwd=${JSON.stringify(dir)},s=await launchClaudePty({cwd,observeScreen:true,timeoutMs:5000}),seen=new Set();let selectedAt=0,matched=false;
try{await s.waitFor('FIXTURE_READY',{timeoutMs:2000,pollMs:20});s.send('/plan-ceo-review\\r');for(let i=0;i<160;i++){await Bun.sleep(20);const screen=await s.currentScreen(),t=readPlanCountTranscript(s.hermeticConfigDir,cwd);if(selectedAt&&hasNativePostAnswerCeoPosture(t,'SCOPE EXPANSION',/\\bexpansion\\b/i,selectedAt)){matched=true;break;}const a=nextCeoModeNavigation(screen,'SCOPE EXPANSION',seen,t.calls.find(c=>!c.answered&&!c.failed),s.visibleText());if(a.kind==='question')s.send(planCountQuestionInput(screen,a.question,'index' in a?a.index:1));if(a.kind==='mode'){selectedAt=Date.now();s.send(planCountQuestionInput(screen,a.question,a.index));}}if(!matched)throw new Error('mode posture absent after prerequisite '+s.visibleText());process.stdout.write('native-mode-confirmed');}finally{await s.close();}
`);
const child=Bun.spawn([process.execPath,worker],{cwd:root,env:{...process.env,BROWSE_TERMINAL_BINARY:fake,EVALS_HERMETIC:'1'},stdout:'pipe',stderr:'pipe'}),timer=setTimeout(()=>child.kill('SIGKILL'),10000);
try{const [code,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);expect(code,stdout+stderr+(fs.existsSync(events)?fs.readFileSync(events,'utf8'):'no fake events')).toBe(0);expect(stdout).toBe('native-mode-confirmed');const rows=fs.readFileSync(events,'utf8').trim().split('\n').map(l=>JSON.parse(l));expect(rows.filter(e=>e.type==='input').map(e=>e.input)).toEqual(['/plan-ceo-review\r',...expected,'3']);expect(()=>process.kill(rows[0].pid,0)).toThrow();}
finally{clearTimeout(timer);child.kill('SIGKILL');cleanupFake(events,fake);fs.rmSync(dir,{recursive:true,force:true});}
},12000);
test.skipIf(process.platform === 'win32')('outer cleanup requires the recorded fake process identity', async () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-mode-owned-cleanup-'));
const fake = path.join(dir, 'fake.ts'), events = path.join(dir, 'events.jsonl');
fs.writeFileSync(fake, 'setInterval(() => {}, 1000);');
const child = Bun.spawn([process.execPath, fake], {stdout:'ignore',stderr:'ignore'});
try {
const identity = ownedIdentity(child.pid);
expect(identity?.includes(fake)).toBe(true);
fs.writeFileSync(events, JSON.stringify({type:'pid',pid:child.pid,identity:'foreign identity'})+'\n');
cleanupFake(events, fake);
expect(() => process.kill(child.pid, 0)).not.toThrow();
fs.writeFileSync(events, JSON.stringify({type:'pid',pid:child.pid,identity})+'\n');
cleanupFake(events, fake);
await child.exited;
expect(() => process.kill(child.pid, 0)).toThrow();
expect(() => cleanupFake(events, fake)).not.toThrow();
} finally {child.kill('SIGKILL');fs.rmSync(dir,{recursive:true,force:true});}
},5000);
+162
View File
@@ -0,0 +1,162 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import captured from './fixtures/ceo-numbered-brief-ak.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const call = (index = 4): any => structuredClone(captured.calls[index]);
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
function edit(c: any, change: (s: string) => string) {
const q = c.questions[0], answer = c.answers[q.question];
q.question = change(q.question); c.answers = { [q.question]: answer };
}
function offered(c: any, change: (o: any, i: number) => void) {
const q = c.questions[0], selected = q.options.findIndex((o: any) => o.label === c.answers[q.question]);
q.options.forEach(change); c.answers = { [q.question]: q.options[selected].label };
}
for (const [index, name] of [[4, 'email ordering'], [5, 'raw SQL'], [6, 'missing automated tests'], [7, 'N+1 read']] as const) {
test(`actual completed ${name} brief starts substantive CEO review`, () => {
expect(ceoFirstReviewAUQ(fp(call(index)))).toBe(true);
});
}
test('the complete captured phase retains routing and factual clarification as setup', () => {
let started = false;
const phases = captured.calls.map(c => {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
return phase.preReview;
});
expect(phases).toEqual([true, true, true, true, false, false, false, false]);
for (const c of captured.calls.slice(0, 4)) expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
});
test('issue ownership survives equivalent separators, optional qids and consistent renumbering', () => {
for (let index = 4; index < 8; index++) {
for (const separator of ['—', '', '-']) {
const c = call(index); edit(c, s => s.replace(/^D\d+ — /, `D12 ${separator} `).replace(/\s*<gstack-qid:[^>]+>\s*$/, ''));
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
const c = call(index), old = index - 2;
edit(c, s => s.replace(`Issue ${old}:`, 'Issue 19:').replace(new RegExp('\\b' + old + '([A-Z])\\b', 'g'), '19$1'));
offered(c, o => { o.label = o.label.replace(/^\d+/, '19'); });
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
c.answers[c.questions[0].question] = c.questions[0].options[2].label;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
});
test('display effort and positive bullet decoration do not change an offered action', () => {
for (let index = 4; index < 8; index++) for (const change of [
(s: string) => s.replace(/Human [^.]+\. /, 'Human 2 days / CC 30 minutes. '),
(s: string) => s.replace(/Human [^.]+\. /, '').replace(/✅ /g, ''),
(s: string) => s.replace(/✅ /g, '✅ '),
]) {
const c = call(index); offered(c, o => { o.description = change(o.description); });
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
const suite = call(6); offered(suite, o => { o.label = o.label.replace('unit + integration', 'unit and integration'); });
expect(ceoFirstReviewAUQ(fp(suite))).toBe(true);
});
test('native completion, the offered answer and exact fingerprint remain mandatory', () => {
for (let index = 4; index < 8; index++) for (const mutate of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.sessionId = ''; },
(c: any) => { c.toolUseId = ''; },
(c: any) => { c.answers = {}; },
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
(c: any) => { c.questions[0].multiSelect = true; },
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
(c: any) => { c.questions[0].header = 'Approach'; },
(c: any) => { c.questions[0].options[1].description = ''; },
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[0].label; },
(c: any) => edit(c, s => s.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-eng-review-finding>')),
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
for (let index = 4; index < 8; index++) {
const f = fp(call(index));
expect(ceoFirstReviewAUQ({ ...f, signature: 'foreign:tool' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...f, nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...f, options: f.options.slice(1) })).toBe(false);
}
});
test('a current issue cannot borrow another issue identity or recommendation', () => {
for (let index = 4; index < 8; index++) for (const mutate of [
(c: any) => edit(c, s => s.replace(/Issue \d+:/, 'Issue 99:')),
(c: any) => { c.questions[0].header = 'Finding 99'; },
(c: any) => { c.questions[0].options[1].label = '99B: Foreign choice'; },
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[1].label.replace(/B:/, 'A:'); },
(c: any) => edit(c, s => s.replace(/^Recommendation: \d+[A-Z]/m, 'Recommendation: 99A')),
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, '')),
(c: any) => edit(c, s => s + '\nRecommendation: 99A'),
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
});
test('source, hypothetical and withdrawn assessments do not start current review', () => {
for (let index = 4; index < 8; index++) for (const change of [
(s: string) => 'Example: ' + s,
(s: string) => '> ' + s,
(s: string) => '```\n' + s + '\n```',
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If approved, '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Suppose '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a quoted source excerpt. '),
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
(s: string) => s.replace(/^ELI10: .+$/m, ''),
(s: string) => s + '\nThis issue has been withdrawn.',
(s: string) => s + `\nIssue ${index - 2} is resolved.`,
(s: string) => s + '\nNo current issue remains.',
]) { const c = call(index); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
for (let index = 4; index < 8; index++) {
const c = call(index); edit(c, s => s + '\nOld note: "This issue has been withdrawn."');
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
});
test('the new declarative findings still need an asserted current defect', () => {
for (const [index, title] of [
[5, 'the lookup does not interpolate request.params.userId into a raw SQL fragment.'],
[5, 'the lookup no longer interpolates request.params.userId into a raw SQL fragment.'],
[5, 'the lookup used to interpolate user input into a raw SQL string.'],
[6, 'automated tests are planned for the new payment handler.'],
[7, 'the handler no longer fetches each order in a loop (N+1).'],
[7, 'the handler reads all orders with one query.'],
] as const) {
const c = call(index);
edit(c, s => s.replace(/^(D\d+ — Issue \d+: ).+$/m, '$1' + title)
.replace(/^ELI10: .+$/m, title.includes('used to')
? 'ELI10: The previous lookup used to interpolate user input into a raw SQL string. The current lookup uses bound parameters and has no injection risk.'
: 'ELI10: The current implementation satisfies the stated contract.'));
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
});
test('only an offered current amendment can supply remedy evidence', () => {
for (let index = 4; index < 8; index++) for (const description of [
'Keep this advisory report for reference.',
'Human ~3h / CC ~15min. ❌ Add bounded error handling.',
'Human ~3h / CC ~15min. ❌ A prior proposal. Add bounded error handling.',
'Human ~3h / CC ~15min. ✅ "Add bounded error handling."',
'Human ~3h / CC ~15min. ✅ If approved, add bounded error handling.',
'Human ~3h / CC ~15min. ✅ Write the completed report.',
'Historical source excerpt: ✅ Add bounded error handling.',
'If approved: ✅ Add bounded error handling.',
'Hypothetical example: ✅ Add bounded error handling.',
'The following is a quoted source excerpt. ✅ Add bounded error handling.',
'The following is a hypothetical example. ✅ Add bounded error handling.',
]) {
const c = call(index);
offered(c, (o, i) => { o.label = `${index - 2}${String.fromCharCode(65 + i)}: Consider candidate ${i}`; o.description = description; });
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
});
test('only the existing CEO finding-count owner selects this regression fixture', () => {
for (const dependency of ['test/ceo-numbered-brief-ak.test.ts', 'test/fixtures/ceo-numbered-brief-ak.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name))
.toEqual(['plan-ceo-finding-count']);
}
});
+160
View File
@@ -0,0 +1,160 @@
import { expect, test } from 'bun:test';
import fixture from './fixtures/ceo-parenthesized-issue-ah.json';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const calls = () => structuredClone(fixture.calls) as NativePlanQuestionCall[];
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
function reanswer(c: NativePlanQuestionCall) {
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
return c;
}
function changed(original: NativePlanQuestionCall, mutate: (c: NativePlanQuestionCall) => void) {
const c = structuredClone(original); mutate(c); return reanswer(c);
}
test('both exact completed Issue questions start review with descriptive headers', () => {
for (const c of calls()) expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
let started = false;
const counts = { setup: 0, review: 0 };
for (const c of [...fixture.setupCalls, ...calls()] as NativePlanQuestionCall[]) {
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted;
counts[phase.preReview ? 'setup' : 'review']++;
}
expect(counts).toEqual({ setup: 2, review: 2 });
expect(fixture.historicalOutcome).toBe('no_review_questions');
});
test('fixture questions and selected answers are exact owned public request/result projections', () => {
for (const c of calls()) {
const requests = fixture.publicEvents.filter(e => e.record.message.content.some(b => 'id' in b && b.id === c.toolUseId));
const results = fixture.publicEvents.filter(e => e.record.message.content.some(b => 'tool_use_id' in b && b.tool_use_id === c.toolUseId));
expect(requests).toHaveLength(1); expect(results).toHaveLength(1);
expect(requests[0]!.record.sessionId).toBe(c.sessionId);
expect(results[0]!.record.sessionId).toBe(c.sessionId);
const request = requests[0]!.record.message.content.find(b => 'id' in b && b.id === c.toolUseId) as any;
expect(request.input.questions).toEqual(c.questions);
const result = results[0]!.record.message.content.find(b => 'tool_use_id' in b && b.tool_use_id === c.toolUseId) as any;
expect(result.is_error).not.toBe(true);
expect(result.content).toContain(`"${c.questions[0]!.question}"="${c.answers![c.questions[0]!.question]}"`);
expect(c.answeredAt).toBe(results[0]!.record.timestamp);
}
});
test('complete current native identity and actual offered answer remain required', () => {
for (const original of calls()) {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.answered = false; },
(c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { c.sessionId = ''; },
(c: NativePlanQuestionCall) => { c.toolUseId = ''; },
(c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.answers = { 'prior question': c.questions[0]!.options[0]!.label }; },
(c: NativePlanQuestionCall) => { c.answers![c.questions[0]!.question] = 'foreign answer'; },
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
]) {
const c = structuredClone(original); mutate(c);
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
expect(ceoFirstReviewAUQ({ ...fp(original), signature: 'foreign:call' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(original), nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...fp(original), options: [] })).toBe(false);
}
});
test('title, recommendation, every option and any numbered header share one issue identity', () => {
for (const original of calls()) {
const n = /\(Issue (\d+)\)/.exec(original.questions[0]!.question)![1]!;
for (const header of [`Finding ${n}`, `Issue ${n}`, `F${n}`]) {
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.header = header; })))).toBe(true);
}
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation:.*\n/m, ''); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation: \d+A/m, 'Recommendation: 99A'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation: \d+A/m, `Recommendation: ${n}Z`); },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = '99B) Different issue'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = c.questions[0]!.options[0]!.label; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = ''; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding 99'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/D\d+ \(Issue \d+\) — /, ''); c.questions[0]!.header = `Issue ${n}`; },
]) expect(ceoFirstReviewAUQ(fp(changed(original, mutate)))).toBe(false);
}
});
test('source, conditional and stale or withdrawn briefs do not start review', () => {
for (const original of calls()) {
for (const prefix of ['Example: ', 'If requested: ', '> ', ' ', '```text\n']) {
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = prefix + c.questions[0]!.question; })))).toBe(false);
}
for (const framing of ['If this hypothetical plan were adopted, ', 'Example: ', 'Historical example only. ', 'Quoted assessment: ']) {
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = c.questions[0]!.question.replace('ELI10: ', `ELI10: ${framing}`); })))).toBe(false);
}
for (const tail of ['No current defect exists.', 'Correction: this issue is already resolved.', 'This question is only an example.', 'I withdraw this finding.']) {
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question += '\n' + tail; })))).toBe(false);
}
for (const prefix of ['> ', ' ', '```\n']) {
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = c.questions[0]!.question.replace(/^ELI10:/m, prefix + 'ELI10:'); })))).toBe(false);
}
// Later attributed source text does not withdraw a present decision.
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question += '\nAn old note said: "No current defect exists."'; })))).toBe(true);
}
});
test('qid and setup exclusions apply before the new identity form', () => {
for (const original of calls()) {
for (const mutate of [
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:plan-ceo-review-extra>'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-eng-review-validation>'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-ceo-review-approach>'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-ceo-review-mode>'); },
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = 'HOLD SCOPE'; },
]) expect(ceoFirstReviewAUQ(fp(changed(original, mutate)))).toBe(false);
}
});
test('identity recognition is independent of observed component and numbering', () => {
for (const original of calls()) {
const c = changed(original, c => {
const q = c.questions[0]!; const n = /\(Issue (\d+)\)/.exec(q.question)![1]!;
q.header = 'Notification state';
q.question = q.question.replace(/^D\d+/, 'D24').replace(`(Issue ${n})`, '(Issue 17)')
.replace(new RegExp(`\\b${n}([ABC])\\b`, 'g'), '17$1')
.replace(/Stripe/g, 'PaymentProvider').replace(/email/g, 'notification').replace(/userId/g, 'accountKey');
q.options.forEach(o => { o.label = o.label.replace(new RegExp(`^${n}`), '17'); });
});
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options.at(-1)!.label };
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
}
});
test('new fixture and controls select only their existing CEO count owner', () => {
for (const file of ['test/ceo-parenthesized-issue-ah.test.ts', 'test/fixtures/ceo-parenthesized-issue-ah.json']) {
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['plan-ceo-finding-count']);
}
});
test('numbered administrative, literal-only, hypothetical and withdrawn briefs are not defects', () => {
for (const original of calls()) {
for (const mutate of [
(q: NativePlanQuestionCall['questions'][number]) => {
const n = /\(Issue (\d+)\)/.exec(q.question)![1]!;
q.question = q.question.replace(/^(D\d+ \(Issue \d+\) — ).*/, '$1How should we archive this completed review?')
.replace(/^ELI10:.*$/m, 'ELI10: The review is complete. This choice only saves the finished report.');
q.options.forEach((o, i) => { o.label = `${n}${String.fromCharCode(65 + i)}) Save report format ${i}`; o.description = 'Store the completed review report.'; });
},
(q: NativePlanQuestionCall['questions'][number]) => { q.question += '\nThere is no defect or unresolved issue; this is a historical example.'; },
(q: NativePlanQuestionCall['questions'][number]) => { q.question = q.question.replace(/^ELI10:.*$/m, 'ELI10: `The plan has no error handling.`'); },
(q: NativePlanQuestionCall['questions'][number]) => { q.question = q.question.replace(/^(D\d+ \(Issue \d+\) — ).*/, '$1What should happen if a hypothetical future handler lacked error handling?'); },
]) {
const c = changed(original, c => mutate(c.questions[0]!));
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
}
});
+165
View File
@@ -0,0 +1,165 @@
import { describe, expect, test } from 'bun:test';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { nextCeoPostureContinuation } from './helpers/ceo-mode-option';
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
const selectedAt = Date.parse('2026-09-09T07:57:52Z');
const questions = ['Autosave', 'View tabs', 'Deep links', 'Team views'].map(header => ({
header, question: `Add ${header.toLowerCase()} to the plan?`, options: [{label:'Add'}, {label:'Defer'}],
}));
function transcript(count = 4, native = false): PlanCountTranscript {
return {status:'ready', assistantMessages:[], calls:[{
sessionId:'mode-session', toolUseId:'mode', answered:true, failed:false,
answeredAt:'2026-09-09T07:57:54Z', unansweredQuestionIndices:[],
questions:[{header:'Review mode',question:'Which mode?',options:[{label:'HOLD SCOPE'},{label:'SCOPE EXPANSION'}]}],
answers:{'Which mode?':'SCOPE EXPANSION'},
}, ...(native ? [{sessionId:'mode-session',toolUseId:'downstream',answered:false,failed:false,questions:structuredClone(questions.slice(0,count))}] : [])]};
}
function screen(next: number, count = 4) {
const bar = '← ' + questions.slice(0,count).map((q,i)=>(i<next?'☒ ':'☐ ')+q.header).join(' ')+' ✔ Submit →\n';
return next === count ? bar+'Review your answers\nReady to submit your answers?\n1.Submit answers\n2.Cancel\n'
: bar+questions[next]!.question+'\n1.Add\n2.Defer\nEnter to select · Tab/Arrow keys to navigate · Esc to cancel\n';
}
const step = (visible: string, t: PlanCountTranscript, seen: Set<string>, started: boolean) =>
nextCeoPostureContinuation(visible,t,'SCOPE EXPANSION',selectedAt,seen,started);
describe('one CEO posture continuation means one complete native call',()=>{
test('two and four tabs complete once with eager or delayed native metadata',()=>{
for(const count of [2,4]) for(const native of [false,true]) {
const t=transcript(count,native),seen=new Set<string>();
for(let i=0;i<count;i++) {
expect(step(screen(i,count),t,seen,i>0)).toBe('question');
expect(step(screen(i,count),t,seen,true)).toBeNull();
}
expect(step(screen(count,count),t,seen,true)).toBe('submission');
expect(step(screen(count,count),t,seen,true)).toBeNull();
expect(step(screen(0,count),t,seen,true)).toBeNull();
expect(step('☐ New issue\nFix another issue?\n1.Fix\n2.Keep',t,seen,true)).toBeNull();
}
});
test('late native metadata binds the existing packet without replaying a tab',()=>{
const seen=new Set<string>();
expect(step(screen(0),transcript(),seen,false)).toBe('question');
expect(step(screen(0),transcript(4,true),seen,true)).toBeNull();
expect(step(screen(1),transcript(4,true),seen,true)).toBe('question');
const foreign=transcript(4,true);foreign.calls[1]!.toolUseId='foreign';
expect(step(screen(2),foreign,seen,true)).toBeNull();
expect(step(screen(2),transcript(4,true),seen,true)).toBe('question');
});
test('changed bars, skipped answers, foreign calls and incomplete panels cannot extend the budget',()=>{
const invalid = [screen(0),screen(2),screen(4),screen(1).replace('Deep links','Other work'),
screen(1).replace('Tab/Arrow keys to navigate','navigate'),
'☐ New issue\nFix another issue?\n1.Fix\n2.Keep'];
for(const visible of invalid) {
const seen=new Set<string>();expect(step(screen(0),transcript(),seen,false)).toBe('question');
expect(step(visible,transcript(),seen,true)).toBeNull();
}
for(const mutate of [
(t:PlanCountTranscript)=>{t.calls[1]!.sessionId='foreign';},
(t:PlanCountTranscript)=>{t.calls[1]!.questions[1]!.question='Unrelated request?';},
(t:PlanCountTranscript)=>{t.calls[0]!.failed=true;},
(t:PlanCountTranscript)=>{t.calls[0]!.answers={'Which mode?':'HOLD SCOPE'};},
]) {
const seen=new Set<string>();expect(step(screen(0),transcript(),seen,false)).toBe('question');
const t=transcript(4,true);mutate(t);
expect(step(screen(1),t,seen,true)).toBeNull();
}
expect(step(screen(1),transcript(),new Set(),false)).toBeNull();
expect(step(screen(0),transcript(),new Set(),true)).toBeNull();
});
test('late metadata at Submit must match every tab and remain pending in the selected session',()=>{
const mutations: Array<(t: PlanCountTranscript)=>void> = [
t=>{t.calls[1]!.sessionId='foreign';},
t=>{t.calls[1]!.questions[0]!.header='Foreign header';},
t=>{t.calls[1]!.questions[0]!.question='Unrelated request?';},
t=>{t.calls[1]!.answered=true;},
t=>{t.calls[1]!.failed=true;},
];
for(const mutate of mutations) {
const seen=new Set<string>();
for(let i=0;i<4;i++)expect(step(screen(i),transcript(),seen,i>0)).toBe('question');
const t=transcript(4,true);mutate(t);
expect(step(screen(4),t,seen,true)).toBeNull();
}
const seen=new Set<string>();
for(let i=0;i<4;i++)expect(step(screen(i),transcript(),seen,i>0)).toBe('question');
expect(step(screen(4),transcript(4,true),seen,true)).toBe('submission');
});
test('single-question continuation remains one question even if the caller flag is stale',()=>{
const t=transcript(),seen=new Set<string>();
const one='☐ Architecture\nGuard the lookup?\n1.Add guard\n2.Defer';
expect(step(one,t,seen,false)).toBe('question');
expect(step(one.replace('lookup','cache'),t,seen,false)).toBeNull();
expect(step(screen(0),t,seen,false)).toBeNull();
});
});
test.skipIf(process.platform==='win32')('real fake CLI flushes native posture only after all four tabs and Submit',async()=>{
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'ceo-posture-packet-'));
const fake=path.join(dir,'fake-claude'),worker=path.join(dir,'worker.ts'),events=path.join(dir,'events.jsonl'),result=path.join(dir,'result.json');
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
import fs from 'node:fs';import path from 'node:path';
const item=JSON.parse(process.env.POSTURE_PACKET);const sid='mode-session';
const file=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture',sid+'.jsonl');fs.mkdirSync(path.dirname(file),{recursive:true});
const native=(role,content,extra={})=>fs.appendFileSync(file,JSON.stringify({cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date().toISOString(),message:{role,content},...extra})+'\n');
const log=e=>fs.appendFileSync(item.events,JSON.stringify(e)+'\n');
log({type:'start',pid:process.pid});
native('assistant',[{type:'tool_use',id:'mode',name:'AskUserQuestion',input:{questions:[{header:'Mode',question:'Which mode?',options:[{label:'HOLD SCOPE'},{label:'SCOPE EXPANSION'}]}]}}]);
native('user',[{type:'tool_result',tool_use_id:'mode',content:'Answered'}],{toolUseResult:{answers:{'Which mode?':'SCOPE EXPANSION'}}});
let index=0,done=false;
const render=()=>process.stdout.write('\x1b[2J\x1b[H'+item.screens[index].replaceAll('\n','\r\n'));
process.stdin.setRawMode?.(true);process.stdin.on('data',data=>{
const input=data.toString();log({type:'input',input,index});
if(done)throw Error('input after completion');
if(index<4){if(input!=='1')throw Error('native shortcut must not queue Enter');index++;render();return;}
if(input!=='\r')throw Error('expected Submit');
native('assistant',[{type:'text',text:'I will explore expansion opportunities that improve saved project views.'},{type:'tool_use',id:'downstream',name:'AskUserQuestion',input:{questions:item.questions}}]);
native('user',[{type:'tool_result',tool_use_id:'downstream',content:'Answered'}],{toolUseResult:{answers:Object.fromEntries(item.questions.map(q=>[q.question,'Add']))}});
done=true;process.stdout.write('\x1b[2J\x1b[HPOSTURE_FLUSHED\r\n');
});process.on('SIGINT',()=>process.exit(0));process.stdin.resume();render();
`);fs.chmodSync(fake,0o755);
const module=(name:string)=>pathToFileURL(path.join(import.meta.dir,'helpers',name)).href;
fs.writeFileSync(worker,`
import {launchClaudePty,capturePlanCountQuestion,planCountQuestionInput} from ${JSON.stringify(module('claude-pty-runner.ts'))};
import {readPlanCountTranscript} from ${JSON.stringify(module('plan-count-transcript.ts'))};
import {nextCeoPostureContinuation,hasNativePostAnswerCeoPosture} from ${JSON.stringify(module('ceo-mode-option.ts'))};
const start=Date.now(),seen=new Set();let started=false;
const session=await launchClaudePty({cwd:${JSON.stringify(dir)},timeoutMs:7000,observeScreen:true,env:{POSTURE_PACKET:${JSON.stringify(JSON.stringify({events,questions,screens:Array.from({length:5},(_,i)=>screen(i))}))}}});
try {
await session.waitFor('Autosave',{timeoutMs:2000,pollMs:20});
const read=()=>readPlanCountTranscript(session.hermeticConfigDir,${JSON.stringify(dir)});
const before=hasNativePostAnswerCeoPosture(read(),'SCOPE EXPANSION',/expansion/,start);
while(Date.now()-start<5000){
const t=read();if(hasNativePostAnswerCeoPosture(t,'SCOPE EXPANSION',/expansion/,start))break;
const visible=await session.currentScreen();
const action=nextCeoPostureContinuation(visible,t,'SCOPE EXPANSION',start,seen,started,session.visibleText());
if(action==='question'){started=true;const pending=t.calls.find(c=>!c.answered&&!c.failed);const fp=capturePlanCountQuestion(visible,new Set(),0,false,pending);session.send(planCountQuestionInput(visible,fp,1));}
else if(action==='submission')session.send('\\r');
await Bun.sleep(25);
}
const t=read();await Bun.write(${JSON.stringify(result)},JSON.stringify({before,after:hasNativePostAnswerCeoPosture(t,'SCOPE EXPANSION',/expansion/,start),calls:t.calls}));
} finally {await session.close();}
`);
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,EVALS_HERMETIC:'1'},stdout:'pipe',stderr:'pipe'});
const timer=setTimeout(()=>child.kill('SIGKILL'),10000);
try{
const [code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
expect(code,out+err).toBe(0);const r=JSON.parse(fs.readFileSync(result,'utf8'));
expect(r.before).toBe(false);expect(r.after).toBe(true);
expect(r.calls).toHaveLength(2);expect(r.calls[1].unansweredQuestionIndices).toEqual([]);
const rows=fs.readFileSync(events,'utf8').trim().split('\n').map(s=>JSON.parse(s));
expect(rows.filter(r=>r.type==='input').map(r=>r.input)).toEqual(['1','1','1','1','\r']);
expect(()=>process.kill(rows[0].pid,0)).toThrow();
} finally {
clearTimeout(timer);child.kill('SIGKILL');await child.exited;
if(fs.existsSync(events))try{
const pid=JSON.parse(fs.readFileSync(events,'utf8').split('\n')[0]!).pid;
const argv=process.platform==='linux'?fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0'):Bun.spawnSync(['ps','-p',String(pid),'-o','command='], { timeout: 1_000 }).stdout.toString().trim().split(/\s+/);
if(argv.includes(fake))process.kill(pid,'SIGKILL');
}catch{}
fs.rmSync(dir,{recursive:true,force:true});
}
},12000);
+75
View File
@@ -0,0 +1,75 @@
import {expect,test} from 'bun:test';
import {capturePlanCountQuestion,nativePlanCallFingerprint,planCountPrerequisitePick,planCountQuestionInput} from './helpers/claude-pty-runner';
import {nextCeoModeNavigation} from './helpers/ceo-mode-option';
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
import captured from './fixtures/ceo-prerequisite-ad-v2.json';
function pending(){const c=structuredClone(captured.completedCall) as NativePlanQuestionCall;c.answered=false;delete c.answers;delete c.answeredAt;delete c.unansweredQuestionIndices;return c;}
// Native identities and questions are exact; pending panes are synthetic projections.
function pane(c:NativePlanQuestionCall,index:number){const q=c.questions[index]!;return [
c.questions.length>1?'← '+c.questions.map((v,i)=>`${i<index?'☒':'☐'} ${v.header}`).join(' ')+' ✔ Submit →':'☐ '+q.header,
q.question,...q.options.map((v,i)=>`${i?' ':''} ${i+1}. ${v.label}`),
`Enter to select · ${c.questions.length>1?'Tab/Arrow keys':'↑/↓'} to navigate · Esc to cancel`].join('\n');}
function frame(c:NativePlanQuestionCall,index:number){const visible=pane(c,index);return {visible,active:capturePlanCountQuestion(visible,new Set(),0,true,c)!,routing:nativePlanCallFingerprint(c,0,true)};}
test('AD v2 actual comma prerequisite selects standard review on its active native tab',()=>{
const actual=captured.completedCall,q=actual.questions[2]!;
expect(actual.answered).toBe(true);expect(actual.failed).toBe(false);expect(actual.answers[q.question]).toBe('Run /office-hours now');
const c=pending(),f=frame(c,2);expect(f.active.nativeQuestionIndex).toBe(2);
expect(planCountPrerequisitePick(f.routing,f.active)).toBe(2);
const a=nextCeoModeNavigation(f.visible,'HOLD SCOPE',new Set(),c);expect(a.kind).toBe('question');
if(a.kind==='question')expect(planCountQuestionInput(f.visible,a.question,a.index)).toBe('2');
});
test('AD v2 prerequisite presentation and actual order do not choose the action',()=>{
for(const header of ['Office hours','Design doc','Prerequisite'])for(const reverse of [false,true]){
const c=pending();c.questions[2]!.header=header;c.questions[2]!.question=c.questions[2]!.question.replace(/^D3 — /,'D41: ');
if(reverse)c.questions[2]!.options.reverse();const f=frame(c,2);
expect(planCountPrerequisitePick(f.routing,f.active)).toBe(reverse?1:2);
for(const index of [0,1]){const other=frame(c,index);expect(planCountPrerequisitePick(other.routing,other.active)).toBeNull();}
}
const c=pending();c.questions=[c.questions[2]!];let f=frame(c,0);expect(planCountPrerequisitePick(f.routing,f.active)).toBe(2);
c.questions[0]!.question='Run /office-hours now or proceed with standard review?\nNo design doc exists for the current feature. The scoped review can begin on the supplied plan.';
c.questions[0]!.options[0]!.description='Create the design document first; then resume standard review.';
for(const description of ['Proceed with standard review.','Proceed straight to Step 0 of the review.']){
c.questions[0]!.options[1]!.description=description;f=frame(c,0);expect(planCountPrerequisitePick(f.routing,f.active)).toBe(2);
}
});
test('AD v2 prerequisite declines no other task or conditional action',()=>{
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
c=>{c.questions[2]!.question=c.questions[2]!.question.replace(/^.*\n/,'Should we deploy the feature now?\n');},
c=>{c.questions[2]!.question='Example: '+c.questions[2]!.question;},
c=>{c.questions[2]!.question='> '+c.questions[2]!.question;},
c=>{c.questions[2]!.question='```text\n'+c.questions[2]!.question+'\n```';},
c=>{c.questions[2]!.question=c.questions[2]!.question.replace('Run /office-hours first, or proceed with the standard review?','Should we remove authorization? Run /office-hours first, or proceed with the standard review?');},
c=>{c.questions[2]!.question+=' Approve production deployment?';},
c=>{c.questions[2]!.question+=' You must run /office-hours first.';},
c=>{c.questions[2]!.question+=' Standard review is forbidden until /office-hours completes.';},
c=>{c.questions[2]!.options[1]!.label+=' if the tests pass';},
c=>{c.questions[2]!.options[0]!.label+=' and rewrite the API';},
c=>{c.questions[2]!.options[1]!.description='Proceed with standard review after completing /office-hours.';},
c=>{c.questions[2]!.options[1]!.description='No review will run.';},
c=>{c.questions[2]!.options[1]!.description='Proceed directly to Step 0 of the CEO review. Remove CI.';},
c=>{c.questions[2]!.options[0]!.description='Do not run /office-hours.';},
c=>{c.questions[2]!.options[0]!.description='Build a design doc first, then resume the review. Deploy to production.';},
c=>{c.questions[2]!.options[1]!.description='';},
c=>{c.questions[2]!.options.push({label:'Approve deployment'});},
c=>{c.questions[2]!.multiSelect=true;},
];
for(const change of changes){const c=pending();change(c);const f=frame(c,2);expect(planCountPrerequisitePick(f.routing,f.active)).toBeNull();}
});
test('AD v2 prerequisite requires the active native packet identity',()=>{
const c=pending(),f=frame(c,2);
for(const active of [{...f.active,preReview:false},{...f.active,signature:'foreign:tool:question:2'},
{...f.active,nativeQuestionIndex:0},{...f.active,promptSnippet:'Unrelated question'},
{...f.active,nativeCall:undefined},{...f.active,options:[...f.active.options].reverse()}])
expect(planCountPrerequisitePick(f.routing,active)).toBeNull();
expect(planCountPrerequisitePick({...f.active,nativeCall:undefined})).toBeNull();
for(const delta of [{answered:true},{failed:true},{sessionId:''},{toolUseId:''}]){const call={...pending(),...delta};const x=frame(call,2);expect(planCountPrerequisitePick(x.routing,x.active)).toBeNull();}
});
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
test('AD v2 prerequisite regression selects its existing mode workflow',()=>{
for(const file of ['test/ceo-prerequisite-ad-v2.test.ts','test/fixtures/ceo-prerequisite-ad-v2.json'])
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['plan-ceo-mode-routing']);
});
+228
View File
@@ -0,0 +1,228 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import captured from './fixtures/ceo-section-choice-ai.json';
import metadataCaptured from './fixtures/ceo-metadata-brief-ax.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
function call(index = 4): any {
const source = structuredClone(captured.calls[index]!);
return { sessionId: source.sessionId, toolUseId: source.toolUseId, questions: source.questions,
answered: true, failed: false, unansweredQuestionIndices: [], answeredAt: source.answeredAt,
answers: Object.fromEntries(source.questions.map((q, i) => [q.question, source.answers[i]])) };
}
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
function edit(c: any, change: (text: string) => string) {
const q = c.questions[0], answer = c.answers[q.question];
q.question = change(q.question); c.answers = { [q.question]: answer };
}
test('exact captured section choices start review; preceding actual setup does not', () => {
let started = false; const classified: boolean[] = [];
for (let i = 0; i < captured.calls.length; i++) {
const question = fp(call(i));
expect(ceoFirstReviewAUQ(question)).toBe(captured.calls[i]!.expectedFirstReview);
const phase = planCountQuestionPhase(question, started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted; classified.push(phase.preReview);
}
expect(classified).toEqual([true, true, true, true, false, false, false, false, false]);
});
test('an offered alternative and a quoted historical withdrawal retain current review identity', () => {
const alternative = call(); alternative.answers[alternative.questions[0].question] = alternative.questions[0].options[1].label;
expect(ceoFirstReviewAUQ(fp(alternative))).toBe(true);
const quoted = call(); edit(quoted, s => s + '\nHistorical quote: "This issue has been resolved."');
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
});
test.each([
['pending', (c: any) => { c.answered = false; }],
['failed', (c: any) => { c.failed = true; }],
['unanswered index', (c: any) => { c.unansweredQuestionIndices = [0]; }],
['missing session', (c: any) => { c.sessionId = ''; }],
['missing tool id', (c: any) => { c.toolUseId = ''; }],
['unoffered answer', (c: any) => { c.answers[c.questions[0].question] = 'Not offered'; }],
['missing answer', (c: any) => { c.answers = {}; }],
['mixed packet', (c: any) => { c.questions.push(structuredClone(c.questions[0])); }],
['multi-select', (c: any) => { c.questions[0].multiSelect = true; }],
['duplicate options', (c: any) => { c.questions[0].options[1] = structuredClone(c.questions[0].options[0]); }],
['missing description', (c: any) => { c.questions[0].options[1].description = ''; }],
['option identity', (c: any) => { c.questions[0].options[1].label = '1B) Other'; }],
['section mismatch', (c: any) => edit(c, s => s.replace('Section 1 Architecture.', 'Section 2 Architecture.'))],
['recommendation mismatch', (c: any) => edit(c, s => s.replace('Recommendation: A', 'Recommendation: B'))],
['missing stakes', (c: any) => edit(c, s => s.replace(/^Stakes if we pick wrong:.*$/m, ''))],
['duplicate assessment', (c: any) => edit(c, s => s + '\nELI10: A second competing assessment.')],
['quoted assessment', (c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'))],
['fenced context', (c: any) => edit(c, s => s.replace(/^(Project\/branch\/task:.*)$/m, '```\n$1\n```'))],
['Suppose assessment', (c: any) => edit(c, s => s.replace('ELI10: The plan', 'ELI10: Suppose the plan'))],
['single quoted assessment', (c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, "ELI10: '$1'"))],
['current withdrawal', (c: any) => edit(c, s => s + '\nThis issue is withdrawn.')],
['completed withdrawal', (c: any) => edit(c, s => s + '\nWe have withdrawn this finding.')],
['conditional assessment', (c: any) => edit(c, s => s.replace('ELI10: The plan', 'ELI10: If the plan'))],
['withdrawn issue', (c: any) => edit(c, s => s + '\nWe withdraw this finding.')],
['resolved issue', (c: any) => edit(c, s => s + '\nThis issue has been resolved.')],
['administrative report', (c: any) => edit(c, s => s.replace(/^.*\n/, '1A — Should the completed review report be saved?\n'))],
['setup header', (c: any) => { c.questions[0].header = 'Setup'; }],
['borrowed qid', (c: any) => edit(c, s => s + '\n<gstack-qid:plan-ceo-review-example>')],
])('rejects %s despite numbered review prose', (_name, mutate) => {
const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
});
test('fingerprints cannot borrow another native call or its options', () => {
const original = fp(call());
expect(ceoFirstReviewAUQ({ ...original, signature: 'other:tool' })).toBe(false);
expect(ceoFirstReviewAUQ({ ...original, nativeCall: undefined })).toBe(false);
expect(ceoFirstReviewAUQ({ ...original, options: original.options.slice(1) })).toBe(false);
});
test('the regression inputs belong to the existing paid CEO case', () => {
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/ceo-section-choice-ai.test.ts');
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/fixtures/ceo-section-choice-ai.json');
});
test('coherent finished-note destination is administrative, despite matching section and choice', () => {
const c = call(), q = c.questions[0];
q.header = 'Destination';
q.question = '1A — Which storage location should hold these notes?\nProject/branch/task: main, Stripe payment webhook plan, Section 1 Architecture.\nELI10: The review is finished; these notes can be saved in either folder for convenience.\nStakes if we pick wrong: People may have to look in a second folder.\nRecommendation: A because the existing folder is easier to find.';
q.options = [{label:'A) Save beside the plan',description:'Keeps the finished notes together.'},{label:'B) Save in another folder',description:'Keeps finished notes separate.'}];
c.answers = {[q.question]: q.options[0].label};
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
});
test('conditional stakes remain valid when the assessment asserts the current gap', () => {
const c = call(); edit(c, s => s.replace('Stakes if we pick wrong:', 'Stakes if we pick wrong: Suppose there were an issue.'));
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
});
test('negated gaps and administrative missing fields cannot borrow review identity', () => {
const negated = call(); edit(negated, s => s.replace(/^ELI10: .+$/m, 'ELI10: The transaction order is not unspecified. The plan guarantees commit before email.'));
expect(ceoFirstReviewAUQ(fp(negated))).toBe(false);
const admin = call(), q = admin.questions[0];
q.header = 'Destination';
q.question = '1A — Which storage location should hold these notes?\nProject/branch/task: main, Stripe payment webhook plan, Section 1 Architecture.\nELI10: These finished notes have a missing storage location.\nStakes if we pick wrong: People may look in the wrong folder.\nRecommendation: A because a notes folder is easy to find.';
q.options = [{label:'A) Add a notes folder',description:'Save the finished notes together.'},{label:'B) Use the existing folder',description:'No new folder.'}];
admin.answers = {[q.question]: q.options[0].label};
expect(ceoFirstReviewAUQ(fp(admin))).toBe(false);
});
function metadataCall(): any {
const c = structuredClone(metadataCaptured.call);
return { ...c, answered: true, failed: false, unansweredQuestionIndices: [],
answers: { [c.questions[0]!.question]: metadataCaptured.answer } };
}
test('AX ordinary D-number question keeps its exact completed review identity', () => {
const c = metadataCall(), before = JSON.stringify(c);
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ)).toMatchObject({ preReview: false, reviewStarted: true });
expect(JSON.stringify(c)).toBe(before);
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/fixtures/ceo-metadata-brief-ax.json');
});
test('decision counter, review name and an alternative selection do not dictate the finding', () => {
const c = metadataCall(); edit(c, s => s.replace(/^D5 /, 'D17 ').replace('Section 2 (Error & Rescue Map)', 'Section 3 (Failure Handling)'));
c.answers[c.questions[0].question] = c.questions[0].options[1].label;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
edit(c, s => s + '\nHistorical quote: "This finding is withdrawn."');
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
});
test('metadata cannot replace native completion, current context or a real defect', () => {
for (const mutate of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.answeredAt = 'not a date'; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.answers = { other: metadataCaptured.answer }; },
(c: any) => { c.questions[0].header = 'Setup'; },
(c: any) => { c.questions[0].header = 'Section 9'; },
(c: any) => edit(c, s => s.replace('Section 2 (Error & Rescue Map)', 'Section 2 (Error & Rescue Map), Section 3 (Security)')),
(c: any) => edit(c, s => s.replace('of the CEO review', 'of an earlier CEO review')),
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task: (.*)$/m, 'Project/branch/task: If approved, $1')),
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task:.*\n/m, '')),
(c: any) => edit(c, s => s.replace(/^ELI10: (.*)$/m, 'ELI10: "$1"')),
(c: any) => edit(c, s => s.replace(/^ELI10: .+$/m, 'ELI10: The handler has no current defect and needs no amendment.')),
(c: any) => edit(c, s => s.replace(/^ELI10: .+$/m, 'ELI10: The handler commits before mail and already rescues every required error.')),
(c: any) => edit(c, s => s.replace('ELI10: After', 'ELI10: Hypothetical example: after')),
(c: any) => edit(c, s => s + '\nThis finding is withdrawn.'),
(c: any) => edit(c, s => s + '; This finding is `no longer current`.'),
(c: any) => edit(c, s => s + '\nThis finding is unproven.'),
]) {
const c = metadataCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
expect(ceoFirstReviewAUQ({ ...fp(metadataCall()), signature: 'foreign:call' })).toBe(false);
});
test('a missing or withdrawn offered remedy cannot borrow metadata or an old assessment', () => {
for (const mutate of [
(c: any) => { c.questions[0].options = [{ label: 'A: Keep the current handler', description: 'No code change.' }, { label: 'B: Save the review notes', description: 'Archive the current report.' }]; c.answers = { [c.questions[0].question]: c.questions[0].options[0].label }; },
(c: any) => { c.questions[0].options[0].description += '; This option is `withdrawn`.'; c.questions[0].options[2].description += '\nThis option is withdrawn.'; },
(c: any) => { c.questions[0].options.forEach((o: any) => { o.description = 'Hypothetical example. ' + o.description; }); },
]) {
const c = metadataCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
});
function metadataRetryCall(): any {
const c = structuredClone(metadataCaptured.retry.call);
return { ...c, answered: true, failed: false, unansweredQuestionIndices: [],
answers: { [c.questions[0]!.question]: metadataCaptured.retry.answer } };
}
test('the separately failed AX retry binds its Issue annotation, bare choices and named plan', () => {
const c = metadataRetryCall(), before = JSON.stringify(c);
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
.toMatchObject({ preReview: false, reviewStarted: true });
expect(JSON.stringify(c)).toBe(before);
c.answers[c.questions[0].question] = c.questions[0].options[2].label;
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
});
test('a reviewed filename and section identity can be consistently renamed', () => {
const c = metadataRetryCall();
edit(c, s => s.replace(/^D4 /, 'D12 ').replace(/Issue 2\.1/, 'Issue 8.3')
.replace('Section 2 (Error & Rescue Map)', 'Section 8 (Notification Handling)')
.replace(/PLAN\.md/g, 'plans/checkout-flow.md'));
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
edit(c, s => s + '\nHistorical quote: "This issue is withdrawn."');
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
edit(c, s => s.replace(/\s*<gstack-qid:[^>]+>/, ''));
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
});
test('retry metadata cannot borrow a foreign section, plan, source or incomplete native call', () => {
for (const [index, mutate] of [
(c: any) => { c.answered = false; },
(c: any) => { c.failed = true; },
(c: any) => { c.answeredAt = 'unknown'; },
(c: any) => { c.unansweredQuestionIndices = [0]; },
(c: any) => { c.questions[0].header = 'Issue 8.1'; },
(c: any) => edit(c, s => s.replace('Issue 2.1', 'Issue 3.1')),
(c: any) => edit(c, s => s.replace('Section 2 (Error & Rescue Map)', 'Section 3 (Security)')),
(c: any) => edit(c, s => s.replace('CEO review of PLAN.md,', 'CEO review of DIFFERENT.md,')),
(c: any) => edit(c, s => s.replace("PLAN.md says 'no error handling on the email leg'", "OTHER.md says 'no error handling on the email leg'")),
(c: any) => edit(c, s => s.replace('CEO review of PLAN.md,', 'Historical CEO review of PLAN.md,')),
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task: (.+)$/m, 'Project/branch/task: If approved, $1')),
(c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"')),
(c: any) => edit(c, s => s.replace('ELI10: The handler', 'ELI10: Source excerpt: the handler')),
(c: any) => edit(c, s => s.replace('PLAN.md says', 'If approved, PLAN.md says')),
(c: any) => edit(c, s => s.replace('PLAN.md says', 'PLAN.md does not say')),
(c: any) => edit(c, s => s.replace('plan-ceo-review-mail-rescue', 'plan-ceo-review-setup')),
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-ceo-review-other>'),
].entries()) {
const c = metadataRetryCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c)), `retry mutation ${index}`).toBe(false);
}
});
test('current withdrawal and a withdrawn offered amendment override the retry brief', () => {
for (const change of [
(s: string) => s + '\nThis issue is withdrawn.',
(s: string) => s + '; This finding is `no longer current`.',
(s: string) => s + '\nIssue 2.1 is withdrawn.',
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The handler has no current defect and needs no amendment.'),
]) {
const c = metadataRetryCall(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
}
const c = metadataRetryCall(); c.questions[0].options[0].description += '; This option is `withdrawn`.';
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
});
+32
View File
@@ -0,0 +1,32 @@
import {describe,test,expect} from 'bun:test';
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
import fixture from './fixtures/ceo-section-declarative-ar.json';
const calls=()=>structuredClone(fixture.calls) as NativePlanQuestionCall[],first=()=>calls()[2]!;
const fp=(c=first())=>nativePlanCallFingerprint(c,0,true),classify=(c=first())=>ceoFirstReviewAUQ(fp(c));
const mutate=(fn:(c:NativePlanQuestionCall)=>void)=>{const c=first();fn(c);return c;};
const text=(fn:(s:string)=>string)=>mutate(c=>{const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};});
const allOptions=(fn:(label:string,description:string)=>{label:string;description:string})=>mutate(c=>{const q=c.questions[0]!,selected=q.options.findIndex(o=>o.label===c.answers![q.question]);q.options=q.options.map(o=>fn(o.label,o.description??''));c.answers={[q.question]:q.options[selected]!.label};});
describe('AR completed declarative Section finding',()=>{
test('exact public calls enter review after genuine setup',()=>{let started=false;const phases=calls().map(c=>{const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ);started=p.reviewStarted;return p.preReview;});expect(phases).toEqual([true,true,false,false,false]);expect(classify(calls()[0])).toBe(false);expect(classify(calls()[1])).toBe(false);expect(classify(calls()[2])).toBe(true);expect(classify(calls()[3])).toBe(true);});
test('comma and question punctuation are presentation',()=>{for(const c of [first(),text(s=>s.replace('Section 6, finding','Section 6 finding')),text(s=>s.replace('receipt is truthy\n','receipt is truthy?\n')),text(s=>s.replace('Section 6, finding','Section 6 finding').replace('receipt is truthy\n','receipt is truthy?\n'))])expect(classify(c)).toBe(true);expect(classify(text(s=>s.replace('D3 — Section 6, finding 1:','D8 — Section 2, finding 3:')))).toBe(true);});
test('native successful completion and exact ownership stay required',()=>{
for(const fn of [(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.answeredAt;},(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers!['foreign']='foreign';},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered'};},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;}])expect(classify(mutate(fn))).toBe(false);
for(const f of [{...fp(),signature:'foreign:tool'},{...fp(),nativeQuestionIndex:1},{...fp(),options:fp().options.toReversed()}])expect(ceoFirstReviewAUQ(f)).toBe(false);
});
test('malformed identities and setup headers stay closed',()=>{for(const [a,b] of [['D3 —','D0 —'],['D3 —','D03 —'],['Section 6,','Section 06,'],['finding 1:','finding 0:'],['finding 1:','finding 1.2:'],['Section 6,','Section 6,,']])expect(classify(text(s=>s.replace(a,b)))).toBe(false);for(const h of ['Section 7','Finding 9','Routing','Approach'])expect(classify(mutate(c=>{c.questions[0]!.header=h;}))).toBe(false);});
test('source, conditional and duplicate assessment metadata stay closed',()=>{for(const field of ['Project/branch/task: ','ELI10: '])for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, '])expect(classify(text(s=>s.replace(field,field+p)))).toBe(false);for(const prefix of ['Source:\n','Earlier review assessment:\n','Project/branch/task: duplicate\n','ELI10: duplicate\n'])expect(classify(text(s=>s.replace('ELI10:',prefix+'ELI10:')))).toBe(false);expect(classify(text(s=>s.replace(/^Project\/branch\/task:.*\n/m,'')))).toBe(false);expect(classify(text(s=>'```\n'+s+'\n```'))).toBe(false);});
test('withdrawn assessment or source-only options cannot establish a current decision',()=>{expect(classify(text(s=>s+'\nThis finding is withdrawn.'))).toBe(false);expect(classify(text(s=>s+'\nCorrection: this finding is "withdrawn".'))).toBe(false);for(const p of ['Source excerpt: ','If approved, '])expect(classify(allOptions((label,description)=>({label:label.replace(/^([A-Z]\) )/,'$1'+p),description:p+description})))).toBe(false);expect(classify(allOptions((label,description)=>({label,description:description+' This amendment is withdrawn.'})))).toBe(false);expect(classify(allOptions((label)=>({label:label.replace(/^([A-Z]\) ).*/,'$1Keep current assertion'),description:'Leave the current assertion unchanged.'})))).toBe(false);});
test('superseded or conditional findings and offered actions are not current',()=>{
for(const status of ['superseded','"superseded"','no longer current','"no longer current"']){
expect(classify(text(s=>s+'\nThis finding is '+status+'.'))).toBe(false);
expect(classify(allOptions((label,description)=>({label,description:description+' This amendment is '+status+'.'})))).toBe(false);
}
for(const prefix of ['Assuming approval, ','Provided approval, ']){
expect(classify(text(s=>s.replace('ELI10: ','ELI10: '+prefix)))).toBe(false);
expect(classify(allOptions((label,description)=>({label,description:prefix+description})))).toBe(false);
}
for(const history of ['> This finding is superseded.','Archived note: "This finding is superseded."','Archived note: "This finding is no longer current."','~~~\nThis finding is superseded.\n~~~'])expect(classify(text(s=>s+'\n'+history))).toBe(true);
});
test('quoted archive and selected opposed option remain valid',()=>{expect(classify(text(s=>s+'\nArchived note: "This finding is withdrawn."'))).toBe(true);expect(classify(mutate(c=>{const q=c.questions[0]!;c.answers={[q.question]:q.options[2]!.label};}))).toBe(true);});
});
+478
View File
@@ -0,0 +1,478 @@
import { describe, expect, test } from 'bun:test';
import {
CACHE_READ_WRITE_SKETCH,
CEO_SECTION_CACHE_PLAN,
hasStaleFillRaceFinding,
} from './helpers/ceo-section-loading-fixture';
describe('future-reader vocabulary in the actual AA finding', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-aa-report.md'), 'utf8');
const paragraph = report.slice(report.indexOf('**S4-1 (CRITICAL'), report.indexOf('\n\nNo UI scope.', report.indexOf('**S4-1 (CRITICAL')));
test('recognizes the exact complete report and its same-paragraph post-write consequence', () => {
expect(paragraph).toContain('A read in-flight when a write');
expect(paragraph).toContain('stale snapshot after `cache.delete` fires');
expect(paragraph).toContain('future callers with stale data');
expect(hasStaleFillRaceFinding(paragraph)).toBe(true);
expect(hasStaleFillRaceFinding(report)).toBe(true);
expect(hasStaleFillRaceFinding(paragraph.replace('future callers', 'subsequent callers'))).toBe(true);
});
test.each(['future reads', 'future requests', 'future callers'])('recognizes a later consumer: %s', reader => {
expect(hasStaleFillRaceFinding(`An in-flight read inserts a stale snapshot after write invalidation, leaving ${reader} with stale data.`)).toBe(true);
});
test.each([
'An in-flight read inserts a stale snapshot after write invalidation. Future work documents the cache.',
'An in-flight read inserts a fresh snapshot after write invalidation, leaving future callers with fresh data.',
'An in-flight read inserts a stale snapshot before write invalidation, leaving future callers with stale data.',
'A completed read inserts a stale snapshot after write invalidation, leaving future callers with stale data.',
'An in-flight read returns a stale snapshot after write invalidation to its original pending caller.',
'An in-flight read inserts a stale snapshot after write invalidation. Future callers seeing stale data is allowed behavior.',
'An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data. This is the accepted consistency model.',
'An in-flight read cannot refill stale data after write invalidation. Future callers observe committed data.',
'An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data. This is not a bug; no guard is required.',
'> An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data.',
'```text\nAn in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data.\n```',
'An in-flight read inserts a stale snapshot after write invalidation.\n\n## A different section\nFuture callers need documentation.',
'* An in-flight read inserts a stale snapshot after write invalidation.\n* Future callers need documentation.',
])('retains ordering, stale-value, source and dismissal boundaries: %s', text => {
expect(hasStaleFillRaceFinding(text)).toBe(false);
});
test('the original-caller exception cannot permit the same stale value for future callers', () => {
expect(hasStaleFillRaceFinding('An in-flight read refills stale data after write invalidation, so new reads see old data. The original pending caller may receive an old snapshot and future callers observe it; this is permitted. Guard cache fills with a generation token.')).toBe(false);
expect(hasStaleFillRaceFinding('The original pending caller may receive an old snapshot; that return is permitted. However, an in-flight read refills stale data after write invalidation, so future callers violate the contract. Guard cache fills with a generation token.')).toBe(true);
});
});
describe('restore vocabulary in the actual Y finding', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-y-report.md'), 'utf8');
const amendment = report.slice(report.indexOf('**AMENDMENT (Finding 1'), report.indexOf('```javascript')).trim();
test('recognizes the delivered report and its explicit original-invariant failure', () => {
expect(amendment).toContain('original pseudocode did not satisfy');
expect(amendment).toContain('in-flight read from restoring a stale cache entry');
expect(hasStaleFillRaceFinding(amendment)).toBe(true);
expect(hasStaleFillRaceFinding(report)).toBe(true);
});
test.each(['can restore', 'restores', 'restored', 'is restoring'])('recognizes the cache-fill verb %s', verb => {
expect(hasStaleFillRaceFinding(`An in-flight read ${verb} stale data after write invalidation. A subsequent read sees the old value, violating the contract.`)).toBe(true);
});
test.each(['cannot restore', "can't restore", 'never restores', 'does not restore', "doesn't restore", 'will not restore', "won't restore", 'did not restore', "didn't restore", 'is not restoring', "isn't restoring", 'was not restoring', 'has not restored', "hasn't restored", 'had not restored'])('rejects a current prevention assertion: %s', denied => {
expect(hasStaleFillRaceFinding(`An in-flight read ${denied} stale data after write invalidation. A subsequent read observes the committed value.`)).toBe(false);
});
test.each([
'An in-flight read restores stale data after write invalidation. This is not a defect; no guard is required.',
'An in-flight read restores stale data after write invalidation. Subsequent stale reads are permitted by the contract.',
'An in-flight read restores stale data after write invalidation. This is the accepted consistency model.',
'An in-flight read restores the committed new value after write invalidation. A subsequent read observes it.',
'The original pending caller receives an old snapshot after the write; that return is permitted.',
'If deletion throws after a write, a restore operation leaves stale cache data. Log and bypass the adapter.',
'> An in-flight read restores stale data after write invalidation; a subsequent read violates the contract.',
'```text\nAn in-flight read restores stale data after write invalidation; a subsequent read violates the contract.\n```',
'* An in-flight read restores stale data after write invalidation.\n* Telemetry has a bug.',
])('preserves dismissal, source and separate-finding boundaries: %s', text => {
expect(hasStaleFillRaceFinding(text)).toBe(false);
});
});
describe('pre-write snapshot vocabulary in the actual U finding', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-u-report.md'), 'utf8');
const paragraph = report.slice(report.indexOf('After T4 the cache correctly reflects'), report.indexOf('**Recommended fix (auto-decided):**')).trim();
test('recognizes the exact delivered report and its complete asserted paragraph independently of the bad remedy', () => {
expect(paragraph).toContain('re-populates the cache with the pre-write snapshot');
expect(paragraph).toContain('This violates the invariant:');
expect(hasStaleFillRaceFinding(paragraph)).toBe(true);
expect(hasStaleFillRaceFinding(report)).toBe(true);
});
test.each(['pre-write snapshot', 'pre write snapshot', 'pre-write value', 'pre-write data', 'pre-write version'])('recognizes an old snapshot synonym: %s', value => {
expect(hasStaleFillRaceFinding(`An in-flight read re-populates the cache with the ${value} after write invalidation. A new reader sees it, violating the contract.`)).toBe(true);
});
test.each([
'A pending read returns the pre-write snapshot to its original caller; that return is permitted.',
'An in-flight read re-populates the cache with the post-write snapshot after invalidation.',
'The pre-write snapshot expires after 30 seconds. The LRU byte cap is adequate.',
'If invalidation throws after a write, the cache retains the pre-write snapshot. Log the failure and bypass the cache.',
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. This is allowed behavior for subsequent reads.',
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. This is the accepted consistency model.',
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation, so a new reader receives that version. This is the accepted consistency model.',
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. It is not a bug; no guard is required.',
'An in-flight read cannot re-populate the cache with the pre-write snapshot after invalidation. No race remains.',
'> An in-flight read re-populates the pre-write snapshot after write invalidation; a new read sees it, violating the contract.',
'```text\nAn in-flight read re-populates the pre-write snapshot after write invalidation; a new read sees it, violating the contract.\n```',
'* An in-flight read re-populates the pre-write snapshot after invalidation.\n* Telemetry retry handling has a bug.',
])('retains original-caller, freshness, dismissal and source boundaries: %s', value => {
expect(hasStaleFillRaceFinding(value)).toBe(false);
});
test('permission for the original caller still cannot excuse a later-reader violation', () => {
expect(hasStaleFillRaceFinding('The original pending caller may receive the pre-write snapshot; that return is permitted. However, an in-flight read refills the cache with the pre-write snapshot after write invalidation, so a new reader violates the contract. Guard cache fills with a generation token.')).toBe(true);
});
});
describe('CEO section-loading cache fixture', () => {
test('the exact proposed wrapper retains a reproducible stale-fill race', async () => {
let releaseRead!: (value: string) => void;
let stored = 'old';
const cache = new Map<string, string>();
const repository = {
read: () => new Promise<string>((resolve) => { releaseRead = resolve; }),
write: async (_key: string, value: string) => { stored = value; return value; },
};
// Execute the same sketch the live reviewer receives, not a second model
// of its ordering. Holding the old read exposes the intended interleaving.
const { readProfile, writeProfile } = new Function('cache', 'repository',
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
const pending = readProfile('tenant:profile');
await writeProfile('tenant:profile', 'new');
releaseRead('old');
const earlierResult = await pending;
expect(stored).toBe('new');
// Returning the earlier snapshot to the already-pending caller is
// explicitly permitted. Reusing it for this new reader is the defect.
expect(earlierResult).toBe('old');
expect(CEO_SECTION_CACHE_PLAN).toContain('Every read begun after that write completes must');
expect(await readProfile('tenant:profile')).toBe('old');
expect(CEO_SECTION_CACHE_PLAN).toContain(CACHE_READ_WRITE_SKETCH);
expect(hasStaleFillRaceFinding(CEO_SECTION_CACHE_PLAN)).toBe(false);
});
test.each([
// Actual finding in the unchanged fixture's successful 38 KB live report.
'**Missing: What happens to in-flight requests during invalidation?** If a write invalidates a key and 10 requests are simultaneously loading it (cache miss, in-flight DB fetch), all 10 will cache the same value after the invalidation. The invalidated key may get re-populated with a stale value if any of those fetches started before the write. No mention of this race.',
'P1: An in-flight read can repopulate stale data after a committed write invalidates the key. Guard fills with a generation token.',
'| Cache fill race | An older value fetched before the write is inserted after eviction, so the next read is stale. | Add a per-key epoch. |',
'**Invalidation race:** the pending fetch stores an outdated snapshot after cache.delete. Serialize the fill with mutation.',
])('recognizes the actual ordering defect: %s', (report) => {
expect(hasStaleFillRaceFinding(report)).toBe(true);
});
test.each([
'The full review is complete. No issues found.',
'A stale value expires after 30 seconds. The LRU byte cap is adequate.',
'If invalidation throws after a write, the cache retains stale data. Log the failure and bypass the cache.',
'Read and write concurrency is covered. No stale data can be returned.',
'| Reads | Coalesced concurrent misses |\n| Writes | Invalidation failure leaves stale data |',
'```javascript\n// An in-flight read can cache stale data after invalidation.\n```',
'> An in-flight read can cache stale data after invalidation.',
])('rejects completion, unrelated text, and quoted source: %s', (report) => {
expect(hasStaleFillRaceFinding(report)).toBe(false);
});
});
const CAPTURED_ACCEPTED_RACE_REPORT = `**Shadow paths:**
1. Nil key: Programming error caught by auth/key-validation before wrapper.
2. Empty key: Same upstream validation gate.
3. Upstream error: Single-flight releases all waiters with the error. Cache
not populated. Next request retries DB. Correct.
4. Concurrent write during read in-flight: The plan documents this explicitly.
The stale read is an accepted invariant, bounded by 30s TTL.
**Async ordering critical race:**
\`\`\`
1. Request A: cache.get(key) miss enters single-flight
2. Request B: cache.get(key) miss joins single-flight (awaiting)
3. fn: repository.read(key) suspend (await)
4. Write commits cache.delete(key) [nothing to delete key not set yet]
5. repository.read(key) returns OLD snapshot (pre-write)
6. cache.set(key, OLD_VALUE) stale value in cache for up to 30s
7. Requests A and B both return OLD_VALUE accepted by plan
\`\`\`
This is the one documented asymmetry. It is not a gap it is a named invariant.
The TTL bounds the stale window to 30 seconds.
`;
describe('CEO concurrency finding requires a violation, not an accepted trace', () => {
test('rejects the captured accepted-invariant report that passed the old keyword oracle', () => {
expect(hasStaleFillRaceFinding(CAPTURED_ACCEPTED_RACE_REPORT)).toBe(false);
});
test.each([
'An in-flight fetch can refill the cache with old data after a write invalidates it. The next read sees that stale snapshot, violating the post-write contract.',
'The pending read stores an older value after invalidation.\n\nGuard cache fills with a version check so a later request cannot observe pre-write state.',
'Returning the old snapshot to the pending caller is permitted. But a late cache.set after concurrent write invalidation exposes stale data to a new reader. Serialize mutation and cache fills.',
'Returning an old snapshot to the original pending caller is an accepted invariant. But an in-flight read can repopulate stale cache data after write invalidation, so a new reader violates the post-write contract. Guard cache fills with a generation token.',
'The original caller may receive the old snapshot; that return is permitted. However, a pending fetch refills stale data after write invalidation, breaking consistency for a later reader. Skip the cache fill when its version changed.',
'An in-flight read can repopulate stale data after write invalidation, so a new reader gets the old value. This is not permitted by the contract. Guard cache fills with a version check.',
'| Late cache fill | A concurrent read repopulates an outdated result after eviction. | Reject the fill when its generation token changed. |',
])('accepts the later-reader consequence or a concrete ordering remedy: %s', report => {
expect(hasStaleFillRaceFinding(report)).toBe(true);
});
test.each([
'An in-flight read repopulates stale data after write invalidation. This is an accepted invariant bounded by the TTL.',
'An in-flight read repopulates stale data after write invalidation. This is permitted by the contract; a later read may be stale for 30 seconds.',
'An in-flight read refills stale data after write invalidation. This is allowed behavior for the next read because the TTL bounds it.',
'The original caller and the new reader may both observe the old snapshot as an accepted invariant. A pending read refills stale data after write invalidation; no guard is required.',
'An in-flight read repopulates stale data after write invalidation.\n\nIt is not a gap. No change is needed.',
'No race: a pending read cannot repopulate stale cache data after write invalidation; the existing version check rejects it.',
'An in-flight read repopulates stale data after write invalidation, but does not violate the contract. No guard is required.',
'1. Cache population after a miss is safe.\n2. Concurrent writes can return an older snapshot to their original pending reader.\n3. Guard unrelated network retries.',
'An in-flight read stores stale data after write invalidation.\n\n**Finding S9:** Guard telemetry delivery with a version token.',
'An in-flight read stores stale data after write invalidation.\n\nGuard unrelated telemetry delivery with a version token.',
'1. An in-flight read repopulates stale data after write invalidation.\n2. Telemetry retry handling has a bug.',
'An in-flight read repopulates stale data after write invalidation.\n\n#2 — Unrelated telemetry delivery bug',
'* An in-flight read repopulates stale data after write invalidation.\n* Telemetry retry handling has a bug.',
'> P1: An in-flight read refills stale data after invalidation; a new read gets the old value.',
'```text\nP1: An in-flight read refills stale data after invalidation; a new read gets the old value.\n```',
])('rejects dismissals, negations, unrelated findings and quoted examples: %s', report => {
expect(hasStaleFillRaceFinding(report)).toBe(false);
});
});
describe('proposed cache-fill prevention remains an unresolved finding', () => {
test('an imperative remedy describes the behavior it must prevent', () => {
expect(hasStaleFillRaceFinding('An in-flight read repopulates stale data after write invalidation. Guard cache fills so pending reads cannot repopulate stale values after invalidation.')).toBe(true);
});
test('an existing guard remains a dismissal, not a proposed fix', () => {
expect(hasStaleFillRaceFinding('An in-flight read cannot repopulate stale data after write invalidation because the existing guard rejects that fill. No race remains.')).toBe(false);
});
test('an imperative does not erase a separate explicit dismissal', () => {
expect(hasStaleFillRaceFinding('An in-flight read repopulates stale data after write invalidation. Guard cache fills so pending reads cannot repopulate stale values after invalidation. This is not a bug; no fix is needed.')).toBe(false);
});
});
// The native report separates an asserted finding, its ordered trace and its
// explicit contract violation. Detection does not certify the offered fix.
describe('structured native stale-fill finding', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-loading-l-report.md'), 'utf8');
const finding = report.slice(report.indexOf('**CRITICAL FINDING — Write-then-read stale-set race**'), report.indexOf('**Required fix:**'));
test('retains the exact positive later-reader finding even though the proposed mitigation is wrong', () => {
expect(hasStaleFillRaceFinding(report)).toBe(true);
expect(hasStaleFillRaceFinding(finding)).toBe(true);
});
test.each([
['standalone trace', finding.slice(finding.indexOf('```'), finding.lastIndexOf('```') + 3)],
['quoted finding', finding.split('\n').map((line: string) => '> ' + line).join('\n')],
['source example', 'Example of report format:\n' + finding],
['outer fenced source', '````text\n' + finding + '\n````'],
['explicit accepted trace', finding.replace('This violates the stated invariant:', 'This is not a gap. The following behavior is accepted:')],
['negated violation', finding.replace('This violates the stated invariant:', 'This does not violate the stated invariant:')],
['separate dismissal', finding + '\nThis is not a defect; no fix is required.\n'],
['no new reader', finding.replace(/T3: readProfile[\s\S]*?\n```/, '```')],
['reverse ordering', finding.replace('cache.delete(key)', 'cache.get(key)')],
['unrelated heading', finding.replace('CRITICAL FINDING', 'EXAMPLE')],
...['~~~', '````'].map(fence => ['nontriple fenced violation', finding.replace(/This violates[\s\S]*$/, text => fence + 'text\n' + text + '\n' + fence)]),
['unclosed fenced violation', finding.replace('This violates', '```text\nThis violates')],
['later named finding', finding.replace('This violates', '**CRITICAL FINDING — unrelated documentation defect**\nThis violates')],
['later heading', finding.replace('This violates', '## Unrelated finding\nThis violates')],
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
});
// Actual Q report amended the original contract to accept later stale reads.
// The oracle must not count that permission paragraph as an unresolved defect.
describe('accepted consistency model is not a stale-fill finding', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-loading-q-report.md'), 'utf8');
const accepted = report.slice(report.indexOf('- **AMENDED (stale-fill race):**'), report.indexOf('\n\n', report.indexOf('- **AMENDED (stale-fill race):**')));
test('rejects the exact amended consistency paragraph', () => {
expect(accepted).toMatch(/accepted consistency\s+model/);
expect(hasStaleFillRaceFinding(accepted)).toBe(false);
});
test('rejects the complete Q report that accepts the late-fill race', () => {
expect(hasStaleFillRaceFinding(report)).toBe(false);
});
test.each([
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is the accepted consistency model; TTL expiry is the consistency deadline.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This remains the documented consistency contract.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is an intentional consistency policy.',
])('rejects a declared consistency allowance: %s', text => {
expect(hasStaleFillRaceFinding(text)).toBe(false);
});
test.each([
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. Every later read must observe the committed version; this is the accepted consistency model. The stale refill violates that contract.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. Every later read must observe the committed version. This is the accepted consistency model. The stale refill violates that contract.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value, violating the accepted consistency model.',
'An in-flight read refills stale data after write invalidation. The accepted consistency model requires a subsequent read to observe the committed version; this violates that contract.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is not the accepted consistency model.',
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is not an accepted consistency model.',
])('retains an unresolved later-reader violation without demanding a correct remedy: %s', text => {
expect(hasStaleFillRaceFinding(text)).toBe(true);
});
});
describe('CEO R report requirement weakening remains rejected', () => {
test('a race trace followed by acceptance and a weaker guarantee is not an unresolved defect', async () => {
// Exact delivered R retry report; the existing oracle already rejects it.
// This pins the policy failure without changing the paid fixture or oracle.
const report = await Bun.file(new URL('./fixtures/ceo-section-r-rejected-report.md', import.meta.url)).text();
expect(report).toContain('Every read begun after that write completes must');
expect(report).toContain('this is an accepted design choice, not a quality gap');
expect(report).toContain('Verify: new reads see stale value until TTL');
expect(hasStaleFillRaceFinding(report)).toBe(false);
});
});
// Exact S output names the original-contract defect but proposes an ineffective
// guard. Detection and remedy correctness remain separate assertions.
describe('S native same-finding ordered trace', () => {
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-s-trace-report.md'), 'utf8');
const finding = report.slice(report.indexOf('### Critical Finding: Stale Re-insertion After Write Invalidation'), report.indexOf('### State Machine: Cache Entry'));
test('recognizes the exact delivered S finding without certifying its remedy', () => {
expect(hasStaleFillRaceFinding(report)).toBe(true);
expect(hasStaleFillRaceFinding(finding)).toBe(true);
expect(report).toContain('if (cache.get(key) === undefined)');
});
test('the same ordered evidence tolerates whitespace and a consistent key name', () => {
expect(hasStaleFillRaceFinding(finding.replaceAll('(key', '(accountKey').replaceAll(' t', ' t'))).toBe(true);
expect(hasStaleFillRaceFinding(finding.replaceAll('(key', '($key'))).toBe(true);
});
test.each([
['optional single-flight label', finding.replace('single-flight → ', '')],
['old/stale value vocabulary', finding.replace('OLD snapshot', 'stale value').replace('OLD_VALUE', 'STALE_VALUE').replace('stale value re-inserted', 'old snapshot refilled').replace('next readProfile', 'subsequent readProfile')],
['prose payload vocabulary', finding.replace('OLD_VALUE', 'old value').replace('returns stale value', 'returns old snapshot')],
['ASCII arrows and compact spacing', finding.replaceAll(' → ', '->').replaceAll(' ← ', '<-')],
['trace keyword case', finding.replace(/t[1-6]:[^\n]*/g, (event: string) => event.toLowerCase())],
['call whitespace and optional suspension annotation', finding.replaceAll('(key)', '( key )').replace('(key, OLD_VALUE)', '( key , OLD_VALUE )').replaceAll(' (suspends)', '')],
])('accepts equivalent %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(true));
test.each([
['only the trace', finding.slice(finding.indexOf('```'), finding.indexOf('```', finding.indexOf('```') + 3) + 3)],
['quoted whole finding', finding.split('\n').map((line: string) => '> ' + line).join('\n')],
['indented whole finding', finding.split('\n').map((line: string) => ' ' + line).join('\n')],
['source format preface', 'Example of report format:\n' + finding],
['outer fenced source', '````text\n' + finding + '\n````'],
['missing independent violation', finding.replace('The proposed wrapper violates this invariant.', '')],
['negated independent violation', finding.replace('The proposed wrapper violates this invariant.', 'The proposed wrapper does not violate this invariant.')],
['unrelated heading', finding.replace('### Critical Finding:', '### Example:')],
['separate named finding', finding.replace('Race sequence', '### Another finding\nRace sequence')],
['separate bold finding', finding.replace('Race sequence', '**HIGH FINDING — unrelated issue**\nRace sequence')],
['missing write completion', finding.replace('DB write completes', 'DB write remains pending')],
['no invalidation', finding.replace('cache.delete(key) → writeProfile returns', 'cache.get(key) → writeProfile returns')],
['missing late old fill', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(key, NEW_VALUE)')],
['different filled key', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(otherKey, OLD_VALUE)')],
['case-distinct filled key', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(KEY, OLD_VALUE)')],
['different later key', finding.replace('next readProfile(key)', 'next readProfile(otherKey)')],
['only original pending reader', finding.replace('next readProfile(key)', 'original pending readProfile(key)')],
['later reader misses', finding.replace('cache HIT → returns stale value', 'cache MISS → returns committed value')],
['nonviolating trace', finding.replace('← INVARIANT VIOLATED', '← INVARIANT PRESERVED')],
['reverse order labels', finding.replace('t3:', 't4:').replace('t4: DB read', 't3: DB read')],
['missing event', finding.replace(/^.*t4:.*\n/m, '')],
['unclosed trace', finding.replace('```\n\nThe plan says', '\nThe plan says')],
['source-code trace fence', finding.replace('```\n t1:', '```javascript\n t1:')],
['split traces', finding.replace(' t4:', '```\n\n```\n t4:')],
['accepted stale trace', finding + '\nThis stale-read behavior is accepted; no guard is required.\n'],
['allowed new-reader consequence', finding + '\nA subsequent stale read is permitted by the amended contract.\n'],
['explicit defect dismissal', finding + '\nThis is not a defect; no fix is required.\n'],
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
});
// Actual v2 SDK review identifies the missing fill/write coordination directly.
// The full delivered report is retained in run evidence; this is its exact finding.
describe('explicit uncoordinated cache-fill freshness violation', () => {
const finding = [
"**[Amended: D2, D3, D4, D5]** The original sketch had no coordination between a",
"cache fill and a write and omitted the single-flight wrapper and the absence",
"sentinel; finding F1 showed that violates the read-after-write rule. The",
"ordering rules below replace it. `flight` is the existing per-key single-flight",
"wrapper extended with `invalidate(key)` and `invalidateAll()`; a fill ticket is",
"`live()` until its key is invalidated. `ProfileNotFound` stands for the",
"repository's existing typed not-found error class.",
].join('\n');
test('accepts the actual finding without requiring its separate execution diagram', () => {
expect(hasStaleFillRaceFinding(finding)).toBe(true);
expect(hasStaleFillRaceFinding('## Proposed wrapper integration\nKeep the current read-through repository interface and shared adapters.\n' + finding)).toBe(true);
});
test.each([
['current wrapper', finding.replace('original sketch had', 'current wrapper has')],
['proposed implementation', finding.replace('original sketch had', 'proposed implementation has')],
['freshness contract', finding.replace('read-after-write rule', 'read-after-write contract')],
])('recognizes equivalent %s evidence', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(true));
test.each([
['no violation asserted', finding.replace('finding F1 showed that violates the read-after-write rule.', '')],
['negated violation', finding.replace('that violates', 'that does not violate')],
['uncertain violation', finding.replace('that violates', 'that might violate')],
['conditional premise', 'If ' + finding],
['coordination exists', finding.replace('had no coordination', 'had coordination')],
['different operations', finding.replace('cache fill and a write', 'cache hit and a read')],
['wrong contract', finding.replace('read-after-write', 'read-before-write')],
['dismissed defect', finding + '\n\nThis is not a defect; no fix is required.'],
['accepted stale consequence', finding + '\n\nA subsequent stale read is permitted by the amended contract.'],
['quoted finding', finding.split('\n').map(line => '> ' + line).join('\n')],
['indented finding', finding.split('\n').map(line => ' ' + line).join('\n')],
['quoted paragraph', '"' + finding + '"'],
['source preface', 'Example of report format:\n' + finding],
['separate source preface', 'Example of report format:\n\n' + finding],
['backtick source fence', '````text\n' + finding + '\n````'],
['tilde source fence', '~~~text\n' + finding + '\n~~~'],
['unclosed source fence', '```text\n' + finding],
['split unrelated paragraphs', finding.replace('sentinel; finding', 'sentinel.\n\n### Separate issue\nFinding')],
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
});
// Peer counterexamples: embedded, uncertain and hypothetical assertions stay closed.
describe('coordination findings require directly asserted premises and conclusions', () => {
const premise = 'The original sketch had no coordination between a cache fill and a write. ';
const claim = 'This violates the read-after-write rule.';
test('accepts a direct assertion', () => expect(hasStaleFillRaceFinding(premise + claim)).toBe(true));
test.each([
['negated embedded conclusion', premise + 'It is false that this violates the read-after-write rule.'],
['unproven conclusion', premise + 'We have not shown that it violates the read-after-write rule.'],
['uncertain conclusion', premise + 'It is unclear whether this violates the read-after-write rule.'],
['question rather than assertion', premise + claim.replace('.', '?')],
['separate issue without blank line', premise + '\n## A different issue\nReplica lag is high. ' + claim],
['hypothetical premise', 'Suppose ' + premise + claim],
['suggested report', 'A suggested report sentence: ' + premise + claim],
['nested unmatched fence', '````text\n' + premise + claim + '\n```\n' + premise + claim + '\n````'],
['mismatched fence', '~~~text\n' + premise + claim + '\n```\n' + premise + claim],
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
});
// AM first SDK report, exact public Write53 acknowledged by native result54.
// The original report remains in run evidence; this is its complete asserted paragraph.
describe('reported original coordination violation with an owned finding citation', () => {
const finding = "`[Amended: F1, F2, F3, F6]` The original sketch stated that no coordination\nbetween a cache fill and a write was proposed. Review showed that sketch\nviolates the retained read-after-write invariant (see F1). The ordering rules\nbelow replace it. They are the complete new read/write ordering rules.";
test('accepts the exact owned paragraph without requiring its separate amended diagram', () => {
expect(hasStaleFillRaceFinding(finding)).toBe(true);
expect(hasStaleFillRaceFinding('## Proposed wrapper integration\n\n' + finding)).toBe(true);
});
test('binds the named violation to its original subject and finding identity', () => {
expect(hasStaleFillRaceFinding(finding.replaceAll('F1', 'F7'))).toBe(true);
expect(hasStaleFillRaceFinding(finding.replaceAll('sketch', 'wrapper'))).toBe(true);
expect(hasStaleFillRaceFinding('## Historical example\n\nA copied example.\n\n## Current review\n\n' + finding)).toBe(true);
expect(hasStaleFillRaceFinding(finding + '\n\n## Assessment of F2\nF2 is rejected.')).toBe(true);
expect(hasStaleFillRaceFinding(finding + '\n\n## Assessment of F1\n> F1 is rejected.')).toBe(true);
});
test.each([
['negated missing coordination', finding.replace('no coordination', 'coordination')],
['unrelated operations', finding.replace('cache fill and a write', 'cache hit and a read')],
['uncertain absence', finding.replace('stated that no', 'might have stated that no')],
['unproven violation', finding.replace('Review showed', 'Review may show')],
['negated violation', finding.replace('sketch\nviolates', 'sketch\ndoes not violate')],
['hypothetical violation', finding.replace('sketch\nviolates', 'sketch\nmight violate')],
['wrong contract', finding.replace('read-after-write', 'read-before-write')],
['different subject', finding.replace('Review showed that sketch', 'Review showed that wrapper')],
['missing finding identity', finding.replace(' (see F1)', '')],
['question rather than conclusion', finding.replace('(see F1).', '(see F1)?')],
['conditional finding', 'If approved: ' + finding],
['historical owner', '## Historical example\n\n' + finding],
['source owner', '## Quoted source\n\n' + finding],
['hypothetical owner', '## Hypothetical example\n\n' + finding],
['explicit source preface', 'The following is a quoted source excerpt.\n\n' + finding],
['quoted paragraph', '"' + finding + '"'],
['block quote', finding.split('\n').map(line => '> ' + line).join('\n')],
['fenced source', '````text\n' + finding + '\n````'],
['literal assertion', finding.replace('The original sketch', '`The original sketch').replace('(see F1).', '(see F1).`')],
['split unrelated section', finding.replace('Review showed', '\n\n## Another finding\nReview showed')],
['direct same-finding withdrawal', finding + '\n\nF1 is rejected.'],
['later named same-finding withdrawal', finding + '\n\n## Assessment of F1\nF1 is withdrawn.'],
['dismissed defect', finding + '\n\nThis is not a defect; no fix is required.'],
['accepted stale consequence', finding + '\n\nA subsequent stale read is permitted by the amended contract.'],
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
});
+38
View File
@@ -0,0 +1,38 @@
import {describe,test,expect} from 'bun:test';
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
import fixture from './fixtures/ceo-section-ordering-aq.json';
const calls=()=>structuredClone(fixture.calls) as NativePlanQuestionCall[];const first=()=>calls()[2]!;
const fp=(c=first())=>nativePlanCallFingerprint(c,0,true);const classify=(c=first())=>ceoFirstReviewAUQ(fp(c));
function mutate(fn:(c:NativePlanQuestionCall)=>void){const c=first();fn(c);return c;}
function text(fn:(s:string)=>string){return mutate(c=>{const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};});}
describe('AQ owned Section architecture ordering brief',()=>{
test('exact seven calls open review at D4 and retain prior setup',()=>{let started=false;const phases=calls().map(c=>{const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ);started=p.reviewStarted;return p.preReview;});expect(phases).toEqual([true,true,false,false,false,false,false]);expect(classify()).toBe(true);});
test('separate counters, choice order and selected option remain valid',()=>{const c=text(s=>s.replace('D4 — Section 1 (Architecture), issue 1:','D9 — Section 3 (Architecture), issue 2:').replace(/\b1([ABC])\b/g,'2$1')),q=c.questions[0]!;for(const o of q.options)o.label=o.label.replace(/^1/,'2');q.options.reverse();for(const o of q.options){c.answers={[q.question]:o.label};expect(classify(c)).toBe(true);}});
test('quoted archive and conditional consequences do not cancel current evidence',()=>{expect(classify(text(s=>s+'\nArchived note: "This finding is withdrawn."'))).toBe(true);expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=' "Earlier review assessment: This remedy is withdrawn."';}))).toBe(true);});
test('native completion and menu ownership remain required',()=>{
for(const fn of [(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.answeredAt;},(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers!['foreign']='foreign';},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered'};},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;}])expect(classify(mutate(fn))).toBe(false);
for(const f of [{...fp(),signature:'foreign:tool'},{...fp(),nativeQuestionIndex:1},{...fp(),options:fp().options.toReversed()}])expect(ceoFirstReviewAUQ(f)).toBe(false);
});
test('malformed or competing identities and setup headers fail closed',()=>{
for(const [a,b] of [['D4 —','D04 —'],['Section 1 (','Section 01 ('],['issue 1:','issue 0:'],['issue 1:','issue 1.2:'],['(Architecture)','(Source excerpt)']])expect(classify(text(s=>s.replace(a,b)))).toBe(false);
for(const h of ['Section 9','Issue 9','Section 01','Routing','Approach'])expect(classify(mutate(c=>{c.questions[0]!.header=h;}))).toBe(false);
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.label=c.questions[0]!.options[0]!.label.replace('1A)','2A)');c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};}))).toBe(false);
});
test('unique current context and assessment cannot come from source or a conditional',()=>{
for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, ','Provided approval, ','Assuming approval, '])for(const field of ['Project/branch/task: ','ELI10: '])expect(classify(text(s=>s.replace(field,field+p)))).toBe(false);
for(const p of ['Source:\n','Earlier review assessment:\n','Project/branch/task: duplicate\n','ELI10: duplicate\n'])expect(classify(text(s=>s.replace('ELI10:',p+'ELI10:')))).toBe(false);
expect(classify(text(s=>s.replace(/^Project\/branch\/task:.*\n/m,'')))).toBe(false);
});
test('current gap and actual commit-then-notify action stay mandatory',()=>{
expect(classify(text(s=>s.replace('but never says whether the email runs inside the database transaction or after it commits','and explicitly specifies that email follows the database commit')))).toBe(false);
for(const [a,b] of [['COMMIT, then call the mail client','call the mail client, then COMMIT'],['Load user and orders, assign payment_status=paid and PaymentIntent ID, COMMIT, then call the mail client','Record this plan as complete'],['Mail failure can never roll back a committed payment','Mail failure can roll back the payment']])expect(classify(mutate(c=>{const o=c.questions[0]!.options[0]!;o.description=o.description!.replace(a,b);}))).toBe(false);
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.label='1A) Save the review';c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};}))).toBe(false);
});
test('direct current status and action cancellation close their owners',()=>{
for(const s of ['withdrawn','superseded','"closed"','“withdrawn”']){expect(classify(text(t=>t+` This finding is ${s}.`))).toBe(false);expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=` This amendment is ${s}.`;}))).toBe(false);}
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=' Correction: do not commit before sending email.';}))).toBe(false);
expect(classify(text(s=>s+' Correction: this ordering gap is resolved.'))).toBe(false);
});
test('source or conditional options cannot supply the amendment',()=>{for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, ','Provided approval, ']){expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description=p+c.questions[0]!.options[0]!.description;}))).toBe(false);expect(classify(mutate(c=>{const q=c.questions[0]!;q.options[0]!.label=q.options[0]!.label.replace('1A) ','1A) '+p);c.answers={[q.question]:q.options[0]!.label};}))).toBe(false);}});
});
+91
View File
@@ -0,0 +1,91 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import captured from './fixtures/ceo-section-parenthesis-at.json';
const call = (): any => structuredClone(captured.calls[1]);
const fp = (value: any) => nativePlanCallFingerprint(value, 0, true);
const matches = (value: any) => ceoFirstReviewAUQ(fp(value));
function edit(value: any, change: (text: string) => string) {
const q = value.questions[0], answer = value.answers[q.question];
q.question = change(q.question); value.answers = { [q.question]: answer };
}
test('the exact completed combined section/finding brief opens the retry review', () => {
const value = call(); expect(matches(value)).toBe(true); expect(value).toEqual(captured.calls[1]);
let started = false;
const phases = captured.calls.map(value => {
const phase = planCountQuestionPhase(fp(value), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted; return phase.preReview;
});
expect(phases).toEqual([true, false, false, false, false, false, false]);
});
test('decision, section, finding and descriptive header retain separate identities', () => {
for (const change of [
(text: string) => text.replace(/^D4/, 'D19'),
(text: string) => text.replace('Section 1, finding 1', 'Section 7, finding 1').replace('review-s1-', 'review-s7-'),
(text: string) => text.replace(') — ', ') - '),
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(true); }
for (const header of ['Receipt rescue', 'Error contract', 'Finding 1', 'Issue 1', 'Section 1', 'Section 1 finding 1']) {
const value = call(); value.questions[0].header = header; expect(matches(value)).toBe(true);
}
for (const option of call().questions[0].options) {
const value = call(); value.answers[value.questions[0].question] = option.label; expect(matches(value)).toBe(true);
}
});
test('conflicting annotation, qid, header and option identities cannot open review', () => {
for (const change of [
(text: string) => text.replace('Section 1, finding 1', 'Section 0, finding 1'),
(text: string) => text.replace('Section 1, finding 1', 'Section 1, finding 0'),
(text: string) => text.replace('Section 1, finding 1', 'Section 1, finding 2'),
(text: string) => text.replace('review-s1-', 'review-s9-'),
(text: string) => text.replace('plan-ceo-review-s1-', 'plan-eng-review-s1-'),
(text: string) => text.replace('Section 1, finding 1', 'Section 1, hypothetical finding 1'),
(text: string) => text.replace(/^Recommendation: 1A/m, 'Recommendation: 9A'),
(text: string) => text + '\n<gstack-qid:plan-ceo-review-s1-other>',
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
for (const header of ['Finding 9', 'Finding one', 'Issue 9', 'Section 9', 'Section 1 finding 9', 'Section one', 'Approach']) {
const value = call(); value.questions[0].header = header; expect(matches(value)).toBe(false);
}
});
test('only current owned assessments and offered amendments supply coverage', () => {
for (const change of [
(text: string) => 'Example: ' + text,
(text: string) => '> ' + text,
(text: string) => '```\n' + text + '\n```',
(text: string) => text.replace('\nProject/branch/task:', '\nSource:\nProject/branch/task:'),
(text: string) => text.replace('Project/branch/task: ', 'Project/branch/task: If approved, '),
(text: string) => text.replace(/^ELI10: /m, 'ELI10: If approved, '),
(text: string) => text.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: This handler has no current defect and needs no amendment.'),
(text: string) => text + '\nThis finding is withdrawn.',
(text: string) => text + '\nThis finding is "withdrawn".',
(text: string) => text + '\nThis finding is no longer current.',
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
for (const prefix of ['Source: ', 'If approved, ', 'This remedy is withdrawn. ']) {
const value = call(); value.questions[0].options.forEach((option: any) => { option.description = prefix + option.description; });
expect(matches(value)).toBe(false);
}
const history = call(); edit(history, text => text + '\nOld note: "This finding is withdrawn."'); expect(matches(history)).toBe(true);
});
test('native completion and exact same-call options remain required', () => {
for (const change of [
(value: any) => { value.answered = false; },
(value: any) => { value.failed = true; },
(value: any) => { value.sessionId = ''; },
(value: any) => { value.toolUseId = ''; },
(value: any) => { value.answeredAt = 'invalid'; },
(value: any) => { value.unansweredQuestionIndices = [0]; },
(value: any) => { value.answers = {}; },
(value: any) => { value.answers[value.questions[0].question] = 'Foreign answer'; },
(value: any) => { value.questions[0].multiSelect = true; },
(value: any) => { value.questions.push(structuredClone(value.questions[0])); },
(value: any) => { value.questions[0].options[1].description = ''; },
(value: any) => { value.questions[0].options[1].label = '9B: Foreign amendment'; },
]) { const value = call(); change(value); expect(matches(value)).toBe(false); }
const original = fp(call());
for (const value of [{ ...original, signature: 'foreign:call' }, { ...original, nativeCall: undefined },
{ ...original, nativeQuestionIndex: 1 }, { ...original, options: original.options.slice(1) }]) expect(ceoFirstReviewAUQ(value)).toBe(false);
});
test('new public artifacts select only CEO finding count', () => {
for (const file of ['test/ceo-section-parenthesis-at.test.ts', 'test/fixtures/ceo-section-parenthesis-at.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name)).toEqual(['plan-ceo-finding-count']);
});
+93
View File
@@ -0,0 +1,93 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
import fixture from './fixtures/ceo-sequence-aq.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
const accepts = (call: any) => ceoFirstReviewAUQ(nativePlanCallFingerprint(call, 0, true));
function changed(edit: (q: any, call: any) => void) {
const call = structuredClone(fixture.calls[2]!), q = call.questions[0]!;
const selected = q.options.findIndex(o => o.label === call.answers[q.question]);
edit(q, call);
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
return call;
}
test('exact completed prefix keeps setup and approach before the current sequence finding', () => {
expect(fixture.calls.map(accepts)).toEqual([false, false, true]);
});
test('equivalent decision identities and explicit current sequencing gaps retain the finding', () => {
for (const gap of [
'the plan never defines the sequence or the transaction boundary.',
'this plan does not specify the order and the commit point.',
]) expect(accepts(changed(q => {
q.question = q.question.replace(/^Project\/branch\/task:.*$/m, 'Project/branch/task: main, PLAN.md; '+gap);
}))).toBe(true);
expect(accepts(changed(q => { q.question=q.question.replace(/^D2 —/, 'd19 -');q.header='d19 Order'; }))).toBe(true);
});
test('native completion, matching identities and selected offered answer remain mandatory', () => {
for (const edit of [
(_q:any,c:any)=>{c.answered=false;}, (_q:any,c:any)=>{c.failed=true;},
(_q:any,c:any)=>{c.unansweredQuestionIndices=[0];}, (_q:any,c:any)=>{c.answeredAt='invalid';},
(q:any)=>{q.header='D3 Sequence';}, (q:any)=>{q.header='D2 Approach';},
(q:any)=>{q.multiSelect=true;}, (q:any)=>{q.question=q.question.replace('Recommendation: A','Recommendation: Z');},
(q:any)=>{q.options[1].label=q.options[1].label.replace('B)','A)');},
]) expect(accepts(changed(edit))).toBe(false);
const noAnswer=changed(()=>{});noAnswer.answers={};expect(accepts(noAnswer)).toBe(false);
const fp=nativePlanCallFingerprint(changed(()=>{}),0,true);
expect(ceoFirstReviewAUQ({...fp,signature:'foreign:call'})).toBe(false);
expect(ceoFirstReviewAUQ({...fp,nativeCall:undefined})).toBe(false);
expect(ceoFirstReviewAUQ({...fp,nativeQuestionIndex:1})).toBe(false);
expect(ceoFirstReviewAUQ({...fp,options:fp.options.map((o,i)=>i===0?{...o,label:'Foreign choice'}:o)})).toBe(false);
});
test('current metadata cannot be replaced by source, history, conditional or duplicate ownership', () => {
for(const prefix of ['Source excerpt: ', 'Earlier review assessment: ', 'If approved, ', 'For historical context, ']) {
expect(accepts(changed(q=>{q.question=q.question.replace('Project/branch/task: ','Project/branch/task: '+prefix);}))).toBe(false);
expect(accepts(changed(q=>{q.question=q.question.replace('ELI10: ','ELI10: '+prefix);}))).toBe(false);
}
for(const edit of [
(q:any)=>{q.question=q.question.replace('the plan lists','the previous plan lists');},
(q:any)=>{q.question=q.question.replace('but never fixes','and now defines');},
(q:any)=>{q.question=q.question.replace(/^Project\/branch\/task:.*$/m,'Project/branch/task: main, PLAN.md; no current sequencing gap.');},
(q:any)=>{q.question=q.question.replace('\nELI10:','\nSource:\nELI10:');},
(q:any)=>{q.question=q.question.replace('\nELI10:','\nProject/branch/task: another plan\nELI10:');},
]) expect(accepts(changed(edit))).toBe(false);
});
test('current withdrawals and a missing commit-first remedy or opposed risk remain setup', () => {
for(const status of ['This finding is withdrawn.','This finding is "closed".','There is no current gap.',
'The gap is resolved.', 'This sequence has been fixed.', 'This transaction boundary is "defined".'])
expect(accepts(changed(q=>{q.question+='\n'+status;}))).toBe(false);
for(const edit of [
(q:any)=>{q.options[0].label='A) Archive the plan (Recommended)';},
(q:any)=>{q.options[0].description='Source excerpt: '+q.options[0].description;},
(q:any)=>{q.options[0].description='Transaction: lookup + update, do not commit. Then receipt send.';},
(q:any)=>{q.options[0].description+=' This remedy is "withdrawn".';},
(q:any)=>{q.options[2].label='C) Save the report';},
(q:any)=>{q.options[2].description='Source excerpt: '+q.options[2].description;},
(q:any)=>{q.options[2].description='Lookup and update with a defined commit point.';},
(q:any)=>{q.options[2].description+=' This option is cancelled.';},
]) expect(accepts(changed(edit))).toBe(false);
});
test('regression paths belong only to the dense CEO owner', () => {
for (const path of ['test/ceo-sequence-aq.test.ts','test/fixtures/ceo-sequence-aq.json'])
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(path)).map(([owner])=>owner)).toEqual(['plan-ceo-finding-count']);
const owner=E2E_TOUCHFILES['plan-ceo-finding-count']!;
for(let i=0;i<owner.length;i++) {
expect(Object.hasOwn(owner,i)).toBe(true);
expect(typeof owner[i]).toBe('string');
}
});
test('source options, conditional metadata and later contract contradictions cannot own sequencing evidence', () => {
for(const edit of [
(q:any)=>{q.options[2].description='> '+q.options[2].description;},
(q:any)=>{q.options[2].description='~~~\n'+q.options[2].description+'\n~~~';},
(q:any)=>{q.question=q.question.replace('Project/branch/task: main','Project/branch/task: Assuming approval, main');},
(q:any)=>{q.question=q.question.replace('Project/branch/task: main','Project/branch/task: Provided approval, main');},
(q:any)=>{q.question+='\nThis decision is "superseded".';},
(q:any)=>{q.question+='\nThis sequence is "cancelled".';},
(q:any)=>{q.question+='\nThis sequence is not current.';},
(q:any)=>{q.options[0].description+='\nCorrection: the receipt is sent before the payment commit.';},
(q:any)=>{q.options[2].description+='\nCorrection: this transaction boundary is now defined.';},
]) expect(accepts(changed(edit))).toBe(false);
expect(accepts(changed(q=>{q.question+='\n"Earlier review assessment: This sequence is cancelled."';}))).toBe(true);
expect(accepts(changed(q=>{q.question+='\nThe archive sequence is cancelled.';}))).toBe(true);
});
+83
View File
@@ -0,0 +1,83 @@
import { expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import fixture from './fixtures/ceo-test-subject-ao.json';
const calls=fixture.fingerprints as AskUserQuestionFingerprint[];
const actual=calls[1]!;
function change(edit:(q:any, call:any, fp:any)=>void) {
const fp=structuredClone(actual),call=fp.nativeCall!,q=call.questions[0]!;
const selected=q.options.findIndex(o=>o.label===call.answers?.[q.question]);
edit(q,call,fp);
call.answers={[q.question]:q.options[selected]?.label??''};
fp.options=q.options.map((o,i)=>({index:i+1,label:o.label}));
return fp;
}
test('the completed affected-Test question starts review from its current ELI10 assertion gap',()=>{
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false,true,false]);
let review=false;
expect(calls.map(fp=>{const phase=planCountQuestionPhase(fp,review,ceoStep0Boundary,ceoFirstReviewAUQ);review=phase.reviewStarted;return phase.preReview;})).toEqual([true,false,false]);
});
test('structural Test identity permits ordinary question and separator variations',()=>{
for(const title of [
'D4 — Test 1: choose its assertion?',
'D4 — Test 1 — which assertion belongs here?',
'D4 - Test 1 (successful charge): what must this test verify?',
'd4 — Test 1 (successful charge): assertion choice?',
'D4 — Test 1 what should the expected result be?',
]) expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace(/^[^\n]+/,title);}))).toBe(true);
expect(ceoFirstReviewAUQ(change(q=>{
q.question=q.question.replace(/^D4/,'D17').replace(/\b4([A-C])\b/g,'17$1');
q.options=q.options.map((o:any)=>({...o,label:o.label.replace(/^4/,'17')}));
}))).toBe(true);
});
test('test headers, competing finding IDs and uniform foreign decision choices cannot borrow the assessment',()=>{
for(const header of ['Test 2','Finding 1','Issue 1','Approach']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 (Finding 2)');}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{
q.question=q.question.replace(/\b4([A-C])\b/g,'8$1');q.options=q.options.map((o:any)=>({...o,label:o.label.replace(/^4/,'8')}));
}))).toBe(false);
});
test('explicit Test and decision identifiers must be anchored integers with one test owner',()=>{
for(const header of ['Test 0','Test 01','Test 1.2']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(false);
for(const decision of ['D0','D04']) expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace(/^D4/,decision);}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 (Test 2)');}))).toBe(false);
for(const header of ['Receipt assertion','Test contract','Test 1: receipt assertion']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(true);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 ("Test 2" is an archive label)');}))).toBe(true);
});
test('the owned weak assertion must remain current and outside quoted or conditional source frames',()=>{
for(const intro of ['Source excerpt: ','Earlier review assessment: ','If approved later, ']) expect(ceoFirstReviewAUQ(change(q=>{
q.question=q.question.replace('The planned test only checks',intro+'The planned test only checks');
}))).toBe(false);
for(const intro of ['Source excerpt follows. ','Earlier review assessment follows. ','If approved later. ']) expect(ceoFirstReviewAUQ(change(q=>{
q.question=q.question.replace('The planned test only checks',intro+'The planned test only checks');
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('The planned test only checks','The planned test no longer only checks');}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('The planned test only checks that the receipt is truthy.','"The planned test only checks that the receipt is truthy."');}))).toBe(false);
});
test('direct current withdrawals stay effective while a quoted historical note stays harmless',()=>{
for(const text of ['This finding is withdrawn.','This assessment is "closed".','Correction: this explanation is not current.']) expect(ceoFirstReviewAUQ(change(q=>{q.question+='\n'+text;}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('\nELI10:','\nArchive note: "Source: this finding is withdrawn."\nELI10:');}))).toBe(true);
});
test('the complete current amendment belongs to an offered option',()=>{
for(const prefix of ['Source excerpt: ','Historical example: ','If approved later: ']) expect(ceoFirstReviewAUQ(change(q=>{
for(const o of q.options)o.description=prefix+o.description;
}))).toBe(false);
for(const text of [' This amendment is withdrawn.',' This amendment is "closed".',' This remedy is a historical example, not the current option.']) expect(ceoFirstReviewAUQ(change(q=>{
for(const o of q.options)o.description+=text;
}))).toBe(false);
expect(ceoFirstReviewAUQ(change(q=>{q.options[0].label='4A: Keep the truthy assertion (recommended)';}))).toBe(false);
});
test('native completion, exact answer, index and menu identity remain required',()=>{
for(const edit of [
(_q:any,c:any)=>{c.answered=false;},(_q:any,c:any)=>{c.failed=true;},
(_q:any,c:any)=>{delete c.answeredAt;},(_q:any,c:any)=>{c.unansweredQuestionIndices=[0];},
(_q:any,_c:any,fp:any)=>{fp.signature='foreign:tool';},
(_q:any,_c:any,fp:any)=>{fp.nativeQuestionIndex=1;},(q:any)=>{q.multiSelect=true;},
]) expect(ceoFirstReviewAUQ(change(edit))).toBe(false);
const wrongAnswer=change(()=>{});wrongAnswer.nativeCall!.answers={};expect(ceoFirstReviewAUQ(wrongAnswer)).toBe(false);
const wrongMenu=change(()=>{});wrongMenu.options[0]!.label='Foreign menu';expect(ceoFirstReviewAUQ(wrongMenu)).toBe(false);
});
test('new inputs belong only to the dense CEO finding owner',()=>{
for(const file of ['test/ceo-test-subject-ao.test.ts','test/fixtures/ceo-test-subject-ao.json']) expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(file)).map(([owner])=>owner)).toEqual(['plan-ceo-finding-count']);
for(const paths of Object.values(E2E_TOUCHFILES))for(let i=0;i<paths.length;i++)expect(Object.hasOwn(paths,i)&&typeof paths[i]==='string').toBe(true);
});
+109
View File
@@ -0,0 +1,109 @@
import { describe, expect, test } from 'bun:test';
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
import fixture from './fixtures/ceo-transaction-contract-ar.json';
const calls = () => structuredClone(fixture.calls) as NativePlanQuestionCall[];
const first = () => calls()[2]!;
const fp = (call = first()) => nativePlanCallFingerprint(call, 0, true);
const classify = (call = first()) => ceoFirstReviewAUQ(fp(call));
const mutate = (fn: (call: NativePlanQuestionCall) => void) => { const call = first(); fn(call); return call; };
const prose = (fn: (text: string) => string) => mutate(call => {
const q = call.questions[0]!, answer = call.answers![q.question]!;
q.question = fn(q.question); call.answers = { [q.question]: answer };
});
const option = (at: number, fn: (o: NativePlanQuestionCall['questions'][number]['options'][number]) => void) => mutate(call => {
const q = call.questions[0]!, selected = q.options.findIndex(o => o.label === call.answers![q.question]);
fn(q.options[at]!); call.answers = { [q.question]: q.options[selected]!.label };
});
describe('AR current transaction decision', () => {
test('the exact transaction decision starts review after setup', () => {
let started = false;
const phases = calls().map(call => {
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
started = phase.reviewStarted; return phase.preReview;
});
expect(phases).toEqual([true, true, false, false, false, false, false, false]);
expect(calls().map(classify)).toEqual([false, false, true, false, false, false, false, false]);
});
test('title wording and ordinal punctuation do not supply semantics', () => {
expect(classify(prose(s => s.replace('Where does the user update commit relative to the email call?', 'When should the update commit before the email call?')))).toBe(true);
expect(classify(mutate(c => { c.questions[0]!.header = 'Transaction boundary'; }))).toBe(true);
expect(classify(mutate(c => {
const q = c.questions[0]!, answer = c.answers![q.question]!;
q.options.forEach(o => { o.label = o.label.replace(/^3([A-Z]) /, '3$1) '); });
c.answers = { [q.question]: answer.replace(/^3([A-Z]) /, '3$1) ') };
}))).toBe(true);
expect(classify(prose(s => s + '\nArchived note: "This finding is withdrawn."'))).toBe(true);
expect(classify(mutate(c => { const q = c.questions[0]!; c.answers = { [q.question]: q.options[1]!.label }; }))).toBe(true);
});
test('a complete owned successful answer is required', () => {
for (const change of [
(c: NativePlanQuestionCall) => { c.sessionId = ''; }, (c: NativePlanQuestionCall) => { c.toolUseId = ''; },
(c: NativePlanQuestionCall) => { c.answered = false; }, (c: NativePlanQuestionCall) => { c.failed = true; },
(c: NativePlanQuestionCall) => { delete c.failed; }, (c: NativePlanQuestionCall) => { delete c.answeredAt; },
(c: NativePlanQuestionCall) => { c.answeredAt = 'invalid'; }, (c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; }, (c: NativePlanQuestionCall) => { c.answers = {}; },
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
]) expect(classify(mutate(change))).toBe(false);
for (const fingerprint of [{ ...fp(), signature: 'foreign' }, { ...fp(), nativeQuestionIndex: 1 }, { ...fp(), options: fp().options.toReversed() }])
expect(ceoFirstReviewAUQ(fingerprint)).toBe(false);
});
test('decision, header, recommendation and offered ordinals agree', () => {
for (const change of [(s: string) => s.replace('D3 —', 'D0 —'), (s: string) => s.replace('D3 —', 'D03 —'), (s: string) => s.replace('Recommendation: 3A', 'Recommendation: 4A')])
expect(classify(prose(change))).toBe(false);
for (const header of ['Routing', 'Approach', 'D4 Txn boundary', 'D03 Txn boundary', 'Source Txn boundary'])
expect(classify(mutate(c => { c.questions[0]!.header = header; }))).toBe(false);
for (const label of ['03A Commit update, then email', '4A Commit update, then email', '3A) 4A Commit update, then email'])
expect(classify(option(0, o => { o.label = label; }))).toBe(false);
});
test('a unique current context and assessment are required', () => {
for (const field of ['Project/branch/task: ', 'ELI10: ']) for (const prefix of ['Source excerpt: ', 'Earlier review assessment: ', 'If approved, ', 'Assuming approval, '])
expect(classify(prose(s => s.replace(field, field + prefix)))).toBe(false);
for (const prefix of ['Source excerpt:\n', 'Project/branch/task: duplicate\n', 'ELI10: duplicate\n'])
expect(classify(prose(s => s.replace('ELI10:', prefix + 'ELI10:')))).toBe(false);
expect(classify(prose(s => s.replace(/^Project\/branch\/task:.*\n/m, '')))).toBe(false);
expect(classify(prose(s => '```\n' + s + '\n```'))).toBe(false);
});
test('the missing boundary must remain current and unresolved', () => {
expect(classify(prose(s => s.replace('but never says whether', 'and explicitly specifies whether')))).toBe(false);
expect(classify(prose(s => s.replace('The plan says', 'Earlier review assessment follows. The plan says')))).toBe(false);
for (const tail of ['This finding is withdrawn.', 'This transaction boundary is now specified.', 'Correction: this transaction boundary is "resolved".'])
expect(classify(prose(s => s + '\n' + tail))).toBe(false);
});
test('one current amendment owns order and rollback safety', () => {
for (const replacement of ['before commit, inside any DB transaction', 'after commit, inside the DB transaction'])
expect(classify(option(0, o => { o.description = o.description!.replace('after commit, outside any DB transaction', replacement); }))).toBe(false);
expect(classify(option(0, o => { o.description = o.description!.replace('can never roll back paid status', 'can roll back paid status'); }))).toBe(false);
expect(classify(option(0, o => { o.description = o.description!.replace('Lookup and update commit in one transaction;', 'No transactional update is planned;'); }))).toBe(false);
expect(classify(option(0, o => { o.label = '3A Write the final report'; }))).toBe(false);
for (const prefix of ['Source excerpt: ', 'If approved, '])
expect(classify(option(0, o => { o.description = prefix + o.description; }))).toBe(false);
for (const tail of ['This amendment is "withdrawn".', 'Correction: do not commit the update before email.'])
expect(classify(option(0, o => { o.description += ' ' + tail; }))).toBe(false);
});
test('new transaction syntax rejects stale and conditional evidence', () => {
for (const prefix of ['Assuming approval, ', 'Provided approval, ']) {
expect(classify(prose(s => s.replace('ELI10: ', 'ELI10: ' + prefix)))).toBe(false);
expect(classify(option(0, o => { o.description = prefix + o.description; }))).toBe(false);
expect(classify(option(1, o => { o.description = prefix + o.description; }))).toBe(false);
}
for (const status of ['superseded', '"superseded"', '"resolved"', 'no longer current', '"no longer current"']) {
expect(classify(prose(s => s + '\nThis finding is ' + status + '.'))).toBe(false);
expect(classify(option(0, o => { o.description += ' This amendment is ' + status + '.'; }))).toBe(false);
expect(classify(option(1, o => { o.description += ' This option is ' + status + '.'; }))).toBe(false);
}
for (const history of ['> This finding is superseded.', 'Archived note: "This finding is superseded."', 'Archived note: "This finding is no longer current."', '~~~\nThis finding is superseded.\n~~~'])
expect(classify(prose(s => s + '\n' + history))).toBe(true);
for (const convert of [(s: string) => '> ' + s, (s: string) => '"' + s + '"', (s: string) => '`' + s + '`']) {
expect(classify(option(0, o => { o.description = convert(o.description!); }))).toBe(false);
expect(classify(option(1, o => { o.description = convert(o.description!); }))).toBe(false);
}
});
test('the opposed option owns the unchanged risk', () => {
expect(classify(option(1, o => { o.description = 'The transaction shape is safe and fully specified.'; }))).toBe(false);
expect(classify(option(1, o => { o.description = 'Source excerpt: ' + o.description; }))).toBe(false);
expect(classify(option(1, o => { o.description += ' This option is withdrawn.'; }))).toBe(false);
});
});
+173
View File
@@ -0,0 +1,173 @@
import { describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
import { migrateClaudeCodeSkills } from '../lib/claude-code-migration';
const ROOT = path.resolve(import.meta.dir, '..');
const banner = '<!-- AUTO-GENERATED from SKILL.md.tmpl — do not edit directly -->\n<!-- Regenerate: bun run gen:skill-docs -->';
const skill = (name: string, host = '') => `---\nname: ${name}\n---\n${banner}\n${host}\n`;
function put(file: string, text: string): void {
fs.mkdirSync(path.dirname(file), { recursive: true }); fs.writeFileSync(file, text);
}
function fixture() {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'claude-rename-test-'));
const root = path.join(dir, 'checkout'); const home = path.join(dir, 'home');
const codex = path.join(home, 'custom-codex', 'skills');
const kiro = path.join(home, '.kiro', 'skills');
for (const rel of ['bin/gstack-claude-code', 'lib/claude-code.ts', 'lib/claude-code-windows-job.ts', 'lib/claude-bin.ts', 'lib/outside-review-result.ts']) put(path.join(root, rel), rel);
const oldRender = path.join(root, '.agents', 'skills', 'gstack-claude');
put(path.join(oldRender, 'SKILL.md'), skill('gstack-claude'));
fs.mkdirSync(codex, { recursive: true }); fs.mkdirSync(kiro, { recursive: true });
const calls: string[] = []; const messages: string[] = [];
const render = (host: string, out: string) => {
calls.push(host);
const subdir = host === 'codex' ? '.agents' : `.${host}`;
// Native generation keeps the unprefixed frontmatter name even though
// the installed directory is namespaced (gstack-claude-code).
put(path.join(out, subdir, 'skills', 'gstack-claude-code', 'SKILL.md'), skill('claude-code', host));
put(path.join(out, subdir, 'skills', 'gstack-review', 'SKILL.md'), skill('gstack-review', `${host} native outside provider`));
put(path.join(out, subdir, 'skills', 'gstack-review', 'sections', 'gate.md'), `${banner}\n${host} native gate`);
put(path.join(out, subdir, 'skills', 'gstack', 'SKILL.md'), skill('gstack', host));
put(path.join(out, subdir, 'skills', 'gstack-office-hours', 'SKILL.md'), skill('gstack-office-hours', `${host} native consultation`));
put(path.join(out, subdir, 'skills', 'gstack-upgrade', 'SKILL.md'), skill('gstack-upgrade', `./setup --host ${host}`));
};
const run = (overrides: Partial<Parameters<typeof migrateClaudeCodeSkills>[0]> = {}) => migrateClaudeCodeSkills({
installDir: root, home, env: { CODEX_HOME: path.dirname(codex) }, render,
log: line => messages.push(line), ...overrides,
});
return { dir, root, home, codex, kiro, oldRender, calls, messages, render, run };
}
describe('Claude wrapper installed-name migration', () => {
test('migrates existing Codex and Kiro installs independently of the selected setup host', () => {
const f = fixture();
try {
fs.symlinkSync(f.oldRender, path.join(f.codex, 'gstack-claude'));
put(path.join(f.kiro, 'gstack-claude', 'SKILL.md'), skill('gstack-claude'));
put(path.join(f.kiro, 'gstack-claude', 'notes.md'), 'user notes');
put(path.join(f.kiro, 'gstack-review', 'SKILL.md'), skill('gstack-review', 'old Codex-shaped Kiro output'));
put(path.join(f.kiro, 'gstack-review', 'sections', 'gate.md'), `${banner}\nold Codex gate`);
put(path.join(f.kiro, 'gstack-review', 'sections', 'notes.md'), 'user section notes');
put(path.join(f.kiro, 'gstack', 'office-hours', 'SKILL.md'), skill('gstack-office-hours', 'old consultation'));
put(path.join(f.kiro, 'gstack', 'gstack-upgrade', 'SKILL.md'), skill('gstack-upgrade', 'old upgrade'));
const result = f.run({ copy: true, render: (host, out) => {
// Both old installations remain readable until that replacement renders.
const old = path.join(host === 'codex' ? f.codex : f.kiro, 'gstack-claude', 'SKILL.md');
expect(fs.existsSync(old)).toBe(true);
f.render(host, out);
} });
expect(result).toEqual({ migrated: 2, pending: [] });
expect(f.calls).toEqual(['codex', 'kiro']);
for (const [host, dir] of [['codex', f.codex], ['kiro', f.kiro]]) {
expect(fs.readFileSync(path.join(dir, 'gstack-claude-code', 'SKILL.md'), 'utf8')).toContain(host);
expect(fs.existsSync(path.join(dir, 'gstack-claude', 'SKILL.md'))).toBe(false);
expect(fs.readFileSync(path.join(dir, 'gstack', 'bin', 'gstack-claude-code'), 'utf8')).toBe('bin/gstack-claude-code');
}
expect(fs.readFileSync(path.join(f.kiro, 'gstack-claude', 'notes.md'), 'utf8')).toBe('user notes');
expect(fs.readFileSync(path.join(f.kiro, 'gstack-review', 'SKILL.md'), 'utf8')).toContain('kiro native outside provider');
expect(fs.readFileSync(path.join(f.kiro, 'gstack-review', 'SKILL.md.before-claude-code'), 'utf8')).toContain('old Codex-shaped');
expect(fs.readFileSync(path.join(f.kiro, 'gstack-review', 'sections', 'gate.md'), 'utf8')).toContain('kiro native gate');
expect(fs.readFileSync(path.join(f.kiro, 'gstack-review', 'sections', 'notes.md'), 'utf8')).toBe('user section notes');
expect(fs.readFileSync(path.join(f.kiro, 'gstack', 'office-hours', 'SKILL.md'), 'utf8')).toContain('kiro native consultation');
expect(fs.readFileSync(path.join(f.kiro, 'gstack', 'gstack-upgrade', 'SKILL.md'), 'utf8')).toContain('./setup --host kiro');
expect(fs.existsSync(path.join(f.codex, 'gstack-review'))).toBe(false); // no new installs
expect(fs.existsSync(f.oldRender)).toBe(false);
expect(f.messages).toHaveLength(1);
expect(f.run()).toEqual({ migrated: 0, pending: [] });
expect(f.messages).toHaveLength(1); // notice is emitted only for a migration
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
});
test('repairs dangling owned links after a standalone build and leaves unrelated Codex output unchanged', () => {
const f = fixture();
try {
fs.symlinkSync(f.oldRender, path.join(f.codex, 'gstack-claude'));
fs.rmSync(f.oldRender, { recursive: true });
const other = path.join(f.root, '.agents', 'skills', 'gstack-review', 'SKILL.md');
put(other, 'Codex Sol profile');
expect(f.run().migrated).toBe(1);
expect(fs.lstatSync(path.join(f.codex, 'gstack-claude-code')).isSymbolicLink()).toBe(true);
expect(fs.readFileSync(other, 'utf8')).toBe('Codex Sol profile');
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
});
test('failed generation preserves the old installed skill and shared render for retry', () => {
const f = fixture();
try {
fs.symlinkSync(f.oldRender, path.join(f.codex, 'gstack-claude'));
const result = f.run({ render: () => { throw new Error('fixture generation failure'); } });
expect(result.pending).toEqual([f.codex]);
expect(fs.readFileSync(path.join(f.codex, 'gstack-claude', 'SKILL.md'), 'utf8')).toContain('gstack-claude');
expect(fs.existsSync(path.join(f.codex, 'gstack-claude-code'))).toBe(false);
expect(f.run().migrated).toBe(1);
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
});
test('foreign old/replacement skills, links, runtime directories and render links are preserved', () => {
for (const conflict of ['old', 'old-one-line-banner', 'old-escaped-link', 'replacement', 'replacement-link', 'runtime', 'render-link']) {
const f = fixture();
try {
const foreign = path.join(f.dir, 'foreign');
put(path.join(foreign, 'SKILL.md'), 'my skill');
const old = path.join(f.codex, 'gstack-claude');
if (conflict === 'old') fs.symlinkSync(foreign, old);
else if (conflict === 'old-one-line-banner') put(path.join(old, 'SKILL.md'), '<!-- AUTO-GENERATED from my own tool -->');
else if (conflict === 'old-escaped-link') {
fs.rmSync(f.oldRender, { recursive: true });
fs.symlinkSync(foreign, f.oldRender);
fs.symlinkSync(f.oldRender, old);
}
else fs.symlinkSync(f.oldRender, old);
if (conflict === 'replacement') put(path.join(f.codex, 'gstack-claude-code', 'SKILL.md'), 'my skill');
if (conflict === 'replacement-link') fs.symlinkSync(foreign, path.join(f.codex, 'gstack-claude-code'));
if (conflict === 'runtime') put(path.join(f.codex, 'gstack', 'SKILL.md'), 'my skill');
if (conflict === 'render-link') fs.symlinkSync(foreign, path.join(f.root, '.agents', 'skills', 'gstack-claude-code'));
const result = f.run();
expect(result.migrated).toBe(0);
expect(fs.existsSync(path.join(old, 'SKILL.md'))).toBe(true);
expect(fs.readFileSync(path.join(foreign, 'SKILL.md'), 'utf8')).toBe('my skill');
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
}
});
test('a missing runtime or malformed replacement cannot retire the original', () => {
for (const failure of ['runtime', 'render']) {
const f = fixture();
try {
fs.symlinkSync(f.oldRender, path.join(f.codex, 'gstack-claude'));
if (failure === 'runtime') fs.rmSync(path.join(f.root, 'lib/claude-code.ts'));
const result = f.run(failure === 'render' ? { render: (_host, out) => put(path.join(out, '.agents/skills/gstack-claude-code/SKILL.md'), skill('wrong')) } : {});
expect(result.pending).toEqual([f.codex]);
expect(fs.existsSync(path.join(f.codex, 'gstack-claude', 'SKILL.md'))).toBe(true);
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
}
});
test('retirement preserves customized copied content without retaining the old command', () => {
const f = fixture();
try {
const old = path.join(f.kiro, 'gstack-claude');
const customized = skill('gstack-claude', 'my extra instructions');
put(path.join(old, 'SKILL.md'), customized);
put(path.join(old, 'SKILL.md.before-claude-code'), 'earlier backup');
expect(f.run().migrated).toBe(1);
expect(fs.existsSync(path.join(old, 'SKILL.md'))).toBe(false);
expect(fs.readFileSync(path.join(old, 'SKILL.md.before-claude-code'), 'utf8')).toBe('earlier backup');
expect(fs.readFileSync(path.join(old, 'SKILL.md.before-claude-code.1'), 'utf8')).toBe(customized);
} finally { fs.rmSync(f.dir, { recursive: true, force: true }); }
});
test('setup runs migration before its first build and defers only the retired Claude render', () => {
const setup = fs.readFileSync(path.join(ROOT, 'setup'), 'utf8');
const migration = setup.indexOf('GSTACK_RENAME_COPY="$IS_WINDOWS" bun_cmd');
expect(migration).toBeGreaterThan(-1);
expect(migration).toBeLessThan(setup.indexOf('bun_cmd run build'));
expect(setup).toContain('export GSTACK_DEFER_CLAUDE_RENAME_PRUNE=1');
expect(setup).toContain('unset GSTACK_DEFER_CLAUDE_RENAME_PRUNE');
expect(setup).toContain('[ "$n" = "gstack-claude" ]');
const version = path.join(ROOT, 'gstack-upgrade/migrations/v1.86.0.0.sh');
expect(fs.statSync(version).mode & 0o111).not.toBe(0);
expect(fs.readFileSync(version, 'utf8')).toContain('gstack-migrate-claude-code');
});
});
+272
View File
@@ -0,0 +1,272 @@
import { afterAll, describe, expect, test } from 'bun:test';
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import path from 'node:path';
import { spawnSync } from 'node:child_process';
import { runClaudeCode } from '../lib/claude-code';
const ROOT = path.resolve(import.meta.dir, '..');
const DIR = mkdtempSync(path.join(tmpdir(), 'claude-code-runner-'));
const FAKE = path.join(DIR, 'fake claude.ts');
const DESCENDANT = path.join(DIR, 'pipe holder.ts');
const CAPTURE = path.join(DIR, 'capture.json');
const PID = path.join(DIR, 'descendant.pid');
// Publish readiness only after the grandchild has initialized and flushed both
// inherited pipes; a PID returned by spawn alone does not establish that state.
writeFileSync(DESCENDANT, `
import { writeFileSync } from 'node:fs';
setInterval(() => {}, 1000);
await new Promise(resolve => process.stdout.write(' ', resolve));
await new Promise(resolve => process.stderr.write(' ', resolve));
writeFileSync(process.env.PID_FILE!, String(process.pid));
`);
writeFileSync(FAKE, `
import { spawn } from 'node:child_process';
import { existsSync, rmSync, writeFileSync } from 'node:fs';
const prompt = await Bun.stdin.text();
writeFileSync(process.env.CAPTURE!, JSON.stringify({args:process.argv.slice(2),prompt,cwd:process.cwd(),model:process.env.ANTHROPIC_MODEL,auth:process.env.ANTHROPIC_API_KEY}));
const mode = process.env.FAKE_MODE;
if (mode === 'timeout' || mode === 'descendant' || mode === 'escaped') {
rmSync(process.env.PID_FILE!, { force: true });
// libuv on Windows kills non-detached children when this fake exits. The
// drain fixture must survive that exit so its inherited pipes remain open.
const child = spawn(process.execPath, [process.env.DESCENDANT!], {stdio:['ignore','inherit','inherit'],
detached:mode === 'escaped' || (process.platform === 'win32' && mode === 'descendant')});
const readyBy = Date.now() + 2000;
while (!existsSync(process.env.PID_FILE!)) {
if (child.exitCode !== null || Date.now() >= readyBy) throw new Error('Descendant did not initialize its inherited pipes');
await Bun.sleep(5);
}
if (mode === 'timeout') await new Promise(() => {});
}
if (mode === 'auth') { process.stderr.write('Not logged in. Please run claude /login.'); process.exit(1); }
if (mode === 'overflow') {
const block = Buffer.alloc(1024 * 1024, 120);
for(let i=0;i<18;i++) { process.stdout.write(block); process.stderr.write(block); }
await new Promise(() => {});
}
if (mode === 'malformed') { process.stdout.write('broken JSON'); process.exit(0); }
if (mode === 'array') { process.stdout.write('[]'); process.exit(0); }
if (mode === 'null') { process.stdout.write('null'); process.exit(0); }
const result = {
result: mode === 'empty' ? ' ' : mode === 'large-success' ? 'x'.repeat(2 * 1024 * 1024) : '[P1] Seeded defect in changed.ts',
is_error:mode === 'error',
session_id:'session-123',
usage:{input_tokens:12,output_tokens:34,cache_read_input_tokens:5},
modelUsage:{'claude-first':{inputTokens:4},'claude-fallback':{inputTokens:8}},
};
await new Promise(resolve => process.stdout.write(JSON.stringify(result),resolve));
process.exit(mode === 'nonzero' ? 7 : 0);
`);
afterAll(() => rmSync(DIR, { recursive: true, force: true }));
function env(mode = 'success'): NodeJS.ProcessEnv {
return { ...process.env, GSTACK_CLAUDE_BIN: process.execPath, GSTACK_CLAUDE_BIN_ARGS: JSON.stringify([FAKE]),
FAKE_MODE: mode, CAPTURE, PID_FILE: PID, DESCENDANT, ANTHROPIC_MODEL: 'configured-model', ANTHROPIC_API_KEY: 'fake-test-credential' };
}
function run(mode = 'success', extra: Partial<Parameters<typeof runClaudeCode>[0]> = {}) {
return runClaudeCode({cwd:DIR, access:'none', timeoutMs:5000, prompt:'review this', env:env(mode), ...extra});
}
function capture() { return JSON.parse(readFileSync(CAPTURE, 'utf8')); }
function running(pid: number): boolean {
try {
process.kill(pid, 0);
// Linux can leave a killed orphan briefly as a zombie before init reaps it.
if (process.platform === 'linux') return readFileSync(`/proc/${pid}/stat`, 'utf8').split(' ')[2] !== 'Z';
return true;
} catch { return false; }
}
async function expectDescendantDead() {
const pid = Number(readFileSync(PID, 'utf8'));
try {
for (let i = 0; i < 20 && running(pid); i++) await Bun.sleep(25);
expect(running(pid)).toBe(false);
} finally {
if (running(pid)) process.kill(pid, 'SIGKILL');
}
}
function cleanupDescendant() {
let pid: number;
try { pid = Number(readFileSync(PID, 'utf8')); } catch { return; }
if (Number.isSafeInteger(pid) && pid > 0 && running(pid)) {
try { process.kill(pid, 'SIGKILL'); } catch { /* Exited after the liveness check. */ }
}
}
describe('Claude Code restricted execution', () => {
test('explicit model override stays one literal argument across access modes and resume', async () => {
const model = 'custom-model "quoted" $(touch /never)';
for (const access of ['none', 'read-only'] as const) {
const result = await run('success', { access, resume: 'session-123', env: { ...env(), GSTACK_CLAUDE_MODEL: model } });
expect(result.status).toBe('completed');
const args = capture().args;
expect(args.filter((arg: string) => arg === '--model')).toHaveLength(1);
expect(args[args.indexOf('--model') + 1]).toBe(model);
expect(args[args.indexOf('--resume') + 1]).toBe('session-123');
expect(capture().auth).toBe('fake-test-credential');
}
await run('success', { env: { ...env(), GSTACK_CLAUDE_MODEL: undefined } });
expect(capture().args).not.toContain('--model');
expect(capture().model).toBe('configured-model');
});
test('stdin stays literal; explicit tools, MCP and hook restrictions preserve configured auth/model', async () => {
const prompt = 'quotes "\' `touch /never` $(touch /never)\nEOF\n--dangerously-skip-permissions';
const result = await run('success', {prompt});
expect(result.status).toBe('completed');
expect(capture().prompt).toBe(prompt);
expect(capture().cwd).toBe(DIR);
expect(capture().model).toBe('configured-model');
expect(capture().auth).toBe('fake-test-credential');
const args = capture().args;
expect(args[args.indexOf('--tools') + 1]).toBe('');
const capability = args[args.indexOf('--append-system-prompt') + 1];
expect(capability).toContain('No tools are available');
expect(capability).toContain('Do not attempt or simulate tool calls');
expect(capability).toContain('If essential context is missing, identify it explicitly');
expect(capability).not.toContain(prompt);
expect(args).toContain('--disable-slash-commands');
expect(args).toContain('--strict-mcp-config');
expect(JSON.parse(args[args.indexOf('--mcp-config') + 1])).toEqual({mcpServers:{}});
expect(JSON.parse(args[args.indexOf('--settings') + 1])).toEqual({disableAllHooks:true});
expect(args[args.indexOf('--disallowedTools') + 1]).toBe('mcp__*');
expect(args).not.toContain('--model');
expect(args).not.toContain('--resume');
expect(args).not.toContain('--dangerously-skip-permissions');
expect(args).not.toContain(prompt);
expect(result.modelUsage).toEqual({'claude-first':{inputTokens:4},'claude-fallback':{inputTokens:8}});
expect(result.usage).toEqual({input_tokens:12,output_tokens:34,cache_read_input_tokens:5});
expect(result.model).toBeUndefined();
expect(result.session_id).toBe('session-123');
});
test('consult exposes only file reading tools and preserves a literal resume argument', async () => {
const resume = 'session " with spaces $(touch /never)';
expect((await run('success', {access:'read-only', resume})).status).toBe('completed');
const args = capture().args;
expect(args[args.indexOf('--tools') + 1]).toBe('Read,Grep,Glob');
expect(args[args.indexOf('--allowedTools') + 1]).toBe('Read,Grep,Glob');
expect(args).not.toContain('--append-system-prompt');
expect(args[args.indexOf('--resume') + 1]).toBe(resume);
expect(args).not.toContain('--no-session-persistence');
});
test('PATH command overrides and argument prefixes resolve in the invocation environment', async () => {
const cliEnv = env();
cliEnv.GSTACK_CLAUDE_BIN = path.basename(process.execPath);
cliEnv.PATH = `${path.dirname(process.execPath)}${path.delimiter}${cliEnv.PATH}`;
expect((await run('success', {env:cliEnv})).status).toBe('completed');
expect(capture().args[0]).toBe('-p');
});
test('WSL-style launcher prefixes remain ordered literal argv entries', async () => {
const cliEnv = env();
cliEnv.GSTACK_CLAUDE_BIN_ARGS = JSON.stringify([FAKE,'claude','--configured-option','a value with spaces']);
expect((await run('success',{env:cliEnv})).status).toBe('completed');
expect(capture().args.slice(0,4)).toEqual(['claude','--configured-option','a value with spaces','-p']);
});
test('missing CLI and broken absolute overrides have distinct actionable failures', async () => {
const missing = await run('success', {env:{PATH:DIR}});
expect(missing.status).toBe('unavailable');
expect(missing.error?.code).toBe('not-found');
expect(missing.error?.message).toContain('GSTACK_CLAUDE_BIN');
const broken = await run('success', {env:{GSTACK_CLAUDE_BIN:path.join(DIR,'missing-cli')}});
expect(broken.status).toBe('unavailable');
expect(broken.error?.code).toBe('spawn');
});
for (const [mode, code] of [
['auth','authentication'], ['nonzero','exit'], ['error','provider-error'],
['malformed','invalid-json'], ['array','invalid-response'], ['null','invalid-response'], ['empty','empty-response'],
]) {
test(`${mode} cannot become a successful review`, async () => {
const result = await run(mode);
expect(result.status).not.toBe('completed');
expect(result.result).toBe('');
expect(result.error?.code).toBe(code);
if (mode === 'auth') expect(result.error?.message).toContain('interactively');
});
}
test('the combined stdout/stderr limit terminates a noisy CLI', async () => {
const result = await run('overflow');
expect(result.status).toBe('error');
expect(result.error?.code).toBe('output-limit');
expect(result.stderr!.length).toBeLessThanOrEqual(16 * 1024);
});
test('timeout kills its descendants and clears process signal listeners', async () => {
rmSync(PID, { force: true });
const before = ['SIGINT','SIGTERM','exit'].map(name => process.listenerCount(name));
const start = Date.now();
try {
const result = await run('timeout', {timeoutMs:500});
expect(result.status).toBe('unavailable');
expect(result.error?.code).toBe('timeout');
expect(Date.now() - start).toBeLessThan(2000);
expect(['SIGINT','SIGTERM','exit'].map(name => process.listenerCount(name))).toEqual(before);
await expectDescendantDead();
} finally { cleanupDescendant(); }
});
test('a child exiting with inherited pipes is unavailable within the drain deadline', async () => {
rmSync(PID, { force: true });
const start = Date.now();
try {
const result = await run('descendant');
expect(result.status).toBe('unavailable');
expect(result.error?.code).toBe('output-drain');
expect(Date.now() - start).toBeLessThan(2000);
// taskkill /T cannot discover an orphan after its parent has exited.
// Windows termination is covered by the timeout case while the parent
// is alive; cleanup below still runs if a drain assertion fails.
if (process.platform !== 'win32') await expectDescendantDead();
} finally { cleanupDescendant(); }
});
test.skipIf(process.platform === 'win32')('an escaped pipe holder cannot wedge draining or count as coverage', async () => {
const start = Date.now();
try {
const result = await run('escaped');
expect(result.status).toBe('unavailable');
expect(result.error?.code).toBe('output-drain');
expect(Date.now() - start).toBeLessThan(2000);
} finally {
const pid = Number(readFileSync(PID, 'utf8'));
if (running(pid)) process.kill(pid, 'SIGKILL');
}
});
test('CLI exits nonzero for malformed responses and argument failures', () => {
const cli = path.join(ROOT, 'bin/gstack-claude-code');
for (const mode of ['success','malformed','error','auth']) {
const result = spawnSync(process.execPath, [cli,'--cwd',DIR,'--access','none','--timeout-ms','5000'], {
env:env(mode), input:'a prompt', encoding:'utf8', timeout:10000,
});
expect(result.status).toBe(mode === 'success' ? 0 : 1);
expect(JSON.parse(result.stdout).status === 'completed').toBe(mode === 'success');
}
const invalid = spawnSync(process.execPath, [cli,'--access','write'], {encoding:'utf8', timeout:5000});
expect(invalid.status).toBe(1);
expect(JSON.parse(invalid.stdout).error.code).toBe('arguments');
});
test('CLI flushes a large successful JSON response before exiting', () => {
const result = spawnSync(process.execPath, [path.join(ROOT,'bin/gstack-claude-code'),'--cwd',DIR,'--access','none','--timeout-ms','5000'], {
env:env('large-success'), input:'prompt', encoding:'utf8', timeout:10000, maxBuffer:8 * 1024 * 1024,
});
expect(result.status).toBe(0);
const parsed = JSON.parse(result.stdout);
expect(parsed.status).toBe('completed');
expect(parsed.result.length).toBe(2 * 1024 * 1024);
});
});
+206
View File
@@ -0,0 +1,206 @@
/** Execute each complete generated wrapper fence in its own fresh shell. */
import { afterAll, beforeAll, describe, expect, test } from 'bun:test';
import * as fs from 'node:fs';
import path from 'node:path';
import os from 'node:os';
import { spawnSync } from 'node:child_process';
const ROOT = path.resolve(import.meta.dir, '..');
const DIR = fs.mkdtempSync(path.join(os.tmpdir(), 'claude-code-skill-'));
const RENDER = path.join(DIR, 'render');
const REPO = path.join(DIR, 'repo');
const SCRATCH = path.join(DIR, 'scratch');
const RUNTIME = path.join(DIR, 'runtime " with $(touch NEVER)');
const BAD_RUNTIME = path.join(DIR, 'bad-runtime');
const PROMPT = path.join(DIR, "prompt ' with $(touch NEVER).txt");
const PROMPT_TEXT = 'Review literal text: $(touch NEVER) `touch NEVER` "\'\n';
const FAKE = path.join(DIR, 'claude.ts');
const CAPTURE = path.join(DIR, 'capture.json');
const q = (text: string) => `'${text.replaceAll("'", "'\\''")}'`;
type Mode = 'review' | 'challenge' | 'consult';
const fences = {} as Record<Mode, string>;
const ENV = { ...process.env, GSTACK_CLAUDE_BIN:process.execPath, GSTACK_CLAUDE_BIN_ARGS:JSON.stringify([FAKE]), CAPTURE,
CLAUDECODE:'', CODEX_THREAD_ID:'test-codex', CODEX_SANDBOX:'', GSTACK_ACTIVE_HOST:'codex',
GIT_AUTHOR_NAME:'Test', GIT_AUTHOR_EMAIL:'test@example.invalid', GIT_COMMITTER_NAME:'Test', GIT_COMMITTER_EMAIL:'test@example.invalid',
TMPDIR:SCRATCH,
};
function git(args: string[], cwd = REPO) {
const result = spawnSync('git', args, {cwd, env:ENV, encoding:'utf8', timeout:5000});
if (result.status !== 0) throw new Error(result.stderr);
}
beforeAll(() => {
for (const dir of [REPO, SCRATCH, path.join(BAD_RUNTIME, 'bin')]) fs.mkdirSync(dir, {recursive:true});
fs.symlinkSync(ROOT, RUNTIME, 'dir');
fs.symlinkSync(path.join(ROOT, 'lib'), path.join(BAD_RUNTIME, 'lib'), 'dir');
fs.writeFileSync(path.join(BAD_RUNTIME, 'bin/gstack-claude-code'), '#!/usr/bin/env bash\nprintf "%s\\n" "$RAW_CAPTURE"\n', {mode:0o755});
fs.writeFileSync(FAKE, `
import {writeFileSync} from 'node:fs';
const prompt = await Bun.stdin.text();
writeFileSync(process.env.CAPTURE!,JSON.stringify({args:process.argv.slice(2),prompt}));
if(process.env.FAKE_ERROR) {process.stdout.write('{broken');process.exit(0);}
console.log(JSON.stringify({result:process.env.FAKE_TEXT || '[P1] Review found a defect.',session_id:'consult-session',modelUsage:{'model-a':{},'model-b':{}}}));
`);
git(['init','-b','main']);
fs.writeFileSync(path.join(REPO, 'changed.txt'), 'baseline\n');
git(['add','changed.txt']);
git(['commit','-m','baseline']);
git(['remote','add','origin','.']);
git(['checkout','-b','work']);
fs.appendFileSync(path.join(REPO, 'changed.txt'), 'committed change\n');
git(['commit','-am','change']);
fs.appendFileSync(path.join(REPO, 'changed.txt'), 'working tree change\n');
// Output-only generation: no mutation of the checkout's active host renders.
const generated = spawnSync(process.execPath, ['run', 'scripts/gen-skill-docs.ts', '--host', 'codex', '--out-dir', RENDER], {
cwd:ROOT, encoding:'utf8', timeout:120000,
});
if (generated.status !== 0) throw new Error(generated.stderr);
const content = fs.readFileSync(path.join(RENDER, '.agents/skills/gstack-claude-code/SKILL.md'), 'utf8');
for (const mode of ['review','challenge','consult'] as const) {
const heading = `## ${mode[0].toUpperCase() + mode.slice(1)} mode`;
const start = content.indexOf(heading);
const end = content.indexOf('\n## ', start + heading.length);
if (start === -1) throw new Error(`Missing generated ${mode} section`);
const section = content.slice(start, end === -1 ? undefined : end);
const matches = [...section.matchAll(/```bash\n([\s\S]*?)\n```/g)];
if (matches.length !== 1) throw new Error(`${mode} must have exactly one executable fence, found ${matches.length}`);
fences[mode] = matches[0][1];
}
});
afterAll(() => fs.rmSync(DIR, {recursive:true, force:true}));
function shell(mode: Mode, extra: NodeJS.ProcessEnv = {}, options: {resume?: boolean; runtime?: string; cwd?: string} = {}) {
fs.rmSync(CAPTURE, {force:true});
fs.writeFileSync(PROMPT, PROMPT_TEXT, {mode:0o600});
// These are precisely the literal substitutions the skill requests. Do not
// prepend setup, merge fences, or supply state from a prior shell invocation.
const script = fences[mode]
.replace("'<prepared-prompt-file>'", q(PROMPT))
.replace("'<gstack-runtime-root>'", q(options.runtime ?? RUNTIME))
.replace("'<base>'", q('main'))
.replace("'<fresh-or-resume>'", q(options.resume ? 'resume' : 'fresh'));
return spawnSync('bash', ['-c',script], {cwd:options.cwd ?? REPO, env:{...ENV,...extra}, encoding:'utf8',timeout:10000});
}
function captured() { return JSON.parse(fs.readFileSync(CAPTURE,'utf8')); }
function expectTempsCleaned() {
expect(fs.existsSync(PROMPT)).toBe(false);
expect(fs.readdirSync(SCRATCH)).toEqual([]);
expect(fs.existsSync(path.join(REPO,'NEVER'))).toBe(false);
}
function saveSession(id: string) {
fs.mkdirSync(path.join(REPO,'.context'),{recursive:true});
fs.writeFileSync(path.join(REPO,'.context/claude-session-id'),id + '\n');
}
function savedSession() { return fs.readFileSync(path.join(REPO,'.context/claude-session-id'),'utf8'); }
describe('complete generated Claude Code wrapper modes', () => {
for (const mode of ['review','challenge','consult'] as const) {
test(`${mode} stale wrapper refuses and cleans its prompt before spawning`, () => {
const result = shell(mode,{CLAUDECODE:'1',CODEX_THREAD_ID:'',GSTACK_ACTIVE_HOST:'claude'});
expect(result.status).toBe(78);
expect(fs.existsSync(CAPTURE)).toBe(false);
expect(result.stderr).toContain('setup --host claude');
expectTempsCleaned();
});
}
for (const mode of ['review','challenge'] as const) {
test(`${mode} independently captures committed and working-tree context with no tools`, () => {
const result = shell(mode);
expect(result.status).toBe(0);
expect(captured().prompt).toStartWith(PROMPT_TEXT);
expect(captured().prompt).toContain('+committed change');
expect(captured().prompt).toContain('+working tree change');
const args = captured().args;
expect(args[args.indexOf('--tools') + 1]).toBe('');
expect(args).not.toContain('--resume');
expect(result.stdout).toContain('[P1] Review found a defect.');
expect(result.stdout).toContain('model-a');
expect(result.stdout).toContain('model-b');
expectTempsCleaned();
});
for (const response of ['I cannot review this request. NO_FINDINGS', 'Several observations without a severity.']) {
test(`${mode} rejects refusal or missing markers after a valid runner completion: ${response}`, () => {
const result = shell(mode, {FAKE_TEXT:response});
expect(result.status).toBe(1);
expect(result.stderr).toContain('missing outside coverage');
expectTempsCleaned();
});
}
}
test('fresh and continued consult each run independently and preserve a literal session argument', () => {
const first = shell('consult', {FAKE_TEXT:'The configuration lives in settings.ts.'});
expect(first.status).toBe(0);
expect(captured().prompt).toBe(PROMPT_TEXT);
const args = captured().args;
expect(args[args.indexOf('--tools') + 1]).toBe('Read,Grep,Glob');
expect(args).not.toContain('--resume');
expect(savedSession()).toBe('consult-session\n');
expectTempsCleaned();
const previous = 'session " with spaces $(touch NEVER)';
saveSession(previous);
const continued = shell('consult', {}, {resume:true});
expect(continued.status).toBe(0);
const resumedArgs = captured().args;
expect(resumedArgs[resumedArgs.indexOf('--resume') + 1]).toBe(previous);
expect(savedSession()).toBe('consult-session\n');
expectTempsCleaned();
});
test('failed resumed invocation cleans captures without overwriting its previous session', () => {
saveSession('previous-session');
const result = shell('consult', {FAKE_ERROR:'1'}, {resume:true});
expect(result.status).toBe(1);
expect(result.stdout).toContain('invalid-json');
expect(savedSession()).toBe('previous-session\n');
expectTempsCleaned();
});
test('complete mode fences reject malformed, false-success and empty runner captures', () => {
for (const raw of ['{broken','[]','{"status":"error","result":"NO_FINDINGS"}','{"status":"completed","is_error":true,"result":"NO_FINDINGS"}','{"status":"completed","result":" "}']) {
for (const mode of ['review','challenge','consult'] as const) {
saveSession('previous-session');
const result = shell(mode, {RAW_CAPTURE:raw}, {runtime:BAD_RUNTIME});
expect(result.status).toBe(1);
expect(result.stderr).toContain('CLAUDE_CODE_ERROR');
expect(savedSession()).toBe('previous-session\n');
expectTempsCleaned();
}
}
});
test('an explicit clean review preserves unknown model identity', () => {
const result = shell('review', {RAW_CAPTURE:'{"status":"completed","result":"NO_FINDINGS"}'}, {runtime:BAD_RUNTIME});
expect(result.status).toBe(0);
expect(result.stdout).toContain('NO_FINDINGS');
expect(result.stdout).toContain('Model: unknown');
expectTempsCleaned();
});
test('missing resume session stops before a provider call', () => {
fs.rmSync(path.join(REPO,'.context/claude-session-id'), {force:true});
const result = shell('consult', {}, {resume:true});
expect(result.status).toBe(1);
expect(result.stderr).toContain('no saved Claude Code session');
expect(fs.existsSync(CAPTURE)).toBe(false);
expectTempsCleaned();
});
test('an empty branch diff stops before invoking Claude Code', () => {
const clean = path.join(DIR, 'clean-repo');
fs.mkdirSync(clean);
git(['init','-b','main'], clean);
fs.writeFileSync(path.join(clean, 'file.txt'), 'unchanged\n');
git(['add','file.txt'], clean);
git(['commit','-m','baseline'], clean);
for (const mode of ['review','challenge'] as const) {
const result = shell(mode, {}, {cwd:clean});
expect(result.status).toBe(0);
expect(result.stdout).toContain('Nothing to review');
expect(fs.existsSync(CAPTURE)).toBe(false);
expectTempsCleaned();
}
});
});
+162
View File
@@ -0,0 +1,162 @@
import { afterAll, describe, expect, test } from 'bun:test';
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import path from 'node:path';
import { spawn, spawnSync } from 'node:child_process';
const ROOT = path.resolve(import.meta.dir, '..');
const DIR = mkdtempSync(path.join(tmpdir(), 'claude-windows-job-'));
const FAKE = path.join(DIR, 'fake claude.ts');
const DESCENDANT = path.join(DIR, 'pipe holder.ts');
const PID_FILE = path.join(DIR, 'descendant.pid');
const PROVIDER_PID_FILE = path.join(DIR, 'provider.pid');
const CLI = path.join(ROOT, 'bin/gstack-claude-code');
const FAILED_JOB = path.join(DIR, 'failed job.ts');
// Exercise the initialization failure at the actual CLI boundary on every OS,
// without adding a production bypass flag for this safety requirement.
writeFileSync(FAILED_JOB, `
Object.defineProperty(process, 'platform', { value: 'win32' });
// OpenProcess(0) is invalid on Windows. Other OSes fail earlier opening the
// Windows DLL; both must fail closed before spawning the configured reviewer.
Object.defineProperty(process, 'pid', { value: 0 });
`);
// Publish readiness only after the grandchild has initialized and flushed both
// inherited pipes; a PID returned by spawn alone does not establish that state.
writeFileSync(DESCENDANT, `
import { writeFileSync } from 'node:fs';
setInterval(() => {}, 1000);
await new Promise(resolve => process.stdout.write(' ', resolve));
await new Promise(resolve => process.stderr.write(' ', resolve));
writeFileSync(process.env.PID_FILE!, String(process.pid));
`);
writeFileSync(FAKE, `
import { spawn } from 'node:child_process';
import { existsSync, rmSync, writeFileSync } from 'node:fs';
await Bun.stdin.text();
writeFileSync(process.env.PROVIDER_PID_FILE!, String(process.pid));
rmSync(process.env.PID_FILE!, { force: true });
const child = spawn(process.execPath, [process.env.DESCENDANT!], {
stdio: ['ignore', 'inherit', 'inherit'],
// Bypass this fake's libuv auto-kill job so the pipe holder survives it.
// DETACHED does not request CREATE_BREAKAWAY_FROM_JOB: gstack's enclosing
// job still owns the descendant, which the assertions below require dead.
detached: process.platform === 'win32' && process.env.FAKE_MODE === 'descendant',
});
const readyBy = Date.now() + 2000;
while (!existsSync(process.env.PID_FILE!)) {
if (child.exitCode !== null || Date.now() >= readyBy) throw new Error('Descendant did not initialize its inherited pipes');
await Bun.sleep(5);
}
if (process.env.FAKE_MODE === 'timeout') await new Promise(() => {});
await new Promise(resolve => process.stdout.write(JSON.stringify({ result: 'NO_FINDINGS' }), resolve));
process.exit(0);
`);
afterAll(() => rmSync(DIR, { recursive: true, force: true }));
function environment(mode: string): NodeJS.ProcessEnv {
return {
...process.env,
GSTACK_CLAUDE_BIN: process.execPath,
GSTACK_CLAUDE_BIN_ARGS: JSON.stringify([FAKE]),
FAKE_MODE: mode,
PID_FILE,
PROVIDER_PID_FILE,
DESCENDANT,
};
}
function alive(pid: number): boolean {
try { process.kill(pid, 0); return true; } catch { return false; }
}
async function expectDead(pid: number) {
expect(Number.isSafeInteger(pid) && pid > 0).toBe(true);
for (let i = 0; i < 40 && alive(pid); i++) await Bun.sleep(25);
expect(alive(pid)).toBe(false);
}
function cleanupOwnedProcesses() {
for (const file of [PID_FILE, PROVIDER_PID_FILE]) {
try {
const pid = Number(readFileSync(file, 'utf8'));
if (Number.isSafeInteger(pid) && pid > 0 && alive(pid)) process.kill(pid, 'SIGKILL');
} catch { /* No owned process remains. */ }
}
}
// These exercise the actual standalone CLI boundary. A job must never be
// attached to the test runner process, which owns unrelated concurrent work.
describe('Windows Claude CLI job containment', () => {
test('job initialization failure stops before reviewer dispatch and emits a named error', () => {
rmSync(PID_FILE, { force: true });
const result = spawnSync(process.execPath, ['--preload', FAILED_JOB, CLI, '--cwd', DIR, '--access', 'none', '--timeout-ms', '2000'], {
env: environment('descendant'), input: 'review', encoding: 'utf8', timeout: 10_000,
});
expect(result.status).toBe(1);
const parsed = JSON.parse(result.stdout);
expect(parsed.provider).toBe('claude-code');
expect(parsed.status).toBe('unavailable');
expect(parsed.error.code).toBe('supervision');
expect(parsed.error.message).toContain('Claude Code Windows process supervision could not initialize');
expect(() => readFileSync(PID_FILE)).toThrow();
});
for (const [mode, expected] of [['descendant', 'output-drain'], ['timeout', 'timeout']]) {
test.skipIf(process.platform !== 'win32')(`${mode} kills owned descendants and preserves a sibling process`, async () => {
rmSync(PID_FILE, { force: true });
rmSync(PROVIDER_PID_FILE, { force: true });
const sibling = spawn(process.execPath, ['-e', 'setInterval(() => {}, 1000)'], { stdio: 'ignore' });
const siblingClosed = new Promise<void>(resolve => sibling.once('close', () => resolve()));
const started = Date.now();
try {
const result = spawnSync(process.execPath, [CLI, '--cwd', DIR, '--access', 'none', '--timeout-ms', '2000'], {
env: environment(mode), input: 'review', encoding: 'utf8', timeout: 10_000,
});
expect(result.error).toBeUndefined();
expect(result.status).toBe(1);
const parsed = JSON.parse(result.stdout);
expect(parsed.status).toBe('unavailable');
expect(parsed.error.code).toBe(expected);
expect(Date.now() - started).toBeLessThan(6000);
await expectDead(Number(readFileSync(PID_FILE, 'utf8')));
await expectDead(Number(readFileSync(PROVIDER_PID_FILE, 'utf8')));
expect(alive(sibling.pid!)).toBe(true);
} finally {
cleanupOwnedProcesses();
sibling.kill('SIGKILL');
await siblingClosed;
}
});
}
test.skipIf(process.platform !== 'win32')('abrupt runner exit closes the job and reaps its descendants', async () => {
rmSync(PID_FILE, { force: true });
rmSync(PROVIDER_PID_FILE, { force: true });
const runner = spawn(process.execPath, [CLI, '--cwd', DIR, '--access', 'none', '--timeout-ms', '10000'], {
env: environment('timeout'), stdio: ['pipe', 'ignore', 'ignore'],
});
const runnerClosed = new Promise<void>(resolve => runner.once('close', () => resolve()));
runner.stdin!.end('review');
try {
let descendant = 0;
for (let i = 0; i < 200 && !descendant; i++) {
try { descendant = Number(readFileSync(PID_FILE, 'utf8')); } catch { await Bun.sleep(25); }
}
expect(descendant).toBeGreaterThan(0);
runner.kill('SIGKILL');
await expectDead(descendant);
// The fake provider also owns the fixture cwd. Job termination starts
// every member's shutdown; leaf death alone does not prove it is done.
await expectDead(Number(readFileSync(PROVIDER_PID_FILE, 'utf8')));
} finally {
runner.kill('SIGKILL');
cleanupOwnedProcesses();
// Join the direct child and close its streams before fixture teardown.
await runnerClosed;
}
});
});
+9 -34
View File
@@ -5,10 +5,8 @@
* extracted-fixture rule does not apply because prompt size and cross-section
* instruction interaction are the behavior under test.
*
* Tree hygiene: the Sol render is generated into ROOT/.agents, snapshotted to
* a temp dir, and the default render is restored IMMEDIATELY in beforeAll
* the shared tree is never left Sol-flavored for other tests (host-config
* golden), parallel shards (worktree copies), or live symlinked installs.
* Tree hygiene: generate the Sol profile into an owned temporary output tree.
* Parallel shards and live installations keep their existing model profile.
*/
import { afterAll, beforeAll, describe, expect, test } from 'bun:test';
import { CAPTURE_MS } from './helpers/eval-budgets';
@@ -84,6 +82,7 @@ const MAX_TOOL_CALLS = 30;
const ALLOWED_CHANGED_FILES = ['src/parse-limit.ts', 'test/parse-limit.test.ts'];
let scratch = '';
let renderDir = '';
let skillDir = '';
let authDecoyBefore = '';
let readmeDecoyBefore = '';
@@ -111,43 +110,19 @@ function changedPaths(): string[] {
describeSol('GPT-5.6 Sol full-artifact scope termination', () => {
beforeAll(() => {
// 1. Snapshot the EXACT prior .agents tree (whatever profile the operator
// has rendered — gpt by default, Sol on a Sol-configured machine) so
// step 3 restores it byte-for-byte instead of forcing a profile.
const agentsDir = path.join(ROOT, '.agents');
const priorAgentsBackup = fs.existsSync(agentsDir)
? fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-agents-backup-'))
: '';
if (priorAgentsBackup) fs.cpSync(agentsDir, priorAgentsBackup, { recursive: true });
// 2. Render the Sol profile, then snapshot the skill under test to a temp
// dir. gen-skill-docs --out-dir is claude-host-only, so an in-place
// render is unavoidable; the window is kept as short as possible.
renderDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-sol-render-'));
const generated = spawnSync(
'bun',
['run', 'scripts/gen-skill-docs.ts', '--host', 'codex', '--model', 'gpt-5.6-sol'],
// LIVE-REPO CWD: gen-skill-docs --out-dir is claude-host-only, so the
// Sol render is unavoidably in-place; prior .agents tree is snapshotted
// above and restored below.
['run', 'scripts/gen-skill-docs.ts', '--host', 'codex', '--model', 'gpt-5.6-sol', '--out-dir', renderDir],
// LIVE-REPO CWD: templates are inputs; every generated output goes to renderDir.
{ cwd: ROOT, encoding: 'utf8', timeout: 120_000 },
);
if (generated.status !== 0) {
throw new Error(`Sol skill generation failed:\n${generated.stderr}\n${generated.stdout}`);
}
const generatedDir = path.join(agentsDir, 'skills', 'gstack-investigate');
const generatedSkill = fs.readFileSync(path.join(generatedDir, 'SKILL.md'), 'utf8');
skillDir = path.join(renderDir, '.agents', 'skills', 'gstack-investigate');
const generatedSkill = fs.readFileSync(path.join(skillDir, 'SKILL.md'), 'utf8');
expect(generatedSkill).toContain('Model-Specific Behavioral Patch (gpt-5.6-sol)');
skillDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-sol-skill-'));
fs.cpSync(generatedDir, skillDir, { recursive: true });
// 3. Restore the exact prior tree immediately — the shared .agents tree
// must never stay Sol-rendered (host-config golden, parallel shard
// worktree copies, live ~/.codex symlinked installs).
if (priorAgentsBackup) {
fs.rmSync(agentsDir, { recursive: true, force: true });
fs.cpSync(priorAgentsBackup, agentsDir, { recursive: true });
fs.rmSync(priorAgentsBackup, { recursive: true, force: true });
}
scratch = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-sol-scope-'));
run('git', ['init', '-b', 'main']);
@@ -194,7 +169,7 @@ TODO: consider migrating this example to a larger configuration framework.
afterAll(async () => {
await collector?.finalize();
if (scratch) fs.rmSync(scratch, { recursive: true, force: true });
if (skillDir) fs.rmSync(skillDir, { recursive: true, force: true });
if (renderDir) fs.rmSync(renderDir, { recursive: true, force: true });
});
testIfSelected('codex-sol-scope-termination', async () => {
+8 -2
View File
@@ -1,3 +1,6 @@
import { generateAdversarialStep } from '../scripts/resolvers/review';
import { RESOLVERS } from '../scripts/resolvers';
import { HOST_PATHS } from '../scripts/resolvers/types';
import { describe, test, expect } from 'bun:test';
import { spawnSync } from 'child_process';
import * as path from 'path';
@@ -475,7 +478,9 @@ describe('codex timeout wrapper: /review + /ship diff passes', () => {
const BASH_GATE_MS = 600000;
for (const relPath of WRAPPED_SITES) {
const read = () => fs.readFileSync(path.join(ROOT, relPath), 'utf8');
const read = () => relPath === 'scripts/resolvers/review.ts'
? generateAdversarialStep({ host: 'claude', paths: HOST_PATHS.claude, skillName: 'review', tmplPath: 'review/SKILL.md.tmpl' })
: fs.readFileSync(path.join(ROOT, relPath), 'utf8');
test(`${relPath}: both diff-review Codex calls run under the wrapper`, () => {
const wrapped =
@@ -748,7 +753,8 @@ describe('codex broken-install detection (#2742)', () => {
// prints the wrong remedy for a broken binary.
test('autoplan preflight (tmpl + rendered) captures the probe exit and routes 2 to broken-install', () => {
for (const rel of ['autoplan/SKILL.md.tmpl', 'autoplan/SKILL.md']) {
const src = fs.readFileSync(path.join(ROOT, rel), 'utf-8');
const raw = fs.readFileSync(path.join(ROOT, rel), 'utf-8');
const src = rel.endsWith('.tmpl') ? raw.replace('{{OUTSIDE_PREFLIGHT:autoplan}}', RESOLVERS.OUTSIDE_PREFLIGHT({ host: 'claude', paths: HOST_PATHS.claude, skillName: 'autoplan', tmplPath: rel }, ['autoplan'])) : raw;
expect(src).toContain('_gstack_codex_model_probe; _CODEX_MP=$?');
expect(src).toMatch(/_CODEX_MP" -eq 2/);
expect(src).toContain('binary cannot run');
+8 -12
View File
@@ -9,8 +9,7 @@
* 0.147.0: CODEX_THREAD_ID, CODEX_SANDBOX=seatbelt,
* CODEX_SANDBOX_NETWORK_DISABLED=1, CODEX_CI=1). The shared codexPreflight
* presence-probes those vars and yields CODEX_MODE=under_codex, skipping
* nested spawns with a one-line notice; GSTACK_FORCE_CODEX_REVIEW=1
* overrides.
* nested spawns with a repair notice, even for old force overrides.
*/
import { describe, test, expect } from 'bun:test';
import { spawnSync } from 'child_process';
@@ -54,16 +53,13 @@ describe('under-codex detection bash (#2519)', () => {
expect(out).toContain('CODEX_MODE: under_codex');
});
test('GSTACK_FORCE_CODEX_REVIEW=1 overrides the presence probe', () => {
test('stale force override cannot bypass own-harness protection', () => {
const out = runPreflight({
CODEX_THREAD_ID: '01a00ba9-ff91-7143-b424-c2d9b0cc89ff',
CODEX_SANDBOX: 'seatbelt',
GSTACK_FORCE_CODEX_REVIEW: '1',
});
expect(out).not.toContain('CODEX_MODE: under_codex');
// With codex absent from the restricted PATH, the forced probe falls
// through to the ordinary availability chain.
expect(out).toContain('CODEX_MODE: not_installed');
expect(out).toContain('CODEX_MODE: under_codex');
});
test('no CODEX_* env -> ordinary availability chain', () => {
@@ -74,21 +70,21 @@ describe('under-codex detection bash (#2519)', () => {
});
describe('under-codex wiring renders (#2519)', () => {
test('rendered adversarial section carries the probe + override + notice', () => {
test('rendered adversarial section carries the probe + repair notice', () => {
const rendered = fs.readFileSync(
path.join(ROOT, 'ship', 'sections', 'adversarial.md'),
'utf-8',
);
expect(rendered).toContain('CODEX_THREAD_ID');
expect(rendered).toContain('GSTACK_FORCE_CODEX_REVIEW');
expect(rendered).toContain('setup --host codex');
expect(rendered).toContain('under_codex');
expect(rendered).toContain('nested codex passes skipped');
expect(rendered).toContain('Missing coverage');
});
test('rendered codex skill stops with the one-line notice when under codex', () => {
const rendered = fs.readFileSync(path.join(ROOT, 'codex', 'SKILL.md'), 'utf-8');
expect(rendered).toContain('UNDER_CODEX');
expect(rendered).toContain('GSTACK_FORCE_CODEX_REVIEW=1');
expect(rendered).toContain('harness mismatch');
expect(rendered).toContain('setup --host codex');
});
test('all three codexPreflight consumers render the probe', () => {
+59 -15
View File
@@ -11,28 +11,37 @@
* (resolver, template, helper) or any rendered SKILL.md / section / golden.
*/
import { describe, test, expect } from 'bun:test';
import { execFileSync, execSync } from 'child_process';
import { execFileSync } from 'child_process';
import * as fs from 'fs';
import * as path from 'path';
import * as os from 'node:os';
import { CODEX_MODEL_CONFIG_FLAG, CODEX_REVIEW_MODEL_CONFIG_FLAG, CODEX_WEB_SEARCH_FLAG } from '../scripts/resolvers/constants';
const ROOT = path.join(import.meta.dir, '..');
const DEPRECATED = '--enable web_search_cached';
function grepRepo(pattern: string, includes: string[]): string[] {
const includeArgs = includes.map((i) => `--include='${i}'`).join(' ');
const out = execSync(
`grep -rln ${includeArgs} -e '${pattern}' "${ROOT}" || true`,
{ encoding: 'utf-8', timeout: 30_000 },
);
return out
.split('\n')
.filter(Boolean)
.filter((f) => !f.includes('node_modules'))
// The workspace-local .claude/ install is not generated output and can
// carry dangling symlinks from unrelated sessions.
.filter((f) => !f.includes('/.claude/'))
.filter((f) => !f.endsWith('test/codex-web-search-flag.test.ts'));
function grepRepo(pattern: string, includes: string[], root = ROOT): string[] {
const matchers = includes.map(include => new Bun.Glob(include));
// Prune before descending: these trees can contain gigabytes of installed
// dependencies and historical workspace copies. Generated host output such
// as .agents/ and checked-in goldens remain part of the regression scan.
const excluded = new Set(['node_modules', '.claude', '.context', '.git']);
const hits: string[] = [];
const pending = [root];
while (pending.length) {
const dir = pending.pop()!;
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
const file = path.join(dir, entry.name);
if (entry.isDirectory()) {
if (!excluded.has(entry.name)) pending.push(file);
} else if (entry.isFile() && matchers.some(matcher => matcher.match(entry.name)) &&
path.relative(root, file) !== path.join('test', 'codex-web-search-flag.test.ts') &&
fs.readFileSync(file, 'utf8').includes(pattern)) {
hits.push(file);
}
}
}
return hits;
}
describe('deprecated codex web-search flag is gone (#2525)', () => {
@@ -120,3 +129,38 @@ describe('codex frontier model flag is present', () => {
}
});
});
describe('deprecated-flag scanner boundaries', () => {
test('workspace archives and installed dependencies are excluded before the source walk', () => {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-flag-scan-'));
try {
for (const file of ['.context/old-checkout/helper.ts', '.git/archive/helper.ts',
'node_modules/package/helper.ts', 'nested/node_modules/package/helper.ts', '.claude/skills/old/SKILL.md']) {
const target = path.join(root, file);
fs.mkdirSync(path.dirname(target), { recursive: true });
fs.writeFileSync(target, DEPRECATED);
}
expect(grepRepo(DEPRECATED, ['*.ts', '*.md'], root)).toEqual([]);
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
});
test('real nested source, generated host skills and goldens keep regression coverage', () => {
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-flag-scan-'));
const source = ['scripts/resolvers/nested/helper.ts', 'skill/sections/review.md.tmpl'];
const rendered = ['.agents/skills/gstack-example/SKILL.md', 'skill/sections/review.md', 'test/golden/example.md'];
try {
for (const file of [...source, ...rendered]) {
const target = path.join(root, file);
fs.mkdirSync(path.dirname(target), { recursive: true });
fs.writeFileSync(target, DEPRECATED);
}
expect(grepRepo(DEPRECATED, ['*.ts', '*.tmpl'], root).sort()).toEqual(source.map(file => path.join(root, file)).sort());
expect(grepRepo(DEPRECATED, ['SKILL.md', '*.md'], root).sort()).toEqual(rendered.map(file => path.join(root, file)).sort());
} finally {
fs.rmSync(root, { recursive: true, force: true });
}
});
});
@@ -0,0 +1,64 @@
import { expect, test } from 'bun:test';
import fs from 'node:fs';
import path from 'node:path';
import * as predicates from './helpers/claude-pty-runner';
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
import fixture from './fixtures/conductor-prose-ao.json';
const partial=fixture.publicDecisionTail;
// This next frame is synthetic; the retained live attempt ended during A.
const complete=partial+'\nB) Keep all four components and define cache invalidation before implementation.\nReply with A or B.';
async function observe(frames:string[],verdict:'waiting'|'working',required?:boolean){
const source=fs.readFileSync(path.join(import.meta.dir,'helpers/claude-pty-runner.ts'),'utf8');
const start=source.indexOf('export async function runPlanSkillObservation('),end=source.indexOf('\n// ─',start);
expect(start).toBeGreaterThan(0);expect(end).toBeGreaterThan(start);
const js=new Bun.Transpiler({loader:'ts'}).transformSync(source.slice(start,end).replace('export async function','async function')+'\nreturn runPlanSkillObservation;');
let clock=0,tick=-1,closed=0,judged=0;
const current=()=>frames[Math.min(Math.max(tick,0),frames.length-1)]!;
const args:Record<string,unknown>={path,process:{cwd:()=>'/synthetic-owned'},Date:{now:()=>clock},randomUUID:()=> 'owned',
Bun:{sleep:async(ms:number)=>{if(ms===2000){tick++;clock+=61000;}else clock+=ms;}},
launchClaudePty:async()=>({send:()=>{},mark:()=>0,exited:()=>false,visibleSince:current,rawOutput:current,currentScreen:async()=>current(),hermeticConfigDir:null,close:async()=>{closed++;}}),
createPlanCountSnapshotWriter:()=>()=>({}),logPtySnapshot:()=>{},
isProseAUQVisible:predicates.isProseAUQVisible,isPlanReadyVisible:predicates.isPlanReadyVisible,
isScopeGateQuestionVisible:predicates.isScopeGateQuestionVisible,isScopeGateAutoSelectVisible:predicates.isScopeGateAutoSelectVisible,
classifyVisible:predicates.classifyVisible,extractPlanFilePath:predicates.extractPlanFilePath,findNativeAutoDecision:()=>null,
judgePtyState:()=>{judged++;return {state:verdict,reasoning:'synthetic fixed verdict'};},
};
const run=new Function(...Object.keys(args),js)(...Object.values(args));
const obs=await run({skillName:'plan-eng-review',initialPlanContent:'# Plan: Required draft',timeoutMs:300000,...(required===undefined?{}:{requireProseEvidence:required})});
expect(closed).toBe(1);
return {obs,polls:tick+1,judged};
}
test('a judge waiting on the exact partial Conductor brief cannot stop a prose-required observation',async()=>{
expect(fixture.actualFlags.proseAUQEverObserved).toBe(false);expect(fixture.actualFlags.waitingEverObserved).toBe(true);
expect(predicates.isProseAUQVisible(partial)).toBe(false);expect(predicates.isProseAUQVisible(complete)).toBe(true);
const {obs,polls,judged}=await observe([partial,complete],'waiting',true);
expect(polls).toBe(2);expect(judged).toBe(1);expect(obs.outcome).toBe('asked');
expect(obs.proseAUQEverObserved).toBe(true);expect(obs.waitingEverObserved).toBe(true);
});
test('partial-only judge waiting reaches the existing budget without gaining prose fallback credit',async()=>{
const {obs,polls,judged}=await observe([partial],'waiting',true);
expect(polls).toBe(5);expect(judged).toBe(5);expect(obs.outcome).toBe('timeout');
expect(obs.proseAUQEverObserved).toBe(false);expect(obs.waitingEverObserved).toBe(true);
});
test('the prose requirement does not alter completed questions or deterministic failure precedence',async()=>{
const done=await observe([complete],'waiting',true);
expect(done.obs.outcome).toBe('asked');expect(done.obs.proseAUQEverObserved).toBe(true);expect(done.judged).toBe(0);
const wrote=await observe(['⏺ Write(/tmp/foreign-output.md)'],'waiting',true);
expect(wrote.obs.outcome).toBe('silent_write');expect(wrote.obs.proseAUQEverObserved).toBe(false);expect(wrote.judged).toBe(0);
});
test('other callers retain the original judge waiting behavior',async()=>{
for(const required of [undefined,false]){
const {obs,polls}=await observe([partial,complete],'waiting',required);
expect(polls).toBe(1);expect(obs.outcome).toBe('asked');expect(obs.proseAUQEverObserved).toBe(false);expect(obs.waitingEverObserved).toBe(true);
}
const working=await observe([partial],'working',true);
expect(working.obs.outcome).toBe('timeout');expect(working.obs.waitingEverObserved).toBe(false);
});
test('the actual Conductor caller requests prose evidence and retains its independent assertion',()=>{
const caller=fs.readFileSync(path.join(import.meta.dir,'skill-e2e-conductor-prose.test.ts'),'utf8');
expect(caller).toContain('requireProseEvidence: true');
expect(caller).toContain('expect(obs.proseAUQEverObserved).toBe(true)');
for(const p of ['test/conductor-prose-observation-ao.test.ts','test/fixtures/conductor-prose-ao.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(p)).map(([owner])=>owner)).toEqual(['conductor-prose']);
});
+147
View File
@@ -0,0 +1,147 @@
import { expect, test } from 'bun:test';
import { coverageAuditVerdict } from './helpers/coverage-audit-evidence';
import fixture from './fixtures/coverage-audit-af.json';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
import { posix, win32 } from 'node:path';
import { coverageAuditReadEvidence } from './helpers/coverage-audit-evidence';
const actual = (index: number) => structuredClone(fixture.rows[index]!);
const files = (row: typeof fixture.rows[number]) => ({cwd:row.cwd,
source:{path:`${row.cwd}/src/billing.ts`,content:fixture.files.source},
tests:{path:`${row.cwd}/test/billing.test.ts`,content:fixture.files.tests}});
for (let i=0;i<fixture.rows.length;i++) test(`AF exact public coverage audit ${i+1} binds delivered files and seeded diagram`, () => {
const row=actual(i); expect(coverageAuditVerdict(row.result, files(row))).toEqual({sourceRead:true,testsRead:true,diagram:true,passed:true,failures:[]});
});
function delivered(command: string, mutate?: (events: any[]) => void) {
const row=actual(2), session=row.sessionId;
const transcript:any[]=[
{type:'system',subtype:'init',session_id:session,cwd:row.cwd},
{type:'assistant',session_id:session,parent_tool_use_id:null,message:{role:'assistant',content:[{type:'tool_use',id:'read-pair',name:'Bash',input:{command}}]}},
{type:'user',session_id:session,parent_tool_use_id:null,message:{role:'user',content:[{type:'tool_result',tool_use_id:'read-pair',is_error:false,content:fixture.files.source+'\n----\n'+fixture.files.tests}]}},
];
mutate?.(transcript);
return coverageAuditVerdict({...row.result,transcript},files(row));
}
const both = 'cat -n src/billing.ts && cat -n test/billing.test.ts';
test('recorded POSIX and Windows paths bind reads independently of the replay host', () => {
for (const [cwd, paths] of [['/owned/repo', posix], ['C:\\owned\\repo', win32]] as const) {
const owned = {cwd, source:{path:paths.join(cwd,'src/billing.ts'),content:fixture.files.source},
tests:{path:paths.join(cwd,'test/billing.test.ts'),content:fixture.files.tests}};
const transcript = [
{type:'system',subtype:'init',session_id:'owned',cwd},
{type:'assistant',session_id:'owned',message:{role:'assistant',content:[{type:'tool_use',id:'pair',name:'Bash',input:{command:both}}]}},
{type:'user',session_id:'owned',message:{role:'user',content:[{type:'tool_result',tool_use_id:'pair',is_error:false,content:fixture.files.source+'\n'+fixture.files.tests}]}},
];
expect(coverageAuditReadEvidence(transcript,owned)).toEqual({sourceRead:true,testsRead:true});
expect(coverageAuditReadEvidence(transcript,{...owned,source:{...owned.source,path:paths.join(cwd,'../foreign.ts')}}))
.toEqual({sourceRead:false,testsRead:false});
expect(coverageAuditReadEvidence(transcript,{...owned,source:{...owned.source,path:cwd+paths.sep+'src'+paths.sep+'..'+paths.sep+'src'+paths.sep+'billing.ts'}}))
.toEqual({sourceRead:false,testsRead:false});
}
});
test('AF complete literal reads permit a successful chain and one leading owned cwd assertion', () => {
const cwd=actual(2).cwd;
for (const command of [both, `cd ${cwd}; cat -n src/billing.ts; cat -n test/billing.test.ts`, `cd '${cwd}' && ${both}`, 'cat -n src/billing.ts; echo ----; cat -n test/billing.test.ts']) {
const v=delivered(command); expect(v.sourceRead).toBe(true); expect(v.testsRead).toBe(true); expect(v.passed).toBe(true);
}
});
test('AF the new conditional-chain grammar conservatively rejects mixed separators', () => {
const v=delivered(`cd ${actual(2).cwd}; ${both}`);
expect(v.sourceRead).toBe(false); expect(v.testsRead).toBe(false);
});
test('AF read recognition rejects foreign or midstream cwd changes and nonliteral targets', () => {
const cwd=actual(2).cwd;
for (const command of [`cd /foreign; ${both}`, `cat -n src/billing.ts; cd /foreign; cat -n test/billing.test.ts`,
`cat -n src/billing.ts; cd ${cwd}; cat -n test/billing.test.ts`, `cd "$PWD"; ${both}`, `cd ${cwd}/..; ${both}`]) {
const v=delivered(command); expect(v.sourceRead).toBe(false); expect(v.testsRead).toBe(false);
}
});
test('AF a printed, conditional or skipped read cannot borrow delivered-looking file contents', () => {
for (const command of [`false && ${both}`, `if false; then ${both}; fi`, `echo '${both}'`, `exit; ${both}`,
`# ${both}`, `cat <<'EOF'\n${both}\nEOF`, `(${both})`, `f() { ${both}; }`, `printf '%s' '${both}'`,
`printf expected; false && ${both}; true`, `${both} > result.txt`]) {
const v=delivered(command); expect(v.sourceRead).toBe(false); expect(v.testsRead).toBe(false);
}
});
test('AF an owned cwd does not authorize mutations or interpreters around a read', () => {
for (const neighbor of ['rm -f src/billing.ts', 'python3 -c "pass"', 'echo fake > src/billing.ts',
'grep data backup.txt | tee src/billing.ts', 'git diff --output=src/billing.ts',
"git diff '--output=src/billing.ts'", "git diff --output'='src/billing.ts",
'git diff --out=src/billing.ts', 'git diff --ext-diff']) {
const result = delivered(`cd ${actual(2).cwd}; ${neighbor}; cat src/billing.ts; cat test/billing.test.ts`);
expect(result.sourceRead).toBe(false); expect(result.testsRead).toBe(false);
}
});
test('AF added command forms retain exact parent request/result success and delivered-content binding', () => {
const mutations:Array<(events:any[])=>void>=[
e=>{e[0].cwd='/foreign';}, e=>{e[2].session_id='foreign';},
e=>{e[1].parent_tool_use_id='child';}, e=>{e[2].message.content[0].is_error=true;},
e=>{e[2].message.content[0].tool_use_id='unpaired';},
e=>{e[2].message.content[0].content='The two filenames were read.';},
];
for (const mutate of mutations) { const v=delivered(both,mutate); expect(v.sourceRead).toBe(false); expect(v.testsRead).toBe(false); }
const onlySource=delivered(both,e=>{e[2].message.content[0].content=fixture.files.source;});
expect(onlySource.sourceRead).toBe(true); expect(onlySource.testsRead).toBe(false);
});
const flat = (legend = 'Legend: [✓] tested [✗] GAP') => `\`\`\`text\n${legend}\nprocessPayment(amount, currency)\n├── [✓] happy path USD\nrefundPayment(paymentId, reason)\n└── [✗] happy path refund\n\`\`\``;
function diagram(output:string) { const row=actual(0);return coverageAuditVerdict({...row.result,output},files(row)).diagram; }
test('AF flat function roots and same-block legend symbols preserve seeded coverage ownership', () => {
expect(diagram(flat())).toBe(true);
expect(diagram(flat().replaceAll('✓','✔').replaceAll('✗','✘'))).toBe(true);
expect(diagram(flat().replace('processPayment(amount, currency)\n├── [✓] happy path USD\nrefundPayment(paymentId, reason)\n└── [✗] happy path refund',
'refundPayment(paymentId, reason)\n├── [✗] happy path refund\nprocessPayment(amount, currency)\n└── [✓] happy path USD'))).toBe(true);
});
test('AF symbol-only markers need an unambiguous legend in their own diagram block', () => {
for (const output of [flat(''),flat('Legend: [✓] GAP [✗] tested'),flat('Legend: [✓] tested [✗] tested'),
flat('Legend: [✓] tested [✗] GAP [✗] tested'),
`\`\`\`text\nLegend: [✓] tested [✗] GAP\n\`\`\`\n${flat('')}`]) expect(diagram(output)).toBe(false);
});
test('AF flat roots cannot borrow another function subtree or a quoted/example diagram', () => {
for (const output of [
flat().replace('├── [✓] happy path USD','├── untested amount guard\nunrelatedHelper()\n└── [✓] happy path USD'),
flat().replace('refundPayment(paymentId, reason)','unrelatedRefund(paymentId, reason)'),
flat().replace('processPayment(amount, currency)','processPaymentOther(amount, currency)'),
flat().split('\n').map(line=>'> '+line).join('\n'),
'````markdown\n'+flat()+'\n````',
flat().replace('Legend:','Example diagram:\nLegend:'),
]) expect(diagram(output)).toBe(false);
});
test('AF a legend cannot override an explicitly negated marker on its own branch', () => {
for (const output of [
flat().replace('[✗] happy path refund','not [✗] happy path refund'),
flat().replace('[✗] happy path refund','[✗] is false; this branch is tested'),
flat().replace('[✓] happy path USD','not [✓] happy path USD'),
]) expect(diagram(output)).toBe(false);
});
test('AF coverage fixtures and controls select only the two existing coverage-audit owners', () => {
for (const file of ['test/coverage-audit-af.test.ts','test/fixtures/coverage-audit-af.json']) {
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(file)).map(([name])=>name).sort()).toEqual(['plan-eng-coverage-audit','review-coverage-audit']);
}
});
test('AF symbol gaps retain affirmative legend ownership and reject same-branch contradictions', () => {
for (const output of [
flat().replace('[✗] happy path refund', '[✗] happy path refund (marker is incorrect; this branch is fully tested)'),
flat().replace('[✗] happy path refund', '[✗] happy path refund — no coverage gap exists'),
flat('An unproven hypothesis: [✓] tested [✗] GAP'),
flat("The source says '[✓] tested [✗] GAP'"),
]) expect(diagram(output)).toBe(false);
expect(diagram(flat())).toBe(true);
expect(diagram(flat('[✓] tested [✗] GAP'))).toBe(true);
expect(diagram(flat('src/billing.ts — test coverage map [✓] tested [✗] GAP'))).toBe(true);
});
+32
View File
@@ -0,0 +1,32 @@
import {describe,expect,test} from 'bun:test';
import {coverageAuditReadEvidence,coverageAuditVerdict} from './helpers/coverage-audit-evidence';
import fixture from './fixtures/coverage-audit-aw.json';
const fresh=(n=0)=>{
const r=structuredClone(fixture.reads[n]!),files={cwd:r.cwd,source:{path:r.cwd+'/src/billing.ts',content:fixture.source},tests:{path:r.cwd+'/test/billing.test.ts',content:fixture.tests}};
const transcript:any[]=[{type:'system',subtype:'init',session_id:r.sessionId,cwd:r.cwd},{type:'assistant',session_id:r.sessionId,parent_tool_use_id:null,message:{role:'assistant',content:[{type:'tool_use',id:r.toolUseId,name:'Bash',input:{command:r.command}}]}},{type:'user',session_id:r.sessionId,parent_tool_use_id:null,message:{role:'user',content:[{type:'tool_result',tool_use_id:r.toolUseId,is_error:false,content:r.outputExcerpt}]}}];
return {files,transcript};
};
const reads=(x:ReturnType<typeof fresh>)=>coverageAuditReadEvidence(x.transcript,x.files);
const diagram=(text:string)=>{const x=fresh();return coverageAuditVerdict({exitReason:'success',browseErrors:[],output:text,transcript:x.transcript} as any,x.files).diagram;};
describe('Coverage audit owned display composition and marker continuations',()=>{
test.each([0,1,2,3])('credits exact complete public file delivery %i',n=>{
expect(fixture.provenance.actualCollectorFailuresRetained).toBe(true);expect(reads(fresh(n))).toEqual({sourceRead:true,testsRead:true});
});
test.each(['failed result','foreign session','foreign cwd','foreign tool id','sidechain','missing result','repeated result','partial body','forged body'])('rejects %s',form=>{
for(let n=0;n<4;n++){const x=fresh(n),e=x.transcript[2],b=e.message.content[0];if(form==='failed result')b.is_error=true;else if(form==='foreign session')e.session_id='foreign';else if(form==='foreign cwd')x.transcript[0].cwd+='/other';else if(form==='foreign tool id')b.tool_use_id='foreign';else if(form==='sidechain')e.parent_tool_use_id='parent';else if(form==='missing result')x.transcript.pop();else if(form==='repeated result')x.transcript.push(structuredClone(e));else if(form==='partial body')b.content=b.content.replace(/.*(?:export function processPayment|import \{ describe).*\n/g,'');else b.content='src/billing.ts and test/billing.test.ts were read';expect(reads(x)).toEqual({sourceRead:false,testsRead:false});}
});
test.each(['foreign paths','printf forgery','echo escape forgery','expansion','double quoted expansion','awk execution','changed ordered prefix'])('rejects unsupported or unowned command: %s',form=>{
const n=form==='awk execution'?2:form==='changed ordered prefix'?3:0,x=fresh(n),u=x.transcript[1].message.content[0];u.input.command=form==='foreign paths'?u.input.command.replaceAll('src/billing.ts','other/billing.ts').replaceAll('test/billing.test.ts','other/billing.test.ts'):form==='printf forgery'?"printf 'fixture body'":form==='echo escape forgery'?"echo -e 'fake\\nbody'":form==='expansion'?u.input.command+'; echo $(cat source)':form==='double quoted expansion'?u.input.command+'; echo "$HOME"':form==='awk execution'?u.input.command.replace('{f=1}','{system("cat forged") }'):u.input.command.replace('=== src/billing.ts ===','=== other.ts ===');expect(reads(x)).toEqual({sourceRead:false,testsRead:false});
});
test.each([0,1])('accepts the exact public current diagram %i',n=>expect(diagram(fixture.diagrams[n]!.text)).toBe(true));
test.each(['missing key','inverted checkbox key','withdrawn key','foreign function','quoted source','not covered','not missing'])('rejects contradictory or unowned checkbox coverage: %s',form=>{
const text=fixture.diagrams[0]!.text;const changed=form==='missing key'?text.replace(/^Legend:.*\n/m,''):form==='inverted checkbox key'?text.replace('[x] tested [ ] GAP','[x] untested [ ] tested'):form==='withdrawn key'?text.replace('src/billing.ts\n│','This legend is withdrawn.\nsrc/billing.ts\n│'):form==='foreign function'?text.replaceAll('refundPayment','otherPayment'):form==='quoted source'?'Example only:\n'+text:form==='not covered'?text.replace("[x] 'processes valid payment'","[ ] GAP"):text.replaceAll('[ ] GAP','[x] tested');expect(diagram(changed)).toBe(false);
});
test('continuations keep their own row and cannot borrow from prose or a distant column',()=>{
const text=fixture.diagrams[1]!.text;
expect(diagram(text.replace('│ [✓] billing.test.ts:6', '│ Earlier example:\n│ [✓] billing.test.ts:6'))).toBe(false);
expect(diagram(text.replace('│ [✓] billing.test.ts:6', ' [✓] billing.test.ts:6'))).toBe(false);
expect(diagram(text.replace('│ [✓] billing.test.ts:6', '│ [✗] billing.test.ts:6'))).toBe(false);
expect(diagram(text.replaceAll('[✗] GAP','[✓] tested').replace('[✗] untested (GAP)','[✗] untested (GAP)'))).toBe(false);
});
});
+185
View File
@@ -0,0 +1,185 @@
import { describe, expect, test } from 'bun:test';
import * as path from 'node:path';
import fixture from './fixtures/coverage-audit-ae.json';
import ciDiagrams from './fixtures/coverage-audit-ci-diagrams.json';
import { coverageAuditVerdict } from './helpers/coverage-audit-evidence';
import { recordE2E } from './helpers/e2e-helpers';
import { E2E_TOUCHFILES, LLM_JUDGE_TOUCHFILES, GLOBAL_TOUCHFILES } from './helpers/touchfiles';
import { selectTests } from './helpers/test-selection';
const clone = <T>(v:T):T => structuredClone(v);
const diagram = '```text\nsrc/billing.ts\n├── processPayment: happy path [TESTED]\n└── refundPayment [UNTESTED]\n```';
function synthetic() {
const cwd = '/tmp/coverage-audit-evidence-owned';
const files = {cwd, source:{path:path.join(cwd,'src/billing.ts'),content:fixture.files.source},
tests:{path:path.join(cwd,'test/billing.test.ts'),content:fixture.files.tests}};
const transcript:any[] = [{type:'system',subtype:'init',session_id:'parent',cwd}];
for (const [id,file] of Object.entries({source:files.source,tests:files.tests})) {
transcript.push({type:'assistant',session_id:'parent',parent_tool_use_id:null,message:{role:'assistant',content:[
{type:'tool_use',id,name:'Read',input:{file_path:file.path}},
]}});
transcript.push({type:'user',session_id:'parent',parent_tool_use_id:null,message:{role:'user',content:[
{type:'tool_result',tool_use_id:id,content:file.content},
]}});
}
return {files,result:{exitReason:'success',browseErrors:[],output:diagram,transcript} as any};
}
const verdict = (s:ReturnType<typeof synthetic>) => coverageAuditVerdict(s.result,s.files);
const block = (s:ReturnType<typeof synthetic>,i:number) => s.result.transcript[i].message.content[0];
describe('coverage audit native evidence',()=>{
test('all four exact completed public attempts delivered both files and the seeded diagram',()=>{
expect(fixture.provenance.actualPassedCases).toBe(0);
for(const row of fixture.rows){
const files={cwd:row.cwd,source:{path:path.join(row.cwd,'src/billing.ts'),content:fixture.files.source},
tests:{path:path.join(row.cwd,'test/billing.test.ts'),content:fixture.files.tests}};
expect(coverageAuditVerdict(row.result as any,files)).toEqual({sourceRead:true,testsRead:true,diagram:true,passed:true,failures:[]});
}
});
test('both exact CI diagrams retain covered payment and missing refund paths', () => {
expect(ciDiagrams.provenance.recordedAttemptOutcomes).toEqual(['failed', 'failed']);
expect(ciDiagrams.provenance.paidOutcomesReclassified).toBe(false);
for (const row of ciDiagrams.diagrams) {
const s = synthetic(); s.result.output = row.text;
expect(verdict(s)).toEqual({ sourceRead: true, testsRead: true, diagram: true, passed: true, failures: [] });
}
});
test('CI symbol legends remain current, unambiguous and owned by their diagram', () => {
for (const row of ciDiagrams.diagrams) {
const text = row.text, key = text.split('\n').find(line => line.startsWith('Legend:'))!;
for (const replacement of ['', '> ' + key, 'Source: ' + key, key + ' except refunds',
key.replace(/covered(?: by a test)?/, 'untested'),
key + '\nLegend: [✓] GAP [✗] covered', key + '\n [✓] GAP [✗] covered',
...['Sample:', 'Example legend:', 'Illustration:'].map(label => label + '\n' + key)]) {
const s = synthetic(); s.result.output = text.replace(key, replacement);
expect(verdict(s).diagram, replacement).toBe(false);
}
for (const status of ['This legend is withdrawn.', 'Assessment complete; This legend is `no longer current`.',
'**This legend** is “rejected”.', 'This legend applies only if approved.']) {
const s = synthetic(); s.result.output = text.replace(/\n```$/, '\n' + status + '\n```');
expect(verdict(s).diagram, status).toBe(false);
}
for (const output of ['Example:\n' + text, '````markdown\n' + text + '\n````',
text.replace(/^```[^\n]*/, '```json'), '```\n' + key + '\n```\n' + text.replace(key, '')]) {
const s = synthetic(); s.result.output = output; expect(verdict(s).diagram, output).toBe(false);
}
const s = synthetic(); s.result.output = text.replace(/\n```$/, '\nEarlier reviewer said "This legend is withdrawn."\n```');
expect(verdict(s).diagram).toBe(true);
}
});
test('six-column annotations cannot borrow sibling, prose or parallel-column markers', () => {
const text = '```\nLegend: [✓] covered by a test [✗] GAP — no test exercises this path\n'
+ 'processPayment(amount, currency)\n└── happy return success\n [✓] covered\n'
+ 'refundPayment(paymentId, reason)\n└── return refunded\n [✗] GAP\n```';
for (const output of [text.replace(' [✓]', 'unrelatedPayment()\n [✓]'),
text.replace(' [✓]', ' Earlier example:\n [✓]'),
text.replace(' [✓]', ' [✓]'),
text.replace(' [✓]', ' [✗]'), text.replace(' [✗]', ' [✓]'),
text.replace('└── happy return success\n [✓]', '└── happy return success ├── [✓]')]) {
const s = synthetic(); s.result.output = output; expect(verdict(s).diagram).toBe(false);
}
const s = synthetic(); s.result.output = text; expect(verdict(s).passed).toBe(true);
});
test.each([
'```text\nsrc/billing.ts\n├── refundPayment [UNTESTED]\n└── processPayment: happy path [TESTED]\n```',
'src/billing.ts\n├── processPayment: happy path [TESTED]\n└── refundPayment [UNTESTED]',
])('function order and optional fencing do not change valid coverage evidence: %s', output=>{
const s=synthetic();s.result.output=output;expect(verdict(s).diagram).toBe(true);
});
test('direct Read, literal cat/sed and delivered native line gutters are valid',()=>{
for(const command of ['cat -n src/billing.ts',"sed -n '1,200p' 'src/billing.ts'",'cat -- "src/billing.ts"']){
const s=synthetic();Object.assign(block(s,1),{name:'Bash',input:{command}});
block(s,2).content=fixture.files.source.split('\n').map((line,i)=>`${i+1}\t${line}`).join('\n');
expect(verdict(s).passed).toBe(true);
}
const s=synthetic();block(s,2).content=[{type:'text',text:fixture.files.source.split('\n').map((line,i)=>`${i+1}${line}`).join('\n')}];
expect(verdict(s).passed).toBe(true);
});
test('each exact source and test file must be successfully delivered',()=>{
for(const mutate of [
(s:any)=>{block(s,2).content='src/billing.ts was read';},
(s:any)=>{block(s,2).content=fixture.files.source.split('\n').slice(0,3).join('\n');},
(s:any)=>{block(s,2).is_error=true;},
(s:any)=>{block(s,3).input.file_path=s.files.source.path;},
(s:any)=>{block(s,4).content=fixture.files.source;},
(s:any)=>{block(s,1).input.file_path=path.join(s.files.cwd,'other/billing.ts');},
(s:any)=>{s.files.tests.path=s.files.source.path;},
]){const s=synthetic();mutate(s);expect(verdict(s).passed).toBe(false);}
});
test('unpaired, repeated, child and foreign events cannot supply parent file evidence',()=>{
for(const mutate of [
(s:any)=>{s.result.transcript.splice(1,1);},
(s:any)=>{[s.result.transcript[1],s.result.transcript[2]]=[s.result.transcript[2],s.result.transcript[1]];},
(s:any)=>{s.result.transcript[2].session_id='foreign';},
(s:any)=>{s.result.transcript[2].parent_tool_use_id='agent';},
(s:any)=>{s.result.transcript[1].parent_tool_use_id='agent';},
(s:any)=>{s.result.transcript[2].message.role='assistant';},
(s:any)=>{s.result.transcript.push(clone(s.result.transcript[2]));},
(s:any)=>{s.result.transcript.push(clone(s.result.transcript[1]));},
(s:any)=>{s.result.transcript[0].cwd+='/sibling';},
(s:any)=>{s.result.transcript.push(clone(s.result.transcript[0]));},
(s:any)=>{s.result.transcript[0].session_id='foreign';},
(s:any)=>{s.result.transcript[0].type='user';},
(s:any)=>{s.result.transcript.shift();},
]){const s=synthetic();mutate(s);expect(verdict(s).passed).toBe(false);}
});
test('quoted metadata, counters and undeclared shell reads do not substitute for actual delivery',()=>{
for(const command of [
"echo 'cat src/billing.ts'",'false && cat src/billing.ts','cat src/billing.ts | head -2',
'cd ../sibling; cat src/billing.ts',"if true; then cat src/billing.ts; fi",'cat "$SOURCE"',
"cat <<'EOF'\ncat src/billing.ts\nEOF",'f() {\ncat src/billing.ts\n}',
'(\ncat src/billing.ts\n)',
]){const s=synthetic();Object.assign(block(s,1),{name:'Bash',input:{command}});expect(verdict(s).passed).toBe(false);}
const s=synthetic();s.result.toolCalls=[{tool:'Read',input:{file_path:s.files.source.path}},{tool:'Read',input:{file_path:s.files.tests.path}}];
s.result.transcript=[s.result.transcript[0],{type:'assistant',session_id:'parent',message:{role:'assistant',content:[{type:'text',text:JSON.stringify(s.result.transcript.slice(1))}]}}];
expect(verdict(s).sourceRead).toBe(false);expect(verdict(s).testsRead).toBe(false);
});
test('a commented read cannot borrow printed bytes; quoted hash paths remain literal',()=>{
const s=synthetic();
const command=`printf '${Buffer.from(fixture.files.source).toString('base64')}' | base64 -d; # only printed bytes; cat src/billing.ts`;
Object.assign(block(s,1),{name:'Bash',input:{command}});
expect(verdict(s).sourceRead).toBe(false);
const quoted=synthetic();quoted.files.source.path=path.join(quoted.files.cwd,'src/billing#branch.ts');
Object.assign(block(quoted,1),{name:'Bash',input:{command:"cat 'src/billing#branch.ts'"}});
expect(verdict(quoted).sourceRead).toBe(true);
});
test('completion and tool errors remain final gate failures despite genuine delivery',()=>{
for(const exitReason of ['timeout','exit_code_1']){const s=synthetic();s.result.exitReason=exitReason;expect(verdict(s).passed).toBe(false);}
const s=synthetic();s.result.browseErrors=['read failed'];expect(verdict(s).passed).toBe(false);
});
test('coverage markers must belong to the seeded payment and refund functions',()=>{
for(const output of [
diagram.replace('[TESTED]','[UNTESTED]'),diagram.replace('[UNTESTED]','[TESTED]'),
diagram.replace('processPayment','processPaymentExample'),diagram.replace('refundPayment','refundPaymentExample'),
'```\n├── processPayment: happy path [TESTED]\n└── refundPayment [TESTED]\n└── unrelatedPayment [UNTESTED]\n```',
'```\n├── processPayment: happy path [TESTED]\n```\n```\n└── refundPayment [UNTESTED]\n```',
]){const s=synthetic();s.result.output=output;expect(verdict(s).diagram).toBe(false);}
});
test('quoted and nested source examples are not the generated coverage diagram',()=>{
for(const output of [diagram.split('\n').map(line=>'> '+line).join('\n'), '````markdown\n'+diagram+'\n````']){
const s=synthetic();s.result.output=output;expect(verdict(s).diagram).toBe(false);
}
});
test.each([
diagram.replace('[UNTESTED]','[NOT UNTESTED]'),
diagram.replace('[UNTESTED]','[UNTESTED] is false; this function is fully covered.'),
'Example only; this diagram is not the audit result.\n'+diagram,
])('negated gaps and explicitly labeled examples are not audit findings: %s', output=>{
const s=synthetic();s.result.output=output;expect(verdict(s).diagram).toBe(false);
});
test('collector receives exactly the asserted verdict even when process exit succeeded',()=>{
for(const valid of [true,false]){
const s=synthetic();if(!valid)s.result.output='No diagram produced.';
Object.assign(s.result,{toolCalls:[],duration:1,costEstimate:{estimatedCost:0,turnsUsed:1,estimatedTokens:1}});
const v=verdict(s),entries:any[]=[];
recordE2E({addTest:(entry:any)=>entries.push(entry)} as any,'coverage','fixture',s.result,{passed:v.passed,error:v.failures.length?v.failures.join('; '):undefined});
expect(entries).toHaveLength(1);expect(entries[0].passed).toBe(valid);expect(entries[0].error).toBe(valid?undefined:v.failures.join('; '));
}
});
test('coverage evidence files select their exact registered consumers',()=>{
for(const file of ['test/helpers/coverage-audit-evidence.ts','test/coverage-audit-evidence.test.ts','test/fixtures/coverage-audit-ae.json','test/fixtures/coverage-audit-ci-diagrams.json']){
expect(selectTests([file],E2E_TOUCHFILES,GLOBAL_TOUCHFILES).selected.sort()).toEqual(file === 'test/helpers/coverage-audit-evidence.ts' ? ['plan-eng-coverage-audit','review-coverage-audit','ship-coverage-audit'] : ['plan-eng-coverage-audit','review-coverage-audit']);
expect(selectTests([file],LLM_JUDGE_TOUCHFILES,GLOBAL_TOUCHFILES).selected).toEqual([]);
}
});
});

Some files were not shown because too many files have changed in this diff Show More