mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
test: retire the finding-count cluster and trim its helpers (C)
- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
budget, never on skill behavior; delete them, their touchfile/tier ids,
AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
(11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
closure, and delete the free replay tests whose assertions exercised only
that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
only as input for a live subject keep their assertions: the multiSelect
default moved to plan-review-decisions, runner PTY tests use inline caller
policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
(its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
1 parent
53e7f3212f
commit
6415690a18
317 files changed
+315
-54111
No files matched your search
@@ -4,7 +4,7 @@ name: Periodic Evals
|
||||
# tests can't rot invisibly — the class where the autoplan-dual-voice E2E was
|
||||
# silently broken for months until a lucky local diff selected it. Engine:
|
||||
# scripts/test-paid-shards.ts (the same runner local eval:bg:periodic uses):
|
||||
# one planner manifest, 6 ordinary slices plus overlay and Autoplan slices, and a FAIL-CLOSED report — a slice
|
||||
# one planner manifest, 6 ordinary slices plus an overlay slice, and a FAIL-CLOSED report — a slice
|
||||
# whose artifact never landed is a failure, not an absence. The gate-census
|
||||
# job is the weekly EVALS_ALL backstop for the gate tier (PR lanes are
|
||||
# diff-billed, so without it the full gate census might never execute
|
||||
@@ -96,7 +96,7 @@ jobs:
|
||||
- name: Emit run manifest (ALL periodic tests minus reasoned excludes)
|
||||
env:
|
||||
EVALS_ALL: "1"
|
||||
run: EVALS_TIER=periodic bun --no-install run scripts/test-paid-shards.ts --tier periodic --emit-plan /tmp/paid-plan/manifest.json --slices 8 --autoplan-slice
|
||||
run: EVALS_TIER=periodic bun --no-install run scripts/test-paid-shards.ts --tier periodic --emit-plan /tmp/paid-plan/manifest.json --slices 7
|
||||
|
||||
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
||||
with:
|
||||
@@ -118,8 +118,8 @@ jobs:
|
||||
eval-slices:
|
||||
runs-on: ubicloud-standard-8
|
||||
needs: [build-image, plan-slices]
|
||||
# Eight slices retain every registered case and retry. The complete
|
||||
# census needs at most 338 minutes per slice, plus 20 minutes setup/upload.
|
||||
# Seven slices retain every registered case and retry. The complete
|
||||
# census needs at most 251 minutes per slice, plus 20 minutes setup/upload.
|
||||
timeout-minutes: 358
|
||||
permissions:
|
||||
contents: read
|
||||
@@ -133,7 +133,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
slice: [1, 2, 3, 4, 5, 6, 7, 8]
|
||||
slice: [1, 2, 3, 4, 5, 6, 7]
|
||||
steps:
|
||||
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
||||
with:
|
||||
@@ -171,7 +171,7 @@ jobs:
|
||||
name: paid-plan
|
||||
path: /tmp/paid-plan
|
||||
|
||||
- name: Run slice ${{ matrix.slice }}/8
|
||||
- name: Run slice ${{ matrix.slice }}/7
|
||||
env:
|
||||
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
||||
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
||||
|
||||
@@ -825,6 +825,19 @@ audit trail lives in Aside.
|
||||
|
||||
## Test infrastructure
|
||||
|
||||
### P3: No paid eval runs the full /autoplan chain
|
||||
|
||||
**What:** `skill-e2e-autoplan-chain` was retired (it never reached a product
|
||||
verdict: launch failures, then 85-minute budget overruns). Phase order is still
|
||||
enforced by `autoplan/bin/phase-publication-hook.ts` and pinned by the free
|
||||
`test/autoplan-publication-guard.test.ts`, and `skill-e2e-autoplan-dual-voice`
|
||||
covers CEO Phase 1 dispatch. Nothing proves a live model completes
|
||||
CEO → Design → DX → Eng or reads the required phase sections
|
||||
(`CARVE_GUARDS.autoplan` is `behavioral: 'none'`).
|
||||
|
||||
**Re-entry:** a chain eval that fits the ordinary PTY tiers, for example one that
|
||||
runs the no-UI, no-DX path (CEO then Eng) and asserts the section reads.
|
||||
|
||||
### P3: CI-unrunnable paid evals
|
||||
|
||||
**What:** Seven paid files cannot execute in the CI image (no `codex` CLI, no
|
||||
@@ -1972,6 +1985,13 @@ plus a TTL so abandoned PTYs eventually exit.
|
||||
**Priority:** P2.
|
||||
**Effort:** S (CC: ~30 min once fixture exists). Captured from v1.21.1.0 plan-eng-review D2.
|
||||
|
||||
**Status (2026-09):** The four `skill-e2e-plan-*-finding-count` evals were retired
|
||||
after eight red weekly runs whose failures were harness and budget, not skill
|
||||
behavior. The `*-finding-floor` evals assert at least one AskUserQuestion, not one
|
||||
per finding, so this contract has no paid coverage today. Re-entry test: a
|
||||
qid-keyed per-finding count on a multi-finding fixture with `QUESTION_TUNING: true`
|
||||
(the `<gstack-qid:…>` markers only appear with tuning on).
|
||||
|
||||
---
|
||||
|
||||
## P3: Honor env vars in gstack-config (so QUESTION_TUNING/EXPLAIN_LEVEL actually isolate tests)
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
"$schema": "https://gstack.dev/schemas/section-manifest.json",
|
||||
"skill": "autoplan",
|
||||
"version": 1,
|
||||
"note": "PASSIVE registry (v2 plan T9 / CM2). Fields are IDs, file paths, human titles, and human-readable trigger text ONLY. The skeleton's phase sequencing (Sequential Execution + the Phase 0 UI/DX scope detection) is the ONLY place that decides WHEN to read a section \u2014 Phase 2 and Phase 2.5 are conditional and their sections must NOT be read when their scope is absent; required section reads are checked by test/skill-e2e-autoplan-chain.test.ts (auditAutoplanMethodReads). No machine predicate here \u2014 see docs/designs/v2_PLAN.md:663.",
|
||||
"note": "PASSIVE registry (v2 plan T9 / CM2). Fields are IDs, file paths, human titles, and human-readable trigger text ONLY. The skeleton's phase sequencing (Sequential Execution + the Phase 0 UI/DX scope detection) is the ONLY place that decides WHEN to read a section \u2014 Phase 2 and Phase 2.5 are conditional and their sections must NOT be read when their scope is absent; no paid eval checks the required section reads since the autoplan chain eval was retired (TODOS.md). No machine predicate here \u2014 see docs/designs/v2_PLAN.md:663.",
|
||||
"sections": [
|
||||
{
|
||||
"id": "ceo-phase",
|
||||
|
||||
+12
-30
@@ -41,7 +41,7 @@ Seeded planning sessions also receive an isolated runtime home through
|
||||
to the working tree under test. Explicit per-test home overrides remain intact.
|
||||
Autoplan resolves each review skill from its own installed host registry.
|
||||
|
||||
**Interactive planning evidence.** Finding-count and autoplan-chain drivers use
|
||||
**Interactive planning evidence.** Native plan-review count drivers use
|
||||
`observeScreen: true` and await `currentScreen()` before choosing an input. The
|
||||
existing xterm dependency interprets cursor moves and erases; old menus in the
|
||||
raw stream cannot establish a current prompt. Snapshots preserve
|
||||
@@ -300,28 +300,11 @@ archaeology.
|
||||
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
|
||||
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall
|
||||
minus overhead and ratchets raw literals. Budget above the wall is fiction.
|
||||
The registered four-phase exception is `AUTOPLAN_CHAIN_BUDGET` for
|
||||
`test/skill-e2e-autoplan-chain.test.ts`: 80 minutes of work (four `PTY_LONG`
|
||||
allocations), an 84-minute session watchdog, an 85-minute Bun test deadline,
|
||||
and a 172-minute supervised shard wall. The unchanged retry count of one
|
||||
permits two 85-minute attempts plus two minutes for cleanup. This is a
|
||||
**specified allocation for the stronger four-phase contract**, not a measured
|
||||
calibration or statistical upper bound. The historical 900-second failures
|
||||
remain failures. Models, fixtures, phase assertions and production review
|
||||
caller timeouts are unchanged; this explicitly changes eval latency/cost policy.
|
||||
No paid test may exceed the ordinary tiers.
|
||||
|
||||
The Autoplan chain explicitly enables native `PreToolUse` approval for edits to
|
||||
its owned temporary review artifacts. Approval starts with the `/autoplan`
|
||||
command and requires the exact parent session, prior successful file history,
|
||||
and a current request digest. Other recorder callers remain observational.
|
||||
A rejected artifact edit fails the test instead of falling through to terminal
|
||||
permission input. Approval itself supplies no edit success or phase credit:
|
||||
the native tool result and all four completed review phases are still required.
|
||||
|
||||
`FINDING_RETRY_BUDGETS` also registers six finding files. Each retains its
|
||||
25-minute case deadline and one retry: the two-case CEO finding-count file has
|
||||
a 102-minute shard wall, and the five single-case files have 52-minute walls,
|
||||
including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
`FINDING_RETRY_BUDGETS` also registers the CEO split-overflow and Eng
|
||||
multi-finding batching files. Each retains its 25-minute case deadline and one
|
||||
retry in a 52-minute shard wall, including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
have a 1,830-second minimum shard wall and run without Bun retries; see the
|
||||
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
|
||||
|
||||
@@ -332,27 +315,26 @@ recording inside a ten-second Bun grace; the other 11 retain their existing
|
||||
120-second Bun timeout. Late responses cannot create records or cache passes.
|
||||
|
||||
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
|
||||
Autoplan, each registered finding file, and each overlay wrapper require their
|
||||
Each registered finding file and each overlay wrapper requires its
|
||||
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
|
||||
jobs are rejected so ordinary files retain their configured retries. An explicit
|
||||
CLI `--timeout`, `EVALS_SHARD_TIMEOUT_MS`, or API `timeoutMs` still wins for these
|
||||
policies, including a lower cap; overlay overrides below their minimum are rejected.
|
||||
Planner entries and execution results record the effective wall,
|
||||
its source and policy identifier. Custom drivers must resolve each job instead
|
||||
of passing their ordinary 1800-second default as an explicit Autoplan cap;
|
||||
of passing their ordinary 1800-second default as an explicit cap;
|
||||
their outer controller/detach wall must also cover the allocated work and cleanup.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
|
||||
wrapper covers a full-gate fallback at its default two workers. The broad gate
|
||||
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
|
||||
tiers. Legacy monolithic
|
||||
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
|
||||
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
|
||||
promise every registered retry; use the sharded periodic path for this policy.
|
||||
|
||||
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
|
||||
When overlays are selected, the seventh is reserved for their serial wrappers;
|
||||
registered finding files are distributed across the remaining ordinary slices
|
||||
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
|
||||
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
|
||||
Periodic CI plans `--slices 7`. When overlays are selected, the seventh is
|
||||
reserved for their serial wrappers; registered finding files are distributed
|
||||
across the remaining ordinary slices by their supervised walls. Each slice job
|
||||
has a 358-minute cap. Reconciliation rejects missing, duplicated or misplaced
|
||||
registered work and absent budget records. The weekly gate census has a
|
||||
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
|
||||
verify these bounds against the complete current census, configured retries,
|
||||
|
||||
@@ -809,7 +809,7 @@ function hasCompleteCiSummary(outcome: FreeShardOutcome): boolean {
|
||||
|
||||
export const QUICK_CORE = [
|
||||
'test/strict-output.test.ts', 'test/gen-skill-docs.test.ts',
|
||||
'test/skill-check-driver.test.ts', 'test/ceo-native-ledger-replay.test.ts',
|
||||
'test/skill-check-driver.test.ts',
|
||||
'test/skill-ceo-section-ordering.test.ts',
|
||||
];
|
||||
|
||||
|
||||
+14
-68
@@ -65,7 +65,7 @@ import {
|
||||
} from './test-strict-output';
|
||||
import { PAID_TEST_GLOBS, isPaidTestFile } from '../test/helpers/paid-test-set';
|
||||
import { PERIODIC_CI_EXCLUDE } from '../test/helpers/periodic-exclude-data';
|
||||
import { AUTOPLAN_CHAIN_BUDGET, FILE_RETRY_BUDGETS, STRICT_RETRY_CASE_BUDGETS } from '../test/helpers/eval-budgets';
|
||||
import { FILE_RETRY_BUDGETS, STRICT_RETRY_CASE_BUDGETS } from '../test/helpers/eval-budgets';
|
||||
import { getProjectEvalDir, getClaudeCliVersion, isFinalizedEvalResultFile, evalEntryOutcome } from '../test/helpers/eval-store';
|
||||
import { manualReviewProblem } from '../test/helpers/cookie-workflow-manual-review';
|
||||
import { preflightAnthropicApi } from '../test/helpers/anthropic-preflight';
|
||||
@@ -464,7 +464,7 @@ export function planPaidShards(
|
||||
const shards: string[][] = [];
|
||||
let pending: string[] = [];
|
||||
for (const file of unique) {
|
||||
if (isOverlayTestFile(file) || file === AUTOPLAN_CHAIN_BUDGET.file || FILE_RETRY_BUDGETS.some(budget => budget.file === file)) {
|
||||
if (isOverlayTestFile(file) || FILE_RETRY_BUDGETS.some(budget => budget.file === file)) {
|
||||
if (pending.length) shards.push(pending);
|
||||
pending = [];
|
||||
shards.push([file]);
|
||||
@@ -485,8 +485,6 @@ export interface PaidShardBudget {
|
||||
|
||||
/** Explicit caller limits win; registered supervision preserves existing attempts. */
|
||||
export function resolvePaidShardBudget(files: string[], overrideMs?: number): PaidShardBudget {
|
||||
const autoplan = files.map(normalizeRelativePath).includes(AUTOPLAN_CHAIN_BUDGET.file);
|
||||
if (autoplan && files.length !== 1) throw new Error('Autoplan budget requires its own shard');
|
||||
const finding = FILE_RETRY_BUDGETS.find(budget => files.map(normalizeRelativePath).includes(budget.file));
|
||||
if (finding && files.length !== 1) throw new Error('Registered retry budget requires its own shard');
|
||||
if (overrideMs !== undefined && (!Number.isSafeInteger(overrideMs) || overrideMs <= 0 || overrideMs > 2_147_483_647)) {
|
||||
@@ -498,9 +496,9 @@ export function resolvePaidShardBudget(files: string[], overrideMs?: number): Pa
|
||||
throw new Error(`Overlay shard requires at least ${OVERLAY_MIN_FILE_WALL_MS}ms; explicit wall ${overrideMs}ms cannot preserve its work and finalization budget`);
|
||||
}
|
||||
return {
|
||||
timeoutMs: overrideMs ?? (autoplan ? AUTOPLAN_CHAIN_BUDGET.shardMs : finding ? finding.shardMs : overlay ? OVERLAY_MIN_FILE_WALL_MS : DEFAULT_SHARD_TIMEOUT_MS),
|
||||
source: overrideMs !== undefined ? 'explicit' : autoplan || finding ? 'registered' : 'default',
|
||||
policyId: autoplan ? AUTOPLAN_CHAIN_BUDGET.id : finding?.id ?? null,
|
||||
timeoutMs: overrideMs ?? (finding ? finding.shardMs : overlay ? OVERLAY_MIN_FILE_WALL_MS : DEFAULT_SHARD_TIMEOUT_MS),
|
||||
source: overrideMs !== undefined ? 'explicit' : finding ? 'registered' : 'default',
|
||||
policyId: finding?.id ?? null,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -602,8 +600,6 @@ export function paidShardWallUpperBoundMs(files: string[], jobs: number, overrid
|
||||
|
||||
export interface RunShardsOptions {
|
||||
timeoutMs?: number;
|
||||
/** Legacy Autoplan allocation; callers may supply registered per-file allocations. */
|
||||
autoplanBudget?: PaidShardBudget;
|
||||
registeredBudgets?: Record<string, PaidShardBudget>;
|
||||
jobs?: number;
|
||||
/** bun --max-concurrency inside each shard (EVALS_CONCURRENCY). */
|
||||
@@ -660,8 +656,7 @@ export async function runPaidShard(
|
||||
): Promise<ShardOutcome> {
|
||||
if (files.length === 0) throw new Error('Cannot run an empty paid-test shard.');
|
||||
const rootDir = options.rootDir ?? ROOT;
|
||||
const planned = options.registeredBudgets?.[normalizeRelativePath(files[0]!)] ??
|
||||
(files.map(normalizeRelativePath).includes(AUTOPLAN_CHAIN_BUDGET.file) ? options.autoplanBudget : undefined);
|
||||
const planned = options.registeredBudgets?.[normalizeRelativePath(files[0]!)];
|
||||
const budget = resolvePaidShardBudget(files, options.timeoutMs ??
|
||||
(planned?.source === 'explicit' ? planned.timeoutMs : undefined));
|
||||
const timeoutMs = budget.timeoutMs;
|
||||
@@ -981,7 +976,7 @@ export interface ManifestEntry {
|
||||
slice: number;
|
||||
status: 'planned' | 'skipped-by-diff' | 'excluded';
|
||||
reason?: string;
|
||||
/** Required when the registered Autoplan workflow is planned. */
|
||||
/** Required when a registered retry-budget file is planned. */
|
||||
budget?: PaidShardBudget;
|
||||
}
|
||||
|
||||
@@ -995,8 +990,6 @@ export interface PaidRunManifest {
|
||||
profile?: PaidProfile;
|
||||
selection?: PaidCaseSelection;
|
||||
prCoverage?: PrProfileSelection;
|
||||
/** Dedicated last slice; preceding slices retain ordinary round-robin work. */
|
||||
autoplanSlice?: number;
|
||||
entries: ManifestEntry[];
|
||||
}
|
||||
|
||||
@@ -1063,7 +1056,6 @@ export function buildRunManifest(opts: {
|
||||
profile?: PaidProfile;
|
||||
sliceCount: number;
|
||||
evalsAll: boolean;
|
||||
dedicatedAutoplanSlice?: boolean;
|
||||
timeoutMs?: number;
|
||||
discovered?: string[];
|
||||
env?: NodeJS.ProcessEnv;
|
||||
@@ -1075,9 +1067,6 @@ export function buildRunManifest(opts: {
|
||||
if (!Number.isInteger(opts.sliceCount) || opts.sliceCount <= 0) {
|
||||
throw new Error(`--slices needs a positive integer. Received: ${opts.sliceCount}`);
|
||||
}
|
||||
if (opts.dedicatedAutoplanSlice && (opts.tier !== 'periodic' || opts.sliceCount < 2)) {
|
||||
throw new Error('Dedicated Autoplan slice requires periodic tier and at least two total slices');
|
||||
}
|
||||
const rootDir = opts.rootDir ?? ROOT;
|
||||
const env = opts.env ?? process.env;
|
||||
const profile = opts.profile ?? validatedProfile(env.EVALS_PROFILE, 'EVALS_PROFILE');
|
||||
@@ -1095,16 +1084,14 @@ export function buildRunManifest(opts: {
|
||||
}
|
||||
|
||||
const entries: ManifestEntry[] = [];
|
||||
const overlaySlice = opts.sliceCount - (opts.dedicatedAutoplanSlice ? 1 : 0);
|
||||
const overlaySlice = opts.sliceCount;
|
||||
const reserveOverlaySlice = overlaySlice > 1 && runnable.some(files => files.some(isOverlayTestFile));
|
||||
const ordinarySlices = overlaySlice - Number(reserveOverlaySlice);
|
||||
// Spread registered long files by supervised load. Keep one ordinary-only
|
||||
// lane when possible, so every lane does not inherit a long-workflow tail.
|
||||
// Reserved overlay and dedicated Autoplan slices retain their ownership.
|
||||
const ordinary = runnable.filter(files => !files.some(isOverlayTestFile) &&
|
||||
!(opts.dedicatedAutoplanSlice && files[0] === AUTOPLAN_CHAIN_BUDGET.file));
|
||||
const registered = ordinary.filter(files => files[0] === AUTOPLAN_CHAIN_BUDGET.file ||
|
||||
FILE_RETRY_BUDGETS.some(budget => budget.file === files[0]));
|
||||
// The reserved overlay slice retains its ownership.
|
||||
const ordinary = runnable.filter(files => !files.some(isOverlayTestFile));
|
||||
const registered = ordinary.filter(files => FILE_RETRY_BUDGETS.some(budget => budget.file === files[0]));
|
||||
const allocations = new Map<string, number>();
|
||||
if (registered.length && ordinarySlices > 1) {
|
||||
const loads = Array<number>(ordinarySlices).fill(0);
|
||||
@@ -1173,12 +1160,9 @@ export function buildRunManifest(opts: {
|
||||
return new Map(lanes.flatMap((files, lane) => files.map(file => [file, lane + 1] as const)));
|
||||
}
|
||||
runnable.forEach((files) => {
|
||||
const autoplan = files[0] === AUTOPLAN_CHAIN_BUDGET.file;
|
||||
const slice = opts.dedicatedAutoplanSlice && autoplan ? opts.sliceCount
|
||||
: files.some(isOverlayTestFile) ? overlaySlice
|
||||
: (packed ?? allocations).get(files[0])!;
|
||||
const slice = files.some(isOverlayTestFile) ? overlaySlice : (packed ?? allocations).get(files[0])!;
|
||||
entries.push({ file: files[0], slice, status: 'planned',
|
||||
...(autoplan || FILE_RETRY_BUDGETS.some(budget => budget.file === files[0])
|
||||
...(FILE_RETRY_BUDGETS.some(budget => budget.file === files[0])
|
||||
? { budget: resolvePaidShardBudget(files, opts.timeoutMs) } : {}) });
|
||||
});
|
||||
for (const s of skipped) entries.push({ file: s.files[0], slice: 0, status: 'skipped-by-diff', reason: s.reason });
|
||||
@@ -1194,7 +1178,6 @@ export function buildRunManifest(opts: {
|
||||
profile,
|
||||
selection: cases.selection,
|
||||
...(cases.coverage ? { prCoverage: cases.coverage } : {}),
|
||||
...(opts.dedicatedAutoplanSlice ? { autoplanSlice: opts.sliceCount } : {}),
|
||||
entries,
|
||||
};
|
||||
return parseRunManifest(JSON.stringify(manifest));
|
||||
@@ -1254,7 +1237,7 @@ export function parseRunManifest(raw: string): PaidRunManifest {
|
||||
}
|
||||
}
|
||||
}
|
||||
const overlaySlice = parsed.sliceCount - (parsed.autoplanSlice !== undefined ? 1 : 0);
|
||||
const overlaySlice = parsed.sliceCount;
|
||||
const plannedOverlays = parsed.entries.filter(entry => entry.status === 'planned' && isOverlayTestFile(entry.file));
|
||||
if (plannedOverlays.some(entry => entry.slice !== overlaySlice)) {
|
||||
throw new Error('Overlay manifest entries must share the final ordinary slice to preserve one-process API admission');
|
||||
@@ -1263,23 +1246,6 @@ export function parseRunManifest(raw: string): PaidRunManifest {
|
||||
entry.status === 'planned' && !isOverlayTestFile(entry.file) && entry.slice === overlaySlice)) {
|
||||
throw new Error('The final ordinary manifest slice is reserved for overlay files');
|
||||
}
|
||||
const autoplan = parsed.entries.filter(entry => normalizeRelativePath(entry.file) === AUTOPLAN_CHAIN_BUDGET.file);
|
||||
if (autoplan.length > 1) throw new Error('Duplicate Autoplan manifest entry');
|
||||
if (parsed.autoplanSlice !== undefined) {
|
||||
if (parsed.tier !== 'periodic' || parsed.autoplanSlice !== parsed.sliceCount || parsed.sliceCount < 2 || autoplan.length !== 1 || autoplan[0].status !== 'planned') {
|
||||
throw new Error('Dedicated Autoplan slice is missing or malformed');
|
||||
}
|
||||
for (const entry of parsed.entries.filter(entry => entry.status === 'planned')) {
|
||||
if ((entry.file === AUTOPLAN_CHAIN_BUDGET.file) !== (entry.slice === parsed.autoplanSlice)) {
|
||||
throw new Error('Dedicated Autoplan slice contains missing or unrelated work');
|
||||
}
|
||||
}
|
||||
}
|
||||
for (const entry of autoplan.filter(entry => entry.status === 'planned')) {
|
||||
if (!entry.budget) throw new Error('Autoplan manifest needs an explicit budget record; emit a fresh plan');
|
||||
const expected = resolvePaidShardBudget([entry.file], entry.budget.source === 'explicit' ? entry.budget.timeoutMs : undefined);
|
||||
if (!sameBudget(entry.budget, expected)) throw new Error('Autoplan manifest budget differs from declared policy');
|
||||
}
|
||||
for (const budget of FILE_RETRY_BUDGETS) {
|
||||
const entries = parsed.entries.filter(entry => normalizeRelativePath(entry.file) === budget.file);
|
||||
if (entries.length > 1) throw new Error(`Duplicate registered manifest entry: ${budget.file}`);
|
||||
@@ -1334,9 +1300,6 @@ export function verifySliceResults(
|
||||
const reported = new Map<string, { slice: number; status: ShardStatus }>();
|
||||
for (const result of results) {
|
||||
for (const outcome of result.outcomes) {
|
||||
if (outcome.files.map(normalizeRelativePath).includes(AUTOPLAN_CHAIN_BUDGET.file) && outcome.files.length !== 1) {
|
||||
problems.push('Autoplan result must report its own shard');
|
||||
}
|
||||
if (outcome.files.some(file => FILE_RETRY_BUDGETS.some(budget => budget.file === normalizeRelativePath(file))) && outcome.files.length !== 1) {
|
||||
problems.push('Registered result must report its own shard');
|
||||
}
|
||||
@@ -1370,17 +1333,6 @@ export function verifySliceResults(
|
||||
if (!sameBudget(outcome.budget, expected)) problems.push(`Registered effective result budget differs from its planned/explicit allocation: ${file}`);
|
||||
} catch { problems.push(`Invalid registered effective result budget: ${file}`); }
|
||||
}
|
||||
if (file === AUTOPLAN_CHAIN_BUDGET.file) {
|
||||
if (outcome.exitCode !== 0 || outcome.executedTests !== 1 || outcome.skippedTests !== 0) {
|
||||
problems.push('Autoplan must execute exactly one unskipped case with exit zero');
|
||||
}
|
||||
try {
|
||||
const planned = manifest.entries.find(entry => entry.file === file)?.budget;
|
||||
const expected = resolvePaidShardBudget([file], result.timeoutOverrideMs ??
|
||||
(planned?.source === 'explicit' ? planned.timeoutMs : undefined));
|
||||
if (!sameBudget(outcome.budget, expected)) problems.push('Autoplan effective result budget differs from its planned/explicit allocation');
|
||||
} catch { problems.push('Invalid Autoplan effective result budget'); }
|
||||
}
|
||||
}
|
||||
}
|
||||
for (const entry of manifest.entries) {
|
||||
@@ -1432,7 +1384,6 @@ type CliOptions = {
|
||||
listOnly: boolean;
|
||||
timeoutMs: number;
|
||||
timeoutExplicit: boolean;
|
||||
dedicatedAutoplanSlice: boolean;
|
||||
jobs: number;
|
||||
withinShardConcurrency: number;
|
||||
maxFilesPerShard: number;
|
||||
@@ -1481,7 +1432,6 @@ export function parseCliOptions(argv: string[], env: NodeJS.ProcessEnv = process
|
||||
profileExplicit: !!env.EVALS_PROFILE,
|
||||
listOnly: false,
|
||||
timeoutExplicit: !!env.EVALS_SHARD_TIMEOUT_MS,
|
||||
dedicatedAutoplanSlice: false,
|
||||
timeoutMs: env.EVALS_SHARD_TIMEOUT_MS
|
||||
? parsePositiveInt(env.EVALS_SHARD_TIMEOUT_MS, 'EVALS_SHARD_TIMEOUT_MS')
|
||||
: DEFAULT_SHARD_TIMEOUT_MS,
|
||||
@@ -1517,7 +1467,6 @@ export function parseCliOptions(argv: string[], env: NodeJS.ProcessEnv = process
|
||||
options.profile = validatedProfile(value, '--profile'); options.profileExplicit = true; continue;
|
||||
}
|
||||
if (arg === '--timeout') { options.timeoutMs = parsePositiveInt(argv[index += 1], '--timeout') * 1000; options.timeoutExplicit = true; continue; }
|
||||
if (arg === '--autoplan-slice') { options.dedicatedAutoplanSlice = true; continue; }
|
||||
if (arg === '--jobs') { options.jobs = parsePositiveInt(argv[index += 1], '--jobs'); continue; }
|
||||
if (arg === '--files-per-shard') { options.maxFilesPerShard = parsePositiveInt(argv[index += 1], '--files-per-shard'); continue; }
|
||||
if (arg === '--emit-plan') {
|
||||
@@ -1541,7 +1490,6 @@ export function parseCliOptions(argv: string[], env: NodeJS.ProcessEnv = process
|
||||
throw new Error(`Unknown argument: ${arg}`);
|
||||
}
|
||||
if (options.writeDurations && !options.reportDir) throw new Error('--write-durations requires --report');
|
||||
if (options.dedicatedAutoplanSlice && !options.emitPlanPath) throw new Error('--autoplan-slice requires --emit-plan');
|
||||
if (options.profile === 'pr' && options.tier !== 'gate') throw new Error('PR profile requires gate tier');
|
||||
if (options.profile === 'pr' && options.maxFilesPerShard !== 1) throw new Error('PR profile requires one file per shard to preserve case accounting');
|
||||
return options;
|
||||
@@ -1557,7 +1505,6 @@ async function main(): Promise<number> {
|
||||
tier: options.tier,
|
||||
profile: options.profile,
|
||||
sliceCount: options.slices,
|
||||
dedicatedAutoplanSlice: options.dedicatedAutoplanSlice,
|
||||
timeoutMs: options.timeoutExplicit ? options.timeoutMs : undefined,
|
||||
evalsAll: process.env.EVALS_ALL === '1',
|
||||
});
|
||||
@@ -1724,7 +1671,6 @@ async function main(): Promise<number> {
|
||||
timeoutMs: options.timeoutExplicit ? options.timeoutMs : undefined,
|
||||
jobs: options.jobs,
|
||||
withinShardConcurrency: options.withinShardConcurrency,
|
||||
autoplanBudget: mine.find(entry => entry.file === AUTOPLAN_CHAIN_BUDGET.file)?.budget,
|
||||
registeredBudgets: Object.fromEntries(mine.filter(entry => entry.budget).map(entry => [normalizeRelativePath(entry.file), entry.budget!])),
|
||||
...(manifest.prCoverage?.mode === 'pr' ? {
|
||||
expectedCases: Object.fromEntries(mine.map(entry => [entry.file, expectedPrCaseCount(entry.file, manifest.selection!)])),
|
||||
|
||||
@@ -167,9 +167,9 @@ describe('owned Autoplan pending artifact metadata recorder',()=>{
|
||||
test('recorder disposal removes owned state and shared recorder inputs select both paid owners',()=>{
|
||||
const f=fixture();f.write(f.event());f.dispose();expect(fs.existsSync(f.recorder.file)).toBe(false);
|
||||
for(const file of ['test/helpers/autoplan-artifact-recorder.ts','test/autoplan-artifact-recorder.test.ts'])
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected.sort()).toEqual(['autoplan-chain-pty','plan-eng-finding-count']);
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected.sort()).toEqual([]);
|
||||
for(const file of ['test/autoplan-pending-artifact.test.ts','test/fixtures/autoplan-pending-artifact-ae.json'])
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
@@ -1,84 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { readFileSync, existsSync, readdirSync } from 'node:fs';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import { createNativeReviewState } from './helpers/plan-count-fixture';
|
||||
import { getHermeticDirs } from './helpers/hermetic-env';
|
||||
import { resolve } from 'node:path';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
const root = resolve(import.meta.dir, '..');
|
||||
const read = (file: string) => readFileSync(resolve(root, file), 'utf8');
|
||||
const fixture = 'test/fixtures/plans/autoplan-dashboard.md';
|
||||
|
||||
test('the chain fixture retains the complete original UI/API scope', () => {
|
||||
// The design fixture adds proposed implementation contracts after the shared
|
||||
// scope. The chain supplies its own existing contracts for independent review.
|
||||
const original = read('test/fixtures/plans/ui-heavy-feature.md')
|
||||
.split('\n## Planned implementation contracts')[0]!.trimEnd();
|
||||
const complete = read(fixture);
|
||||
expect(complete.startsWith(original + '\n')).toBe(true);
|
||||
// This supplements dependency facts; it does not supply a completed review,
|
||||
// prescribe its decisions, or pre-build the feature exercised by the chain.
|
||||
expect(complete).not.toMatch(/Phase \d|GSTACK REVIEW REPORT|AUTO-DECIDE|all findings resolved/i);
|
||||
expect(complete).toContain('there are no dashboard-specific tests yet');
|
||||
expect(complete).toContain('not completed work');
|
||||
});
|
||||
|
||||
test('the new fixture is isolated to the chain and its selection dependencies', () => {
|
||||
expect(read('test/skill-e2e-autoplan-chain.test.ts')).toContain("'plans', 'autoplan-dashboard.md'");
|
||||
expect(read('test/skill-e2e-plan-design-with-ui.test.ts')).toContain("'plans', 'ui-heavy-feature.md'");
|
||||
for (const file of [fixture, 'test/autoplan-chain-fixture.test.ts']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
|
||||
}
|
||||
expect(selectTests(['test/fixtures/plans/ui-heavy-feature.md'], E2E_TOUCHFILES).selected)
|
||||
.toEqual(['plan-design-with-ui-scope']);
|
||||
});
|
||||
|
||||
|
||||
test('native sequencing config reaches the real CLI reader without changing shared state', () => {
|
||||
const shared = getHermeticDirs().gstackHome;
|
||||
const before = readFileSync(resolve(shared, 'config.yaml'), 'utf8');
|
||||
const first = createNativeReviewState();
|
||||
const second = createNativeReviewState();
|
||||
try {
|
||||
expect(first.env.GSTACK_HOME).not.toBe(shared);
|
||||
expect(first.env.GSTACK_HOME).not.toBe(second.env.GSTACK_HOME);
|
||||
expect(first.env.GSTACK_STATE_ROOT).toBe(first.env.GSTACK_HOME);
|
||||
const result = spawnSync('bash', [resolve(root, 'bin/gstack-config'), 'get', 'codex_reviews'], {
|
||||
cwd: root, env: { ...process.env, ...first.env }, encoding: 'utf8', timeout: 5000,
|
||||
});
|
||||
expect(result.status, result.stderr).toBe(0);
|
||||
expect(result.stdout.trim()).toBe('disabled');
|
||||
for (const marker of readdirSync(shared).filter(name => name === '.activated' ||
|
||||
/^\..*(?:-seen|-prompted|-shown)$/.test(name) || name.startsWith('.feature-prompted-'))) {
|
||||
expect(readFileSync(resolve(first.env.GSTACK_HOME!, marker), 'utf8'))
|
||||
.toBe(readFileSync(resolve(shared, marker), 'utf8'));
|
||||
}
|
||||
first.cleanup();
|
||||
first.cleanup();
|
||||
expect(existsSync(first.env.GSTACK_HOME!)).toBe(false);
|
||||
expect(existsSync(second.env.GSTACK_HOME!)).toBe(true);
|
||||
expect(readFileSync(resolve(shared, 'config.yaml'), 'utf8')).toBe(before);
|
||||
} finally {
|
||||
first.cleanup();
|
||||
second.cleanup();
|
||||
}
|
||||
expect(existsSync(second.env.GSTACK_HOME!)).toBe(false);
|
||||
});
|
||||
|
||||
test('the UI/API chain requires all four native phases and registers its config dependency', () => {
|
||||
const source = read('test/skill-e2e-autoplan-chain.test.ts');
|
||||
const plan = read(fixture);
|
||||
expect(plan).toContain('## UI Scope');
|
||||
expect(plan).toContain('New REST endpoint `GET /api/dashboard`');
|
||||
expect(source).toContain('env: nativeState.env');
|
||||
expect(source).toContain('if (!ceo || !design || !dx || !eng)');
|
||||
expect(source).toContain('expect(ceo.ts).toBeLessThan(design.ts)');
|
||||
expect(source).toContain('expect(design.ts).toBeLessThan(dx.ts)');
|
||||
expect(source).toContain('expect(dx.ts).toBeLessThan(eng.ts)');
|
||||
expect(source).toContain('nativeState?.cleanup()');
|
||||
for (const file of ['test/helpers/plan-count-fixture.ts', 'test/plan-count-fixture.test.ts', 'bin/gstack-config']) {
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain(file);
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('autoplan-chain-pty');
|
||||
}
|
||||
});
|
||||
@@ -2,5 +2,5 @@ import {test,expect,afterEach} from 'bun:test';
|
||||
import {E2E_TOUCHFILES} from './helpers/touchfiles-data';
|
||||
const cleanup:Array<()=>void>=[];afterEach(()=>{for(const f of cleanup.splice(0))f()});
|
||||
test('new regression files register only the actual Autoplan owner',()=>{
|
||||
for(const p of ['test/autoplan-clipped-suffix-aq.test.ts','test/fixtures/autoplan-clipped-suffix-aq.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes(p)).map(([owner])=>owner)).toEqual(['autoplan-chain-pty']);
|
||||
for(const p of ['test/autoplan-clipped-suffix-aq.test.ts','test/fixtures/autoplan-clipped-suffix-aq.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes(p)).map(([owner])=>owner)).toEqual([]);
|
||||
});
|
||||
@@ -1,95 +0,0 @@
|
||||
import {expect, test} from 'bun:test';
|
||||
import {readFileSync} from 'node:fs';
|
||||
import {autoplanBlockingQuestionBoundary, autoplanSetupDecision} from './helpers/autoplan-setup-question';
|
||||
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
|
||||
import {E2E_TOUCHFILES, GLOBAL_TOUCHFILES} from './helpers/touchfiles';
|
||||
import capture from './fixtures/autoplan-cropped-gate-av.json';
|
||||
const fixture = (): {screen:string; context:Parameters<typeof autoplanBlockingQuestionBoundary>[1]} => ({screen:capture.screen,
|
||||
context:{commandStartedAt:capture.commandStartedAt,viewportCapturedAt:capture.viewportCapturedAt,
|
||||
transcript:{status:'ready',calls:[structuredClone(capture.call)],assistantMessages:[]},publicTools:[structuredClone(capture.publicUse)]}});
|
||||
const call=(f:ReturnType<typeof fixture>)=>f.context.transcript.calls[0]!;
|
||||
const detect=(f=fixture())=>autoplanBlockingQuestionBoundary(f.screen,f.context);
|
||||
const expected={sessionId:capture.call.sessionId,toolUseId:capture.call.toolUseId,source:'native'};
|
||||
const rebind=(f:ReturnType<typeof fixture>)=>{f.context.publicTools[0]!.input!.questions=structuredClone(call(f).questions);};
|
||||
type Change=(f:ReturnType<typeof fixture>)=>void;
|
||||
|
||||
test('exact AV crop proves a human wait without answer, phase credit or evidence mutation',()=>{
|
||||
const f=fixture(),before=JSON.stringify(f);expect(detect(f)).toEqual(expected);
|
||||
expect(autoplanSetupDecision(f.screen,new Set(),call(f))).toEqual({kind:'unrelated'});
|
||||
expect(autoplanPhaseCompletions(f.context.transcript,f.context.commandStartedAt)).toEqual([]);
|
||||
expect(call(f).answered).toBe(false);expect(call(f).failed).toBe(false);expect(JSON.stringify(f)).toBe(before);
|
||||
});
|
||||
test('wrapping and crop position may vary while the owned excerpt and choices remain exact',()=>{
|
||||
const controls:Change[]=[
|
||||
f=>{f.screen=f.screen.replace(/\n/g,'\r\n');},f=>{f.screen=f.screen.replace(/^│ /gm,'┃ ');},
|
||||
f=>{f.screen=f.screen.replace('wall…','wall-clock time');},
|
||||
f=>{f.screen=f.screen.replace('│ Pros / cons:\n','│ Pros /\n│ cons:\n');},
|
||||
f=>{f.screen=f.screen.slice(f.screen.indexOf('│ Stakes if'));},
|
||||
f=>{call(f).questions[0]!.header='Final approval gate';call(f).questions[0]!.question=call(f).questions[0]!.question.replace('D1 — Final Approval Gate: approve the reviewed plan?','D8 — Final Approval: approve the amended plan?');rebind(f);},
|
||||
// A native human wait stays real even if the question body retracts approval.
|
||||
f=>{call(f).questions[0]!.question+='\nThis final approval gate is withdrawn.';rebind(f);},
|
||||
];for(const [i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toEqual(expected);}
|
||||
});
|
||||
test('native identity, public use, no acknowledgment, current session and time remain mandatory',()=>{
|
||||
const controls:Change[]=[
|
||||
f=>{f.context.transcript.status='missing';},f=>{f.context.transcript.status='error';},f=>{f.context.transcript.calls=[];},f=>{f.context.publicTools=[];},
|
||||
f=>{call(f).answered=true;},f=>{call(f).failed=true;},f=>{call(f).toolUseId='foreign';},f=>{call(f).sessionId='foreign';},
|
||||
f=>{f.context.publicTools[0]!.toolUseId='foreign';},f=>{f.context.publicTools[0]!.sessionId='foreign';},f=>{f.context.publicTools[0]!.name='Read';},
|
||||
f=>{f.context.publicTools[0]!.timestamp='bad';},f=>{f.context.publicTools[0]!.timestamp=new Date(f.context.viewportCapturedAt+1).toISOString();},
|
||||
f=>{f.context.commandStartedAt=Date.parse(capture.publicUse.timestamp)+1;},f=>{f.context.commandStartedAt=NaN;},f=>{f.context.viewportCapturedAt=Infinity;},
|
||||
f=>{f.context.publicTools[0]!.input!.questions=[];},f=>{f.context.publicTools[0]!.input!.questions=[{header:'Foreign',question:'Other?'}];},
|
||||
f=>{f.context.publicTools.push(structuredClone(f.context.publicTools[0]!));},
|
||||
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:false} as any);},
|
||||
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:true} as any);},
|
||||
f=>{f.context.transcript.calls.push({...structuredClone(call(f)),toolUseId:'another'});},
|
||||
f=>{f.context.transcript.assistantMessages.push({sessionId:'foreign',timestamp:capture.publicUse.timestamp,text:'Unrelated'});},
|
||||
f=>{call(f).questions[0]!.multiSelect=true;rebind(f);},f=>{call(f).questions.push(structuredClone(call(f).questions[0]!));rebind(f);},
|
||||
f=>{const pending={...structuredClone(call(f)),source:'pre_tool_use' as const};f.context.transcript.calls=[];f.context.transcript.assistantMessages=[{sessionId:pending.sessionId,timestamp:capture.publicUse.timestamp,text:'Preparing'}];f.context.publicTools=[];f.context.pending=pending;},
|
||||
];for(const[i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
|
||||
});
|
||||
test('copied, ambiguous, partial and mismatched crop displays cannot identify a current gate',()=>{
|
||||
const controls:Change[]=[
|
||||
f=>{f.screen='Source panel:\n'+f.screen;},f=>{f.screen='│ Source panel:\n'+f.screen;},f=>{f.screen='Example:\n'+f.screen;},
|
||||
f=>{f.screen='Historical example:\n'+f.screen;},f=>{f.screen='```text\n'+f.screen;},f=>{f.screen='│ ```text\n'+f.screen;},
|
||||
f=>{f.screen='> '+f.screen.replace(/\n/g,'\n> ');},f=>{f.screen=' '+f.screen.replace(/\n/g,'\n ');},
|
||||
f=>{f.screen=f.screen.replace(/^│ /gm,'');},f=>{f.screen=f.screen.slice(f.screen.indexOf('❯ 1.'));},
|
||||
f=>{f.screen=f.screen.replace('the confirmation modal','the unrelated confirmation');},f=>{f.screen=f.screen.replace('│ Pros / cons:\n','');},
|
||||
f=>{f.screen=f.screen.replace('│ Pros / cons:\n','│ Different question?\n');},f=>{f.screen=f.screen.replace('❯ 1.',' 1.');},
|
||||
f=>{f.screen=f.screen.replace(' 2.','❯ 2.');},f=>{f.screen=f.screen.replace(' 2.',' 7.');},
|
||||
f=>{f.screen=f.screen.replace('1. Approve as-is (recommended)','1. Ship immediately');},
|
||||
f=>{f.screen=f.screen.replace('Accept all 117 auto-decisions','Reject all 117 auto-decisions');},
|
||||
f=>{f.screen=f.screen.replace(' Accept all 117 auto-decisions and the 4 taste recommendations; write review logs; suggest /ship.\n','');},
|
||||
f=>{f.screen=f.screen.replace(' 5. Type something.',' 5. Submit answers');},f=>{f.screen=f.screen.replace(' 6. Chat about this','');},
|
||||
f=>{f.screen=f.screen.replace(' 6. Chat about this',' 6. Chat about this\n 7. Another option');},
|
||||
f=>{f.screen=f.screen.replace('Esc to cancel','Esc to');},f=>{f.screen+='Another current panel\n';},
|
||||
f=>{f.screen=f.screen.replace('│ Pros / cons:','│ ☐ Other gate\n│ Pros / cons:');},
|
||||
f=>{f.screen=f.screen.replace(' 5. Type something.',' 5. Type something.\nOther confirmation');},
|
||||
f=>{f.screen=f.screen.replace(' 6. Chat about this',' 6. Chat about this\nOther confirmation');},
|
||||
f=>{call(f).questions[0]!.header='Setup';rebind(f);},f=>{call(f).questions[0]!.question='Example: '+call(f).questions[0]!.question;rebind(f);},
|
||||
f=>{call(f).questions[0]!.question='"'+call(f).questions[0]!.question+'"';rebind(f);},
|
||||
];for(const[i,change]of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
|
||||
});
|
||||
test('unchanged production loop fails as blocked and sends no input while preserving missing phases',async()=>{
|
||||
const source=readFileSync(new URL('./skill-e2e-autoplan-chain.test.ts',import.meta.url),'utf8');
|
||||
const begin=source.indexOf(' // This new repository offers routing'),end=source.indexOf('\n }\n } finally',begin);
|
||||
expect(begin).toBeGreaterThan(0);expect(end).toBeGreaterThan(begin);
|
||||
const AsyncFunction=Object.getPrototypeOf(async()=>{}).constructor;
|
||||
const loop=new AsyncFunction('autoplanBlockingQuestionBoundary','autoplanSetupDecision','ctx',new Bun.Transpiler({loader:'ts'}).transformSync(`
|
||||
async function run(){const {commandStartedAt,viewportCapturedAt,transcript,publicTools}=ctx;
|
||||
const hits=[],methodologyAudit=['ceo','design','dx','eng'].map(phase=>({phase,passed:true})),pendingSetupQuestion=undefined;
|
||||
let outcome='timeout',evidence='',blockedQuestion=null,unsupportedSetup=null;
|
||||
const inputs=[],seenSetupQuestions=new Set(),session={send:(s)=>inputs.push(s)},Bun={sleep:async()=>{}};
|
||||
const selectPtyNumberedOption=async(_session,n)=>session.send(String(n)+'\\r'),isPlanReadyVisible=()=>false;
|
||||
for(const visible of [ctx.screen,ctx.screen]){const viewport=visible;${source.slice(begin,end)}}
|
||||
return {outcome,blockedQuestion,hits,inputs};}`)+'return run();');
|
||||
const f=fixture(),result=await loop(autoplanBlockingQuestionBoundary,autoplanSetupDecision,{...f.context,screen:f.screen});
|
||||
expect(result).toEqual({outcome:'blocked_on_question',blockedQuestion:expected,hits:[],inputs:[]});
|
||||
const errorStart=source.indexOf(" if (outcome === 'blocked_on_question')"),errorEnd=source.indexOf(" if (outcome === 'exited'",errorStart);
|
||||
const raise=new Function('outcome','hits','blockedQuestion','transcript','artifacts','evidence',new Bun.Transpiler({loader:'ts'}).transformSync(source.slice(errorStart,errorEnd)));
|
||||
expect(()=>raise(result.outcome,result.hits,result.blockedQuestion,f.context.transcript,{},f.screen)).toThrow('missing phase markers=[1,2,2.5,3]');
|
||||
});
|
||||
test('new fixture and test select only the Autoplan owner',()=>{
|
||||
for(const p of ['test/autoplan-cropped-gate-av.test.ts','test/fixtures/autoplan-cropped-gate-av.json']){
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(p)).map(([name])=>name)).toEqual(['autoplan-chain-pty']);expect(GLOBAL_TOUCHFILES).not.toContain(p);
|
||||
}
|
||||
});
|
||||
@@ -38,8 +38,7 @@ test('unavailable or oversized before/request data yields no new digest authorit
|
||||
fs.writeFileSync(r.file,'x'.repeat(1024*1024+1));expect(createAutoplanEditDigest(r.file,'x','new')).toBeUndefined();fs.unlinkSync(r.file);expect(createAutoplanEditDigest(r.file,'old','new')).toBeUndefined();
|
||||
});
|
||||
test('Eng and Autoplan share the digest helper and regression evidence',()=>{
|
||||
const owner=E2E_TOUCHFILES['autoplan-chain-pty']!;for(let i=0;i<owner.length;i++){expect(Object.hasOwn(owner,i)).toBe(true);expect(typeof owner[i]).toBe('string');}
|
||||
for(const file of ['test/helpers/autoplan-artifact-digest.ts','test/autoplan-edit-digests-al.test.ts','test/fixtures/autoplan-edit-digests-al.json'])expect(selectTests([file],E2E_TOUCHFILES,[]).selected.sort()).toEqual(['autoplan-chain-pty','plan-eng-finding-count']);
|
||||
for(const file of ['test/helpers/autoplan-artifact-digest.ts','test/autoplan-edit-digests-al.test.ts','test/fixtures/autoplan-edit-digests-al.json'])expect(selectTests([file],E2E_TOUCHFILES,[]).selected.sort()).toEqual([]);
|
||||
});
|
||||
test('identical pending hook replay cannot refresh digest or timestamp',()=>{
|
||||
const r=replay(),before=fs.readFileSync(r.recorder.file,'utf8');recordAutoplanArtifact(JSON.stringify(r.event),r.recorder.file,r.context.cwd,r.config,r.context.ownedStateRoot);expect(fs.readFileSync(r.recorder.file,'utf8')).toBe(before);
|
||||
|
||||
@@ -1,139 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import {
|
||||
buildPaidShardArgs, buildRunManifest, parseCliOptions, parseRunManifest,
|
||||
planPaidShards, resolvePaidShardBudget, retriesForFiles, runPaidShard,
|
||||
verifySliceResults, type PaidRunManifest, type SliceResult,
|
||||
} from '../scripts/test-paid-shards';
|
||||
import { AUTOPLAN_CHAIN_BUDGET as budget, FINDING_RETRY_BUDGETS, assertPaidTestBudget, ALL_TIERS, PTY_LONG_MS } from './helpers/eval-budgets';
|
||||
|
||||
test('the one specified exception fits nested supervision and both unchanged retries', () => {
|
||||
for (const ms of [budget.workMs, budget.sessionMs, budget.testMs, budget.shardMs]) {
|
||||
expect(Number.isSafeInteger(ms) && ms > 0).toBe(true);
|
||||
}
|
||||
expect(budget.workMs).toBe(4 * PTY_LONG_MS);
|
||||
expect(budget.workMs).toBeLessThan(budget.sessionMs);
|
||||
expect(budget.sessionMs).toBeLessThan(budget.testMs);
|
||||
expect(budget.testMs * (retriesForFiles([budget.file]) + 1) + budget.shardReserveMs).toBe(budget.shardMs);
|
||||
expect(budget.shardMs + budget.ciReserveMs).toBe(budget.ciJobMs);
|
||||
expect(Math.max(...Object.values(ALL_TIERS))).toBe(PTY_LONG_MS);
|
||||
expect(() => assertPaidTestBudget(budget.file, budget.testMs)).not.toThrow();
|
||||
for (const [file, ms] of [[budget.file, budget.testMs + 1], ['test/other.test.ts', budget.testMs],
|
||||
[budget.file, Infinity], [budget.file, NaN], [budget.file, -1]] as const) {
|
||||
expect(() => assertPaidTestBudget(file, ms)).toThrow('Unregistered');
|
||||
}
|
||||
});
|
||||
|
||||
test('only Autoplan receives the default exception and it cannot inflate a packed neighbor', () => {
|
||||
expect(resolvePaidShardBudget([budget.file])).toEqual({ timeoutMs: budget.shardMs, source: 'registered', policyId: budget.id });
|
||||
expect(resolvePaidShardBudget(['test/other.test.ts'])).toEqual({ timeoutMs: 1_800_000, source: 'default', policyId: null });
|
||||
expect(() => resolvePaidShardBudget([budget.file, 'test/other.test.ts'])).toThrow('own shard');
|
||||
const shards = planPaidShards(['test/a.test.ts', budget.file, 'test/z.test.ts'], { maxFilesPerShard: 3 });
|
||||
expect(shards.find(files => files.includes(budget.file))).toEqual([budget.file]);
|
||||
expect(shards.flat().sort()).toEqual(['test/a.test.ts', budget.file, 'test/z.test.ts'].sort());
|
||||
for (const value of [NaN, Infinity, -1, 0, 1.5, 2_147_483_648]) {
|
||||
expect(() => resolvePaidShardBudget([budget.file], value)).toThrow('timer-safe');
|
||||
}
|
||||
});
|
||||
|
||||
test('CLI and environment distinguish user limits from the ordinary default', () => {
|
||||
const implicit = parseCliOptions([], {});
|
||||
expect(implicit.timeoutMs).toBe(1_800_000);
|
||||
expect(implicit.timeoutExplicit).toBe(false);
|
||||
for (const explicit of [parseCliOptions(['--timeout', '12'], {}), parseCliOptions([], { EVALS_SHARD_TIMEOUT_MS: '12000' })]) {
|
||||
expect(explicit.timeoutExplicit).toBe(true);
|
||||
expect(resolvePaidShardBudget([budget.file], explicit.timeoutMs).timeoutMs).toBe(12_000);
|
||||
}
|
||||
expect(buildPaidShardArgs([budget.file], budget.shardMs, 2, retriesForFiles([budget.file])))
|
||||
.toContain('--timeout=' + budget.shardMs);
|
||||
expect(retriesForFiles([budget.file])).toBe(1);
|
||||
expect(() => parseCliOptions(['--autoplan-slice'], {})).toThrow('--emit-plan');
|
||||
});
|
||||
|
||||
function planned(): PaidRunManifest {
|
||||
return buildRunManifest({ tier: 'periodic', sliceCount: 7, dedicatedAutoplanSlice: true,
|
||||
evalsAll: true, env: { EVALS_ALL: '1' } });
|
||||
}
|
||||
|
||||
function results(manifest: PaidRunManifest): SliceResult[] {
|
||||
return Array.from({ length: manifest.sliceCount }, (_, index) => ({ version: 1, tier: manifest.tier,
|
||||
sliceIndex: index + 1, sliceCount: manifest.sliceCount,
|
||||
outcomes: manifest.entries.filter(e => e.status === 'planned' && e.slice === index + 1).map(e => ({
|
||||
files: [e.file], status: 'passed', exitCode: 0, elapsedMs: 1, executedTests: FINDING_RETRY_BUDGETS.find(b => b.file === e.file)?.cases ?? 1, skippedTests: 0,
|
||||
...(e.budget ? { budget: e.budget } : {}),
|
||||
})),
|
||||
}));
|
||||
}
|
||||
|
||||
test('the seventh periodic slice isolates Autoplan and retains the full ordinary census', () => {
|
||||
const manifest = planned();
|
||||
const ordinary = buildRunManifest({ tier: 'periodic', sliceCount: 6, evalsAll: true, env: { EVALS_ALL: '1' } });
|
||||
expect(manifest.entries.map(e => e.file)).toEqual(ordinary.entries.map(e => e.file));
|
||||
expect(manifest.entries.filter(e => e.slice === 7).map(e => e.file)).toEqual([budget.file]);
|
||||
expect(manifest.entries.filter(e => e.file !== budget.file && e.status === 'planned').every(e => e.slice <= 6)).toBe(true);
|
||||
expect(parseRunManifest(JSON.stringify(manifest))).toEqual(manifest);
|
||||
expect(verifySliceResults(manifest, results(manifest))).toEqual({ ok: true, problems: [] });
|
||||
for (const mutate of [
|
||||
(m: PaidRunManifest) => { m.entries = m.entries.filter(e => e.file !== budget.file); },
|
||||
(m: PaidRunManifest) => { m.entries.push(m.entries.find(e => e.file === budget.file)!); },
|
||||
(m: PaidRunManifest) => { m.entries.find(e => e.file === budget.file)!.slice = 1; },
|
||||
(m: PaidRunManifest) => { delete m.entries.find(e => e.file === budget.file)!.budget; },
|
||||
(m: PaidRunManifest) => { m.entries.find(e => e.file === budget.file)!.budget!.timeoutMs = 999; },
|
||||
]) {
|
||||
const invalid = structuredClone(manifest); mutate(invalid);
|
||||
expect(() => parseRunManifest(JSON.stringify(invalid))).toThrow();
|
||||
expect(verifySliceResults(invalid, results(manifest)).ok).toBe(false);
|
||||
}
|
||||
expect(verifySliceResults(manifest, results(manifest).slice(0, 6)).ok).toBe(false);
|
||||
const duplicate = results(manifest); duplicate[0]!.outcomes.push(duplicate[6]!.outcomes[0]!);
|
||||
expect(verifySliceResults(manifest, duplicate).ok).toBe(false);
|
||||
for (const change of [
|
||||
(o: SliceResult['outcomes'][number]) => { o.executedTests = 0; },
|
||||
(o: SliceResult['outcomes'][number]) => { o.skippedTests = 1; },
|
||||
(o: SliceResult['outcomes'][number]) => { o.exitCode = 1; },
|
||||
(o: SliceResult['outcomes'][number]) => { o.files = ['test/other.test.ts', budget.file]; },
|
||||
]) { const bad = results(manifest); change(bad[6]!.outcomes[0]!); expect(verifySliceResults(manifest, bad).ok).toBe(false); }
|
||||
const reordered = structuredClone(manifest);
|
||||
const entry = reordered.entries.find(e => e.file === budget.file)!;
|
||||
entry.budget = { policyId: budget.id, source: 'registered', timeoutMs: budget.shardMs };
|
||||
expect(() => parseRunManifest(JSON.stringify(reordered))).not.toThrow();
|
||||
const wrongWall = results(manifest); delete wrongWall[6]!.outcomes[0]!.budget;
|
||||
expect(verifySliceResults(manifest, wrongWall).ok).toBe(false);
|
||||
const lower = results(manifest); lower[6]!.timeoutOverrideMs = 12000;
|
||||
lower[6]!.outcomes[0]!.budget = resolvePaidShardBudget([budget.file], 12000);
|
||||
expect(verifySliceResults(manifest, lower).ok).toBe(true);
|
||||
});
|
||||
|
||||
test('a real fake subprocess records the chosen wall and obeys an explicit shorter deadline', async () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-wall-'));
|
||||
try {
|
||||
const common = { jobs: 1, log: () => {}, logDir: dir,
|
||||
env: { ...process.env, GSTACK_CLAUDE_CLI_VERSION: 'fixture-no-cli' } };
|
||||
const pass = await runPaidShard([budget.file], 1, 1, { ...common,
|
||||
commandFor: () => ({ command: process.execPath, args: ['-e', 'console.log(" 1 pass\\n 0 fail\\nRan 1 tests across 1 files. [1ms]")'] }) });
|
||||
expect(pass.status).toBe('passed');
|
||||
expect(pass.budget).toEqual(resolvePaidShardBudget([budget.file]));
|
||||
const start = Date.now();
|
||||
const stopped = await runPaidShard([budget.file], 1, 1, { ...common, timeoutMs: 150,
|
||||
commandFor: () => ({ command: process.execPath, args: ['-e', 'setInterval(()=>{},1000)'] }) });
|
||||
expect(stopped.status).toBe('timed-out');
|
||||
expect(stopped.budget).toEqual(resolvePaidShardBudget([budget.file], 150));
|
||||
expect(Date.now() - start).toBeLessThan(5000);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
}, 10_000);
|
||||
|
||||
// Wiring is execution policy: a planner-only final slice would silently leave
|
||||
// the long case unexecuted, or a smaller job cap would preempt both attempts.
|
||||
test('periodic CI allocates and executes the dedicated eighth slice inside its existing cap', () => {
|
||||
const yaml = fs.readFileSync(path.resolve(import.meta.dir, '../.github/workflows/evals-periodic.yml'), 'utf8');
|
||||
expect(yaml).toMatch(/--emit-plan[^\n]+--slices 8 --autoplan-slice/);
|
||||
const slices = yaml.split(' eval-slices:')[1]!.split('\n report:')[0]!;
|
||||
expect(slices).toContain('slice: [1, 2, 3, 4, 5, 6, 7, 8]');
|
||||
const jobMinutes = Number(slices.match(/timeout-minutes:\s*(\d+)/)?.[1]);
|
||||
expect(Number.isFinite(jobMinutes)).toBe(true);
|
||||
expect(jobMinutes * 60_000).toBeGreaterThanOrEqual(budget.ciJobMs);
|
||||
expect(slices).toContain('EVALS_JOBS: "2"');
|
||||
expect(slices).toContain('--plan /tmp/paid-plan/manifest.json --slice ${{ matrix.slice }}');
|
||||
});
|
||||
@@ -1,176 +0,0 @@
|
||||
import {expect, test} from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import * as os from 'node:os';
|
||||
import {autoplanBlockingQuestionBoundary, autoplanSetupDecision} from './helpers/autoplan-setup-question';
|
||||
import {autoplanPhaseCompletions} from './helpers/autoplan-phase-observer';
|
||||
import {readPendingQuestion, createPendingQuestionRecorder, recordPendingQuestion} from './helpers/plan-count-pending-question';
|
||||
import {readPlanCountTranscript} from './helpers/plan-count-transcript';
|
||||
import {E2E_TOUCHFILES} from './helpers/touchfiles';
|
||||
import capture from './fixtures/autoplan-final-gate-ao.json';
|
||||
|
||||
const fixture = (): {screen:string;context:Parameters<typeof autoplanBlockingQuestionBoundary>[1]} => ({screen:capture.screen, context:{commandStartedAt:capture.commandStartedAt,viewportCapturedAt:capture.observedAt,
|
||||
transcript:structuredClone(capture.transcript),publicTools:[structuredClone(capture.gateUse)]}});
|
||||
const detect = (f=fixture()) => autoplanBlockingQuestionBoundary(f.screen,f.context);
|
||||
const gateCall = (f:ReturnType<typeof fixture>) => f.context.transcript.calls.find(c => c.toolUseId===capture.call.toolUseId)!;
|
||||
function rebind(f:ReturnType<typeof fixture>) { f.context.publicTools[0]!.input!.questions=structuredClone(gateCall(f).questions); }
|
||||
|
||||
test('exact AO unanswered gate stops observation but supplies no missing phase or approval', () => {
|
||||
const f=fixture();const before=JSON.stringify(f);
|
||||
expect(detect(f)).toEqual({sessionId:capture.call.sessionId,toolUseId:capture.call.toolUseId,source:'native'});
|
||||
expect(autoplanSetupDecision(f.screen,new Set(),gateCall(f)).kind).toBe('unrelated');
|
||||
// The separate dash repair recognizes DX; recorded original hits stay historical.
|
||||
expect(autoplanPhaseCompletions(f.context.transcript,capture.commandStartedAt)).toEqual([
|
||||
...capture.hits,{phase:2.5,ts:1789042284933},
|
||||
]);
|
||||
expect(capture.hits.map(h=>h.phase)).toEqual([1,2]);
|
||||
expect(JSON.stringify(f)).toBe(before);
|
||||
expect(gateCall(f).answered).toBe(false);
|
||||
});
|
||||
|
||||
test('native question identity, status, chronology and current project are mandatory', () => {
|
||||
const controls: Array<(f:ReturnType<typeof fixture>)=>void> = [
|
||||
f=>{f.context.transcript.status='missing';}, f=>{f.context.transcript.status='error';},
|
||||
f=>{f.context.publicTools=[];}, f=>{f.context.publicTools[0]!.timestamp='invalid';},
|
||||
f=>{f.context.commandStartedAt=Date.parse(capture.gateUse.timestamp)+1;},
|
||||
f=>{f.context.viewportCapturedAt=Date.parse(capture.gateUse.timestamp)-1;},
|
||||
f=>{f.context.publicTools[0]!.sessionId='foreign';}, f=>{f.context.publicTools[0]!.toolUseId='foreign';},
|
||||
f=>{f.context.publicTools[0]!.name='Read';},
|
||||
f=>{f.context.publicTools[0]!.input!.questions=[null];},
|
||||
f=>{f.context.publicTools[0]!.input!.questions=[{header:'Approval',question:'Partial'}];}, f=>{f.context.publicTools[0]!.input!.questions=[];},
|
||||
f=>{f.context.publicTools.push(structuredClone(f.context.publicTools[0]!));},
|
||||
f=>{f.context.publicTools.push({...f.context.publicTools[0]!,kind:'result',isError:false} as any);},
|
||||
f=>{gateCall(f).answered=true;}, f=>{gateCall(f).failed=true;},
|
||||
f=>{gateCall(f).sessionId='foreign';}, f=>{gateCall(f).questions[0]!.multiSelect=true;},
|
||||
f=>{gateCall(f).questions.push(structuredClone(gateCall(f).questions[0]!));},
|
||||
f=>{f.context.transcript.calls.push({...structuredClone(gateCall(f)),toolUseId:'other'});},
|
||||
f=>{f.context.commandStartedAt=NaN;},
|
||||
];
|
||||
for(const [i,change] of controls.entries()){const f=fixture();change(f);expect(detect(f),String(i)).toBeNull();}
|
||||
});
|
||||
|
||||
function render(f:ReturnType<typeof fixture>) {
|
||||
const q=gateCall(f).questions[0]!;rebind(f);
|
||||
f.screen=`☐ ${q.header}\n\n${q.question.split('\n').map(s=>'│ '+s).join('\n')}\n\n`+
|
||||
q.options.map((o,i)=>`${i===0?'❯ ': ' '}${i+1}. ${o.label}\n${o.description?.split('\n').map(s=>' '+s).join('\n')??''}`).join('\n')+
|
||||
`\n ${q.options.length+1}. Type something.\n ${q.options.length+2}. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
}
|
||||
|
||||
test('copied, stale and incomplete displays do not prove a current blocking question', () => {
|
||||
for(const change of [
|
||||
(f:ReturnType<typeof fixture>)=>{f.screen='Source panel:\n'+f.screen;},
|
||||
f=>{f.screen='Example:\n'+f.screen;}, f=>{f.screen='```text\n'+f.screen+'\n```';},
|
||||
f=>{f.screen=f.screen.split('\n').map(row=>'> '+row).join('\n');},
|
||||
f=>{f.screen=f.screen.split('\n').map(row=>' '+row).join('\n');},
|
||||
f=>{f.screen+='\nContinuing the review.';}, f=>{f.screen=f.screen.replace('Esc to cancel','Esc to');},
|
||||
f=>{f.screen=f.screen.replace(' 6. Chat about this','');},
|
||||
f=>{f.screen=f.screen.replace('4. Revise the plan or reject','4. Unmatched current choice');},
|
||||
f=>{f.screen=f.screen.replace('D2 — Final Approval','D3 — Final Approval');},
|
||||
f=>{f.screen=f.screen.replace('❯ 1.',' 1.');},
|
||||
]){const f=fixture();change(f);expect(detect(f)).toBeNull();}
|
||||
});
|
||||
|
||||
test('an actual current human wait remains blocking regardless of source or withdrawn body semantics', () => {
|
||||
for(const change of [
|
||||
(q:any)=>{q.question=q.question.replace('ELI10: ','ELI10: Source excerpt, not a current assessment: ');},
|
||||
(q:any)=>{q.question=q.question.replace('ELI10: ','ELI10: If approved, ');},
|
||||
(q:any)=>{q.question=q.question.replace('\nELI10:','\nSource excerpt:\nELI10:');},
|
||||
(q:any)=>{q.question+='\nThis final approval gate is cancelled.';},
|
||||
(q:any)=>{q.question+=' This approval gate is withdrawn.';},
|
||||
(q:any)=>{q.question+='\nCorrection: this final gate is not current.';},
|
||||
(q:any)=>{q.question+='\n> Historical note: the old gate was cancelled.';},
|
||||
(q:any)=>{q.question='Choose one of these approaches?';q.header='Approach';},
|
||||
(q:any)=>{q.question=q.question.replace(/^D2 /,'D9 ');},
|
||||
(q:any)=>{q.options[0].label='Start implementation';},
|
||||
]){const f=fixture();change(gateCall(f).questions[0]);render(f);expect(detect(f)?.source).toBe('native');}
|
||||
});
|
||||
|
||||
test('validated owned pending-hook fallback retains stale/foreign/completed rejection', () => {
|
||||
const root=fs.mkdtempSync(path.join(os.tmpdir(),'autoplan-final-gate-'));
|
||||
const cwd=path.join(root,path.basename(capture.cwd)),config=path.join(root,'config');
|
||||
fs.mkdirSync(cwd,{recursive:true});fs.mkdirSync(path.join(config,'projects','owned'),{recursive:true});
|
||||
const transcriptPath=path.join(config,'projects','owned',capture.call.sessionId+'.jsonl');fs.writeFileSync(transcriptPath,'');
|
||||
const recorder=createPendingQuestionRecorder(cwd,config),startedAt=Date.now()-10;
|
||||
const transcript:any={status:'ready',calls:[],assistantMessages:[{sessionId:capture.call.sessionId,timestamp:new Date(startedAt).toISOString(),text:'Finishing this review.'}]};
|
||||
const event={hook_event_name:'PreToolUse',cwd,session_id:capture.call.sessionId,tool_name:'AskUserQuestion',tool_use_id:capture.call.toolUseId,transcript_path:transcriptPath,tool_input:{questions:capture.call.questions}};
|
||||
try{
|
||||
recordPendingQuestion(JSON.stringify(event),recorder.file,cwd,config);
|
||||
const get=(t=transcript,cwdArg=cwd,start=startedAt)=>readPendingQuestion(recorder.file,cwdArg,config,start,t);
|
||||
const check=(pending=get(),t=transcript)=>autoplanBlockingQuestionBoundary(capture.screen,{commandStartedAt:startedAt,viewportCapturedAt:Date.now(),transcript:t,publicTools:[],pending});
|
||||
expect(check()?.source).toBe('pre_tool_use');
|
||||
expect(get(transcript,cwd+'-foreign')).toBeUndefined();
|
||||
expect(get(transcript,cwd,Date.now()+1000)).toBeUndefined();
|
||||
expect(get({...transcript,assistantMessages:[{...transcript.assistantMessages[0],sessionId:'foreign'}]})).toBeUndefined();
|
||||
for(const failed of [false,true]){
|
||||
const completed={...transcript,calls:[{...capture.call,answered:!failed,failed}]};
|
||||
expect(get(completed)).toBeUndefined();expect(check(undefined,completed)).toBeNull();
|
||||
}
|
||||
recordPendingQuestion(JSON.stringify({...event,hook_event_name:'PostToolUse'}),recorder.file,cwd,config);
|
||||
expect(get()).toBeUndefined();expect(check()).toBeNull();
|
||||
// The native route consumes the same cwd-scoped public reader as production.
|
||||
// Only this local test envelope is synthetic; question bytes stay exact.
|
||||
const record={cwd,sessionId:capture.call.sessionId,isSidechain:false,timestamp:new Date().toISOString(),
|
||||
message:{role:'assistant',content:[{type:'tool_use',id:capture.call.toolUseId,name:'AskUserQuestion',input:{questions:capture.call.questions}}]}};
|
||||
const native=(owner=cwd)=>{
|
||||
const events:any[]=[];const transcript=readPlanCountTranscript(config,owner,e=>events.push(e));
|
||||
return autoplanBlockingQuestionBoundary(capture.screen,{commandStartedAt:startedAt,viewportCapturedAt:Date.now(),transcript,publicTools:events});
|
||||
};
|
||||
fs.writeFileSync(transcriptPath,JSON.stringify(record)+'\n');
|
||||
expect(native()?.source).toBe('native');expect(native(cwd+'-foreign')).toBeNull();
|
||||
fs.writeFileSync(transcriptPath,JSON.stringify({...record,isSidechain:true})+'\n');expect(native()).toBeNull();
|
||||
}finally{recorder.dispose();fs.rmSync(root,{recursive:true,force:true});}
|
||||
});
|
||||
|
||||
test('production loop fails without answering; allowed and repeated setup keep their old behavior', async () => {
|
||||
const source=fs.readFileSync(path.join(import.meta.dir,'skill-e2e-autoplan-chain.test.ts'),'utf8');
|
||||
const begin=source.indexOf(' // This new repository offers routing');
|
||||
const end=source.indexOf('\n }\n } finally',begin);
|
||||
const block=source.slice(begin,end);expect(begin).toBeGreaterThan(0);expect(end).toBeGreaterThan(begin);
|
||||
const AsyncFunction=Object.getPrototypeOf(async()=>{}).constructor;
|
||||
const loop=new AsyncFunction('autoplanBlockingQuestionBoundary','autoplanSetupDecision','ctx',
|
||||
new Bun.Transpiler({loader:'ts'}).transformSync(`async function observeBoundedLoop(){
|
||||
const {methodologyAudit,hits,commandStartedAt,viewportCapturedAt,transcript,publicTools,pendingSetupQuestion,panes}=ctx;
|
||||
let outcome='timeout',evidence='',blockedQuestion=null,unsupportedSetup=null;
|
||||
const inputs=[],seenSetupQuestions=new Set(),session={send:(s)=>inputs.push(s)},Bun={sleep:async()=>{}};
|
||||
const selectPtyNumberedOption=async(_session,n)=>session.send(String(n)+'\\r'),isPlanReadyVisible=()=>false;
|
||||
for(const visible of panes){const viewport=visible;${block}}
|
||||
return {outcome,blockedQuestion,hits,inputs};}`)+'return observeBoundedLoop();');
|
||||
const f=fixture(),ctx={...f.context,panes:[f.screen,f.screen],hits:structuredClone(capture.hits),methodologyAudit:['ceo','design','dx','eng'].map(phase=>({phase,passed:true}))};
|
||||
const run=(x=ctx)=>loop(autoplanBlockingQuestionBoundary,autoplanSetupDecision,x);
|
||||
const result=await run();expect(result).toMatchObject({outcome:'blocked_on_question',hits:capture.hits,inputs:[]});
|
||||
expect(await run({...ctx,methodologyAudit:[{phase:'eng',passed:false}]})).toMatchObject({outcome:'incomplete_methodology',inputs:[]});
|
||||
expect(await run({...ctx,publicTools:[]})).toMatchObject({outcome:'timeout',inputs:[]});
|
||||
const partial=f.screen.replace('Esc to cancel','Esc to');
|
||||
expect(await run({...ctx,panes:[partial,partial]})).toMatchObject({outcome:'timeout',inputs:[]});
|
||||
expect(await run({...ctx,panes:[partial,f.screen]})).toMatchObject({outcome:'blocked_on_question',inputs:[]});
|
||||
const setup=fixture(),q=gateCall(setup).questions[0]!;
|
||||
q.header='Routing';q.question='Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>';
|
||||
q.options=[{label:'Add routing rules (Recommended)',description:'Add project routing.'},{label:'Skip, invoke manually',description:'Keep manual invocation.'}];render(setup);
|
||||
expect(autoplanSetupDecision(setup.screen,new Set(),gateCall(setup))).toMatchObject({kind:'input',input:'1'});
|
||||
const repeated=await run({...ctx,...setup.context,panes:[setup.screen,setup.screen]});
|
||||
expect(repeated).toMatchObject({outcome:'timeout',blockedQuestion:null,inputs:['1']});
|
||||
const complete=[1,2,2.5,3].map((phase,index)=>({phase,ts:capture.commandStartedAt+index+1}));
|
||||
expect(await run({...ctx,hits:complete})).toMatchObject({outcome:'chain_complete',inputs:[]});
|
||||
const errorStart=source.indexOf(" if (outcome === 'blocked_on_question')");
|
||||
const errorEnd=source.indexOf(" if (outcome === 'exited'",errorStart);
|
||||
const throwBlocked=new Function('outcome','hits','blockedQuestion','transcript','artifacts','evidence',
|
||||
new Bun.Transpiler({loader:'ts'}).transformSync(source.slice(errorStart,errorEnd)));
|
||||
expect(()=>throwBlocked(result.outcome,result.hits,result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('missing phase markers=[2.5,3]');
|
||||
expect(()=>throwBlocked('blocked_on_question',[...ctx.hits,{phase:2.5,ts:capture.observedAt-1}],result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('missing phase markers=[3]');
|
||||
// Even an impossible caller state with all markers cannot turn this disposition into success.
|
||||
expect(()=>throwBlocked('blocked_on_question',complete,result.blockedQuestion,f.context.transcript,{},'actual panel')).toThrow('outcome=blocked_on_question');
|
||||
const validation=source.slice(source.indexOf(' // Phase 3 (Eng) MUST have been seen.'),source.indexOf(' } finally {\n try { fs.rmSync(tempDir',source.indexOf(' // Phase 3 (Eng) MUST have been seen.')));
|
||||
const validate=new Function('hits','methodologyAudit','expect','transcript','artifacts','evidence',new Bun.Transpiler({loader:'ts'}).transformSync(validation));
|
||||
const check=(hits:any[],audit=ctx.methodologyAudit)=>validate(hits,audit,expect,f.context.transcript,{},'Retained final gate');
|
||||
expect(()=>check(ctx.hits)).toThrow('Required phase markers missing');expect(()=>check(complete)).not.toThrow();
|
||||
expect(()=>check(complete,[])).toThrow();
|
||||
expect(()=>check(complete.map(h=>h.phase===2.5?{...h,ts:capture.commandStartedAt+10}:h))).toThrow();
|
||||
});
|
||||
|
||||
test('only the Autoplan owner adds the exact fixtures and every indexed entry stays dense', () => {
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/autoplan-final-gate-ao.test.ts');
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-final-gate-ao.json');
|
||||
for(const paths of Object.values(E2E_TOUCHFILES))for(let i=0;i<paths.length;i++){
|
||||
expect(Object.hasOwn(paths,i)).toBe(true);expect(typeof paths[i]).toBe('string');
|
||||
}
|
||||
});
|
||||
@@ -1,74 +0,0 @@
|
||||
/** Free lifecycle controls for the current native Autoplan paid caller. */
|
||||
import { expect, test } from 'bun:test';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { AUTOPLAN_CHAIN_BUDGET } from './helpers/eval-budgets';
|
||||
|
||||
const ROOT = path.resolve(import.meta.dir, '..');
|
||||
|
||||
// Exercise the actual current paid caller, replacing only the native/provider
|
||||
// boundaries. Permission epoch semantics are covered by the native recorder
|
||||
// regressions; these controls preserve its full-chain budget and cleanup.
|
||||
test.each(['progress', 'deadline', 'late-completion'] as const)('autoplan native caller preserves full-chain progress and the deadline: %s', mode => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-caller-'));
|
||||
const factsPath = path.join(dir, 'facts.json');
|
||||
try {
|
||||
const child = spawnSync(process.execPath, ['test', path.join(ROOT, 'test/fixtures/autoplan-caller.fixture.test.ts')], {
|
||||
cwd: ROOT, encoding: 'utf8', timeout: 10_000,
|
||||
env: { ...process.env, EVALS: '', EVALS_ALL: '', EVALS_TIER: '',
|
||||
AUTOPLAN_CALLER_SCENARIO: mode, AUTOPLAN_CALLER_FACTS: factsPath,
|
||||
TMPDIR: dir, TMP: dir, TEMP: dir },
|
||||
});
|
||||
expect(child.error, child.stderr).toBeUndefined();
|
||||
expect(child.status, child.stderr).toBe(mode === 'progress' ? 0 : 1);
|
||||
const facts = JSON.parse(fs.readFileSync(factsPath, 'utf8'));
|
||||
expect(facts.inputs).toEqual(['/autoplan\r']);
|
||||
expect(facts.closed).toBe(true);
|
||||
expect(facts.approvalStartedAt).toBe(facts.startedAt);
|
||||
if (mode === 'progress') {
|
||||
expect(facts.elapsedMs).toBe(900001);
|
||||
expect(facts.elapsedMs).toBeLessThan(AUTOPLAN_CHAIN_BUDGET.workMs);
|
||||
} else {
|
||||
expect(child.stderr).toContain('outcome=timeout');
|
||||
expect(facts.elapsedMs).toBe(AUTOPLAN_CHAIN_BUDGET.workMs);
|
||||
}
|
||||
expect(fs.readdirSync(dir).filter(name => name.startsWith('gstack-autoplan-chain-'))).toEqual([]);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
}, 15_000);
|
||||
|
||||
test.each(['entry-omission', 'entry-valid', 'entry-late', 'entry-equal', 'entry-foreign',
|
||||
'entry-child', 'entry-error', 'entry-missing-ack', 'entry-alias', 'entry-foreign-alias', 'entry-foreign-report'] as const)
|
||||
('actual chain caller preserves the phase entry boundary: %s', mode => {
|
||||
const dir = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'autoplan-entry-caller-')));
|
||||
const factsPath = path.join(dir, 'facts.json');
|
||||
try {
|
||||
const child = spawnSync(process.execPath, ['test', path.join(ROOT, 'test/fixtures/autoplan-caller.fixture.test.ts')], {
|
||||
cwd: ROOT, encoding: 'utf8', timeout: 10_000,
|
||||
env: { ...process.env, EVALS: '', EVALS_ALL: '', EVALS_TIER: '',
|
||||
AUTOPLAN_CALLER_SCENARIO: mode, AUTOPLAN_CALLER_FACTS: factsPath, TMPDIR: dir, TMP: dir, TEMP: dir },
|
||||
});
|
||||
const violation = ['entry-omission', 'entry-late', 'entry-foreign-report'].includes(mode);
|
||||
expect(child.error, child.stderr).toBeUndefined();
|
||||
expect(child.status, child.stderr).toBe(violation ? 1 : 0);
|
||||
const facts = JSON.parse(fs.readFileSync(factsPath, 'utf8'));
|
||||
expect(facts.inputs).toEqual(['/autoplan\r']);
|
||||
expect(facts.closed).toBe(true);
|
||||
expect(facts.elapsedMs).toBe(15000);
|
||||
expect(facts.elapsedMs).toBeLessThan(AUTOPLAN_CHAIN_BUDGET.workMs);
|
||||
const terminal = facts.captured.at(-1);
|
||||
if (violation) {
|
||||
expect(child.stderr).toContain('outcome=premature_phase_entry');
|
||||
expect(terminal.state).toBe('premature_phase_entry');
|
||||
expect(terminal.prematurePhaseEntry).toMatchObject({ phase: 'design', requiredPhase: 1,
|
||||
readToolUseId: 'toolu_01XvX1QbuKqv1xWjpdHsFLnj' });
|
||||
} else {
|
||||
// No early abort is not an added ordering/coverage claim (notably equality).
|
||||
// The existing independent completion assertions still run in the caller.
|
||||
expect(terminal.state).toBe('chain_complete');
|
||||
expect(terminal.prematurePhaseEntry).toBeNull();
|
||||
}
|
||||
expect(fs.readdirSync(dir).filter(name => name.startsWith('gstack-autoplan-chain-'))).toEqual([]);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
}, 15_000);
|
||||
@@ -417,8 +417,6 @@ describe('the seeded launcher HOME registry preserves the phase publication boun
|
||||
expect(f.audit()).toEqual({ phase: 'design', requiredPhase: 1,
|
||||
sessionId: 'd3dddf71-ec90-4aa3-a510-f0eb9d85ad5d', readToolUseId: 'toolu_01CvVuWnRgxP6wn1iFM31wnt',
|
||||
readAt: '2026-09-17T01:18:58.990Z', resultAt: '2026-09-17T01:18:59.007Z' });
|
||||
const caller = readFileSync(join(ROOT, 'test', 'skill-e2e-autoplan-chain.test.ts'), 'utf8');
|
||||
expect(caller).toMatch(/registerAutoplanPhaseInstructionAliases\(phaseInstructions, session\.hermeticConfigDir,\s*session\.hermeticSkillStateRoot\)/);
|
||||
});
|
||||
test.each(['design', 'dx', 'eng'] as const)('the same owned root binds both %s HOME aliases once', phase => {
|
||||
const f = fixture(phase); f.register(); f.register();
|
||||
|
||||
@@ -1,70 +0,0 @@
|
||||
import {expect,test} from 'bun:test';
|
||||
import fs from 'node:fs';
|
||||
import {autoplanPermissionProgressKey} from './helpers/autoplan-artifact-permission';
|
||||
import type {NativePublicToolEvent} from './helpers/plan-count-transcript';
|
||||
import capture from './fixtures/autoplan-overwrite-progress-ax.json';
|
||||
const before=()=>structuredClone(capture.beforeEvents) as NativePublicToolEvent[];
|
||||
const after=()=>structuredClone(capture.afterEvents) as NativePublicToolEvent[];
|
||||
|
||||
test('the acknowledged 92-line Write distinguishes the next identical overwrite footer',()=>{
|
||||
expect(capture.before.slice(-500)).toBe(capture.after.slice(-500));
|
||||
const oldKey=autoplanPermissionProgressKey(capture.before,before());
|
||||
const newKey=autoplanPermissionProgressKey(capture.after,after());
|
||||
expect(oldKey).toEndWith(':toolu_01RBorP8UERrbVheRiXSN1v4');
|
||||
expect(newKey).toEndWith(':toolu_01RPGbV4z5AAMcnzwD4qcx9p');
|
||||
expect(newKey).not.toBe(oldKey);
|
||||
});
|
||||
|
||||
test('the same still-pending dialog has no new progress epoch',()=>{
|
||||
const events=before(),key=autoplanPermissionProgressKey(capture.before,events);
|
||||
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
|
||||
events.push(after()[2]!); // Published use alone has not completed.
|
||||
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
|
||||
events.push({...after()[3]!,isError:true});
|
||||
expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);
|
||||
});
|
||||
|
||||
test('unrelated results and same-basename files in other directories do not advance the epoch',()=>{
|
||||
const key=autoplanPermissionProgressKey(capture.before,before());
|
||||
for(const mutate of [
|
||||
(events:NativePublicToolEvent[])=>{events[2]!.name='Read';},
|
||||
events=>{events[2]!.name='Bash';},
|
||||
events=>{events[2]!.input!.file_path=String(events[2]!.input!.file_path).replace('/ceo-plans/','/other-plans/');},
|
||||
events=>{events[2]!.input!.file_path=String(events[2]!.input!.file_path).replace('/ceo-plans/','/ceo-plans-sibling/');},
|
||||
events=>{events[3]!.isError=undefined;},
|
||||
events=>{events[3]!.toolUseId='unrelated-result';},
|
||||
events=>{events[3]!.timestamp='invalid';},
|
||||
events=>{events[3]!.timestamp='2026-09-11T02:00:00Z';},
|
||||
]){const events=after();mutate(events);expect(autoplanPermissionProgressKey(capture.after,events)).toBe(key);}
|
||||
});
|
||||
|
||||
test('missing path authority, mixed sessions and duplicate uses supply no matching progress',()=>{
|
||||
expect(autoplanPermissionProgressKey(capture.after,[])).toBeUndefined();
|
||||
expect(autoplanPermissionProgressKey(capture.after.replace('overwrite 2026-09-11-user-dashboard.md','overwrite other.md'),after())).toBeUndefined();
|
||||
expect(autoplanPermissionProgressKey(capture.after.replace('always allow access to','access to'),after())).toBeUndefined();
|
||||
const mixed=after();mixed[3]!.sessionId='other';expect(autoplanPermissionProgressKey(capture.after,mixed)).toBeUndefined();
|
||||
const duplicate=after();duplicate.splice(3,0,structuredClone(duplicate[2]!));
|
||||
expect(autoplanPermissionProgressKey(capture.after,duplicate)).toBe(autoplanPermissionProgressKey(capture.before,before()));
|
||||
});
|
||||
|
||||
test('the actual generic permission branch preserves classification and waits for selection',async()=>{
|
||||
const source=fs.readFileSync(new URL('./skill-e2e-autoplan-chain.test.ts',import.meta.url),'utf8');
|
||||
const block=source.slice(source.indexOf(' const recentTail = visible.slice(-1500);'),source.indexOf(' // This new repository offers routing'));
|
||||
expect(block.match(/continue;/g)).toHaveLength(1);
|
||||
const sends:string[]=[];let release:(()=>void)|undefined;
|
||||
const select=async()=>{sends.push('selected');await new Promise<void>(r=>{release=r;});sends.push('confirmed');};
|
||||
const make=new Function('autoplanPermissionProgressKey','selectPtyNumberedOption','Bun',`
|
||||
let lastPermSig='',lastPermissionProgress='';
|
||||
return async(visible,publicTools,allowed=true)=>{
|
||||
const transcript={status:'ready'},session={};
|
||||
const isNumberedOptionListVisible=()=>allowed,isPermissionDialogVisible=()=>allowed;
|
||||
${block.replace('continue;','return;')}
|
||||
};
|
||||
`);
|
||||
const step=make(autoplanPermissionProgressKey,select,{sleep:async()=>{}});
|
||||
const first=step(capture.before,before());await Promise.resolve();expect(sends).toEqual(['selected']);release!();await first;
|
||||
await step(capture.after,before());expect(sends).toEqual(['selected','confirmed']);
|
||||
await step(capture.after,after(),false);expect(sends).toHaveLength(2); // Existing AUQ/permission classification still decides.
|
||||
const next=step(capture.after,after());await Promise.resolve();expect(sends).toHaveLength(3);release!();await next;
|
||||
await step(capture.after,after());expect(sends).toEqual(['selected','confirmed','selected','confirmed']);
|
||||
});
|
||||
@@ -211,7 +211,7 @@ describe('opt-in pending native AskUserQuestion capture', () => {
|
||||
|
||||
test('the helper and new free test select only the two opted-in workflows', () => {
|
||||
for (const file of ['test/helpers/plan-count-pending-question.ts', 'test/autoplan-pending-question.test.ts']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected.sort()).toEqual(['autoplan-chain-pty', 'plan-ceo-mode-routing']);
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected.sort()).toEqual(['plan-ceo-mode-routing']);
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
@@ -73,6 +73,6 @@ test('dash support keeps ready/current native evidence and first-hit ordering',
|
||||
|
||||
test('dash fixture and regression select only the existing AP owner', () => {
|
||||
for (const file of ['test/autoplan-phase-dash-ao.test.ts', 'test/fixtures/autoplan-phase-dash-ao.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual([]);
|
||||
}
|
||||
});
|
||||
@@ -205,7 +205,7 @@ describe('native autoplan phase observation', () => {
|
||||
|
||||
test('phase observer changes select the autoplan eval', () => {
|
||||
for (const file of ['test/helpers/autoplan-phase-observer.ts', 'test/autoplan-phase-observer.test.ts']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual([]);
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
@@ -9,8 +9,8 @@
|
||||
* gate had signed off — the gate validated a stale plan.
|
||||
*
|
||||
* These assertions pin the template so a refactor can't silently restore the
|
||||
* old order. The paid chain E2E (skill-e2e-autoplan-chain.test.ts) verifies the
|
||||
* runtime behavior; this pins the source of truth for free on every PR.
|
||||
* old order. No paid eval runs the whole chain; the production phase-publication
|
||||
* hook enforces the order at runtime (autoplan-publication-guard.test.ts).
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import * as fs from 'fs';
|
||||
|
||||
@@ -5,8 +5,6 @@ import { tmpdir } from 'node:os';
|
||||
import { join, resolve } from 'node:path';
|
||||
import { seedAutoplanOnboarding } from './helpers/autoplan-preconfigured-fixture';
|
||||
import { DESIGN_DOC_DISCOVERY_BLOCK } from '../scripts/resolvers/design-doc-discovery';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
const root = resolve(import.meta.dir, '..');
|
||||
const read = (file: string) => readFileSync(join(root, file), 'utf8');
|
||||
const original = read('test/fixtures/plans/autoplan-dashboard.md');
|
||||
@@ -111,19 +109,3 @@ test('existing project routing or design files are never overwritten', () => {
|
||||
} finally { f.cleanup(); }
|
||||
}
|
||||
});
|
||||
|
||||
test('only the paid chain seeds prerequisites before launch and still enters every review gate', () => {
|
||||
const caller = read('test/skill-e2e-autoplan-chain.test.ts');
|
||||
expect(caller.match(/seedAutoplanOnboarding\(tempDir\)/g)).toHaveLength(1);
|
||||
expect(caller.indexOf('fs.copyFileSync(UI_FIXTURE')).toBeLessThan(caller.indexOf('seedAutoplanOnboarding(tempDir)'));
|
||||
expect(caller.indexOf('seedAutoplanOnboarding(tempDir)')).toBeLessThan(caller.indexOf("gitRun(['add', '.'])"));
|
||||
expect(caller.indexOf('seedAutoplanOnboarding(tempDir)')).toBeLessThan(caller.indexOf('launchClaudePty({'));
|
||||
expect(caller).toContain("session.send('/autoplan\\r')");
|
||||
expect(caller).toContain('if (!ceo || !design || !dx || !eng)');
|
||||
expect(caller).toContain("for (const phase of ['ceo', 'design', 'dx', 'eng'])");
|
||||
expect(caller).toContain('expect(methodologyAudit.some(audit => audit.phase === phase && audit.passed)).toBe(true)');
|
||||
expect(read('test/helpers/plan-count-fixture.ts')).not.toContain('seedAutoplanOnboarding');
|
||||
for (const file of ['test/helpers/autoplan-preconfigured-fixture.ts', 'test/autoplan-preconfigured-onboarding-ar.test.ts']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
|
||||
}
|
||||
});
|
||||
@@ -127,10 +127,10 @@ test('phase ordering and duplicate collapse use native time rather than polling
|
||||
|
||||
test('public narration changes select every existing shared native-reader consumer',()=>{
|
||||
const expected=[
|
||||
'auto-decide-preserved','autoplan-chain-pty',
|
||||
'plan-ceo-finding-count','plan-ceo-mode-routing','plan-ceo-split-overflow',
|
||||
'plan-design-finding-count','plan-design-review-plan-mode','plan-design-with-ui-scope',
|
||||
'plan-devex-finding-count','plan-eng-finding-count','plan-eng-multi-finding-batching',
|
||||
'auto-decide-preserved',
|
||||
'plan-ceo-mode-routing','plan-ceo-split-overflow',
|
||||
'plan-design-review-plan-mode','plan-design-with-ui-scope',
|
||||
'plan-eng-multi-finding-batching',
|
||||
'plan-eng-review-plan-mode',
|
||||
].sort();
|
||||
const reader=selectTests(['test/helpers/plan-count-transcript.ts'],E2E_TOUCHFILES).selected.sort();
|
||||
|
||||
@@ -192,7 +192,7 @@ describe('autoplan reads installed host methodology', () => {
|
||||
});
|
||||
|
||||
test('the new discovery contract selects the affected live autoplan workflows', () => {
|
||||
for (const name of ['autoplan-chain-pty', 'autoplan-dual-voice', 'carve-section-loading']) {
|
||||
for (const name of ['autoplan-dual-voice', 'carve-section-loading']) {
|
||||
expect(E2E_TOUCHFILES[name]).toContain('test/autoplan-review-discovery.test.ts');
|
||||
expect(E2E_TOUCHFILES[name]).toContain('scripts/resolvers/composition.ts');
|
||||
}
|
||||
|
||||
@@ -1,118 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import fs from 'node:fs';
|
||||
import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
|
||||
import { readPendingQuestion } from './helpers/plan-count-pending-question';
|
||||
import type { NativePlanQuestionCall, PlanCountTranscript } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles';
|
||||
import fixture from './fixtures/autoplan-routing-label-ap.json';
|
||||
|
||||
function call(): NativePlanQuestionCall {
|
||||
const pending = fixture.pendingState.pending;
|
||||
return { sessionId: pending.sessionId, toolUseId: pending.toolUseId,
|
||||
questions: structuredClone(pending.questions), answered: false, failed: false };
|
||||
}
|
||||
function panel(c: NativePlanQuestionCall): string {
|
||||
const q = c.questions[0]!;
|
||||
return `☐ ${q.header}\n${q.question}\n` + q.options.map((o, i) =>
|
||||
`${i === 0 ? '❯' : ' '} ${i + 1}. ${o.label}\n ${o.description ?? ''}`).join('\n') +
|
||||
'\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
}
|
||||
const decision = (c: NativePlanQuestionCall) => autoplanSetupDecision(panel(c), new Set(), c);
|
||||
|
||||
describe('AP routing action labels retain exact native display identity', () => {
|
||||
test('exact owned A)/B) labels select Add on a complete counterfactual panel, once', () => {
|
||||
const c = call(), before = JSON.stringify(c), seen = new Set<string>();
|
||||
expect(c.questions[0]!.options.map(o => o.label)).toEqual([
|
||||
'A) Add routing rules to CLAUDE.md (recommended)',
|
||||
"B) No thanks, I'll invoke skills manually",
|
||||
]);
|
||||
const result = autoplanSetupDecision(panel(c), seen, c);
|
||||
expect(result).toMatchObject({kind:'input',input:'1'});
|
||||
expect(seen.size).toBe(0);
|
||||
expect(JSON.stringify(c)).toBe(before);
|
||||
if (result.kind !== 'input') throw Error('Expected the allowed Add action');
|
||||
result.signatures.forEach(signature => seen.add(signature));
|
||||
expect(autoplanSetupDecision(panel(c), seen, c).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('the exact observed damaged display still waits; action normalization does not repair it', () => {
|
||||
expect(autoplanSetupDecision(fixture.observedScreen, new Set(), call()).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('reordered actions select the native numeric position, with corresponding letters', () => {
|
||||
const c = call(), q = c.questions[0]!;
|
||||
q.options.reverse();
|
||||
q.options = q.options.map((o, i) => ({...o, label:String.fromCharCode(65 + i) + ') ' + o.label.slice(3)}));
|
||||
expect(decision(c)).toMatchObject({kind:'input',input:'2'});
|
||||
const lower = call(); lower.questions[0]!.options.forEach(o => { o.label = o.label[0]!.toLowerCase() + o.label.slice(1); });
|
||||
expect(decision(lower)).toMatchObject({kind:'input',input:'1'});
|
||||
const plain = call(); plain.questions[0]!.options.forEach(o => { o.label = o.label.slice(3); });
|
||||
expect(decision(plain)).toMatchObject({kind:'input',input:'1'});
|
||||
});
|
||||
|
||||
test('one marker cannot hide another marker, noncorresponding ordinal or unrelated action', () => {
|
||||
for (const prefix of ['B) ', 'AA) ', 'A)) ', 'A) B) ', 'A) A) ', 'A.', '1) ', 'Option A) ', 'A)Source excerpt: ', 'A) If approved, ', 'A) Do not ']) {
|
||||
const c = call(); c.questions[0]!.options[0]!.label = prefix + c.questions[0]!.options[0]!.label.slice(3);
|
||||
expect(decision(c).kind, prefix).not.toBe('input');
|
||||
}
|
||||
for (const label of ['A) Add product routes', 'A) Add routing rules to README.md', 'A) Add routing rules to CLAUDE.md and deploy', 'A) Add routing rules to CLAUDE.md (recommended) then delete the plan']) {
|
||||
const c = call(); c.questions[0]!.options[0]!.label = label;
|
||||
expect(decision(c).kind, label).not.toBe('input');
|
||||
}
|
||||
const unsupported = call(); unsupported.questions[0]!.options[1]!.label = 'B) Ask me after this review';
|
||||
expect(decision(unsupported).kind).toBe('unsupported_setup');
|
||||
});
|
||||
|
||||
test('normalization never changes full label, question, status or menu binding', () => {
|
||||
const original = call(), display = panel(original);
|
||||
for (const mutate of [
|
||||
(c:NativePlanQuestionCall) => { c.questions[0]!.options.reverse(); },
|
||||
(c:NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = c.questions[0]!.options[0]!.label.slice(3); },
|
||||
(c:NativePlanQuestionCall) => { c.questions[0]!.question = 'A different routing question?'; },
|
||||
(c:NativePlanQuestionCall) => { c.questions[0]!.header = 'Foreign routing'; },
|
||||
(c:NativePlanQuestionCall) => { c.answered = true; },
|
||||
(c:NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c:NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
]) {
|
||||
const c = call(); mutate(c);
|
||||
expect(autoplanSetupDecision(display,new Set(),c).kind).not.toBe('input');
|
||||
}
|
||||
for (const screen of [display.replace(' 2. B)', ' 2. A)'), display.replace(' 2. B)', ' 2. '),
|
||||
display.replace('Esc to cancel','Esc to'), 'Source example panel:\n' + display,
|
||||
'```text\n' + display + '\n```']) {
|
||||
expect(autoplanSetupDecision(screen,new Set(),original).kind).not.toBe('input');
|
||||
}
|
||||
});
|
||||
|
||||
test('existing owned pending reader rejects foreign, stale and completed requests before action selection', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(),'routing-label-reader-'));
|
||||
try {
|
||||
const cwd=path.join(dir,'repo'),config=path.join(dir,'config'),state=structuredClone(fixture.pendingState);
|
||||
state.cwd=cwd; state.configDir=config;
|
||||
state.pending.transcriptPath=path.join(config,'projects','owned',`${state.sessionId}.jsonl`);
|
||||
fs.mkdirSync(path.dirname(state.pending.transcriptPath),{recursive:true});
|
||||
fs.writeFileSync(state.pending.transcriptPath,'');
|
||||
const file=path.join(dir,'state.json'); fs.writeFileSync(file,JSON.stringify(state));
|
||||
const transcript=structuredClone(fixture.nativeTranscript) as PlanCountTranscript;
|
||||
const read=(c=cwd,cf=config,t=fixture.commandLowerBound,n=transcript) => readPendingQuestion(file,c,cf,t,n);
|
||||
const owned=read(); expect(owned).toBeDefined();
|
||||
expect(autoplanSetupDecision(panel(owned!),new Set(),owned)).toMatchObject({kind:'input',input:'1'});
|
||||
expect(read(cwd+'-foreign')).toBeUndefined();
|
||||
expect(read(cwd,config+'-foreign')).toBeUndefined();
|
||||
expect(read(cwd,config,Date.parse(state.pending.timestamp)+1)).toBeUndefined();
|
||||
const foreign=structuredClone(transcript);foreign.assistantMessages[0]!.sessionId='foreign';
|
||||
expect(read(cwd,config,fixture.commandLowerBound,foreign)).toBeUndefined();
|
||||
const completed=structuredClone(transcript);completed.calls.push({...call(),answered:true});
|
||||
expect(read(cwd,config,fixture.commandLowerBound,completed)).toBeUndefined();
|
||||
} finally { fs.rmSync(dir,{recursive:true,force:true}); }
|
||||
});
|
||||
|
||||
test('the new regression and exact public fixture are mapped without sparse owner entries', () => {
|
||||
const owner=E2E_TOUCHFILES['autoplan-chain-pty'];
|
||||
expect(owner).toContain('test/autoplan-routing-label-ap.test.ts');
|
||||
expect(owner).toContain('test/fixtures/autoplan-routing-label-ap.json');
|
||||
for(let i=0;i<owner.length;i++){expect(Object.hasOwn(owner,i)).toBe(true);expect(typeof owner[i]).toBe('string');}
|
||||
});
|
||||
});
|
||||
@@ -1,94 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
// Exact public unanswered call and final screen from AC's first attempt.
|
||||
// Mutated panels below are synthetic controls, not historical execution.
|
||||
const fixture = JSON.parse(readFileSync(new URL('./fixtures/autoplan-routing-manual-skills-ac.json', import.meta.url), 'utf8'));
|
||||
const native = (): NativePlanQuestionCall => structuredClone(fixture.call);
|
||||
function panel(call: NativePlanQuestionCall): string {
|
||||
const q = call.questions[0]!;
|
||||
return `☐ ${q.header}\n${q.question}\n` + q.options.map((option, index) =>
|
||||
`${index === 0 ? '❯ ' : ' '}${index + 1}. ${option.label}\n ${option.description ?? ''}`).join('\n') +
|
||||
'\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
}
|
||||
|
||||
describe('AC routing manual-skills option', () => {
|
||||
test('helper, regression test and retained fixture each select only the native Autoplan chain', () => {
|
||||
for (const file of [
|
||||
'test/helpers/autoplan-setup-question.ts',
|
||||
'test/autoplan-routing-manual-skills.test.ts',
|
||||
'test/fixtures/autoplan-routing-manual-skills-ac.json',
|
||||
]) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected, file).toEqual(['autoplan-chain-pty']);
|
||||
}
|
||||
});
|
||||
|
||||
test('the exact retained native call and renderer frame preserve the existing Add action once', () => {
|
||||
const seen = new Set<string>();
|
||||
const call = native();
|
||||
expect(call.answered).toBe(false);
|
||||
expect(call.questions[0]!.options[1]!.label).toBe('No thanks, manual skills');
|
||||
const decision = autoplanSetupDecision(fixture.visible, seen, call);
|
||||
expect(decision).toMatchObject({ kind: 'input', input: '1' });
|
||||
expect(seen.size).toBe(0);
|
||||
if (decision.kind !== 'input') throw new Error('Expected recognized routing setup');
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
expect(autoplanSetupDecision(fixture.visible, seen, call).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('synthetic option reversal retains the Add choice without depending on its index', () => {
|
||||
const call = native();
|
||||
call.questions[0]!.options.reverse();
|
||||
expect(autoplanSetupDecision(panel(call), new Set(), call)).toMatchObject({ kind: 'input', input: '2' });
|
||||
});
|
||||
|
||||
test('the same whole manual-skills action accepts existing courtesy and only modifiers', () => {
|
||||
for (const label of ['Manual skills', 'Manual skills only', 'No thanks, manual skills', 'Skip — manual skills only']) {
|
||||
const call = native(); call.questions[0]!.options[1]!.label = label;
|
||||
expect(autoplanSetupDecision(panel(call), new Set(), call), label).toMatchObject({ kind: 'input', input: '1' });
|
||||
}
|
||||
});
|
||||
|
||||
test('other manual workflows and extra actions remain unsupported', () => {
|
||||
for (const label of [
|
||||
'Manual deployment skills', 'Manual billing skills', 'Manual skills after deleting CLAUDE.md',
|
||||
'No thanks, manual skills then skip the review', 'No thanks, manual skills and ship now',
|
||||
'No thanks, manual skills approval', 'Manual skills only after removing CI',
|
||||
]) {
|
||||
const call = native(); call.questions[0]!.options[1]!.label = label;
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanSetupDecision(panel(call), seen, call).kind, label).not.toBe('input');
|
||||
expect(seen.size).toBe(0);
|
||||
}
|
||||
});
|
||||
|
||||
test('unrelated product choices and additional Add actions do not borrow routing setup', () => {
|
||||
for (const question of [
|
||||
'Which product API routing design should we choose? <gstack-qid:routing-injection>',
|
||||
'The plan quotes gstack skill routing rules in CLAUDE.md. Should we expand the feature? <gstack-qid:routing-injection>',
|
||||
]) {
|
||||
const call = native(); call.questions[0]!.question = question;
|
||||
expect(autoplanSetupDecision(panel(call), new Set(), call).kind).not.toBe('input');
|
||||
}
|
||||
const call = native(); call.questions[0]!.options[0]!.label = 'Add routing rules and delete the CI gate';
|
||||
expect(autoplanSetupDecision(panel(call), new Set(), call).kind).not.toBe('input');
|
||||
});
|
||||
|
||||
test('answered, failed, mismatched and mixed native identities remain non-actionable', () => {
|
||||
for (const change of [
|
||||
(call: NativePlanQuestionCall) => { call.answered = true; },
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'Other'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Manual skills only'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions.push(structuredClone(call.questions[0]!)); },
|
||||
]) {
|
||||
const call = native(); change(call);
|
||||
expect(autoplanSetupDecision(fixture.visible, new Set(), call).kind).not.toBe('input');
|
||||
}
|
||||
expect(autoplanSetupDecision(fixture.visible + '\nContinuing the review.', new Set(), native()).kind).not.toBe('input');
|
||||
});
|
||||
});
|
||||
@@ -1,159 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles';
|
||||
|
||||
const frame = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-routing-o-screen.txt'), 'utf8');
|
||||
const question = {
|
||||
header: 'Routing rules',
|
||||
question: "gstack works best when your project's CLAUDE.md includes skill routing rules. Add them now?",
|
||||
options: [{ label: 'Add routing rules (Recommended)' }, { label: 'Skip for now' }],
|
||||
};
|
||||
const native = () => ({ sessionId: 'o-routing', toolUseId: 'routing', answered: false, failed: false, questions: [structuredClone(question)] });
|
||||
|
||||
describe('complete routing panel with a temporary decline', () => {
|
||||
test('the exact O panel chooses Add once before native persistence and after matching persistence', () => {
|
||||
expect(frame).toContain('Invoke skills manually going forward.');
|
||||
for (const pending of [undefined, native()]) {
|
||||
const seen = new Set<string>();
|
||||
const decision = autoplanSetupDecision(frame, seen, pending);
|
||||
expect(decision).toMatchObject({ kind: 'input', input: '1' });
|
||||
expect(seen.size).toBe(0);
|
||||
if (decision.kind !== 'input') throw Error('Expected input');
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
expect(autoplanSetupDecision(frame, seen, pending).kind).toBe('waiting');
|
||||
expect(autoplanSetupDecision(frame, seen, native()).kind).toBe('waiting');
|
||||
}
|
||||
});
|
||||
|
||||
test('the unambiguous opposed decline does not depend on its description or choice order', () => {
|
||||
const withoutDescription = frame.replace(/^\s+Invoke skills manually going forward\..*$/m, '');
|
||||
for (const label of ['Skip for now', 'Skip for now (Recommended)', 'SKIP FOR NOW']) {
|
||||
const current = withoutDescription.replace('2. Skip for now', '2. ' + label);
|
||||
expect(autoplanSetupDecision(current, new Set())).toMatchObject({ kind: 'input', input: '1' });
|
||||
const reversed = current.replace('1. Add routing rules (Recommended)', '1. ' + label)
|
||||
.replace('2. ' + label, '2. Add routing rules (Recommended)');
|
||||
expect(autoplanSetupDecision(reversed, new Set())).toMatchObject({ kind: 'input', input: '2' });
|
||||
}
|
||||
for (const label of ['Skip', 'No thanks', 'Skip — invoke skills manually', 'Manual only']) {
|
||||
expect(autoplanSetupDecision(frame.replace('Skip for now', label), new Set())).toMatchObject({ kind: 'input', input: '1' });
|
||||
}
|
||||
});
|
||||
|
||||
test('extra actions, unrelated questions and ambiguous offered choices do not acquire input', () => {
|
||||
for (const label of ['Skip for now and delete CLAUDE.md', 'Skip for now, implement the feature', 'Skip the review for now', 'Skip for now unless the API changes', 'Ask me after this review']) {
|
||||
expect(autoplanSetupDecision(frame.replace('2. Skip for now', '2. ' + label), new Set()).kind, label).not.toBe('input');
|
||||
}
|
||||
for (const changed of [
|
||||
frame.replace(question.question, 'Which product API routing design should we choose?'),
|
||||
frame.replace(question.question, 'The plan quotes gstack skill routing rules in CLAUDE.md. Should we build an API router?'),
|
||||
frame.replace('1. Add routing rules (Recommended)', '1. Implement routing (Recommended)'),
|
||||
frame.replace('2. Skip for now', '2. Add routing rules'),
|
||||
frame.replace('3. Type something.', '3. Skip for now\n 4. Type something.').replace('4. Chat about this', '5. Chat about this'),
|
||||
frame.replace('3. Type something.', '3. Implement the feature\n 4. Type something.').replace('4. Chat about this', '5. Chat about this'),
|
||||
]) expect(autoplanSetupDecision(changed, new Set()).kind, changed).not.toBe('input');
|
||||
});
|
||||
|
||||
test('only the complete current native panel can supply this additional label', () => {
|
||||
const panel = frame.slice(frame.indexOf(' ☐ Routing rules'));
|
||||
for (const changed of [
|
||||
'Example panel:\n' + panel, 'Quoted source:\n' + panel, '```text\n' + panel, '~~~~text\n' + panel,
|
||||
panel.split('\n').map(line => ' ' + line).join('\n'), panel.split('\n').map(line => '> ' + line).join('\n'),
|
||||
panel + '\n● Continuing the review.', panel.replace('Esc to cancel', 'Esc to'),
|
||||
panel.replace(' 4. Chat about this', ''), panel.replace(' 3. Type something.', ''),
|
||||
panel.replace('❯ 1.', ' 1.'), panel.replace(' 2.', '❯ 2.'),
|
||||
panel.replace('1. Add', '1. [ ] Add'), panel.replace(' ☐ Routing rules', '← ☐ Routing rules ✔ Submit →'),
|
||||
]) expect(autoplanSetupDecision(changed, new Set()).kind, changed).not.toBe('input');
|
||||
expect(autoplanSetupDecision('```text\nold code\n```\n' + panel, new Set())).toMatchObject({kind:'input',input:'1'});
|
||||
});
|
||||
|
||||
test('present metadata cannot be replaced by the visible decline label', () => {
|
||||
for (const mutate of [
|
||||
(call:any) => {call.failed=true;}, (call:any) => {call.answered=true;},
|
||||
(call:any) => {call.questions=[];}, (call:any) => {call.questions.push(structuredClone(question));},
|
||||
(call:any) => {call.questions[0].multiSelect=true;}, (call:any) => {call.questions[0].header='Other';},
|
||||
(call:any) => {call.questions[0].question='Different question';},
|
||||
(call:any) => {call.questions[0].options[1].label='Different choice';},
|
||||
]) {const call=native();mutate(call);expect(autoplanSetupDecision(frame,new Set(),call).kind).not.toBe('input');}
|
||||
});
|
||||
});
|
||||
|
||||
test('routing regression inputs remain paid-selection dependencies', () => {
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/autoplan-routing-o.test.ts');
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-routing-o-screen.txt');
|
||||
});
|
||||
|
||||
test.skipIf(process.platform === 'win32')('real PTY temporary routing decline advances after readiness with exactly one Add digit', async () => {
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-routing-o-'));
|
||||
const fake=path.join(dir,'fake-claude');const worker=path.join(dir,'worker.ts');const output=path.join(dir,'result.json');
|
||||
const cases=[false,true].map(early=>({name:early?'early':'deferred',early,frame,question,
|
||||
cwd:path.join(dir,early?'early':'deferred'),events:path.join(dir,early?'early.jsonl':'deferred.jsonl')}));
|
||||
for(const item of cases)fs.mkdirSync(item.cwd);
|
||||
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
|
||||
import * as fs from 'node:fs';import * as path from 'node:path';
|
||||
const item=JSON.parse(process.env.ROUTING_CASE);const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
|
||||
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
|
||||
const file=path.join(folder,item.name+'.jsonl');
|
||||
const persist=value=>fs.appendFileSync(file,JSON.stringify({sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),...value})+'\n');
|
||||
const use=()=>persist({type:'assistant',message:{role:'assistant',content:[{type:'tool_use',id:'routing',name:'AskUserQuestion',input:{questions:[item.question]}}]}});
|
||||
event({kind:'startup',pid:process.pid});if(item.early)use();
|
||||
process.stdin.setRawMode?.(true);process.stdin.resume();
|
||||
process.stdin.on('data',data=>{event({kind:'input',data:data.toString()});if(!item.early)use();
|
||||
persist({type:'user',toolUseResult:{answers:{[item.question.question]:'Add routing rules (Recommended)'}},message:{role:'user',content:[{type:'tool_result',tool_use_id:'routing',content:'User has answered your questions: "'+item.question.question+'"="Add routing rules (Recommended)". You can now continue with the user\'s answers in mind.'}]}});
|
||||
process.stdout.write('\r\nROUTING_ACCEPTED\r\n');});
|
||||
process.stdout.write('\x1b[2J\x1b[H'+item.frame.replace(/\n/g,'\r\n'));
|
||||
process.on('SIGINT',()=>process.exit(0));
|
||||
`);fs.chmodSync(fake,0o755);
|
||||
const url=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
|
||||
fs.writeFileSync(worker,`
|
||||
import * as fs from 'node:fs';
|
||||
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(url('claude-pty-runner.ts'))};
|
||||
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
|
||||
import {autoplanSetupDecision} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
|
||||
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
|
||||
const results=[];
|
||||
for(const item of ${JSON.stringify(cases)}){
|
||||
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:15000,env:{ROUTING_CASE:JSON.stringify(item)}});
|
||||
try{
|
||||
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
|
||||
const screen=await session.currentScreen();const before=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
|
||||
const pending=before.calls.find(call=>!call.answered&&!call.failed);
|
||||
if(Boolean(pending)!==item.early)throw Error('Incorrect readiness metadata');
|
||||
const seen=new Set();const decision=autoplanSetupDecision(screen,seen,pending);
|
||||
if(decision.kind!=='input'||decision.input!=='1')throw Error('Expected Add input: '+JSON.stringify(decision));
|
||||
session.send(decision.input);for(const signature of decision.signatures)seen.add(signature);
|
||||
await session.waitFor('ROUTING_ACCEPTED',{timeoutMs:3000,pollMs:20});
|
||||
const after=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
|
||||
results.push({name:item.name,decision,after,redraw:autoplanSetupDecision(screen,seen,pending).kind,
|
||||
answered:autoplanSetupDecision(screen,new Set(),after.calls[0]).kind});
|
||||
}finally{await session.close();}
|
||||
}
|
||||
fs.writeFileSync(${JSON.stringify(output)},JSON.stringify(results));
|
||||
`);
|
||||
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake},stdout:'pipe',stderr:'pipe'});
|
||||
const killer=setTimeout(()=>child.kill('SIGKILL'),25000);
|
||||
try{
|
||||
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
|
||||
expect(exit,stdout+stderr).toBe(0);
|
||||
const results=JSON.parse(fs.readFileSync(output,'utf8'));
|
||||
expect(results.length).toBe(2);
|
||||
for(const [index,result]of results.entries()){
|
||||
expect(result.decision).toMatchObject({kind:'input',input:'1'});expect(result.redraw).toBe('waiting');expect(result.answered).toBe('waiting');
|
||||
expect(result.after.calls.length).toBe(1);expect(result.after.calls[0].answered).toBe(true);
|
||||
expect(result.after.calls[0].answers[question.question]).toBe('Add routing rules (Recommended)');
|
||||
const events=fs.readFileSync(cases[index]!.events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
|
||||
expect(events.filter(event=>event.kind==='input')).toEqual([{kind:'input',data:'1'}]);
|
||||
expect(()=>process.kill(events[0].pid,0)).toThrow();
|
||||
}
|
||||
}finally{
|
||||
clearTimeout(killer);child.kill('SIGKILL');
|
||||
for(const item of cases)if(fs.existsSync(item.events)){
|
||||
const pid=JSON.parse(fs.readFileSync(item.events,'utf8').split('\n')[0]!).pid;
|
||||
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
|
||||
}
|
||||
fs.rmSync(dir,{recursive:true,force:true});
|
||||
}
|
||||
},30000);
|
||||
@@ -1,365 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
import { autoplanSetupDecision } from './helpers/autoplan-setup-question';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
const captured = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-packet-o-screen.txt'), 'utf8');
|
||||
const original = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-packet-o-call.json'), 'utf8')) as NativePlanQuestionCall;
|
||||
const zPacket = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-z-packet.json'), 'utf8')) as {pendingCall: NativePlanQuestionCall; screen: string};
|
||||
const footer = 'Enter to select · Tab/Arrow keys to navigate · Esc to cancel';
|
||||
function pane(call: NativePlanQuestionCall, index: number, answered: number[] = []) {
|
||||
const bar = '← ' + call.questions.map((q,i) => (answered.includes(i) ? '☒ ' : '☐ ') + q.header).join(' ') + ' ✔ Submit →';
|
||||
if (index === call.questions.length) return `${bar}\nReview your answers\nReady to submit your answers?\n❯ 1. Submit answers\n 2. Cancel\n${footer}\n`;
|
||||
const q = call.questions[index]!;
|
||||
return `${bar}\n│ ${q.question}\n` + q.options.map((option,i) => `${i===0?'❯':' '} ${i+1}. ${option.label}`).join('\n') +
|
||||
`\n 3. Type something.\n 4. Chat about this\n${footer}\n`;
|
||||
}
|
||||
function commit(screen: string, seen: Set<string>, call: NativePlanQuestionCall, expected: string) {
|
||||
const before = [...seen]; const action = autoplanSetupDecision(screen, seen, call);
|
||||
expect([...seen]).toEqual(before); expect(action).toMatchObject({kind:'input',input:expected});
|
||||
if (action.kind !== 'input') throw Error('Expected input');
|
||||
for (const key of action.signatures) seen.add(key);
|
||||
expect(autoplanSetupDecision(screen,seen,call).kind).toBe('waiting');
|
||||
return action;
|
||||
}
|
||||
|
||||
describe('native routing and prerequisite setup packet', () => {
|
||||
test('exact O active pane then prerequisite each receive one bound choice, followed by one Submit', () => {
|
||||
const seen=new Set<string>();
|
||||
commit(captured,seen,original,'1');
|
||||
expect(autoplanSetupDecision(pane(original,2,[0,1]),seen,original).kind).toBe('waiting');
|
||||
commit(pane(original,1,[0]),seen,original,'1');
|
||||
commit(pane(original,2,[0,1]),seen,original,'\r');
|
||||
expect(autoplanSetupDecision(captured,new Set(),{...original,answered:true}).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('question and option order may change without changing the existing choices', () => {
|
||||
for (const reverseQuestions of [false,true]) for (const reverseOptions of [false,true]) {
|
||||
const call=structuredClone(original);if(reverseQuestions)call.questions.reverse();
|
||||
if(reverseOptions)for(const question of call.questions)question.options.reverse();
|
||||
const seen=new Set<string>();
|
||||
for(let index=0;index<2;index++)commit(pane(call,index,index?[0]:[]),seen,call,reverseOptions?'2':'1');
|
||||
commit(pane(call,2,[0,1]),seen,call,'\r');
|
||||
}
|
||||
});
|
||||
|
||||
test('metadata may persist late, but no tab is answered before the complete packet is known', () => {
|
||||
const seen=new Set<string>();
|
||||
expect(autoplanSetupDecision(captured,seen).kind).toBe('waiting');expect(seen.size).toBe(0);
|
||||
expect(autoplanSetupDecision(pane(original,0).split('\n').slice(1).join('\n'),seen).kind).toBe('waiting');
|
||||
expect(autoplanSetupDecision(captured,seen,{...original,questions:[original.questions[0]!]}).kind).toBe('waiting');
|
||||
commit(captured,seen,original,'1');
|
||||
expect(autoplanSetupDecision(pane(original,2,[0,1]),new Set(),original).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('every native question must be one unambiguous setup offer', () => {
|
||||
for (const mutate of [
|
||||
(call:any)=>{call.failed=true;},(call:any)=>{call.answered=true;},(call:any)=>{call.questions[1].multiSelect=true;},
|
||||
(call:any)=>{call.questions.push(structuredClone(call.questions[0]));},
|
||||
(call:any)=>{call.questions[1]=structuredClone(call.questions[0]);},
|
||||
(call:any)=>{call.questions[1].question='Which user experience should the API provide?';},
|
||||
(call:any)=>{call.questions[1].question='No design doc exists for /office-hours integration. Should we build X or defer Y?';},
|
||||
(call:any)=>{call.questions[0].question+=' Should we delete the archived invoices?';},
|
||||
(call:any)=>{call.questions[0].question+=' Also approve deleting the archived invoices before continuing.';},
|
||||
(call:any)=>{call.questions[1].question='Should we delete the archived invoices? '+call.questions[1].question;},
|
||||
(call:any)=>{call.questions[1].question=call.questions[1].question.replace('— sharper input','and also approve deleting the archived invoices — sharper input');},
|
||||
(call:any)=>{call.questions[1].question='No design doc found for this branch. /office-hours produces a design doc — also archive the invoices. Run it first or proceed with standard review?';},
|
||||
(call:any)=>{call.questions[1].options[1].label='Run /office-hours first then implement';},
|
||||
(call:any)=>{call.questions[1].options[0].label='Skip — implement the feature';},
|
||||
(call:any)=>{call.questions[0].question='The plan quotes gstack skill routing rules in CLAUDE.md. Should we build an API router?';},
|
||||
(call:any)=>{call.questions[0].options[1].label='No thanks, delete CLAUDE.md';},
|
||||
(call:any)=>{call.questions[0].options.push({label:'Implement the feature'});},
|
||||
]) {const call=structuredClone(original);mutate(call);const seen=new Set<string>();
|
||||
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
|
||||
});
|
||||
|
||||
test('current tab, full offered labels and active panel context must all agree', () => {
|
||||
const first=pane(original,0);
|
||||
for (const changed of [
|
||||
'Example panel:\n'+first, 'Example:\n'+first, 'Quoted source:\n'+first, '```text\n'+first, '~~~~text\n'+first,
|
||||
first.split('\n').map(line=>' '+line).join('\n'), first.split('\n').map(line=>'> '+line).join('\n'),
|
||||
first+'\n● Continuing the review.', first+first, first.replace('Esc to cancel','Esc to'),
|
||||
first.replace('Prerequisite doc','Other tab'), first.replace(original.questions[0]!.question,'Unrelated question'),
|
||||
first.replace('← ', '← Different call '),
|
||||
first.replace('1. Add','1. Delete'), first.replace('1. Add','1. [ ] Add'),
|
||||
first.replace(' 3. Type something.',''), first.replace(' 4. Chat about this',''),
|
||||
first.replace('❯ 1.',' 1.'),first.replace(' 2.','❯ 2.'),
|
||||
first.replace(' 3. Type something.',' 3. Implement the feature\n 4. Type something.').replace(' 4. Chat about this',' 5. Chat about this'),
|
||||
]) expect(autoplanSetupDecision(changed,new Set(),original).kind,changed).toBe('waiting');
|
||||
expect(autoplanSetupDecision('```text\nearlier code\n```\n'+first,new Set(),original)).toMatchObject({kind:'input',input:'1'});
|
||||
expect(autoplanSetupDecision(pane(original,0,[0]),new Set(),original).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('Submit requires each actual sent identity, checked tabs, unchanged packet and a current Submit panel', () => {
|
||||
const seen=new Set<string>();commit(pane(original,0),seen,original,'1');commit(pane(original,1,[0]),seen,original,'1');
|
||||
const submit=pane(original,2,[0,1]);
|
||||
for (const changed of [
|
||||
pane(original,2,[0]), 'Example panel:\n'+submit,'Example:\n'+submit,'```text\n'+submit,submit+'\n● Finished.',
|
||||
submit.replace('Submit answers','Accept implementation'),submit.replace('Ready to submit your answers?','Implement the feature?'),
|
||||
submit.replace(' 2. Cancel',' 2. Cancel\n 3. Deploy'),submit.replace('Esc to cancel','Esc to'),
|
||||
]) expect(autoplanSetupDecision(changed,seen,original).kind,changed).toBe('waiting');
|
||||
for (const change of ['session','tool','question','description']) {
|
||||
const call=structuredClone(original);
|
||||
if(change==='session')call.sessionId+='-other';if(change==='tool')call.toolUseId+='-other';
|
||||
if(change==='question')call.questions[0]!.question+=' ';
|
||||
if(change==='description')call.questions[0]!.options[0]!.description='Changed';
|
||||
expect(autoplanSetupDecision(submit,seen,call).kind).toBe('waiting');
|
||||
}
|
||||
commit(submit,seen,original,'\r');
|
||||
});
|
||||
|
||||
test('packet captures and regression remain paid-selection dependencies', () => {
|
||||
for(const file of ['test/autoplan-setup-packet-o.test.ts','test/fixtures/autoplan-setup-packet-o-screen.txt','test/fixtures/autoplan-setup-packet-o-call.json'])
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain(file);
|
||||
});
|
||||
});
|
||||
|
||||
describe('numbered native setup packet preserves the full existing review', () => {
|
||||
test('exact Z packet advances both bound tabs and only then submits once', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanSetupDecision(zPacket.screen, seen).kind).toBe('waiting');
|
||||
commit(zPacket.screen, seen, zPacket.pendingCall, '1');
|
||||
expect(autoplanSetupDecision(pane(zPacket.pendingCall, 2, [0,1]), seen, zPacket.pendingCall).kind).toBe('waiting');
|
||||
commit(pane(zPacket.pendingCall, 1, [0]), seen, zPacket.pendingCall, '1');
|
||||
commit(pane(zPacket.pendingCall, 2, [0,1]), seen, zPacket.pendingCall, '\r');
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-setup-z-packet.json');
|
||||
});
|
||||
|
||||
test('numbering and offered order vary while picks keep their exact native identities', () => {
|
||||
for (const reverseQuestions of [false, true]) for (const reverseOptions of [false, true]) {
|
||||
const call = structuredClone(zPacket.pendingCall);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('D1 —', 'D17:');
|
||||
call.questions[1]!.question = call.questions[1]!.question.replace('D2 —', 'D23 –').replace('this branch', 'the project').replace('the review input', 'this review input');
|
||||
if (reverseQuestions) call.questions.reverse();
|
||||
if (reverseOptions) call.questions.forEach(question => question.options.reverse());
|
||||
const seen = new Set<string>();
|
||||
commit(pane(call,0), seen, call, reverseOptions ? '2' : '1');
|
||||
commit(pane(call,1,[0]), seen, call, reverseOptions ? '2' : '1');
|
||||
commit(pane(call,2,[0,1]), seen, call, '\r');
|
||||
}
|
||||
});
|
||||
|
||||
test('all new question and option description clauses must remain setup only', () => {
|
||||
const mutations: Array<(call: NativePlanQuestionCall) => void> = [
|
||||
call => { call.questions[0]!.question += ' Also remove account-owner authorization.'; },
|
||||
call => { call.questions[1]!.question += ' Approve dropping the audit tests?'; },
|
||||
call => { call.questions[0]!.question = 'The plan quotes ' + call.questions[0]!.question; },
|
||||
call => { call.questions[1]!.question = call.questions[1]!.question.replace('sharpen the review input', 'approve the proposed changes'); },
|
||||
call => { call.questions[1]!.options[0]!.description = call.questions[1]!.options[0]!.description!.replace('CEO → Design → DX → Eng', 'CEO → Eng'); },
|
||||
call => { call.questions[1]!.options[0]!.description = call.questions[1]!.options[0]!.description!.replace('plan as-is', 'plan after removing authorization'); },
|
||||
call => { call.questions[1]!.options[0]!.label += ' and implement'; },
|
||||
call => { call.questions[1]!.options[1]!.label += ' then ship'; },
|
||||
];
|
||||
for (let question = 0; question < 2; question++) for (let option = 0; option < 2; option++) {
|
||||
mutations.push(call => { call.questions[question]!.options[option]!.description += ' Also delete the account-owner check.'; });
|
||||
mutations.push(call => { call.questions[question]!.options[option]!.description = undefined; });
|
||||
}
|
||||
for (const mutate of mutations) {
|
||||
const call = structuredClone(zPacket.pendingCall); mutate(call);
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanSetupDecision(pane(call,0), seen, call).kind).toBe('waiting');
|
||||
expect(seen.size).toBe(0);
|
||||
}
|
||||
});
|
||||
|
||||
test('new forms require complete pending native identity and the same intact active pane', () => {
|
||||
const mutations: Array<(call: any) => void> = [
|
||||
call => { delete call.answered; }, call => { delete call.failed; }, call => { call.answered = true; }, call => { call.failed = true; },
|
||||
call => { delete call.sessionId; }, call => { delete call.toolUseId; },
|
||||
call => { call.questions[0].question = call.questions[0].question.replace('routing-injection', 'other-question'); },
|
||||
call => { call.questions[0].question += ' <gstack-qid:routing-injection>'; },
|
||||
call => { call.questions[1].question = call.questions[1].question.replace('D2', 'D0'); },
|
||||
call => { call.questions[1].multiSelect = true; },
|
||||
call => { call.questions.push(structuredClone(call.questions[0])); },
|
||||
call => { call.questions[1] = structuredClone(call.questions[0]); },
|
||||
];
|
||||
for (const mutate of mutations) { const call = structuredClone(zPacket.pendingCall); mutate(call);
|
||||
expect(autoplanSetupDecision(pane(call,0),new Set(),call).kind).toBe('waiting'); }
|
||||
const first = pane(zPacket.pendingCall,0);
|
||||
for (const screen of ['Example panel:\n'+first, '```text\n'+first, first+'\nProceeding.', first.replace('Esc to cancel','Esc to'),
|
||||
first.replace('Design doc','Other tab'), first.replace('1. Add','1. Delete'), first.replace('← ','← Unrelated packet '),
|
||||
first.split('\n').map(line => '> '+line).join('\n')]) {
|
||||
expect(autoplanSetupDecision(screen,new Set(),zPacket.pendingCall).kind).toBe('waiting');
|
||||
}
|
||||
expect(autoplanSetupDecision(pane(zPacket.pendingCall,2,[0,1]),new Set(),zPacket.pendingCall).kind).toBe('waiting');
|
||||
});
|
||||
});
|
||||
|
||||
test.skipIf(process.platform==='win32')('real PTY native setup packet waits for metadata, answers each visible tab once and submits without a stray digit',async()=>{
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-setup-packet-o-'));const fake=path.join(dir,'fake-claude');
|
||||
const worker=path.join(dir,'worker.ts');const resultFile=path.join(dir,'results.json');
|
||||
const cases=[{name:'o',call:original,first:captured},{name:'z',call:zPacket.pendingCall,first:zPacket.screen}].flatMap(packet =>
|
||||
[false,true].map(late=>({name:packet.name+(late?'-late':'-early'),late,cwd:path.join(dir,packet.name+(late?'-late':'-early')),
|
||||
events:path.join(dir,packet.name+(late?'-late.jsonl':'-early.jsonl')),release:path.join(dir,packet.name+(late?'-late.release':'-early.release')),call:packet.call,first:packet.first})));
|
||||
for(const item of cases)fs.mkdirSync(item.cwd);
|
||||
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
|
||||
import * as fs from 'node:fs';import * as path from 'node:path';
|
||||
const item=JSON.parse(process.env.PACKET_CASE);const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
|
||||
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
|
||||
const file=path.join(folder,item.call.sessionId+'.jsonl');let index=0,answers={},published=false,done=false;
|
||||
const persist=value=>fs.appendFileSync(file,JSON.stringify({sessionId:item.call.sessionId,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),...value})+'\n');
|
||||
const publish=()=>{if(published)return;published=true;persist({type:'assistant',message:{role:'assistant',content:[{type:'tool_use',id:item.call.toolUseId,name:'AskUserQuestion',input:{questions:item.call.questions}}]}});event({kind:'metadata'});process.stdout.write('\r\nMETADATA_READY\r\n');render();};
|
||||
function render(){const q=item.call.questions;let screen=item.first;
|
||||
if(index>0){const bar='← '+q.map(question=>(answers[question.question]?'☒ ':'☐ ')+question.header).join(' ')+' ✔ Submit →';
|
||||
screen=index<q.length?bar+'\n│ '+q[index].question+'\n'+q[index].options.map((o,i)=>(i===0?'❯':' ')+' '+(i+1)+'. '+o.label).join('\n')+'\n 3. Type something.\n 4. Chat about this':bar+'\nReview your answers\nReady to submit your answers?\n❯ 1. Submit answers\n 2. Cancel';
|
||||
screen+='\nEnter to select · Tab/Arrow keys to navigate · Esc to cancel\n';}
|
||||
process.stdout.write('\x1b[2J\x1b[H'+screen.replace(/\n/g,'\r\n'));}
|
||||
event({kind:'startup',pid:process.pid});process.stdin.setRawMode?.(true);process.stdin.resume();
|
||||
process.stdin.on('data',data=>{const input=data.toString();event({kind:'input',input,index,published});if(!published||done)throw Error('Unexpected input lifecycle');
|
||||
if(index<2){if(!/^[12]$/.test(input))throw Error('One native digit required');answers[item.call.questions[index].question]=item.call.questions[index].options[Number(input)-1].label;index++;render();}
|
||||
else{if(input!=='\r')throw Error('Raw Submit required');done=true;persist({type:'user',toolUseResult:{answers},message:{role:'user',content:[{type:'tool_result',tool_use_id:item.call.toolUseId,content:'Answered.'}]}});event({kind:'submitted',answers});process.stdout.write('\x1b[2J\x1b[HNATIVE_PACKET_COMPLETE\r\n');}});
|
||||
render();if(!item.late)publish();const timer=setInterval(()=>{if(item.late&&fs.existsSync(item.release))publish();},10);
|
||||
process.on('SIGINT',()=>{clearInterval(timer);process.exit(0);});
|
||||
`);fs.chmodSync(fake,0o755);
|
||||
const url=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
|
||||
fs.writeFileSync(worker,`
|
||||
import * as fs from 'node:fs';
|
||||
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(url('claude-pty-runner.ts'))};
|
||||
import {autoplanSetupDecision} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
|
||||
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
|
||||
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
|
||||
const results=[];
|
||||
for(const item of ${JSON.stringify(cases)}){
|
||||
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:15000,env:{PACKET_CASE:JSON.stringify(item)}});
|
||||
try{
|
||||
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});const seen=new Set();
|
||||
if(item.late){const pending=readPlanCountTranscript(session.hermeticConfigDir,item.cwd).calls[0];if(pending)throw Error('Expected missing native packet');
|
||||
if(autoplanSetupDecision(await session.currentScreen(),seen,pending).kind!=='waiting'||seen.size)throw Error('Guessed before native identity');
|
||||
fs.writeFileSync(item.release,'release');}
|
||||
await session.waitFor('METADATA_READY',{timeoutMs:3000,pollMs:20});
|
||||
const inputs=[];
|
||||
for(let step=0;step<3;step++){
|
||||
const current=await session.currentScreen();const call=readPlanCountTranscript(session.hermeticConfigDir,item.cwd).calls[0];
|
||||
const action=autoplanSetupDecision(current,seen,call);
|
||||
if(action.kind!=='input')throw Error('Expected input at '+step+': '+JSON.stringify({action,current,call}));
|
||||
session.send(action.input);inputs.push(action.input);for(const signature of action.signatures)seen.add(signature);
|
||||
if(autoplanSetupDecision(current,seen,call).kind!=='waiting')throw Error('Repeated input on unchanged pane');
|
||||
await session.waitFor(step===0?'☒ '+item.call.questions[0].header:step===1?'Ready to submit your answers?':'NATIVE_PACKET_COMPLETE',{timeoutMs:3000,pollMs:20});
|
||||
}
|
||||
const transcript=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);results.push({name:item.name,inputs,transcript});
|
||||
}finally{await session.close();}}
|
||||
fs.writeFileSync(${JSON.stringify(resultFile)},JSON.stringify(results));
|
||||
`);
|
||||
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake},stdout:'pipe',stderr:'pipe'});
|
||||
const killer=setTimeout(()=>child.kill('SIGKILL'),26000);
|
||||
try{
|
||||
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);expect(exit,stdout+stderr).toBe(0);
|
||||
const results=JSON.parse(fs.readFileSync(resultFile,'utf8'));expect(results.length).toBe(4);
|
||||
for(const [index,result]of results.entries()){
|
||||
expect(result.inputs).toEqual(['1','1','\r']);expect(result.transcript.calls.length).toBe(1);expect(result.transcript.calls[0].answered).toBe(true);
|
||||
expect(result.transcript.calls[0].answers).toEqual(Object.fromEntries(cases[index]!.call.questions.map(q=>[q.question,q.options[0]!.label])));
|
||||
const events=fs.readFileSync(cases[index]!.events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
|
||||
expect(events.filter(e=>e.kind==='input').map(e=>({input:e.input,index:e.index,published:e.published}))).toEqual([
|
||||
{input:'1',index:0,published:true},{input:'1',index:1,published:true},{input:'\r',index:2,published:true}]);
|
||||
expect(events.filter(e=>e.kind==='submitted').length).toBe(1);expect(()=>process.kill(events[0].pid,0)).toThrow();
|
||||
}
|
||||
}finally{
|
||||
clearTimeout(killer);child.kill('SIGKILL');for(const item of cases)if(fs.existsSync(item.events)){
|
||||
const pid=JSON.parse(fs.readFileSync(item.events,'utf8').split('\n')[0]!).pid;
|
||||
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
|
||||
}fs.rmSync(dir,{recursive:true,force:true});
|
||||
}
|
||||
},30000);
|
||||
|
||||
const adV2Packet = JSON.parse(fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-setup-ad-v2-packet.json'), 'utf8')) as {pendingCall: NativePlanQuestionCall; screen: string};
|
||||
test('AD v2 actual setup packet chooses routing and standard review with the existing native identity',()=>{
|
||||
const seen=new Set<string>(),call=adV2Packet.pendingCall;
|
||||
expect(autoplanSetupDecision(adV2Packet.screen,seen).kind).toBe('waiting');
|
||||
commit(adV2Packet.screen,seen,call,'1');
|
||||
// Only the first pane was retained live. Later panes are explicit native-question projections.
|
||||
commit(pane(call,1,[0]),seen,call,'2');
|
||||
commit(pane(call,2,[0,1]),seen,call,'\r');
|
||||
});
|
||||
|
||||
test('AD v2 setup policy uses the task and actions across presentation and option order',()=>{
|
||||
for(const variant of ['numbered','unprefixed','different explanation'])for(const reverseQuestions of [false,true])for(const reverseOptions of [false,true]){
|
||||
const call=structuredClone(adV2Packet.pendingCall);
|
||||
call.questions.forEach((q,index)=>{
|
||||
q.question=q.question.replace(/^D\d+\s*[—–:-]\s*/,variant==='unprefixed'?'':`D${31+index}: `);
|
||||
if(variant==='different explanation')q.question=q.question.split('\n')[0]+'\nProject/branch/task: disposable review fixture, another branch and release.\nELI10: This setup changes how later sessions find workflow context.\nStakes if we pick wrong: an extra setup step.\nRecommendation: Keep the offered actions explicit.\nNet: setup now versus a direct review.';
|
||||
});
|
||||
if(reverseQuestions)call.questions.reverse();if(reverseOptions)call.questions.forEach(q=>q.options.reverse());
|
||||
const seen=new Set<string>();
|
||||
for(let i=0;i<2;i++){
|
||||
const ordinary=call.questions[i]!.header==='Routing'?1:2;
|
||||
commit(pane(call,i,i===1?[0]:[]),seen,call,String(reverseOptions?3-ordinary:ordinary));
|
||||
}
|
||||
commit(pane(call,2,[0,1]),seen,call,'\r');
|
||||
}
|
||||
});
|
||||
|
||||
test('AD v2 setup cannot borrow a header, subject or adjacent question for a different decision',()=>{
|
||||
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
|
||||
c=>{c.questions[0]!.header='Product router';},
|
||||
c=>{c.questions[1]!.header='Deployment';},
|
||||
c=>{[c.questions[0]!.header,c.questions[1]!.header]=[c.questions[1]!.header,c.questions[0]!.header];},
|
||||
c=>{c.questions[0]!.question=c.questions[0]!.question.replace(/^.*\n/,'D1 — Should the application route requests through a proxy?\n');},
|
||||
c=>{c.questions[1]!.question=c.questions[1]!.question.replace(/^.*\n/,'D2 — Should we add an office-hours page to the product?\n');},
|
||||
c=>{c.questions[0]!.question='The plan quotes: '+c.questions[0]!.question;},
|
||||
c=>{c.questions[1]!.question='```text\n'+c.questions[1]!.question+'\n```';},
|
||||
c=>{c.questions[1]!.question=c.questions[1]!.question.split('\n').map(l=>'> '+l).join('\n');},
|
||||
c=>{c.questions[1]!.question+=' Should we remove the authorization check?';},
|
||||
c=>{c.questions[1]={...structuredClone(c.questions[1]!),question:'Approve deployment to production?',header:'Approval'};},
|
||||
];
|
||||
for(const change of changes){const call=structuredClone(adV2Packet.pendingCall);change(call);const seen=new Set<string>();
|
||||
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
|
||||
});
|
||||
|
||||
test('AD v2 setup rejects conditional, contradictory and ambiguous actions in either tab',()=>{
|
||||
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
|
||||
c=>{c.questions[0]!.options[0]!.description='Do not add routing rules to CLAUDE.md.';},
|
||||
c=>{c.questions[0]!.options[1]!.description='Add routing rules to CLAUDE.md after declining.';},
|
||||
c=>{c.questions[1]!.options[0]!.description='Skip the design doc and begin the review now.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Run /office-hours first, then proceed with standard review.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review after completing /office-hours.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='No review will run.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review?';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review but do not run it.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Review starts now, but not yet.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review when /office-hours completes.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Review starts immediately after completing /office-hours.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review once the design doc is complete.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review if the tests pass.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Skip the CEO review and proceed directly to engineering.';},
|
||||
c=>{c.questions[1]!.options[1]!.description='Proceed with standard review only if the tests pass.';},
|
||||
c=>{c.questions[1]!.question+=' You must run /office-hours before the review.';},
|
||||
c=>{c.questions[1]!.question+=' Standard review is forbidden until /office-hours completes.';},
|
||||
c=>{c.questions[0]!.options[0]!.label+=' and implement the feature';},
|
||||
c=>{c.questions[1]!.options[1]!.label+=' if the tests pass';},
|
||||
c=>{c.questions[0]!.options[0]!.description+=' Also delete the authorization check.';},
|
||||
c=>{c.questions[1]!.options[1]!.description+=' Also deploy to production.';},
|
||||
c=>{c.questions[0]!.options[1]=structuredClone(c.questions[0]!.options[0]!);},
|
||||
c=>{c.questions[1]!.options.push({label:'Skip the remaining review phases'});},
|
||||
];
|
||||
for(const change of changes){const call=structuredClone(adV2Packet.pendingCall);change(call);const seen=new Set<string>();
|
||||
expect(autoplanSetupDecision(pane(call,0),seen,call).kind).toBe('waiting');expect(seen.size).toBe(0);}
|
||||
});
|
||||
|
||||
test('AD v2 setup retains complete native identity and current-pane requirements',()=>{
|
||||
const first=adV2Packet.screen,call=adV2Packet.pendingCall;
|
||||
// The example label must introduce the panel, not precede unrelated earlier transcript rows.
|
||||
for(const screen of ['Example panel:\n'+pane(call,0),'```text\n'+first,first+'\nContinuing.',
|
||||
first.replace('Design doc','Different tab'),first.replace('Esc to cancel','Esc to'),
|
||||
first.replace('Add routing rules to CLAUDE.md (recommended)','Add routing rules to OTHER.md (recommended)')]){
|
||||
expect(screen).not.toBe(first);expect(autoplanSetupDecision(screen,new Set(),call).kind).toBe('waiting');
|
||||
}
|
||||
for(const delta of [{answered:true},{failed:true},{sessionId:''},{toolUseId:''}])
|
||||
expect(autoplanSetupDecision(first,new Set(),{...call,...delta}).kind).toBe('waiting');
|
||||
});
|
||||
|
||||
test('AD v2 setup fixture selects the existing Autoplan paid case only',()=>{
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-setup-ad-v2-packet.json');
|
||||
const owners=Object.entries(E2E_TOUCHFILES).filter(([,files])=>files.includes('test/fixtures/autoplan-setup-ad-v2-packet.json')).map(([name])=>name);
|
||||
expect(owners).toEqual(['autoplan-chain-pty']);
|
||||
});
|
||||
|
||||
test('AD v2 selected review action allows short affirmative descriptions with dynamic tradeoffs',()=>{
|
||||
for(const description of ['Proceed with standard review. The plan already states its goals.', 'Review begins now using the existing plan. No separate design artifact is created.', 'Start the standard review immediately with the supplied context.']){
|
||||
const call=structuredClone(adV2Packet.pendingCall);call.questions[1]!.options[1]!.description=description;
|
||||
expect(autoplanSetupDecision(pane(call,0),new Set(),call).kind).toBe('input');
|
||||
}
|
||||
});
|
||||
@@ -1,916 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { autoplanRoutingSetupInput, autoplanSetupDecision } from './helpers/autoplan-setup-question';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
|
||||
const CLIPPED_ROUTING_N = fs.readFileSync(path.join(import.meta.dir, 'fixtures/autoplan-routing-n-screen.txt'), 'utf8');
|
||||
|
||||
describe('current routing title survives a scrolled native header before metadata flushes', () => {
|
||||
test('exact N frame selects the offered Add action once with a native digit only', () => {
|
||||
expect(CLIPPED_ROUTING_N).not.toMatch(/[☐□]/);
|
||||
const seen = new Set<string>();
|
||||
const decision = autoplanSetupDecision(CLIPPED_ROUTING_N, seen);
|
||||
expect(decision.kind).toBe('input');
|
||||
if (decision.kind !== 'input') throw Error('Expected native setup input');
|
||||
expect(decision.input).toBe('1');
|
||||
expect(seen.size).toBe(0);
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
expect(autoplanSetupDecision(CLIPPED_ROUTING_N, seen).kind).toBe('waiting');
|
||||
expect(E2E_TOUCHFILES['autoplan-chain-pty']).toContain('test/fixtures/autoplan-routing-n-screen.txt');
|
||||
});
|
||||
|
||||
test('equivalent direct title and reordered opposed choices retain picker binding', () => {
|
||||
const frame = CLIPPED_ROUTING_N.replace('D1 — Add skill', 'D9 — Add gstack skill');
|
||||
const swapped = frame.replace('1. Add routing rules (Recommended)', '1. Skip, invoke manually')
|
||||
.replace('2. Skip, invoke manually', '2. Add routing rules (Recommended)');
|
||||
expect(autoplanSetupDecision(frame, new Set())).toMatchObject({kind:'input',input:'1'});
|
||||
expect(autoplanSetupDecision(swapped, new Set())).toMatchObject({kind:'input',input:'2'});
|
||||
});
|
||||
|
||||
test('copied, stale, incomplete, ambiguous and substantive panels cannot borrow the top routing identity', () => {
|
||||
for (const frame of [
|
||||
'Example panel:\n' + CLIPPED_ROUTING_N,
|
||||
'Quoted source:\n' + CLIPPED_ROUTING_N,
|
||||
'```text\n' + CLIPPED_ROUTING_N + '\n```',
|
||||
'~~~~text\n' + CLIPPED_ROUTING_N,
|
||||
CLIPPED_ROUTING_N.split('\n').map(line => ' ' + line).join('\n'),
|
||||
CLIPPED_ROUTING_N.split('\n').map(line => '> ' + line).join('\n'),
|
||||
CLIPPED_ROUTING_N + '\n⏺ Continuing the review.',
|
||||
CLIPPED_ROUTING_N.replace('Esc to cancel', 'Esc to'),
|
||||
CLIPPED_ROUTING_N.replace('❯ 1.', ' 1.'),
|
||||
CLIPPED_ROUTING_N.replace('❯ 1.', ' 1.').replace(' 2.', '❯ 2.'),
|
||||
CLIPPED_ROUTING_N.replace(' 2.', '❯ 2.'),
|
||||
CLIPPED_ROUTING_N.replace('1. Add', '1. [ ] Add'),
|
||||
CLIPPED_ROUTING_N.replace('│\n│ Project', '│ ← ☐ Routing ✔ Submit →\n│ Project'),
|
||||
CLIPPED_ROUTING_N.replace(' 4. Chat about this', ''),
|
||||
CLIPPED_ROUTING_N.replace('2. Skip, invoke manually', '2. Add routing rules (Recommended)'),
|
||||
CLIPPED_ROUTING_N.replace('2. Skip, invoke manually', '2. Delete routing and migrate the product'),
|
||||
CLIPPED_ROUTING_N.replace('routing-injection>', 'product-routing>'),
|
||||
CLIPPED_ROUTING_N.replace('routing-injection>', 'routing-injection'),
|
||||
CLIPPED_ROUTING_N.replace('│ Project/branch:', '│ <gstack-qid:routing-injection>\n│ Project/branch:'),
|
||||
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'Choose the product API router for CLAUDE.md?'),
|
||||
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'The spec quotes Add skill routing rules to CLAUDE.md?'),
|
||||
CLIPPED_ROUTING_N.replace('Add skill routing rules to CLAUDE.md?', 'Add skill routing rules to README.md?'),
|
||||
CLIPPED_ROUTING_N.replace(' <gstack-qid:routing-injection>', '').replace('│ Net:', '│ <gstack-qid:routing-injection> Net:'),
|
||||
CLIPPED_ROUTING_N.replace('│ ELI10:', '│ ```text\n│ ELI10:'),
|
||||
CLIPPED_ROUTING_N.replace('│ ELI10:', '│ > Quoted source:\n│ ELI10:'),
|
||||
]) expect(autoplanSetupDecision(frame, new Set()).kind, frame).not.toBe('input');
|
||||
});
|
||||
|
||||
test('present native metadata keeps its full existing identity binding', () => {
|
||||
const before = CLIPPED_ROUTING_N.split('❯ 1.')[0]!.replace(/^[│┃] ?/gm, '').trim();
|
||||
const call: any = {toolUseId:'n-routing',sessionId:'n',timestamp:'2026-09-09T01:10:05Z',answered:false,failed:false,
|
||||
questions:[{header:'Routing',question:before,options:[{label:'Add routing rules (Recommended)'},{label:'Skip, invoke manually'}]}]};
|
||||
expect(autoplanSetupDecision(CLIPPED_ROUTING_N,new Set(),call)).toMatchObject({kind:'input',input:'1'});
|
||||
for (const mutate of [
|
||||
(q:any) => {q.failed=true;}, (q:any) => {q.answered=true;}, (q:any) => {q.questions=[];},
|
||||
(q:any) => {q.questions.push(structuredClone(q.questions[0]));},
|
||||
(q:any) => {q.questions[0].multiSelect=true;},
|
||||
(q:any) => {q.questions[0].question='Unrelated finding <gstack-qid:routing-injection>';},
|
||||
(q:any) => {q.questions[0].options[1].label='Another choice';},
|
||||
]) {const changed=structuredClone(call);mutate(changed);expect(autoplanSetupDecision(CLIPPED_ROUTING_N,new Set(),changed).kind).not.toBe('input');}
|
||||
const seen=new Set<string>();
|
||||
const early=autoplanSetupDecision(CLIPPED_ROUTING_N,seen);
|
||||
if(early.kind!=='input')throw Error('Expected initial input');
|
||||
for(const signature of early.signatures)seen.add(signature);
|
||||
expect(autoplanSetupDecision(CLIPPED_ROUTING_N,seen,call).kind).toBe('waiting');
|
||||
});
|
||||
});
|
||||
|
||||
test.skipIf(process.platform === 'win32')('real PTY clipped routing advances from the exact current panel with one digit and no Enter', async () => {
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-clipped-routing-'));
|
||||
const fake=path.join(dir,'fake-claude');const events=path.join(dir,'events.jsonl');
|
||||
fs.writeFileSync(fake,`#!${process.execPath}\n`+String.raw`
|
||||
import * as fs from 'node:fs';
|
||||
const emit=value=>fs.appendFileSync(process.env.ROUTING_EVENTS,JSON.stringify(value)+'\n');
|
||||
emit({kind:'started',pid:process.pid});
|
||||
process.stdin.setRawMode?.(true);process.stdin.resume();
|
||||
process.stdin.on('data',data=>{emit({kind:'input',data:data.toString()});process.stdout.write('\r\nNATIVE_SETUP_ACCEPTED\r\n');});
|
||||
process.stdout.write(fs.readFileSync(process.env.ROUTING_SCREEN,'utf8').replace(/\n/g,'\r\n'));
|
||||
process.on('SIGINT',()=>process.exit(0));
|
||||
`);fs.chmodSync(fake,0o755);
|
||||
const worker=path.join(dir,'worker.ts');const resultFile=path.join(dir,'result.json');
|
||||
const helper=(name:string)=>pathToFileURL(path.resolve(import.meta.dir,'helpers',name)).href;
|
||||
fs.writeFileSync(worker,`
|
||||
import * as fs from 'node:fs';
|
||||
import {launchClaudePty,resolveClaudeBinary} from ${JSON.stringify(helper('claude-pty-runner.ts'))};
|
||||
import {autoplanSetupDecision} from ${JSON.stringify(helper('autoplan-setup-question.ts'))};
|
||||
if(resolveClaudeBinary()!==${JSON.stringify(fake)})throw Error('Fake binary binding failed before launch');
|
||||
const session=await launchClaudePty({cwd:${JSON.stringify(dir)},observeScreen:true,timeoutMs:15000,
|
||||
env:{ROUTING_EVENTS:process.env.ROUTING_EVENTS,ROUTING_SCREEN:process.env.ROUTING_SCREEN}});
|
||||
try{
|
||||
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
|
||||
const screen=await session.currentScreen();
|
||||
const decision=autoplanSetupDecision(screen,new Set());
|
||||
if(decision.kind!=='input'||decision.input!=='1')throw Error('Expected current setup: '+JSON.stringify(decision));
|
||||
session.send(decision.input);
|
||||
await session.waitFor('NATIVE_SETUP_ACCEPTED',{timeoutMs:3000,pollMs:20});
|
||||
fs.writeFileSync(${JSON.stringify(resultFile)},JSON.stringify({screen,decision}));
|
||||
}finally{await session.close();}
|
||||
`);
|
||||
const child=Bun.spawn([process.execPath,worker],{env:{...process.env,BROWSE_TERMINAL_BINARY:fake,
|
||||
ROUTING_EVENTS:events,ROUTING_SCREEN:path.join(import.meta.dir,'fixtures/autoplan-routing-n-screen.txt')},stdout:'pipe',stderr:'pipe'});
|
||||
const killer=setTimeout(()=>child.kill('SIGKILL'),17000);
|
||||
try {
|
||||
const [exit,stdout,stderr]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
|
||||
expect(exit,stdout+stderr).toBe(0);
|
||||
const result=JSON.parse(fs.readFileSync(resultFile,'utf8'));
|
||||
expect(result.screen).not.toMatch(/[☐□]/);
|
||||
expect(result.decision).toMatchObject({kind:'input',input:'1'});
|
||||
const recorded=fs.readFileSync(events,'utf8').trim().split('\n').map(line=>JSON.parse(line));
|
||||
expect(recorded.filter(e=>e.kind==='input')).toEqual([{kind:'input',data:'1'}]);
|
||||
expect(()=>process.kill(recorded[0].pid,0)).toThrow();
|
||||
} finally {
|
||||
clearTimeout(killer);child.kill('SIGKILL');
|
||||
if(fs.existsSync(events)){
|
||||
const pid=JSON.parse(fs.readFileSync(events,'utf8').split('\n')[0]!).pid;
|
||||
if(process.platform==='linux')try{if(fs.readFileSync('/proc/'+pid+'/cmdline','utf8').split('\0').includes(fake))process.kill(pid,'SIGKILL');}catch{}
|
||||
}
|
||||
fs.rmSync(dir,{recursive:true,force:true});
|
||||
}
|
||||
},20000);
|
||||
|
||||
// Sanitized terminal frame from the 2026-09-08 autoplan timeout. The qid is
|
||||
// visibly incomplete; the prompt body and explicit choices remain intact.
|
||||
const CAPTURE = [
|
||||
'─'.repeat(120),
|
||||
'Planning: /tmp/hermetic/.claude/plans/modular-bouncing-swing.md',
|
||||
'─'.repeat(120),
|
||||
' ☐ Routing rules',
|
||||
"│ gstack works best when your project's CLAUDE.md includes skill routing rules. Add them now?",
|
||||
'│<gstck-qid:routing-injectin>',
|
||||
'❯1.Addroutingrules(Recommended)',
|
||||
'CreatesCLAUDE.mdwithskillroutingrulessogstackknowswhentoinvoke/office-hours,/autoplan,/ship,/qa,',
|
||||
"etc.automatically.We'lldothisafterthereview.",
|
||||
'2.Nothanks',
|
||||
"Skip—I'llinvokeskillsmanually.Youcanenablethislaterbyrunninggstack-configsetrouting_declinedfalse.",
|
||||
'3.Typesomething.',
|
||||
'4.Chataboutthis',
|
||||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||||
].join('\r\r');
|
||||
|
||||
// Targeted-a stalled on this complete menu for the full test budget. Parsing
|
||||
// retained its identity and choices; the setup helper rejected their wording.
|
||||
const CURRENT_CAPTURE = [
|
||||
' ☐ Routing rules',
|
||||
'',
|
||||
'Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>',
|
||||
'',
|
||||
'❯1.AddtoCLAUDE.md(recommended)',
|
||||
'',
|
||||
'Appendsa##SkillroutingsectiontoCLAUDE.mdandcommitsit.Futuresessionswillauto-invoketherightskill',
|
||||
'(/investigateforbugs,/shipforPRs,/qafortesting,etc.)withoutmanualinvocation.',
|
||||
'',
|
||||
'2.Skip—invokemanually',
|
||||
'',
|
||||
"Nofilechanges.You'llcontinuecallingskillsbyname.Canaddroutingruleslater.",
|
||||
'',
|
||||
'3.Typesomething.',
|
||||
'─'.repeat(120),
|
||||
'4.Chataboutthis',
|
||||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||||
].join('\n');
|
||||
|
||||
// Targeted-b's first attempt stayed on this complete setup menu until its
|
||||
// 15-minute deadline. The parser retained the prompt and both labels, but
|
||||
// the setup selector rejected "No thanks, invoke manually".
|
||||
const B_CAPTURE = [
|
||||
' ☐ CLAUDE.md',
|
||||
'',
|
||||
'│ D1 — Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>',
|
||||
'│',
|
||||
'│ELI10:ThisprojecthasnoCLAUDE.md.Thatfileiswheregstacklooksforroutingrules—instructionstellingClaude',
|
||||
'│Codewhichskilltoauto-invokeforwhichrequest(e.g."ship→/ship","bugs→/investigate").Withoutityoutype',
|
||||
'│theskillnameeverytime.Withit,gstackcanrecognizeyourintentandrouteautomatically.',
|
||||
'│',
|
||||
'│Stakesifweskip:Noauto-routing;youinvokeskillsmanuallyeachsession.',
|
||||
'│',
|
||||
'│Recommendation:A—one-timesetup,saveskeystrokesoneveryfuturesession.',
|
||||
'│Completeness:A=9/10,B=5/10',
|
||||
'',
|
||||
'❯1.AddroutingrulestoCLAUDE.md(Recommended)',
|
||||
'AppendsthestandardgstackroutingblocktoanewCLAUDE.mdandcommitsit.Doneonce,activeforever.',
|
||||
'2.Nothanks,invokemanually',
|
||||
'SkipCLAUDE.mdsetup.Youcontinuecalling/autoplan,/ship,/qa,etc.bynameeachtime.',
|
||||
'3.Typesomething.',
|
||||
'─'.repeat(120),
|
||||
'4.Chataboutthis',
|
||||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||||
].join('\r\r');
|
||||
|
||||
// Fresh broad retry: the complete setup menu uses a noun for the manual
|
||||
// alternative. This is the same opposed setup action as "invoke manually".
|
||||
const FRESH_RETRY_CAPTURE = [
|
||||
'☐Routingsetup',
|
||||
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
|
||||
'❯1.Addroutingrules(Recommended)',
|
||||
'AppendskillroutingrulestoCLAUDE.mdsoClaudeautomaticallyinvokestherightskillforproduct,engineering,',
|
||||
'design,andshipworkflows.Willbedoneafterplanapproval(planmodeisactivenow).',
|
||||
'2.Nothanks,manualinvocation',
|
||||
"Skip—I'llinvokeskillsmanually.Thispromptwon'tappearagain.",
|
||||
'3.Typesomething.',
|
||||
'─'.repeat(120),
|
||||
'4.Chataboutthis',
|
||||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||||
].join('\n');
|
||||
|
||||
describe('autoplan routing setup handling', () => {
|
||||
// Source-F retry's first complete frame preceded damaged terminal redraws.
|
||||
const F_SETUP_CAPTURE = [
|
||||
'Planning: /tmp/hermetic/.claude/plans/deep-coalescing-valiant.md',
|
||||
'☐Skillrouting',
|
||||
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
|
||||
'❯1.AddroutingrulestoCLAUDE.md',
|
||||
'AppendsskillroutingrulestoCLAUDE.mdsogstackauto-invokestherightskillforcommonrequests(review,ship,',
|
||||
'investigate,etc.).Willbecommittedtotherepo.(recommended)',
|
||||
'2.Nothanks,skip',
|
||||
"I'llinvokeskillsmanually.Youcanaddroutinglater.",
|
||||
'3.Typesomething.',
|
||||
'4.Chataboutthis',
|
||||
'Entertoselect·↑/↓tonavigate·Esctocancel',
|
||||
].join('\r\r');
|
||||
|
||||
test('answers the captured combined decline action once, regardless of option order', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE, seen)).toBeNull();
|
||||
const reordered = F_SETUP_CAPTURE.replace('❯1.AddroutingrulestoCLAUDE.md', '❯1.Nothanks,skip')
|
||||
.replace('2.Nothanks,skip', '2.AddroutingrulestoCLAUDE.md');
|
||||
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
|
||||
expect(autoplanRoutingSetupInput(F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip—invoke skills manually'), new Set())).toBe('1');
|
||||
});
|
||||
|
||||
test('does not infer a routing answer from damaged, ambiguous, or unrelated setup choices', () => {
|
||||
for (const frame of [
|
||||
F_SETUP_CAPTURE.replace('Addroutingrules', 'Addrutingrules'),
|
||||
F_SETUP_CAPTURE.replace('Nothanks,skip', 'Nothank,skip'),
|
||||
F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip the review'),
|
||||
F_SETUP_CAPTURE.replace('Nothanks,skip', 'No thanks, skip then delete CLAUDE.md'),
|
||||
F_SETUP_CAPTURE.replace('3.Typesomething.', '3.Skip'),
|
||||
F_SETUP_CAPTURE.replace("gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?", 'Which routing design should the application use?'),
|
||||
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
|
||||
});
|
||||
|
||||
test('answers the fresh retry manual-invocation setup once in either option order', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE, seen)).toBeNull();
|
||||
const reordered = FRESH_RETRY_CAPTURE.replace('❯1.Addroutingrules(Recommended)', '❯1.Nothanks,manualinvocation')
|
||||
.replace('2.Nothanks,manualinvocation', '2.Addroutingrules(Recommended)');
|
||||
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
|
||||
});
|
||||
|
||||
test('requires opposed manual setup actions and rejects ambiguous or unrelated choices', () => {
|
||||
for (const decline of [
|
||||
'No thanks, delete the file manually',
|
||||
'No thanks, manual data migration',
|
||||
'No thanks, invoke the deploy manually',
|
||||
'Manual deployment invocation',
|
||||
'Accept recommendation',
|
||||
'No thanks, manual invocation then delete CLAUDE.md',
|
||||
]) {
|
||||
const frame = FRESH_RETRY_CAPTURE.replace('Nothanks,manualinvocation', decline);
|
||||
expect(autoplanRoutingSetupInput(frame, new Set()), decline).toBeNull();
|
||||
}
|
||||
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE.replace('3.Typesomething.', '3.Add routing rules'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(FRESH_RETRY_CAPTURE.replace('3.Typesomething.', '3.Skip—invoke manually'), new Set())).toBeNull();
|
||||
const review = FRESH_RETRY_CAPTURE.replace(
|
||||
"gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Addthemnow?",
|
||||
'Which product routing design should we ship? <gstack-qid:routing-injection>',
|
||||
);
|
||||
expect(autoplanRoutingSetupInput(review, new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('answers the captured setup once, using the full question identity', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(CAPTURE, seen)).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(CAPTURE.replace('works best', 'works best'), seen)).toBeNull();
|
||||
});
|
||||
|
||||
test('chooses Add routing rules by label when option order changes', () => {
|
||||
const reordered = CAPTURE.replace('❯1.Addroutingrules(Recommended)', '❯1.Nothanks')
|
||||
.replace('2.Nothanks', '2.Add routing rules (Recommended)');
|
||||
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
|
||||
});
|
||||
|
||||
test('accepts the full option labels captured from the subsequent live setup prompt', () => {
|
||||
const fullLabels = CAPTURE.replace('Addroutingrules(Recommended)', 'Add routing rules to CLAUDE.md (Recommended)')
|
||||
.replace('2.Nothanks', "2.No thanks, I'll invoke skills manually");
|
||||
expect(autoplanRoutingSetupInput(fullLabels, new Set())).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(fullLabels.replace('CLAUDE.md (Recommended)', 'product routes (Recommended)'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(fullLabels.replace("I'll invoke skills manually", 'delete the existing rules'), new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('answers the current captured CLAUDE.md setup, including reordered choices, once', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE, seen)).toBeNull();
|
||||
const reordered = CURRENT_CAPTURE.replace('❯1.AddtoCLAUDE.md(recommended)', '❯1.Skip—invokemanually')
|
||||
.replace('2.Skip—invokemanually', '2.AddtoCLAUDE.md(recommended)');
|
||||
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
|
||||
expect(autoplanRoutingSetupInput(CURRENT_CAPTURE.replace('to CLAUDE.md?', "to this project's CLAUDE.md?"), new Set())).toBe('1');
|
||||
});
|
||||
|
||||
test('the current wording still requires both explicit setup choices and the CLAUDE.md target', () => {
|
||||
for (const frame of [
|
||||
CURRENT_CAPTURE.replace('to CLAUDE.md?', 'to the application API?'),
|
||||
CURRENT_CAPTURE.replace('AddtoCLAUDE.md(recommended)', 'Acceptrecommendation'),
|
||||
CURRENT_CAPTURE.replace('Skip—invokemanually', 'Deferthisfinding'),
|
||||
CURRENT_CAPTURE.replace('AddtoCLAUDE.md(recommended)', 'Deletetheexistingroutingrules'),
|
||||
CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Should we expand the current feature?'),
|
||||
]) expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('recognizes the native A retry packet with its abbreviated manual-decline label', () => {
|
||||
const retry = CAPTURE.replace('Addroutingrules(Recommended)', 'Add to CLAUDE.md (Recommended)')
|
||||
.replace('2.Nothanks', '2.No thanks, manual');
|
||||
expect(autoplanRoutingSetupInput(retry, new Set())).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(retry.replace('No thanks, manual', 'No thanks, delete it'), new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('answers the exact B timeout menu by its routing label, in either order', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(B_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(B_CAPTURE, seen)).toBeNull();
|
||||
const reordered = B_CAPTURE.replace('❯1.AddroutingrulestoCLAUDE.md(Recommended)', '❯1.Nothanks,invokemanually')
|
||||
.replace('2.Nothanks,invokemanually', '2.AddroutingrulestoCLAUDE.md(Recommended)');
|
||||
expect(autoplanRoutingSetupInput(reordered, new Set())).toBe('2');
|
||||
expect(autoplanRoutingSetupInput(B_CAPTURE.replace('Nothanks,invokemanually', 'Nothanks,deletethefilemanually'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(B_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Which routing design should the application use?'), new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('recognizes the setup premise without depending on its closing sentence', () => {
|
||||
const openings = [
|
||||
"gstack works best when your project's CLAUDE.md includes skill routing rules. Would you like to add them?",
|
||||
"gstack works best when your project's CLAUDE.md includes skill routing rules. Enable them for this repository?",
|
||||
'Should we configure skill routing rules for gstack in CLAUDE.md?',
|
||||
'Set up gstack skill routing rules in CLAUDE.md.',
|
||||
];
|
||||
for (const opening of openings) {
|
||||
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>', opening);
|
||||
expect(autoplanRoutingSetupInput(frame, new Set()), opening).toBe('1');
|
||||
}
|
||||
});
|
||||
|
||||
test('recognizes an intact setup qid with an explicit CLAUDE.md action and opposed manual decline', () => {
|
||||
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md?', 'Configure this project’s CLAUDE.md?');
|
||||
expect(autoplanRoutingSetupInput(frame, new Set())).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(frame.replace('gstack-qid:routing-injection', 'gstack-qid:product-routing'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(frame.replace('AddtoCLAUDE.md(recommended)', 'Acceptrecommendation'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(frame.replace('Skip—invokemanually', 'Deferthisfinding'), new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('keeps generic review, quoted premises and different routing targets out of setup handling', () => {
|
||||
for (const question of [
|
||||
'Which dashboard layout should we ship?',
|
||||
'Add routing rules to the application API? <gstack-qid:product-routing>',
|
||||
'The plan quotes gstack CLAUDE.md skill routing rules. Which API design should we use?',
|
||||
'The document references gstack skill routing rules in CLAUDE.md. Should we expand the feature?',
|
||||
]) {
|
||||
const frame = CURRENT_CAPTURE.replace('Add gstack skill routing rules to CLAUDE.md? <gstack-qid:routing-injection>', question);
|
||||
expect(autoplanRoutingSetupInput(frame, new Set()), question).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('waits for complete recognized choices rather than guessing a default', () => {
|
||||
expect(autoplanRoutingSetupInput(CAPTURE.replace('2.Nothanks', '2.Ask me later'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(CAPTURE.replace('Addroutingrules(Recommended)', 'Accept recommendation'), new Set())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput('❯1.Addroutingrules(Recommended)\r2.Nothanks', new Set())).toBeNull();
|
||||
});
|
||||
|
||||
test('never answers review or taste questions, even with a routing qid or the same choices', () => {
|
||||
const prompts = [
|
||||
'Which visual direction should this settings page use?',
|
||||
'Should the payment handler bypass the existing dispatcher?',
|
||||
'Add routing rules to the product API now? <gstack-qid:routing-injection>',
|
||||
'The plan quotes CLAUDE.md skill routing rules. Should we change this feature?',
|
||||
];
|
||||
for (const prompt of prompts) {
|
||||
const frame = `☐ Review decision\r${prompt}\r❯1.Addroutingrules(Recommended)\r2.Nothanks`;
|
||||
expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('setup helper and captured-frame changes select the autoplan eval only', () => {
|
||||
for (const file of ['test/helpers/autoplan-setup-question.ts', 'test/autoplan-setup-question.test.ts', 'test/fixtures/autoplan-routing-n-screen.txt']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['autoplan-chain-pty']);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// Source-G's retry remained at this actual captured menu until shard timeout.
|
||||
// The action is intact; cumulative ANSI stripping loses the courtesy's 'o'.
|
||||
// A real xterm replay retains it in the prior screen cell.
|
||||
const G_ROUTING_CAPTURE = [
|
||||
'☐Routingrules',
|
||||
"│gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Wouldyouliketoaddthem?",
|
||||
'❯1.AddroutingrulestoCLAUDE.md',
|
||||
'AppendsstandardskillroutingrulestoCLAUDE.md(creatingitifabsent)andcommits.Meansgstackskillslike',
|
||||
'/autoplan,/ship,/qaetc.getinvokedautomaticallywhenthetaskmatches.(recommended)',
|
||||
"2. N thanks, I'll invokeskillsmanually",
|
||||
'Skiprouting setup. You can re-enable later by removing the routing_declined flag.',
|
||||
'3.Typesomething.',
|
||||
'4.Chataboutthis',
|
||||
'Enter toselect · ↑/↓ to navigate · Esc to cancel',
|
||||
].join('\r');
|
||||
|
||||
describe('autoplan routing action survives courtesy repaint', () => {
|
||||
test('selects the explicit Add action once in the captured G menu, in both orders', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE, seen)).toBeNull();
|
||||
const reversed = G_ROUTING_CAPTURE.replace('❯1.AddroutingrulestoCLAUDE.md', "❯1.N thanks, I'll invokeskillsmanually")
|
||||
.replace("2. N thanks, I'll invokeskillsmanually", '2.AddroutingrulestoCLAUDE.md');
|
||||
expect(autoplanRoutingSetupInput(reversed, new Set())).toBe('2');
|
||||
});
|
||||
|
||||
test('the actual manual-invocation action needs no courtesy formula', () => {
|
||||
for (const action of ['Manual invocation', 'Invoke skills manually', "I'll invoke skills manually", 'Thanks, invoke manually']) {
|
||||
expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE.replace("N thanks, I'll invokeskillsmanually", action), new Set()), action).toBe('1');
|
||||
}
|
||||
});
|
||||
|
||||
test('still requires exact opposed setup actions and a genuine routing premise', () => {
|
||||
for (const label of [
|
||||
'N thanks', 'Invoke the deployment manually', 'N thanks, manual data migration',
|
||||
'Delete CLAUDE.md, invoke skills manually', 'No thanks, invoke skills manually then delete CLAUDE.md',
|
||||
'Skip the review, invoke skills manually', 'Skip the review thanks, invoke skills manually',
|
||||
]) expect(autoplanRoutingSetupInput(G_ROUTING_CAPTURE.replace("N thanks, I'll invokeskillsmanually", label), new Set()), label).toBeNull();
|
||||
for (const frame of [
|
||||
G_ROUTING_CAPTURE.replace("gstackworksbestwhenyourproject'sCLAUDE.mdincludesskillroutingrules.Wouldyouliketoaddthem?", 'Which application router should we implement?'),
|
||||
G_ROUTING_CAPTURE.replace('AddroutingrulestoCLAUDE.md', 'AddrutingrulestoCLAUDE.md'),
|
||||
G_ROUTING_CAPTURE.replace('3.Typesomething.', '3.Invoke skills manually'),
|
||||
G_ROUTING_CAPTURE.replace('3.Typesomething.', '3.Add routing rules'),
|
||||
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
const PREREQUISITE_CAPTURE = " ☐ Design doc\n\n│ No design doc found for this branch. /office-hours produces a structured problem statement, premise challenge, and\n│ explored alternatives — it gives this review much sharper input to work with. Takes about 10 minutes. The design doc\n│ is per-feature, not per-product — it captures the thinking behind this specific change. Run /office-hours first?\n\n❯ 1. Run /office-hours now\n Runs /office-hours to produce a design doc first, then picks up the full autoplan review right after. (~10 min)\n 2. Skip — proceed with standard review\n Skips /office-hours and runs the autoplan review pipeline now using the existing plan file as input.\n 3. Type something.\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n 4. Chat about this\n\nEnter to select · ↑/↓ to navigate · Esc to cancel\n";
|
||||
const prerequisiteQuestion = {
|
||||
header: 'Design doc',
|
||||
question: "No design doc found for this branch. /office-hours produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input to work with. Takes about 10 minutes. The design doc is per-feature, not per-product — it captures the thinking behind this specific change. Run /office-hours first?",
|
||||
options: [{ label: 'Run /office-hours now' }, { label: 'Skip — proceed with standard review' }],
|
||||
};
|
||||
const prerequisiteCall = () => ({
|
||||
sessionId: 'prerequisite-session', toolUseId: 'prerequisite-call',
|
||||
answered: false, failed: false, questions: [structuredClone(prerequisiteQuestion)],
|
||||
});
|
||||
function prerequisiteMenu(reverse = false) {
|
||||
if (!reverse) return PREREQUISITE_CAPTURE;
|
||||
return PREREQUISITE_CAPTURE
|
||||
.replace('1. Run /office-hours now', '1. Skip — proceed with standard review')
|
||||
.replace('2. Skip — proceed with standard review', '2. Run /office-hours now');
|
||||
}
|
||||
|
||||
describe('autoplan optional design-doc prerequisite', () => {
|
||||
test('the exact K native screen declines the optional prerequisite by label', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const frame = prerequisiteMenu(reverse);
|
||||
expect(autoplanRoutingSetupInput(frame, new Set())).toBe(reverse ? '1' : '2');
|
||||
const native = prerequisiteCall(); if (reverse) native.questions[0]!.options.reverse();
|
||||
expect(autoplanRoutingSetupInput(frame, new Set(), native)).toBe(reverse ? '1' : '2');
|
||||
}
|
||||
});
|
||||
|
||||
test('quoted panels and menus followed by new output are not active input', () => {
|
||||
for (const frame of [
|
||||
'Example panel:\n```text\n' + PREREQUISITE_CAPTURE + '\n```\n',
|
||||
'Example panel:\n~~~text\n' + PREREQUISITE_CAPTURE,
|
||||
'Example panel:\n' + PREREQUISITE_CAPTURE,
|
||||
PREREQUISITE_CAPTURE.split('\n').map(line => ' ' + line).join('\n'),
|
||||
'The document quotes this panel:\n────────────────────\n' + PREREQUISITE_CAPTURE,
|
||||
PREREQUISITE_CAPTURE + '\n⏺ Continuing the review without office hours.\n',
|
||||
PREREQUISITE_CAPTURE + '\n❯ 1. A new menu\n 2. Another choice\n',
|
||||
]) for (const native of [undefined, prerequisiteCall()]) {
|
||||
expect(autoplanRoutingSetupInput(frame, new Set(), native)).toBeNull();
|
||||
}
|
||||
expect(autoplanRoutingSetupInput('```text\nearlier real code\n```\n────────────────────\n' + PREREQUISITE_CAPTURE, new Set())).toBe('2');
|
||||
});
|
||||
|
||||
test('late native identity does not re-answer the retained menu', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen)).toBe('2');
|
||||
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, prerequisiteCall())).toBeNull();
|
||||
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen)).toBeNull();
|
||||
});
|
||||
|
||||
test('unrelated, failed, mixed and checkbox native calls do not borrow the setup menu', () => {
|
||||
for (const mutate of [
|
||||
(call: ReturnType<typeof prerequisiteCall>) => { call.questions[0]!.question = 'Should we change the dashboard design?'; },
|
||||
(call: ReturnType<typeof prerequisiteCall>) => { call.failed = true; },
|
||||
(call: ReturnType<typeof prerequisiteCall>) => { call.answered = true; },
|
||||
(call: ReturnType<typeof prerequisiteCall>) => { call.questions.push({ header:'Finding', question:'Fix missing auth?', options:[{label:'Fix it'},{label:'Defer'}] }); },
|
||||
(call: ReturnType<typeof prerequisiteCall>) => { Object.assign(call.questions[0]!, {multiSelect:true}); },
|
||||
]) {
|
||||
const native = prerequisiteCall(); mutate(native);
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, native)).toBeNull();
|
||||
// Waiting for correct metadata must not mark an unanswered UI as sent.
|
||||
expect(autoplanRoutingSetupInput(PREREQUISITE_CAPTURE, seen, prerequisiteCall())).toBe('2');
|
||||
}
|
||||
});
|
||||
|
||||
test('arbitrary skip, outside offers, mixed actions and prose examples remain unanswered', () => {
|
||||
for (const frame of [
|
||||
PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Skip this security check'),
|
||||
PREREQUISITE_CAPTURE.replaceAll('/office-hours', '/codex'),
|
||||
PREREQUISITE_CAPTURE.replace('3. Type something.', '3. Fix the missing authorization check'),
|
||||
PREREQUISITE_CAPTURE.replace('No design doc found for this branch.', 'A dashboard design issue was found.'),
|
||||
PREREQUISITE_CAPTURE.replace(' ☐ Design doc', 'Example choices:').replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
|
||||
]) expect(autoplanRoutingSetupInput(frame, new Set())).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
test.skipIf(process.platform === 'win32')('real PTY prerequisite answer survives early and deferred native records without a second key', async () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-autoplan-prereq-'));
|
||||
const fake = path.join(dir, 'fake-claude');
|
||||
const worker = path.join(dir, 'worker.ts');
|
||||
const resultFile = path.join(dir, 'result.json');
|
||||
const cases = [false, true].flatMap(early => [false, true].map(reverse => {
|
||||
const name = `${early ? 'early' : 'deferred'}-${reverse ? 'reversed' : 'original'}`;
|
||||
const q = structuredClone(prerequisiteQuestion); if (reverse) q.options.reverse();
|
||||
return { name, early, cwd: path.join(dir, name), record: path.join(dir, name + '.jsonl'),
|
||||
question: q, frame: prerequisiteMenu(reverse), expected: reverse ? '1' : '2' };
|
||||
}));
|
||||
for (const item of cases) fs.mkdirSync(item.cwd);
|
||||
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
const item = JSON.parse(process.env.PREREQUISITE_REPLAY);
|
||||
const record = event => fs.appendFileSync(item.record, JSON.stringify(event) + '\n');
|
||||
record({type:'startup',pid:process.pid});
|
||||
const folder = path.join(process.env.CLAUDE_CONFIG_DIR, 'projects', 'fixture');
|
||||
fs.mkdirSync(folder, {recursive:true});
|
||||
const transcript = path.join(folder, item.name + '.jsonl');
|
||||
let logged = false;
|
||||
function writeCall() {
|
||||
if (logged) return; logged = true;
|
||||
fs.appendFileSync(transcript, JSON.stringify({type:'assistant',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),
|
||||
message:{role:'assistant',content:[{type:'tool_use',id:'prerequisite',name:'AskUserQuestion',input:{questions:[item.question]}}]}})+'\n');
|
||||
}
|
||||
if (item.early) writeCall();
|
||||
process.stdin.setRawMode?.(true);
|
||||
let answered = false;
|
||||
process.stdin.on('data', data => {
|
||||
record({type:'input',data:data.toString()});
|
||||
for (const key of data.toString()) if (/^[12]$/.test(key) && !answered) {
|
||||
answered = true; writeCall();
|
||||
const label = item.question.options[Number(key)-1].label;
|
||||
fs.appendFileSync(transcript, JSON.stringify({type:'user',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),
|
||||
toolUseResult:{answers:{[item.question.question]:label}},
|
||||
message:{role:'user',content:[{type:'tool_result',tool_use_id:'prerequisite',content:'answered'}]}})+'\n');
|
||||
process.stdout.write('\x1b[2J\x1b[H'+item.frame+'\nSETUP_ANSWERED\n');
|
||||
}
|
||||
});
|
||||
process.stdout.write('\x1b[2J\x1b[H'+item.frame);
|
||||
process.on('SIGINT', () => process.exit(0));
|
||||
process.stdin.resume();
|
||||
`);
|
||||
fs.chmodSync(fake, 0o755);
|
||||
const moduleUrl = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
|
||||
fs.writeFileSync(worker, `
|
||||
import {launchClaudePty} from ${JSON.stringify(moduleUrl('claude-pty-runner.ts'))};
|
||||
import {autoplanRoutingSetupInput} from ${JSON.stringify(moduleUrl('autoplan-setup-question.ts'))};
|
||||
import {readPlanCountTranscript} from ${JSON.stringify(moduleUrl('plan-count-transcript.ts'))};
|
||||
const results = await Promise.all(${JSON.stringify(cases)}.map(async item => {
|
||||
const session = await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:20000,env:{PREREQUISITE_REPLAY:JSON.stringify(item)}});
|
||||
try {
|
||||
await session.waitFor('Enter to select', {timeoutMs:10000,pollMs:20});
|
||||
const screen = await session.currentScreen();
|
||||
const before = readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
|
||||
const pending = before.calls.find(call => !call.answered && !call.failed);
|
||||
if (Boolean(pending) !== item.early) throw Error('Wrong initial native persistence state');
|
||||
const seen = new Set();
|
||||
const input = autoplanRoutingSetupInput(screen,seen,pending);
|
||||
if (input !== item.expected) throw Error('Expected skip input '+item.expected+', got '+JSON.stringify(input));
|
||||
session.send(input);
|
||||
await session.waitFor('SETUP_ANSWERED', {timeoutMs:10000,pollMs:20});
|
||||
const after = readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
|
||||
const call = after.calls[0];
|
||||
if (after.calls.length !== 1 || !call.answered) throw Error('Native answer was not persisted');
|
||||
const retained = await session.currentScreen();
|
||||
return {name:item.name,input,answer:call.answers[item.question.question],
|
||||
redraw:autoplanRoutingSetupInput(retained,seen),
|
||||
delayedIdentity:autoplanRoutingSetupInput(screen,seen,{...call,answered:false})};
|
||||
} finally {await session.close();}
|
||||
}));
|
||||
await Bun.write(${JSON.stringify(resultFile)},JSON.stringify(results));
|
||||
`);
|
||||
const child = Bun.spawn([process.execPath, worker], {
|
||||
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' },
|
||||
stdout: 'pipe', stderr: 'pipe',
|
||||
});
|
||||
const killer = setTimeout(() => child.kill('SIGKILL'), 25000);
|
||||
try {
|
||||
const [exit, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
|
||||
expect(exit, stdout + stderr).toBe(0);
|
||||
expect(JSON.parse(fs.readFileSync(resultFile, 'utf8'))).toEqual(cases.map(item => ({
|
||||
name:item.name,input:item.expected,answer:'Skip — proceed with standard review',redraw:null,delayedIdentity:null,
|
||||
})));
|
||||
for (const item of cases) {
|
||||
const events = fs.readFileSync(item.record, 'utf8').trim().split('\n').map(line => JSON.parse(line));
|
||||
expect(events.filter(event => event.type === 'input').map(event => event.data).join('')).toBe(item.expected);
|
||||
expect(() => process.kill(events[0].pid, 0)).toThrow();
|
||||
}
|
||||
} finally {
|
||||
clearTimeout(killer); child.kill('SIGKILL');
|
||||
for (const item of cases) {
|
||||
if (!fs.existsSync(item.record)) continue;
|
||||
const first = JSON.parse(fs.readFileSync(item.record, 'utf8').split('\n')[0]!);
|
||||
try { process.kill(first.pid, 'SIGKILL'); } catch { /* already reaped */ }
|
||||
}
|
||||
fs.rmSync(dir, {recursive:true,force:true});
|
||||
}
|
||||
}, 30000);
|
||||
|
||||
|
||||
// Exact current viewport from source-M's routing stall. Owned temporary paths
|
||||
// are retained as display text; no fixture path is accessed by this replay.
|
||||
const M_ROUTING_CAPTURE = "\n\n❯ /autoplan\n\n● Starting the autoplan pipeline — running the preamble first.\n\n● Bash(_SS=\"$HOME/.claude/skills/gstack/bin/gstack-skill-start\"\n [ -x \"$_SS\" ] || _SS=\".claude/skills/gstack/bin/gstack-skill-start\"…)\n ⎿ SKILL_START_PROTO: 1\n BRANCH: main\n PROACTIVE: true \n … +54 lines (ctrl+o to expand)\n ⎿ Allowed by auto mode classifier\n\n● The preamble ran. SESSION_KIND is interactive, SESSION_ID is 1144263-1788912944-701e8cc4. There's a one-time routing\n instruction to handle first.\n\n Let me check if CLAUDE.md exists and explore the repo before presenting the routing question.\n\n Read 1 file, listed 1 directory (ctrl+o to expand)\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\nPlanning:\n/tmp/gstack-paid-shard-2DwzUD/tmp/gstack-hermetic-1144068-Ep9FFb/with-skills/.claude/plans/scalable-bouncing-moth.md\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n ☐ Skill routing\n\n│ gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?\n│ <gstack-qid:routing-injection>\n\n❯ 1. Add routing rules (Recommended)\n Append skill routing rules to CLAUDE.md and commit it — /autoplan, /ship, /qa, and other skills will be suggested\n automatically when relevant.\n 2. No thanks, manual only\n Skip for now; you can invoke skills manually anytime. You won't be asked again.\n 3. Type something.\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n 4. Chat about this\n\nEnter to select · ↑/↓ to navigate · Esc to cancel\n";
|
||||
|
||||
describe('M routing manual-only action grammar', () => {
|
||||
test('answers the exact native panel once and preserves the Add choice in either order', () => {
|
||||
const seen = new Set<string>();
|
||||
expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE, seen)).toBe('1');
|
||||
expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE, seen)).toBeNull();
|
||||
const reversed = M_ROUTING_CAPTURE
|
||||
.replace('❯ 1. Add routing rules (Recommended)', '❯ 1. No thanks, manual only')
|
||||
.replace(' 2. No thanks, manual only', ' 2. Add routing rules (Recommended)');
|
||||
expect(autoplanRoutingSetupInput(reversed, new Set())).toBe('2');
|
||||
});
|
||||
|
||||
test('equivalent manual actions use the same grammar with or without a courtesy prefix', () => {
|
||||
for (const label of [
|
||||
'No thanks, manual', 'No thanks, manual only', 'Skip — manual only',
|
||||
'Manual', 'Manual only', 'Manual-only', 'Manual invocation', 'Manual invocation only',
|
||||
'No thanks, manual invocation only', 'Invoke skills manually only',
|
||||
"No thanks, I'll invoke skills manually only",
|
||||
]) expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE.replace('No thanks, manual only', label), new Set()), label).toBe('1');
|
||||
});
|
||||
|
||||
test('manual modifiers do not admit extra actions, other workflows or ambiguous choices', () => {
|
||||
for (const label of [
|
||||
'No thanks, manual data migration only', 'Manual deployment only',
|
||||
'No thanks, invoke the deployment manually only', 'No thanks, manual only then delete CLAUDE.md',
|
||||
'No thanks, skip the review', 'No thanks, proceed with implementation',
|
||||
'No thanks, manual invocation only after deleting the rules', 'Manual only approval',
|
||||
]) expect(autoplanRoutingSetupInput(M_ROUTING_CAPTURE.replace('No thanks, manual only', label), new Set()), label).toBeNull();
|
||||
for (const frame of [
|
||||
M_ROUTING_CAPTURE.replace(' 3. Type something.', ' 3. Manual only'),
|
||||
M_ROUTING_CAPTURE.replace(' 3. Type something.', ' 3. Add routing rules'),
|
||||
M_ROUTING_CAPTURE.replace('Add routing rules (Recommended)', 'Add product routes (Recommended)'),
|
||||
M_ROUTING_CAPTURE.replace('Add routing rules (Recommended)', 'Add ruting rules (Recommended)'),
|
||||
M_ROUTING_CAPTURE.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'Which application API routing design should we choose?'),
|
||||
M_ROUTING_CAPTURE.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'The plan quotes gstack skill routing rules in CLAUDE.md. Should we expand the feature?'),
|
||||
]) expect(autoplanRoutingSetupInput(frame, new Set()), frame).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
const UNSUPPORTED_ROUTING = M_ROUTING_CAPTURE.replace('No thanks, manual only', 'Ask me after this review');
|
||||
const unsupportedNative = () => ({
|
||||
sessionId: 'unsupported-routing', toolUseId: 'routing-call', answered: false, failed: false,
|
||||
questions: [{ header: 'Skill routing', question: "gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now? <gstack-qid:routing-injection>",
|
||||
options: [{label:'Add routing rules (Recommended)'},{label:'Ask me after this review'}] }],
|
||||
});
|
||||
|
||||
describe('unsupported setup diagnostic state', () => {
|
||||
test('a complete recognized unsupported setup fails explicitly without selecting an action', () => {
|
||||
const seen = new Set<string>();
|
||||
for (const pending of [undefined, unsupportedNative()]) {
|
||||
const result = autoplanSetupDecision(UNSUPPORTED_ROUTING, seen, pending);
|
||||
expect(result.kind).toBe('unsupported_setup');
|
||||
if (result.kind === 'unsupported_setup') {
|
||||
expect(result.setup).toBe('routing');
|
||||
expect(result.options).toEqual([{index:1,label:'Add routing rules (Recommended)'},{index:2,label:'Ask me after this review'}]);
|
||||
expect(result.identitySource).toBe(pending ? 'native-bound' : 'current-native-panel');
|
||||
}
|
||||
expect(seen.size).toBe(0);
|
||||
}
|
||||
expect(autoplanSetupDecision(PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Ask me later'), new Set()).kind).toBe('unsupported_setup');
|
||||
});
|
||||
|
||||
test('supported input is pure until sent; redraw and delayed metadata then wait', () => {
|
||||
const seen = new Set<string>();
|
||||
const decision = autoplanSetupDecision(M_ROUTING_CAPTURE, seen);
|
||||
expect(decision.kind).toBe('input'); expect(seen.size).toBe(0);
|
||||
if (decision.kind !== 'input') throw Error('Expected supported setup');
|
||||
expect(decision.input).toBe('1');
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
expect(autoplanSetupDecision(M_ROUTING_CAPTURE, seen).kind).toBe('waiting');
|
||||
const native = unsupportedNative(); native.questions[0]!.options[1]!.label = 'No thanks, manual only';
|
||||
expect(autoplanSetupDecision(M_ROUTING_CAPTURE, seen, native).kind).toBe('waiting');
|
||||
expect(autoplanSetupDecision(M_ROUTING_CAPTURE + '\n⏺ Continuing…', seen).kind).toBe('waiting');
|
||||
expect(autoplanSetupDecision(PREREQUISITE_CAPTURE, new Set()).kind).toBe('input');
|
||||
});
|
||||
|
||||
test('a substantive product or taste question mentioning office hours is not an unsupported prerequisite', () => {
|
||||
const fullQuestion = prerequisiteQuestion.question;
|
||||
const unsupported = PREREQUISITE_CAPTURE.replace('Skip — proceed with standard review', 'Ask me after this review');
|
||||
for (const [prompt, first, second] of [
|
||||
['No design doc exists for /office-hours integration. Should we build X or defer Y?', 'Build X', 'Defer Y'],
|
||||
['We should produce a design doc for /office-hours. Which visual style should this product use?', 'Minimal', 'Expressive'],
|
||||
['No design doc exists for /office-hours integration. Should we build X or defer Y?', 'Run /office-hours now', 'Defer Y'],
|
||||
['No design doc found. Run /office-hours first?', 'Run /office-hours now and delete the feature', 'Ask me later'],
|
||||
]) {
|
||||
const native = prerequisiteCall();
|
||||
native.questions[0]!.question = prompt!;
|
||||
native.questions[0]!.options = [{label:first!},{label:second!}];
|
||||
// Reconstruct from the actual full native layout, including footer.
|
||||
const frame = unsupported.replace(/│ No design doc[\s\S]*?Run \/office-hours first\?/, prompt!)
|
||||
.replace('1. Run /office-hours now', '1. ' + first)
|
||||
.replace('2. Ask me after this review', '2. ' + second);
|
||||
for (const pending of [undefined, native]) {
|
||||
expect(autoplanSetupDecision(frame, new Set(), pending).kind, prompt).toBe('unrelated');
|
||||
}
|
||||
}
|
||||
// Existing unsupported offer remains positively identified independently
|
||||
// of the unsupported opposite label; no exact question wording is needed.
|
||||
const native = prerequisiteCall();
|
||||
native.questions[0]!.question = fullQuestion.replace('Run /office-hours first?', 'Would you like to run /office-hours now?');
|
||||
native.questions[0]!.options[1]!.label = 'Ask me after this review';
|
||||
expect(autoplanSetupDecision(unsupported.replace('Run /office-hours first?', 'Would you like to run /office-hours now?'), new Set(), native).kind).toBe('unsupported_setup');
|
||||
});
|
||||
|
||||
test('routing identity still needs its explicit setup action before an unsupported failure', () => {
|
||||
for (const [first, second] of [['React', 'Vue'], ['Accept recommendation', 'Defer finding'], ['Add routing rules (Recommended)', 'Add routing rules (Recommended)']]) {
|
||||
const frame = UNSUPPORTED_ROUTING.replace('1. Add routing rules (Recommended)', '1. ' + first)
|
||||
.replace('2. Ask me after this review', '2. ' + second);
|
||||
const native = unsupportedNative();
|
||||
native.questions[0]!.options = [{label:first!},{label:second!}];
|
||||
for (const pending of [undefined,native]) expect(autoplanSetupDecision(frame,new Set(),pending).kind).toBe('waiting');
|
||||
}
|
||||
});
|
||||
|
||||
test('incomplete, stale, quoted, indented or mixed UI cannot establish unsupported setup', () => {
|
||||
const panel = UNSUPPORTED_ROUTING.slice(UNSUPPORTED_ROUTING.indexOf(' ☐ Skill routing'));
|
||||
for (const frame of [
|
||||
panel.replace('Enter to select · ↑/↓ to navigate · Esc to cancel', ''),
|
||||
panel.replace(' 2. Ask me after this review', ''),
|
||||
panel.replace(' 4. Chat about this', ''),
|
||||
panel.replace('❯ 1.', ' 1.'),
|
||||
panel.replace(' 2.', '❯ 2.'),
|
||||
panel.replace('1. Add', '1. [ ] Add'),
|
||||
panel.replace(' ☐ Skill routing', '← ☐ Skill routing ✔ Submit →'),
|
||||
panel + '\n⏺ Continuing the review now.',
|
||||
panel + '\n❯ 1. Different menu\n 2. Other choice',
|
||||
'Example panel:\n' + panel,
|
||||
'Quoted source:\n' + panel,
|
||||
'```text\n' + panel,
|
||||
'~~~~text\n```\n' + panel,
|
||||
panel.split('\n').map(line => ' ' + line).join('\n'),
|
||||
panel.split('\n').map(line => '> ' + line).join('\n'),
|
||||
]) expect(autoplanSetupDecision(frame, new Set()).kind, frame).not.toBe('unsupported_setup');
|
||||
expect(autoplanSetupDecision('```text\nearlier code\n```\n' + panel, new Set()).kind).toBe('unsupported_setup');
|
||||
const product = panel.replace("gstack works best when your project's CLAUDE.md includes skill routing rules. Should I add them now?", 'Which product API router should we use?');
|
||||
expect(autoplanSetupDecision(product, new Set()).kind).toBe('unrelated');
|
||||
});
|
||||
|
||||
test('mismatched, failed, answered, empty and multi-question metadata cannot diagnose this panel', () => {
|
||||
for (const mutate of [
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.failed = true; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.answered = true; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions = []; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions.push(structuredClone(call.questions[0]!)); },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { Object.assign(call.questions[0]!, {multiSelect:true}); },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.header = 'Other question'; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.question = 'Different question <gstack-qid:routing-injection>'; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.options[1]!.label = 'Different choice'; },
|
||||
(call: ReturnType<typeof unsupportedNative>) => { call.questions[0]!.options[1]!.label = 'No thanks, manual only'; },
|
||||
]) {
|
||||
const native = unsupportedNative(); mutate(native);
|
||||
expect(autoplanSetupDecision(UNSUPPORTED_ROUTING, new Set(), native).kind).not.toBe('unsupported_setup');
|
||||
}
|
||||
});
|
||||
|
||||
test('a supported native question clipped by the actual viewport preserves its existing input policy', async () => {
|
||||
const {createPtyScreen} = await import('./helpers/pty-screen');
|
||||
const {matchesNativePlanQuestion} = await import('./helpers/claude-pty-runner');
|
||||
const native = unsupportedNative();
|
||||
native.questions[0]!.question += '\n' + Array.from({length:41}, (_,i) =>
|
||||
`Routing context line ${i+1}: keep current project conventions and existing commands.`).join('\n');
|
||||
native.questions[0]!.options[1]!.label = 'No thanks, invoke manually';
|
||||
const frame = `☐ Skill routing\n${native.questions[0]!.question}\n❯ 1. Add routing rules (Recommended)\n 2. No thanks, invoke manually\n 3. Type something.\n 4. Chat about this\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const screen = await createPtyScreen(120,40);
|
||||
try {
|
||||
screen.write(frame.replace(/\n/g,'\r\n'));
|
||||
const visible = await screen.read();
|
||||
expect(visible).not.toContain('☐ Skill routing');
|
||||
expect(matchesNativePlanQuestion(visible,native)).toBe(true);
|
||||
const seen = new Set<string>();
|
||||
const decision = autoplanSetupDecision(visible,seen,native);
|
||||
expect(decision.kind).toBe('input');
|
||||
if (decision.kind !== 'input') throw new Error('Expected supported native input');
|
||||
expect(decision.input).toBe('1');
|
||||
expect(seen.size).toBe(0);
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
expect(autoplanSetupDecision(visible,seen,native).kind).toBe('waiting');
|
||||
} finally { await screen.dispose(); }
|
||||
});
|
||||
});
|
||||
|
||||
test.skipIf(process.platform === 'win32')('real PTY unsupported setup fails after ready with zero input and durable parsed evidence', async () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-unsupported-setup-'));
|
||||
const fake = path.join(dir, 'fake-claude');
|
||||
const worker = path.join(dir, 'worker.ts');
|
||||
const resultFile = path.join(dir, 'result.json');
|
||||
const cases = [false, true].map(early => ({
|
||||
name: early ? 'early' : 'deferred', early, cwd: path.join(dir, early ? 'early' : 'deferred'),
|
||||
events: path.join(dir, early ? 'early.jsonl' : 'deferred.jsonl'),
|
||||
evalDir: path.join(dir, early ? 'early-artifacts' : 'deferred-artifacts'),
|
||||
frame: UNSUPPORTED_ROUTING, native: unsupportedNative(),
|
||||
}));
|
||||
for (const item of cases) fs.mkdirSync(item.cwd);
|
||||
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
const item=JSON.parse(process.env.SETUP_DIAGNOSTIC_CASE);
|
||||
const event=value=>fs.appendFileSync(item.events,JSON.stringify(value)+'\n');
|
||||
event({kind:'startup',pid:process.pid});
|
||||
if(item.early){
|
||||
const folder=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','fixture');fs.mkdirSync(folder,{recursive:true});
|
||||
fs.writeFileSync(path.join(folder,item.name+'.jsonl'),JSON.stringify({type:'assistant',sessionId:item.name,isSidechain:false,cwd:process.cwd(),timestamp:new Date().toISOString(),message:{role:'assistant',content:[{type:'tool_use',id:'setup',name:'AskUserQuestion',input:{questions:item.native.questions}}]}})+'\n');
|
||||
}
|
||||
process.stdin.setRawMode?.(true);
|
||||
process.stdin.on('data',data=>event({kind:'input',data:data.toString()}));
|
||||
process.stdout.write('\x1b[2J\x1b[H'+item.frame);
|
||||
process.on('SIGINT',()=>process.exit(0));process.stdin.resume();
|
||||
`);
|
||||
fs.chmodSync(fake, 0o755);
|
||||
const url = (name: string) => pathToFileURL(path.resolve(import.meta.dir, 'helpers', name)).href;
|
||||
fs.writeFileSync(worker, `
|
||||
import * as fs from 'node:fs';
|
||||
import {launchClaudePty} from ${JSON.stringify(url('claude-pty-runner.ts'))};
|
||||
import {autoplanSetupDecision,autoplanRoutingSetupInput} from ${JSON.stringify(url('autoplan-setup-question.ts'))};
|
||||
import {readPlanCountTranscript} from ${JSON.stringify(url('plan-count-transcript.ts'))};
|
||||
import {createPlanCountSnapshotWriter} from ${JSON.stringify(url('plan-count-artifacts.ts'))};
|
||||
const results=[];
|
||||
for(const item of ${JSON.stringify(cases)}){
|
||||
const session=await launchClaudePty({cwd:item.cwd,observeScreen:true,timeoutMs:20000,env:{SETUP_DIAGNOSTIC_CASE:JSON.stringify(item)}});
|
||||
const result={name:item.name,config:session.hermeticConfigDir};
|
||||
try{
|
||||
await session.waitFor('Enter to select',{timeoutMs:10000,pollMs:20});
|
||||
const viewport=await session.currentScreen();
|
||||
const native=readPlanCountTranscript(session.hermeticConfigDir,item.cwd);
|
||||
const pending=native.calls.find(call=>!call.answered&&!call.failed);
|
||||
if(Boolean(pending)!==item.early)throw Error('Readiness did not establish expected metadata state');
|
||||
result.legacyInput=autoplanRoutingSetupInput(viewport,new Set(),pending);
|
||||
const decision=autoplanSetupDecision(viewport,new Set(),pending);
|
||||
if(decision.kind==='input')throw Error('Unexpected guessed input');
|
||||
if(decision.kind!=='unsupported_setup')throw Error('Expected unsupported_setup, got '+decision.kind);
|
||||
const save=createPlanCountSnapshotWriter({EVALS_RUN_ID:item.name,GSTACK_EVAL_DIR:item.evalDir});
|
||||
Object.assign(result,save({skillName:'autoplan',cwd:item.cwd,claudeConfigDir:session.hermeticConfigDir,raw:session.rawOutput(),visible:session.visibleText(),viewport,
|
||||
observation:{state:'unsupported_setup',unsupportedSetup:decision,native,retention:'UI and parsed metadata only; full parent JSONL not guaranteed.'}}));
|
||||
throw Error('UNSUPPORTED_SETUP_DIAGNOSTIC: '+decision.prompt);
|
||||
}catch(error){result.failed=true;result.error=String(error);}
|
||||
finally{await session.close();fs.rmSync(item.cwd,{recursive:true,force:true});}
|
||||
results.push(result);
|
||||
}
|
||||
await Bun.write(${JSON.stringify(resultFile)},JSON.stringify(results));
|
||||
process.exitCode=results.some(result=>result.failed)?1:0;
|
||||
`);
|
||||
const child = Bun.spawn([process.execPath, worker], {
|
||||
env: { ...process.env, BROWSE_TERMINAL_BINARY: fake, EVALS_HERMETIC: '1' }, stdout: 'pipe', stderr: 'pipe',
|
||||
});
|
||||
const killer = setTimeout(() => child.kill('SIGKILL'), 25000);
|
||||
try {
|
||||
const [exit, stdout, stderr] = await Promise.all([child.exited, new Response(child.stdout).text(), new Response(child.stderr).text()]);
|
||||
expect(exit, stdout + stderr).toBe(1);
|
||||
const results = JSON.parse(fs.readFileSync(resultFile, 'utf8'));
|
||||
expect(results.length).toBe(2);
|
||||
for (const [index, result] of results.entries()) {
|
||||
const item = cases[index]!;
|
||||
expect(result.failed).toBe(true);
|
||||
expect(result.error).toContain('UNSUPPORTED_SETUP_DIAGNOSTIC:');
|
||||
expect(result.legacyInput).toBeNull();
|
||||
expect(result.artifactError).toBeUndefined();
|
||||
const artifact = JSON.parse(fs.readFileSync(path.join(result.artifactDir, 'observation.json'), 'utf8'));
|
||||
expect(artifact.state).toBe('unsupported_setup');
|
||||
expect(artifact.native.calls.length).toBe(item.early ? 1 : 0);
|
||||
expect(artifact.retention).toContain('full parent JSONL not guaranteed');
|
||||
expect(fs.readFileSync(path.join(result.artifactDir, 'terminal.screen.log'), 'utf8')).toContain('Ask me after this review');
|
||||
expect(fs.readFileSync(path.join(result.artifactDir, 'terminal.raw.log'), 'utf8')).toContain('routing-injection');
|
||||
expect(fs.existsSync(item.cwd)).toBe(false);
|
||||
expect(fs.existsSync(result.config)).toBe(false);
|
||||
const events = fs.readFileSync(item.events, 'utf8').trim().split('\n').map(line => JSON.parse(line));
|
||||
expect(events.filter(event => event.kind === 'input')).toEqual([]);
|
||||
expect(() => process.kill(events[0].pid, 0)).toThrow();
|
||||
}
|
||||
} finally {
|
||||
clearTimeout(killer); child.kill('SIGKILL');
|
||||
for (const item of cases) {
|
||||
if (!fs.existsSync(item.events)) continue;
|
||||
const pid = JSON.parse(fs.readFileSync(item.events, 'utf8').split('\n')[0]!).pid;
|
||||
if (process.platform === 'linux') {
|
||||
try { if (fs.readFileSync('/proc/' + pid + '/cmdline', 'utf8').split('\0').includes(fake)) process.kill(pid, 'SIGKILL'); }
|
||||
catch { /* already reaped or PID no longer belongs to this fixture */ }
|
||||
}
|
||||
}
|
||||
fs.rmSync(dir, {recursive:true,force:true});
|
||||
}
|
||||
}, 30000);
|
||||
@@ -501,7 +501,7 @@ describe('installed snapshot helper in fresh shells', () => {
|
||||
});
|
||||
|
||||
test('all affected live workflow selectors include the executable continuity contract', () => {
|
||||
for (const name of ['autoplan-chain-pty', 'autoplan-dual-voice', 'carve-section-loading']) {
|
||||
for (const name of ['autoplan-dual-voice', 'carve-section-loading']) {
|
||||
expect(E2E_TOUCHFILES[name]).toContain('bin/gstack-autoplan-snapshot.ts');
|
||||
expect(E2E_TOUCHFILES[name]).toContain('test/autoplan-snapshot.test.ts');
|
||||
}
|
||||
|
||||
@@ -95,7 +95,7 @@ test('native readiness, timestamp, duplicate and observed-order rules remain int
|
||||
|
||||
test('the regression and exact public message select only the existing Autoplan workflow',()=>{
|
||||
for(const file of ['test/autoplan-with-result-au.test.ts','test/fixtures/autoplan-with-result-au.json'])
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['autoplan-chain-pty']);
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual([]);
|
||||
});
|
||||
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ describe('carved-skill cases each get a complete paid process budget', () => {
|
||||
expect(isPaidTestFile('test/' + file)).toBe(true);
|
||||
return calls.map(match => match[1]);
|
||||
});
|
||||
expect(covered.sort()).toEqual(Object.values(CARVE_GUARDS).filter(guard => guard.behavioral !== 'external').map(guard => guard.skill).sort());
|
||||
expect(covered.sort()).toEqual(Object.values(CARVE_GUARDS).filter(guard => guard.behavioral === 'plan' || guard.behavioral === 'prompt').map(guard => guard.skill).sort());
|
||||
expect(new Set(covered).size).toBe(covered.length);
|
||||
expect(selectPaidTestFiles(files.map(file => 'test/' + file), 'periodic').selected).toHaveLength(files.length);
|
||||
expect(selectPaidTestFiles(files.map(file => 'test/' + file), 'gate').selected).toHaveLength(0);
|
||||
|
||||
@@ -1,218 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import captured from './fixtures/ceo-annotation-aj.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const call = (index = 3): any => structuredClone(captured.cases.paired.calls[index]);
|
||||
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
|
||||
function edit(c: any, change: (s: string) => string) {
|
||||
const q = c.questions[0], answer = c.answers[q.question];
|
||||
q.question = change(q.question); c.answers = { [q.question]: answer };
|
||||
}
|
||||
|
||||
test('completed native receipt and retry findings retain identity through section annotations', () => {
|
||||
for (const index of [3, 4]) expect(ceoFirstReviewAUQ(fp(call(index)))).toBe(true);
|
||||
});
|
||||
|
||||
test('actual setup remains excluded before the two completed assertion findings', () => {
|
||||
let started = false;
|
||||
const phases = captured.cases.paired.calls.map(c => {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, true, false, false]);
|
||||
const approach = structuredClone(captured.cases.distinct.calls[2]);
|
||||
expect(ceoFirstReviewAUQ(fp(approach))).toBe(false);
|
||||
});
|
||||
|
||||
test('the new captured inputs belong only to the existing CEO count owner', () => {
|
||||
for (const dependency of ['test/ceo-annotation-aj.test.ts', 'test/fixtures/ceo-annotation-aj.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name))
|
||||
.toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
});
|
||||
|
||||
test('section references do not replace finding or native option identity', () => {
|
||||
for (const index of [3, 4]) {
|
||||
const c = call(index);
|
||||
edit(c, s => s.replace(/\(Sections? [^)]+\)/, '(Sections 3, 5 and 8, Error Handling)'));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
c.answers[c.questions[0].question] = c.questions[0].options[1].label;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
const renamed = call(index), q = renamed.questions[0], old = String(index - 2);
|
||||
edit(renamed, s => s.replace(new RegExp('Finding F' + old), 'Finding F9')
|
||||
.replace(new RegExp('^Recommendation: ' + old, 'm'), 'Recommendation: 9')
|
||||
.replace(new RegExp('^' + old + '([A-Z][)])', 'gm'), '9$1'));
|
||||
q.header = q.header.replace(/^F\d+/, 'F9');
|
||||
q.options.forEach((o: any) => { o.label = o.label.replace(/^\d+/, '9'); });
|
||||
renamed.answers = { [q.question]: q.options[0].label };
|
||||
expect(ceoFirstReviewAUQ(fp(renamed))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('source frames and conditional or missing assessments cannot supply a current finding', () => {
|
||||
for (const index of [3, 4]) for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```\n' + s + '\n```',
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: `$1`'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Suppose '),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, ''),
|
||||
(s: string) => s + '\nELI10: A second contradictory assessment.',
|
||||
(s: string) => s.replace(/: the (success|repeated)/, ': the hypothetical $1'),
|
||||
(s: string) => s.replace(/\(Sections? [^)]+\)/, '(Section 6, Quoted Source)'),
|
||||
(s: string) => s.replace(/\(Sections? [^)]+\)/, '(Section 6, Historical Example)'),
|
||||
]) { const c = call(index); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
});
|
||||
|
||||
test('same-brief withdrawals defeat a finding while attributed historical quotes do not', () => {
|
||||
for (const index of [3, 4]) {
|
||||
for (const tail of ['This issue is withdrawn.', 'We have withdrawn this finding.',
|
||||
'There is no current defect or unresolved issue.', `F${index - 2} is rejected.`]) {
|
||||
const c = call(index); edit(c, s => s + '\n' + tail); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
const quoted = call(index); edit(quoted, s => s + '\nOld note: "This issue is withdrawn."');
|
||||
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('administrative options and stale action rows do not amend the current contract', () => {
|
||||
for (const index of [3, 4]) {
|
||||
for (const wording of ['Start review', 'Pause', 'Write the completed report']) {
|
||||
const stale = call(index), native = stale.questions[0];
|
||||
native.options.forEach((o: any, i: number) => {
|
||||
o.label = `${index - 2}${String.fromCharCode(65 + i)}: ${wording}`;
|
||||
o.description = wording;
|
||||
});
|
||||
stale.answers = { [native.question]: native.options[0].label };
|
||||
expect(ceoFirstReviewAUQ(fp(stale))).toBe(false);
|
||||
}
|
||||
const c = call(index), q = c.questions[0];
|
||||
q.options.forEach((o: any, i: number) => {
|
||||
o.label = `${index - 2}${String.fromCharCode(65 + i)}: Archive the completed report ${i}`;
|
||||
o.description = 'Save the completed review for reference.';
|
||||
});
|
||||
c.answers = { [q.question]: q.options[0].label };
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
const noGap = call(index);
|
||||
edit(noGap, s => s.replace(/^D\d+[^\n]+/, `D4 — Finding F${index - 2} (Sections 2 and 6): where should the completed report be stored?`)
|
||||
.replace(/^ELI10: .+$/m, 'ELI10: The review is complete and all assertions already enforce the full contract.'));
|
||||
expect(ceoFirstReviewAUQ(fp(noGap))).toBe(false);
|
||||
const negated = call(index);
|
||||
edit(negated, s => s.replace(/asserts only/, 'does not assert only')
|
||||
.replace(/^ELI10: .+$/m, 'ELI10: The assertions enforce the complete receipt and retry contracts.'));
|
||||
expect(ceoFirstReviewAUQ(fp(negated))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('completion, recommendation, identity and actual offered options remain required', () => {
|
||||
for (const index of [3, 4]) for (const mutate of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.sessionId = ''; },
|
||||
(c: any) => { c.answers = {}; },
|
||||
(c: any) => { c.answers[c.questions[0].question] = 'Foreign answer'; },
|
||||
(c: any) => { c.questions[0].multiSelect = true; },
|
||||
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
|
||||
(c: any) => { c.questions[0].header = 'Approach'; },
|
||||
(c: any) => { c.questions[0].header = 'Finding 99'; },
|
||||
(c: any) => { c.questions[0].options[1].description = ''; },
|
||||
(c: any) => { c.questions[0].options[1].label = '99B: Different finding'; },
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, 'Recommendation: 99Z')),
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, '')),
|
||||
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-eng-review-finding>'),
|
||||
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
for (const index of [3, 4]) {
|
||||
const original = fp(call(index));
|
||||
expect(ceoFirstReviewAUQ({ ...original, signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...original, nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...original, options: original.options.slice(1) })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('completed dotted issue briefs retain their full identity and section option binding', () => {
|
||||
for (const source of captured.cases.distinct.calls.slice(4)) {
|
||||
expect(ceoFirstReviewAUQ(fp(source))).toBe(true);
|
||||
}
|
||||
let started = false;
|
||||
const phases = captured.cases.distinct.calls.map(c => {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, true, true, false, false, false, false, false]);
|
||||
});
|
||||
|
||||
test('dotted issue syntax never supplies missing current defect or remedy evidence', () => {
|
||||
for (const source of captured.cases.distinct.calls.slice(4)) {
|
||||
const identity = /\(Issue ([\d.]+)\)/.exec(source.questions[0]!.question)![1]!;
|
||||
for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Historical example: '),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: `$1`'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, ''),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The completed review has no current defect or unresolved issue.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The existing behavior satisfies every contract and needs no change.'),
|
||||
(s: string) => s + '\nThis issue has been resolved.',
|
||||
(s: string) => s + `\nIssue ${identity} is rejected.`,
|
||||
(s: string) => s.replace(/^Recommendation: \d+[A-Z]/m, 'Recommendation: 99Z'),
|
||||
(s: string) => s + '\n<gstack-qid:plan-eng-review-finding>',
|
||||
]) { const c = structuredClone(source); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
const sourceOnly = structuredClone(source), q = sourceOnly.questions[0]!;
|
||||
q.options.forEach((o, i) => { o.label = `${identity.split('.')[0]}${String.fromCharCode(65 + i)}: Archive the completed report ${i}`; o.description = 'Save the completed review.'; });
|
||||
sourceOnly.answers = { [q.question]: q.options[0]!.label };
|
||||
expect(ceoFirstReviewAUQ(fp(sourceOnly))).toBe(false);
|
||||
const quoted = structuredClone(source); edit(quoted, s => s + `\nOld note: "Issue ${identity} is rejected."`);
|
||||
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('dotted identities remain complete while the option prefix names the containing section', () => {
|
||||
for (const source of captured.cases.distinct.calls.slice(4)) {
|
||||
const identity = /\(Issue ([\d.]+)\)/.exec(source.questions[0]!.question)![1]!;
|
||||
for (const mutate of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.answers = {}; },
|
||||
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
|
||||
(c: any) => { c.questions[0].header = 'Approach'; },
|
||||
(c: any) => { c.questions[0].header = `Finding ${identity.split('.')[0]}`; },
|
||||
(c: any) => { c.questions[0].header = 'Issue 99.1'; },
|
||||
(c: any) => { c.questions[0].options[1].label = '99B: Borrowed option'; },
|
||||
(c: any) => { c.questions[0].options[1].description = ''; },
|
||||
(c: any) => { c.questions[0].multiSelect = true; },
|
||||
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
|
||||
(c: any) => edit(c, s => s.replace(`Issue ${identity}`, 'Issue 1.0')),
|
||||
]) { const c = structuredClone(source); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
const header = structuredClone(source); header.questions[0]!.header = `Finding ${identity}`;
|
||||
expect(ceoFirstReviewAUQ(fp(header))).toBe(true);
|
||||
const localQid = structuredClone(source); edit(localQid, s => s + '\n<gstack-qid:plan-ceo-review-handler-correctness>');
|
||||
expect(ceoFirstReviewAUQ(fp(localQid))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(source), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(source), nativeCall: undefined })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('an owning assessment declaration cannot relabel source or hypothetical prose as a current finding', () => {
|
||||
for (const source of [...captured.cases.paired.calls.slice(3), ...captured.cases.distinct.calls.slice(4)]) {
|
||||
for (const frame of [
|
||||
'The following is a quoted source excerpt.',
|
||||
'The following is a hypothetical example.',
|
||||
'This assessment is only a historical example.',
|
||||
]) {
|
||||
const c = structuredClone(source);
|
||||
edit(c, s => s.replace(/^ELI10: /m, 'ELI10: ' + frame + ' '));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
const quoted = structuredClone(source);
|
||||
edit(quoted, s => s.replace(/^(ELI10: .+)$/m, '$1 Old note: "The following is a hypothetical example."'));
|
||||
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
|
||||
}
|
||||
});
|
||||
@@ -1,139 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
import captured from './fixtures/ceo-annotation-header-at.json';
|
||||
|
||||
const call = (): any => structuredClone(captured.calls[2]);
|
||||
const fp = (value: any) => nativePlanCallFingerprint(value, 0, true);
|
||||
const matches = (value: any) => ceoFirstReviewAUQ(fp(value));
|
||||
function edit(value: any, change: (text: string) => string) {
|
||||
const q = value.questions[0], answer = value.answers[q.question];
|
||||
q.question = change(q.question);
|
||||
value.answers = { [q.question]: answer };
|
||||
}
|
||||
|
||||
test('the exact completed section-annotated mail rescue finding opens review', () => {
|
||||
const original = call();
|
||||
expect(matches(original)).toBe(true);
|
||||
expect(original).toEqual(captured.calls[2]);
|
||||
let started = false;
|
||||
const phases = captured.calls.map(value => {
|
||||
const phase = planCountQuestionPhase(fp(value), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, false, false, false, false, false, false]);
|
||||
});
|
||||
|
||||
test('descriptive headers and matching finding renames preserve the same rich decision', () => {
|
||||
for (const header of ['Email rescue', 'Mail failure', 'Receipt retry', 'Finding 2', 'Issue 2', 'F2 rescue']) {
|
||||
const value = call(); value.questions[0].header = header;
|
||||
expect(matches(value)).toBe(true);
|
||||
}
|
||||
for (const label of call().questions[0].options.map((option: any) => option.label)) {
|
||||
const value = call(); value.answers[value.questions[0].question] = label;
|
||||
expect(matches(value)).toBe(true);
|
||||
}
|
||||
const renamed = call();
|
||||
edit(renamed, text => text.replace('Finding 2 (Section 2', 'Finding 9 (Section 2')
|
||||
.replace(/^Recommendation: 2A/m, 'Recommendation: 9A'));
|
||||
renamed.questions[0].options.forEach((option: any) => { option.label = option.label.replace(/^2/, '9'); });
|
||||
renamed.answers = { [renamed.questions[0].question]: renamed.questions[0].options[0].label };
|
||||
expect(matches(renamed)).toBe(true);
|
||||
const decision = call(); edit(decision, text => text.replace(/^D2/, 'D19'));
|
||||
expect(matches(decision)).toBe(true);
|
||||
});
|
||||
|
||||
test('section metadata cannot override conflicting or malformed identities', () => {
|
||||
for (const header of ['Finding 9', 'Issue 9', 'F9 rescue', 'Finding 2.1', 'Finding zero', 'Section 9', 'Section 2']) {
|
||||
const value = call(); value.questions[0].header = header;
|
||||
expect(matches(value)).toBe(false);
|
||||
}
|
||||
for (const change of [
|
||||
(text: string) => text.replace('Finding 2 (Section 2, CRITICAL GAP)', 'Finding 0 (Section 2, CRITICAL GAP)'),
|
||||
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 0, CRITICAL GAP)'),
|
||||
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, Historical Example)'),
|
||||
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, Quoted Source)'),
|
||||
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2, CRITICAL GAP) (Section 9)'),
|
||||
(text: string) => text.replace('Finding 2 (Section 2, CRITICAL GAP)', 'Finding 2 and Finding 9 (Section 2, CRITICAL GAP)'),
|
||||
(text: string) => text.replace('(Section 2, CRITICAL GAP)', '(Section 2)'),
|
||||
(text: string) => text.replace(/^Recommendation: 2A/m, 'Recommendation: 9A'),
|
||||
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
|
||||
});
|
||||
|
||||
test('source, hypothetical, withdrawn or missing assessments do not open review', () => {
|
||||
for (const change of [
|
||||
(text: string) => 'Example: ' + text,
|
||||
(text: string) => '> ' + text,
|
||||
(text: string) => '```\n' + text + '\n```',
|
||||
(text: string) => text.replace('\nProject/branch/task:', '\nSource:\nProject/branch/task:'),
|
||||
(text: string) => text.replace(/^ELI10: /m, 'ELI10: If approved, '),
|
||||
(text: string) => text.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
|
||||
(text: string) => text.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: This handler has no current defect and needs no amendment.'),
|
||||
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: The plan needs no change.'),
|
||||
(text: string) => text.replace(/^ELI10: .+$/m, ''),
|
||||
(text: string) => text + '\nELI10: Another assessment.',
|
||||
(text: string) => text + '\nThis finding is withdrawn.',
|
||||
(text: string) => text + '\nThis finding is "withdrawn".',
|
||||
(text: string) => text + '\nThis finding is hypothetical.',
|
||||
(text: string) => text + '\nThis finding is not current.',
|
||||
(text: string) => text + '\nThis finding is no longer current.',
|
||||
(text: string) => text + '\nThis finding is "no longer current".',
|
||||
(text: string) => text + '\nThis finding is superseded.',
|
||||
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
|
||||
const historical = call(); edit(historical, text => text + '\nOld note: "This finding is withdrawn."');
|
||||
expect(matches(historical)).toBe(true);
|
||||
const resolvedHistory = call();
|
||||
edit(resolvedHistory, text => text.replace(/^(ELI10: .+)$/m, '$1 Old note: "This handler has no current defect and needs no amendment."'));
|
||||
expect(matches(resolvedHistory)).toBe(true);
|
||||
});
|
||||
|
||||
test('only current offered remedies can supply the amendment', () => {
|
||||
for (const prefix of ['Source: ', 'If approved, ', 'This remedy is withdrawn. ', 'This remedy is "withdrawn". ']) {
|
||||
const value = call();
|
||||
value.questions[0].options.forEach((option: any) => { option.description = prefix + option.description; });
|
||||
expect(matches(value)).toBe(false);
|
||||
}
|
||||
const report = call();
|
||||
report.questions[0].options.forEach((option: any, index: number) => {
|
||||
option.label = `2${String.fromCharCode(65 + index)}) Archive the completed report ${index}`;
|
||||
option.description = 'Save the completed review for reference.';
|
||||
});
|
||||
report.answers = { [report.questions[0].question]: report.questions[0].options[0].label };
|
||||
expect(matches(report)).toBe(false);
|
||||
});
|
||||
|
||||
test('the completed native identity, offered choice and answer remain required', () => {
|
||||
for (const change of [
|
||||
(value: any) => { value.answered = false; },
|
||||
(value: any) => { value.failed = true; },
|
||||
(value: any) => { value.sessionId = ''; },
|
||||
(value: any) => { value.toolUseId = ''; },
|
||||
(value: any) => { value.unansweredQuestionIndices = [0]; },
|
||||
(value: any) => { value.answeredAt = 'invalid'; },
|
||||
(value: any) => { value.answers = {}; },
|
||||
(value: any) => { value.answers[value.questions[0].question] = 'Foreign answer'; },
|
||||
(value: any) => { value.questions[0].multiSelect = true; },
|
||||
(value: any) => { value.questions.push(structuredClone(value.questions[0])); },
|
||||
(value: any) => { value.questions[0].header = 'Approach'; },
|
||||
(value: any) => { value.questions[0].options[1].description = ''; },
|
||||
(value: any) => { value.questions[0].options[1].label = '9B) Borrowed amendment'; },
|
||||
(value: any) => edit(value, text => text.replace(/^Recommendation: .+$/m, '')),
|
||||
(value: any) => edit(value, text => text + '\n<gstack-qid:plan-eng-review-finding>'),
|
||||
]) { const value = call(); change(value); expect(matches(value)).toBe(false); }
|
||||
const original = fp(call());
|
||||
for (const changed of [
|
||||
{ ...original, signature: 'foreign:call' },
|
||||
{ ...original, nativeCall: undefined },
|
||||
{ ...original, nativeQuestionIndex: 1 },
|
||||
{ ...original, options: original.options.slice(1) },
|
||||
]) expect(ceoFirstReviewAUQ(changed)).toBe(false);
|
||||
});
|
||||
|
||||
test('new retained inputs belong only to the CEO finding-count workflow', () => {
|
||||
for (const file of ['test/ceo-annotation-header-at.test.ts', 'test/fixtures/ceo-annotation-header-at.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name))
|
||||
.toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -1,261 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { capturePlanCountQuestion, nativePlanCallFingerprint, planCountQuestionInput } from './helpers/claude-pty-runner';
|
||||
import { pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import { pickCeoCountQuestion, pickCeoRecommendedApproach } from './helpers/ceo-approach-pick';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import recorded from './fixtures/ceo-approach-q-call.json';
|
||||
import pairedRecorded from './fixtures/ceo-approach-q-paired-call.json';
|
||||
import handoffs from './fixtures/ceo-completion-handoff-m-call.json';
|
||||
import recordedY from './fixtures/ceo-approach-y-call.json';
|
||||
import recordedAA from './fixtures/ceo-approach-aa-call.json';
|
||||
|
||||
function pending(source: NativePlanQuestionCall = recorded as NativePlanQuestionCall): NativePlanQuestionCall {
|
||||
const call = structuredClone(source);
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
return call;
|
||||
}
|
||||
const fingerprint = (call: NativePlanQuestionCall, preReview = true) => nativePlanCallFingerprint(call, 0, preReview);
|
||||
|
||||
describe('numbered native approach identity', () => {
|
||||
test('the actual AA question selects its offered recommendation with a projected pending binding', () => {
|
||||
// Only completed live versions survived capture. Preserve the actual A
|
||||
// answer; this projection tests routing, not live metadata availability.
|
||||
const call = pending(recordedAA as NativePlanQuestionCall);
|
||||
const active = capturePlanCountQuestion(screen(call), new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBe(call);
|
||||
expect(pickCeoCountQuestion(fingerprint(call), active)).toBe(2);
|
||||
expect(planCountQuestionInput(screen(call), active, 2)).toBe('2');
|
||||
expect(recordedAA.answers[recordedAA.questions[0]!.question]).toBe(recordedAA.questions[0]!.options[0]!.label);
|
||||
expect(pickCeoCountQuestion(fingerprint(recordedAA as NativePlanQuestionCall))).toBeNull();
|
||||
});
|
||||
|
||||
test('decision numbers and option positions may change together without changing policy', () => {
|
||||
for (const decision of ['2', '37']) {
|
||||
const call = pending(recordedAA as NativePlanQuestionCall);
|
||||
const q = call.questions[0]!;
|
||||
q.question = q.question.replace(/^D1/, `D${decision}`).replace('approach-d1>', `approach-d${decision}>`);
|
||||
q.options.reverse();
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
|
||||
q.options.unshift(q.options.pop()!);
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(3);
|
||||
}
|
||||
});
|
||||
|
||||
test('numbered identities must agree with the explicit decision and remain a supported approach id', () => {
|
||||
for (const id of ['plan-ceo-review-approach-d2', 'plan-ceo-review-approach-d0',
|
||||
'plan-ceo-review-approach-d01', 'plan-ceo-review-approach-d1-extra',
|
||||
'plan-eng-review-approach-d1', 'plan-ceo-review-mode-d1']) {
|
||||
const call = pending(recordedAA as NativePlanQuestionCall);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-review-approach-d1', id);
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
|
||||
}
|
||||
for (const prefix of ['', 'D2 — ', 'Example: D1 — ', '> D1 — ']) {
|
||||
const call = pending(recordedAA as NativePlanQuestionCall);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace(/^D1 — /, prefix);
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('numbered ids retain the native binding, phase, question and sole recommendation guards', () => {
|
||||
const call = pending(recordedAA as NativePlanQuestionCall);
|
||||
const fp = fingerprint(call);
|
||||
const unbound = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
|
||||
expect(pickCeoCountQuestion(fp, unbound)).toBeNull();
|
||||
expect(pickCeoCountQuestion({...fp, preReview: false})).toBeNull();
|
||||
expect(pickCeoCountQuestion({...fp, signature: 'foreign:call'})).toBeNull();
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('should this plan use?', 'should this plan not use?'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Mode'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
]) {
|
||||
const changed = pending(recordedAA as NativePlanQuestionCall); mutate(changed);
|
||||
expect(pickCeoRecommendedApproach(fingerprint(changed))).toBeNull();
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('Y named component approach menu', () => {
|
||||
const actualScreen = readFileSync(join(import.meta.dir, 'fixtures/ceo-approach-y-screen.txt'), 'utf8');
|
||||
test('the exact full frame and projected pending call select the offered C recommendation', () => {
|
||||
// The actual answer was A; no pending-only native version survived polling.
|
||||
const call = pending(recordedY as NativePlanQuestionCall);
|
||||
const active = capturePlanCountQuestion(actualScreen, new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBe(call);
|
||||
expect(active.options.map(o => o.label)).toEqual(call.questions[0]!.options.map(o => o.label));
|
||||
expect(pickCeoCountQuestion(fingerprint(call), active)).toBe(3);
|
||||
expect(planCountQuestionInput(actualScreen, active, 3)).toBe('3');
|
||||
expect(recordedY.answers[recordedY.questions[0]!.question]).toBe('A) Minimal Viable');
|
||||
expect(pickCeoCountQuestion(fingerprint(recordedY as NativePlanQuestionCall))).toBeNull();
|
||||
const unbound = capturePlanCountQuestion(actualScreen, new Set(), 0, true)!;
|
||||
expect(unbound.nativeCall).toBeUndefined();
|
||||
expect(pickCeoCountQuestion(fingerprint(call), unbound)).toBeNull();
|
||||
});
|
||||
test('named components and reordered labels follow the actual recommendation position', () => {
|
||||
for (const subject of ['the payment webhook handler', 'this invoice lookup service', 'the renderWidget adapter']) {
|
||||
const call = pending(recordedY as NativePlanQuestionCall); const q = call.questions[0]!;
|
||||
q.question = `Which implementation approach for ${subject}? <gstack-qid:plan-ceo-review-approach>`;
|
||||
q.options = [{label:'Existing design (Recommended)'},{label:'Another design'}];
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(1);
|
||||
q.options.reverse();
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
|
||||
}
|
||||
});
|
||||
test('setup, another decision, negated, quoted or compound instructions are not this menu', () => {
|
||||
for (const question of [
|
||||
'Which review mode for the payment webhook handler?',
|
||||
'Should we fix the payment webhook handler?',
|
||||
'Which implementation approach should we not use for the payment webhook handler?',
|
||||
'Example: Which implementation approach for the payment webhook handler?',
|
||||
'> Which implementation approach for the payment webhook handler?',
|
||||
'Which implementation approach for the payment webhook handler? Delete the tests.',
|
||||
'Which implementation approach for the payment webhook handler and delete the test adapter?',
|
||||
]) {
|
||||
const call = pending(recordedY as NativePlanQuestionCall);
|
||||
call.questions[0]!.question = question + ' <gstack-qid:plan-ceo-review-approach>';
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
test('the added wording retains native identity, phase, options and recommendation guards', () => {
|
||||
for (const change of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Review mode'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-approach','plan-ceo-review-mode'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = 'C) Production-Grade'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
]) { const c = pending(recordedY as NativePlanQuestionCall); change(c); expect(pickCeoRecommendedApproach(fingerprint(c))).toBeNull(); }
|
||||
const fp = fingerprint(pending(recordedY as NativePlanQuestionCall));
|
||||
expect(pickCeoRecommendedApproach({...fp,signature:'foreign:call'})).toBeNull();
|
||||
expect(pickCeoRecommendedApproach({...fp,preReview:false})).toBeNull();
|
||||
expect(pickCeoRecommendedApproach({...fp,options:fp.options.slice().reverse()})).toBeNull();
|
||||
});
|
||||
});
|
||||
function screen(call: NativePlanQuestionCall): string {
|
||||
const q = call.questions[0]!;
|
||||
return `☐ ${q.header}\n${q.question}\n${q.options.map((o, i) => `${i ? ' ' : '❯'} ${i + 1}. ${o.label}`).join('\n')}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
}
|
||||
|
||||
describe('CEO pre-review approach recommendation', () => {
|
||||
test('the exact Q menu changes the old default risk acceptance to its offered recommendation', () => {
|
||||
const call = pending();
|
||||
const visible = screen(call);
|
||||
const active = capturePlanCountQuestion(visible, new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBe(call);
|
||||
const before = pickCeoCompletionHandoff(fingerprint(call), active) ?? 1;
|
||||
const after = pickCeoCountQuestion(fingerprint(call), active) ?? 1;
|
||||
expect(before).toBe(1);
|
||||
expect(after).toBe(2);
|
||||
expect(planCountQuestionInput(visible, active, after)).toBe('2');
|
||||
expect(recorded.answers[recorded.questions[0]!.question]).toBe(recorded.questions[0]!.options[0]!.label);
|
||||
expect(pickCeoCountQuestion(fingerprint(recorded as NativePlanQuestionCall))).toBeNull();
|
||||
});
|
||||
|
||||
test('the paired first native approach uses the same offered recommendation policy', () => {
|
||||
const call = pending(pairedRecorded as NativePlanQuestionCall);
|
||||
const visible = screen(call);
|
||||
const active = capturePlanCountQuestion(visible, new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBe(call);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), active) ?? 1).toBe(1);
|
||||
const after = pickCeoCountQuestion(fingerprint(call), active) ?? 1;
|
||||
expect(after).toBe(2);
|
||||
expect(planCountQuestionInput(visible, active, after)).toBe('2');
|
||||
expect(pairedRecorded.answers[pairedRecorded.questions[0]!.question]).toBe(pairedRecorded.questions[0]!.options[0]!.label);
|
||||
expect(pickCeoCountQuestion(fingerprint(pairedRecorded as NativePlanQuestionCall))).toBeNull();
|
||||
});
|
||||
|
||||
test('paired approach grammar is function-agnostic and follows reordered options', () => {
|
||||
const call = pending(pairedRecorded as NativePlanQuestionCall);
|
||||
const q = call.questions[0]!;
|
||||
q.question = 'D3 — Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-review-impl-approach>';
|
||||
q.options = [{ label: 'A) Custom renderer' }, { label: 'B) Existing renderer (Recommended)' }];
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(2);
|
||||
q.options.reverse();
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(1);
|
||||
});
|
||||
|
||||
test.each([
|
||||
['non-approach question', 'Should the renderWidget() tests be deleted? <gstack-qid:plan-ceo-review-impl-approach>'],
|
||||
['negated question', 'Which implementation approach should the renderWidget() tests not use? <gstack-qid:plan-ceo-review-impl-approach>'],
|
||||
['negated test subject', 'Which implementation approach for not testing renderWidget()? <gstack-qid:plan-ceo-review-impl-approach>'],
|
||||
['wrong approach identity', 'Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-approach>'],
|
||||
['extra action before question', 'Delete the tests. Which implementation approach for the renderWidget() tests? <gstack-qid:plan-ceo-review-impl-approach>'],
|
||||
])('does not apply paired approach selection to %s', (_name, question) => {
|
||||
const call = pending(pairedRecorded as NativePlanQuestionCall);
|
||||
call.questions[0]!.question = question;
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBeNull();
|
||||
});
|
||||
|
||||
test('recommendation follows actual option position and arbitrary approach content, never seed words', () => {
|
||||
for (const order of [[0, 1, 2], [1, 2, 0], [2, 0, 1]]) {
|
||||
const call = pending();
|
||||
const q = call.questions[0]!;
|
||||
const options = [{ label: 'A) Compare two renderers' }, { label: 'B) Existing renderer (Recommended)' }, { label: 'C) Custom renderer' }];
|
||||
q.options = order.map(index => options[index]!);
|
||||
q.question = 'D1 — Which implementation approach should this plan use? <gstack-qid:plan-ceo-approach>';
|
||||
expect(pickCeoRecommendedApproach(fingerprint(call))).toBe(order.indexOf(1) + 1);
|
||||
}
|
||||
});
|
||||
|
||||
test.each([
|
||||
['no recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Secure Baseline'; }],
|
||||
['duplicate recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' (Recommended)'; }],
|
||||
['duplicate offered label', (c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = c.questions[0]!.options[1]!.label; }],
|
||||
['negated recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Not (Recommended)'; }],
|
||||
['conflicting recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Not recommended here (Recommended)'; }],
|
||||
['description-only recommendation', (c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'B) Secure Baseline'; c.questions[0]!.options[1]!.description = 'Recommended'; }],
|
||||
['unknown qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-approach', 'plan-ceo-security'); }],
|
||||
['missing qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('<gstack-qid:plan-ceo-approach>', ''); }],
|
||||
['malformed extra qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question += '<gstack-qid:broken'; }],
|
||||
['duplicate qid', (c: NativePlanQuestionCall) => { c.questions[0]!.question += '<gstack-qid:plan-ceo-approach>'; }],
|
||||
['negated approach question', (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('should this plan use?', 'should this plan not use?'); }],
|
||||
['non-approach question', (c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we accept this security risk? <gstack-qid:plan-ceo-approach>'; }],
|
||||
['non-approach header', (c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Review Mode'; }],
|
||||
['multi-select', (c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; }],
|
||||
['mixed packet', (c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); }],
|
||||
['failed native call', (c: NativePlanQuestionCall) => { c.failed = true; }],
|
||||
])('keeps the old caller/default policy for %s', (_name, change) => {
|
||||
const call = pending();
|
||||
change(call);
|
||||
const fp = fingerprint(call);
|
||||
expect(pickCeoRecommendedApproach(fp)).toBeNull();
|
||||
expect(pickCeoCountQuestion(fp)).toBe(pickCeoCompletionHandoff(fp));
|
||||
});
|
||||
|
||||
test('requires current native binding and pre-review phase', () => {
|
||||
const call = pending();
|
||||
const fp = fingerprint(call);
|
||||
expect(pickCeoRecommendedApproach({ ...fp, preReview: false })).toBeNull();
|
||||
expect(pickCeoRecommendedApproach({ ...fp, signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoRecommendedApproach({ ...fp, nativeQuestionIndex: 1 })).toBeNull();
|
||||
expect(pickCeoRecommendedApproach({ ...fp, options: fp.options.slice().reverse() })).toBeNull();
|
||||
const visibleOnly = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
|
||||
expect(visibleOnly.nativeCall).toBeUndefined();
|
||||
expect(pickCeoCountQuestion(fp, visibleOnly)).toBeNull();
|
||||
const foreign = '☐ Finding\nShould we add validation?\n❯ 1. Add fix\n 2. Defer\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
const active = capturePlanCountQuestion(foreign, new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBeUndefined();
|
||||
expect(pickCeoCountQuestion(fp, active)).toBeNull();
|
||||
});
|
||||
|
||||
test('the existing completed-review manual picker still runs after approach selection declines', () => {
|
||||
const call = structuredClone(handoffs.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
const fp = fingerprint(call, false);
|
||||
const expected = pickCeoCompletionHandoff(fp);
|
||||
expect(expected).not.toBeNull();
|
||||
expect(pickCeoCountQuestion(fp)).toBe(expected);
|
||||
});
|
||||
|
||||
test('both count callers use the composed picker while leaving first-scope and count predicates intact', () => {
|
||||
const caller = readFileSync(join(import.meta.dir, 'skill-e2e-plan-ceo-finding-count.test.ts'), 'utf8');
|
||||
expect(caller.match(/pickAUQ: pickCeoCountQuestion/g)).toHaveLength(2);
|
||||
expect(caller.match(/isFirstReviewAUQ: ceoFirstReviewAUQ/g)).toHaveLength(2);
|
||||
expect(caller).toContain('firstAUQPick: pickSkipInterview');
|
||||
});
|
||||
});
|
||||
@@ -1,89 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, nativePlanCallFingerprint, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import fixture from './fixtures/ceo-assertion-header-am-calls.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const calls = fixture.calls as AskUserQuestionFingerprint[];
|
||||
const findings = calls.slice(2);
|
||||
function change(fp: AskUserQuestionFingerprint, edit: (q: NonNullable<AskUserQuestionFingerprint['nativeCall']>['questions'][number]) => void) {
|
||||
const call = structuredClone(fp.nativeCall!);
|
||||
const selected = call.questions[0]!.options.findIndex(o => o.label === call.answers?.[call.questions[0]!.question]);
|
||||
edit(call.questions[0]!);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[selected]!.label };
|
||||
return nativePlanCallFingerprint(call, fp.observedAtMs, fp.preReview);
|
||||
}
|
||||
|
||||
for (const [i, fp] of findings.entries()) {
|
||||
test(`the actual completed assertion finding ${i + 1} starts review with its descriptive header`, () => {
|
||||
expect(ceoFirstReviewAUQ(fp)).toBe(true);
|
||||
});
|
||||
}
|
||||
test('routing and implementation layout remain setup', () => {
|
||||
for (const fp of calls.slice(0, 2)) expect(ceoFirstReviewAUQ(fp)).toBe(false);
|
||||
});
|
||||
test('the same current issue is already recognized with an explicit numbered header', () => {
|
||||
findings.forEach((fp, i) => expect(ceoFirstReviewAUQ(change(fp, q => { q.header = `Issue ${i + 1}`; }))).toBe(true));
|
||||
});
|
||||
test('a competing numbered header cannot borrow the title issue', () => {
|
||||
findings.forEach((fp, i) => expect(ceoFirstReviewAUQ(change(fp, q => { q.header = `Issue ${i + 2}`; }))).toBe(false));
|
||||
});
|
||||
test('only the completed owned native decision supplies the finding', () => {
|
||||
for (const fp of findings) {
|
||||
for (const mutate of [
|
||||
(x: AskUserQuestionFingerprint) => { x.nativeCall!.answered = false; },
|
||||
(x: AskUserQuestionFingerprint) => { x.nativeCall!.failed = true; },
|
||||
(x: AskUserQuestionFingerprint) => { x.nativeCall!.unansweredQuestionIndices = [0]; },
|
||||
(x: AskUserQuestionFingerprint) => { x.signature = 'foreign:call'; },
|
||||
(x: AskUserQuestionFingerprint) => { x.nativeCall!.answers = {}; },
|
||||
(x: AskUserQuestionFingerprint) => { x.options[0]!.label = 'different menu'; },
|
||||
]) {
|
||||
const modified = structuredClone(fp); mutate(modified);
|
||||
expect(ceoFirstReviewAUQ(modified)).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
test('descriptive headers and decision ordinals do not replace the issue identity', () => {
|
||||
for (const fp of findings) {
|
||||
for (const titlePrefix of ['D19', 'd4']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/^D\d+/, titlePrefix); }))).toBe(true);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.header = 'Test contract'; }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/^Recommendation: \d+/m, 'Recommendation: 9'); }))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('source, earlier and conditional framing cannot own the current assessment', () => {
|
||||
for (const fp of findings) {
|
||||
for (const prefix of ['Source excerpt:', 'The following assessment is hypothetical.', 'Earlier review assessment:', 'If approved:', 'Source:', 'Example:', 'Historical review:']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', `\n${prefix}\nELI10:`); }))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['Source excerpt. ', 'Previously, ', 'If approved, ']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', `ELI10: ${prefix}`); }))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\nArchived wording: "Source excerpt."\nELI10:'); }))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('the assertion gap and offered remedy must still be current', () => {
|
||||
for (const fp of findings) {
|
||||
for (const correction of ['This finding is withdrawn.', 'No current defect remains.']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + correction; }))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace(/^\d+[A-Z]\)[\s\S]*?(?=^Net:)/m, '');
|
||||
const issue = q.options[0]!.label.match(/^\d+/)![0];
|
||||
q.options.forEach((option, i) => {
|
||||
option.label = `${issue}${String.fromCharCode(65 + i)}: Keep the current assertion`;
|
||||
option.description = 'Leave the assertion unchanged.';
|
||||
});
|
||||
}))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('regression inputs belong only to the existing CEO finding owner without sparse paths', () => {
|
||||
for (const input of ['test/ceo-assertion-header-am.test.ts', 'test/fixtures/ceo-assertion-header-am-calls.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(input)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
|
||||
for (let i = 0; i < paths.length; i++) {
|
||||
expect(Object.hasOwn(paths, i)).toBe(true);
|
||||
expect(typeof paths[i]).toBe('string');
|
||||
}
|
||||
});
|
||||
@@ -1,71 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/ceo-completion-handoff-l-calls.json';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
|
||||
describe('CEO closed review with zero unresolved decisions', () => {
|
||||
test('the actual final handoff leaves all four independent issue and TODO decisions intact', () => {
|
||||
const input = calls();
|
||||
const original = structuredClone(input);
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of input) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 4, review: 4, administrative: 1 });
|
||||
expect(input.filter(c => /TODO/i.test(c.questions[0]!.header)).every(c =>
|
||||
!isCeoCompletionHandoff(fingerprint(c)))).toBe(true);
|
||||
expect(input).toEqual(original);
|
||||
});
|
||||
|
||||
test('the offered manual action binds to the active native menu in either order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = calls().at(-1)!;
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
const q = call.questions[0]!;
|
||||
if (reverse) q.options.reverse();
|
||||
const visible = `☐ ${q.header}\n${q.question}\n` + q.options.map((option, i) =>
|
||||
`${i ? ' ' : '❯'} ${i + 1}. ${option.label}`).join('\n') +
|
||||
'\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
const active = capturePlanCountQuestion(visible, new Set(), 0, false, call)!;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), { ...active, signature: 'other' })).toBeNull();
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('conditional, unresolved, substantive and unconfirmed variants are not handoffs', () => {
|
||||
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 unresolved', '1 unresolved'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 unresolved decisions.', '0 unresolved decisions after fixing receipt assertions.'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete', 'is not complete'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('ceo-next-step-eng-review', 'ceo-security-finding'); },
|
||||
c => { c.questions[0]!.header = 'Receipt gap'; },
|
||||
c => { c.questions[0]!.options.push({ label: 'Add the missing happy-path assertions' }); },
|
||||
c => { c.questions.push(calls()[4]!.questions[0]!); },
|
||||
c => { c.failed = true; },
|
||||
c => { c.answered = false; },
|
||||
c => { c.unansweredQuestionIndices = [0]; },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = calls().at(-1)!;
|
||||
mutate(call);
|
||||
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
const call = calls().at(-1)!;
|
||||
call.answers = { [call.questions[0]!.question]: 'First add another test' };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -2,9 +2,8 @@ import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal,
|
||||
nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import { hasNativePlanTerminal,
|
||||
nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/ceo-completion-handoff-m-call.json';
|
||||
import nextStepCapture from './fixtures/ceo-handoff-n-calls.json';
|
||||
@@ -12,325 +11,17 @@ import nextStepCapture from './fixtures/ceo-handoff-n-calls.json';
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const handoff = () => calls().at(-1)!;
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
|
||||
function reanswer(call: NativePlanQuestionCall) {
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('CEO completion described by a native navigation choice', () => {
|
||||
test('the exact seven-call session preserves three setup and three finding decisions', () => {
|
||||
let reviewStarted = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
const original = calls();
|
||||
for (const call of original) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), reviewStarted, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
reviewStarted = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 3, review: 3, administrative: 1 });
|
||||
expect(original).toEqual(calls());
|
||||
expect(original.slice(3, -1).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false, false]);
|
||||
});
|
||||
|
||||
test('the active pending handoff selects the actual manual option in either order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = handoff();
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
const q = call.questions[0]!;
|
||||
if (reverse) q.options.reverse();
|
||||
const screen = `☐ ${q.header}\n${q.question}\n❯ 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const active = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
|
||||
expect(active.nativeCall?.toolUseId).toBe(call.toolUseId);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(active)).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(capturePlanCountQuestion(screen, new Set(), 0, false)!)).toBeNull();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), { ...active, signature: 'another:call' })).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('completion placement is independent of the next-step wording and option order', () => {
|
||||
const call = handoff();
|
||||
const q = call.questions[0]!;
|
||||
q.question = 'D9 — Next steps: The review is done. Where should we go next? <gstack-qid:plan-ceo-review-next-step>';
|
||||
q.header = 'Next review';
|
||||
q.options[0]!.description = 'Eng review is the required shipping gate.';
|
||||
for (const description of [
|
||||
'CEO review found 3 specification gaps (all resolved). Continue manually.',
|
||||
'The CEO review identified gaps; all findings are resolved. Continue manually.',
|
||||
'CEO review is complete with 0 unresolved decisions. Continue manually.',
|
||||
]) {
|
||||
q.options[1]!.description = description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('conditional, unfinished, quoted and non-CEO recaps cannot supply completion', () => {
|
||||
for (const description of [
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved after adding tests).',
|
||||
'Eng review is the required shipping gate. If CEO review found 3 gaps (all resolved), continue.',
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved); one gap remains.',
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). There is an unresolved test issue.',
|
||||
'Eng review is the required shipping gate. The document says "CEO review found 3 gaps (all resolved)."',
|
||||
'Eng review is the required shipping gate. Design review found 3 gaps (all resolved).',
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps.',
|
||||
'Eng review is the required shipping gate. CEO review found 3 specification gaps (not all resolved).',
|
||||
'Eng review is the required shipping gate. CEO review did not find all gaps resolved.',
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). Also add a new test before proceeding.',
|
||||
'Eng review is the required shipping gate. CEO review found 3 gaps (all resolved). Please fix the new missing authorization check before proceeding.',
|
||||
]) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.options[0]!.description = description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), description).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('native identity, completion, required gate and exclusively administrative choices remain necessary', () => {
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Review complete only after fixing tests. What next? <gstack-qid:plan-ceo-review-next-step>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Should we finish reviewing? <gstack-qid:plan-ceo-review-next-step>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid:plan-ceo-security-issue>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = '<gstack-qid broken> ' + call.questions[0]!.question; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO decision'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add another TODO'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[0]!.label += ' and fix the missing test'; },
|
||||
(call: NativePlanQuestionCall) => { for (const option of call.questions[0]!.options) option.description = option.description?.replaceAll('required', 'optional'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions.push(calls()[3]!.questions[0]!); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
|
||||
(call: NativePlanQuestionCall) => { call.answered = false; },
|
||||
]) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
|
||||
}
|
||||
const addedWork = handoff();
|
||||
addedWork.answers = { [addedWork.questions[0]!.question]: 'First add another payment test' };
|
||||
expect(isCeoCompletionHandoff(fingerprint(addedWork))).toBe(false);
|
||||
});
|
||||
|
||||
test('the actual report and Exit order permits only the administrative freshness exception', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-handoff-'));
|
||||
const report = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(report, captured.report.content);
|
||||
const reportAt = Date.parse(captured.report.successfulResult.timestamp) / 1000;
|
||||
fs.utimesSync(report, reportAt, reportAt);
|
||||
const transcript = { status: 'ready' as const, calls: calls(), assistantMessages: [],
|
||||
planReadyRequests: structuredClone(captured.planReadyRequests) };
|
||||
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
|
||||
.map(call => `${call.sessionId}:${call.toolUseId}`));
|
||||
const startedAt = Date.parse('2026-09-09T00:15:27Z');
|
||||
expect(administrative.size).toBe(1);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
transcript.calls.push({ ...structuredClone(transcript.calls[3]!), toolUseId: 'new-test-obligation',
|
||||
answeredAt: captured.calls.at(-1)!.answeredAt });
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('native next-review navigation with a resolved CEO recap', () => {
|
||||
const retryCalls = () => structuredClone(captured.distinctRetry.calls) as NativePlanQuestionCall[];
|
||||
const retryHandoff = () => retryCalls().at(-1)!;
|
||||
|
||||
test('the captured retry preserves its four actual findings and the unchanged mechanical band', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
const original = retryCalls();
|
||||
for (const call of original) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 4, review: 4, administrative: 1 });
|
||||
expect(original).toEqual(retryCalls());
|
||||
// The transcript contains four individual findings. The unasked dispatcher
|
||||
// remedy remains a separate workflow-quality limitation, never a fifth call.
|
||||
expect(original.slice(4, -1).every(call => !isCeoCompletionHandoff(fingerprint(call)))).toBe(true);
|
||||
});
|
||||
|
||||
test('actual offered manual navigation still requires the matching pending native question', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = retryHandoff();
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
const q = call.questions[0]!;
|
||||
if (reverse) q.options.reverse();
|
||||
const screen = `☐ ${q.header}\n${q.question}\n❯ 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const active = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
|
||||
expect(active.nativeCall?.toolUseId).toBe(call.toolUseId);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), active)).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(active)).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(capturePlanCountQuestion(screen, new Set(), 0, false)!)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('partial, conditional, quoted or still-open recaps never establish this navigation boundary', () => {
|
||||
for (const recap of [
|
||||
'This CEO review resolved some security bugs.',
|
||||
'This CEO review resolved most security bugs.',
|
||||
'This CEO review resolved all but one security bugs.',
|
||||
'This CEO review resolved two of three security bugs.',
|
||||
'This CEO review only resolved the security bugs.',
|
||||
'This CEO review did not resolve the security bugs.',
|
||||
'If this CEO review resolved the security bugs, continue.',
|
||||
'The document says "This CEO review resolved the security bugs."',
|
||||
'This CEO review resolved the security bugs. One issue remains unresolved.',
|
||||
'This CEO review resolved the security bugs; validation of that remedy is still pending.',
|
||||
'This CEO review resolved the security bugs. Please add a new test first.',
|
||||
'This CEO review will resolve the security bugs.',
|
||||
]) {
|
||||
const call = retryHandoff();
|
||||
call.questions[0]!.options[0]!.description = 'Eng review is the required shipping gate. ' + recap;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), recap).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
}
|
||||
for (const question of [
|
||||
'Should we add a missing authorization test as the next step after this CEO review?',
|
||||
'The CEO review did not finish. What is the next step after this CEO review?',
|
||||
'Can you first fix the missing authorization check as the next step after this CEO review?',
|
||||
'If the CEO review finishes, what is the next step after this CEO review?',
|
||||
'Example: What is the next step after this CEO review?',
|
||||
]) {
|
||||
const call = retryHandoff();
|
||||
call.questions[0]!.question = question + ' <gstack-qid:plan-ceo-next-step>';
|
||||
reanswer(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), question).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call)), question).toBeNull();
|
||||
}
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Choose a fix for the missing test <gstack-qid:plan-ceo-next-step>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-next-step', 'plan-ceo-test-gap'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace(/ <gstack-qid:[^>]+>/, ''); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add a missing receipt assertion'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions.push(retryCalls()[4]!.questions[0]!); },
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
|
||||
]) {
|
||||
const call = retryHandoff();
|
||||
mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the final native report edit precedes handoff and still covers every real answer', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-retry-handoff-'));
|
||||
const report = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(report, captured.distinctRetry.reportContent);
|
||||
const reportAt = Date.parse(captured.distinctRetry.reportUpdate.at(-1)!.timestamp) / 1000;
|
||||
fs.utimesSync(report, reportAt, reportAt);
|
||||
const transcript = { status: 'ready' as const, calls: retryCalls(), assistantMessages: [],
|
||||
planReadyRequests: structuredClone(captured.distinctRetry.planReadyRequests) };
|
||||
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
|
||||
.map(call => `${call.sessionId}:${call.toolUseId}`));
|
||||
const startedAt = Date.parse('2026-09-09T00:23:30Z');
|
||||
expect(administrative.size).toBe(1);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
transcript.calls.push({ ...structuredClone(transcript.calls[4]!), toolUseId: 'new-independent-finding',
|
||||
answeredAt: transcript.calls.at(-1)!.answeredAt });
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('CEO completed next-step identity in native option order', () => {
|
||||
const input = () => structuredClone(nextStepCapture.calls) as NativePlanQuestionCall[];
|
||||
const actual = () => input().at(-1)!;
|
||||
|
||||
test('the complete native sequence retains two setup and four real issue decisions', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
const native = input();
|
||||
const original = structuredClone(native);
|
||||
for (const call of native) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 2, review: 4, administrative: 1 });
|
||||
expect(native).toEqual(original);
|
||||
});
|
||||
|
||||
test('only the positively bound pending menu selects its offered manual action', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = actual();
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'other:call' })).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('the observed identity cannot excuse unfinished work, a finding or a malformed native call', () => {
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete only after fixing authorization.'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete. One issue remains unresolved.'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('complete.', 'complete. Please fix the missing authorization test.'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'Should we add a missing authorization test before the next review?'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'We should fix the missing authorization test before the next review.'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('What next?', 'We should fix the missing authorization test. What next?'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('required shipping gate', 'optional review'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('ceo-plan-next-steps', 'ceo-plan-test-gap'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid:ceo-plan-next-steps>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'TODO'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.description = 'Proceed to fix the missing authorization test before Eng review.'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[0]!.label += ' and add a missing test'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions.push(input()[2]!.questions[0]!); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
|
||||
]) {
|
||||
const call = actual();
|
||||
mutate(call);
|
||||
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
}
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
|
||||
(call: NativePlanQuestionCall) => { call.answers = {}; },
|
||||
(call: NativePlanQuestionCall) => { call.answers = { [call.questions[0]!.question]: 'Build another feature' }; },
|
||||
]) {
|
||||
const call = actual();
|
||||
mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the actual report precedes handoff but the captured absent Exit remains incomplete', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-native-next-step-'));
|
||||
const report = path.join(dir, 'plan.md');
|
||||
|
||||
@@ -2,249 +2,19 @@ import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import { hasNativePlanTerminal } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/ceo-completion-handoff-o-call.json';
|
||||
import capturedQ from './fixtures/ceo-completion-handoff-q-call.json';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const handoff = () => calls().at(-1)!;
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
const reanswer = (call: NativePlanQuestionCall) => {
|
||||
const question = call.questions[0]!;
|
||||
call.answers = { [question.question]: question.options[0]!.label };
|
||||
return call;
|
||||
};
|
||||
|
||||
describe('closed CEO navigation with the native review-prefixed identity', () => {
|
||||
test('the actual six-call sequence preserves setup and both independent findings', () => {
|
||||
const original = calls();
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of original) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 3, review: 2, administrative: 1 });
|
||||
expect(original).toEqual(calls());
|
||||
expect(original.slice(3, 5).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false]);
|
||||
});
|
||||
|
||||
test('the offered manual action needs the matching pending native call in either order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = handoff();
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'foreign:call' })).toBeNull();
|
||||
}
|
||||
expect(pickCeoCompletionHandoff(fingerprint(handoff()))).toBeNull();
|
||||
});
|
||||
|
||||
test('closed navigation semantics are shared across the bounded review identity family', () => {
|
||||
for (const id of ['ceo-review-next-step', 'ceo-review-next-steps', 'ceo-review-next-review', 'ceo-plan-next-steps']) {
|
||||
for (const completion of ['done', 'complete', 'cleared']) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.question = call.questions[0]!.question
|
||||
.replace('ceo-review-next-steps', id).replace('CEO review done.', `CEO review ${completion}.`);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
|
||||
}
|
||||
}
|
||||
const sequencing = handoff();
|
||||
sequencing.questions[0]!.options[1]!.description = 'Once implementation is finished, run /plan-eng-review. After Eng review is complete, proceed to shipping.';
|
||||
expect(isCeoCompletionHandoff(fingerprint(sequencing))).toBe(true);
|
||||
});
|
||||
|
||||
test('a completed heading cannot conceal unresolved work or a substantive question', () => {
|
||||
for (const text of [
|
||||
'CEO review is not done. What\'s next?',
|
||||
'CEO review done only after fixing the missing authorization test. What\'s next?',
|
||||
'CEO review done. Should we add a missing authorization test before Eng?',
|
||||
'CEO review done. We should fix the missing authorization test. What\'s next?',
|
||||
'CEO review done. Do you want me to fix the missing authorization test? What\'s next?',
|
||||
'CEO review done. One contrast issue remains. What\'s next?',
|
||||
'CEO review done. Validation is still pending. What\'s next?',
|
||||
'CEO review done. Not all findings are resolved. What\'s next?',
|
||||
'CEO review done. One test issue is still open. What\'s next?',
|
||||
'CEO review done. There are not 0 unresolved decisions. What\'s next?',
|
||||
'CEO review done. If the tests pass, what\'s next?',
|
||||
'CEO review done. What\'s next? Once the tests pass, all decisions are resolved.',
|
||||
'CEO review done. What\'s next? After the authorization tests pass, the review is complete.',
|
||||
'CEO review done. What\'s next? The review is complete when authorization tests pass.',
|
||||
'CEO review done. What\'s next? Once the tests pass, all decisions will be resolved.',
|
||||
'CEO review done. What\'s next? All findings become resolved after the tests pass.',
|
||||
'Example: CEO review done. What\'s next?',
|
||||
]) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace("CEO review done. What's next?", text);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call))), text).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call)), text).toBeNull();
|
||||
}
|
||||
for (const description of [
|
||||
'Proceed to fix the missing authorization test before Eng.',
|
||||
'The contrast gap remains unresolved; handle it manually.',
|
||||
'Please add a new regression test before implementation.',
|
||||
'We could add a missing regression test before Eng.',
|
||||
'Do you want to add a new test before the next review?',
|
||||
]) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.options[1]!.description = description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), description).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('failed, partial, malformed, unrelated or mixed native calls remain substantive', () => {
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.multiSelect = true; },
|
||||
(call: NativePlanQuestionCall) => { call.questions.push(calls()[3]!.questions[0]!); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'Test gap'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('ceo-review-next-steps', 'ceo-review-test-gap'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question += ' <gstack-qid'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = call.questions[0]!.question.replace('required shipping gate', 'optional review'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.options[1]!.label = 'Add a missing receipt assertion'; },
|
||||
]) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
|
||||
}
|
||||
const freeform = handoff();
|
||||
freeform.answers![freeform.questions[0]!.question] = 'Please add another test first';
|
||||
expect(isCeoCompletionHandoff(fingerprint(freeform))).toBe(false);
|
||||
});
|
||||
|
||||
test('actual report edits precede the handoff and retain the strict native Exit and freshness checks', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-closed-navigation-'));
|
||||
const report = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(report, captured.reportContent);
|
||||
const reportAt = Date.parse(captured.reportUpdate.at(-1)!.timestamp) / 1000;
|
||||
fs.utimesSync(report, reportAt, reportAt);
|
||||
const transcript = { status: 'ready' as const, calls: calls(), assistantMessages: [],
|
||||
planReadyRequests: structuredClone(captured.planReadyRequests) };
|
||||
const administrative = new Set(transcript.calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
|
||||
.map(call => `${call.sessionId}:${call.toolUseId}`));
|
||||
const startedAt = Date.parse('2026-09-09T01:46:12Z');
|
||||
expect(administrative.size).toBe(1);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
transcript.planReadyRequests = [];
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
transcript.planReadyRequests = structuredClone(captured.planReadyRequests);
|
||||
transcript.calls.push({ ...structuredClone(transcript.calls[3]!), toolUseId: 'new-real-finding',
|
||||
answeredAt: transcript.calls.at(-1)!.answeredAt });
|
||||
expect(hasNativePlanTerminal(transcript, report, startedAt, 'plan_ready', administrative)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('CEO completion recap after native project metadata', () => {
|
||||
const qCalls = () => structuredClone(capturedQ.calls) as NativePlanQuestionCall[];
|
||||
const qHandoff = () => qCalls().at(-1)!;
|
||||
|
||||
test('the exact Q sequence keeps all three substantive calls and four setup calls', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of qCalls()) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, ceoStep0Boundary,
|
||||
ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 4, review: 3, administrative: 1 });
|
||||
expect(qCalls().slice(4, 7).map(call => isCeoCompletionHandoff(fingerprint(call)))).toEqual([false, false, false]);
|
||||
expect(isCeoCompletionHandoff(fingerprint(qHandoff()))).toBe(true);
|
||||
});
|
||||
|
||||
test('only the current offered manual option is selected, including reordered choices', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = qHandoff();
|
||||
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
|
||||
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'foreign:call' })).toBeNull();
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
expect(pickCeoCompletionHandoff(fingerprint(qHandoff()))).toBeNull();
|
||||
});
|
||||
|
||||
test('unconditional line recaps support ordinary completion wording and Eng sequencing', () => {
|
||||
for (const state of ['done and clear', 'done', 'complete', 'cleared']) {
|
||||
const call = qHandoff();
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('done and clear', state);
|
||||
call.questions[0]!.options[1]!.description = 'Once implementation is finished, run /plan-eng-review. After Eng review is complete, proceed to shipping.';
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('the recap cannot hide contradictory, conditional, quoted or new work in question or choices', () => {
|
||||
for (const extra of [
|
||||
'CEO review is not complete.', 'The review remains incomplete.', 'Not all decisions are resolved.',
|
||||
'One test gap remains.', 'Validation is still pending.', 'There are unresolved findings.',
|
||||
'Once tests pass, the CEO review will be complete.', 'All decisions resolved after tests pass.',
|
||||
'We should fix a missing authorization test.', 'We could repair a missing authorization check.',
|
||||
'Repair the missing authorization test.', 'Recommendation: repair the missing authorization test.',
|
||||
'We may repair the missing authorization test.', 'We might fix the missing authorization test.',
|
||||
'Proceed to add a new regression.', 'Do you want to add a missing test?',
|
||||
'```text\nCEO review is complete.', '> CEO review is complete.', 'Example: CEO review is complete.',
|
||||
]) {
|
||||
for (const target of ['question', 'description']) {
|
||||
const call = qHandoff();
|
||||
if (target === 'question') call.questions[0]!.question += `\n${extra}`;
|
||||
else call.questions[0]!.options[1]!.description += ` ${extra}`;
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call))), `${target}: ${extra}`).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call)), `${target}: ${extra}`).toBeNull();
|
||||
}
|
||||
}
|
||||
for (const first of [
|
||||
'Should we add a missing authorization test as the next step after this CEO review?',
|
||||
'The CEO review did not finish. What is next after this CEO review?',
|
||||
'Can you first fix authorization? What is next after this CEO review?',
|
||||
]) {
|
||||
const call = qHandoff();
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace("What's next after this CEO review?", first);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('native failures, mixed choices, absent gates and source copies cannot become administrative', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(qCalls()[4]!.questions[0]!); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Missing tests'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-next-step', 'plan-ceo-new-test'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate', 'optional check'); c.questions[0]!.options[0]!.description = 'Optional check.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix the missing assertion'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('ELI10:', ' ELI10:'); },
|
||||
]) {
|
||||
const call = qHandoff(); mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(reanswer(call)))).toBe(false);
|
||||
}
|
||||
const call = qHandoff();
|
||||
call.answers![call.questions[0]!.question] = 'Please fix another gap first';
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
});
|
||||
|
||||
test('actual full report and Exit chronology retain last substantive-answer freshness', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-metadata-navigation-'));
|
||||
const report = path.join(dir, 'plan.md');
|
||||
|
||||
@@ -1,949 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { capturePlanCountQuestion, ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captures from './fixtures/ceo-completion-handoff-calls.json';
|
||||
import currentHandoffs from './fixtures/ceo-completion-handoff-j-calls.json';
|
||||
import kHandoffs from './fixtures/ceo-completion-handoff-k-calls.json';
|
||||
import rCalls from './fixtures/ceo-completion-handoff-r-calls.json';
|
||||
import tHandoff from './fixtures/ceo-completion-handoff-t-call.json';
|
||||
import uHandoff from './fixtures/ceo-completion-handoff-u-call.json';
|
||||
import vHandoff from './fixtures/ceo-completion-handoff-v-call.json';
|
||||
import wHandoff from './fixtures/ceo-completion-handoff-w-call.json';
|
||||
|
||||
type CapturedCall = typeof captures.cases[number]['calls'][number];
|
||||
function nativeCall(record: CapturedCall, sessionId = 'native-capture'): NativePlanQuestionCall {
|
||||
return {
|
||||
sessionId, toolUseId: record.toolUseId, answered: true, failed: false,
|
||||
questions: [{ header: record.header, question: record.question,
|
||||
options: record.options.map(label => ({ label })), multiSelect: false }],
|
||||
answers: { [record.question]: record.answer }, unansweredQuestionIndices: [],
|
||||
};
|
||||
}
|
||||
const handoff = () => nativeCall(captures.cases[0]!.calls.at(-1)!);
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
|
||||
describe('W unconditional CLEAR recap and required Eng pronoun navigation', () => {
|
||||
const actual = () => structuredClone(wHandoff.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pending = (call: NativePlanQuestionCall) => {
|
||||
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
|
||||
return fingerprint(copy);
|
||||
};
|
||||
const answer = (call: NativePlanQuestionCall) => {
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
return call;
|
||||
};
|
||||
test('exact seven calls retain two issue decisions and select the offered manual action', () => {
|
||||
const calls = structuredClone(wHandoff.calls) as NativePlanQuestionCall[];
|
||||
expect(replay(calls, false, ceoFirstReviewAUQ))
|
||||
.toMatchObject({ step0Count: 4, reviewCount: 2, administrativeCount: 1 });
|
||||
expect(isCeoCompletionHandoff(fingerprint(actual()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(actual()))).toBe(2);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(actual()))).toBeNull();
|
||||
expect(calls).toEqual(wHandoff.calls);
|
||||
});
|
||||
test('case, gap count and pure navigation option order do not change the meaning', () => {
|
||||
const call = actual(); const q = call.questions[0]!;
|
||||
q.question = q.question.toLowerCase().replace(' — ', ' - ');
|
||||
q.options[0]!.description = q.options[0]!.description!.replace('2 assertion gaps', '12 assertion gaps');
|
||||
q.options.reverse(); answer(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
|
||||
});
|
||||
test('conditional, negated, quoted or additional question text is not a closed handoff', () => {
|
||||
const source = actual().questions[0]!.question;
|
||||
for (const question of [
|
||||
source.replace('is CLEAR.', 'is not CLEAR.'), source.replace('is CLEAR.', 'will be CLEAR.'),
|
||||
source.replace('is CLEAR.', 'is CLEAR after tests pass.'), 'Once ' + source,
|
||||
source.replace('required shipping gate', 'optional shipping check'),
|
||||
source.replace('Eng review', 'Design review'), source.replace('run it next?', 'repair its findings next?'),
|
||||
source + ' Remove the failing test.', source + ' Should we change the error contract?',
|
||||
'> ' + source, 'Example: ' + source, '`' + source + '`',
|
||||
source + ' <gstack-qid:ceo-plan-next-steps>',
|
||||
]) {
|
||||
const call = actual(); call.questions[0]!.question = question; answer(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), question).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(call)), question).toBeNull();
|
||||
}
|
||||
});
|
||||
test('every description sentence must be closed navigation, including unknown action verbs', () => {
|
||||
for (const extra of [
|
||||
'Delete the authorization test.', 'Grant access to all accounts.', 'Repair the missing assertion.',
|
||||
'One gap remains unresolved.', 'The CEO review is CLEAR only if we change the contract.',
|
||||
'The CEO review will be CLEAR after another fix.', 'Should we add another test?',
|
||||
'Quoted source: CEO review is CLEAR.',
|
||||
]) {
|
||||
for (const index of [0, 1]) {
|
||||
const call = actual(); call.questions[0]!.options[index]!.description += ' ' + extra;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call)), extra).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(call)), extra).toBeNull();
|
||||
}
|
||||
}
|
||||
for (const description of ['', 'This CEO review held scope and resolved some assertion gaps — eng review verifies the test structure is sound.',
|
||||
'This CEO review held scope and resolved 2 assertion gaps after changing the contract — eng review verifies the test structure is sound.']) {
|
||||
const call = actual(); call.questions[0]!.options[0]!.description = description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
test('native identity, complete answers, Eng/manual choices and a single question remain required', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix another issue' }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'New finding'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = 'Run /plan-design-review'; answer(c); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix remaining issues manually'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[1]!)); },
|
||||
]) { const call = actual(); mutate(call); expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false); }
|
||||
expect(pickCeoCompletionHandoff({ ...pending(actual()), signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(actual()), nativeCall: undefined })).toBeNull();
|
||||
});
|
||||
test('controlled report time excludes the handoff but still rejects a later real issue answer', () => {
|
||||
expect(wHandoff.provenance.reportMtimeMs).toBeNull(); // No historical filesystem-time claim.
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-w-handoff-'));
|
||||
try {
|
||||
const calls = structuredClone(wHandoff.calls) as NativePlanQuestionCall[];
|
||||
const issueAt = Date.parse(calls.at(-2)!.answeredAt!);
|
||||
const navigationAt = Date.parse(calls.at(-1)!.answeredAt!);
|
||||
const syntheticWritten = Math.floor((issueAt + navigationAt) / 2);
|
||||
const file = path.join(dir, 'report.md'); fs.writeFileSync(file, wHandoff.reportContent);
|
||||
fs.utimesSync(file, syntheticWritten / 1000, syntheticWritten / 1000);
|
||||
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: wHandoff.planReadyRequests };
|
||||
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
|
||||
const start = Date.parse('2026-09-09T09:28:55Z');
|
||||
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', new Set())).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', admin)).toBe(true);
|
||||
calls.at(-2)!.answeredAt = new Date(syntheticWritten + 1000).toISOString();
|
||||
expect(hasNativePlanTerminal(transcript, file, start, 'plan_ready', admin)).toBe(false);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
});
|
||||
});
|
||||
|
||||
describe('V closed CEO recap with a resolved-gap count', () => {
|
||||
const actual = () => structuredClone(vHandoff.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pending = (call: NativePlanQuestionCall) => {
|
||||
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
return fingerprint(call);
|
||||
};
|
||||
test('actual navigation stays outside the two issue decisions and selects manual', () => {
|
||||
expect(replay(structuredClone(vHandoff.calls) as NativePlanQuestionCall[], false, ceoFirstReviewAUQ))
|
||||
.toMatchObject({ step0Count: 3, reviewCount: 2, administrativeCount: 1 });
|
||||
expect(isCeoCompletionHandoff(fingerprint(actual()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(actual())) ?? 1).toBe(2);
|
||||
const reordered = actual(); reordered.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(pending(reordered))).toBe(1);
|
||||
});
|
||||
test('the actual report is fresh after issue decisions but before this navigation', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-v-handoff-'));
|
||||
try {
|
||||
const report = path.join(dir, 'report.md'); fs.writeFileSync(report, vHandoff.reportContent);
|
||||
const written = vHandoff.provenance.reportMtimeMs / 1000; fs.utimesSync(report, written, written);
|
||||
const calls = structuredClone(vHandoff.calls) as NativePlanQuestionCall[];
|
||||
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: vHandoff.planReadyRequests };
|
||||
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
|
||||
const start = Date.parse('2026-09-09T08:42:53Z');
|
||||
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', admin)).toBe(true);
|
||||
calls.at(-2)!.answeredAt = new Date(vHandoff.provenance.reportMtimeMs + 1).toISOString();
|
||||
expect(hasNativePlanTerminal(transcript, report, start, 'plan_ready', admin)).toBe(false);
|
||||
} finally { fs.rmSync(dir, {recursive:true,force:true}); }
|
||||
});
|
||||
test('the new recap cannot hide incomplete review, another remedy or altered gate', () => {
|
||||
const edits: Array<(call: NativePlanQuestionCall) => void> = [
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('0 critical gaps','1 critical gap'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('gaps resolved','gaps unresolved'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete','is complete only after tests pass'); },
|
||||
c => { c.questions[0]!.question += ' Repair the missing authorization test.'; },
|
||||
c => { c.questions[0]!.question += ' Should we remove the owner check?'; },
|
||||
c => { c.questions[0]!.options[0]!.description += ' Delete the failing test.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = 'The CEO review is NOT CLEARED until its gaps are resolved.'; },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate','optional review'); },
|
||||
c => { c.questions[0]!.header = 'New finding'; },
|
||||
c => { c.questions[0]!.options[1]!.label = 'Implement a new feature'; },
|
||||
];
|
||||
for (const edit of edits) {
|
||||
const c = actual(); edit(c); c.answers = {[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
|
||||
}
|
||||
});
|
||||
test('completed identity, offered answer and single question remain required', () => {
|
||||
for (const edit of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {[c.questions[0]!.question]:'Add a new task'}; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
]) { const c=actual();edit(c);expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
|
||||
expect(pickCeoCompletionHandoff({...pending(actual()),signature:'foreign:call'})).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({...pending(actual()),nativeCall:undefined})).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('U completed CEO metadata navigation with scoped review explanations', () => {
|
||||
const captured = () => structuredClone(uHandoff.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pending = (call: NativePlanQuestionCall) => {
|
||||
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
|
||||
return fingerprint(copy);
|
||||
};
|
||||
test('the actual six-call stream retains two issues and selects the offered manual stop', () => {
|
||||
const calls = structuredClone(uHandoff.calls) as NativePlanQuestionCall[];
|
||||
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 2, administrativeCount: 1 });
|
||||
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
|
||||
expect(calls).toEqual(uHandoff.calls);
|
||||
});
|
||||
test('native identity, completed answer and real option order remain required', () => {
|
||||
const call = captured(); call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix one more issue first' }; },
|
||||
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
|
||||
});
|
||||
test('unfinished, conditional, quoted and additional-work descriptions remain substantive', () => {
|
||||
for (const extra of [
|
||||
'Delete the failing regression test before Eng.', 'Remove the owner check before Eng.',
|
||||
'Change the guarantee to permit old results.', 'Rewrite the acceptance criteria before shipping.',
|
||||
'Repair the missing authorization test.', 'We may repair the missing authorization test.',
|
||||
'All findings become resolved after the tests pass.', 'There is an outstanding authorization gap.',
|
||||
'Should we add another test before Eng?', 'Stakes if we pick wrong: delete the owner check.',
|
||||
'No UI scope was detected, so the CEO review is not complete.',
|
||||
]) {
|
||||
const c = captured(); c.questions[0]!.options[1]!.description += ' ' + extra;
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
|
||||
}
|
||||
for (const [from, to] of [
|
||||
['The CEO review is done.', 'The CEO review is not done.'],
|
||||
['The CEO review is done.', 'The CEO review is done if tests pass.'],
|
||||
['Two assertion spec gaps were caught and resolved.', 'Not all assertion spec gaps were resolved.'],
|
||||
['Two assertion spec gaps were caught and resolved.', 'Two assertion spec gaps remain unresolved.'],
|
||||
['No UI scope was detected, so a design review is not needed.', 'The CEO review is not needed.'],
|
||||
['No UI scope was detected, so a design review is not needed.', 'No UI scope was detected, so a design review is not complete.'],
|
||||
['Stakes if we pick wrong:', 'The CEO review is complete only if we pick correctly:'],
|
||||
]) {
|
||||
const c = captured(); const q = c.questions[0]!; const old = q.question; q.question = old.replace(from!, to!);
|
||||
c.answers = { [q.question]: c.answers![old]! };
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
|
||||
}
|
||||
for (const prefix of ['> ', '```text\n', 'Example: ']) {
|
||||
const c = captured(); c.questions[0]!.options[1]!.description = prefix + c.questions[0]!.options[1]!.description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we fix the missing authorization check?'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Authorization gap'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add a new TODO' }); },
|
||||
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); expect(pickCeoCompletionHandoff(pending(c))).toBeNull(); }
|
||||
});
|
||||
test('metadata headings cannot shelter an extra obligation or conditional completion', () => {
|
||||
for (const extra of ['Delete the owner check.', 'Remove the failing regression.', 'Change the guarantee.',
|
||||
'Rewrite the acceptance criteria.', 'All decisions are resolved after the tests pass.',
|
||||
'Should we approve one more issue?', 'The CEO review is not complete.']) {
|
||||
const c = captured(); const q = c.questions[0]!; const old = q.question;
|
||||
q.question += ' ' + extra; c.answers = { [q.question]: c.answers![old]! };
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
|
||||
}
|
||||
});
|
||||
test('the retained pending Exit and report still require fresh substantive decisions', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-u-handoff-')); const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(file, uHandoff.reportContent);
|
||||
fs.utimesSync(file, uHandoff.reportAtMs / 1000, uHandoff.reportAtMs / 1000);
|
||||
const calls = structuredClone(uHandoff.calls) as NativePlanQuestionCall[];
|
||||
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: [{
|
||||
sessionId: uHandoff.pendingExit.sessionId, toolUseId: uHandoff.pendingExit.toolUseId,
|
||||
timestamp: uHandoff.pendingExit.timestamp, failed: false, source: 'pre_tool_use' as const,
|
||||
}] };
|
||||
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
|
||||
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
transcript.planReadyRequests[0]!.sessionId = 'foreign-session';
|
||||
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.sessionId = uHandoff.pendingExit.sessionId;
|
||||
calls[3]!.answeredAt = new Date(uHandoff.reportAtMs + 1000).toISOString();
|
||||
expect(hasNativePlanTerminal(transcript, file, uHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
});
|
||||
});
|
||||
|
||||
describe('T completed CEO next-review navigation', () => {
|
||||
const captured = () => structuredClone(tHandoff.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pending = (call: NativePlanQuestionCall) => {
|
||||
const copy = structuredClone(call); copy.answered = false; delete copy.answers; delete copy.unansweredQuestionIndices;
|
||||
return fingerprint(copy);
|
||||
};
|
||||
test('the actual nine-call stream retains five issues and selects the offered manual stop', () => {
|
||||
const calls = structuredClone(tHandoff.calls) as NativePlanQuestionCall[];
|
||||
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 5, administrativeCount: 1 });
|
||||
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
|
||||
expect(calls).toEqual(tHandoff.calls);
|
||||
});
|
||||
test('native identity, completed answer and real option order remain required', () => {
|
||||
const call = captured(); call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Fix one more issue first' }; },
|
||||
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); }
|
||||
});
|
||||
test('unfinished, conditional, quoted and additional-work descriptions remain substantive', () => {
|
||||
for (const extra of [
|
||||
'Delete the failing regression test before Eng.', 'Remove the owner check before Eng.',
|
||||
'Change the guarantee to permit old results.', 'Rewrite the acceptance criteria before shipping.',
|
||||
'Repair the missing authorization test.', 'We may repair the missing authorization test.',
|
||||
'All findings become resolved after the tests pass.', 'There is an outstanding authorization gap.',
|
||||
'Should we add another test before Eng?',
|
||||
]) {
|
||||
const c = captured(); c.questions[0]!.options[1]!.description += ' ' + extra;
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBeNull();
|
||||
}
|
||||
for (const replacement of ['resolved some findings', 'did not resolve all findings', 'will resolve all findings after tests pass']) {
|
||||
const c = captured(); c.questions[0]!.options[1]!.description = c.questions[0]!.options[1]!.description!.replace('resolved all findings', replacement);
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['> ', '```text\n', 'Example: ']) {
|
||||
const c = captured(); c.questions[0]!.options[1]!.description = prefix + c.questions[0]!.options[1]!.description;
|
||||
expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false);
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Should we fix the missing authorization check?'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Authorization gap'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add a new TODO' }); },
|
||||
]) { const c = captured(); mutate(c); expect(isCeoCompletionHandoff(fingerprint(c))).toBe(false); expect(pickCeoCompletionHandoff(pending(c))).toBeNull(); }
|
||||
});
|
||||
test('the retained pending Exit and report still require fresh substantive decisions', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-t-handoff-')); const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(file, tHandoff.reportContent);
|
||||
fs.utimesSync(file, tHandoff.reportAtMs / 1000, tHandoff.reportAtMs / 1000);
|
||||
const calls = structuredClone(tHandoff.calls) as NativePlanQuestionCall[];
|
||||
const transcript = { status: 'ready' as const, calls, assistantMessages: [], planReadyRequests: [{
|
||||
sessionId: tHandoff.pendingExit.sessionId, toolUseId: tHandoff.pendingExit.toolUseId,
|
||||
timestamp: tHandoff.pendingExit.timestamp, failed: false, source: 'pre_tool_use' as const,
|
||||
}] };
|
||||
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
|
||||
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
transcript.planReadyRequests[0]!.sessionId = 'foreign-session';
|
||||
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.sessionId = tHandoff.pendingExit.sessionId;
|
||||
calls[3]!.answeredAt = new Date(tHandoff.reportAtMs + 1000).toISOString();
|
||||
expect(hasNativePlanTerminal(transcript, file, tHandoff.startedAtMs, 'plan_ready', admin)).toBe(false);
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
});
|
||||
});
|
||||
|
||||
describe('native direct Eng/manual handoff with described CEO closure', () => {
|
||||
const captured = () => structuredClone(rCalls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pending = (call: NativePlanQuestionCall) => {
|
||||
const copy = structuredClone(call); copy.answered = false; delete copy.answers;
|
||||
delete copy.unansweredQuestionIndices;
|
||||
return fingerprint(copy);
|
||||
};
|
||||
test('actual R calls retain zero findings and choose offered manual instead of starting Eng', () => {
|
||||
const calls = structuredClone(rCalls) as NativePlanQuestionCall[];
|
||||
expect(replay(calls, false, ceoFirstReviewAUQ)).toMatchObject({ step0Count: 3, reviewCount: 0, administrativeCount: 1, reviewStarted: true });
|
||||
expect(replay(calls, false, ceoFirstReviewAUQ).reviewCount).toBeLessThan(2); // Existing paired floor still fails.
|
||||
expect(isCeoCompletionHandoff(fingerprint(captured()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(captured())) ?? 1).toBe(2);
|
||||
expect(calls).toEqual(rCalls);
|
||||
});
|
||||
test('manual choice follows real option order and still requires pending native identity', () => {
|
||||
const call = captured(); call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBe(1);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), signature: 'foreign-call' })).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...pending(call), nativeCall: undefined })).toBeNull();
|
||||
call.failed = true;
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
|
||||
});
|
||||
test('same native menu retains every incomplete, conditional, quoted or substantive obligation', () => {
|
||||
const changes: Array<(c: NativePlanQuestionCall) => void> = [
|
||||
c => { c.questions[0]!.question = 'Should we fix the missing authorization test before the next review?'; },
|
||||
c => { c.questions[0]!.question += ' First repair the missing assertion.'; },
|
||||
c => { c.questions[0]!.question = 'The review did not finish. ' + c.questions[0]!.question; },
|
||||
c => { c.questions[0]!.header = 'Authorization gap'; },
|
||||
c => { c.questions[0]!.question += ' <gstack-qid:ceo-security-finding>'; },
|
||||
c => { c.questions[0]!.options[1]!.label = 'Skip'; },
|
||||
c => { c.questions[0]!.options[1]!.label = 'Repair authorization before Eng'; },
|
||||
c => { c.questions[0]!.options.push({ ...c.questions[0]!.options[1]! }); },
|
||||
c => { c.questions[0]!.options.push({ label: 'Run /plan-design-review' }); },
|
||||
c => { c.questions[0]!.options[1]!.description = 'The CEO review is not clear.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = 'The CEO review remains incomplete.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = 'The CEO review is clear once tests pass.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = 'Once tests pass, the CEO review will be clear.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' All findings become resolved after tests pass.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' The contrast gap remains unresolved.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Not all decisions are resolved.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Repair the missing authorization test.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Recommendation: repair the missing assertion.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' We may repair the missing assertion.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' We must add the authorization test.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Delete the failing regression test before Eng.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Remove the owner check before Eng.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Change the guarantee to permit old results.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Rewrite the acceptance criteria before shipping.'; },
|
||||
c => { c.questions[0]!.options[1]!.description += ' Do you want me to fix the missing test?'; },
|
||||
c => { c.questions[0]!.options[1]!.description = 'Example: The CEO review is clear.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = '> The CEO review is clear.'; },
|
||||
c => { c.questions[0]!.options[1]!.description = '```text\nThe CEO review is clear.'; },
|
||||
];
|
||||
for (const change of changes) {
|
||||
const call = captured(); change(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
test('an unconditional completed recap permits next Eng sequencing but no failed or free-form answer', () => {
|
||||
const call = captured();
|
||||
call.questions[0]!.options[1]!.description = 'The CEO review is complete. Run /plan-eng-review after implementation and before shipping.';
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(pending(call))).toBe(2);
|
||||
call.answers = { [call.questions[0]!.question]: 'First fix the missing receipt assertion' };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
call.unansweredQuestionIndices = [0];
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
function replay(calls: NativePlanQuestionCall[], reviewStarted = true, firstReview = (_fp: ReturnType<typeof fingerprint>) => true) {
|
||||
const counts = { step0Count: 0, reviewCount: 0, administrativeCount: 0 };
|
||||
const classifications = [];
|
||||
for (const call of calls) {
|
||||
const fp = fingerprint(call);
|
||||
const phase = planCountQuestionPhase(fp, reviewStarted, ceoStep0Boundary,
|
||||
// A completion summary can mention defects; even a broad positive
|
||||
// first-finding predicate must not promote a handoff into coverage.
|
||||
firstReview, undefined, isCeoCompletionHandoff);
|
||||
if (phase.administrative) counts.administrativeCount++;
|
||||
else if (phase.preReview) counts.step0Count++;
|
||||
else counts.reviewCount++;
|
||||
reviewStarted = phase.reviewStarted;
|
||||
classifications.push(phase);
|
||||
}
|
||||
return { ...counts, reviewStarted, classifications };
|
||||
}
|
||||
|
||||
describe('CEO completion handoff classification and selection', () => {
|
||||
test('captured first attempts keep every finding/TODO and exclude only the handoff; substantive retry still fails its band', () => {
|
||||
for (const scenario of captures.cases) {
|
||||
const calls = scenario.calls.map(c => nativeCall(c, scenario.sessionId));
|
||||
const original = structuredClone(calls);
|
||||
const result = replay(calls);
|
||||
expect(result.reviewCount).toBe(scenario.expectedReviewCount);
|
||||
expect(result.administrativeCount).toBe(scenario.name === 'five-retry' ? 0 : 1);
|
||||
expect(result.step0Count).toBe(0);
|
||||
expect(calls).toEqual(original); // Classification never discards or rewrites native evidence.
|
||||
for (const [i, call] of calls.entries()) {
|
||||
if (/TODO/i.test(call.questions[0]!.header)) expect(result.classifications[i]!.administrative).toBeUndefined();
|
||||
}
|
||||
}
|
||||
expect(replay(captures.cases[2]!.calls.map(c => nativeCall(c))).reviewCount).toBeGreaterThan(7);
|
||||
});
|
||||
test('handoff-only replay adds no findings or setup and cannot establish a first finding', () => {
|
||||
const result = replay([handoff()], false);
|
||||
expect(result).toMatchObject({ step0Count: 0, reviewCount: 0, administrativeCount: 1, reviewStarted: false });
|
||||
expect(result.classifications[0]).toEqual({ preReview: false, reviewStarted: false, administrative: 'completion-handoff' });
|
||||
});
|
||||
test('manual/done action is selected in either option order only while the matching native question is pending', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = handoff(); call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
const fp = fingerprint(call);
|
||||
expect(pickCeoCompletionHandoff(fp)).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(fp)).toBe(false);
|
||||
}
|
||||
expect(pickCeoCompletionHandoff(fingerprint(handoff()))).toBeNull();
|
||||
});
|
||||
test('substantive choices mentioning another review retain the normal choice and finding count', () => {
|
||||
const call = nativeCall(captures.cases[0]!.calls[0]!);
|
||||
call.questions[0]!.question += ' Run /plan-eng-review next after deciding how to fix this issue.';
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(replay([call]).reviewCount).toBe(1);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
});
|
||||
test('mixed packets and unknown action choices are not classified as an administrative handoff', () => {
|
||||
const mixed = handoff();
|
||||
const finding = nativeCall(captures.cases[0]!.calls[0]!);
|
||||
mixed.questions.push(finding.questions[0]!);
|
||||
mixed.answers = { ...mixed.answers, ...finding.answers };
|
||||
expect(isCeoCompletionHandoff(fingerprint(mixed))).toBe(false);
|
||||
expect(replay([mixed]).reviewCount).toBe(1);
|
||||
mixed.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(mixed))).toBeNull();
|
||||
const unknown = handoff(); unknown.questions[0]!.options.push({ label: 'Add another payment test before continuing' });
|
||||
expect(isCeoCompletionHandoff(fingerprint(unknown))).toBe(false);
|
||||
});
|
||||
test('unknown identities and generic skip choices remain counted', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:plan-ceo-security-finding>'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Test gap'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Skip'; },
|
||||
]) {
|
||||
const call = handoff(); mutate(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(replay([call]).reviewCount).toBe(1);
|
||||
}
|
||||
const call = handoff(); call.answered = false;
|
||||
const mismatched = { ...fingerprint(call), signature: 'another-native-call' };
|
||||
expect(pickCeoCompletionHandoff(mismatched)).toBeNull();
|
||||
});
|
||||
test('pending, failed, partial and free-form answers never create an exclusion', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'First add a refund test' }; },
|
||||
]) {
|
||||
const call = handoff(); mutate(call);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('UI-only and unrelated pending metadata cannot steer the active menu', () => {
|
||||
const pending = handoff(); pending.answered = false; delete pending.answers;
|
||||
const q = pending.questions[0]!;
|
||||
const active = `☐ ${q.header}\n${q.question}\n❯ 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const bound = capturePlanCountQuestion(active, new Set(), 0, false, pending)!;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(pending), bound)).toBe(2);
|
||||
const uiOnly = capturePlanCountQuestion(active, new Set(), 0, false)!;
|
||||
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
|
||||
const issue = '☐ Security finding\nChoose how to parameterize the SQL query.\n❯ 1. Fix query\n 2. Add a TODO\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
const unbound = capturePlanCountQuestion(issue, new Set(), 0, false, pending)!;
|
||||
expect(unbound.nativeCall).toBeUndefined();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(pending), unbound)).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('completed CEO handoff with native next-step identity', () => {
|
||||
function capturedHandoff(): NativePlanQuestionCall {
|
||||
const question = 'D7 — CEO review is complete. Run /plan-eng-review next (the required shipping gate)? <gstack-qid:plan-ceo-review-next-step>';
|
||||
return {
|
||||
sessionId: 'e10cf0b4-525b-442d-9c2a-7a48d6b39f50',
|
||||
toolUseId: 'toolu_01FmkkRpoE3s6Y93KX6zLN1q',
|
||||
answered: true,
|
||||
failed: false,
|
||||
questions: [{
|
||||
question,
|
||||
header: 'Next review',
|
||||
multiSelect: false,
|
||||
options: [
|
||||
{ label: 'Run /plan-eng-review next (recommended)' },
|
||||
{ label: "Skip — I'll handle reviews manually" },
|
||||
],
|
||||
}],
|
||||
answers: { [question]: 'Run /plan-eng-review next (recommended)' },
|
||||
unansweredQuestionIndices: [],
|
||||
};
|
||||
}
|
||||
|
||||
test('captured completed-review menu is administrative and retains every independent finding and TODO', () => {
|
||||
const calls = captures.cases[1]!.calls.slice(0, -1).map(c => nativeCall(c));
|
||||
const result = replay([...calls, capturedHandoff()]);
|
||||
expect(result).toMatchObject({ reviewCount: 4, administrativeCount: 1, step0Count: 0 });
|
||||
expect(result.classifications.slice(0, -1).every(p => !p.administrative)).toBe(true);
|
||||
});
|
||||
|
||||
test('only the positively bound pending handoff selects manual, in either option order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = capturedHandoff();
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('incomplete review, missing gate, findings, mixed choices, and unoffered answers stay substantive', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('is complete', 'has an unresolved test gap'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('required shipping gate', 'optional follow-up'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-ceo-review-next-step', 'plan-ceo-security-finding'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'TODO: email queue'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add missing staging validation to this plan' }); },
|
||||
]) {
|
||||
const call = capturedHandoff();
|
||||
mutate(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(replay([call]).reviewCount).toBe(1);
|
||||
}
|
||||
const call = capturedHandoff();
|
||||
call.answers = { [call.questions[0]!.question]: 'First add the missing retry test' };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
const CAPTURED_PAIRED_RETRY_CALLS: NativePlanQuestionCall[] = [
|
||||
{
|
||||
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
|
||||
"toolUseId": "toolu_01CG8hh817d7CvFk9kH5ZFW4",
|
||||
"questions": [
|
||||
{
|
||||
"question": "D6 — Section 2 finding: the 502 failure path test's assertion is under-specified. What does 'fails clean' mean as an observable outcome? <gstack-qid:plan-ceo-fails-clean>",
|
||||
"header": "502 failure mode",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Specify the exception type in the plan (Recommended)",
|
||||
"description": "Update the plan to name the exception class processPayment() raises after 502 exhaustion (e.g. 'assert raises Stripe::APIConnectionError' or 'assert raises PaymentFailedError'). The test must assert a concrete observable: the exception class, not just 'something goes wrong.' Effort: add 1 line to the plan. Verify: test fails with wrong exception type.",
|
||||
"preview": "REMEDY:\n Plan change: add to item 2 under ## Tests:\n 'The 502 test must assert the specific exception class\n (or nil return, or error struct) processPayment() raises\n after retry exhaustion. The test factory already exposes\n mock call history; the test should also assert exactly 2\n charge attempts and 1 backoff sleep call.'\n\nWhy: without this, the implementer will write\n expect { processPayment() }.not_to raise_error\nwhich passes on the wrong behavior (swallowed exception)."
|
||||
},
|
||||
{
|
||||
"label": "Accept 'fails clean' as implementation-determined",
|
||||
"description": "Trust the implementer to look at processPayment() and assert whatever behavior they find. The test is still useful. Risk: if processPayment() silently swallows the error (no raise, no return value), the test will pass even when payment silently fails."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"answered": true,
|
||||
"failed": false,
|
||||
"answers": {
|
||||
"D6 — Section 2 finding: the 502 failure path test's assertion is under-specified. What does 'fails clean' mean as an observable outcome? <gstack-qid:plan-ceo-fails-clean>": "Specify the exception type in the plan (Recommended)"
|
||||
},
|
||||
"unansweredQuestionIndices": [],
|
||||
"answeredAt": "2026-09-08T20:57:40.308Z"
|
||||
},
|
||||
{
|
||||
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
|
||||
"toolUseId": "toolu_01Bc1mwoqXgNQK7NVx8MA21L",
|
||||
"questions": [
|
||||
{
|
||||
"question": "D7 — Section 4 finding: the happy path assertion 'correct receipt is generated' needs to be field-specific to be a correctness test. <gstack-qid:plan-ceo-receipt-fields>",
|
||||
"header": "Receipt assertion",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Add field-level assertion requirement to the plan (Recommended)",
|
||||
"description": "Update the plan: the happy path test must assert specific receipt fields (at minimum: amount matches charged amount, stripe_charge_id matches the mock's returned charge ID). Prevents the test from being just a nil-check smoke test. Effort: add 1 line to the plan. Verify: test fails if receipt has wrong charge ID.",
|
||||
"preview": "REMEDY:\n Plan change: add to item 1 under ## Tests:\n 'The happy path test must assert field-level receipt\n correctness: at minimum, the receipt amount equals the\n charged amount and the receipt stripe_charge_id matches\n the charge ID returned by the Stripe mock.\n assert receipt.amount == expected_amount\n assert receipt.stripe_charge_id == mock_charge.id'\n\nWhy: 'assert receipt is generated' is a smoke test.\n It passes even if receipt contains wrong amount or\n no charge ID, which is the correctness we care about."
|
||||
},
|
||||
{
|
||||
"label": "Accept 'correct receipt generated' as-is",
|
||||
"description": "Leave it to the implementer. Risk: a future refactor that breaks receipt field accuracy (e.g., amount in wrong currency unit) could pass the test silently."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"answered": true,
|
||||
"failed": false,
|
||||
"answers": {
|
||||
"D7 — Section 4 finding: the happy path assertion 'correct receipt is generated' needs to be field-specific to be a correctness test. <gstack-qid:plan-ceo-receipt-fields>": "Add field-level assertion requirement to the plan (Recommended)"
|
||||
},
|
||||
"unansweredQuestionIndices": [],
|
||||
"answeredAt": "2026-09-08T20:58:04.465Z"
|
||||
},
|
||||
{
|
||||
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
|
||||
"toolUseId": "toolu_014MYMCNbEGQfNYwqFkQQmKm",
|
||||
"questions": [
|
||||
{
|
||||
"question": "D8 — TODO candidate: test the orphaned-payment scenario (receipt-builder fails after Stripe charge succeeds). Add to TODOS.md? <gstack-qid:plan-ceo-todo-orphaned-payment>",
|
||||
"header": "TODO: orphaned payment",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Add to TODOS.md (Recommended)",
|
||||
"description": "Stripe charge succeeds, then receipt-builder throws — processPayment() returns an error to the caller. If the caller retries, the card gets charged twice. The existing receipt-builder regression tests cover receipt-builder itself, but nothing tests processPayment()'s behavior in this specific sequence. P2, M effort (human: ~1h / CC: ~8min).",
|
||||
"preview": "TODO entry:\n What: Test orphaned-payment scenario in processPayment()\n Why: Stripe charge succeeds, receipt-builder throws,\n caller retries → double charge. No test covers this.\n Where: payment test factory already supports this setup.\n Effort: M (human ~1h / CC ~8min)\n Priority: P2\n Depends on: this PR (test infra in place)"
|
||||
},
|
||||
{
|
||||
"label": "Skip — not valuable enough",
|
||||
"description": "The receipt-builder regression tests provide sufficient coverage. Double-charge scenario is handled by idempotency keys at the Stripe level."
|
||||
},
|
||||
{
|
||||
"label": "Build it now in this PR",
|
||||
"description": "Add a third test case to this PR: receipt-builder throws after successful charge — assert processPayment() returns the expected error and Stripe mock shows only 1 charge attempt (no retry on receipt failure). Expands scope from HOLD SCOPE decision."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"answered": true,
|
||||
"failed": false,
|
||||
"answers": {
|
||||
"D8 — TODO candidate: test the orphaned-payment scenario (receipt-builder fails after Stripe charge succeeds). Add to TODOS.md? <gstack-qid:plan-ceo-todo-orphaned-payment>": "Add to TODOS.md (Recommended)"
|
||||
},
|
||||
"unansweredQuestionIndices": [],
|
||||
"answeredAt": "2026-09-08T20:59:02.919Z"
|
||||
},
|
||||
{
|
||||
"sessionId": "eaedca8a-f52b-4739-a559-3f330e10b3c6",
|
||||
"toolUseId": "toolu_019ppgizjxzRiJd2QXPV7rYQ",
|
||||
"questions": [
|
||||
{
|
||||
"question": "D9 — CEO review complete. Run /plan-eng-review next? <gstack-qid:plan-ceo-next-review>",
|
||||
"header": "Next review",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Run /plan-eng-review next (Recommended)",
|
||||
"description": "Eng review is the required shipping gate. It covers architecture, code quality, and test correctness at the code level — what the CEO review doesn't dig into. The 2 spec gaps found here (exception type, receipt fields) should be verified at the code level too."
|
||||
},
|
||||
{
|
||||
"label": "Skip — handle reviews manually",
|
||||
"description": "Proceed without running eng review now. You can run it later with /plan-eng-review. Note: eng review is the only gate that blocks shipping by default."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"answered": true,
|
||||
"failed": false,
|
||||
"answers": {
|
||||
"D9 — CEO review complete. Run /plan-eng-review next? <gstack-qid:plan-ceo-next-review>": "Run /plan-eng-review next (Recommended)"
|
||||
},
|
||||
"unansweredQuestionIndices": [],
|
||||
"answeredAt": "2026-09-08T21:03:34.802Z"
|
||||
}
|
||||
];
|
||||
|
||||
|
||||
describe('completed CEO next-review declaration and final report order', () => {
|
||||
test('canonical identity alone never replaces actual completion and the next-review header', () => {
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.question = 'Should we finish reviewing? <gstack-qid:plan-ceo-next-steps>'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[0]!.header = 'New security issue'; },
|
||||
]) {
|
||||
const call = structuredClone(CAPTURED_PAIRED_RETRY_CALLS.at(-1)!);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('plan-ceo-next-review', 'plan-ceo-next-steps');
|
||||
call.questions[0]!.options[1]!.label = "Skip — I'll handle reviews manually";
|
||||
mutate(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the captured retry keeps its three findings/TODOs and recognizes only the completed handoff', () => {
|
||||
expect(replay(structuredClone(CAPTURED_PAIRED_RETRY_CALLS))).toMatchObject({
|
||||
reviewCount: 3, administrativeCount: 1, step0Count: 0,
|
||||
});
|
||||
const call = structuredClone(CAPTURED_PAIRED_RETRY_CALLS.at(-1)!);
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(2);
|
||||
call.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBe(1);
|
||||
});
|
||||
|
||||
test('a report written before the administrative handoff can reach the real plan-approval gate', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-handoff-order-'));
|
||||
const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(file, '# Plan\n\n## GSTACK REVIEW REPORT\n\n' +
|
||||
'| Review | Runs | Status | Findings |\n|---|---|---|---|\n| CEO | 1 | COMPLETE | 3 |\n\n' +
|
||||
'VERDICT: CEO CLEARED\n\nNO UNRESOLVED DECISIONS\n');
|
||||
// Native Write succeeded at this time, before the final handoff. The
|
||||
// live inode was cleaned up; this fixture replays that observed order.
|
||||
const reportAt = Date.parse('2026-09-08T21:01:38.295Z') / 1000;
|
||||
fs.utimesSync(file, reportAt, reportAt);
|
||||
const calls = structuredClone(CAPTURED_PAIRED_RETRY_CALLS);
|
||||
const transcript = {
|
||||
status: 'ready' as const,
|
||||
calls,
|
||||
assistantMessages: [],
|
||||
planReadyRequests: [{
|
||||
sessionId: calls[0]!.sessionId,
|
||||
toolUseId: 'toolu_01XK7amzoCx4VTm1r2bHdtsH',
|
||||
timestamp: '2026-09-08T21:03:46.725Z',
|
||||
failed: false,
|
||||
}],
|
||||
};
|
||||
const admin = new Set(calls.filter(call => isCeoCompletionHandoff(fingerprint(call)))
|
||||
.map(call => `${call.sessionId}:${call.toolUseId}`));
|
||||
const startedAt = Date.parse('2026-09-08T20:51:50Z');
|
||||
expect(admin.size).toBe(1);
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
// A new substantive answer after the Write remains a freshness boundary.
|
||||
calls.splice(-1, 0, { ...structuredClone(calls[0]!), toolUseId: 'later-substantive-fix',
|
||||
answeredAt: '2026-09-08T21:03:00.000Z' });
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('captured CEO next-step prefixes and immediate review menus', () => {
|
||||
test('next-step prefixes and a CLEAN declaration still identify only the completed handoff', () => {
|
||||
for (const scenario of currentHandoffs.cases) {
|
||||
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
|
||||
const before = structuredClone(call);
|
||||
expect(replay([call])).toMatchObject({ reviewCount: 0, administrativeCount: 1, step0Count: 0 });
|
||||
expect(call).toEqual(before);
|
||||
}
|
||||
});
|
||||
|
||||
test('the bound pending menu selects the offered manual action in either order', () => {
|
||||
for (const scenario of currentHandoffs.cases) for (const reverse of [false, true]) {
|
||||
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
|
||||
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
const q = call.questions[0]!;
|
||||
const active = `☐ ${q.header}\n${q.question}\n❯ 1. ${q.options[0]!.label}\n 2. ${q.options[1]!.label}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const bound = capturePlanCountQuestion(active, new Set(), 0, false, call)!;
|
||||
expect(bound.nativeCall?.toolUseId).toBe(call.toolUseId);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), bound)).toBe(reverse ? 1 : 2);
|
||||
expect(isCeoCompletionHandoff(bound)).toBe(false);
|
||||
const uiOnly = capturePlanCountQuestion(active, new Set(), 0, false)!;
|
||||
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('conditional completion, substantive actions, and mismatched identities still cannot authorize a handoff', () => {
|
||||
for (const scenario of currentHandoffs.cases) for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: If the CEO review is complete, should we run the next review? Eng review is the required shipping gate.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: The CEO review is not complete. Eng review is the required shipping gate.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Next steps: CEO review is CLEAN only after fixing this security gap. Eng review is the required shipping gate.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Security finding'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' and implement the fixes'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Add missing retry coverage to TODOS.md' }); },
|
||||
]) {
|
||||
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
|
||||
mutate(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(replay([call]).reviewCount).toBe(1);
|
||||
call.answered = false; delete call.answers;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
}
|
||||
for (const scenario of currentHandoffs.cases) {
|
||||
const call = structuredClone(scenario.nativeCall) as NativePlanQuestionCall;
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff({ ...fingerprint(call), signature: 'other-session:other-call' })).toBeNull();
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('native CEO completed handoffs with deferred implementation', () => {
|
||||
test('captured full sessions keep all substantive questions and classify only the final handoff', () => {
|
||||
for (const scenario of kHandoffs.cases) {
|
||||
const calls = structuredClone(scenario.calls) as NativePlanQuestionCall[];
|
||||
const original = structuredClone(calls);
|
||||
const result = replay(calls, false, ceoFirstReviewAUQ);
|
||||
expect(result).toMatchObject({ step0Count: scenario.expectedSetupCount,
|
||||
reviewCount: scenario.expectedReviewCount, administrativeCount: 1 });
|
||||
expect(result.classifications.slice(0, -1).every(p => !p.administrative)).toBe(true);
|
||||
expect(calls).toEqual(original);
|
||||
}
|
||||
});
|
||||
|
||||
test('active native handoffs choose manual in either order, never implementation or another review', () => {
|
||||
for (const scenario of kHandoffs.cases) for (const reverse of [false, true]) {
|
||||
const call = structuredClone(scenario.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
call.answered = false; delete call.answers; delete call.unansweredQuestionIndices;
|
||||
const q = call.questions[0]!;
|
||||
if (reverse) q.options.reverse();
|
||||
const options = q.options.map((option, i) => `${i === 0 ? '❯' : ' '} ${i + 1}. ${option.label}`).join('\n');
|
||||
const screen = `☐ ${q.header}\n${q.question}\n${options}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
const bound = capturePlanCountQuestion(screen, new Set(), 0, false, call)!;
|
||||
expect(bound.nativeCall?.toolUseId).toBe(call.toolUseId);
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call), bound)).toBe(q.options.findIndex(o => /handle.*manually/i.test(o.label)) + 1);
|
||||
expect(isCeoCompletionHandoff(bound)).toBe(false);
|
||||
const uiOnly = capturePlanCountQuestion(screen, new Set(), 0, false)!;
|
||||
expect(pickCeoCompletionHandoff(uiOnly)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('conditional declarations and new implementation obligations remain substantive', () => {
|
||||
for (const scenario of kHandoffs.cases) for (const question of [
|
||||
'ELI10: If the CEO review is done and the plan is cleared, choose the next step.',
|
||||
'ELI10: The CEO review is done only after resolving the test gap.',
|
||||
'ELI10: The CEO review is done and the plan is cleared after you add retry tests.',
|
||||
'ELI10: The CEO review is not done and the plan is not cleared.',
|
||||
]) {
|
||||
const call = structuredClone(scenario.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
call.questions[0]!.question = question + ' The required shipping gate is an Eng Review.';
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
for (const option of [
|
||||
{ label: 'Implement now, eng review later', description: 'Add the missing receipt test, then implement.' },
|
||||
{ label: 'Implement now, eng review later', description: 'Implement the approved tasks and add a new receipt assertion before the next review.' },
|
||||
{ label: 'Implement now, eng review later', description: 'The plan has no approved tasks; decide the missing error contract during implementation.' },
|
||||
{ label: 'Implement new retry behavior now, eng review later', description: 'The plan already has approved tasks.' },
|
||||
{ label: 'Add another TODO before implementing', description: 'Use the approved plan.' },
|
||||
]) {
|
||||
const call = structuredClone(kHandoffs.cases[0]!.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
call.questions[0]!.options[1] = option;
|
||||
expect(isCeoCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
call.answered = false;
|
||||
expect(pickCeoCompletionHandoff(fingerprint(call))).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('a real native approval after the completed report still requires all substantive answers in that report', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-k-handoff-'));
|
||||
const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
for (const scenario of kHandoffs.cases) {
|
||||
fs.writeFileSync(file, '# Plan\n\n## GSTACK REVIEW REPORT\n\n' +
|
||||
'| Review | Runs | Status | Findings |\n|---|---|---|---|\n| CEO | 1 | COMPLETE | 4 |\n\n' +
|
||||
'VERDICT: CEO CLEARED\n\nNO UNRESOLVED DECISIONS\n');
|
||||
fs.utimesSync(file, scenario.reportAtMs / 1000, scenario.reportAtMs / 1000);
|
||||
const calls = structuredClone(scenario.calls) as NativePlanQuestionCall[];
|
||||
const transcript = { status: 'ready' as const, calls, assistantMessages: [],
|
||||
planReadyRequests: structuredClone(scenario.planReadyRequests) };
|
||||
const admin = new Set(calls.filter(c => isCeoCompletionHandoff(fingerprint(c))).map(c => `${c.sessionId}:${c.toolUseId}`));
|
||||
const startedAt = Date.parse('2026-09-08T22:17:54Z');
|
||||
expect(admin.size).toBe(1);
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
calls.splice(-1, 0, { ...structuredClone(calls[2]!), toolUseId: 'new-substantive-answer',
|
||||
answeredAt: new Date(scenario.reportAtMs + 1000).toISOString() });
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready', admin)).toBe(false);
|
||||
}
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
});
|
||||
});
|
||||
@@ -1,89 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { createHash } from 'node:crypto';
|
||||
import fixture from './fixtures/ceo-conditional-option-facts-c6fc.json';
|
||||
import { createCeoPaymentFindingCounter, ceoPaymentFinding } from './helpers/ceo-payment-findings';
|
||||
import { ceoFirstReviewAUQ, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
|
||||
const originalCons = 'if the prior lookup helper is library-adapter-owned it may need a small extraction into app code.';
|
||||
const replaceOnce = (text: string, before: string, after: string) => {
|
||||
expect(text.split(before)).toHaveLength(2);
|
||||
return text.replace(before, after);
|
||||
};
|
||||
const cons = (plan: string, text: string) => replaceOnce(plan, `Cons: ${originalCons}`, `Cons: ${text}`);
|
||||
const fingerprint = (index: number) => {
|
||||
const capture = fixture.captures[index]!;
|
||||
return capture.fingerprint ? structuredClone(capture.fingerprint)
|
||||
: nativePlanCallFingerprint(structuredClone(capture.nativeCall), capture.observedAtMs, false);
|
||||
};
|
||||
type Fingerprint = ReturnType<typeof fingerprint>;
|
||||
function count(plan = fixture.captures[1]!.savedPlan, change?: (fp: Fingerprint) => void) {
|
||||
let saved = fixture.captures[0]!.savedPlan;
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => saved, ceoFirstReviewAUQ);
|
||||
const first = fingerprint(0), second = fingerprint(1);
|
||||
expect(counter.isReviewAUQ(first, [])).toBe(true);
|
||||
expect(counter.trace).toEqual([{signature: first.signature, kind: 'recorded-decision', ledgerId: 'D1', phase: 'currentDecision: D1'}]);
|
||||
saved = plan;
|
||||
change?.(second);
|
||||
const result = counter.isReviewAUQ(second, [first.nativeCall!]);
|
||||
return {result, trace: counter.trace, second};
|
||||
}
|
||||
|
||||
test('original c6fc D1 then D2: complete conditional option risk counts without a SQL synonym', () => {
|
||||
for (const capture of fixture.captures)
|
||||
expect(createHash('sha256').update(capture.savedPlan).digest('hex')).toBe(capture.savedSha256);
|
||||
const {result, trace, second} = count();
|
||||
expect(result).toBe(true);
|
||||
expect(ceoPaymentFinding(second, fixture.seed, fixture.captures[1]!.savedPlan)).toBeNull();
|
||||
expect(ceoFirstReviewAUQ(second)).toBe(false);
|
||||
expect(trace).toEqual([
|
||||
{signature: fingerprint(0).signature, kind: 'recorded-decision', ledgerId: 'D1', phase: 'currentDecision: D1'},
|
||||
{signature: second.signature, kind: 'recorded-decision', ledgerId: 'D2', phase: 'currentDecision: D2'},
|
||||
]);
|
||||
});
|
||||
|
||||
for (const text of [
|
||||
originalCons,
|
||||
'If the prior lookup helper remains library-adapter-owned, extracting it may cost extra work.',
|
||||
'unless the prior lookup helper is already application-owned, a small extraction into app code may be needed.',
|
||||
'a small extraction into app code may be needed if the prior lookup helper is library-adapter-owned.',
|
||||
'a small extraction into app code may be needed unless the prior lookup helper is already application-owned.',
|
||||
'when the prior lookup helper remains library-adapter-owned, a small extraction may be needed.',
|
||||
]) test(`a current option can state its conditional cost: ${text}`, () => {
|
||||
expect(count(cons(fixture.captures[1]!.savedPlan, text)).result).toBe(true);
|
||||
});
|
||||
|
||||
for (const [name, change] of Object.entries({
|
||||
'withdrawn option': (p: string) => cons(p, originalCons + ' This option is withdrawn.'),
|
||||
'resolved decision': (p: string) => cons(p, originalCons + ' This decision is resolved.'),
|
||||
'conditional clause cannot shelter withdrawal': (p: string) => cons(p, 'if the helper needs extraction, this option is no longer current.'),
|
||||
'quoted withdrawal remains active when explicitly attributed': (p: string) => cons(p, originalCons + ' This option is now "withdrawn".'),
|
||||
'historical fact': (p: string) => cons(p, 'Previously the helper needed extraction.'),
|
||||
'conditional historical fact': (p: string) => cons(p, 'if previously the helper needed extraction.'),
|
||||
'missing current comparison': (p: string) => p.slice(0, p.indexOf('## currentDecision: D2')),
|
||||
'wrong current comparison identity': (p: string) => replaceOnce(p, '## currentDecision: D2', '## currentDecision: OTHER'),
|
||||
'historical comparison': (p: string) => replaceOnce(p, '## currentDecision: D2', '## Historical currentDecision: D2'),
|
||||
'foreign source': (p: string) => p.replaceAll('PLAN.md', 'OTHER.md'),
|
||||
'missing cons field': (p: string) => replaceOnce(p, `Cons: ${originalCons}`, `Notes: ${originalCons}`),
|
||||
'missing risk field': (p: string) => replaceOnce(p, 'Risk low. Pros: injection impossible', 'Exposure low. Pros: injection impossible'),
|
||||
'duplicated effort field': (p: string) => replaceOnce(p, 'Risk low. Pros: injection impossible', 'Effort S. Risk low. Pros: injection impossible'),
|
||||
'conditional risk scalar': (p: string) => replaceOnce(p, 'Risk low. Pros: injection impossible', 'Risk if approved, low. Pros: injection impossible'),
|
||||
'conditional effort scalar': (p: string) => replaceOnce(p, 'Effort S (human ~1 hour / CC ~5 min). Risk low.', 'Effort if approved, S. Risk low.'),
|
||||
'conditional benefit claim': (p: string) => replaceOnce(p, 'Pros: injection impossible by construction;', 'Pros: if approved, injection impossible by construction;'),
|
||||
})) test(`a conditional cost cannot validate ${name}`, () => {
|
||||
expect(() => count(change(fixture.captures[1]!.savedPlan))).toThrow();
|
||||
});
|
||||
|
||||
for (const [name, change] of Object.entries({
|
||||
'failed ACK': (fp: Fingerprint) => { fp.nativeCall!.failed = true; },
|
||||
'missing ACK': (fp: Fingerprint) => { fp.nativeCall!.answered = false; fp.nativeCall!.answers = {}; },
|
||||
'foreign signature': (fp: Fingerprint) => { fp.signature = 'foreign:call'; },
|
||||
'unoffered selection': (fp: Fingerprint) => { fp.nativeCall!.answers = {[fp.nativeCall!.questions[0]!.question]: 'Other'}; },
|
||||
'foreign option contract': (fp: Fingerprint) => {
|
||||
const q = fp.nativeCall!.questions[0]!;
|
||||
q.options[0]!.label = 'A) Publish account credentials';
|
||||
q.options[0]!.description = 'Effort S, risk high. ✅ Easier access. ✅ Fewer prompts. ❌ Exposes accounts.';
|
||||
fp.options[0]!.label = q.options[0]!.label; fp.nativeCall!.answers = {[q.question]: q.options[0]!.label};
|
||||
},
|
||||
})) test(`current conditional costs preserve ${name} rejection`, () => {
|
||||
expect(() => count(fixture.captures[1]!.savedPlan, change)).toThrow();
|
||||
});
|
||||
@@ -1,181 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/ceo-contract-assertions-ag.json';
|
||||
import retry from './fixtures/ceo-contract-assertions-ag-retry.json';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
function reanswer(call: NativePlanQuestionCall) {
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
|
||||
test('actual declarative assertion defects start review after routing and approach', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0 };
|
||||
for (const call of calls()) {
|
||||
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
counts[phase.preReview ? 'setup' : 'review']++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 2, review: 2 });
|
||||
for (const call of calls().slice(2)) expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
|
||||
// Correct classification cannot retroactively complete the original paid run.
|
||||
expect(captured.observedOutcome).toBe('no_review_questions');
|
||||
expect(captured.observedReviewCount).toBe(0);
|
||||
});
|
||||
|
||||
test('assertion briefs still require completed native identity and their actual remedy', () => {
|
||||
for (const original of calls().slice(2)) {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Issue 99'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/Recommendation: \d[A-Z]/, 'Recommendation: 99Z'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.forEach((o, i) => { o.label = `${i + 1}A) Keep`; o.description = 'Keep the saved report.'; }); },
|
||||
]) {
|
||||
const call = structuredClone(original); mutate(call);
|
||||
if (call.answers && Object.keys(call.answers).length) reanswer(call);
|
||||
expect(ceoFirstReviewAUQ(fp(call))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ({ ...fp(original), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(original), options: [] })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('historical, hypothetical, quoted and withdrawn assertion problems are not current findings', () => {
|
||||
for (const original of calls().slice(2)) {
|
||||
for (const prefix of ['If ', 'Example: ', 'Whether ', 'Unless ']) {
|
||||
const call = structuredClone(original);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace(/(Issue \d+: )/, `$1${prefix}`);
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
for (const replacement of [
|
||||
'test 2 can detect all retry or backoff regressions',
|
||||
'test 2 previously could not detect retry or backoff regressions',
|
||||
'test 1 does not accept any truthy value as a correct receipt',
|
||||
'"test 2 cannot detect retry or backoff regressions"',
|
||||
]) {
|
||||
const call = structuredClone(original);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace(/(Issue \d+: )[^\n]+/, `$1${replacement}`);
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
const withdrawn = structuredClone(original);
|
||||
withdrawn.questions[0]!.question = withdrawn.questions[0]!.question.replace(/(ELI10:[^\n]+)/, '$1 No current defect exists.');
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(withdrawn)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the captured assertion regression selects the existing CEO count eval', () => {
|
||||
for (const file of ['test/ceo-contract-assertions-ag.test.ts', 'test/fixtures/ceo-contract-assertions-ag.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('actual retry contract wording recognizes its first repair and counts three review decisions', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0 };
|
||||
for (const call of structuredClone(retry.calls) as NativePlanQuestionCall[]) {
|
||||
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
counts[phase.preReview ? 'setup' : 'review']++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 2, review: 3 });
|
||||
for (const original of retry.calls.slice(2, 4)) {
|
||||
const call = structuredClone(original) as NativePlanQuestionCall;
|
||||
expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options.at(-1)!.label };
|
||||
expect(ceoFirstReviewAUQ(fp(call))).toBe(true);
|
||||
}
|
||||
expect(retry.observedOutcome).toBe('no_review_questions');
|
||||
expect(retry.observedReviewCount).toBe(0);
|
||||
expect(selectTests(['test/fixtures/ceo-contract-assertions-ag-retry.json'], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
});
|
||||
|
||||
test('already complete assertions and layout-only choices do not invent a defect', () => {
|
||||
const cases = [
|
||||
[2, 'D2 — Issue 1: test 2 cannot detect retry regressions (historical assessment)', 'The assertion gap was fixed yesterday. The current test pins the retry count and delay; this choice only arranges the already complete tests.'],
|
||||
[3, 'D3 — Issue 2: test 1 accepts any truthy value as specified by its success contract', 'The contract intentionally accepts every truthy success marker. The current assertion covers the contract completely; this choice only arranges the existing test.'],
|
||||
] as const;
|
||||
for (const [index, title, explanation] of cases) {
|
||||
const call = calls()[index]!;
|
||||
const q = call.questions[0]!;
|
||||
q.question = `${title}\nELI10: ${explanation}\nRecommendation: A`;
|
||||
q.options = [{ label: 'A) Use a table-driven layout', description: 'Use a table-driven layout for the existing assertions.' }, { label: 'B) Keep the existing layout', description: 'Keep the existing assertions in place.' }];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('retry assertion brief keeps native identity, exact contract and repair requirements', () => {
|
||||
for (const original of retry.calls.slice(2, 4)) {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding 99'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Example: ' + c.questions[0]!.question; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('but the contract is', 'but there is no contract for'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^ELI10:.*$/m, 'ELI10: The current assertion covers the contract completely.'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('Fix the assertion?', 'Save the report?'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.forEach(o => { o.label = o.label.replace(/\).*/, ') Use the existing layout'); o.description = 'Use the existing layout.'; }); },
|
||||
]) {
|
||||
const call = structuredClone(original) as NativePlanQuestionCall;
|
||||
mutate(call);
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('one offered option must repair the assertion rather than borrow report and layout actions', () => {
|
||||
for (const original of [...captured.calls.slice(2), ...retry.calls.slice(2, 4)]) {
|
||||
for (const administrative of ['Verify the saved report', 'Assert the full report', 'Pin the exact saved plan', 'Verify the expected layout']) {
|
||||
const call = structuredClone(original) as NativePlanQuestionCall;
|
||||
const q = call.questions[0]!;
|
||||
const prefix = /^([1-9]\d*)?[A-Z]/.exec(q.options[0]!.label)![1] ?? '';
|
||||
q.options = [
|
||||
{ label: `${prefix}A) ${administrative}`, description: administrative + '.' },
|
||||
{ label: `${prefix}B) Use a table-driven layout`, description: 'Use a table-driven layout for the existing assertions.' },
|
||||
];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('administrative report qualifiers cannot strengthen the unchanged assertion clause', () => {
|
||||
for (const suffix of [' and include a full report.', '; write an exact report.', '. Save the complete plan.']) {
|
||||
const call = calls()[2]!;
|
||||
const q = call.questions[0]!;
|
||||
q.options = [
|
||||
{ label: '1A) Assert the error class only', description: 'Assert the error class only' + suffix },
|
||||
{ label: '1B) Keep the current test', description: 'Leave the current rejection-only assertion unchanged.' },
|
||||
];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('each assertion clause owns its strong qualifier and actual assertion target', () => {
|
||||
for (const suffix of [' and verify the full report.', ' and check the full report.', ' with a full report.', ' with a complete saved plan.']) {
|
||||
const call = calls()[2]!;
|
||||
call.questions[0]!.options = [
|
||||
{ label: '1A) Assert the error class only', description: 'Assert the error class only' + suffix },
|
||||
{ label: '1B) Keep the current test', description: 'Leave the current rejection-only assertion unchanged.' },
|
||||
];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(false);
|
||||
}
|
||||
for (const description of ['Assert the rejection class and exactly two Stripe attempts.', 'Assert the error class only and assert exactly two Stripe attempts.']) {
|
||||
const call = calls()[2]!;
|
||||
call.questions[0]!.options[0]!.label = '1A) Strengthen the assertions';
|
||||
call.questions[0]!.options[0]!.description = description;
|
||||
call.questions[0]!.options = [call.questions[0]!.options[0]!, { label: '1B) Keep the current test', description: 'Leave the rejection-only assertion unchanged.' }];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(call)))).toBe(true);
|
||||
}
|
||||
});
|
||||
@@ -1,245 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import fixture from './fixtures/ceo-contract-question-an.json';
|
||||
import sectionFixture from './fixtures/ceo-section-finding-an.json';
|
||||
import contractFixture from './fixtures/ceo-current-contract-an.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
|
||||
const findings = calls.slice(2);
|
||||
const sectionCalls = sectionFixture.fingerprints as AskUserQuestionFingerprint[];
|
||||
const sectionFindings = sectionCalls.slice(4, 6);
|
||||
function change(fp: AskUserQuestionFingerprint, edit: (q: any, call: any, fp: any) => void) {
|
||||
const copy = structuredClone(fp), call = copy.nativeCall!, q = call.questions[0]!;
|
||||
const answerIndex = q.options.findIndex(o => o.label === call.answers?.[q.question]);
|
||||
edit(q, call, copy);
|
||||
call.answers = { [q.question]: q.options[answerIndex]?.label ?? '' };
|
||||
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return copy;
|
||||
}
|
||||
test('both actual completed contract questions start review; routing and test layout remain setup', () => {
|
||||
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, true, true]);
|
||||
});
|
||||
test('the decision ordinal, punctuation and form of the remedy question do not carry the finding', () => {
|
||||
for (const fp of findings) for (const title of [
|
||||
'd19 — Test 1 checks only truthiness; what should the exact assertion verify?',
|
||||
'D4 - Test 1 asserts only truthiness. How should the test check the full contract?',
|
||||
'D7 — Test 1 checks only truthiness: assert the contract or keep this check?',
|
||||
]) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace(q.question.split('\n')[0], title);
|
||||
q.header = 'Test contract';
|
||||
}))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('a competing test header or explicit foreign issue cannot borrow a test assertion', () => {
|
||||
for (const header of ['Test 99 assert', 'Finding 3', 'Issue 1'])
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.header = header; }))).toBe(false);
|
||||
});
|
||||
test('a title alone or an administrative response does not establish a review finding', () => {
|
||||
for (const fp of findings) {
|
||||
expect(ceoFirstReviewAUQ({ ...fp, nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.options = [
|
||||
{ label: 'A) Keep the current assertion (recommended)', description: 'Leave the test unchanged.' },
|
||||
{ label: 'B) Archive the review', description: 'Save the existing report without changing tests.' },
|
||||
];
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.header = 'Approach'; }))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('the full native identity, selected answer and completed result remain required', () => {
|
||||
for (const fp of findings) {
|
||||
for (const edit of [
|
||||
(_q: any, c: any) => { c.answered = false; },
|
||||
(_q: any, c: any) => { c.failed = true; },
|
||||
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(_q: any, _c: any, f: any) => { f.signature = 'foreign:tool'; },
|
||||
(q: any) => { q.multiSelect = true; },
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
|
||||
const answer = change(fp, () => {}); answer.nativeCall!.answers = {};
|
||||
expect(ceoFirstReviewAUQ(answer)).toBe(false);
|
||||
const menu = change(fp, () => {}); menu.options[0]!.label = 'Foreign selection';
|
||||
expect(ceoFirstReviewAUQ(menu)).toBe(false);
|
||||
}
|
||||
});
|
||||
test('source and conditional frames cannot own the current assertion assessment', () => {
|
||||
for (const intro of ['Source:', 'Example:', 'Earlier review assessment:', 'The following assessment is hypothetical.'])
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('\nELI10:', '\n' + intro + '\nELI10:'); }))).toBe(false);
|
||||
for (const intro of ['Source excerpt: ', 'Previously, ', 'If approved, ', 'The following is a hypothetical example. '])
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + intro); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: If approved, '); }))).toBe(false);
|
||||
});
|
||||
test('literal titles and withdrawn current findings supply no first-review credit', () => {
|
||||
for (const fp of findings) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { const lines = q.question.split('\n'); lines[0] = '`' + lines[0] + '`'; q.question = lines.join('\n'); }))).toBe(false);
|
||||
for (const statement of [
|
||||
'Correction: this finding is withdrawn.',
|
||||
'Correction: this finding is "withdrawn".',
|
||||
'Correction: this explanation is not current.',
|
||||
'There is no current gap.',
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + statement; }))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('quoted historical notes cannot withdraw the current finding', () => {
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => {
|
||||
q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:');
|
||||
}))).toBe(true);
|
||||
});
|
||||
test('uniform recommendation and option identities remain required', () => {
|
||||
for (const edit of [
|
||||
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: Z'); },
|
||||
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', '9B)'); },
|
||||
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', 'A)'); },
|
||||
]) expect(ceoFirstReviewAUQ(change(findings[0]!, edit))).toBe(false);
|
||||
});
|
||||
test('the new regression inputs belong only to the dense CEO finding owner', () => {
|
||||
for (const name of ['test/ceo-contract-question-an.test.ts', 'test/fixtures/ceo-contract-question-an.json', 'test/fixtures/ceo-section-finding-an.json', 'test/fixtures/ceo-current-contract-an.json'])
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(name)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
|
||||
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
|
||||
for (let i = 0; i < paths.length; i++) {
|
||||
expect(Object.hasOwn(paths, i)).toBe(true);
|
||||
expect(typeof paths[i]).toBe('string');
|
||||
}
|
||||
});
|
||||
test('owned Section finding briefs establish review through their current defect and remedy', () => {
|
||||
expect(sectionCalls.map(ceoFirstReviewAUQ)).toEqual([false, false, false, false, true, true, false]);
|
||||
for (const fp of sectionFindings) for (const separator of [':', '—', '-']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace(/^D\d+ — Section 2 finding (\d):/, `d19 — Section 7 finding $1 ${separator}`);
|
||||
q.header = 'Section 7';
|
||||
}))).toBe(true);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ(change(sectionFindings[0]!, q => {
|
||||
q.question = q.question.replace('the lookup reads request.params.userId into a raw SQL fragment', 'the query reads payload.accountId into a raw SQL string');
|
||||
}))).toBe(true);
|
||||
for (const term of ['“no error handling”', "'no error handling'", 'no error handling'])
|
||||
expect(ceoFirstReviewAUQ(change(sectionFindings[1]!, q => {
|
||||
q.question = q.question.replace('"no error handling"', term);
|
||||
}))).toBe(true);
|
||||
});
|
||||
test('Section dispatch requires an exact completed native question and consistent finding identity', () => {
|
||||
for (const fp of sectionFindings) for (const edit of [
|
||||
(_q: any, c: any) => { delete c.answeredAt; },
|
||||
(_q: any, c: any) => { c.answeredAt = 'not-a-time'; },
|
||||
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(_q: any, c: any) => { c.answered = false; },
|
||||
(_q: any, c: any) => { c.failed = true; },
|
||||
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; },
|
||||
(q: any) => { q.header = 'Section 8'; },
|
||||
(q: any) => { q.header = 'Finding 99'; },
|
||||
(q: any) => { q.header = 'Section 2 finding 99'; },
|
||||
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: 99A'); },
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
|
||||
});
|
||||
test('Section declarations cannot borrow source, historical, conditional or negated defects', () => {
|
||||
for (const fp of sectionFindings) for (const prefix of ['Source: ', 'Previously, ', 'If approved, ', 'The hypothetical example: ', 'Earlier review assessment: ', 'For historical context, ']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/(Section 2 finding \d: )/, '$1' + prefix); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
|
||||
}
|
||||
for (const fp of sectionFindings) for (const prefix of ['Source:', 'Earlier review assessment:', 'The following is a hypothetical example.'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(sectionFindings[0]!, q => {
|
||||
q.question = q.question.replace('the lookup reads', 'the lookup no longer reads');
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(sectionFindings[1]!, q => {
|
||||
q.question = q.question.replace('the receipt email has "no error handling"', 'the receipt email no longer has "no error handling"');
|
||||
}))).toBe(false);
|
||||
});
|
||||
test('Section review requires a current offered amendment and an unwithdrawn assessment', () => {
|
||||
for (const fp of sectionFindings) {
|
||||
for (const status of ['This finding is withdrawn.', 'Correction: this finding is "withdrawn".', 'This explanation is not current.', 'There is no current gap.'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + status; }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.options = [
|
||||
{ label: 'A) Keep the existing implementation', description: 'Leave all behavior unchanged.' },
|
||||
{ label: 'B) Archive the report', description: 'Export the report.' },
|
||||
];
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
for (const option of q.options) option.description = 'Source excerpt: ' + option.description;
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
for (const option of q.options) option.description += '\nThis amendment is withdrawn.';
|
||||
}))).toBe(false);
|
||||
for (const status of [' This amendment is withdrawn.', ' This remedy is a historical example, not the current option.'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
for (const option of q.options) option.description += status;
|
||||
}))).toBe(false);
|
||||
for (const prefix of ['Source excerpt: ', 'If approved later: '])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
for (const option of q.options) option.label = option.label.replace(/^([A-C]\)) /, '$1 ' + prefix);
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:');
|
||||
}))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('the current plan contract can establish the gap in a later ELI10 sentence', () => {
|
||||
const fp = contractFixture.fingerprints[2] as AskUserQuestionFingerprint;
|
||||
expect(ceoFirstReviewAUQ(fp)).toBe(true);
|
||||
for (const clause of [
|
||||
"The current plan states 'no error handling on the email leg'.",
|
||||
'The plan specifies “no error handling on the email leg”.',
|
||||
'This plan requires "no error handling on the email leg".',
|
||||
'The plan says no error handling on the email leg.',
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace("The plan says 'no error handling on the email leg'.", clause);
|
||||
}))).toBe(true);
|
||||
});
|
||||
test('later contract declarations retain source, currentness and remedy ownership', () => {
|
||||
const fp = contractFixture.fingerprints[2] as AskUserQuestionFingerprint;
|
||||
for (const clause of [
|
||||
"The old plan said 'no error handling on the email leg'.",
|
||||
"If approved, the plan says 'no error handling on the email leg'.",
|
||||
"Source excerpt: the plan says 'no error handling on the email leg'.",
|
||||
'"The plan says no error handling on the email leg."',
|
||||
"The plan no longer says 'no error handling on the email leg'.",
|
||||
"The plan says 'no error handling on the email leg' only in a historical example.",
|
||||
"The plan says 'no error handling on the email leg”.",
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question = q.question.replace("The plan says 'no error handling on the email leg'.", clause);
|
||||
}))).toBe(false);
|
||||
for (const edit of [
|
||||
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: Earlier review assessment: '); },
|
||||
(q: any) => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: Source excerpt: '); },
|
||||
(q: any) => { q.question += '\nThis finding is "withdrawn".'; },
|
||||
(q: any) => { for (const o of q.options) o.description += ' This amendment is withdrawn.'; },
|
||||
(q: any) => { for (const o of q.options) o.description = 'Source excerpt: ' + o.description; },
|
||||
(_q: any, c: any) => { delete c.answeredAt; },
|
||||
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(q: any) => { q.header = 'Finding 99'; },
|
||||
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "Source excerpt follows. The plan says 'no error handling"); },
|
||||
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "Earlier review assessment follows. The plan says 'no error handling"); },
|
||||
(q: any) => { q.question = q.question.replace("The plan says 'no error handling", "If approved later. The plan says 'no error handling"); },
|
||||
(q: any) => { q.question = q.question.replace("'no error handling on the email leg'.", "'no error handling on the email leg'. This no-error-handling contract is withdrawn."); },
|
||||
(q: any) => { q.question = q.question.replace("'no error handling on the email leg'.", "'no error handling on the email leg'. This contract is a historical example, not the current plan."); },
|
||||
(q: any) => { q.question += '\nThis finding is "resolved".'; },
|
||||
(q: any) => { for (const o of q.options) o.description += '\nThis amendment is "closed".'; },
|
||||
(q: any) => { for (const o of q.options) o.description += ' This amendment is "closed".'; },
|
||||
]) expect(ceoFirstReviewAUQ(change(fp, edit))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => {
|
||||
q.question += '\nArchive note: "This finding is withdrawn."';
|
||||
}))).toBe(true);
|
||||
});
|
||||
test('the assertion assessment and strengthening action retain their own current authority', () => {
|
||||
for (const edit of [
|
||||
(_q: any, c: any) => { delete c.answeredAt; },
|
||||
(_q: any, c: any) => { c.answeredAt = 'invalid'; },
|
||||
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(q: any) => { q.question = q.question.replace('ELI10: The plan states', 'ELI10: The historical plan stated'); },
|
||||
(q: any) => { q.question = q.question.replace('But the planned test only checks', 'But the planned test no longer only checks'); },
|
||||
(q: any) => { q.options[0].label = q.options[0].label.replace('Assert deep equality with', 'Assert truthiness for'); },
|
||||
(q: any) => { q.options[0].description = 'Source excerpt:\n' + q.options[0].description; },
|
||||
(q: any) => { q.options[0].description = 'Earlier review assessment:\n' + q.options[0].description; },
|
||||
(q: any) => { q.options[0].description += '\nThis amendment is withdrawn.'; },
|
||||
(q: any) => { q.options[0].description += '\nThis amendment is "withdrawn".'; },
|
||||
(q: any) => { q.options[0].description += '\nThis amendment is “withdrawn”.'; },
|
||||
(q: any) => { q.options[0].description += '\nThis remedy is a historical example, not the current option.'; },
|
||||
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: Earlier review assessment: '); },
|
||||
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: For historical context, '); },
|
||||
]) expect(ceoFirstReviewAUQ(change(findings[0]!, edit))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(findings[0]!, q => {
|
||||
q.options[1] = { label: 'B) Export documentation', description: 'Export the report.' };
|
||||
}))).toBe(true);
|
||||
});
|
||||
@@ -1,423 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/ceo-count-ac-calls.json';
|
||||
import later from './fixtures/ceo-count-ac-later-calls.json';
|
||||
import alias from './fixtures/ceo-finding-alias-af.json';
|
||||
import numberedBrief from './fixtures/ceo-numbered-brief-af.json';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
|
||||
const finding = () => calls()[2]!;
|
||||
const handoff = () => calls()[3]!;
|
||||
function reanswer(c: NativePlanQuestionCall) {
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
return c;
|
||||
}
|
||||
function pending(c = handoff()) {
|
||||
c.answered = false; delete c.answers; delete c.answeredAt;
|
||||
c.unansweredQuestionIndices = [0]; return c;
|
||||
}
|
||||
|
||||
test('the actual paired attempt has one finding and remains below its two-finding floor', () => {
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const c of calls()) {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ,
|
||||
undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
counts[phase.administrative ? 'administrative' : phase.preReview ? 'setup' : 'review']++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 2, review: 1, administrative: 1 });
|
||||
expect(counts.review).toBeLessThan(2);
|
||||
expect(calls()[1]!.answers).toEqual(captured.calls[1]!.answers);
|
||||
});
|
||||
|
||||
test('qidless explicit Findings need a completed matching native decision', () => {
|
||||
expect(ceoFirstReviewAUQ(fp(finding()))).toBe(true);
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered answer' }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
]) {
|
||||
const c = finding(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ({ ...fp(finding()), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(finding()), nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(finding()), options: [] })).toBe(false);
|
||||
});
|
||||
|
||||
test('setup recaps, quoted titles and foreign qids cannot start a review', () => {
|
||||
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
|
||||
const c = finding(); c.questions[0]!.question = prefix + c.questions[0]!.question;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
for (const header of ['Approach', 'Mode', 'Next review', 'Setup']) {
|
||||
const c = finding(); c.questions[0]!.header = header;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
for (const id of ['plan-eng-review-finding', 'plan-ceo-review-mode', 'broken']) {
|
||||
const c = finding(); c.questions[0]!.question += ` <gstack-qid:${id}>`;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ(fp(calls()[1]!))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(fp(handoff()))).toBe(false);
|
||||
});
|
||||
|
||||
test('the exact administrative menu chooses manual without awarding completion coverage', () => {
|
||||
expect(isCeoCompletionHandoff(fp(handoff()))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(fp(pending()))).toBe(2);
|
||||
const c = pending(); c.questions[0]!.options.reverse();
|
||||
expect(pickCeoCompletionHandoff(fp(c))).toBe(1);
|
||||
expect(isCeoCompletionHandoff(fp(c))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(fp(handoff()))).toBeNull();
|
||||
});
|
||||
|
||||
test('appended obligations and altered navigation context remain substantive', () => {
|
||||
for (const extra of [' Also add another test.', ' Fix the missing auth check.',
|
||||
' Once the outstanding gap is resolved.', ' Decide whether to add retry support?',
|
||||
' The CEO review is not complete.']) {
|
||||
for (const target of ['question', 'run', 'manual']) {
|
||||
const c = handoff(), q = c.questions[0]!;
|
||||
if (target === 'question') q.question += extra;
|
||||
else q.options[target === 'run' ? 0 : 1]!.description += extra;
|
||||
expect(isCeoCompletionHandoff(fp(reanswer(c)))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(fp(pending(c)))).toBeNull();
|
||||
}
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('complete and clean', 'not complete'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Add another test'; },
|
||||
]) { const c = handoff(); mutate(c); expect(isCeoCompletionHandoff(fp(reanswer(c)))).toBe(false); }
|
||||
expect(pickCeoCompletionHandoff({ ...fp(pending()), signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoCompletionHandoff({ ...fp(pending()), options: [] })).toBeNull();
|
||||
});
|
||||
|
||||
test('the paired transcript regression remains selected from both new files', () => {
|
||||
for (const file of ['test/ceo-count-ac.test.ts', 'test/fixtures/ceo-count-ac-calls.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
test('later actual calls count explicit Issue and sectioned Finding titles without crediting a terminal', () => {
|
||||
for (const [key, expected] of [['distinct', { setup: 4, review: 5 }], ['pairedRetry', { setup: 4, review: 4 }]] as const) {
|
||||
let started = false;
|
||||
const count = { setup: 0, review: 0 };
|
||||
for (const c of structuredClone(later[key].nativeCalls) as NativePlanQuestionCall[]) {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
count[phase.preReview ? 'setup' : 'review']++;
|
||||
}
|
||||
expect(count).toEqual(expected);
|
||||
}
|
||||
// These attempts were stalled on file permission; count correction supplies
|
||||
// no terminal, written report or complete methodology evidence.
|
||||
expect(later.distinct.observedOutcome).toBe('timeout');
|
||||
expect(later.pairedRetry.observedOutcome).toBe('running');
|
||||
});
|
||||
|
||||
test('numbered Issue/sectioned Finding titles must agree with their native header', () => {
|
||||
for (const source of [later.distinct.nativeCalls[4]!, later.pairedRetry.nativeCalls[4]!]) {
|
||||
const c = structuredClone(source) as NativePlanQuestionCall;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
for (const header of ['Issue 7.2', 'Finding 9', 'Mode', 'Next review']) {
|
||||
c.questions[0]!.header = header;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
expect(selectTests(['test/fixtures/ceo-count-ac-later-calls.json'], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
});
|
||||
|
||||
function remedyCall(header: string, title: string, qid?: string) {
|
||||
const c = finding();
|
||||
c.questions[0]!.header = header;
|
||||
c.questions[0]!.question = title + (qid ? `\n<gstack-qid:${qid}>` : '');
|
||||
c.questions[0]!.options = [{ label: 'Repair the plan' }, { label: 'Keep the plan' }];
|
||||
return reanswer(c);
|
||||
}
|
||||
|
||||
function assertionCall(qid?: string) {
|
||||
const c = remedyCall('Receipt shape', 'D2 — Test 1 asserts only that the receipt is truthy, but the plan states the exact receipt contract. Pin the full receipt?', qid);
|
||||
c.questions[0]!.options = [
|
||||
{ label: 'A) Assert the exact receipt', description: 'Deep equality against the complete stated receipt.' },
|
||||
{ label: 'B) Keep truthy-only assertion', description: 'Leave the weaker planned assertion unchanged.' },
|
||||
];
|
||||
return reanswer(c);
|
||||
}
|
||||
|
||||
test('an explicit exact-contract assertion gap does not depend on a Finding header or question tuning', () => {
|
||||
for (const qid of [undefined, 'plan-ceo-review-receipt-contract']) {
|
||||
const c = assertionCall(qid);
|
||||
for (const option of c.questions[0]!.options) {
|
||||
c.answers = { [c.questions[0]!.question]: option.label };
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('assertion-gap evidence needs a direct contract mismatch and opposed assertion choices', () => {
|
||||
for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => s.replace('Test 1 asserts', 'If Test 1 asserts'),
|
||||
(s: string) => s.replace('the exact receipt contract', 'no required receipt shape'),
|
||||
(s: string) => s.replace('the exact receipt contract', 'the exact receipt contract is already covered'),
|
||||
]) {
|
||||
const c = assertionCall(); c.questions[0]!.question = change(c.questions[0]!.question);
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Skip this review'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
]) { const c = assertionCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
});
|
||||
|
||||
test('completed native remedy headers and numbered Issue titles start CEO review', () => {
|
||||
for (const c of [
|
||||
remedyCall('F1 remedy', 'D2 — Test 1: assert the full receipt, or keep the truthy-only assertion?'),
|
||||
remedyCall('F2 remedy', 'D3 — Test 2: assert attempt count and backoff, or only the rejection?'),
|
||||
remedyCall('Email leg', 'D4 — Issue 1: where does the notification run relative to commit?', 'plan-ceo-review-email-leg'),
|
||||
]) {
|
||||
for (const option of c.questions[0]!.options) {
|
||||
c.answers = { [c.questions[0]!.question]: option.label };
|
||||
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
|
||||
.toEqual({ preReview: false, reviewStarted: true });
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('a remedy header requires a matching completed decision and consistent finding identity', () => {
|
||||
const source = remedyCall('F1 remedy', 'D2 — Assert the complete receipt?');
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
]) {
|
||||
const c = structuredClone(source); mutate(c);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ({ ...fp(source), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(source), nativeCall: undefined })).toBe(false);
|
||||
for (const title of ['D2 — Issue 2: Assert the receipt?', 'D2 — Issue 0: Assert the receipt?',
|
||||
'D2 — Issue 1.0: Assert the receipt?', 'Example: D2 — Assert the receipt?',
|
||||
'> D2 — Assert the receipt?', '```\nD2 — Assert the receipt?']) {
|
||||
expect(ceoFirstReviewAUQ(fp(remedyCall('F1 remedy', title)))).toBe(false);
|
||||
}
|
||||
for (const header of ['Approach', 'F1', 'Remedy', 'F0 remedy', 'Next review']) {
|
||||
expect(ceoFirstReviewAUQ(fp(remedyCall(header, 'D2 — Assert the receipt?')))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('numbered Issue titles cannot bypass setup, provider or native-answer checks', () => {
|
||||
const title = 'D4 — Issue 1: where does the notification run relative to commit?';
|
||||
for (const qid of ['plan-ceo-review-scope', 'plan-ceo-review-next-steps', 'plan-eng-review-email', 'foreign']) {
|
||||
expect(ceoFirstReviewAUQ(fp(remedyCall('Email leg', title, qid)))).toBe(false);
|
||||
}
|
||||
for (const header of ['Setup', 'Approach', 'Mode', 'Next steps', 'Issue 2']) {
|
||||
expect(ceoFirstReviewAUQ(fp(remedyCall(header, title, 'plan-ceo-review-email')))).toBe(false);
|
||||
}
|
||||
for (const suffix of ['<gstack-qid:plan-ceo-review-email', '<gstack-qid:plan-ceo-review-email:foreign>',
|
||||
'<gstack-qid:plan-ceo-review-email> <gstack-qid:plan-eng-review-email>']) {
|
||||
const c = remedyCall('Email leg', title + '\n' + suffix);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ ...c.questions[0]!.options[0]! }); },
|
||||
]) {
|
||||
const c = remedyCall('Email leg', title, 'plan-ceo-review-email'); mutate(c);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
const aliasCalls = () => alias.rows.map(row => structuredClone(row.call) as NativePlanQuestionCall);
|
||||
|
||||
test('AF exact native Finding headers and same-number Issue titles start review', () => {
|
||||
for (const c of aliasCalls()) {
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
|
||||
.toEqual({ preReview: false, reviewStarted: true });
|
||||
}
|
||||
expect(alias.provenance.partial).toBe(true);
|
||||
expect(alias.provenance.paidCoverageCredit).toBe(false);
|
||||
});
|
||||
|
||||
test('AF Issue and Finding aliases compare the complete native number, not the decision counter', () => {
|
||||
for (const titleKind of ['Issue', 'Finding']) for (const headerKind of ['Issue', 'Finding']) {
|
||||
for (const number of ['1', '2.1', '27.3']) {
|
||||
const c = aliasCalls()[0]!, q = c.questions[0]!;
|
||||
q.question = q.question.replace(/ <gstack-qid:[^>]+>/, '').replace(/^D4 — Issue 1:/, `D87 — ${titleKind} ${number}:`);
|
||||
q.header = `${headerKind} ${number}`;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
|
||||
for (const wrong of ['9', `${number}.2`]) {
|
||||
q.header = `${headerKind} ${wrong}`;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('AF aliases preserve section and parenthesized issue header requirements', () => {
|
||||
for (const title of ['D87 — Issue 2.1 (Section 4): Which assertion should be used?',
|
||||
'D87 (issue 2.1) — Which assertion should be used?']) {
|
||||
const c = aliasCalls()[0]!; c.questions[0]!.question = title;
|
||||
for (const header of ['Issue 2.1', 'Finding 2.1']) {
|
||||
c.questions[0]!.header = header;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
|
||||
}
|
||||
for (const header of ['Receipt assertion', 'Issue 2', 'Finding 2.2']) {
|
||||
c.questions[0]!.header = header;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('AF aliases retain native completion, answer, qid and setup boundaries', () => {
|
||||
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
|
||||
c => { c.answered = false; }, c => { c.failed = true; },
|
||||
c => { c.answers = {}; }, c => { c.unansweredQuestionIndices = [0]; },
|
||||
c => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
|
||||
c => { c.questions[0]!.multiSelect = true; },
|
||||
c => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
c => { c.questions[0]!.options[1]!.label = c.questions[0]!.options[0]!.label; },
|
||||
];
|
||||
for (const mutate of mutations) for (const c of aliasCalls()) {
|
||||
mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
for (const c of aliasCalls()) {
|
||||
expect(ceoFirstReviewAUQ({ ...fp(c), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(c), nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(c), options: [] })).toBe(false);
|
||||
for (const header of ['Issue 9', 'Finding 9', 'Setup', 'Approach', 'Mode', 'Next steps']) {
|
||||
const changed = structuredClone(c); changed.questions[0]!.header = header;
|
||||
expect(ceoFirstReviewAUQ(fp(changed))).toBe(false);
|
||||
}
|
||||
for (const qid of ['plan-ceo-review-setup', 'plan-eng-review-finding']) {
|
||||
const changed = structuredClone(c); changed.questions[0]!.question = changed.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, `<gstack-qid:${qid}>`);
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(changed)))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
|
||||
const changed = structuredClone(c); changed.questions[0]!.question = prefix + changed.questions[0]!.question;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(changed)))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('AF alias evidence remains registered only to the CEO count workflow', () => {
|
||||
const file = 'test/fixtures/ceo-finding-alias-af.json';
|
||||
const owners = Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name);
|
||||
expect(owners).toEqual(['plan-ceo-finding-count']);
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
});
|
||||
|
||||
|
||||
test('AF complete numbered native briefs identify the three remaining first decisions', () => {
|
||||
for (const row of numberedBrief.rows) {
|
||||
expect(ceoFirstReviewAUQ(fp(structuredClone(row.call) as NativePlanQuestionCall))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('AF complete finding identities permit F notation but never contradict the native header', () => {
|
||||
for (const title of ['D7 — Finding F2.1: Which implementation should be used?', 'D7 — Issue 2.1: Which implementation should be used?']) {
|
||||
const c = structuredClone(numberedBrief.rows[1]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.question = title;
|
||||
for (const header of ['F2.1 remedy', 'Issue F2.1', 'Finding 2.1']) {
|
||||
c.questions[0]!.header = header; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(true);
|
||||
}
|
||||
for (const header of ['F2 remedy', 'Finding 2.1.1', 'Issue F2.1.0', 'F2.1 and F3']) {
|
||||
c.questions[0]!.header = header; expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
const c = aliasCalls()[0]!; c.questions[0]!.question = c.questions[0]!.question.replace('Issue 1:', 'Finding F1:');
|
||||
c.questions[0]!.header = 'Finding 2'; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
});
|
||||
|
||||
test('AF a declarative numbered brief needs current problem evidence and an actual amendment choice', () => {
|
||||
for (const body of [
|
||||
'ELI10: The handler has error handling.\nRecommendation: A because it is ready.',
|
||||
'ELI10: The handler has no current defect.\nRecommendation: A because it is ready.',
|
||||
'ELI10: If the handler has no error handling, we would repair it.\nRecommendation: A because this is a hypothetical.',
|
||||
'ELI10: Example: the handler has no error handling.\nRecommendation: A because this is an example.',
|
||||
'ELI10: "The handler has no error handling."\nRecommendation: A because this quotes the old plan.',
|
||||
'ELI10: The error contract is not missing.\nRecommendation: A because it is ready.',
|
||||
'ELI10: The email failure is no longer unhandled.\nRecommendation: A because it is ready.',
|
||||
]) {
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.question = 'D9 — 1.1 Email leg: transaction boundary and failure handling\n' + body;
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
for (const labels of [['Start review', 'Pause'], ['Write the completed report', 'Save the reviewed plan']]) {
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.options = labels.map((label,i) => ({label:`${i ? 'B' : 'A'}: ${label}`, description:label}));
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('AF new brief form preserves setup, native answer, quotation and subject binding', () => {
|
||||
for (const row of numberedBrief.rows) for (const mutation of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Next steps'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
]) {
|
||||
const c = structuredClone(row.call) as NativePlanQuestionCall; mutation(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['Example: ', '> ', '"', '```\n']) {
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.question = prefix + c.questions[0]!.question; expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
}
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.header = 'SQL lookup'; expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes('test/fixtures/ceo-numbered-brief-af.json')).map(([name])=>name))
|
||||
.toEqual(['plan-ceo-finding-count']);
|
||||
});
|
||||
|
||||
test('AF a resolved historical gap and completed-review log check cannot start current review', () => {
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.header = 'Finding 1';
|
||||
c.questions[0]!.question = 'D4 — Issue 1: Validation was missing in the prior review.\nELI10: The old gap is already resolved. Current validation is complete; this choice only checks the completed review log.\nRecommendation: A because it checks the record.';
|
||||
c.questions[0]!.options = [
|
||||
{label:'A) Check the prior review log',description:'Check the prior review log.'},
|
||||
{label:'B) Keep current report',description:'Keep the current completed report.'},
|
||||
];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(false);
|
||||
});
|
||||
|
||||
test('AF currentness uses the whole explanation and saved-log actions remain administrative', () => {
|
||||
for (const [subject, explanation, action, expected] of [
|
||||
['Historical missing validation', 'Validation is complete. This task only verifies the stored review log; there is no current defect.', 'Validate the saved review log', false],
|
||||
['Required validation is missing', 'The required validation is missing.', 'Validate the saved review log', false],
|
||||
['Historical missing validation', 'The prior review omitted a note; there is no current defect.', 'Validate the incoming request', false],
|
||||
['Required validation is missing', 'A previous log says "there is no current defect." The current plan still lacks validation.', 'Validate the incoming request', true],
|
||||
] as const) {
|
||||
const c = structuredClone(numberedBrief.rows[0]!.call) as NativePlanQuestionCall;
|
||||
c.questions[0]!.header = 'Issue 1';
|
||||
c.questions[0]!.question = `D1 — Issue 1: ${subject}\nELI10: ${explanation}\nRecommendation: A because it addresses this decision.`;
|
||||
c.questions[0]!.options = [
|
||||
{label:`A) ${action}`, description:`${action}.`},
|
||||
{label:'B) Keep the current report', description:'Leave the stored report unchanged.'},
|
||||
];
|
||||
expect(ceoFirstReviewAUQ(fp(reanswer(c)))).toBe(expected);
|
||||
}
|
||||
});
|
||||
@@ -1,125 +1,7 @@
|
||||
import {expect,test} from 'bun:test';
|
||||
import fs from 'node:fs';import os from 'node:os';import path from 'node:path';
|
||||
import fixture from './fixtures/ceo-count-ad-v2.json';
|
||||
import {readPlanCountTranscript,type NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
|
||||
import {isCeoCompletionHandoff,pickCeoCompletionHandoff} from './helpers/ceo-completion-handoff';
|
||||
import {selectTests, E2E_TOUCHFILES} from './helpers/touchfiles';
|
||||
const fp=(c:NativePlanQuestionCall)=>nativePlanCallFingerprint(c,0,true);
|
||||
const get=(which:'distinct'|'paired'|'pairedRetry',index:number)=>structuredClone(fixture.cases[which].calls[index]) as NativePlanQuestionCall;
|
||||
const realFindings=()=>[get('distinct',4),get('paired',4),get('paired',5),get('pairedRetry',2),get('pairedRetry',3)];
|
||||
const answer=(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};return c;};
|
||||
function count(calls:NativePlanQuestionCall[]){let started=false;const n={setup:0,review:0,administrative:0};for(const c of calls){const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=p.reviewStarted;n[p.administrative?'administrative':p.preReview?'setup':'review']++;}return n;}
|
||||
|
||||
import {readPlanCountTranscript} from './helpers/plan-count-transcript';
|
||||
test('exact public native requests and successful replies reconstruct the captured calls once',()=>{
|
||||
for(const which of ['distinct','paired','pairedRetry'] as const){const c=fixture.cases[which],dir=fs.mkdtempSync(path.join(os.tmpdir(),'ceo-count-public-'));try{const project=path.join(dir,'projects','owned');fs.mkdirSync(project,{recursive:true});const records=c.nativeRecords.map(r=>JSON.stringify(r)).join('\n')+'\n';fs.writeFileSync(path.join(project,c.calls[0]!.sessionId+'.jsonl'),records+records);expect(readPlanCountTranscript(dir,c.observation.capture.cwd).calls).toEqual(c.calls);for(const a of c.timeAnchors){expect(Date.parse(a.requestAt)).toBeLessThanOrEqual(Date.parse(a.replyAt));expect(Date.parse(a.replyAt)).toBeLessThanOrEqual(Date.parse(c.observation.capture.at));}}finally{fs.rmSync(dir,{recursive:true,force:true});}}
|
||||
});
|
||||
for(const [which,index] of [['distinct',4],['paired',4],['paired',5]] as const)test(`actual ${which} issue ${index} starts review from a completed native decision`,()=>expect(ceoFirstReviewAUQ(fp(get(which,index)))).toBe(true));
|
||||
test('exact snapshots keep real issue counts and separate the administrative handoff',()=>{
|
||||
expect(count(fixture.cases.distinct.calls as NativePlanQuestionCall[])).toEqual({setup:4,review:1,administrative:0});
|
||||
expect(count(fixture.cases.paired.calls as NativePlanQuestionCall[])).toEqual({setup:4,review:2,administrative:1});
|
||||
expect(fixture.cases.distinct.observation.state).toBe('in_progress');expect(fixture.cases.paired.observation.state).toBe('in_progress');
|
||||
expect(count(fixture.cases.distinct.calls as NativePlanQuestionCall[]).review).toBeLessThan(4);
|
||||
});
|
||||
test('finding numbering and matching header identity are presentation, not extra findings',()=>{
|
||||
for(const c of realFindings()){
|
||||
const q=c.questions[0]!,oldTitle=q.question.split('\n')[0]!;q.question=q.question.replace(/^D\d+\s*[—–-]\s*/,'');answer(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
q.question=q.question.replace(/^(Finding|Issue)\s+[\d.]+:/,'$1 27.3:');if(/^(Finding|Issue)\s+[\d.]+$/i.test(q.header))q.header=q.header.replace(/[\d.]+/,'27.3');answer(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
expect(oldTitle).toContain('?');
|
||||
}
|
||||
});
|
||||
test('native completion, identity, offered answer and unambiguous single issue remain mandatory',()=>{
|
||||
const mutations:Array<(c:NativePlanQuestionCall)=>void>=[c=>{c.answered=false;},c=>{c.failed=true;},c=>{c.answers={};},c=>{c.unansweredQuestionIndices=[0];},c=>{c.questions[0]!.multiSelect=true;},c=>{c.answers={[c.questions[0]!.question]:'unoffered'};},c=>{c.questions.push(structuredClone(c.questions[0]!));},c=>{c.questions[0]!.options[1]!.label=c.questions[0]!.options[0]!.label;answer(c);}];
|
||||
for(const mutate of mutations)for(const c of realFindings()){mutate(c);expect(ceoFirstReviewAUQ(fp(c))).toBe(false);}
|
||||
for(const c of realFindings()){expect(ceoFirstReviewAUQ({...fp(c),signature:'foreign:call'})).toBe(false);expect(ceoFirstReviewAUQ({...fp(c),nativeCall:undefined})).toBe(false);expect(ceoFirstReviewAUQ({...fp(c),options:[]})).toBe(false);}
|
||||
});
|
||||
test('setup, quoted examples, foreign qids and contradictory numbered headers cannot start review',()=>{
|
||||
for(const prefix of ['Example: ','> ','"','```\n'])for(const c of realFindings()){c.questions[0]!.question=prefix+c.questions[0]!.question;expect(ceoFirstReviewAUQ(fp(answer(c)))).toBe(false);}
|
||||
for(const header of ['Approach','Mode','Next review','Setup','Finding 88','Issue 88'])for(const c of realFindings()){c.questions[0]!.header=header;expect(ceoFirstReviewAUQ(fp(c))).toBe(false);}
|
||||
for(const c of realFindings()){c.questions[0]!.question+=' <gstack-qid:plan-eng-review-finding>';expect(ceoFirstReviewAUQ(fp(answer(c)))).toBe(false);}
|
||||
for(const which of ['distinct','paired'] as const)for(const c of fixture.cases[which].calls.slice(0,4))expect(ceoFirstReviewAUQ(fp(c as NativePlanQuestionCall))).toBe(false);
|
||||
});
|
||||
test('a completed pure next-review menu is administrative without granting pending input permission',()=>{
|
||||
const c=get('paired',6);expect(isCeoCompletionHandoff(fp(c))).toBe(true);expect(ceoFirstReviewAUQ(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();c.answered=false;delete c.answers;delete c.answeredAt;c.unansweredQuestionIndices=[0];expect(isCeoCompletionHandoff(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();
|
||||
});
|
||||
test('new work or uncertain closure in the next-review choice stays substantive',()=>{
|
||||
for(const suffix of ['\nFix the missing authentication check.','\nDelete the CI gate.','\nShip the new endpoint now.','\nWhich new endpoint should we add?']){const c=get('paired',6);c.questions[0]!.question+=suffix;expect(isCeoCompletionHandoff(fp(answer(c)))).toBe(false);}
|
||||
for(const change of ['CEO review is not complete.','CEO review will be complete.','Example: CEO review complete.']){const c=get('paired',6);c.questions[0]!.question=c.questions[0]!.question.replace('CEO review complete.',change);expect(isCeoCompletionHandoff(fp(answer(c)))).toBe(false);}
|
||||
for(const mutation of [c=>{c.failed=true;},c=>{c.questions[0].multiSelect=true;},c=>{c.answers={[c.questions[0].question]:'Fix the bug first'};},c=>{c.questions[0].options.push({label:'Fix the security issue',description:'Add a new check.'});}] as Array<(c:NativePlanQuestionCall)=>void>){const c=get('paired',6);mutation(c);expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
|
||||
});
|
||||
|
||||
// The next-gate explanation must never turn conditional CEO closure into a
|
||||
// completed review. Its narrow normalization is for counting only.
|
||||
test('next Eng gate timing cannot supply conditional CEO completion', () => {
|
||||
for (const replacement of [
|
||||
'The CEO review is complete until someone runs it later.',
|
||||
'The CEO review is complete if someone runs it later.',
|
||||
'The CEO review will be complete after someone runs it later.',
|
||||
'The CEO review still has unresolved findings.',
|
||||
]) {
|
||||
const call = get('paired', 6);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace(
|
||||
'The CEO review cleared scope and strengthened both test assertions.', replacement);
|
||||
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('actual evidence and its regression select the paid CEO counting test', () => {
|
||||
for (const file of ['test/ceo-count-ad-v2.test.ts', 'test/fixtures/ceo-count-ad-v2.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toContain('plan-ceo-finding-count');
|
||||
}
|
||||
});
|
||||
|
||||
test('the actual completed retry keeps two findings and its body-closure handoff administrative', () => {
|
||||
const calls = fixture.cases.pairedRetry.calls as NativePlanQuestionCall[];
|
||||
expect(count(calls)).toEqual({setup: 2, review: 2, administrative: 1});
|
||||
expect(fixture.cases.pairedRetry.observation.outcome).toBe('no_review_questions');
|
||||
expect(fixture.cases.pairedRetry.observation.completionCredit).toBe(false);
|
||||
expect(isCeoCompletionHandoff(fp(get('pairedRetry', 4)))).toBe(true);
|
||||
expect(pickCeoCompletionHandoff(fp(get('pairedRetry', 4)))).toBeNull();
|
||||
});
|
||||
|
||||
test('body closure and echoed choices cannot hide new work or uncertain CEO closure', () => {
|
||||
for (const text of ['Fix the missing authentication check.', 'Delete the CI gate.', 'Ship the new endpoint now.', 'Which endpoint should we add?']) {
|
||||
const call = get('pairedRetry', 4);
|
||||
call.questions[0]!.question += '\n' + text;
|
||||
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
|
||||
}
|
||||
for (const text of ['The CEO review is not done', 'The CEO review will be done', 'The CEO review is done if the fixes land', 'Example: The CEO review is done']) {
|
||||
const call = get('pairedRetry', 4);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('The CEO review is done', text);
|
||||
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
|
||||
}
|
||||
const pending = get('pairedRetry', 4); pending.answered = false; delete pending.answers; delete pending.answeredAt; pending.unansweredQuestionIndices = [0];
|
||||
expect(isCeoCompletionHandoff(fp(pending))).toBe(false);
|
||||
expect(pickCeoCompletionHandoff(fp(pending))).toBeNull();
|
||||
expect(isCeoCompletionHandoff({...fp(get('pairedRetry', 4)), signature: 'foreign:call'})).toBe(false);
|
||||
});
|
||||
|
||||
test('every offered navigation clause rejects a new repair rather than hiding it under a valid recap', () => {
|
||||
for (const [which, index] of [['paired', 6], ['pairedRetry', 4]] as const) {
|
||||
for (const extra of ['Delete the CI gate.', 'Repair the retry assertion.', 'Disable authentication.', 'Please rewrite the endpoint.']) {
|
||||
for (const optionIndex of [0, 1]) {
|
||||
const call = get(which, index);
|
||||
call.questions[0]!.options[optionIndex]!.description += ' ' + extra;
|
||||
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
|
||||
}
|
||||
for (const where of ['before-net', 'inside-eli10'] as const) {
|
||||
const call = get(which, index);
|
||||
call.questions[0]!.question = where === 'before-net'
|
||||
? call.questions[0]!.question.replace('\nNet:', '\n' + extra + '\nNet:')
|
||||
: call.questions[0]!.question.replace('\nStakes if', ' ' + extra + '\nStakes if');
|
||||
expect(isCeoCompletionHandoff(fp(answer(call)))).toBe(false);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('timing annotations cannot conceal substantive instructions', () => {
|
||||
for (const [which, index] of [['paired', 6], ['pairedRetry', 4]] as const) {
|
||||
for (const text of [' (human: Delete the CI gate)', ' (human: ~2 min / CC: Disable authentication)', ' (human: ~2 min / CC: ~1 min; repair the retry assertion)']) {
|
||||
const c = get(which, index); c.questions[0]!.options[0]!.description += text;
|
||||
expect(isCeoCompletionHandoff(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -1,88 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { capturePlanCountQuestion, nativePlanCallFingerprint, planCountQuestionInput } from './helpers/claude-pty-runner';
|
||||
import { pickCeoCountQuestion } from './helpers/ceo-approach-pick';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import recorded from './fixtures/ceo-count-mode-ab-call.json';
|
||||
|
||||
function pending(): NativePlanQuestionCall {
|
||||
const call = structuredClone(recorded) as NativePlanQuestionCall;
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
delete call.answeredAt;
|
||||
return call;
|
||||
}
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
function screen(call: NativePlanQuestionCall): string {
|
||||
const q = call.questions[0]!;
|
||||
return `☐ ${q.header}\n${q.question}\n${q.options.map((o, i) => `${i ? ' ' : '❯'} ${i + 1}. ${o.label}`).join('\n')}\nEnter to select · ↑/↓ to navigate · Esc to cancel`;
|
||||
}
|
||||
|
||||
describe('fixed-scope CEO finding-count mode', () => {
|
||||
test('the AB native mode menu selects HOLD SCOPE rather than its first expansion option', () => {
|
||||
// The retained live record already contains the expansion answer. The
|
||||
// pending state and frame are projections, not proof of live availability.
|
||||
const call = pending();
|
||||
const active = capturePlanCountQuestion(screen(call), new Set(), 0, true, call)!;
|
||||
expect(active.nativeCall).toBe(call);
|
||||
const selected = pickCeoCountQuestion(fp(call), active) ?? 1;
|
||||
expect(selected).toBe(3);
|
||||
expect(planCountQuestionInput(screen(call), active, selected)).toBe('3');
|
||||
const q = recorded.questions[0]!;
|
||||
expect(recorded.answers[q.question]).toBe(q.options[0]!.label);
|
||||
expect(pickCeoCountQuestion(fp(recorded as NativePlanQuestionCall))).toBeNull();
|
||||
});
|
||||
|
||||
test('every offered position chooses the same fixed scope, independent of the recommendation', () => {
|
||||
for (let shift = 0; shift < 4; shift++) {
|
||||
const call = pending();
|
||||
const q = call.questions[0]!;
|
||||
q.options = [...q.options.slice(shift), ...q.options.slice(0, shift)];
|
||||
q.options.forEach(o => { o.label = o.label.replace(/ \(Recommended\)$/, ''); });
|
||||
q.options.find(o => o.label.startsWith('SCOPE EXPANSION'))!.label += ' (Recommended)';
|
||||
expect(pickCeoCountQuestion(fp(call))).toBe(q.options.findIndex(o => o.label.startsWith('HOLD SCOPE')) + 1);
|
||||
}
|
||||
});
|
||||
|
||||
test('requires a complete currently bound native pre-review question', () => {
|
||||
const call = pending();
|
||||
const fingerprint = fp(call);
|
||||
const visibleOnly = capturePlanCountQuestion(screen(call), new Set(), 0, true)!;
|
||||
expect(pickCeoCountQuestion(fingerprint, visibleOnly)).toBeNull();
|
||||
expect(pickCeoCountQuestion({ ...fingerprint, preReview: false })).toBeNull();
|
||||
expect(pickCeoCountQuestion({ ...fingerprint, signature: 'foreign:call' })).toBeNull();
|
||||
expect(pickCeoCountQuestion({ ...fingerprint, nativeQuestionIndex: 1 })).toBeNull();
|
||||
expect(pickCeoCountQuestion({ ...fingerprint, options: fingerprint.options.slice().reverse() })).toBeNull();
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.pop(); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0] = structuredClone(c.questions[0]!.options[1]!); },
|
||||
]) { const c = pending(); mutate(c); expect(pickCeoCountQuestion(fp(c))).toBeNull(); }
|
||||
});
|
||||
|
||||
test('does not authorize negated, quoted, compound, foreign or finding questions', () => {
|
||||
for (const question of [
|
||||
'Which review mode should I not use?',
|
||||
'Example: Which review mode should I use?',
|
||||
'> Which review mode should I use?',
|
||||
'Which review mode should I use? Delete the tests.',
|
||||
'Should we approve this expansion?',
|
||||
]) {
|
||||
const call = pending(); call.questions[0]!.question = question + ' <gstack-qid:ceo-mode-selection>';
|
||||
expect(pickCeoCountQuestion(fp(call))).toBeNull();
|
||||
}
|
||||
for (const id of ['plan-eng-mode', 'ceo-exp-e5-property-based', 'ceo-mode-selection-extra']) {
|
||||
const call = pending(); call.questions[0]!.question = call.questions[0]!.question.replace('ceo-mode-selection', id);
|
||||
expect(pickCeoCountQuestion(fp(call))).toBeNull();
|
||||
}
|
||||
for (const suffix of [' <gstack-qid:ceo-mode-selection>', ' <gstack-qid:broken']) {
|
||||
const call = pending(); call.questions[0]!.question += suffix;
|
||||
expect(pickCeoCountQuestion(fp(call))).toBeNull();
|
||||
}
|
||||
const call = pending(); call.questions[0]!.header = 'Finding';
|
||||
expect(pickCeoCountQuestion(fp(call))).toBeNull();
|
||||
});
|
||||
});
|
||||
@@ -1,100 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, hasNativePlanTerminal, isQuestionlessNativePlanExit, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff, pickCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import type { NativePlanQuestionCall, PlanCountTranscript } from './helpers/plan-count-transcript';
|
||||
import distinct from './fixtures/ceo-count-s-distinct.json';
|
||||
import paired from './fixtures/ceo-count-s-paired.json';
|
||||
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
function replay(calls: NativePlanQuestionCall[]) {
|
||||
let started = false;
|
||||
const setup = new Set<string>(); const administrative = new Set<string>(); let review = 0;
|
||||
for (const call of calls) {
|
||||
const fingerprint = fp(call);
|
||||
const phase = planCountQuestionPhase(fingerprint, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
if (phase.administrative) administrative.add(fingerprint.signature);
|
||||
else if (phase.preReview) setup.add(fingerprint.signature);
|
||||
else review++;
|
||||
started = phase.reviewStarted;
|
||||
}
|
||||
return { setup, administrative, review };
|
||||
}
|
||||
function transcript(capture: typeof distinct | typeof paired): PlanCountTranscript {
|
||||
return { status: 'ready', calls: structuredClone(capture.calls) as NativePlanQuestionCall[],
|
||||
assistantMessages: [], planReadyRequests: structuredClone(capture.planReadyRequests) };
|
||||
}
|
||||
function withReport(capture: typeof distinct | typeof paired, run: (file: string, start: number) => void) {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-s-terminal-'));
|
||||
const file = path.join(dir, 'report.md');
|
||||
fs.writeFileSync(file, capture.report);
|
||||
fs.utimesSync(file, capture.reportMtimeMs / 1000, capture.reportMtimeMs / 1000);
|
||||
try { run(file, capture.reportMtimeMs - 1000); }
|
||||
finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
}
|
||||
|
||||
describe('captured S native CEO completion gates', () => {
|
||||
test('setup-only exit fails promptly without relaxing report freshness for positive coverage', () => {
|
||||
const t = transcript(distinct); const result = replay(t.calls);
|
||||
expect(result.setup.size).toBe(4); expect(result.review).toBe(0);
|
||||
withReport(distinct, (file, start) => {
|
||||
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen, result.setup)).toBe(true);
|
||||
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen)).toBe(false);
|
||||
expect(hasNativePlanTerminal(t, file, start, 'plan_ready')).toBe(false);
|
||||
for (const mutate of [
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.answered = false; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.failed = true; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.sessionId = 'foreign'; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.answeredAt = 'invalid'; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.answeredAt = v.planReadyRequests![0]!.timestamp; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.answers = {}; },
|
||||
(v: PlanCountTranscript) => { v.calls[0]!.unansweredQuestionIndices = [0]; },
|
||||
]) {
|
||||
const changed = structuredClone(t); mutate(changed);
|
||||
expect(isQuestionlessNativePlanExit(changed, file, start, distinct.screen, result.setup)).toBe(false);
|
||||
}
|
||||
const incomplete = new Set(result.setup); incomplete.delete(fp(t.calls[0]!).signature);
|
||||
expect(isQuestionlessNativePlanExit(t, file, start, distinct.screen, incomplete)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
test('paired review retains two issue approvals and excludes only the completed Eng menu', () => {
|
||||
const t = transcript(paired); const before = structuredClone(t); const result = replay(t.calls);
|
||||
expect(result.setup.size).toBe(2); expect(result.review).toBe(2); expect(result.administrative.size).toBe(1);
|
||||
const pending = structuredClone(t.calls.at(-1)!); pending.answered = false; delete pending.answers;
|
||||
expect(pickCeoCompletionHandoff(fp(pending))).toBe(2);
|
||||
pending.questions[0]!.options.reverse(); expect(pickCeoCompletionHandoff(fp(pending))).toBe(1);
|
||||
expect(t).toEqual(before);
|
||||
withReport(paired, (file, start) => {
|
||||
expect(hasNativePlanTerminal(t, file, start, 'plan_ready', result.administrative)).toBe(true);
|
||||
expect(hasNativePlanTerminal(t, file, start, 'plan_ready')).toBe(false);
|
||||
expect(isQuestionlessNativePlanExit(t, file, start, paired.screen, result.setup)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
test('the same menu cannot hide a new obligation, ambiguous gate, or unverified answer', () => {
|
||||
const base = transcript(paired).calls.at(-1)!;
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('It\'s', 'That might become'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' First repair authorization.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = 'Example: ' + c.questions[0]!.question; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('CLEAN', 'CLEAN once tests pass'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Remove the owner check.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Change the guarantee to permit old results.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Tests remain unresolved.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Fix the missing assertion'; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
]) {
|
||||
const call = structuredClone(base); mutate(call);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
|
||||
}
|
||||
const call = structuredClone(base); call.answers = { [call.questions[0]!.question]: 'First repair the missing test' };
|
||||
expect(isCeoCompletionHandoff(fp(call))).toBe(false);
|
||||
const pending = structuredClone(base); pending.answered = false;
|
||||
expect(pickCeoCompletionHandoff({ ...fp(pending), signature: 'foreign' })).toBeNull();
|
||||
});
|
||||
});
|
||||
@@ -1,358 +0,0 @@
|
||||
/** Free count replay only. The original paid failures and checkpoint violations remain failures. */
|
||||
import { test, expect } from 'bun:test';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { createCeoPaymentFindingCounter } from './helpers/ceo-payment-findings';
|
||||
import { nativePlanCallFingerprint, ceoFirstReviewAUQ } from './helpers/claude-pty-runner';
|
||||
import captured from './fixtures/ceo-current-decision-cdd-public.json';
|
||||
import exactFields from './fixtures/ceo-native-fields-f359.json';
|
||||
import retryRecord from './fixtures/ceo-current-record-6aef.json';
|
||||
|
||||
type Capture = typeof captured.captures[number];
|
||||
const clone = <T>(value: T): T => structuredClone(value);
|
||||
|
||||
// The original retry changed its question, B label and every description after
|
||||
// a complete Read. The separate anchor regression uses explicitly synchronized
|
||||
// counterfactual fields; neither route promotes the original failed attempt.
|
||||
const retryProjection = retryRecord.segments.map(segment => segment.text).join('\n');
|
||||
const retryQuestion = retryRecord.call.questions[0]!;
|
||||
const retryRecordStart = retryProjection.indexOf('## currentDecision (R4)');
|
||||
const retryFieldsStart = retryProjection.indexOf('Question:', retryRecordStart);
|
||||
const retryExactFields = `Question: ${retryQuestion.question}\nHeader: ${retryQuestion.header}\n` +
|
||||
retryQuestion.options.map((option, index) =>
|
||||
`${/^[A-D][).:]\s/.test(option.label) ? '' : `${'ABCD'[index]}) `}${option.label}\n${option.description}`).join('\n') + '\n';
|
||||
const retrySynchronized = retryProjection.slice(0, retryFieldsStart) + retryExactFields;
|
||||
function countRetryRecord(plan: string, call = clone(retryRecord.call)) {
|
||||
const counter = createCeoPaymentFindingCounter(retryRecord.seed, () => plan, () => false);
|
||||
const counted = counter.isReviewAUQ(nativePlanCallFingerprint(call, 1, false));
|
||||
return { counted, trace: counter.trace };
|
||||
}
|
||||
test('6aef retry literal source projection preserves actual native drift rejection', () => {
|
||||
for (const segment of retryRecord.segments)
|
||||
expect(createHash('sha256').update(segment.text).digest('hex')).toBe(segment.sha256);
|
||||
expect(retryRecordStart).toBeGreaterThan(0); expect(retryFieldsStart).toBeGreaterThan(retryRecordStart);
|
||||
expect(retryRecord.call.answered).toBe(true);
|
||||
expect(() => countRetryRecord(retryProjection)).toThrow(/Unsupported/);
|
||||
expect(countRetryRecord(retrySynchronized)).toMatchObject({ counted: true });
|
||||
expect(countRetryRecord(retrySynchronized).trace.at(-1)).toMatchObject({ kind: 'recorded-decision', ledgerId: 'R4' });
|
||||
});
|
||||
for (const heading of [
|
||||
'### Per-item coverage (pending R4)', '### Test coverage for R4', '### TODO follow-up (R4)',
|
||||
'### R4 section notes', '### R4 implementation tasks',
|
||||
]) test(`an incidental current row heading does not own a second record: ${heading}`, () => {
|
||||
expect(countRetryRecord(retrySynchronized.replace('### Per-item coverage (pending R4)', heading)).counted).toBe(true);
|
||||
});
|
||||
test('a contextual parent row heading does not borrow the nested record fields', () => {
|
||||
const plan = retrySynchronized.replace('## currentDecision (R4)', '## R4 coverage context\n\n### currentDecision (R4)');
|
||||
expect(countRetryRecord(plan).counted).toBe(true);
|
||||
});
|
||||
for (const heading of ['## R4 decision', '## Pending R4 options', '## R4 comparison'])
|
||||
test(`a generic owned heading can introduce complete native fields: ${heading}`, () => {
|
||||
expect(countRetryRecord(retrySynchronized.replace('## currentDecision (R4)', heading)).counted).toBe(true);
|
||||
});
|
||||
// A declaration owns a record regardless of row/name order or whether its
|
||||
// fields have been filled yet; incompleteness cannot remove an ambiguity.
|
||||
for (const kind of ['decision', 'review', 'options', 'approaches', 'comparison'])
|
||||
for (const heading of [`## ${kind} R4`, `## R4 ${kind}`, `## Pending R4 ${kind}`, `## Current ${kind} for R4`])
|
||||
for (const body of ['', '\n\nStatus: pending'])
|
||||
test(`an explicit record declaration competes before its fields exist: ${heading} ${body}`, () => {
|
||||
expect(() => countRetryRecord(retrySynchronized + '\n\n' + heading + body)).toThrow(/Unsupported/);
|
||||
});
|
||||
|
||||
for (const [name, record] of Object.entries({
|
||||
'empty named heading': '## currentDecision (R4)',
|
||||
'explicit decision status': '## Decision R4\n\nStatus: pending',
|
||||
'explicit review state': '## Review R4\n\nState: current',
|
||||
'empty decision declaration': '## Decision R4',
|
||||
'empty review declaration': '## Review R4',
|
||||
'incomplete named heading': '## currentDecision (R4)\n\nQuestion: incomplete',
|
||||
'empty named paragraph': '**currentDecision: R4**',
|
||||
'explicit options declaration': 'Options for R4:',
|
||||
'question fields': '## R4 other record\n\nQuestion: another question',
|
||||
'header fields': '## R4 other record\n\nHeader: another question',
|
||||
'option paragraph': '## R4 other record\n\nA) Another option\nB) Another choice',
|
||||
'option list': '## R4 other record\n\n- A) Another option\n- B) Another choice',
|
||||
'option comparison table': '## R4 other record\n\n| Option | Effort |\n| --- | --- |\n| A | S |\n| B | M |',
|
||||
'column comparison table': '## R4 other record\n\n| Commitment | A | B |\n| --- | --- | --- |\n| Work | fixed | changed |',
|
||||
'literal comparison grid': '## R4 other record\n\n```text\nCommitment | A | B\nWork | fixed | changed\n```',
|
||||
'complete duplicate': '## currentDecision (R4)\n\n' + retryExactFields,
|
||||
})) test(`a competing current record remains ambiguous: ${name}`, () => {
|
||||
const plan = retrySynchronized + '\n\n' + record + '\n';
|
||||
expect(() => countRetryRecord(plan)).toThrow(/Unsupported/);
|
||||
});
|
||||
for (const example of [
|
||||
'> Question: example only', '```text\nQuestion: example only\nHeader: example\n```',
|
||||
'"Question: example only"', '`Question: example only`',
|
||||
]) test(`quoted field examples do not own another current record: ${JSON.stringify(example)}`, () => {
|
||||
expect(countRetryRecord(retrySynchronized + '\n\n## R4 explanatory notes\n\n' + example).counted).toBe(true);
|
||||
});
|
||||
for (const [name, change] of Object.entries({
|
||||
question: (s: string) => s.replace(retryQuestion.question, retryQuestion.question + ' Changed.'),
|
||||
label: (s: string) => s.replace(retryQuestion.options[1]!.label, 'B) Changed choice'),
|
||||
description: (s: string) => s.replace(retryQuestion.options[0]!.description!, 'Shortened description.'),
|
||||
source: (s: string) => s.replaceAll('PLAN.md', 'other/PLAN.md'),
|
||||
row: (s: string) => s.replace('## currentDecision (R4)', '## currentDecision (R99)'),
|
||||
})) test(`incidental headings cannot bypass native or source identity: ${name}`, () => {
|
||||
const plan = change(retrySynchronized); expect(plan).not.toBe(retrySynchronized);
|
||||
expect(() => countRetryRecord(plan)).toThrow(/Unsupported/);
|
||||
});
|
||||
test('a complete saved record still needs an actual answer', () => {
|
||||
const call = clone(retryRecord.call); call.answered = false;
|
||||
expect(() => countRetryRecord(retrySynchronized, call)).toThrow(/Unsupported|Invalid/);
|
||||
});
|
||||
|
||||
const paired = captured.captures[0]!, distinct = captured.captures[1]!, retry = captured.captures[2]!;
|
||||
function replay(row: Capture, plan = row.savedPlan, calls = clone(row.calls)) {
|
||||
const counter = createCeoPaymentFindingCounter(row.source, () => plan, ceoFirstReviewAUQ);
|
||||
const counted = calls.map((call, index) => counter.isReviewAUQ(nativePlanCallFingerprint(call, 1, false), calls.slice(0, index)));
|
||||
return { counted, trace: counter.trace };
|
||||
}
|
||||
function reject(row: Capture, plan: string, calls = clone(row.calls)) {
|
||||
expect(() => replay(row, plan, calls)).toThrow(/Unsupported|Invalid/);
|
||||
}
|
||||
for (const row of captured.captures) test(`${row.name}: exact public calls and saved record receive count credit, never paid PASS credit`, () => {
|
||||
expect(createHash('sha256').update(row.source).digest('hex')).toBe(row.sourceSha256);
|
||||
expect(createHash('sha256').update(row.savedPlan).digest('hex')).toBe(row.savedSha256);
|
||||
expect(row.originalOutcome).toBe('FAIL'); expect(row.paidPassCredit).toBe(0);
|
||||
const result = replay(row);
|
||||
expect(result.counted).toEqual(row.calls.map((_, i) => i === row.calls.length - 1));
|
||||
expect(result.trace.at(-1)).toMatchObject({ kind: 'recorded-decision', ledgerId: row === paired ? 'D1' : 'R1' });
|
||||
});
|
||||
|
||||
const pairedMarker = paired.savedPlan.match(/^\*\*(currentDecision: D1[^\n]+)\*\*$/m)![1]!;
|
||||
for (const marker of [pairedMarker, `**${pairedMarker}**`, `### ${pairedMarker}`, `#### ${pairedMarker}`])
|
||||
test(`current comparison marker retains Markdown presentation ${marker.slice(0, 20)}`, () => {
|
||||
expect(replay(paired, paired.savedPlan.replace(`**${pairedMarker}**`, marker)).counted.at(-1)).toBe(true);
|
||||
});
|
||||
const rowMarker = distinct.savedPlan.match(/^\*\*(Row R1[^\n]+)\*\*$/m)![1]!;
|
||||
for (const marker of [rowMarker, `**${rowMarker}**`, `### ${rowMarker}`])
|
||||
test(`row marker under currentDecision retains Markdown presentation ${marker.slice(0, 14)}`, () => {
|
||||
expect(replay(distinct, distinct.savedPlan.replace(`**${rowMarker}**`, marker)).counted.at(-1)).toBe(true);
|
||||
});
|
||||
for (const row of [paired, distinct]) {
|
||||
const marker = row === paired ? pairedMarker : rowMarker;
|
||||
const id = row === paired ? 'D1' : 'R1';
|
||||
for (const [name, change] of Object.entries({
|
||||
'quoted marker': (s: string) => s.replace(`**${marker}**`, `> **${marker}**`),
|
||||
'fenced marker': (s: string) => s.replace(`**${marker}**`, '```text\n'+marker+'\n```'),
|
||||
'different row marker': (s: string) => s.replace(`**${marker}**`, `**${marker.replace(id, 'R999')}**`),
|
||||
'duplicated current marker': (s: string) => s.replace(`**${marker}**`, `**${marker}**\n\n**${marker}**`),
|
||||
'withdrawn current marker': (s: string) => s.replace(`**${marker}**`, `**${marker}**\nThis decision is withdrawn.`),
|
||||
'historical comparison': (s: string) => s.replace(`**${marker}**`, `## Historical comparison\n\n**${marker}**`),
|
||||
'foreign source': (s: string) => s.replaceAll('PLAN.md', 'other/PLAN.md'),
|
||||
'missing source': (s: string) => s.replaceAll('PLAN.md', 'input'),
|
||||
'missing current row': (s: string) => s.replace(new RegExp('^\\| '+id+'(?:\\s|\\|)[^\\n]+\\n','m'), ''),
|
||||
'missing option risk': (s: string) => s.replace('Risk low.', ''),
|
||||
'invalid option risk': (s: string) => s.replace('Risk low.', 'Risk unknown.'),
|
||||
'invalid option effort': (s: string) => s.replace('Effort S ', 'Effort XS '),
|
||||
'withdrawn option': (s: string) => s.replace('Pros:', 'Pros: This option is withdrawn.'),
|
||||
})) test(`${row.name}: current paragraph rejects ${name}`, () => {
|
||||
const changed = change(row.savedPlan); expect(changed !== row.savedPlan).toBe(true); reject(row, changed);
|
||||
});
|
||||
}
|
||||
test('bare Row marker cannot borrow a non-currentDecision heading', () => {
|
||||
reject(distinct, distinct.savedPlan.replace('## currentDecision', '## Unrelated notes'));
|
||||
});
|
||||
|
||||
for (const verb of ['Keep', 'Retain', 'Preserve']) for (const form of ['suffix', 'prefix', 'description']) test(`saved and offered ${verb} baseline resolve symmetrically (${form})`, () => {
|
||||
const calls = clone(retry.calls), q = calls.at(-1)!.questions[0]!;
|
||||
q.options[2]!.label = q.options[2]!.label.replace('Keep', verb);
|
||||
const caption = form === 'suffix' ? `C) ${verb} truthy only (as planned).`
|
||||
: form === 'prefix' ? `**C) As planned: ${verb} truthy only.**` : `**C) ${verb} truthy only** (as planned) —`;
|
||||
const plan = retry.savedPlan.replace('C) Keep truthy only (as planned).', caption);
|
||||
expect(replay(retry, plan, calls).counted.at(-1)).toBe(true);
|
||||
});
|
||||
for (const [name, caption] of Object.entries({
|
||||
'added action': 'Keep truthy only and delete records',
|
||||
'changed negation': 'Do not keep truthy only',
|
||||
'narrowed scope': 'Keep truthy only for admins',
|
||||
'different baseline': 'Keep rejection only',
|
||||
})) test(`same-letter saved baseline rejects ${name}`, () => {
|
||||
reject(retry, retry.savedPlan.replace('C) Keep truthy only (as planned).', `C) ${caption} (as planned).`));
|
||||
});
|
||||
|
||||
// Exercise the existing strict exact-native-fields path with the new marker
|
||||
// presentations. This is distinct from the older complete-prose count path.
|
||||
const q = exactFields.call.questions[0]!;
|
||||
const begin = exactFields.savedPlan.indexOf('### currentDecision (D1)');
|
||||
const end = exactFields.savedPlan.indexOf('## NOT in scope', begin);
|
||||
const fields = ['Question: '+q.question, 'Header: '+q.header,
|
||||
...q.options.map(o => o.label+'\n'+o.description)].join('\n\n');
|
||||
function exactPlan(marker: string, body = fields) {
|
||||
return exactFields.savedPlan.slice(0, begin)+marker+'\n\n'+body+'\n\n'+exactFields.savedPlan.slice(end);
|
||||
}
|
||||
function exactCount(plan: string, call = clone(exactFields.call)) {
|
||||
return createCeoPaymentFindingCounter(exactFields.seed, () => plan, ceoFirstReviewAUQ)
|
||||
.isReviewAUQ(nativePlanCallFingerprint(call, 1, false));
|
||||
}
|
||||
for (const marker of ['**Row D1 — current question**', '**currentDecision (D1)**']) {
|
||||
const heading = '### currentDecision (D1)';
|
||||
test(`one exact record retains its heading plus immediate paragraph marker ${marker}`, () => {
|
||||
expect(exactCount(exactPlan(heading+'\n\n'+marker))).toBe(true);
|
||||
});
|
||||
test(`same-row heading continuation cannot hide a second full record ${marker}`, () => {
|
||||
expect(() => exactCount(exactPlan(heading+'\n\n'+marker, fields+'\n\n'+heading+'\n\n'+marker+'\n\n'+fields))).toThrow(/Unsupported/);
|
||||
});
|
||||
test(`same-row heading continuation cannot hide a later paragraph record ${marker}`, () => {
|
||||
expect(() => exactCount(exactPlan(heading+'\n\n'+marker, fields+'\n\n'+marker+'\n\n'+fields))).toThrow(/Unsupported/);
|
||||
});
|
||||
}
|
||||
for (const marker of ['### currentDecision (D1)', '**currentDecision (D1)**', 'currentDecision (D1)']) {
|
||||
test(`full native fields count with ${marker}`, () => expect(exactCount(exactPlan(marker))).toBe(true));
|
||||
for (const [name, change] of Object.entries({
|
||||
'missing Question': (s: string) => s.replace('Question: '+q.question, ''),
|
||||
'mismatched Header': (s: string) => s.replace('Header: '+q.header, 'Header: Another decision'),
|
||||
'missing option description': (s: string) => s.replace(q.options[0]!.description!, ''),
|
||||
'invalid effort domain': (s: string) => s.replace('Effort S', 'Effort XS'),
|
||||
'invalid risk domain': (s: string) => s.replace(/Risk (?:low|medium|high)/i, 'Risk unknown'),
|
||||
})) test(`${marker}: strict native fields reject ${name}`, () => {
|
||||
expect(() => exactCount(exactPlan(marker, change(fields)))).toThrow(/Unsupported/);
|
||||
});
|
||||
}
|
||||
for (const row of captured.captures) for (const defect of ['missing ACK', 'failed ACK', 'unoffered answer', 'foreign identity'])
|
||||
test(`${row.name}: paragraph normalization retains ${defect} rejection`, () => {
|
||||
const calls = clone(row.calls), call = calls.at(-1)!;
|
||||
if (defect === 'missing ACK') call.answered = false;
|
||||
if (defect === 'failed ACK') call.failed = true;
|
||||
if (defect === 'unoffered answer') call.answers = { [call.questions[0]!.question]: 'Not offered' };
|
||||
if (defect === 'foreign identity') call.sessionId = '';
|
||||
reject(row, row.savedPlan, calls);
|
||||
});
|
||||
|
||||
test('distinct retry retains its actual preceding D2 count and rejects D3 without an owned ledger row', () => {
|
||||
const row = captured.rejectedMissingRow;
|
||||
expect(row.originalOutcome).toBe('FAIL'); expect(row.paidPassCredit).toBe(0);
|
||||
expect(createHash('sha256').update(row.source).digest('hex')).toBe(row.sourceSha256);
|
||||
row.plans.forEach((plan, i) => {
|
||||
expect(createHash('sha256').update(plan).digest('hex')).toBe(row.planSha256[i]);
|
||||
expect(Date.parse(row.snapshotTimes[i]!)).toBeLessThan(Date.parse(row.questionTimes[i]!));
|
||||
});
|
||||
let plan = row.plans[0]!;
|
||||
const counter = createCeoPaymentFindingCounter(row.source, () => plan, ceoFirstReviewAUQ);
|
||||
expect(counter.isReviewAUQ(nativePlanCallFingerprint(clone(row.calls[0]!), 1, false))).toBe(false);
|
||||
expect(counter.isReviewAUQ(nativePlanCallFingerprint(clone(row.calls[1]!), 1, false), row.calls.slice(0, 1))).toBe(true);
|
||||
plan = row.plans[1]!;
|
||||
expect(plan).toContain('### currentDecision (D3, owner Section 2)');
|
||||
expect(/^\| D3\b/m.test(plan)).toBe(false);
|
||||
expect(() => counter.isReviewAUQ(nativePlanCallFingerprint(clone(row.calls[2]!), 1, false), row.calls.slice(0, 2))).toThrow(/Unsupported/);
|
||||
expect(counter.trace).toHaveLength(2);
|
||||
});
|
||||
|
||||
test('the actual CEO save layout preserves the full native payload and separates prior records', () => {
|
||||
const template = readFileSync(`${import.meta.dir}/../plan-ceo-review/SKILL.md.tmpl`, 'utf8');
|
||||
const layout = template.match(/```text\n( ## currentDecision \(ROW-ID\)[\s\S]+?)\n ```/);
|
||||
expect(layout).not.toBeNull();
|
||||
const grid = exactFields.savedPlan.slice(begin, end).match(/```text\n[\s\S]+?\n```/);
|
||||
expect(grid).not.toBeNull();
|
||||
// Fill the actual source example with the existing captured native fields;
|
||||
// do not reconstruct a more permissive format or promote its original FAIL.
|
||||
const record = layout![1]!.replace(/^ /gm, '')
|
||||
.replace('ROW-ID', 'D1').replace('<complete grid>', '\n\n'+grid![0])
|
||||
.replace('<complete currentDecision.question>', q.question)
|
||||
.replace('<exact currentDecision.header>', q.header)
|
||||
.replace('A) <exact first option label>', q.options[0]!.label)
|
||||
.replace('<full first option description>', q.options[0]!.description!)
|
||||
.replace('B) <exact second option label>', q.options[1]!.label)
|
||||
.replace('<full second option description; repeat for all offered options>',
|
||||
q.options[1]!.description!+'\n'+q.options[2]!.label+'\n'+q.options[2]!.description!);
|
||||
const saved = (section: string) => exactFields.savedPlan.slice(0, begin)+section+'\n\n'+exactFields.savedPlan.slice(end);
|
||||
expect(exactCount(saved(record))).toBe(true);
|
||||
const prior = '## Answered decision D0\nExact approval: prior answer A, scope unchanged.\n'+fields.replaceAll('D1', 'D0');
|
||||
expect(exactCount(saved(prior+'\n\n'+record))).toBe(true);
|
||||
for (const changed of [
|
||||
record.replace(q.question, q.question.split('\n')[0]!),
|
||||
record.replace(q.question.split('\n')[0]!, q.question.split('\n')[0]!+' (changed title)'),
|
||||
record.replace('Header: '+q.header, 'Header: Another decision'),
|
||||
record.replace(q.options[0]!.label, 'A) Delete every test'),
|
||||
record+'\n\n'+fields.replaceAll('D1', 'D0'),
|
||||
record+'\n\n'+record,
|
||||
'```text\n'+record+'\n```',
|
||||
record.replace('Question: ', 'Question:\n'),
|
||||
record.replace('Header: '+q.header, 'Header: '+q.header+'\nOptions:'),
|
||||
]) {
|
||||
expect(changed).not.toBe(record);
|
||||
expect(() => exactCount(saved(changed))).toThrow(/Unsupported/);
|
||||
}
|
||||
});
|
||||
|
||||
test('a reopened row has one current comparison alongside its answered decision history', () => {
|
||||
const oldFields = fields.replace(q.question, q.question.replace(/^D1 — /, 'D0 — D1: '));
|
||||
const currentRecord = '### currentDecision (D1)\n'+fields;
|
||||
const prior = (heading: string) => heading+'\n\nAnswer: A; prior choice retained in history.\n\n'+oldFields;
|
||||
const replaceRecord = (record: string) => exactFields.savedPlan.slice(0, begin)+record+'\n\n'+exactFields.savedPlan.slice(end);
|
||||
for (const heading of ['### Answered decision (D1) — D0', '### Answered decisions for D1']) {
|
||||
expect(exactCount(replaceRecord(prior(heading)+'\n\n'+currentRecord))).toBe(true);
|
||||
// An answered record cannot supply the missing current comparison.
|
||||
expect(() => exactCount(replaceRecord(prior(heading)))).toThrow(/Unsupported/);
|
||||
// A second current record still conflicts; history does not hide it.
|
||||
expect(() => exactCount(replaceRecord(prior(heading)+'\n\n'+currentRecord+'\n\n'+currentRecord))).toThrow(/Unsupported/);
|
||||
}
|
||||
for (const heading of ['### Unanswered decision (D1)', '### Not answered decision (D1)', '### currentDecision (D1)']) {
|
||||
expect(() => exactCount(replaceRecord(prior(heading)+'\n\n'+currentRecord))).toThrow(/Unsupported/);
|
||||
}
|
||||
});
|
||||
|
||||
test('prepared native identity distinguishes the question number from its ledger row before saving', () => {
|
||||
const template = readFileSync(`${import.meta.dir}/../plan-ceo-review/SKILL.md.tmpl`, 'utf8');
|
||||
const titleLayout = template.match(/`(D<N> — <ROW-ID>: <one-line question>)`/)?.[1];
|
||||
expect(titleLayout).toBeDefined();
|
||||
const withoutId = q.question.replace(/^D1 — /, 'D7 — ');
|
||||
const title = titleLayout!.replace('<N>', '7').replace('<ROW-ID>', 'D1')
|
||||
.replace('<one-line question>', q.question.split('\n')[0]!.replace(/^D1 — /, ''));
|
||||
const prepared = withoutId.replace(withoutId.split('\n')[0]!, title);
|
||||
const callWithQuestion = (question: string) => {
|
||||
const call = clone(exactFields.call);
|
||||
call.questions[0]!.question = question;
|
||||
// Counterfactual native questions need their matching answer key too.
|
||||
// This does not alter or approve an original captured question.
|
||||
call.answers = { [question]: Object.values(call.answers)[0]! } as typeof call.answers;
|
||||
return call;
|
||||
};
|
||||
const payload = (call: typeof exactFields.call) => {
|
||||
const current = call.questions[0]!;
|
||||
return ['Question: '+current.question, 'Header: '+current.header,
|
||||
...current.options.map(option => option.label+'\n'+option.description)].join('\n');
|
||||
};
|
||||
const saved = (call: typeof exactFields.call) => exactPlan('### currentDecision (D1)', payload(call));
|
||||
const missing = callWithQuestion(withoutId), ready = callWithQuestion(prepared);
|
||||
// 749df paired retry copied every field and read them all, but omitted its
|
||||
// row ID. The distinct attempt added the ID only after the saved Read.
|
||||
expect(() => exactCount(saved(missing), missing)).toThrow(/Unsupported/);
|
||||
expect(() => exactCount(saved(missing), ready)).toThrow(/Unsupported/);
|
||||
expect(() => exactCount(saved(ready), missing)).toThrow(/Unsupported/);
|
||||
expect(exactCount(saved(ready), ready)).toBe(true);
|
||||
const foreign = callWithQuestion(prepared.replace('D7 — D1:', 'D7 — R999:'));
|
||||
expect(() => exactCount(saved(foreign), foreign)).toThrow(/Unsupported/);
|
||||
expect(() => exactCount(saved(ready)+'\n\n### currentDecision (D1)\n'+payload(ready), ready)).toThrow(/Unsupported/);
|
||||
|
||||
// A late recommended suffix or a brief-only tradeoff list cannot stand in
|
||||
// for the final saved native labels and complete option descriptions.
|
||||
expect(() => exactCount(saved(ready).replace(q.options[0]!.label,
|
||||
q.options[0]!.label.replace(' (recommended)', '')), ready)).toThrow(/Unsupported/);
|
||||
const briefOnly = callWithQuestion(prepared+'\nPros / cons:\n'+q.options.map(option =>
|
||||
option.label+'\n'+option.description!.split('\n').slice(1).join('\n')).join('\n'));
|
||||
for (const option of briefOnly.questions[0]!.options)
|
||||
option.description = option.description!.replaceAll('✅', 'Pros:').replaceAll('❌', 'Cons:');
|
||||
expect(() => exactCount(saved(briefOnly), briefOnly)).toThrow(/Unsupported/);
|
||||
expect(exactCount(saved(ready), ready)).toBe(true);
|
||||
});
|
||||
|
||||
// The 749df R2 evidence used "punctuation/Unicode" as ordinary prose. This
|
||||
// must not become a foreign source, while actual cited paths remain closed.
|
||||
const withEvidence = (text: string) => exactPlan('### currentDecision (D1)')
|
||||
.replace('Evidence: PLAN.md lines 18-23 state the exact contracts;',
|
||||
`Evidence: PLAN.md lines 18-23 state the exact contracts; ${text};`);
|
||||
for (const compound of ['punctuation/Unicode', 'read/write', 'success/failure', 'input/output', 'request/response'])
|
||||
test(`current native record permits ordinary slash prose ${compound}`, () => {
|
||||
expect(exactCount(withEvidence(`The contract preserves ${compound} behavior`))).toBe(true);
|
||||
});
|
||||
for (const reference of [
|
||||
'other/PLAN.md', 'other/handler.ts', '/PLAN', '/elsewhere/PLAN', './PLAN', '../PLAN', '~/PLAN',
|
||||
'C:\\other\\PLAN', 'C:/other/PLAN', '\\\\host\\share\\PLAN',
|
||||
'`other/PLAN`', '"other/PLAN"', '[source](other/PLAN)', '<other/PLAN>',
|
||||
'Source: other/PLAN', 'file: other/PLAN', 'see other/PLAN', 'according to other/PLAN',
|
||||
'other/PLAN:21', 'other/PLAN#L21',
|
||||
'"read other/PLAN for the current external source contract"',
|
||||
]) test(`slash prose cannot conceal an explicit foreign reference ${reference}`, () => {
|
||||
expect(() => exactCount(withEvidence(`The contract preserves read/write behavior; ${reference}`))).toThrow(/Unsupported/);
|
||||
});
|
||||
@@ -1,78 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
import fixture from './fixtures/ceo-current-omission-ap.json';
|
||||
|
||||
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
|
||||
const first = calls[3]!;
|
||||
const originalClause = 'The plan also does not say whether the email runs inside or after the DB transaction.';
|
||||
function change(edit: (q: any, call: any, fp: any) => void) {
|
||||
const copy = structuredClone(first), call = copy.nativeCall!, q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers?.[q.question]);
|
||||
edit(q, call, copy);
|
||||
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
|
||||
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return copy;
|
||||
}
|
||||
test('the exact failed retry begins review at its current missing transaction contract', () => {
|
||||
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, false, true, false, false, false]);
|
||||
let started = false;
|
||||
const phases = calls.map(fp => { const p = planCountQuestionPhase(fp, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff); started = p.reviewStarted; return p; });
|
||||
expect(fixture.actualCounts).toEqual({ setup: 7, review: 0 });
|
||||
expect(phases.map(p => p.preReview)).toEqual([true, true, true, false, false, false, false]);
|
||||
expect(phases.filter(p => !p.preReview && !p.administrative)).toHaveLength(4);
|
||||
});
|
||||
test('optional also and equivalent present-tense current owners preserve omission meaning', () => {
|
||||
for (const phrase of ['The plan does not say whether', "This plan also doesn't say whether", 'This plan does not say whether'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('The plan also does not say whether', phrase); }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('D2 —', 'D19 —'); }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:'); }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, '"Source: this finding is withdrawn." ' + originalClause); }))).toBe(true);
|
||||
});
|
||||
test('source, conditional and historical declarations cannot supply the missing contract', () => {
|
||||
for (const prefix of ['Source: ', 'If approved, ', 'Previously, ', 'Earlier review assessment: ', 'The following is hypothetical. ']) {
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, prefix + originalClause); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: ' + prefix); }))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['Source:', 'Earlier review assessment:', 'If approved:'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
|
||||
for (const wrapped of ['"' + originalClause + '"', '`' + originalClause + '`', '> ' + originalClause, '```' + originalClause + '```'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace(originalClause, wrapped); }))).toBe(false);
|
||||
for (const replacement of ['The previous plan also did not say whether', 'The plan now says whether', 'The example plan also does not say whether'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question = q.question.replace('The plan also does not say whether', replacement); }))).toBe(false);
|
||||
});
|
||||
test('the current omission and offered amendment must remain in force', () => {
|
||||
for (const status of ['This finding is withdrawn.', 'This issue is "rejected".', 'Correction: this contract is not current.', 'There is no current gap.'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.question += '\n' + status; }))).toBe(false);
|
||||
for (const prefix of ['Source excerpt: ', 'If approved later: ', 'Earlier review assessment: '])
|
||||
expect(ceoFirstReviewAUQ(change(q => { for (const option of q.options) option.description = prefix + option.description; }))).toBe(false);
|
||||
for (const status of ['This amendment is withdrawn.', 'This remedy is "cancelled".'])
|
||||
expect(ceoFirstReviewAUQ(change(q => { for (const option of q.options) option.description += '\n' + status; }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q => { q.options = [{ label: 'A) Keep existing behavior', description: 'Leave the implementation unchanged.' }, { label: 'B) Archive the report', description: 'Save the review text.' }]; }))).toBe(false);
|
||||
});
|
||||
test('a current native completion, selected answer and consistent decision are still required', () => {
|
||||
for (const edit of [
|
||||
(_q: any, c: any) => { c.answered = false; }, (_q: any, c: any) => { c.failed = true; },
|
||||
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; }, (_q: any, c: any) => { delete c.answeredAt; },
|
||||
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; }, (_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(q: any) => { q.multiSelect = true; }, (q: any) => { q.header = 'Approach'; }, (q: any) => { q.header = 'Issue 99'; },
|
||||
(q: any) => { q.question = q.question.replace('D2 —', 'D02 —'); },
|
||||
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: Z'); },
|
||||
(q: any) => { q.question = q.question.replace('Recommendation: A', 'Recommendation: 9A'); for (const option of q.options) option.label = '9' + option.label; },
|
||||
(q: any) => { q.options[1].label = q.options[1].label.replace('B)', '3B)'); },
|
||||
(q: any) => { q.question = q.question.replace('plan-ceo-review-email-rescue', 'plan-ceo-review-setup'); },
|
||||
(q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10 omitted: '); },
|
||||
(q: any) => { q.question = q.question.replace(/\nProject\/branch\/task:[^\n]+/, ''); },
|
||||
]) expect(ceoFirstReviewAUQ(change(edit))).toBe(false);
|
||||
const noAnswer = structuredClone(first); noAnswer.nativeCall!.answers = {}; expect(ceoFirstReviewAUQ(noAnswer)).toBe(false);
|
||||
const menu = structuredClone(first); menu.options[0]!.label = 'Unowned'; expect(ceoFirstReviewAUQ(menu)).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...first, nativeCall: undefined })).toBe(false);
|
||||
});
|
||||
test('only the existing dense CEO finding owner selects the retry regression', () => {
|
||||
for (const path of ['test/ceo-current-omission-ap.test.ts', 'test/fixtures/ceo-current-omission-ap.json'])
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, files]) => files.includes(path)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
|
||||
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
|
||||
for (let i = 0; i < paths.length; i++) { expect(Object.hasOwn(paths, i)).toBe(true); expect(typeof paths[i]).toBe('string'); }
|
||||
});
|
||||
@@ -1,63 +0,0 @@
|
||||
import {expect, test} from 'bun:test';
|
||||
import {ceoFirstReviewAUQ, nativePlanCallFingerprint} from './helpers/claude-pty-runner';
|
||||
import fixture from './fixtures/ceo-decision-prefix-al.json';
|
||||
import {E2E_TOUCHFILES} from './helpers/touchfiles-data';
|
||||
const call=(n=0):any=>structuredClone(fixture.calls[n]);
|
||||
const accepts=(c:any)=>ceoFirstReviewAUQ(nativePlanCallFingerprint(c,0,true));
|
||||
function text(c:any,fn:(s:string)=>string){const q=c.questions[0],a=c.answers[q.question];q.question=fn(q.question);c.answers={[q.question]:a};}
|
||||
function menu(c:any,fn:(o:any,i:number)=>void){const q=c.questions[0],i=q.options.findIndex((o:any)=>o.label===c.answers[q.question]);q.options.forEach(fn);c.answers={[q.question]:q.options[i].label};}
|
||||
test('the actual completed email finding uses decision-prefixed options and a bare recommendation',()=>expect(accepts(call())).toBe(true));
|
||||
test('the actual completed SQL finding includes a raw SQL qualifier',()=>expect(accepts(call(1))).toBe(true));
|
||||
test('decision and finding identifiers remain independent when consistently renamed',()=>{
|
||||
for(const n of [0,1]){
|
||||
for(const dotted of [false,true]){const c=call(n);text(c,s=>s.replace(/^D\d+/,'D27').replace(/\(Finding \d+\)/,`(Finding ${dotted?'8.3':'8'})`));menu(c,o=>{o.label=o.label.replace(/^\d+/,'27')});expect(accepts(c)).toBe(true);}
|
||||
const c=call(n);menu(c,o=>{o.label=o.label.replace(/^\d+/,'')});text(c,s=>s.replace("'no error handling on the email leg'",'no error handling on the email leg').replace('a raw SQL fragment','a SQL fragment'));expect(accepts(c)).toBe(true);
|
||||
const q=call(n);text(q,s=>s+'\nOld note: "This finding is withdrawn."');expect(accepts(q)).toBe(true);
|
||||
const lower=call(n);text(lower,s=>s.replace(/^D/,'d'));expect(accepts(lower)).toBe(true);
|
||||
}
|
||||
});
|
||||
test('native ownership, offered answers and unambiguous decision identities are mandatory',()=>{
|
||||
for(const mutate of [
|
||||
(c:any)=>{c.answered=false},(c:any)=>{c.failed=true},(c:any)=>{c.unansweredQuestionIndices=[0]},(c:any)=>{c.sessionId=''},
|
||||
(c:any)=>{c.answers={}},(c:any)=>{c.answers[c.questions[0].question]='A'},(c:any)=>{c.questions[0].multiSelect=true},
|
||||
(c:any)=>{c.questions[0].header='Finding 9'},(c:any)=>{c.questions[0].header='Approach'},
|
||||
(c:any)=>text(c,s=>s.replace(/^D4/,'D9')),
|
||||
(c:any)=>menu(c,o=>{o.label=o.label.replace(/^4/,'9')}),
|
||||
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4/,'9')}),
|
||||
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4/,'')}),
|
||||
(c:any)=>text(c,s=>s.replace(/^Recommendation: A/m,'Recommendation: 9A')),
|
||||
(c:any)=>text(c,s=>s.replace(/^Recommendation: A/m,'Recommendation: Z')),
|
||||
(c:any)=>menu(c,(o,i)=>{if(i===1)o.label=o.label.replace(/^4B/,'4A')}),
|
||||
(c:any)=>{c.questions[0].options[1].description=''},
|
||||
]){const c=call();mutate(c);expect(accepts(c)).toBe(false);}
|
||||
const f=nativePlanCallFingerprint(call(),0,true);expect(ceoFirstReviewAUQ({...f,signature:'foreign:tool'})).toBe(false);
|
||||
});
|
||||
test('embedded quoted contract terms cannot supply a hypothetical, historical or withdrawn assessment',()=>{
|
||||
for(const n of [0,1])for(const fn of [
|
||||
(s:string)=>'Source excerpt: '+s,(s:string)=>'> '+s,(s:string)=>'```\n'+s+'\n```',
|
||||
(s:string)=>s.replace(/^ELI10: (.+)$/m,'ELI10: "$1"'),
|
||||
(s:string)=>s.replace(/^ELI10: /m,'ELI10: If approved, '),
|
||||
(s:string)=>s.replace(/^ELI10: /m,'ELI10: Source excerpt. '),
|
||||
(s:string)=>s.replace(/^ELI10: /m,'ELI10: The following is a historical source excerpt. '),
|
||||
(s:string)=>s.replace(/^ELI10: /m,'ELI10: Previously, '),
|
||||
(s:string)=>s+'\nThis finding is withdrawn.',
|
||||
(s:string)=>s+'\nNo current defect remains.',
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan sends the email inline with \'no error handling\' only in a historical example.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan does not send the email inline with \'no error handling\'.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan used to paste the user ID straight into a raw SQL fragment. The current query is parameterized.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: A proposed example pastes the user ID string straight into a raw SQL fragment.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The historical example pastes the user ID string straight into a raw SQL fragment.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: A template pastes the user ID string straight into a raw SQL fragment.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: An unrelated example pastes the user ID string straight into a raw SQL fragment.'),
|
||||
(s:string)=>s.replace(/^ELI10: .+$/m,'ELI10: The plan pastes the user ID string straight into a raw SQL fragment only in a hypothetical example.'),
|
||||
]){const c=call(n);text(c,fn);expect(accepts(c)).toBe(false);}
|
||||
});
|
||||
test('a substantive current assessment still needs an offered technical amendment',()=>{
|
||||
for(const description of ['Archive this report.','If approved: ✅ Rescue named mail exceptions.','Source excerpt: ✅ Rescue named mail exceptions.','❌ Rescue named mail exceptions.','✅ "Rescue named mail exceptions."']){
|
||||
const c=call();menu(c,(o,i)=>{o.label=`4${String.fromCharCode(65+i)}: Consider candidate ${i}`;o.description=description});expect(accepts(c)).toBe(false);
|
||||
}
|
||||
});
|
||||
test('only the existing CEO count owner selects the captured regression',()=>{
|
||||
for(const dependency of E2E_TOUCHFILES['plan-ceo-finding-count']) expect(typeof dependency).toBe('string');
|
||||
for(const d of ['test/ceo-decision-prefix-al.test.ts','test/fixtures/ceo-decision-prefix-al.json'])expect(Object.entries(E2E_TOUCHFILES).filter(([,v])=>v.includes(d)).map(([k])=>k)).toEqual(['plan-ceo-finding-count']);
|
||||
});
|
||||
@@ -1,111 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import { isCeoCompletionHandoff } from './helpers/ceo-completion-handoff';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
import fixture from './fixtures/ceo-declarative-premise-ap.json';
|
||||
|
||||
const calls = fixture.fingerprints as AskUserQuestionFingerprint[];
|
||||
const first = calls[2]!;
|
||||
function change(fp: AskUserQuestionFingerprint, edit: (q: any, call: any, fp: any) => void) {
|
||||
const copy = structuredClone(fp), call = copy.nativeCall!, q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers?.[q.question]);
|
||||
edit(q, call, copy);
|
||||
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
|
||||
copy.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return copy;
|
||||
}
|
||||
|
||||
test('the exact completed defect premises start review; later calls use unchanged phase continuation', () => {
|
||||
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false, false, true, true, false, false]);
|
||||
let started = false;
|
||||
const phases = calls.map(fp => {
|
||||
const phase = planCountQuestionPhase(fp, started, ceoStep0Boundary, ceoFirstReviewAUQ, undefined, isCeoCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(fixture.actualCounts).toEqual({ setup: 6, review: 0 });
|
||||
expect(phases.map(p => p.preReview)).toEqual([true, true, false, false, false, false]);
|
||||
expect(phases.filter(p => !p.preReview && !p.administrative)).toHaveLength(4);
|
||||
});
|
||||
|
||||
test('current metadata, premise and explanation cannot borrow quoted, historical or conditional authority', () => {
|
||||
for (const fp of calls.slice(2, 4)) {
|
||||
for (const prefix of ['Source:', 'Earlier review assessment:', 'If approved:', 'Example:'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\n' + prefix + '\nELI10:'); }))).toBe(false);
|
||||
for (const prefix of ['If approved, ', 'Source excerpt: ', 'Earlier review assessment: ', 'The following is a hypothetical example. ', 'Previously, ', 'Formerly, ']) {
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('ELI10: ', 'ELI10: ' + prefix); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: ' + prefix); }))).toBe(false);
|
||||
}
|
||||
for (const replacement of ['Source: The ', 'If approved, the ', 'The historical ', 'The quoted ', 'The previously '])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('— The ', '— ' + replacement); }))).toBe(false);
|
||||
for (const wrap of [(s: string) => `"${s}"`, (s: string) => '`' + s + '`', (s: string) => '> ' + s])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { const title = q.question.split('\n')[0]; q.question = q.question.replace(title, wrap(title)); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace(/\nProject\/branch\/task:[^\n]+/, ''); }))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('a current finding and its offered amendments cannot be withdrawn', () => {
|
||||
for (const fp of calls.slice(2, 4)) {
|
||||
for (const status of ['This finding is withdrawn.', 'This issue is "rejected".', 'Correction: this assessment is not current.', 'There is no current gap.'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question += '\n' + status; }))).toBe(false);
|
||||
for (const prefix of ['Source excerpt: ', 'If approved later: ', 'Previously, ', 'Formerly, '])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description = prefix + option.description; }))).toBe(false);
|
||||
for (const status of ['This amendment is withdrawn.', 'This remedy is "cancelled".'])
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description += '\n' + status; }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { for (const option of q.options) option.description = JSON.stringify(option.description); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(fp, q => { q.question = q.question.replace('\nELI10:', '\nArchive note: "Source: this finding is withdrawn."\nELI10:'); }))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('statement-only and administrative menus are not review decisions', () => {
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace(' How should the handler treat a mail failure?', ''); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace(' How should the handler treat a mail failure?', ' Record this in the report.'); }))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(first, q => {
|
||||
q.options = [
|
||||
{ label: '2A) Keep the existing implementation', description: 'Leave current behavior unchanged.' },
|
||||
{ label: '2B) Archive the report', description: 'Save the existing review text.' },
|
||||
];
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.header = 'Approach'; }))).toBe(false);
|
||||
});
|
||||
|
||||
test('the native completion, selected answer and issue identities stay bound', () => {
|
||||
for (const edit of [
|
||||
(_q: any, c: any) => { c.answered = false; },
|
||||
(_q: any, c: any) => { c.failed = true; },
|
||||
(_q: any, c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(_q: any, c: any) => { delete c.answeredAt; },
|
||||
(_q: any, c: any) => { c.answeredAt = 'not-a-time'; },
|
||||
(_q: any, _c: any, f: any) => { f.signature = 'foreign:call'; },
|
||||
(_q: any, _c: any, f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(q: any) => { q.multiSelect = true; },
|
||||
(q: any) => { q.header = 'Issue 99'; },
|
||||
(q: any) => { q.header = 'Issue 0'; },
|
||||
(q: any) => { q.header = 'Issue 02'; },
|
||||
(q: any) => { q.question = q.question.replace('(Issue 2)', '(Issue 0)'); },
|
||||
(q: any) => { q.question = q.question.replace('D4 (', 'D04 ('); },
|
||||
(q: any) => { q.question = q.question.replace('Recommendation: 2A', 'Recommendation: 9A'); },
|
||||
(q: any) => { q.options[1].label = q.options[1].label.replace('2B)', '3B)'); },
|
||||
]) expect(ceoFirstReviewAUQ(change(first, edit))).toBe(false);
|
||||
const wrongAnswer = structuredClone(first); wrongAnswer.nativeCall!.answers = {};
|
||||
expect(ceoFirstReviewAUQ(wrongAnswer)).toBe(false);
|
||||
const wrongMenu = structuredClone(first); wrongMenu.options[0]!.label = 'Foreign';
|
||||
expect(ceoFirstReviewAUQ(wrongMenu)).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...first, nativeCall: undefined })).toBe(false);
|
||||
});
|
||||
|
||||
test('equivalent current wording and descriptive or matching issue headers preserve the decision', () => {
|
||||
for (const header of ['Email contract', 'Issue 2', 'Finding 2'])
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.header = header; }))).toBe(true);
|
||||
for (const separator of ['—', '–', '-'])
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace('D4 (Issue 2) —', `D19 (Issue 2) ${separator}`); }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(first, q => { q.question = q.question.replace('How should the handler treat a mail failure?', 'Which handling should the current implementation use?'); }))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(calls[3]!, q => { q.question = q.question.replace('request.params.userId', 'payload.accountId'); }))).toBe(true);
|
||||
});
|
||||
|
||||
test('the regression fixture is registered only to the dense CEO finding owner', () => {
|
||||
for (const path of ['test/ceo-declarative-premise-ap.test.ts', 'test/fixtures/ceo-declarative-premise-ap.json'])
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, files]) => files.includes(path)).map(([owner]) => owner)).toEqual(['plan-ceo-finding-count']);
|
||||
const paths = E2E_TOUCHFILES['plan-ceo-finding-count']!;
|
||||
for (let i = 0; i < paths.length; i++) { expect(Object.hasOwn(paths, i)).toBe(true); expect(typeof paths[i]).toBe('string'); }
|
||||
});
|
||||
@@ -1,129 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import captured from './fixtures/ceo-finding-brief-ak.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
const call = (index = 4): any => structuredClone(captured.calls[index]);
|
||||
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
|
||||
function edit(c: any, change: (s: string) => string) {
|
||||
const q = c.questions[0], answer = c.answers[q.question];
|
||||
q.question = change(q.question); c.answers = { [q.question]: answer };
|
||||
}
|
||||
function offered(c: any, change: (o: any, i: number) => void) {
|
||||
const q = c.questions[0], selected = q.options.findIndex((o: any) => o.label === c.answers[q.question]);
|
||||
q.options.forEach(change); c.answers = { [q.question]: q.options[selected].label };
|
||||
}
|
||||
test('the completed parenthesized finding with letter-only choices starts current review', () => {
|
||||
expect(ceoFirstReviewAUQ(fp(call()))).toBe(true);
|
||||
});
|
||||
test('the exact retry phase preserves four setup calls then six substantive choices', () => {
|
||||
let started = false;
|
||||
const phases = captured.calls.map(c => {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted; return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, true, true, false, false, false, false, false, false]);
|
||||
});
|
||||
|
||||
test('the finding identity is independent of decision number, separator and optional qid', () => {
|
||||
for (const change of [
|
||||
(s: string) => s.replace(/^D5/, 'D19'),
|
||||
(s: string) => s.replace(') — ', ') - '),
|
||||
(s: string) => s.replace('Finding 1.1', 'Finding 9.4'),
|
||||
(s: string) => s.replace('Finding 1.1', 'Finding 1'),
|
||||
(s: string) => s.replace(/\s*<gstack-qid:[^>]+>\s*$/, ''),
|
||||
(s: string) => s.replace(/\s*<gstack-qid:[^>]+>\s*$/, '') + '\n<gstack-qid:plan-ceo-review-finding-mail>',
|
||||
(s: string) => s.replace('lets any mail failure', 'allows any mail failure'),
|
||||
]) { const c = call(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(true); }
|
||||
for (const label of call().questions[0].options.map((o: any) => o.label)) {
|
||||
const c = call(); c.answers[c.questions[0].question] = label;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
const numbered = call(); edit(numbered, s => s.replace(/^Recommendation: A/m, 'Recommendation: 1A'));
|
||||
offered(numbered, o => { o.label = o.label.replace(/^([A-Z])\)/, '1$1)'); });
|
||||
expect(ceoFirstReviewAUQ(fp(numbered))).toBe(true);
|
||||
const header = call(); header.questions[0].header = 'Finding 1.1';
|
||||
expect(ceoFirstReviewAUQ(fp(header))).toBe(true);
|
||||
});
|
||||
|
||||
test('competing finding, section, recommendation and offered choice identities are rejected', () => {
|
||||
for (const mutate of [
|
||||
(c: any) => { c.questions[0].header = 'Finding 1'; },
|
||||
(c: any) => { c.questions[0].header = 'Finding 9.1'; },
|
||||
(c: any) => { c.questions[0].header = 'Approach'; },
|
||||
(c: any) => edit(c, s => s.replace('Finding 1.1', 'Finding 1.0')),
|
||||
(c: any) => edit(c, s => s.replace('Finding 1.1', 'Finding 1.1 and Finding 2.1')),
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: A/m, 'Recommendation: 2A')),
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: A/m, 'Recommendation: Z')),
|
||||
(c: any) => { c.questions[0].options[0].label = '9A) Foreign issue'; c.answers = { [c.questions[0].question]: c.questions[0].options[0].label }; },
|
||||
(c: any) => { c.questions[0].options[1].label = 'A) Same choice letter, different action'; },
|
||||
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[0].label; },
|
||||
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-eng-review-finding-mail>'),
|
||||
]) { const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
});
|
||||
|
||||
test('a current complete assessment cannot come from source, conditions or a withdrawal', () => {
|
||||
for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```\n' + s + '\n```',
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, ''),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If approved, '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a quoted source excerpt. '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
|
||||
(s: string) => s + '\nThis finding is withdrawn.',
|
||||
(s: string) => s + '\nFinding 1.1 is rejected.',
|
||||
(s: string) => s + '\nNo current defect remains.',
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan no longer lets mail failures escape the handler. The current named rescue keeps them contained.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan used to let mail failures escape the handler. That was the prior behavior.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan does not let mail failures escape the handler.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Source excerpt: the plan lets mail failures escape the handler.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Previously, the plan lets mail failures escape the handler.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan lets no mail failure escape the handler.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan allows mail failures to escape only in a historical quoted example.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: Source excerpt. The plan lets mail failures escape the handler.'),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The plan allows mail failures to never escape the handler.'),
|
||||
]) { const c = call(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
const c = call(); edit(c, s => s + '\nOld note: "Finding 1.1 is rejected."');
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
});
|
||||
|
||||
test('current finding prose must offer an actual remedy, not an advisory or hypothetical action', () => {
|
||||
for (const description of [
|
||||
'Archive this review for later.',
|
||||
'Historical source excerpt: ✅ Rescue named mail exceptions.',
|
||||
'The following is a quoted source excerpt. ✅ Rescue named mail exceptions.',
|
||||
'If approved: ✅ Rescue named mail exceptions.',
|
||||
'❌ Rescue named mail exceptions.',
|
||||
'✅ "Rescue named mail exceptions."',
|
||||
'✅ If approved, rescue named mail exceptions.',
|
||||
]) {
|
||||
const c = call(); offered(c, (o, i) => { o.label = `${String.fromCharCode(65 + i)}) Consider candidate ${i}`; o.description = description; });
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the completed native call, exact offered answer and fingerprint remain mandatory', () => {
|
||||
for (const mutate of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.sessionId = ''; },
|
||||
(c: any) => { c.toolUseId = ''; },
|
||||
(c: any) => { c.answers = {}; },
|
||||
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
|
||||
(c: any) => { c.questions[0].multiSelect = true; },
|
||||
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
|
||||
(c: any) => { c.questions[0].options[1].description = ''; },
|
||||
]) { const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
const f = fp(call());
|
||||
expect(ceoFirstReviewAUQ({ ...f, signature: 'foreign:tool' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...f, nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...f, options: f.options.slice(1) })).toBe(false);
|
||||
});
|
||||
|
||||
test('retry fixture and controls select only the existing CEO count owner', () => {
|
||||
for (const dependency of ['test/ceo-finding-brief-ak.test.ts', 'test/fixtures/ceo-finding-brief-ak.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name)).toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -378,100 +378,3 @@ describe('CEO finding fixture establishes scope before launch', () => {
|
||||
} finally { fs.rmSync(root, { recursive: true, force: true }); }
|
||||
});
|
||||
});
|
||||
|
||||
// Main owns both distinct and paired registrations in this file. Select the
|
||||
// actual case and replace only its native count boundary; report/band checks
|
||||
// and the output-directory finally stay live.
|
||||
test.each(['success5', 'success7', 'success-paired', 'below', 'above', 'missing-report', 'trailing-report', 'timeout', 'throw', 'native-error', 'unknown-current'])('native count registration: %s', scenario => {
|
||||
const root = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'ceo-count-body-')));
|
||||
const script = path.join(root, 'registration.test.ts');
|
||||
const factsPath = path.join(root, 'facts.json');
|
||||
fs.writeFileSync(script, `
|
||||
import {describe, expect, mock} from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import {execFileSync} from 'node:child_process';
|
||||
import * as runner from ${JSON.stringify(path.join(ROOT, 'test/helpers/claude-pty-runner.ts'))};
|
||||
import {createPlanCountFixture} from ${JSON.stringify(path.join(ROOT, 'test/helpers/plan-count-fixture.ts'))};
|
||||
const captured = JSON.parse(fs.readFileSync(${JSON.stringify(path.join(ROOT, 'test/fixtures/ceo-payment-ledger-decisions.json'))}, 'utf8'));
|
||||
const original = {...runner}, scenario = ${JSON.stringify(scenario)}, paired = scenario === 'success-paired';
|
||||
let calls = 0;
|
||||
mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/e2e-gate.ts'))}, () => ({describeE2ETier:tier=>{expect(tier).toBe('periodic');return describe;}}));
|
||||
mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/claude-pty-runner.ts'))}, () => ({...original,
|
||||
runPlanSkillCounting:async opts=>{
|
||||
calls++;
|
||||
const target=opts.expectedPlanPath;
|
||||
const facts={calls,target,validated:false};
|
||||
fs.writeFileSync(${JSON.stringify(factsPath)},JSON.stringify(facts));
|
||||
expect(path.dirname(path.dirname(target))).toBe(${JSON.stringify(root)});
|
||||
expect(opts.cwd).toBeUndefined();
|
||||
expect(opts.followUpPrompt).toContain(target);
|
||||
expect(opts.followUpPrompt).toContain('in HOLD SCOPE mode');
|
||||
expect(opts.followUpPrompt).toContain('skip the optional /office-hours prerequisite');
|
||||
expect(opts).toMatchObject({skillName:'plan-ceo-review',slashCommand:'/plan-ceo-review',
|
||||
reviewCountCeiling:paired?5:8,timeoutMs:1500000,env:{QUESTION_TUNING:'false',EXPLAIN_LEVEL:'default'}});
|
||||
for(const key of ['isLastStep0AUQ','isFirstReviewAUQ','isCompletionHandoffAUQ','pickAUQ'])expect(typeof opts[key]).toBe('function');
|
||||
const required=paired?[
|
||||
'assert only','that the returned receipt is truthy','No assertion about the mock call history or virtual sleeper record',
|
||||
'max_retries=1 means two total charge attempts',
|
||||
]:[
|
||||
'bypasses the existing \\x60WebhookDispatcher\\x60','directly into a raw SQL','no error handling on the email leg',
|
||||
"None planned. We'll rely on the existing integration suite catching regressions.",'order in a loop',
|
||||
];
|
||||
for(const finding of required)expect(opts.followUpPrompt).toContain(finding);
|
||||
if(!paired)expect(opts.firstAUQPick({options:[{index:1,label:'Branch diff vs main'},{index:7,label:'Skip interview and plan immediately'}]})).toBe(7);
|
||||
const fixture=createPlanCountFixture(opts.followUpPrompt,{files:opts.fixtureFiles});
|
||||
try {
|
||||
const committed=execFileSync('git',['show','HEAD:PLAN.md'],{cwd:fixture.cwd,encoding:'utf8',timeout:5000});
|
||||
expect(committed).toBe(opts.followUpPrompt);
|
||||
expect(fs.readFileSync(path.join(fixture.cwd,'CLAUDE.md'),'utf8')).toContain(committed);
|
||||
} finally {fixture.cleanup();}
|
||||
facts.validated=true;fs.writeFileSync(${JSON.stringify(factsPath)},JSON.stringify(facts));
|
||||
if(scenario==='throw')throw new Error('controlled count observation failure');
|
||||
if(!paired){
|
||||
expect(typeof opts.isReviewAUQ).toBe('function');
|
||||
const prior=[];
|
||||
for(const [index,item] of captured.captures.entries()){
|
||||
if(item.savedPlan)fs.writeFileSync(target,item.savedPlan);
|
||||
const call=structuredClone(item.call);
|
||||
const fp=original.nativePlanCallFingerprint(call,index,true);
|
||||
expect(opts.isReviewAUQ(fp,prior)).toBe(item.kind==='seeded-remedy'||item.call.questions[0].header==='TODO-1');
|
||||
prior.push(call);
|
||||
}
|
||||
if(scenario==='unknown-current'){
|
||||
const call=structuredClone(captured.captures[2].call),q=call.questions[0];
|
||||
q.question='D99 — Should we change the billing currency?';call.answers={[q.question]:q.options[0].label};call.toolUseId+='-extra';
|
||||
opts.isReviewAUQ(original.nativePlanCallFingerprint(call,99,true),prior);
|
||||
}
|
||||
}
|
||||
if(scenario==='missing-report')fs.rmSync(target,{force:true});
|
||||
if(scenario!=='missing-report')fs.writeFileSync(target,'# Reviewed plan\\n\\n## GSTACK REVIEW REPORT\\nVERDICT: APPROVED\\n'+(scenario==='trailing-report'?'\\n## Unreviewed tail\\n':''));
|
||||
return {outcome:scenario==='timeout'?'timeout':scenario==='native-error'?'transcript_unavailable':'plan_ready',
|
||||
reviewCount:{success5:5,success7:7,'success-paired':2,below:3,above:8}[scenario]??5,
|
||||
step0Count:2,elapsedMs:1000,fingerprints:[],evidence:'controlled native observation'};
|
||||
},
|
||||
}));
|
||||
await import(${JSON.stringify(path.join(ROOT, 'test/skill-e2e-plan-ceo-finding-count.test.ts'))});
|
||||
`);
|
||||
try {
|
||||
const child = spawnSync(process.execPath, ['test', script, '--test-name-pattern', scenario === 'success-paired' ? 'paired-finding positive control' : '5-finding plan'], {
|
||||
cwd: ROOT, encoding: 'utf8', timeout: 10_000,
|
||||
env: {PATH:process.env.PATH ?? '', HOME:root,TMPDIR:root,TEMP:root,TMP:root,GIT_CONFIG_NOSYSTEM:'1',
|
||||
...(process.env.SystemRoot ? {SystemRoot:process.env.SystemRoot} : {})},
|
||||
});
|
||||
const output=child.stdout+child.stderr;
|
||||
expect(child.error,output).toBeUndefined();
|
||||
const facts=JSON.parse(fs.readFileSync(factsPath,'utf8'));
|
||||
expect(facts.calls).toBe(1);
|
||||
expect(facts.validated,output).toBe(true);
|
||||
expect(fs.existsSync(path.dirname(facts.target)),'actual paid finally removes its owned output directory').toBe(false);
|
||||
expect(child.status,output).toBe(scenario.startsWith('success')?0:1);
|
||||
const failures:Record<string,string>={below:'BAND FAIL (below floor)',above:'BAND FAIL (above ceiling)',
|
||||
'missing-report':'D19 FAIL: agent did not produce expected plan file',
|
||||
'trailing-report':'trailing ## heading(s) after GSTACK REVIEW REPORT',
|
||||
timeout:'finding-count FAILED: outcome=timeout',throw:'controlled count observation failure',
|
||||
'native-error':'finding-count FAILED: outcome=transcript_unavailable',
|
||||
'unknown-current':'cannot exclude it from the 4–7 count'};
|
||||
if(failures[scenario])expect(output).toContain(failures[scenario]);
|
||||
} finally {fs.rmSync(root,{recursive:true,force:true});}
|
||||
},20_000);
|
||||
@@ -3,40 +3,10 @@ import fs from 'node:fs';
|
||||
import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import fixture from './fixtures/ceo-handoff-y-call.json';
|
||||
import zFixture from './fixtures/ceo-handoff-z-call.json';
|
||||
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import {ceoFirstReviewAUQ,ceoStep0Boundary,hasNativePlanTerminal,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
|
||||
import {isCeoCompletionHandoff,pickCeoCompletionHandoff} from './helpers/ceo-completion-handoff';
|
||||
const actual=()=>structuredClone(fixture.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
import {hasNativePlanTerminal,nativePlanCallFingerprint} from './helpers/claude-pty-runner';
|
||||
const fp=(c:NativePlanQuestionCall)=>nativePlanCallFingerprint(c,0,false);
|
||||
const pending=(c:NativePlanQuestionCall)=>{c.answered=false;delete c.answers;delete c.unansweredQuestionIndices;return fp(c);};
|
||||
function change(c:NativePlanQuestionCall,fn:(s:string)=>string){const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};return c;}
|
||||
|
||||
describe('Y bare next-Eng navigation is administrative, not completion evidence',()=>{
|
||||
test('exact four issues remain while a closed next-workflow menu cannot start review',()=>{
|
||||
const c=actual();expect(isCeoCompletionHandoff(fp(c))).toBe(true);expect(pickCeoCompletionHandoff(pending(actual()))).toBe(2);
|
||||
expect(pickCeoCompletionHandoff(fp(c))).toBeNull();expect(c.answers![c.questions[0]!.question]).toBe('A) Run /plan-eng-review next (recommended)');
|
||||
let started=false;let setup=0,review=0,admin=0;
|
||||
for(const c of fixture.calls){const p=planCountQuestionPhase(fp(structuredClone(c) as NativePlanQuestionCall),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=p.reviewStarted;if(p.administrative)admin++;else if(p.preReview)setup++;else review++;}
|
||||
expect({setup,review,admin}).toEqual({setup:2,review:4,admin:1});
|
||||
expect(planCountQuestionPhase(fp(actual()),false,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff)).toEqual({preReview:false,reviewStarted:false,administrative:'completion-handoff'});
|
||||
});
|
||||
test('either offered navigation answer and option order preserve administrative meaning',()=>{
|
||||
const c=actual();c.questions[0]!.options.reverse();
|
||||
for(const o of c.questions[0]!.options){c.answers={[c.questions[0]!.question]:o.label};expect(isCeoCompletionHandoff(fp(c))).toBe(true);}
|
||||
expect(pickCeoCompletionHandoff(pending(c))).toBe(1);
|
||||
expect(isCeoCompletionHandoff(fp(change(actual(),s=>s.replace('D7 - Next step: run','D17 — Next review: Run').replace('plan-ceo-review-next-step','plan-ceo-review-next-review'))))).toBe(true);
|
||||
});
|
||||
test('whole question and description boundaries reject added product work and unfinished choices',()=>{
|
||||
for(const fn of [(s:string)=>s.replace('run /plan-eng-review?', 'fix the cache before /plan-eng-review?'),(s:string)=>s.replace('run /plan-eng-review?', 'run /plan-eng-review? Also repair the cache.'),(s:string)=>'> '+s,(s:string)=>'Example: '+s,(s:string)=>s.replace('plan-ceo-review-next-step','foreign-next-step'),(s:string)=>s+' <gstack-qid:plan-ceo-review-next-step>'])expect(isCeoCompletionHandoff(fp(change(actual(),fn)))).toBe(false);
|
||||
for(const i of [0,1])for(const extra of [' Also implement a new cache.',' Resolve the remaining CEO decisions first.',' Should we add another requirement?']){const c=actual();c.questions[0]!.options[i]!.description+=extra;expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
|
||||
for(const text of ['Resume the unfinished CEO review.','Proceed directly to implementation and add the missing test.','Eng review is optional.']){const c=actual();c.questions[0]!.options[1]!.description=text;expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
|
||||
});
|
||||
test('native completion, current offered answer and pending identity remain required',()=>{
|
||||
for(const mutate of [(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions[0]!.header='Issue';},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'Fix another issue'};}]){const c=actual();mutate(c);expect(isCeoCompletionHandoff(fp(c))).toBe(false);}
|
||||
expect(isCeoCompletionHandoff({...fp(actual()),signature:'foreign:call'})).toBe(false);expect(isCeoCompletionHandoff({...fp(actual()),options:[]})).toBe(false);
|
||||
expect(pickCeoCompletionHandoff({...pending(actual()),nativeCall:undefined})).toBeNull();expect(pickCeoCompletionHandoff({...pending(actual()),signature:'foreign:call'})).toBeNull();
|
||||
});
|
||||
test('independent fresh report and native Exit still gate completion; menu alone cannot pass',()=>{
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-handoff-y-free-'));const report=path.join(dir,'report.md');
|
||||
try{fs.writeFileSync(report,fixture.report);const calls=structuredClone(fixture.calls) as NativePlanQuestionCall[];const transcript={status:'ready' as const,calls,assistantMessages:[],planReadyRequests:structuredClone(fixture.planReadyRequests)};const handoff=calls.at(-1)!;const admin=new Set([fp(handoff).signature]);const issueAt=Date.parse(calls.at(-2)!.answeredAt!),handoffAt=Date.parse(handoff.answeredAt!);const started=Date.parse(calls[0]!.answeredAt!)-1000;
|
||||
@@ -49,43 +19,3 @@ describe('Y bare next-Eng navigation is administrative, not completion evidence'
|
||||
}finally{fs.rmSync(dir,{recursive:true,force:true});}
|
||||
});
|
||||
});
|
||||
|
||||
describe('Z completed CEO with an unrun required Eng gate',()=>{
|
||||
const actualZ=()=>structuredClone(zFixture.calls.at(-1)!) as NativePlanQuestionCall;
|
||||
const pendingZ=(c=actualZ())=>{c.answered=false;delete c.answers;delete c.answeredAt;delete c.unansweredQuestionIndices;return c;};
|
||||
const reject=(c:NativePlanQuestionCall)=>{expect(isCeoCompletionHandoff(fp(c))).toBe(false);expect(pickCeoCompletionHandoff(fp(c))).toBeNull();};
|
||||
test('exact six calls preserve two findings; handoff selects the offered manual route',()=>{
|
||||
let started=false;const counts={setup:0,review:0,admin:0};
|
||||
for(const c of zFixture.calls){const phase=planCountQuestionPhase(fp(c as NativePlanQuestionCall),started,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff);started=phase.reviewStarted;counts[phase.administrative?'admin':phase.preReview?'setup':'review']++;}
|
||||
expect(counts).toEqual({setup:3,review:2,admin:1});expect(isCeoCompletionHandoff(fp(actualZ()))).toBe(true);expect(pickCeoCompletionHandoff(fp(pendingZ()))).toBe(2);
|
||||
expect(planCountQuestionPhase(fp(actualZ()),false,ceoStep0Boundary,ceoFirstReviewAUQ,undefined,isCeoCompletionHandoff)).toEqual({preReview:false,reviewStarted:false,administrative:'completion-handoff'});
|
||||
});
|
||||
test('number, typography and option order are not semantic requirements',()=>{
|
||||
const c=change(actualZ(),s=>s.replace('D5 —','D27:').replace("hasn't",'has not').replace("What's",'What is'));c.questions[0]!.options.reverse();
|
||||
for(const option of c.questions[0]!.options){c.answers={[c.questions[0]!.question]:option.label};expect(isCeoCompletionHandoff(fp(c))).toBe(true);}
|
||||
expect(pickCeoCompletionHandoff(fp(pendingZ(c)))).toBe(1);
|
||||
});
|
||||
test('whole question and role-specific descriptions cannot hide new or conditional work',()=>{
|
||||
for(const fn of [(s:string)=>s.replace('CEO Review is CLEAR','If CEO Review is CLEAR'),(s:string)=>s.replace('CEO Review is CLEAR','CEO Review is not CLEAR'),(s:string)=>s.replace('required shipping gate','optional shipping gate'),(s:string)=>s.replace("What's next?","What's next? Also add retries."),(s:string)=>'> '+s,(s:string)=>'Example: '+s,(s:string)=>'```\n'+s+'\n```',(s:string)=>s.replace('plan-ceo-next-review','foreign-next-review'),(s:string)=>s+' <gstack-qid:plan-ceo-next-review>'])reject(change(actualZ(),fn));
|
||||
for(const i of [0,1])for(const extra of [' Also implement the missing checks.',' Rotate credentials.',' Should we add a new requirement?',' Once remaining findings are fixed.']){const c=actualZ();c.questions[0]!.options[i]!.description+=extra;reject(c);}
|
||||
const swapped=actualZ();[swapped.questions[0]!.options[0]!.description,swapped.questions[0]!.options[1]!.description]=[swapped.questions[0]!.options[1]!.description,swapped.questions[0]!.options[0]!.description];reject(swapped);
|
||||
const optional=actualZ();optional.questions[0]!.options[1]!.description=optional.questions[0]!.options[1]!.description!.replace('required before shipping','optional before shipping');reject(optional);
|
||||
});
|
||||
test('new arm requires explicit native completion and exact producer pending state',()=>{
|
||||
const mutations=[(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.questions[0]!.header='Issue';},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.options.push({label:'Add a repair',description:'Add a new requirement.'});}];
|
||||
for(const mutate of mutations){const c=actualZ();mutate(c);reject(c);const p=pendingZ();mutate(p);reject(p);}
|
||||
for(const mutate of [(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'Repair first'};}]){const c=actualZ();mutate(c);reject(c);}
|
||||
for(const mutate of [(c:NativePlanQuestionCall)=>{delete (c as any).answered;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[];},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[1];},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answeredAt='2026-09-09T12:00:00Z';}]){const c=pendingZ();mutate(c);reject(c);}
|
||||
const projected=pendingZ();projected.unansweredQuestionIndices=[0];expect(pickCeoCompletionHandoff(fp(projected))).toBe(2);
|
||||
for(const variant of [{...fp(pendingZ()),signature:'foreign:call'},{...fp(pendingZ()),options:[]},{...fp(pendingZ()),nativeQuestionIndex:1}])expect(pickCeoCompletionHandoff(variant)).toBeNull();
|
||||
});
|
||||
test('retained original mtime passes only with the administrative handoff and real Exit',()=>{
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'gstack-handoff-z-free-'));const report=path.join(dir,'report.md');
|
||||
try{fs.writeFileSync(report,zFixture.report);const calls=structuredClone(zFixture.calls) as NativePlanQuestionCall[];const transcript={status:'ready' as const,calls,assistantMessages:[],planReadyRequests:structuredClone(zFixture.planReadyRequests)};const admin=new Set(calls.filter(c=>isCeoCompletionHandoff(fp(c))).map(c=>fp(c).signature));const mtime=Number(BigInt(zFixture.reportOriginalMtimeNs))/1e6;fs.utimesSync(report,mtime/1000,mtime/1000);
|
||||
expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready')).toBe(false);expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready',admin)).toBe(true);
|
||||
expect(hasNativePlanTerminal({...transcript,planReadyRequests:[]},report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
|
||||
expect(hasNativePlanTerminal({...transcript,calls:[calls.at(-1)!]},report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
|
||||
const stale=Date.parse(calls.at(-2)!.answeredAt!)-1;fs.utimesSync(report,stale/1000,stale/1000);expect(hasNativePlanTerminal(transcript,report,zFixture.startedAt,'plan_ready',admin)).toBe(false);
|
||||
}finally{fs.rmSync(dir,{recursive:true,force:true});}
|
||||
});
|
||||
});
|
||||
@@ -1,52 +0,0 @@
|
||||
/** Free replay only. Both actual paid failures remain rejected; completions are synthetic. */
|
||||
import { test, expect } from 'bun:test';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { createCeoPaymentFindingCounter, ceoPaymentFinding } from './helpers/ceo-payment-findings';
|
||||
import { nativePlanCallFingerprint, ceoFirstReviewAUQ } from './helpers/claude-pty-runner';
|
||||
import fixture from './fixtures/ceo-incomplete-save-b176.json';
|
||||
|
||||
const sha = (value: string) => createHash('sha256').update(value).digest('hex');
|
||||
const replaceOnce = (value: string, from: string, to: string) => {
|
||||
expect(value.split(from)).toHaveLength(2);
|
||||
return value.replace(from, to);
|
||||
};
|
||||
for (const [attemptIndex, capture] of fixture.captures.entries()) {
|
||||
const addSavedNativeFacts = (plan: string, omitLastCons = false) => {
|
||||
const paragraphs = capture.call.questions[0]!.options.map((option, i) => {
|
||||
// Only saved formatting is synthetic. Facts come from actual native descriptions;
|
||||
// effort S / risk low are already present in the original saved comparison.
|
||||
const label = option.label.replace(/^[A-D][.):]\s*/i, '').replace(/\s*\((?:recommended|as planned)\)$/i, '');
|
||||
const [pros, ...cons] = option.description!.split('❌');
|
||||
expect(pros).toContain('✅'); expect(cons).toHaveLength(1);
|
||||
return `**${String.fromCharCode(65 + i)}) ${label}.** Effort S. Risk low. Pros: ${pros!.replaceAll('✅', '').trim()}` +
|
||||
(omitLastCons && i === 2 ? '' : ` Cons: ${cons[0]!.trim()}`);
|
||||
}).join('\n\n');
|
||||
return replaceOnce(plan, '### R2 commitment comparison', paragraphs + '\n\n### R2 commitment comparison');
|
||||
};
|
||||
const citeSource = (plan: string) => plan.replace(/^(\| R1[^|]+\|\s*)([^|]+)(\|)/m,
|
||||
(whole, prefix, evidence, end) => evidence.includes('PLAN.md') ? whole : prefix + 'PLAN.md: ' + evidence + end);
|
||||
const scenarios = [
|
||||
{ name: 'actual incomplete save stays rejected', expected: 'Unsupported', plan: () => capture.savedPlan },
|
||||
{ name: 'synthetic full facts still require row source', expected: attemptIndex === 0 ? 'recorded' : 'Unsupported', plan: () => addSavedNativeFacts(capture.savedPlan) },
|
||||
{ name: 'synthetic source alone cannot replace full facts', expected: 'Unsupported', plan: () => citeSource(capture.savedPlan) },
|
||||
{ name: 'synthetic complete facts and source count the same R1', expected: 'recorded', plan: () => addSavedNativeFacts(citeSource(capture.savedPlan)) },
|
||||
{ name: 'synthetic missing con stays rejected', expected: 'Unsupported', plan: () => addSavedNativeFacts(citeSource(capture.savedPlan), true) },
|
||||
{ name: 'synthetic archived comparison stays rejected', expected: 'Unsupported', plan: () => replaceOnce(addSavedNativeFacts(citeSource(capture.savedPlan)), '### R1 commitment comparison', '### Archived R1 commitment comparison') },
|
||||
{ name: 'synthetic complete save without ACK stays rejected', expected: 'Invalid', missingAck: true, plan: () => addSavedNativeFacts(citeSource(capture.savedPlan)) },
|
||||
];
|
||||
for (const scenario of scenarios) test(`b176 paired attempt ${attemptIndex + 1}: ${scenario.name}`, () => {
|
||||
expect(sha(capture.seed)).toBe(capture.sourceRecord.sha256);
|
||||
expect(sha(capture.savedPlan)).toBe(capture.savedRecord.sha256);
|
||||
const savedPlan = scenario.plan(), call = structuredClone(capture.call);
|
||||
if (scenario.missingAck) call.answered = false;
|
||||
const fp = nativePlanCallFingerprint(call, 1, true);
|
||||
// Existing proposed tests cannot earn the separate seeded "no tests" finding.
|
||||
expect(ceoPaymentFinding(fp, capture.seed, savedPlan)).toBeNull();
|
||||
const counter = createCeoPaymentFindingCounter(capture.seed, () => savedPlan, ceoFirstReviewAUQ);
|
||||
if (scenario.expected === 'recorded') {
|
||||
expect(counter.isReviewAUQ(fp, structuredClone(capture.priorCalls))).toBe(true);
|
||||
expect(counter.trace).toHaveLength(1);
|
||||
expect(counter.trace[0]).toMatchObject({ kind: 'recorded-decision', ledgerId: 'R1' });
|
||||
} else expect(() => counter.isReviewAUQ(fp, structuredClone(capture.priorCalls))).toThrow(scenario.expected);
|
||||
});
|
||||
}
|
||||
@@ -69,7 +69,7 @@ describe('CEO mode option matching', () => {
|
||||
|
||||
test('the shared parser selects all callers while mode-specific regressions stay scoped', () => {
|
||||
expect(selectTests(['test/helpers/ceo-mode-option.ts'], E2E_TOUCHFILES).selected)
|
||||
.toEqual(['plan-ceo-mode-routing', 'plan-ceo-finding-count', 'plan-ceo-split-overflow']);
|
||||
.toEqual(['plan-ceo-mode-routing', 'plan-ceo-split-overflow']);
|
||||
expect(selectTests(['test/ceo-mode-option.test.ts'], E2E_TOUCHFILES).selected)
|
||||
.toEqual(['plan-ceo-mode-routing', 'plan-ceo-split-overflow']);
|
||||
expect(selectTests(['test/pty-option-selection.test.ts'], E2E_TOUCHFILES).selected)
|
||||
|
||||
@@ -1,165 +0,0 @@
|
||||
/** Free exact-field replay. The original incomplete paid report stays rejected. */
|
||||
import { test, expect } from 'bun:test';
|
||||
import fs from 'node:fs';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { createCeoPaymentFindingCounter } from './helpers/ceo-payment-findings';
|
||||
import { nativePlanCallFingerprint, ceoFirstReviewAUQ } from './helpers/claude-pty-runner';
|
||||
import capture from './fixtures/ceo-native-fields-f359.json';
|
||||
import plainCapture from './fixtures/ceo-plain-fields-f359.json';
|
||||
|
||||
const q=capture.call.questions[0]!;
|
||||
const start=capture.savedPlan.indexOf('### currentDecision (D1)');
|
||||
const end=capture.savedPlan.indexOf('## NOT in scope',start);
|
||||
const prefix=capture.savedPlan.slice(0,start),suffix=capture.savedPlan.slice(end);
|
||||
function section(question=q,bold=true) {
|
||||
const field=(name:string,value:string)=>`${bold?'**'+name+':**':name+':'} ${value}`;
|
||||
return ['### currentDecision (D1)','',field('Question',question.question),'',field('Header',question.header),'',
|
||||
...question.options.flatMap((o,i)=>{
|
||||
const label=/^[A-D][).:]\s+/.test(o.label)?o.label:`${String.fromCharCode(65+i)}) ${o.label}`;
|
||||
return [bold?`**${label}**`:label,o.description,''];
|
||||
})].join('\n');
|
||||
}
|
||||
function counter(plan:string,call=structuredClone(capture.call)) {
|
||||
const fp=nativePlanCallFingerprint(call,1,true);
|
||||
return {fp,count:createCeoPaymentFindingCounter(capture.seed,()=>plan,ceoFirstReviewAUQ)};
|
||||
}
|
||||
function reject(plan:string,call=structuredClone(capture.call)) {
|
||||
const {fp,count}=counter(plan,call);expect(()=>count.isReviewAUQ(fp)).toThrow(/Unsupported|Invalid/);expect(count.trace).toHaveLength(0);
|
||||
}
|
||||
const complete=()=>prefix+section()+'\n'+suffix;
|
||||
const replace=(text:string,from:string,to:string)=>{expect(text.split(from)).toHaveLength(2);return text.replace(from,to);};
|
||||
|
||||
test('original f359 paid incomplete report remains rejected with actual successful ACK',()=>{
|
||||
expect(createHash('sha256').update(capture.seed).digest('hex')).toBe(capture.sourceSha256);
|
||||
expect(createHash('sha256').update(capture.savedPlan).digest('hex')).toBe(capture.savedSha256);
|
||||
expect(Date.parse(capture.writeAck)).toBeLessThan(Date.parse(capture.readBackAck));
|
||||
expect(Date.parse(capture.readBackAck)).toBeLessThan(Date.parse(capture.questionAt));
|
||||
expect(Date.parse(capture.questionAt)).toBeLessThan(Date.parse(capture.call.answeredAt));
|
||||
reject(capture.savedPlan);
|
||||
});
|
||||
for(const bold of [false,true])for(const prefixed of [false,true])test(`complete exact native fields: ${bold?'bold':'plain'}, native ${prefixed?'prefixed':'unprefixed'} labels`,()=>{
|
||||
const call=structuredClone(capture.call),question=call.questions[0]!;
|
||||
if(!prefixed){question.options.forEach(o=>o.label=o.label.replace(/^[A-D][).:]\s+/,''));call.answers={[question.question]:question.options[0]!.label};}
|
||||
const {fp,count}=counter(prefix+section(question,bold)+'\n'+suffix,call);
|
||||
expect(count.isReviewAUQ(fp)).toBe(true);expect(count.trace).toMatchObject([{kind:'recorded-decision',ledgerId:'D1'}]);
|
||||
});
|
||||
const fieldMutations:Record<string,(s:string)=>string>={
|
||||
'missing question':s=>replace(s,'**Question:** '+q.question,''),
|
||||
'missing header':s=>replace(s,'**Header:** '+q.header,''),
|
||||
'changed question':s=>replace(s,'**Question:** '+q.question,'**Question:** '+q.question.replace('1000','1001')),
|
||||
'changed header':s=>replace(s,'**Header:** '+q.header,'**Header:** Foreign choice'),
|
||||
'changed label':s=>replace(s,'**'+q.options[0]!.label+'**','**A) Delete every test**'),
|
||||
'changed description':s=>replace(s,q.options[0]!.description!,q.options[0]!.description!.replace('exactly 2','exactly 20')),
|
||||
'missing label':s=>replace(s,'**'+q.options[0]!.label+'**',''),
|
||||
'missing description':s=>replace(s,q.options[0]!.description!,''),
|
||||
'missing final con':s=>replace(s,q.options[2]!.description!,q.options[2]!.description!.split('\n❌')[0]!),
|
||||
'duplicated question':s=>replace(s,'**Header:**','**Question:** '+q.question+'\n\n**Header:**'),
|
||||
'duplicated header':s=>replace(s,'**Header:** '+q.header,'**Header:** '+q.header+'\n\n**Header:** '+q.header),
|
||||
'duplicated option':s=>s+'\n**'+q.options[0]!.label+'**\n'+q.options[0]!.description+'\n',
|
||||
'conflicting field suffix':s=>replace(s,'**Header:** '+q.header,'**Header:** '+q.header+'; delete every job'),
|
||||
'quoted question':s=>replace(s,'**Question:** '+q.question,('**Question:** '+q.question).split('\n').map(l=>'> '+l).join('\n')),
|
||||
'quoted option':s=>replace(s,'**'+q.options[0]!.label+'**\n'+q.options[0]!.description,('**'+q.options[0]!.label+'**\n'+q.options[0]!.description).split('\n').map(l=>'> '+l).join('\n')),
|
||||
'code-only fields':s=>replace(s,s.slice(s.indexOf('**Question:**')),'```text\n'+s.slice(s.indexOf('**Question:**'))+'\n```'),
|
||||
'historical comparison':s=>s.replace('currentDecision','Archived currentDecision'),
|
||||
'unrelated instruction in descriptions':s=>s+'\nDelete all payment records before implementing this option.\n',
|
||||
'label consumes description line':s=>replace(s,'**'+q.options[0]!.label+'**\n','**'+q.options[0]!.label+'** '),
|
||||
};
|
||||
for(const [name,mutation]of Object.entries(fieldMutations))test(`exact native fields reject ${name} through exported counter`,()=>{
|
||||
reject(prefix+mutation(section())+'\n'+suffix);
|
||||
});
|
||||
const planMutations:Record<string,(s:string)=>string>={
|
||||
'missing row source':s=>s.replace(/Evidence: PLAN\.md/g,'Evidence: input').replace(/Factory exposes call history \+ sleeper record \(PLAN\.md lines 12-14\)/g,'Factory exposes call history + sleeper record'),
|
||||
'foreign row source':s=>s.replace('Evidence: PLAN.md','Evidence: OTHER.md'),
|
||||
'foreign document source':s=>s.replace('Source plan: PLAN.md','Source plan: OTHER.md'),
|
||||
'historical ledger ancestor':s=>s.replace('## Decision ledger','## Historical Decision ledger'),
|
||||
'historical row owner':s=>s.replace('| D1 (user) |','| D1 (historical user) |'),
|
||||
'duplicate active comparison':s=>s.replace('## NOT in scope',section()+'\n## NOT in scope'),
|
||||
'duplicate current row':s=>s.replace(/^(\| D1 \(user\).*\n)/m,'$1$1'),
|
||||
'conflicting current row':s=>s.replace(/^(\| D1 \(user\).*\n)/m,match=>match+match.replace('unresolved','declined')),
|
||||
};
|
||||
for(const [name,mutation]of Object.entries(planMutations))test(`exact native fields reject ${name}`,()=>{const plan=complete(),changed=mutation(plan);expect(changed).not.toBe(plan);reject(changed);});
|
||||
for(const kind of ['no ACK','failed ACK','wrong answer','empty header','empty description','inconsistent prefix','double prefix'])test(`exact native fields reject native ${kind}`,()=>{
|
||||
const call=structuredClone(capture.call),question=call.questions[0]!;
|
||||
if(kind==='no ACK')call.answered=false;
|
||||
if(kind==='failed ACK')call.failed=true;
|
||||
if(kind==='wrong answer')call.answers={[question.question]:'unoffered'};
|
||||
if(kind==='empty header')question.header='';
|
||||
if(kind==='empty description')question.options[0]!.description='';
|
||||
if(kind==='inconsistent prefix')question.options[0]!.label=question.options[0]!.label.replace('A)','B)');
|
||||
if(kind==='double prefix')question.options[0]!.label='A) '+question.options[0]!.label;
|
||||
if(kind.includes('prefix'))call.answers={[question.question]:question.options[0]!.label};
|
||||
reject(prefix+section(question)+'\n'+suffix,call);
|
||||
});
|
||||
test('exact native fields retain signature and duplicate ACK guards',()=>{
|
||||
const {fp,count}=counter(complete());fp.signature='foreign';expect(()=>count.isReviewAUQ(fp)).toThrow(/Invalid/);
|
||||
const fresh=counter(complete());expect(()=>fresh.count.isReviewAUQ(fresh.fp,[capture.call])).toThrow(/duplicated/);
|
||||
});
|
||||
|
||||
for(const bold of [false,true])test(`complete exact fields retain Header before Question (${bold?'bold':'plain'})`,()=>{
|
||||
const marker=(name:string)=>bold?`**${name}:**`:`${name}:`;
|
||||
const record=section(q,bold),question=`${marker('Question')} ${q.question}`,header=`${marker('Header')} ${q.header}`;
|
||||
const plan=prefix+replace(record,question+'\n\n'+header,header+'\n\n'+question)+'\n'+suffix;
|
||||
const {fp,count}=counter(plan);expect(count.isReviewAUQ(fp)).toBe(true);
|
||||
});
|
||||
test('complete exact fields preserve unique native selectors in non-positional order',()=>{
|
||||
const call=structuredClone(capture.call),question=call.questions[0]!;
|
||||
question.options=[question.options[2]!,question.options[0]!,question.options[1]!];
|
||||
const {fp,count}=counter(prefix+section(question)+'\n'+suffix,call);expect(count.isReviewAUQ(fp)).toBe(true);
|
||||
});
|
||||
for(const ancestor of ['Archived','Historical'])test(`complete exact fields cannot borrow options from ${ancestor} child`,()=>{
|
||||
reject(prefix+replace(section(),'**'+q.options[1]!.label+'**','#### '+ancestor+' option details\n\n**'+q.options[1]!.label+'**')+'\n'+suffix);
|
||||
});
|
||||
function plainEvaluate(plan=plainCapture.savedPlan) {
|
||||
const fp=nativePlanCallFingerprint(structuredClone(plainCapture.call),1,true);
|
||||
const count=createCeoPaymentFindingCounter(plainCapture.seed,()=>plan,ceoFirstReviewAUQ);
|
||||
return {fp,count};
|
||||
}
|
||||
test('actual f359 plain selector paragraphs count through legacy complete-facts path',()=>{
|
||||
expect(createHash('sha256').update(plainCapture.seed).digest('hex')).toBe(plainCapture.sourceSha256);
|
||||
expect(createHash('sha256').update(plainCapture.savedPlan).digest('hex')).toBe(plainCapture.savedSha256);
|
||||
const {fp,count}=plainEvaluate();expect(count.isReviewAUQ(fp)).toBe(true);
|
||||
expect(count.trace).toMatchObject([{kind:'recorded-decision',ledgerId:'D1'}]);
|
||||
});
|
||||
const plainBlocks=plainCapture.savedPlan.match(/^[A-D]\) .+\n(?: .*(?:\n|$))+/gm)!;
|
||||
const plainMutations:Record<string,(s:string)=>string>={
|
||||
'missing effort':s=>s.replace('Effort S','Work S'),
|
||||
'missing risk':s=>s.replace('Risk low','Exposure low'),
|
||||
'missing pros':s=>s.replace('Pros:','Benefits:'),
|
||||
'missing cons':s=>s.replace('Cons:','Costs:'),
|
||||
'ambiguous duplicate risk':s=>s+' Risk high.\n',
|
||||
'unrelated complete option':()=> 'A) Delete all payment tables\n Summary: remove all customer records. Effort S. Risk high. Pros: reduces storage. Cons: destroys data.\n',
|
||||
'quoted option fields':s=>s.split('\n').map(l=>'> '+l).join('\n'),
|
||||
'code-only option fields':s=>'```text\n'+s+'\n```\n',
|
||||
'detached option fields':s=>s.replace('\n Summary:','\n\nUnrelated record:\n Summary:'),
|
||||
};
|
||||
for(const [name,mutation]of Object.entries(plainMutations))test(`plain selector paragraphs reject ${name}`,()=>{
|
||||
expect(plainBlocks).toHaveLength(3);
|
||||
const plan=replace(plainCapture.savedPlan,plainBlocks[0]!,mutation(plainBlocks[0]!));
|
||||
const {fp,count}=plainEvaluate(plan);expect(()=>count.isReviewAUQ(fp)).toThrow(/Unsupported/);
|
||||
});
|
||||
test('plain selector paragraphs retain unindented continuation and reject duplicated options',()=>{
|
||||
const normalized=plainCapture.savedPlan.replace(/^ /gm,'');
|
||||
const valid=plainEvaluate(normalized);expect(valid.count.isReviewAUQ(valid.fp)).toBe(true);
|
||||
const duplicate=plainEvaluate(replace(normalized,plainBlocks[0]!.replace(/^ /gm,''),plainBlocks[0]!.replace(/^ /gm,'')+'\n'+plainBlocks[0]!.replace(/^ /gm,'')));
|
||||
expect(()=>duplicate.count.isReviewAUQ(duplicate.fp)).toThrow(/Unsupported/);
|
||||
});
|
||||
|
||||
test('plain selector paragraphs cannot borrow facts from an archived child',()=>{
|
||||
const plan=replace(plainCapture.savedPlan,plainBlocks[0]!,'### Archived option details\n\n'+plainBlocks[0]!);
|
||||
const {fp,count}=plainEvaluate(plan);expect(()=>count.isReviewAUQ(fp)).toThrow(/Unsupported/);
|
||||
});
|
||||
|
||||
for (const ancestor of ['Historical', 'Archived']) test(`plain selector list children cannot bypass ${ancestor} ancestry`, () => {
|
||||
let plan=plainCapture.savedPlan;
|
||||
for (const block of plainBlocks) plan=replace(plan,block,'- '+block);
|
||||
plan=replace(plan,'- '+plainBlocks[0]!,`### ${ancestor} option details\n\n- `+plainBlocks[0]!);
|
||||
const {fp,count}=plainEvaluate(plan);expect(()=>count.isReviewAUQ(fp)).toThrow(/Unsupported/);
|
||||
});
|
||||
|
||||
for(const prelude of ['Prepared for this current decision.', 'Status: pending']) test(`exact native record retains neutral prefix: ${prelude}`,()=>{
|
||||
const plan=prefix+replace(section(),'### currentDecision (D1)','### currentDecision (D1)\n\n'+prelude)+'\n'+suffix;
|
||||
const {fp,count}=counter(plan);expect(count.isReviewAUQ(fp)).toBe(true);
|
||||
});
|
||||
for(const prelude of ['This decision is withdrawn.', 'This decision is resolved.', 'Status: withdrawn', 'Status: superseded', 'The decision is not current.']) test(`exact native record rejects inactive prefix: ${prelude}`,()=>{
|
||||
reject(prefix+replace(section(),'### currentDecision (D1)','### currentDecision (D1)\n\n'+prelude)+'\n'+suffix);
|
||||
});
|
||||
File diff suppressed because it is too large.
Load diff
@@ -1,162 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import captured from './fixtures/ceo-numbered-brief-ak.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const call = (index = 4): any => structuredClone(captured.calls[index]);
|
||||
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
|
||||
function edit(c: any, change: (s: string) => string) {
|
||||
const q = c.questions[0], answer = c.answers[q.question];
|
||||
q.question = change(q.question); c.answers = { [q.question]: answer };
|
||||
}
|
||||
function offered(c: any, change: (o: any, i: number) => void) {
|
||||
const q = c.questions[0], selected = q.options.findIndex((o: any) => o.label === c.answers[q.question]);
|
||||
q.options.forEach(change); c.answers = { [q.question]: q.options[selected].label };
|
||||
}
|
||||
|
||||
for (const [index, name] of [[4, 'email ordering'], [5, 'raw SQL'], [6, 'missing automated tests'], [7, 'N+1 read']] as const) {
|
||||
test(`actual completed ${name} brief starts substantive CEO review`, () => {
|
||||
expect(ceoFirstReviewAUQ(fp(call(index)))).toBe(true);
|
||||
});
|
||||
}
|
||||
|
||||
test('the complete captured phase retains routing and factual clarification as setup', () => {
|
||||
let started = false;
|
||||
const phases = captured.calls.map(c => {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, true, true, false, false, false, false]);
|
||||
for (const c of captured.calls.slice(0, 4)) expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
});
|
||||
|
||||
test('issue ownership survives equivalent separators, optional qids and consistent renumbering', () => {
|
||||
for (let index = 4; index < 8; index++) {
|
||||
for (const separator of ['—', '–', '-']) {
|
||||
const c = call(index); edit(c, s => s.replace(/^D\d+ — /, `D12 ${separator} `).replace(/\s*<gstack-qid:[^>]+>\s*$/, ''));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
const c = call(index), old = index - 2;
|
||||
edit(c, s => s.replace(`Issue ${old}:`, 'Issue 19:').replace(new RegExp('\\b' + old + '([A-Z])\\b', 'g'), '19$1'));
|
||||
offered(c, o => { o.label = o.label.replace(/^\d+/, '19'); });
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
c.answers[c.questions[0].question] = c.questions[0].options[2].label;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('display effort and positive bullet decoration do not change an offered action', () => {
|
||||
for (let index = 4; index < 8; index++) for (const change of [
|
||||
(s: string) => s.replace(/Human [^.]+\. /, 'Human 2 days / CC 30 minutes. '),
|
||||
(s: string) => s.replace(/Human [^.]+\. /, '').replace(/✅ /g, ''),
|
||||
(s: string) => s.replace(/✅ /g, '✅ '),
|
||||
]) {
|
||||
const c = call(index); offered(c, o => { o.description = change(o.description); });
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
const suite = call(6); offered(suite, o => { o.label = o.label.replace('unit + integration', 'unit and integration'); });
|
||||
expect(ceoFirstReviewAUQ(fp(suite))).toBe(true);
|
||||
});
|
||||
|
||||
test('native completion, the offered answer and exact fingerprint remain mandatory', () => {
|
||||
for (let index = 4; index < 8; index++) for (const mutate of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.sessionId = ''; },
|
||||
(c: any) => { c.toolUseId = ''; },
|
||||
(c: any) => { c.answers = {}; },
|
||||
(c: any) => { c.answers[c.questions[0].question] = 'Unrelated answer'; },
|
||||
(c: any) => { c.questions[0].multiSelect = true; },
|
||||
(c: any) => { c.questions.push(structuredClone(c.questions[0])); },
|
||||
(c: any) => { c.questions[0].header = 'Approach'; },
|
||||
(c: any) => { c.questions[0].options[1].description = ''; },
|
||||
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[0].label; },
|
||||
(c: any) => edit(c, s => s.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-eng-review-finding>')),
|
||||
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
for (let index = 4; index < 8; index++) {
|
||||
const f = fp(call(index));
|
||||
expect(ceoFirstReviewAUQ({ ...f, signature: 'foreign:tool' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...f, nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...f, options: f.options.slice(1) })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('a current issue cannot borrow another issue identity or recommendation', () => {
|
||||
for (let index = 4; index < 8; index++) for (const mutate of [
|
||||
(c: any) => edit(c, s => s.replace(/Issue \d+:/, 'Issue 99:')),
|
||||
(c: any) => { c.questions[0].header = 'Finding 99'; },
|
||||
(c: any) => { c.questions[0].options[1].label = '99B: Foreign choice'; },
|
||||
(c: any) => { c.questions[0].options[1].label = c.questions[0].options[1].label.replace(/B:/, 'A:'); },
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: \d+[A-Z]/m, 'Recommendation: 99A')),
|
||||
(c: any) => edit(c, s => s.replace(/^Recommendation: .+$/m, '')),
|
||||
(c: any) => edit(c, s => s + '\nRecommendation: 99A'),
|
||||
]) { const c = call(index); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
});
|
||||
|
||||
test('source, hypothetical and withdrawn assessments do not start current review', () => {
|
||||
for (let index = 4; index < 8; index++) for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```\n' + s + '\n```',
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: If approved, '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: Suppose '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a quoted source excerpt. '),
|
||||
(s: string) => s.replace(/^ELI10: /m, 'ELI10: The following is a hypothetical example. '),
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, ''),
|
||||
(s: string) => s + '\nThis issue has been withdrawn.',
|
||||
(s: string) => s + `\nIssue ${index - 2} is resolved.`,
|
||||
(s: string) => s + '\nNo current issue remains.',
|
||||
]) { const c = call(index); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false); }
|
||||
for (let index = 4; index < 8; index++) {
|
||||
const c = call(index); edit(c, s => s + '\nOld note: "This issue has been withdrawn."');
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('the new declarative findings still need an asserted current defect', () => {
|
||||
for (const [index, title] of [
|
||||
[5, 'the lookup does not interpolate request.params.userId into a raw SQL fragment.'],
|
||||
[5, 'the lookup no longer interpolates request.params.userId into a raw SQL fragment.'],
|
||||
[5, 'the lookup used to interpolate user input into a raw SQL string.'],
|
||||
[6, 'automated tests are planned for the new payment handler.'],
|
||||
[7, 'the handler no longer fetches each order in a loop (N+1).'],
|
||||
[7, 'the handler reads all orders with one query.'],
|
||||
] as const) {
|
||||
const c = call(index);
|
||||
edit(c, s => s.replace(/^(D\d+ — Issue \d+: ).+$/m, '$1' + title)
|
||||
.replace(/^ELI10: .+$/m, title.includes('used to')
|
||||
? 'ELI10: The previous lookup used to interpolate user input into a raw SQL string. The current lookup uses bound parameters and has no injection risk.'
|
||||
: 'ELI10: The current implementation satisfies the stated contract.'));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('only an offered current amendment can supply remedy evidence', () => {
|
||||
for (let index = 4; index < 8; index++) for (const description of [
|
||||
'Keep this advisory report for reference.',
|
||||
'Human ~3h / CC ~15min. ❌ Add bounded error handling.',
|
||||
'Human ~3h / CC ~15min. ❌ A prior proposal. Add bounded error handling.',
|
||||
'Human ~3h / CC ~15min. ✅ "Add bounded error handling."',
|
||||
'Human ~3h / CC ~15min. ✅ If approved, add bounded error handling.',
|
||||
'Human ~3h / CC ~15min. ✅ Write the completed report.',
|
||||
'Historical source excerpt: ✅ Add bounded error handling.',
|
||||
'If approved: ✅ Add bounded error handling.',
|
||||
'Hypothetical example: ✅ Add bounded error handling.',
|
||||
'The following is a quoted source excerpt. ✅ Add bounded error handling.',
|
||||
'The following is a hypothetical example. ✅ Add bounded error handling.',
|
||||
]) {
|
||||
const c = call(index);
|
||||
offered(c, (o, i) => { o.label = `${index - 2}${String.fromCharCode(65 + i)}: Consider candidate ${i}`; o.description = description; });
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('only the existing CEO finding-count owner selects this regression fixture', () => {
|
||||
for (const dependency of ['test/ceo-numbered-brief-ak.test.ts', 'test/fixtures/ceo-numbered-brief-ak.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name))
|
||||
.toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -1,160 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import fixture from './fixtures/ceo-parenthesized-issue-ah.json';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
const calls = () => structuredClone(fixture.calls) as NativePlanQuestionCall[];
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
|
||||
function reanswer(c: NativePlanQuestionCall) {
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
return c;
|
||||
}
|
||||
function changed(original: NativePlanQuestionCall, mutate: (c: NativePlanQuestionCall) => void) {
|
||||
const c = structuredClone(original); mutate(c); return reanswer(c);
|
||||
}
|
||||
|
||||
test('both exact completed Issue questions start review with descriptive headers', () => {
|
||||
for (const c of calls()) expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0 };
|
||||
for (const c of [...fixture.setupCalls, ...calls()] as NativePlanQuestionCall[]) {
|
||||
const phase = planCountQuestionPhase(fp(c), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted;
|
||||
counts[phase.preReview ? 'setup' : 'review']++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 2, review: 2 });
|
||||
expect(fixture.historicalOutcome).toBe('no_review_questions');
|
||||
});
|
||||
|
||||
test('fixture questions and selected answers are exact owned public request/result projections', () => {
|
||||
for (const c of calls()) {
|
||||
const requests = fixture.publicEvents.filter(e => e.record.message.content.some(b => 'id' in b && b.id === c.toolUseId));
|
||||
const results = fixture.publicEvents.filter(e => e.record.message.content.some(b => 'tool_use_id' in b && b.tool_use_id === c.toolUseId));
|
||||
expect(requests).toHaveLength(1); expect(results).toHaveLength(1);
|
||||
expect(requests[0]!.record.sessionId).toBe(c.sessionId);
|
||||
expect(results[0]!.record.sessionId).toBe(c.sessionId);
|
||||
const request = requests[0]!.record.message.content.find(b => 'id' in b && b.id === c.toolUseId) as any;
|
||||
expect(request.input.questions).toEqual(c.questions);
|
||||
const result = results[0]!.record.message.content.find(b => 'tool_use_id' in b && b.tool_use_id === c.toolUseId) as any;
|
||||
expect(result.is_error).not.toBe(true);
|
||||
expect(result.content).toContain(`"${c.questions[0]!.question}"="${c.answers![c.questions[0]!.question]}"`);
|
||||
expect(c.answeredAt).toBe(results[0]!.record.timestamp);
|
||||
}
|
||||
});
|
||||
|
||||
test('complete current native identity and actual offered answer remain required', () => {
|
||||
for (const original of calls()) {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.sessionId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.toolUseId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { 'prior question': c.questions[0]!.options[0]!.label }; },
|
||||
(c: NativePlanQuestionCall) => { c.answers![c.questions[0]!.question] = 'foreign answer'; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
]) {
|
||||
const c = structuredClone(original); mutate(c);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ({ ...fp(original), signature: 'foreign:call' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(original), nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...fp(original), options: [] })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('title, recommendation, every option and any numbered header share one issue identity', () => {
|
||||
for (const original of calls()) {
|
||||
const n = /\(Issue (\d+)\)/.exec(original.questions[0]!.question)![1]!;
|
||||
for (const header of [`Finding ${n}`, `Issue ${n}`, `F${n}`]) {
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.header = header; })))).toBe(true);
|
||||
}
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation:.*\n/m, ''); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation: \d+A/m, 'Recommendation: 99A'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/^Recommendation: \d+A/m, `Recommendation: ${n}Z`); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = '99B) Different issue'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = c.questions[0]!.options[0]!.label; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding 99'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Finding'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/D\d+ \(Issue \d+\) — /, ''); c.questions[0]!.header = `Issue ${n}`; },
|
||||
]) expect(ceoFirstReviewAUQ(fp(changed(original, mutate)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('source, conditional and stale or withdrawn briefs do not start review', () => {
|
||||
for (const original of calls()) {
|
||||
for (const prefix of ['Example: ', 'If requested: ', '> ', ' ', '```text\n']) {
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = prefix + c.questions[0]!.question; })))).toBe(false);
|
||||
}
|
||||
for (const framing of ['If this hypothetical plan were adopted, ', 'Example: ', 'Historical example only. ', 'Quoted assessment: ']) {
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = c.questions[0]!.question.replace('ELI10: ', `ELI10: ${framing}`); })))).toBe(false);
|
||||
}
|
||||
for (const tail of ['No current defect exists.', 'Correction: this issue is already resolved.', 'This question is only an example.', 'I withdraw this finding.']) {
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question += '\n' + tail; })))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['> ', ' ', '```\n']) {
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question = c.questions[0]!.question.replace(/^ELI10:/m, prefix + 'ELI10:'); })))).toBe(false);
|
||||
}
|
||||
// Later attributed source text does not withdraw a present decision.
|
||||
expect(ceoFirstReviewAUQ(fp(changed(original, c => { c.questions[0]!.question += '\nAn old note said: "No current defect exists."'; })))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('qid and setup exclusions apply before the new identity form', () => {
|
||||
for (const original of calls()) {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question += ' <gstack-qid:plan-ceo-review-extra>'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-eng-review-validation>'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-ceo-review-approach>'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-ceo-review-mode>'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Approach'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = 'HOLD SCOPE'; },
|
||||
]) expect(ceoFirstReviewAUQ(fp(changed(original, mutate)))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('identity recognition is independent of observed component and numbering', () => {
|
||||
for (const original of calls()) {
|
||||
const c = changed(original, c => {
|
||||
const q = c.questions[0]!; const n = /\(Issue (\d+)\)/.exec(q.question)![1]!;
|
||||
q.header = 'Notification state';
|
||||
q.question = q.question.replace(/^D\d+/, 'D24').replace(`(Issue ${n})`, '(Issue 17)')
|
||||
.replace(new RegExp(`\\b${n}([ABC])\\b`, 'g'), '17$1')
|
||||
.replace(/Stripe/g, 'PaymentProvider').replace(/email/g, 'notification').replace(/userId/g, 'accountKey');
|
||||
q.options.forEach(o => { o.label = o.label.replace(new RegExp(`^${n}`), '17'); });
|
||||
});
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options.at(-1)!.label };
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('new fixture and controls select only their existing CEO count owner', () => {
|
||||
for (const file of ['test/ceo-parenthesized-issue-ah.test.ts', 'test/fixtures/ceo-parenthesized-issue-ah.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES).selected).toEqual(['plan-ceo-finding-count']);
|
||||
}
|
||||
});
|
||||
|
||||
test('numbered administrative, literal-only, hypothetical and withdrawn briefs are not defects', () => {
|
||||
for (const original of calls()) {
|
||||
for (const mutate of [
|
||||
(q: NativePlanQuestionCall['questions'][number]) => {
|
||||
const n = /\(Issue (\d+)\)/.exec(q.question)![1]!;
|
||||
q.question = q.question.replace(/^(D\d+ \(Issue \d+\) — ).*/, '$1How should we archive this completed review?')
|
||||
.replace(/^ELI10:.*$/m, 'ELI10: The review is complete. This choice only saves the finished report.');
|
||||
q.options.forEach((o, i) => { o.label = `${n}${String.fromCharCode(65 + i)}) Save report format ${i}`; o.description = 'Store the completed review report.'; });
|
||||
},
|
||||
(q: NativePlanQuestionCall['questions'][number]) => { q.question += '\nThere is no defect or unresolved issue; this is a historical example.'; },
|
||||
(q: NativePlanQuestionCall['questions'][number]) => { q.question = q.question.replace(/^ELI10:.*$/m, 'ELI10: `The plan has no error handling.`'); },
|
||||
(q: NativePlanQuestionCall['questions'][number]) => { q.question = q.question.replace(/^(D\d+ \(Issue \d+\) — ).*/, '$1What should happen if a hypothetical future handler lacked error handling?'); },
|
||||
]) {
|
||||
const c = changed(original, c => mutate(c.questions[0]!));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -1,304 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import fixture from './fixtures/ceo-payment-ledger-decisions.json';
|
||||
import { ceoPaymentFinding, createCeoPaymentFindingCounter } from './helpers/ceo-payment-findings';
|
||||
import { nativePlanCallFingerprint, ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
|
||||
const clone = <T>(value: T): T => structuredClone(value);
|
||||
const seeded = fixture.captures.filter(c => c.kind === 'seeded-remedy');
|
||||
const fingerprint = (capture = seeded[0]!) => nativePlanCallFingerprint(clone(capture.call), 1, true);
|
||||
const saved = (i = 0) => seeded[i]!.savedPlan!;
|
||||
const recognize = (fp = fingerprint(), plan = saved(), seed = fixture.seed) => ceoPaymentFinding(fp, seed, plan);
|
||||
const reanswer = (fp: ReturnType<typeof fingerprint>) => {
|
||||
const q = fp.nativeCall!.questions[0]!;
|
||||
fp.nativeCall!.answers = { [q.question]: q.options[0]!.label };
|
||||
fp.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
};
|
||||
|
||||
test('actual CLI 2.1.251 capture: five independently acknowledged 0D remedies have saved seed linkage', () => {
|
||||
expect(fixture.originalOutcome).toEqual({ outcome: 'no_review_questions', reviewCount: 0, step0Count: 8 });
|
||||
expect(seeded).toHaveLength(5);
|
||||
expect(seeded.map(c => ceoFirstReviewAUQ(fingerprint(c)))).toEqual([false, false, false, false, false]);
|
||||
expect(seeded.map(c => ceoPaymentFinding(fingerprint(c), fixture.seed, c.savedPlan!)?.seed))
|
||||
.toEqual(['dispatcher', 'lookup', 'email', 'tests', 'orders']);
|
||||
for (const c of seeded) {
|
||||
const found = ceoPaymentFinding(fingerprint(c), fixture.seed, c.savedPlan!);
|
||||
expect(found?.phase).toBe('Step 0D. Alternatives (pending)');
|
||||
expect(found?.signature).toBe(`${c.call.sessionId}:${c.call.toolUseId}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('all eight captured calls retain phase provenance, exclude onboarding and count the actual TODO toward the upper bound', () => {
|
||||
let plan = '', boundary = false, count = 0;
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => plan, ceoFirstReviewAUQ);
|
||||
const prior: any[] = [];
|
||||
for (const c of fixture.captures) {
|
||||
if (c.savedPlan) plan = c.savedPlan;
|
||||
const fp = fingerprint(c);
|
||||
const phase = planCountQuestionPhase(fp, boundary, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
count += Number(counter.isReviewAUQ(fp, prior));
|
||||
boundary = phase.reviewStarted;
|
||||
prior.push(c.call);
|
||||
}
|
||||
expect(count).toBe(6);
|
||||
expect(boundary).toBe(false);
|
||||
expect(counter.trace.filter(t => 'seed' in t)).toHaveLength(5);
|
||||
expect(counter.trace.filter(t => 'phase' in t).every(t => t.phase.startsWith('Step 0D'))).toBe(true);
|
||||
});
|
||||
|
||||
for (const [name, mutate] of Object.entries({
|
||||
unanswered: (fp: any) => { fp.nativeCall.answered = false; fp.nativeCall.answers = {}; },
|
||||
'failed tool result': (fp: any) => { fp.nativeCall.failed = true; },
|
||||
'unanswered question index': (fp: any) => { fp.nativeCall.unansweredQuestionIndices = [0]; },
|
||||
'foreign signature': (fp: any) => { fp.signature = 'other:tool'; },
|
||||
'wrong question answer identity': (fp: any) => { fp.nativeCall.answers = { other: fp.options[0].label }; },
|
||||
'not an offered answer': (fp: any) => { fp.nativeCall.answers[fp.nativeCall.questions[0].question] = 'not offered'; },
|
||||
multiselect: (fp: any) => { fp.nativeCall.questions[0].multiSelect = true; },
|
||||
'duplicate native labels': (fp: any) => { fp.nativeCall.questions[0].options[1].label = fp.options[0].label; reanswer(fp); },
|
||||
'stale visible option': (fp: any) => { fp.options[0].label = 'other'; },
|
||||
'missing acknowledgment time': (fp: any) => { delete fp.nativeCall.answeredAt; },
|
||||
'quoted current question': (fp: any) => { fp.nativeCall.questions[0].question = fp.nativeCall.questions[0].question.split('\n').map((l: string) => '> ' + l).join('\n'); reanswer(fp); },
|
||||
'copied question in code': (fp: any) => { fp.nativeCall.questions[0].question = '```\n' + fp.nativeCall.questions[0].question + '\n```'; reanswer(fp); },
|
||||
'wrong defect': (fp: any) => { fp.nativeCall.questions[0].question = fp.nativeCall.questions[0].question.replace(/^ELI10: .+$/m, 'ELI10: The plan has a missing loading spinner.'); reanswer(fp); },
|
||||
'ordinary approach only': (fp: any) => { fp.nativeCall.questions[0].question = fp.nativeCall.questions[0].question.replace(/^ELI10: .+$/m, 'ELI10: Choose the overall project approach; all current obligations are already satisfied.'); reanswer(fp); },
|
||||
})) test(`does not credit ${name}`, () => { const fp = fingerprint(); mutate(fp); expect(recognize(fp)).toBeNull(); });
|
||||
|
||||
test('unrelated, quoted, duplicated or unresolved-without-comparison ledgers do not bind', () => {
|
||||
expect(recognize(fingerprint(), saved().replaceAll('R1', 'OTHER'))).toBeNull();
|
||||
expect(recognize(fingerprint(), saved().split('\n').map(l => '> ' + l).join('\n'))).toBeNull();
|
||||
expect(recognize(fingerprint(), '```md\n' + saved() + '\n```')).toBeNull();
|
||||
expect(recognize(fingerprint(), saved() + '\n' + saved())).toBeNull();
|
||||
expect(recognize(fingerprint(), saved().slice(0, saved().indexOf('### R1.')))).toBeNull();
|
||||
expect(recognize(fingerprint(), saved().replaceAll('PLAN.md', 'unrelated-project.md'))).toBeNull();
|
||||
expect(recognize(fingerprint(), saved().replace('Bypass `WebhookDispatcher` with standalone class', 'Existing dispatcher routing is correct'))).toBeNull();
|
||||
expect(recognize(fingerprint(), saved(), '# Unrelated plan\nBuild a loading spinner.')).toBeNull();
|
||||
});
|
||||
|
||||
test('a changed baseline cannot borrow an obsolete seeded defect', () => {
|
||||
const fp = fingerprint(seeded[1]!);
|
||||
fp.nativeCall!.questions[0]!.question = fp.nativeCall!.questions[0]!.question.replace(/^ELI10: .+$/m,
|
||||
'ELI10: This finding is resolved. The current plan uses a bound parameter and has no current SQL defect.'); reanswer(fp);
|
||||
expect(ceoPaymentFinding(fp, fixture.seed, saved(1))).toBeNull();
|
||||
expect(recognize(fingerprint(), saved().replace('Bypass `WebhookDispatcher` with standalone class', 'Register through the existing dispatcher'))).toBeNull();
|
||||
});
|
||||
|
||||
test('ledger IDs are bound values, not literal R1/R2 labels; saved phase remains accurate', () => {
|
||||
const fp = fingerprint(); fp.nativeCall!.questions[0]!.question = fp.nativeCall!.questions[0]!.question.replaceAll('R1', 'PAYMENT-9'); reanswer(fp);
|
||||
expect(recognize(fp, saved().replaceAll('R1', 'PAYMENT-9'))?.ledgerId).toBe('PAYMENT-9');
|
||||
expect(recognize(fingerprint(), saved().replace('Step 0D. Alternatives (pending)', 'Section 1. Architecture'))?.phase).toBe('Section 1. Architecture');
|
||||
});
|
||||
|
||||
test('duplicate native callbacks never earn credit and repeated real questions still count toward the ceiling', () => {
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => saved(), ceoFirstReviewAUQ);
|
||||
const fp = fingerprint();
|
||||
expect(counter.isReviewAUQ(fp)).toBe(true);
|
||||
expect(() => counter.isReviewAUQ(fp, [fp.nativeCall!])).toThrow(/duplicated/);
|
||||
let count = 1;
|
||||
for (let i = 1; i < 8; i++) {
|
||||
const repeated = fingerprint(); repeated.nativeCall!.toolUseId += `-${i}`; repeated.signature += `-${i}`;
|
||||
count += Number(counter.isReviewAUQ(repeated));
|
||||
}
|
||||
expect(count).toBe(8); // unchanged hard cap: above the accepted ceiling of 7
|
||||
expect(counter.trace.filter(t => 'seed' in t)).toHaveLength(8);
|
||||
});
|
||||
|
||||
test('a later mode/setup question stays excluded and unknown extra decisions fail rather than disappear', () => {
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => saved(), ceoFirstReviewAUQ);
|
||||
expect(counter.isReviewAUQ(fingerprint())).toBe(true);
|
||||
const mode = fingerprint(); const q = mode.nativeCall!.questions[0]!;
|
||||
q.header = 'Mode'; q.question = 'D9 — Which review mode should I use?';
|
||||
q.options = ['HOLD SCOPE', 'SELECTIVE EXPANSION', 'SCOPE EXPANSION', 'SCOPE REDUCTION'].map(label => ({ label })); reanswer(mode);
|
||||
expect(counter.isReviewAUQ(mode)).toBe(false);
|
||||
const approach = fingerprint(); approach.nativeCall!.questions[0]!.header = 'Approach';
|
||||
approach.nativeCall!.questions[0]!.question = 'D10 — Which overall approach should we choose?'; reanswer(approach);
|
||||
expect(counter.isReviewAUQ(approach)).toBe(false);
|
||||
const extra = fingerprint(); extra.nativeCall!.questions[0]!.question = 'D11 — Should the project change its billing currency?'; reanswer(extra);
|
||||
expect(() => counter.isReviewAUQ(extra)).toThrow(/cannot exclude it from the 4–7 count/);
|
||||
});
|
||||
|
||||
|
||||
test('source-required ledger meanings survive reordered columns, renamed heading and different nesting', () => {
|
||||
const plan = saved().replace('## Decision ledger', '# Choices').replace('## Step 0D.', '## Initial choices: Step 0D.').replace('### R1.', '#### R1.');
|
||||
const lines = plan.split('\n').map(line => {
|
||||
if (!line.startsWith('|')) return line;
|
||||
const cells = line.split('|');
|
||||
if (cells.length !== 8) return line;
|
||||
return '|'+[cells[3],cells[1],cells[5],cells[4],cells[2],cells[6]].join('|')+'|';
|
||||
});
|
||||
expect(recognize(fingerprint(), lines.join('\n'))?.seed).toBe('dispatcher');
|
||||
});
|
||||
|
||||
test('native labels, decision title syntax and chosen alternative are not metric protocols', () => {
|
||||
const fp = fingerprint(seeded[1]!); const q = fp.nativeCall!.questions[0]!;
|
||||
q.question = q.question.replace('D4 (ledger R2) —', 'Resolve R2:');
|
||||
q.header = 'Safe lookup'; q.options[0]!.label = 'Keep the DB interface';
|
||||
q.options[0]!.description = 'Bind the external ID as a database parameter. ' + q.options[0]!.description;
|
||||
reanswer(fp);
|
||||
expect(ceoPaymentFinding(fp, fixture.seed, saved(1))?.seed).toBe('lookup');
|
||||
fp.nativeCall!.answers = { [q.question]: q.options[1]!.label };
|
||||
expect(ceoPaymentFinding(fp, fixture.seed, saved(1))?.seed).toBe('lookup');
|
||||
});
|
||||
|
||||
test('an operative inline Proposed field needs no separately named comparison table', () => {
|
||||
const plan = saved().slice(0,saved().indexOf('## Step 0D.')).replace('see 0D', 'Register the handler through WebhookDispatcher; preserve its class name');
|
||||
expect(recognize(fingerprint(),plan)?.seed).toBe('dispatcher');
|
||||
});
|
||||
|
||||
|
||||
test('declared onboarding subjects and option semantics survive numbering and punctuation changes', () => {
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => saved(), ceoFirstReviewAUQ);
|
||||
for (const c of fixture.captures.slice(0, 2)) {
|
||||
const fp = fingerprint(c), q = fp.nativeCall!.questions[0]!;
|
||||
q.question = q.question.replace(/^D[0-9]+ — /, 'D37: ');
|
||||
q.options = q.options.map((o, i) => ({ ...o, label: `${i + 1}. ${o.label}` })); reanswer(fp);
|
||||
expect(counter.isReviewAUQ(fp)).toBe(false);
|
||||
}
|
||||
const mode = fingerprint(), q = mode.nativeCall!.questions[0]!;
|
||||
q.header = 'Review preference'; q.question = 'Select a review posture?';
|
||||
q.options = ['SCOPE REDUCTION — narrowest deliverable', 'HOLD SCOPE (recommended)', 'SELECTIVE EXPANSION — cherry-pick', 'SCOPE EXPANSION — dream big'].map(label => ({ label })); reanswer(mode);
|
||||
expect(counter.isReviewAUQ(mode)).toBe(false);
|
||||
});
|
||||
|
||||
test('a TODO label cannot hide an actual question, and the existing completion predicate remains the administrative owner', () => {
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => saved(4), ceoFirstReviewAUQ);
|
||||
const todo = fixture.captures.at(-1)!;
|
||||
expect(counter.isReviewAUQ(fingerprint(todo))).toBe(true);
|
||||
expect(counter.trace.at(-1)).toMatchObject({ kind: 'additional-current-decision' });
|
||||
const informational = fingerprint(todo); informational.nativeCall!.questions[0]!.options = [{label:'Read the example'}, {label:'Show the same example'}]; reanswer(informational);
|
||||
expect(() => counter.isReviewAUQ(informational)).toThrow(/cannot exclude/);
|
||||
});
|
||||
|
||||
|
||||
import zeroAbsenceFixture from './fixtures/ceo-zero-test-absence-6f6730f4.json';
|
||||
const zeroAbsenceFingerprint = (replacement = 'zero automated tests') => {
|
||||
const call = structuredClone(zeroAbsenceFixture.call);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('zero automated tests', replacement);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label } as typeof call.answers;
|
||||
return nativePlanCallFingerprint(call, 1, true);
|
||||
};
|
||||
const zeroAbsenceFinding = (question = zeroAbsenceFingerprint(), plan = zeroAbsenceFixture.savedPlan) =>
|
||||
ceoPaymentFinding(question, zeroAbsenceFixture.seed, plan);
|
||||
|
||||
test('captured numeric-zero D4 question binds its authenticated unchanged ledger row', () => {
|
||||
// The final full report was not retained; this tests the observed lexical
|
||||
// blocker using the complete earlier report and unchanged D4 row only.
|
||||
expect(zeroAbsenceFinding()).toMatchObject({ seed: 'tests', ledgerId: 'D4' });
|
||||
});
|
||||
for (const absence of ['no automated tests', 'zero automated tests', '0 automated tests',
|
||||
'no tests', 'zero tests', '0 tests', 'no automated coverage', 'zero automated coverage', '0 automated coverage'])
|
||||
test(`current test absence: ${absence}`, () => {
|
||||
expect(zeroAbsenceFinding(zeroAbsenceFingerprint(absence))?.seed).toBe('tests');
|
||||
});
|
||||
for (const claim of ['not zero automated tests', 'not 0 automated tests', 'more than zero automated tests',
|
||||
'more than 0 automated tests', 'greater than zero automated tests', 'at least zero automated tests',
|
||||
'not exactly zero automated tests', 'no longer zero automated tests', '"zero automated tests"', '`zero automated tests`'])
|
||||
test(`test absence rejects ${claim}`, () => {
|
||||
expect(zeroAbsenceFinding(zeroAbsenceFingerprint(claim))).toBeNull();
|
||||
});
|
||||
for (const intro of ['Previously the plan shipped', 'The prior plan shipped', 'The old version shipped', 'A historical example shipped'])
|
||||
test(`test absence rejects historical claim: ${intro}`, () => {
|
||||
const question = zeroAbsenceFingerprint(), call = question.nativeCall!;
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('The plan ships a new payment handler', intro + ' a payment handler');
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(zeroAbsenceFinding(question)).toBeNull();
|
||||
});
|
||||
for (const [name, mutate] of Object.entries({
|
||||
'missing row': (s: string) => s.replace(/^\| D4 .*\n/m, ''),
|
||||
'foreign source': (s: string) => s.replaceAll('PLAN.md', 'other.md'),
|
||||
'quoted-only row absence': (s: string) => s.replace('No automated coverage of new handler', '"No automated coverage of new handler"'),
|
||||
'code-only row absence': (s: string) => s.replace('No automated coverage of new handler', '`No automated coverage of new handler`'),
|
||||
'negated row absence': (s: string) => s.replace('No automated coverage of new handler', 'not zero automated coverage of new handler'),
|
||||
})) test(`test absence keeps ${name} rejected`, () => {
|
||||
expect(zeroAbsenceFinding(zeroAbsenceFingerprint(), mutate(zeroAbsenceFixture.savedPlan))).toBeNull();
|
||||
});
|
||||
test('test absence never substitutes for a native acknowledgment', () => {
|
||||
const question = zeroAbsenceFingerprint(); question.nativeCall!.answered = false;
|
||||
expect(zeroAbsenceFinding(question)).toBeNull();
|
||||
});
|
||||
|
||||
for (const [claim, expected] of [
|
||||
['The plan has not currently zero automated tests', false],
|
||||
['The plan does not have zero automated tests', false],
|
||||
['The plan does not currently have zero automated tests', false],
|
||||
['The plan does not have exactly 0 automated tests', false],
|
||||
['The plan does not yet contain 0 automated tests', false],
|
||||
['The number of automated tests is not currently zero automated tests', false],
|
||||
['The plan doesn’t have zero automated tests', false],
|
||||
["The plan doesn't currently provide 0 automated tests", false],
|
||||
['The plan has more than currently zero automated tests', false],
|
||||
['The plan currently ships a new payment handler with zero automated tests', true],
|
||||
['The plan is not ready because it ships a new payment handler with zero automated tests', true],
|
||||
['The plan ships a new payment handler with 0 automated tests', true],
|
||||
] as const) test(`test absence respects quantified negation: ${claim}`, () => {
|
||||
const question = zeroAbsenceFingerprint(), call = question.nativeCall!;
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace(
|
||||
'The plan ships a new payment handler with zero automated tests', claim);
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
expect(zeroAbsenceFinding(question)?.seed === 'tests').toBe(expected);
|
||||
});
|
||||
|
||||
import onboardingFixture from './fixtures/ceo-onboarding-packet-90f.json';
|
||||
const onboarding = () => nativePlanCallFingerprint(clone(onboardingFixture.call), 0, true);
|
||||
const setupCounter = () => createCeoPaymentFindingCounter('', () => { throw new Error('setup must not read a report'); }, ceoFirstReviewAUQ);
|
||||
const refreshPacket = (fp: ReturnType<typeof onboarding>) => {
|
||||
fp.nativeCall!.answers = Object.fromEntries(fp.nativeCall!.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
fp.options = fp.nativeCall!.questions.flatMap(q => q.options.map((o, i) => ({ index: i + 1, label: o.label })));
|
||||
};
|
||||
|
||||
test('actual acknowledged two-question onboarding packet is excluded only after complete packet validation', () => {
|
||||
const fp = onboarding(), counter = setupCounter();
|
||||
expect(fp.nativeCall!.questions.map(q => q.header)).toEqual(['Routing', 'Learnings']);
|
||||
expect(counter.isReviewAUQ(fp)).toBe(false);
|
||||
expect(counter.trace).toEqual([{ signature: fp.signature, kind: 'setup' }]);
|
||||
expect(() => counter.isReviewAUQ(fp, [fp.nativeCall!])).toThrow(/duplicated/);
|
||||
expect(ceoPaymentFinding(fp, fixture.seed, saved())).toBeNull(); // never a single review record
|
||||
for (const q of fp.nativeCall!.questions) {
|
||||
const call = { ...clone(fp.nativeCall!), questions: [q], answers: { [q.question]: fp.nativeCall!.answers[q.question]! } };
|
||||
expect(setupCounter().isReviewAUQ(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('native onboarding accepts complete packets through four questions regardless of tab order', () => {
|
||||
const fp = onboarding(); fp.nativeCall!.questions.reverse(); refreshPacket(fp);
|
||||
expect(setupCounter().isReviewAUQ(fp)).toBe(false);
|
||||
for (const header of ['Scope', 'Mode']) {
|
||||
const q = clone(fp.nativeCall!.questions[0]!);
|
||||
q.header = header; q.question = header === 'Scope' ? 'D8 — Which review target should we use?' : 'D9 — Which review mode should we use?';
|
||||
q.options = (header === 'Scope' ? ['Skip interview and plan immediately', 'Describe the idea inline'] :
|
||||
['HOLD SCOPE', 'SELECTIVE EXPANSION', 'SCOPE EXPANSION', 'SCOPE REDUCTION']).map(label => ({ label, description: '' }));
|
||||
fp.nativeCall!.questions.push(q); refreshPacket(fp);
|
||||
expect(setupCounter().isReviewAUQ(fp)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
for (const [name, mutate] of Object.entries({
|
||||
'unacknowledged packet': (fp: any) => { fp.nativeCall.answered = false; },
|
||||
'failed packet': (fp: any) => { fp.nativeCall.failed = true; },
|
||||
'foreign signature': (fp: any) => { fp.signature = 'foreign:call'; },
|
||||
'missing session': (fp: any) => { fp.nativeCall.sessionId = ''; fp.signature = ':' + fp.nativeCall.toolUseId; },
|
||||
'missing tool identity': (fp: any) => { fp.nativeCall.toolUseId = ''; fp.signature = fp.nativeCall.sessionId + ':'; },
|
||||
'unfinished second tab': (fp: any) => { fp.nativeCall.unansweredQuestionIndices = [1]; },
|
||||
'missing unanswered inventory': (fp: any) => { delete fp.nativeCall.unansweredQuestionIndices; },
|
||||
'missing acknowledgment time': (fp: any) => { delete fp.nativeCall.answeredAt; },
|
||||
'invalid acknowledgment time': (fp: any) => { fp.nativeCall.answeredAt = 'invalid'; },
|
||||
'tab-only index': (fp: any) => { fp.nativeQuestionIndex = 0; },
|
||||
'out-of-bounds tab index': (fp: any) => { fp.nativeQuestionIndex = 9; },
|
||||
'missing second answer': (fp: any) => { delete fp.nativeCall.answers[fp.nativeCall.questions[1].question]; },
|
||||
'foreign answer key': (fp: any) => { const q = fp.nativeCall.questions[1]; delete fp.nativeCall.answers[q.question]; fp.nativeCall.answers.other = q.options[0].label; },
|
||||
'extra answer': (fp: any) => { fp.nativeCall.answers.other = 'extra'; },
|
||||
'unoffered second answer': (fp: any) => { fp.nativeCall.answers[fp.nativeCall.questions[1].question] = 'not offered'; },
|
||||
'duplicate question identity': (fp: any) => { fp.nativeCall.questions[1].question = fp.nativeCall.questions[0].question; refreshPacket(fp); },
|
||||
'multiselect second tab': (fp: any) => { fp.nativeCall.questions[1].multiSelect = true; },
|
||||
'duplicate second-tab options': (fp: any) => { fp.nativeCall.questions[1].options[1].label = fp.nativeCall.questions[1].options[0].label; refreshPacket(fp); },
|
||||
'one second-tab option': (fp: any) => { fp.nativeCall.questions[1].options.pop(); refreshPacket(fp); },
|
||||
'five second-tab options': (fp: any) => { for (const label of ['other3', 'other4', 'other5']) fp.nativeCall.questions[1].options.push({ label }); refreshPacket(fp); },
|
||||
'stale second-tab label': (fp: any) => { fp.options.at(-1).label = 'stale'; },
|
||||
'stale second-tab index': (fp: any) => { fp.options.at(-1).index = 4; },
|
||||
'missing second-tab options': (fp: any) => { fp.options.splice(2); },
|
||||
'mixed setup and review': (fp: any) => { fp.nativeCall.questions[1] = clone(seeded[0]!.call.questions[0]!); refreshPacket(fp); },
|
||||
'two review questions': (fp: any) => { fp.nativeCall.questions = [clone(seeded[0]!.call.questions[0]!), clone(seeded[1]!.call.questions[0]!)]; refreshPacket(fp); },
|
||||
'five native questions': (fp: any) => { for (let i = 0; i < 3; i++) { const q = clone(fp.nativeCall.questions[0]); q.question += ' ' + i; fp.nativeCall.questions.push(q); } refreshPacket(fp); },
|
||||
})) test(`onboarding packet rejects ${name} without reading or counting a review`, () => {
|
||||
const fp = onboarding(), counter = setupCounter(); mutate(fp);
|
||||
expect(() => counter.isReviewAUQ(fp)).toThrow(/Invalid or duplicated completed native decision/);
|
||||
expect(counter.trace).toEqual([]);
|
||||
});
|
||||
@@ -1,228 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import captured from './fixtures/ceo-section-choice-ai.json';
|
||||
import metadataCaptured from './fixtures/ceo-metadata-brief-ax.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
function call(index = 4): any {
|
||||
const source = structuredClone(captured.calls[index]!);
|
||||
return { sessionId: source.sessionId, toolUseId: source.toolUseId, questions: source.questions,
|
||||
answered: true, failed: false, unansweredQuestionIndices: [], answeredAt: source.answeredAt,
|
||||
answers: Object.fromEntries(source.questions.map((q, i) => [q.question, source.answers[i]])) };
|
||||
}
|
||||
const fp = (c: any) => nativePlanCallFingerprint(c, 0, true);
|
||||
function edit(c: any, change: (text: string) => string) {
|
||||
const q = c.questions[0], answer = c.answers[q.question];
|
||||
q.question = change(q.question); c.answers = { [q.question]: answer };
|
||||
}
|
||||
|
||||
test('exact captured section choices start review; preceding actual setup does not', () => {
|
||||
let started = false; const classified: boolean[] = [];
|
||||
for (let i = 0; i < captured.calls.length; i++) {
|
||||
const question = fp(call(i));
|
||||
expect(ceoFirstReviewAUQ(question)).toBe(captured.calls[i]!.expectedFirstReview);
|
||||
const phase = planCountQuestionPhase(question, started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted; classified.push(phase.preReview);
|
||||
}
|
||||
expect(classified).toEqual([true, true, true, true, false, false, false, false, false]);
|
||||
});
|
||||
|
||||
test('an offered alternative and a quoted historical withdrawal retain current review identity', () => {
|
||||
const alternative = call(); alternative.answers[alternative.questions[0].question] = alternative.questions[0].options[1].label;
|
||||
expect(ceoFirstReviewAUQ(fp(alternative))).toBe(true);
|
||||
const quoted = call(); edit(quoted, s => s + '\nHistorical quote: "This issue has been resolved."');
|
||||
expect(ceoFirstReviewAUQ(fp(quoted))).toBe(true);
|
||||
});
|
||||
|
||||
test.each([
|
||||
['pending', (c: any) => { c.answered = false; }],
|
||||
['failed', (c: any) => { c.failed = true; }],
|
||||
['unanswered index', (c: any) => { c.unansweredQuestionIndices = [0]; }],
|
||||
['missing session', (c: any) => { c.sessionId = ''; }],
|
||||
['missing tool id', (c: any) => { c.toolUseId = ''; }],
|
||||
['unoffered answer', (c: any) => { c.answers[c.questions[0].question] = 'Not offered'; }],
|
||||
['missing answer', (c: any) => { c.answers = {}; }],
|
||||
['mixed packet', (c: any) => { c.questions.push(structuredClone(c.questions[0])); }],
|
||||
['multi-select', (c: any) => { c.questions[0].multiSelect = true; }],
|
||||
['duplicate options', (c: any) => { c.questions[0].options[1] = structuredClone(c.questions[0].options[0]); }],
|
||||
['missing description', (c: any) => { c.questions[0].options[1].description = ''; }],
|
||||
['option identity', (c: any) => { c.questions[0].options[1].label = '1B) Other'; }],
|
||||
['section mismatch', (c: any) => edit(c, s => s.replace('Section 1 Architecture.', 'Section 2 Architecture.'))],
|
||||
['recommendation mismatch', (c: any) => edit(c, s => s.replace('Recommendation: A', 'Recommendation: B'))],
|
||||
['missing stakes', (c: any) => edit(c, s => s.replace(/^Stakes if we pick wrong:.*$/m, ''))],
|
||||
['duplicate assessment', (c: any) => edit(c, s => s + '\nELI10: A second competing assessment.')],
|
||||
['quoted assessment', (c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'))],
|
||||
['fenced context', (c: any) => edit(c, s => s.replace(/^(Project\/branch\/task:.*)$/m, '```\n$1\n```'))],
|
||||
['Suppose assessment', (c: any) => edit(c, s => s.replace('ELI10: The plan', 'ELI10: Suppose the plan'))],
|
||||
['single quoted assessment', (c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, "ELI10: '$1'"))],
|
||||
['current withdrawal', (c: any) => edit(c, s => s + '\nThis issue is withdrawn.')],
|
||||
['completed withdrawal', (c: any) => edit(c, s => s + '\nWe have withdrawn this finding.')],
|
||||
['conditional assessment', (c: any) => edit(c, s => s.replace('ELI10: The plan', 'ELI10: If the plan'))],
|
||||
['withdrawn issue', (c: any) => edit(c, s => s + '\nWe withdraw this finding.')],
|
||||
['resolved issue', (c: any) => edit(c, s => s + '\nThis issue has been resolved.')],
|
||||
['administrative report', (c: any) => edit(c, s => s.replace(/^.*\n/, '1A — Should the completed review report be saved?\n'))],
|
||||
['setup header', (c: any) => { c.questions[0].header = 'Setup'; }],
|
||||
['borrowed qid', (c: any) => edit(c, s => s + '\n<gstack-qid:plan-ceo-review-example>')],
|
||||
])('rejects %s despite numbered review prose', (_name, mutate) => {
|
||||
const c = call(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
});
|
||||
|
||||
test('fingerprints cannot borrow another native call or its options', () => {
|
||||
const original = fp(call());
|
||||
expect(ceoFirstReviewAUQ({ ...original, signature: 'other:tool' })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...original, nativeCall: undefined })).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({ ...original, options: original.options.slice(1) })).toBe(false);
|
||||
});
|
||||
|
||||
test('the regression inputs belong to the existing paid CEO case', () => {
|
||||
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/ceo-section-choice-ai.test.ts');
|
||||
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/fixtures/ceo-section-choice-ai.json');
|
||||
});
|
||||
|
||||
test('coherent finished-note destination is administrative, despite matching section and choice', () => {
|
||||
const c = call(), q = c.questions[0];
|
||||
q.header = 'Destination';
|
||||
q.question = '1A — Which storage location should hold these notes?\nProject/branch/task: main, Stripe payment webhook plan, Section 1 Architecture.\nELI10: The review is finished; these notes can be saved in either folder for convenience.\nStakes if we pick wrong: People may have to look in a second folder.\nRecommendation: A because the existing folder is easier to find.';
|
||||
q.options = [{label:'A) Save beside the plan',description:'Keeps the finished notes together.'},{label:'B) Save in another folder',description:'Keeps finished notes separate.'}];
|
||||
c.answers = {[q.question]: q.options[0].label};
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
});
|
||||
|
||||
test('conditional stakes remain valid when the assessment asserts the current gap', () => {
|
||||
const c = call(); edit(c, s => s.replace('Stakes if we pick wrong:', 'Stakes if we pick wrong: Suppose there were an issue.'));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
});
|
||||
|
||||
test('negated gaps and administrative missing fields cannot borrow review identity', () => {
|
||||
const negated = call(); edit(negated, s => s.replace(/^ELI10: .+$/m, 'ELI10: The transaction order is not unspecified. The plan guarantees commit before email.'));
|
||||
expect(ceoFirstReviewAUQ(fp(negated))).toBe(false);
|
||||
const admin = call(), q = admin.questions[0];
|
||||
q.header = 'Destination';
|
||||
q.question = '1A — Which storage location should hold these notes?\nProject/branch/task: main, Stripe payment webhook plan, Section 1 Architecture.\nELI10: These finished notes have a missing storage location.\nStakes if we pick wrong: People may look in the wrong folder.\nRecommendation: A because a notes folder is easy to find.';
|
||||
q.options = [{label:'A) Add a notes folder',description:'Save the finished notes together.'},{label:'B) Use the existing folder',description:'No new folder.'}];
|
||||
admin.answers = {[q.question]: q.options[0].label};
|
||||
expect(ceoFirstReviewAUQ(fp(admin))).toBe(false);
|
||||
});
|
||||
|
||||
function metadataCall(): any {
|
||||
const c = structuredClone(metadataCaptured.call);
|
||||
return { ...c, answered: true, failed: false, unansweredQuestionIndices: [],
|
||||
answers: { [c.questions[0]!.question]: metadataCaptured.answer } };
|
||||
}
|
||||
|
||||
test('AX ordinary D-number question keeps its exact completed review identity', () => {
|
||||
const c = metadataCall(), before = JSON.stringify(c);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ)).toMatchObject({ preReview: false, reviewStarted: true });
|
||||
expect(JSON.stringify(c)).toBe(before);
|
||||
expect(E2E_TOUCHFILES['plan-ceo-finding-count']).toContain('test/fixtures/ceo-metadata-brief-ax.json');
|
||||
});
|
||||
|
||||
test('decision counter, review name and an alternative selection do not dictate the finding', () => {
|
||||
const c = metadataCall(); edit(c, s => s.replace(/^D5 /, 'D17 ').replace('Section 2 (Error & Rescue Map)', 'Section 3 (Failure Handling)'));
|
||||
c.answers[c.questions[0].question] = c.questions[0].options[1].label;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
edit(c, s => s + '\nHistorical quote: "This finding is withdrawn."');
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
});
|
||||
|
||||
test('metadata cannot replace native completion, current context or a real defect', () => {
|
||||
for (const mutate of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.answeredAt = 'not a date'; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.answers = { other: metadataCaptured.answer }; },
|
||||
(c: any) => { c.questions[0].header = 'Setup'; },
|
||||
(c: any) => { c.questions[0].header = 'Section 9'; },
|
||||
(c: any) => edit(c, s => s.replace('Section 2 (Error & Rescue Map)', 'Section 2 (Error & Rescue Map), Section 3 (Security)')),
|
||||
(c: any) => edit(c, s => s.replace('of the CEO review', 'of an earlier CEO review')),
|
||||
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task: (.*)$/m, 'Project/branch/task: If approved, $1')),
|
||||
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task:.*\n/m, '')),
|
||||
(c: any) => edit(c, s => s.replace(/^ELI10: (.*)$/m, 'ELI10: "$1"')),
|
||||
(c: any) => edit(c, s => s.replace(/^ELI10: .+$/m, 'ELI10: The handler has no current defect and needs no amendment.')),
|
||||
(c: any) => edit(c, s => s.replace(/^ELI10: .+$/m, 'ELI10: The handler commits before mail and already rescues every required error.')),
|
||||
(c: any) => edit(c, s => s.replace('ELI10: After', 'ELI10: Hypothetical example: after')),
|
||||
(c: any) => edit(c, s => s + '\nThis finding is withdrawn.'),
|
||||
(c: any) => edit(c, s => s + '; This finding is `no longer current`.'),
|
||||
(c: any) => edit(c, s => s + '\nThis finding is unproven.'),
|
||||
]) {
|
||||
const c = metadataCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
expect(ceoFirstReviewAUQ({ ...fp(metadataCall()), signature: 'foreign:call' })).toBe(false);
|
||||
});
|
||||
|
||||
test('a missing or withdrawn offered remedy cannot borrow metadata or an old assessment', () => {
|
||||
for (const mutate of [
|
||||
(c: any) => { c.questions[0].options = [{ label: 'A: Keep the current handler', description: 'No code change.' }, { label: 'B: Save the review notes', description: 'Archive the current report.' }]; c.answers = { [c.questions[0].question]: c.questions[0].options[0].label }; },
|
||||
(c: any) => { c.questions[0].options[0].description += '; This option is `withdrawn`.'; c.questions[0].options[2].description += '\nThis option is withdrawn.'; },
|
||||
(c: any) => { c.questions[0].options.forEach((o: any) => { o.description = 'Hypothetical example. ' + o.description; }); },
|
||||
]) {
|
||||
const c = metadataCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
function metadataRetryCall(): any {
|
||||
const c = structuredClone(metadataCaptured.retry.call);
|
||||
return { ...c, answered: true, failed: false, unansweredQuestionIndices: [],
|
||||
answers: { [c.questions[0]!.question]: metadataCaptured.retry.answer } };
|
||||
}
|
||||
|
||||
test('the separately failed AX retry binds its Issue annotation, bare choices and named plan', () => {
|
||||
const c = metadataRetryCall(), before = JSON.stringify(c);
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
expect(planCountQuestionPhase(fp(c), false, ceoStep0Boundary, ceoFirstReviewAUQ))
|
||||
.toMatchObject({ preReview: false, reviewStarted: true });
|
||||
expect(JSON.stringify(c)).toBe(before);
|
||||
c.answers[c.questions[0].question] = c.questions[0].options[2].label;
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
});
|
||||
|
||||
test('a reviewed filename and section identity can be consistently renamed', () => {
|
||||
const c = metadataRetryCall();
|
||||
edit(c, s => s.replace(/^D4 /, 'D12 ').replace(/Issue 2\.1/, 'Issue 8.3')
|
||||
.replace('Section 2 (Error & Rescue Map)', 'Section 8 (Notification Handling)')
|
||||
.replace(/PLAN\.md/g, 'plans/checkout-flow.md'));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
edit(c, s => s + '\nHistorical quote: "This issue is withdrawn."');
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
edit(c, s => s.replace(/\s*<gstack-qid:[^>]+>/, ''));
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(true);
|
||||
});
|
||||
|
||||
test('retry metadata cannot borrow a foreign section, plan, source or incomplete native call', () => {
|
||||
for (const [index, mutate] of [
|
||||
(c: any) => { c.answered = false; },
|
||||
(c: any) => { c.failed = true; },
|
||||
(c: any) => { c.answeredAt = 'unknown'; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: any) => { c.questions[0].header = 'Issue 8.1'; },
|
||||
(c: any) => edit(c, s => s.replace('Issue 2.1', 'Issue 3.1')),
|
||||
(c: any) => edit(c, s => s.replace('Section 2 (Error & Rescue Map)', 'Section 3 (Security)')),
|
||||
(c: any) => edit(c, s => s.replace('CEO review of PLAN.md,', 'CEO review of DIFFERENT.md,')),
|
||||
(c: any) => edit(c, s => s.replace("PLAN.md says 'no error handling on the email leg'", "OTHER.md says 'no error handling on the email leg'")),
|
||||
(c: any) => edit(c, s => s.replace('CEO review of PLAN.md,', 'Historical CEO review of PLAN.md,')),
|
||||
(c: any) => edit(c, s => s.replace(/^Project\/branch\/task: (.+)$/m, 'Project/branch/task: If approved, $1')),
|
||||
(c: any) => edit(c, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"')),
|
||||
(c: any) => edit(c, s => s.replace('ELI10: The handler', 'ELI10: Source excerpt: the handler')),
|
||||
(c: any) => edit(c, s => s.replace('PLAN.md says', 'If approved, PLAN.md says')),
|
||||
(c: any) => edit(c, s => s.replace('PLAN.md says', 'PLAN.md does not say')),
|
||||
(c: any) => edit(c, s => s.replace('plan-ceo-review-mail-rescue', 'plan-ceo-review-setup')),
|
||||
(c: any) => edit(c, s => s + '\n<gstack-qid:plan-ceo-review-other>'),
|
||||
].entries()) {
|
||||
const c = metadataRetryCall(); mutate(c); expect(ceoFirstReviewAUQ(fp(c)), `retry mutation ${index}`).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('current withdrawal and a withdrawn offered amendment override the retry brief', () => {
|
||||
for (const change of [
|
||||
(s: string) => s + '\nThis issue is withdrawn.',
|
||||
(s: string) => s + '; This finding is `no longer current`.',
|
||||
(s: string) => s + '\nIssue 2.1 is withdrawn.',
|
||||
(s: string) => s.replace(/^ELI10: .+$/m, 'ELI10: The handler has no current defect and needs no amendment.'),
|
||||
]) {
|
||||
const c = metadataRetryCall(); edit(c, change); expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
}
|
||||
const c = metadataRetryCall(); c.questions[0].options[0].description += '; This option is `withdrawn`.';
|
||||
expect(ceoFirstReviewAUQ(fp(c))).toBe(false);
|
||||
});
|
||||
@@ -1,32 +0,0 @@
|
||||
import {describe,test,expect} from 'bun:test';
|
||||
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
|
||||
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import fixture from './fixtures/ceo-section-declarative-ar.json';
|
||||
const calls=()=>structuredClone(fixture.calls) as NativePlanQuestionCall[],first=()=>calls()[2]!;
|
||||
const fp=(c=first())=>nativePlanCallFingerprint(c,0,true),classify=(c=first())=>ceoFirstReviewAUQ(fp(c));
|
||||
const mutate=(fn:(c:NativePlanQuestionCall)=>void)=>{const c=first();fn(c);return c;};
|
||||
const text=(fn:(s:string)=>string)=>mutate(c=>{const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};});
|
||||
const allOptions=(fn:(label:string,description:string)=>{label:string;description:string})=>mutate(c=>{const q=c.questions[0]!,selected=q.options.findIndex(o=>o.label===c.answers![q.question]);q.options=q.options.map(o=>fn(o.label,o.description??''));c.answers={[q.question]:q.options[selected]!.label};});
|
||||
describe('AR completed declarative Section finding',()=>{
|
||||
test('exact public calls enter review after genuine setup',()=>{let started=false;const phases=calls().map(c=>{const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ);started=p.reviewStarted;return p.preReview;});expect(phases).toEqual([true,true,false,false,false]);expect(classify(calls()[0])).toBe(false);expect(classify(calls()[1])).toBe(false);expect(classify(calls()[2])).toBe(true);expect(classify(calls()[3])).toBe(true);});
|
||||
test('comma and question punctuation are presentation',()=>{for(const c of [first(),text(s=>s.replace('Section 6, finding','Section 6 finding')),text(s=>s.replace('receipt is truthy\n','receipt is truthy?\n')),text(s=>s.replace('Section 6, finding','Section 6 finding').replace('receipt is truthy\n','receipt is truthy?\n'))])expect(classify(c)).toBe(true);expect(classify(text(s=>s.replace('D3 — Section 6, finding 1:','D8 — Section 2, finding 3:')))).toBe(true);});
|
||||
test('native successful completion and exact ownership stay required',()=>{
|
||||
for(const fn of [(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.answeredAt;},(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers!['foreign']='foreign';},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered'};},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;}])expect(classify(mutate(fn))).toBe(false);
|
||||
for(const f of [{...fp(),signature:'foreign:tool'},{...fp(),nativeQuestionIndex:1},{...fp(),options:fp().options.toReversed()}])expect(ceoFirstReviewAUQ(f)).toBe(false);
|
||||
});
|
||||
test('malformed identities and setup headers stay closed',()=>{for(const [a,b] of [['D3 —','D0 —'],['D3 —','D03 —'],['Section 6,','Section 06,'],['finding 1:','finding 0:'],['finding 1:','finding 1.2:'],['Section 6,','Section 6,,']])expect(classify(text(s=>s.replace(a,b)))).toBe(false);for(const h of ['Section 7','Finding 9','Routing','Approach'])expect(classify(mutate(c=>{c.questions[0]!.header=h;}))).toBe(false);});
|
||||
test('source, conditional and duplicate assessment metadata stay closed',()=>{for(const field of ['Project/branch/task: ','ELI10: '])for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, '])expect(classify(text(s=>s.replace(field,field+p)))).toBe(false);for(const prefix of ['Source:\n','Earlier review assessment:\n','Project/branch/task: duplicate\n','ELI10: duplicate\n'])expect(classify(text(s=>s.replace('ELI10:',prefix+'ELI10:')))).toBe(false);expect(classify(text(s=>s.replace(/^Project\/branch\/task:.*\n/m,'')))).toBe(false);expect(classify(text(s=>'```\n'+s+'\n```'))).toBe(false);});
|
||||
test('withdrawn assessment or source-only options cannot establish a current decision',()=>{expect(classify(text(s=>s+'\nThis finding is withdrawn.'))).toBe(false);expect(classify(text(s=>s+'\nCorrection: this finding is "withdrawn".'))).toBe(false);for(const p of ['Source excerpt: ','If approved, '])expect(classify(allOptions((label,description)=>({label:label.replace(/^([A-Z]\) )/,'$1'+p),description:p+description})))).toBe(false);expect(classify(allOptions((label,description)=>({label,description:description+' This amendment is withdrawn.'})))).toBe(false);expect(classify(allOptions((label)=>({label:label.replace(/^([A-Z]\) ).*/,'$1Keep current assertion'),description:'Leave the current assertion unchanged.'})))).toBe(false);});
|
||||
test('superseded or conditional findings and offered actions are not current',()=>{
|
||||
for(const status of ['superseded','"superseded"','no longer current','"no longer current"']){
|
||||
expect(classify(text(s=>s+'\nThis finding is '+status+'.'))).toBe(false);
|
||||
expect(classify(allOptions((label,description)=>({label,description:description+' This amendment is '+status+'.'})))).toBe(false);
|
||||
}
|
||||
for(const prefix of ['Assuming approval, ','Provided approval, ']){
|
||||
expect(classify(text(s=>s.replace('ELI10: ','ELI10: '+prefix)))).toBe(false);
|
||||
expect(classify(allOptions((label,description)=>({label,description:prefix+description})))).toBe(false);
|
||||
}
|
||||
for(const history of ['> This finding is superseded.','Archived note: "This finding is superseded."','Archived note: "This finding is no longer current."','~~~\nThis finding is superseded.\n~~~'])expect(classify(text(s=>s+'\n'+history))).toBe(true);
|
||||
});
|
||||
test('quoted archive and selected opposed option remain valid',()=>{expect(classify(text(s=>s+'\nArchived note: "This finding is withdrawn."'))).toBe(true);expect(classify(mutate(c=>{const q=c.questions[0]!;c.answers={[q.question]:q.options[2]!.label};}))).toBe(true);});
|
||||
});
|
||||
@@ -1,38 +0,0 @@
|
||||
import {describe,test,expect} from 'bun:test';
|
||||
import {ceoFirstReviewAUQ,ceoStep0Boundary,nativePlanCallFingerprint,planCountQuestionPhase} from './helpers/claude-pty-runner';
|
||||
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import fixture from './fixtures/ceo-section-ordering-aq.json';
|
||||
const calls=()=>structuredClone(fixture.calls) as NativePlanQuestionCall[];const first=()=>calls()[2]!;
|
||||
const fp=(c=first())=>nativePlanCallFingerprint(c,0,true);const classify=(c=first())=>ceoFirstReviewAUQ(fp(c));
|
||||
function mutate(fn:(c:NativePlanQuestionCall)=>void){const c=first();fn(c);return c;}
|
||||
function text(fn:(s:string)=>string){return mutate(c=>{const q=c.questions[0]!,a=c.answers![q.question]!;q.question=fn(q.question);c.answers={[q.question]:a};});}
|
||||
describe('AQ owned Section architecture ordering brief',()=>{
|
||||
test('exact seven calls open review at D4 and retain prior setup',()=>{let started=false;const phases=calls().map(c=>{const p=planCountQuestionPhase(fp(c),started,ceoStep0Boundary,ceoFirstReviewAUQ);started=p.reviewStarted;return p.preReview;});expect(phases).toEqual([true,true,false,false,false,false,false]);expect(classify()).toBe(true);});
|
||||
test('separate counters, choice order and selected option remain valid',()=>{const c=text(s=>s.replace('D4 — Section 1 (Architecture), issue 1:','D9 — Section 3 (Architecture), issue 2:').replace(/\b1([ABC])\b/g,'2$1')),q=c.questions[0]!;for(const o of q.options)o.label=o.label.replace(/^1/,'2');q.options.reverse();for(const o of q.options){c.answers={[q.question]:o.label};expect(classify(c)).toBe(true);}});
|
||||
test('quoted archive and conditional consequences do not cancel current evidence',()=>{expect(classify(text(s=>s+'\nArchived note: "This finding is withdrawn."'))).toBe(true);expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=' "Earlier review assessment: This remedy is withdrawn."';}))).toBe(true);});
|
||||
test('native completion and menu ownership remain required',()=>{
|
||||
for(const fn of [(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},(c:NativePlanQuestionCall)=>{delete c.answeredAt;},(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers!['foreign']='foreign';},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered'};},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;}])expect(classify(mutate(fn))).toBe(false);
|
||||
for(const f of [{...fp(),signature:'foreign:tool'},{...fp(),nativeQuestionIndex:1},{...fp(),options:fp().options.toReversed()}])expect(ceoFirstReviewAUQ(f)).toBe(false);
|
||||
});
|
||||
test('malformed or competing identities and setup headers fail closed',()=>{
|
||||
for(const [a,b] of [['D4 —','D04 —'],['Section 1 (','Section 01 ('],['issue 1:','issue 0:'],['issue 1:','issue 1.2:'],['(Architecture)','(Source excerpt)']])expect(classify(text(s=>s.replace(a,b)))).toBe(false);
|
||||
for(const h of ['Section 9','Issue 9','Section 01','Routing','Approach'])expect(classify(mutate(c=>{c.questions[0]!.header=h;}))).toBe(false);
|
||||
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.label=c.questions[0]!.options[0]!.label.replace('1A)','2A)');c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};}))).toBe(false);
|
||||
});
|
||||
test('unique current context and assessment cannot come from source or a conditional',()=>{
|
||||
for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, ','Provided approval, ','Assuming approval, '])for(const field of ['Project/branch/task: ','ELI10: '])expect(classify(text(s=>s.replace(field,field+p)))).toBe(false);
|
||||
for(const p of ['Source:\n','Earlier review assessment:\n','Project/branch/task: duplicate\n','ELI10: duplicate\n'])expect(classify(text(s=>s.replace('ELI10:',p+'ELI10:')))).toBe(false);
|
||||
expect(classify(text(s=>s.replace(/^Project\/branch\/task:.*\n/m,'')))).toBe(false);
|
||||
});
|
||||
test('current gap and actual commit-then-notify action stay mandatory',()=>{
|
||||
expect(classify(text(s=>s.replace('but never says whether the email runs inside the database transaction or after it commits','and explicitly specifies that email follows the database commit')))).toBe(false);
|
||||
for(const [a,b] of [['COMMIT, then call the mail client','call the mail client, then COMMIT'],['Load user and orders, assign payment_status=paid and PaymentIntent ID, COMMIT, then call the mail client','Record this plan as complete'],['Mail failure can never roll back a committed payment','Mail failure can roll back the payment']])expect(classify(mutate(c=>{const o=c.questions[0]!.options[0]!;o.description=o.description!.replace(a,b);}))).toBe(false);
|
||||
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.label='1A) Save the review';c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};}))).toBe(false);
|
||||
});
|
||||
test('direct current status and action cancellation close their owners',()=>{
|
||||
for(const s of ['withdrawn','superseded','"closed"','“withdrawn”']){expect(classify(text(t=>t+` This finding is ${s}.`))).toBe(false);expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=` This amendment is ${s}.`;}))).toBe(false);}
|
||||
expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=' Correction: do not commit before sending email.';}))).toBe(false);
|
||||
expect(classify(text(s=>s+' Correction: this ordering gap is resolved.'))).toBe(false);
|
||||
});
|
||||
test('source or conditional options cannot supply the amendment',()=>{for(const p of ['Source excerpt: ','Earlier review assessment: ','If approved, ','Provided approval, ']){expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description=p+c.questions[0]!.options[0]!.description;}))).toBe(false);expect(classify(mutate(c=>{const q=c.questions[0]!;q.options[0]!.label=q.options[0]!.label.replace('1A) ','1A) '+p);c.answers={[q.question]:q.options[0]!.label};}))).toBe(false);}});
|
||||
});
|
||||
@@ -1,91 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
import captured from './fixtures/ceo-section-parenthesis-at.json';
|
||||
const call = (): any => structuredClone(captured.calls[1]);
|
||||
const fp = (value: any) => nativePlanCallFingerprint(value, 0, true);
|
||||
const matches = (value: any) => ceoFirstReviewAUQ(fp(value));
|
||||
function edit(value: any, change: (text: string) => string) {
|
||||
const q = value.questions[0], answer = value.answers[q.question];
|
||||
q.question = change(q.question); value.answers = { [q.question]: answer };
|
||||
}
|
||||
test('the exact completed combined section/finding brief opens the retry review', () => {
|
||||
const value = call(); expect(matches(value)).toBe(true); expect(value).toEqual(captured.calls[1]);
|
||||
let started = false;
|
||||
const phases = captured.calls.map(value => {
|
||||
const phase = planCountQuestionPhase(fp(value), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted; return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, false, false, false, false, false, false]);
|
||||
});
|
||||
test('decision, section, finding and descriptive header retain separate identities', () => {
|
||||
for (const change of [
|
||||
(text: string) => text.replace(/^D4/, 'D19'),
|
||||
(text: string) => text.replace('Section 1, finding 1', 'Section 7, finding 1').replace('review-s1-', 'review-s7-'),
|
||||
(text: string) => text.replace(') — ', ') - '),
|
||||
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(true); }
|
||||
for (const header of ['Receipt rescue', 'Error contract', 'Finding 1', 'Issue 1', 'Section 1', 'Section 1 finding 1']) {
|
||||
const value = call(); value.questions[0].header = header; expect(matches(value)).toBe(true);
|
||||
}
|
||||
for (const option of call().questions[0].options) {
|
||||
const value = call(); value.answers[value.questions[0].question] = option.label; expect(matches(value)).toBe(true);
|
||||
}
|
||||
});
|
||||
test('conflicting annotation, qid, header and option identities cannot open review', () => {
|
||||
for (const change of [
|
||||
(text: string) => text.replace('Section 1, finding 1', 'Section 0, finding 1'),
|
||||
(text: string) => text.replace('Section 1, finding 1', 'Section 1, finding 0'),
|
||||
(text: string) => text.replace('Section 1, finding 1', 'Section 1, finding 2'),
|
||||
(text: string) => text.replace('review-s1-', 'review-s9-'),
|
||||
(text: string) => text.replace('plan-ceo-review-s1-', 'plan-eng-review-s1-'),
|
||||
(text: string) => text.replace('Section 1, finding 1', 'Section 1, hypothetical finding 1'),
|
||||
(text: string) => text.replace(/^Recommendation: 1A/m, 'Recommendation: 9A'),
|
||||
(text: string) => text + '\n<gstack-qid:plan-ceo-review-s1-other>',
|
||||
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
|
||||
for (const header of ['Finding 9', 'Finding one', 'Issue 9', 'Section 9', 'Section 1 finding 9', 'Section one', 'Approach']) {
|
||||
const value = call(); value.questions[0].header = header; expect(matches(value)).toBe(false);
|
||||
}
|
||||
});
|
||||
test('only current owned assessments and offered amendments supply coverage', () => {
|
||||
for (const change of [
|
||||
(text: string) => 'Example: ' + text,
|
||||
(text: string) => '> ' + text,
|
||||
(text: string) => '```\n' + text + '\n```',
|
||||
(text: string) => text.replace('\nProject/branch/task:', '\nSource:\nProject/branch/task:'),
|
||||
(text: string) => text.replace('Project/branch/task: ', 'Project/branch/task: If approved, '),
|
||||
(text: string) => text.replace(/^ELI10: /m, 'ELI10: If approved, '),
|
||||
(text: string) => text.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(text: string) => text.replace(/^ELI10: .+$/m, 'ELI10: This handler has no current defect and needs no amendment.'),
|
||||
(text: string) => text + '\nThis finding is withdrawn.',
|
||||
(text: string) => text + '\nThis finding is "withdrawn".',
|
||||
(text: string) => text + '\nThis finding is no longer current.',
|
||||
]) { const value = call(); edit(value, change); expect(matches(value)).toBe(false); }
|
||||
for (const prefix of ['Source: ', 'If approved, ', 'This remedy is withdrawn. ']) {
|
||||
const value = call(); value.questions[0].options.forEach((option: any) => { option.description = prefix + option.description; });
|
||||
expect(matches(value)).toBe(false);
|
||||
}
|
||||
const history = call(); edit(history, text => text + '\nOld note: "This finding is withdrawn."'); expect(matches(history)).toBe(true);
|
||||
});
|
||||
test('native completion and exact same-call options remain required', () => {
|
||||
for (const change of [
|
||||
(value: any) => { value.answered = false; },
|
||||
(value: any) => { value.failed = true; },
|
||||
(value: any) => { value.sessionId = ''; },
|
||||
(value: any) => { value.toolUseId = ''; },
|
||||
(value: any) => { value.answeredAt = 'invalid'; },
|
||||
(value: any) => { value.unansweredQuestionIndices = [0]; },
|
||||
(value: any) => { value.answers = {}; },
|
||||
(value: any) => { value.answers[value.questions[0].question] = 'Foreign answer'; },
|
||||
(value: any) => { value.questions[0].multiSelect = true; },
|
||||
(value: any) => { value.questions.push(structuredClone(value.questions[0])); },
|
||||
(value: any) => { value.questions[0].options[1].description = ''; },
|
||||
(value: any) => { value.questions[0].options[1].label = '9B: Foreign amendment'; },
|
||||
]) { const value = call(); change(value); expect(matches(value)).toBe(false); }
|
||||
const original = fp(call());
|
||||
for (const value of [{ ...original, signature: 'foreign:call' }, { ...original, nativeCall: undefined },
|
||||
{ ...original, nativeQuestionIndex: 1 }, { ...original, options: original.options.slice(1) }]) expect(ceoFirstReviewAUQ(value)).toBe(false);
|
||||
});
|
||||
test('new public artifacts select only CEO finding count', () => {
|
||||
for (const file of ['test/ceo-section-parenthesis-at.test.ts', 'test/fixtures/ceo-section-parenthesis-at.json'])
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(file)).map(([name]) => name)).toEqual(['plan-ceo-finding-count']);
|
||||
});
|
||||
@@ -1,93 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import fixture from './fixtures/ceo-sequence-aq.json';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const accepts = (call: any) => ceoFirstReviewAUQ(nativePlanCallFingerprint(call, 0, true));
|
||||
function changed(edit: (q: any, call: any) => void) {
|
||||
const call = structuredClone(fixture.calls[2]!), q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers[q.question]);
|
||||
edit(q, call);
|
||||
call.answers = { [q.question]: q.options[selected]?.label ?? '' };
|
||||
return call;
|
||||
}
|
||||
test('exact completed prefix keeps setup and approach before the current sequence finding', () => {
|
||||
expect(fixture.calls.map(accepts)).toEqual([false, false, true]);
|
||||
});
|
||||
test('equivalent decision identities and explicit current sequencing gaps retain the finding', () => {
|
||||
for (const gap of [
|
||||
'the plan never defines the sequence or the transaction boundary.',
|
||||
'this plan does not specify the order and the commit point.',
|
||||
]) expect(accepts(changed(q => {
|
||||
q.question = q.question.replace(/^Project\/branch\/task:.*$/m, 'Project/branch/task: main, PLAN.md; '+gap);
|
||||
}))).toBe(true);
|
||||
expect(accepts(changed(q => { q.question=q.question.replace(/^D2 —/, 'd19 -');q.header='d19 Order'; }))).toBe(true);
|
||||
});
|
||||
test('native completion, matching identities and selected offered answer remain mandatory', () => {
|
||||
for (const edit of [
|
||||
(_q:any,c:any)=>{c.answered=false;}, (_q:any,c:any)=>{c.failed=true;},
|
||||
(_q:any,c:any)=>{c.unansweredQuestionIndices=[0];}, (_q:any,c:any)=>{c.answeredAt='invalid';},
|
||||
(q:any)=>{q.header='D3 Sequence';}, (q:any)=>{q.header='D2 Approach';},
|
||||
(q:any)=>{q.multiSelect=true;}, (q:any)=>{q.question=q.question.replace('Recommendation: A','Recommendation: Z');},
|
||||
(q:any)=>{q.options[1].label=q.options[1].label.replace('B)','A)');},
|
||||
]) expect(accepts(changed(edit))).toBe(false);
|
||||
const noAnswer=changed(()=>{});noAnswer.answers={};expect(accepts(noAnswer)).toBe(false);
|
||||
const fp=nativePlanCallFingerprint(changed(()=>{}),0,true);
|
||||
expect(ceoFirstReviewAUQ({...fp,signature:'foreign:call'})).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({...fp,nativeCall:undefined})).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({...fp,nativeQuestionIndex:1})).toBe(false);
|
||||
expect(ceoFirstReviewAUQ({...fp,options:fp.options.map((o,i)=>i===0?{...o,label:'Foreign choice'}:o)})).toBe(false);
|
||||
});
|
||||
test('current metadata cannot be replaced by source, history, conditional or duplicate ownership', () => {
|
||||
for(const prefix of ['Source excerpt: ', 'Earlier review assessment: ', 'If approved, ', 'For historical context, ']) {
|
||||
expect(accepts(changed(q=>{q.question=q.question.replace('Project/branch/task: ','Project/branch/task: '+prefix);}))).toBe(false);
|
||||
expect(accepts(changed(q=>{q.question=q.question.replace('ELI10: ','ELI10: '+prefix);}))).toBe(false);
|
||||
}
|
||||
for(const edit of [
|
||||
(q:any)=>{q.question=q.question.replace('the plan lists','the previous plan lists');},
|
||||
(q:any)=>{q.question=q.question.replace('but never fixes','and now defines');},
|
||||
(q:any)=>{q.question=q.question.replace(/^Project\/branch\/task:.*$/m,'Project/branch/task: main, PLAN.md; no current sequencing gap.');},
|
||||
(q:any)=>{q.question=q.question.replace('\nELI10:','\nSource:\nELI10:');},
|
||||
(q:any)=>{q.question=q.question.replace('\nELI10:','\nProject/branch/task: another plan\nELI10:');},
|
||||
]) expect(accepts(changed(edit))).toBe(false);
|
||||
});
|
||||
test('current withdrawals and a missing commit-first remedy or opposed risk remain setup', () => {
|
||||
for(const status of ['This finding is withdrawn.','This finding is "closed".','There is no current gap.',
|
||||
'The gap is resolved.', 'This sequence has been fixed.', 'This transaction boundary is "defined".'])
|
||||
expect(accepts(changed(q=>{q.question+='\n'+status;}))).toBe(false);
|
||||
for(const edit of [
|
||||
(q:any)=>{q.options[0].label='A) Archive the plan (Recommended)';},
|
||||
(q:any)=>{q.options[0].description='Source excerpt: '+q.options[0].description;},
|
||||
(q:any)=>{q.options[0].description='Transaction: lookup + update, do not commit. Then receipt send.';},
|
||||
(q:any)=>{q.options[0].description+=' This remedy is "withdrawn".';},
|
||||
(q:any)=>{q.options[2].label='C) Save the report';},
|
||||
(q:any)=>{q.options[2].description='Source excerpt: '+q.options[2].description;},
|
||||
(q:any)=>{q.options[2].description='Lookup and update with a defined commit point.';},
|
||||
(q:any)=>{q.options[2].description+=' This option is cancelled.';},
|
||||
]) expect(accepts(changed(edit))).toBe(false);
|
||||
});
|
||||
test('regression paths belong only to the dense CEO owner', () => {
|
||||
for (const path of ['test/ceo-sequence-aq.test.ts','test/fixtures/ceo-sequence-aq.json'])
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(path)).map(([owner])=>owner)).toEqual(['plan-ceo-finding-count']);
|
||||
const owner=E2E_TOUCHFILES['plan-ceo-finding-count']!;
|
||||
for(let i=0;i<owner.length;i++) {
|
||||
expect(Object.hasOwn(owner,i)).toBe(true);
|
||||
expect(typeof owner[i]).toBe('string');
|
||||
}
|
||||
});
|
||||
|
||||
test('source options, conditional metadata and later contract contradictions cannot own sequencing evidence', () => {
|
||||
for(const edit of [
|
||||
(q:any)=>{q.options[2].description='> '+q.options[2].description;},
|
||||
(q:any)=>{q.options[2].description='~~~\n'+q.options[2].description+'\n~~~';},
|
||||
(q:any)=>{q.question=q.question.replace('Project/branch/task: main','Project/branch/task: Assuming approval, main');},
|
||||
(q:any)=>{q.question=q.question.replace('Project/branch/task: main','Project/branch/task: Provided approval, main');},
|
||||
(q:any)=>{q.question+='\nThis decision is "superseded".';},
|
||||
(q:any)=>{q.question+='\nThis sequence is "cancelled".';},
|
||||
(q:any)=>{q.question+='\nThis sequence is not current.';},
|
||||
(q:any)=>{q.options[0].description+='\nCorrection: the receipt is sent before the payment commit.';},
|
||||
(q:any)=>{q.options[2].description+='\nCorrection: this transaction boundary is now defined.';},
|
||||
]) expect(accepts(changed(edit))).toBe(false);
|
||||
expect(accepts(changed(q=>{q.question+='\n"Earlier review assessment: This sequence is cancelled."';}))).toBe(true);
|
||||
expect(accepts(changed(q=>{q.question+='\nThe archive sequence is cancelled.';}))).toBe(true);
|
||||
});
|
||||
@@ -1,123 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { createHash } from 'node:crypto';
|
||||
import fixture from './fixtures/ceo-source-attribution-6aef.json';
|
||||
import { ceoPaymentFinding, createCeoPaymentFindingCounter } from './helpers/ceo-payment-findings';
|
||||
import type { AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
|
||||
const savedPlan = fixture.savedPlanSegments.map(segment => segment.text).join('\n');
|
||||
const sourceLine = savedPlan.split('\n').find(line => line.startsWith('Source under review:'))!;
|
||||
const declaration = (value: string) => savedPlan.replace(sourceLine, value);
|
||||
const fingerprint = (): AskUserQuestionFingerprint => structuredClone(fixture.fingerprint);
|
||||
const count = (plan = savedPlan, fp = fingerprint()) => {
|
||||
const counter = createCeoPaymentFindingCounter(fixture.seed, () => plan, () => false);
|
||||
const result = counter.isReviewAUQ(fp);
|
||||
return { result, counter };
|
||||
};
|
||||
|
||||
test('captured source, full ledger and literal native packet retain the actual R2 ownership', () => {
|
||||
for (const segment of fixture.savedPlanSegments) {
|
||||
expect(createHash('sha256').update(segment.text).digest('hex')).toBe(segment.sha256);
|
||||
}
|
||||
const call = fixture.fingerprint.nativeCall;
|
||||
expect(fixture.fingerprint.signature).toBe(`${call.sessionId}:${call.toolUseId}`);
|
||||
expect(call.answered).toBe(true);
|
||||
expect(call.failed).toBe(false);
|
||||
const question = call.questions[0]!;
|
||||
expect(savedPlan).toContain(`Question: ${question.question}\nHeader: ${question.header}`);
|
||||
for (const option of question.options) expect(savedPlan).toContain(`${option.label}\n${option.description}`);
|
||||
expect(savedPlan).toContain('| raw SQL fragment | pending |');
|
||||
expect(ceoPaymentFinding(fixture.fingerprint, fixture.seed, savedPlan)).toBeNull();
|
||||
const { result, counter } = count();
|
||||
expect(result).toBe(true);
|
||||
expect(counter.trace).toMatchObject([{ kind: 'recorded-decision', ledgerId: 'R2' }]);
|
||||
expect(() => counter.isReviewAUQ(fingerprint(), [call])).toThrow(/duplicated/);
|
||||
});
|
||||
|
||||
// These labels all assert one current source. They must share both acceptance
|
||||
// and foreign/ambiguous source rules; a label-specific exception is insufficient.
|
||||
const labels = ['Source', 'Source plan', 'Source under review', 'Source plan under review',
|
||||
'Source document', 'Source file under review', 'Plan under review', 'Document under review',
|
||||
'File under review', 'Reviewed plan', 'Review target plan', 'Input plan'];
|
||||
for (const label of labels) {
|
||||
test(`current source declaration accepts ${label}`, () => {
|
||||
expect(count(declaration(`${label}: \`PLAN.md\` (repo root, commit e4bae55).`)).result).toBe(true);
|
||||
});
|
||||
for (const [name, value] of Object.entries({
|
||||
foreign: `${label}: OTHER.md.`,
|
||||
duplicate: `${label}: PLAN.md.\n\nSource under review: PLAN.md.`,
|
||||
conflict: `Source plan: PLAN.md.\n\n${label}: OTHER.md.`,
|
||||
'quoted conflicting field': `Source plan: PLAN.md.\n\n${label}: "OTHER.md".`,
|
||||
'missing conflicting field': `Source plan: PLAN.md.\n\n${label}:`,
|
||||
'negated conflicting field': `Source plan: PLAN.md.\n\n${label}: not PLAN.md.`,
|
||||
quoted: `> ${label}: PLAN.md.\n`,
|
||||
literal: `"${label}: PLAN.md."`,
|
||||
code: `\`\`\`md\n${label}: PLAN.md.\n\`\`\``,
|
||||
historical: `## History\n\n${label}: PLAN.md.\n\n## Current review`,
|
||||
withdrawn: `## Withdrawn attribution\n\n${label}: PLAN.md.\n\n## Current review`,
|
||||
conditional: `${label}: PLAN.md if the user approves it.`,
|
||||
inactive: `${label}: PLAN.md, but this source is no longer current.`,
|
||||
})) test(`${label} rejects ${name} attribution`, () => {
|
||||
expect(() => count(declaration(value))).toThrow(/cannot exclude/);
|
||||
});
|
||||
}
|
||||
|
||||
for (const [name, value] of Object.entries({
|
||||
'paragraph metadata after a sentence': 'Working plan for the current CEO review. Source under review: PLAN.md (repo root).',
|
||||
'multiple metadata lines': 'Working plan for the current CEO review.\nSource under review: PLAN.md (repo root).\nMode: HOLD SCOPE.',
|
||||
'inline source formatting': '**Source under review:** `PLAN.md` (repo root).',
|
||||
'copied source metadata': 'Source under review: PLAN.md (copied into CLAUDE.md as the session request).',
|
||||
'byte-identical source copy metadata': 'Source under review: PLAN.md (byte-identical to the plan embedded in CLAUDE.md).',
|
||||
'prior source in separate inactive scope': '## History\n\nSource under review: OTHER.md.\n\n## Current source\n\nSource under review: PLAN.md.',
|
||||
'nested inactive scope closes': '## Metadata\n\n### Archived source\n\nSource plan: OTHER.md.\n\n### Current source\n\nSource under review: PLAN.md.',
|
||||
})) test(`current attribution supports ${name}`, () => expect(count(declaration(value)).result).toBe(true));
|
||||
|
||||
for (const value of [
|
||||
'PLAN.md or OTHER.md', 'PLAN.md and OTHER.md', 'PLAN.md versus OTHER.md',
|
||||
'PLAN.md / OTHER.md', 'PLAN.md, OTHER.md', 'PLAN.md; OTHER.md',
|
||||
'PLAN.md (repo root) or OTHER.md', 'PLAN.md rather than OTHER.md',
|
||||
'PLAN.md instead of OTHER.md', 'PLAN.md or PLAN.md',
|
||||
'PLAN.md & OTHER.md', 'PLAN.md + OTHER.md', 'PLAN.md vs. OTHER.md',
|
||||
'PLAN.md (repo root; or OTHER.md)', 'PLAN.md at repo root & OTHER.md',
|
||||
'PLAN.md (copied into CLAUDE.md or OTHER.md)',
|
||||
]) test(`a compound current source is not reduced to its first filename: ${value}`, () => {
|
||||
expect(() => count(declaration(`Source under review: ${value}.`))).toThrow(/cannot exclude/);
|
||||
});
|
||||
|
||||
for (const [name, value] of Object.entries({
|
||||
absent: '',
|
||||
'unrelated filename': 'The review happens to mention PLAN.md.',
|
||||
'quoted source filename': 'Source under review: "PLAN.md".',
|
||||
'conditional prefix': 'If approved, Source under review: PLAN.md.',
|
||||
'historical paragraph prefix': 'Historical metadata. Source under review: PLAN.md.',
|
||||
'history paragraph prefix': 'History: earlier review. Source under review: PLAN.md.',
|
||||
'negative prefix': 'Not the Source under review: PLAN.md.',
|
||||
'negated source': 'Source under review: not PLAN.md.',
|
||||
'conditional suffix': 'Source under review: PLAN.md would be used after approval.',
|
||||
'current source withdrawn later in paragraph': 'Source under review: PLAN.md. This source is withdrawn.',
|
||||
'foreign declaration later in paragraph': 'Source plan: PLAN.md. Source under review: OTHER.md.',
|
||||
'duplicate declaration later in paragraph': 'Source under review: PLAN.md. Input plan: PLAN.md.',
|
||||
})) test(`pending R2 rejects ${name}`, () => expect(() => count(declaration(value))).toThrow(/cannot exclude/));
|
||||
|
||||
for (const [name, mutate] of Object.entries({
|
||||
'foreign row source': (plan: string) => plan.replace('from `request.params.userId` (PLAN.md:16-31, 110-112)', 'from `request.params.userId` (OTHER.md:16-31, 110-112)'),
|
||||
'missing row source': (plan: string) => plan.replace('from `request.params.userId` (PLAN.md:16-31, 110-112)', 'from `request.params.userId` (no evidence)'),
|
||||
'withdrawn row': (plan: string) => plan.replace('| R2 (backend owner)', '| R2 (withdrawn backend owner)'),
|
||||
'compound row status': (plan: string) => plan.replace('| raw SQL fragment | pending |', '| raw SQL fragment | pending / approved |'),
|
||||
'quoted row status': (plan: string) => plan.replace('| raw SQL fragment | pending |', '| raw SQL fragment | "pending" |'),
|
||||
'historical currentDecision': (plan: string) => plan.replace('## currentDecision (R2)', '## Historical currentDecision (R2)'),
|
||||
'different saved header': (plan: string) => plan.replace('Header: Lookup query', 'Header: Foreign lookup'),
|
||||
'different saved question': (plan: string) => plan.replace('Question: D2 — R2:', 'Question: D2 — R3:'),
|
||||
'missing full saved option': (plan: string) => plan.replace(fixture.fingerprint.nativeCall.questions[0]!.options[1]!.description, 'Summary only.'),
|
||||
})) test(`source attribution does not weaken ${name}`, () => expect(() => count(mutate(savedPlan))).toThrow(/cannot exclude/));
|
||||
|
||||
for (const [name, mutate] of Object.entries({
|
||||
signature: (fp: ReturnType<typeof fingerprint>) => { fp.signature = 'foreign'; },
|
||||
unanswered: (fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.answered = false; },
|
||||
failed: (fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.failed = true; },
|
||||
'missing answer': (fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.answers = {}; },
|
||||
'pending answer': (fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.unansweredQuestionIndices = [0]; },
|
||||
'foreign answer': (fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.answers = { foreign: 'A' }; },
|
||||
})) test(`source attribution retains native ${name} ownership rejection`, () => {
|
||||
const fp = fingerprint(); mutate(fp);
|
||||
expect(() => count(savedPlan, fp)).toThrow();
|
||||
});
|
||||
@@ -1,83 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, planCountQuestionPhase, type AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
import fixture from './fixtures/ceo-test-subject-ao.json';
|
||||
const calls=fixture.fingerprints as AskUserQuestionFingerprint[];
|
||||
const actual=calls[1]!;
|
||||
function change(edit:(q:any, call:any, fp:any)=>void) {
|
||||
const fp=structuredClone(actual),call=fp.nativeCall!,q=call.questions[0]!;
|
||||
const selected=q.options.findIndex(o=>o.label===call.answers?.[q.question]);
|
||||
edit(q,call,fp);
|
||||
call.answers={[q.question]:q.options[selected]?.label??''};
|
||||
fp.options=q.options.map((o,i)=>({index:i+1,label:o.label}));
|
||||
return fp;
|
||||
}
|
||||
test('the completed affected-Test question starts review from its current ELI10 assertion gap',()=>{
|
||||
expect(calls.map(ceoFirstReviewAUQ)).toEqual([false,true,false]);
|
||||
let review=false;
|
||||
expect(calls.map(fp=>{const phase=planCountQuestionPhase(fp,review,ceoStep0Boundary,ceoFirstReviewAUQ);review=phase.reviewStarted;return phase.preReview;})).toEqual([true,false,false]);
|
||||
});
|
||||
test('structural Test identity permits ordinary question and separator variations',()=>{
|
||||
for(const title of [
|
||||
'D4 — Test 1: choose its assertion?',
|
||||
'D4 — Test 1 — which assertion belongs here?',
|
||||
'D4 - Test 1 (successful charge): what must this test verify?',
|
||||
'd4 — Test 1 (successful charge): assertion choice?',
|
||||
'D4 — Test 1 what should the expected result be?',
|
||||
]) expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace(/^[^\n]+/,title);}))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{
|
||||
q.question=q.question.replace(/^D4/,'D17').replace(/\b4([A-C])\b/g,'17$1');
|
||||
q.options=q.options.map((o:any)=>({...o,label:o.label.replace(/^4/,'17')}));
|
||||
}))).toBe(true);
|
||||
});
|
||||
test('test headers, competing finding IDs and uniform foreign decision choices cannot borrow the assessment',()=>{
|
||||
for(const header of ['Test 2','Finding 1','Issue 1','Approach']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 (Finding 2)');}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{
|
||||
q.question=q.question.replace(/\b4([A-C])\b/g,'8$1');q.options=q.options.map((o:any)=>({...o,label:o.label.replace(/^4/,'8')}));
|
||||
}))).toBe(false);
|
||||
});
|
||||
test('explicit Test and decision identifiers must be anchored integers with one test owner',()=>{
|
||||
for(const header of ['Test 0','Test 01','Test 1.2']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(false);
|
||||
for(const decision of ['D0','D04']) expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace(/^D4/,decision);}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 (Test 2)');}))).toBe(false);
|
||||
for(const header of ['Receipt assertion','Test contract','Test 1: receipt assertion']) expect(ceoFirstReviewAUQ(change(q=>{q.header=header;}))).toBe(true);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('Test 1 (successful charge)','Test 1 ("Test 2" is an archive label)');}))).toBe(true);
|
||||
});
|
||||
test('the owned weak assertion must remain current and outside quoted or conditional source frames',()=>{
|
||||
for(const intro of ['Source excerpt: ','Earlier review assessment: ','If approved later, ']) expect(ceoFirstReviewAUQ(change(q=>{
|
||||
q.question=q.question.replace('The planned test only checks',intro+'The planned test only checks');
|
||||
}))).toBe(false);
|
||||
for(const intro of ['Source excerpt follows. ','Earlier review assessment follows. ','If approved later. ']) expect(ceoFirstReviewAUQ(change(q=>{
|
||||
q.question=q.question.replace('The planned test only checks',intro+'The planned test only checks');
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('The planned test only checks','The planned test no longer only checks');}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('The planned test only checks that the receipt is truthy.','"The planned test only checks that the receipt is truthy."');}))).toBe(false);
|
||||
});
|
||||
test('direct current withdrawals stay effective while a quoted historical note stays harmless',()=>{
|
||||
for(const text of ['This finding is withdrawn.','This assessment is "closed".','Correction: this explanation is not current.']) expect(ceoFirstReviewAUQ(change(q=>{q.question+='\n'+text;}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.question=q.question.replace('\nELI10:','\nArchive note: "Source: this finding is withdrawn."\nELI10:');}))).toBe(true);
|
||||
});
|
||||
test('the complete current amendment belongs to an offered option',()=>{
|
||||
for(const prefix of ['Source excerpt: ','Historical example: ','If approved later: ']) expect(ceoFirstReviewAUQ(change(q=>{
|
||||
for(const o of q.options)o.description=prefix+o.description;
|
||||
}))).toBe(false);
|
||||
for(const text of [' This amendment is withdrawn.',' This amendment is "closed".',' This remedy is a historical example, not the current option.']) expect(ceoFirstReviewAUQ(change(q=>{
|
||||
for(const o of q.options)o.description+=text;
|
||||
}))).toBe(false);
|
||||
expect(ceoFirstReviewAUQ(change(q=>{q.options[0].label='4A: Keep the truthy assertion (recommended)';}))).toBe(false);
|
||||
});
|
||||
test('native completion, exact answer, index and menu identity remain required',()=>{
|
||||
for(const edit of [
|
||||
(_q:any,c:any)=>{c.answered=false;},(_q:any,c:any)=>{c.failed=true;},
|
||||
(_q:any,c:any)=>{delete c.answeredAt;},(_q:any,c:any)=>{c.unansweredQuestionIndices=[0];},
|
||||
(_q:any,_c:any,fp:any)=>{fp.signature='foreign:tool';},
|
||||
(_q:any,_c:any,fp:any)=>{fp.nativeQuestionIndex=1;},(q:any)=>{q.multiSelect=true;},
|
||||
]) expect(ceoFirstReviewAUQ(change(edit))).toBe(false);
|
||||
const wrongAnswer=change(()=>{});wrongAnswer.nativeCall!.answers={};expect(ceoFirstReviewAUQ(wrongAnswer)).toBe(false);
|
||||
const wrongMenu=change(()=>{});wrongMenu.options[0]!.label='Foreign menu';expect(ceoFirstReviewAUQ(wrongMenu)).toBe(false);
|
||||
});
|
||||
test('new inputs belong only to the dense CEO finding owner',()=>{
|
||||
for(const file of ['test/ceo-test-subject-ao.test.ts','test/fixtures/ceo-test-subject-ao.json']) expect(Object.entries(E2E_TOUCHFILES).filter(([,paths])=>paths.includes(file)).map(([owner])=>owner)).toEqual(['plan-ceo-finding-count']);
|
||||
for(const paths of Object.values(E2E_TOUCHFILES))for(let i=0;i<paths.length;i++)expect(Object.hasOwn(paths,i)&&typeof paths[i]==='string').toBe(true);
|
||||
});
|
||||
@@ -1,109 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { ceoFirstReviewAUQ, ceoStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import fixture from './fixtures/ceo-transaction-contract-ar.json';
|
||||
|
||||
const calls = () => structuredClone(fixture.calls) as NativePlanQuestionCall[];
|
||||
const first = () => calls()[2]!;
|
||||
const fp = (call = first()) => nativePlanCallFingerprint(call, 0, true);
|
||||
const classify = (call = first()) => ceoFirstReviewAUQ(fp(call));
|
||||
const mutate = (fn: (call: NativePlanQuestionCall) => void) => { const call = first(); fn(call); return call; };
|
||||
const prose = (fn: (text: string) => string) => mutate(call => {
|
||||
const q = call.questions[0]!, answer = call.answers![q.question]!;
|
||||
q.question = fn(q.question); call.answers = { [q.question]: answer };
|
||||
});
|
||||
const option = (at: number, fn: (o: NativePlanQuestionCall['questions'][number]['options'][number]) => void) => mutate(call => {
|
||||
const q = call.questions[0]!, selected = q.options.findIndex(o => o.label === call.answers![q.question]);
|
||||
fn(q.options[at]!); call.answers = { [q.question]: q.options[selected]!.label };
|
||||
});
|
||||
|
||||
describe('AR current transaction decision', () => {
|
||||
test('the exact transaction decision starts review after setup', () => {
|
||||
let started = false;
|
||||
const phases = calls().map(call => {
|
||||
const phase = planCountQuestionPhase(fp(call), started, ceoStep0Boundary, ceoFirstReviewAUQ);
|
||||
started = phase.reviewStarted; return phase.preReview;
|
||||
});
|
||||
expect(phases).toEqual([true, true, false, false, false, false, false, false]);
|
||||
expect(calls().map(classify)).toEqual([false, false, true, false, false, false, false, false]);
|
||||
});
|
||||
test('title wording and ordinal punctuation do not supply semantics', () => {
|
||||
expect(classify(prose(s => s.replace('Where does the user update commit relative to the email call?', 'When should the update commit before the email call?')))).toBe(true);
|
||||
expect(classify(mutate(c => { c.questions[0]!.header = 'Transaction boundary'; }))).toBe(true);
|
||||
expect(classify(mutate(c => {
|
||||
const q = c.questions[0]!, answer = c.answers![q.question]!;
|
||||
q.options.forEach(o => { o.label = o.label.replace(/^3([A-Z]) /, '3$1) '); });
|
||||
c.answers = { [q.question]: answer.replace(/^3([A-Z]) /, '3$1) ') };
|
||||
}))).toBe(true);
|
||||
expect(classify(prose(s => s + '\nArchived note: "This finding is withdrawn."'))).toBe(true);
|
||||
expect(classify(mutate(c => { const q = c.questions[0]!; c.answers = { [q.question]: q.options[1]!.label }; }))).toBe(true);
|
||||
});
|
||||
test('a complete owned successful answer is required', () => {
|
||||
for (const change of [
|
||||
(c: NativePlanQuestionCall) => { c.sessionId = ''; }, (c: NativePlanQuestionCall) => { c.toolUseId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; }, (c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; }, (c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.answeredAt = 'invalid'; }, (c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; }, (c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
]) expect(classify(mutate(change))).toBe(false);
|
||||
for (const fingerprint of [{ ...fp(), signature: 'foreign' }, { ...fp(), nativeQuestionIndex: 1 }, { ...fp(), options: fp().options.toReversed() }])
|
||||
expect(ceoFirstReviewAUQ(fingerprint)).toBe(false);
|
||||
});
|
||||
test('decision, header, recommendation and offered ordinals agree', () => {
|
||||
for (const change of [(s: string) => s.replace('D3 —', 'D0 —'), (s: string) => s.replace('D3 —', 'D03 —'), (s: string) => s.replace('Recommendation: 3A', 'Recommendation: 4A')])
|
||||
expect(classify(prose(change))).toBe(false);
|
||||
for (const header of ['Routing', 'Approach', 'D4 Txn boundary', 'D03 Txn boundary', 'Source Txn boundary'])
|
||||
expect(classify(mutate(c => { c.questions[0]!.header = header; }))).toBe(false);
|
||||
for (const label of ['03A Commit update, then email', '4A Commit update, then email', '3A) 4A Commit update, then email'])
|
||||
expect(classify(option(0, o => { o.label = label; }))).toBe(false);
|
||||
});
|
||||
test('a unique current context and assessment are required', () => {
|
||||
for (const field of ['Project/branch/task: ', 'ELI10: ']) for (const prefix of ['Source excerpt: ', 'Earlier review assessment: ', 'If approved, ', 'Assuming approval, '])
|
||||
expect(classify(prose(s => s.replace(field, field + prefix)))).toBe(false);
|
||||
for (const prefix of ['Source excerpt:\n', 'Project/branch/task: duplicate\n', 'ELI10: duplicate\n'])
|
||||
expect(classify(prose(s => s.replace('ELI10:', prefix + 'ELI10:')))).toBe(false);
|
||||
expect(classify(prose(s => s.replace(/^Project\/branch\/task:.*\n/m, '')))).toBe(false);
|
||||
expect(classify(prose(s => '```\n' + s + '\n```'))).toBe(false);
|
||||
});
|
||||
test('the missing boundary must remain current and unresolved', () => {
|
||||
expect(classify(prose(s => s.replace('but never says whether', 'and explicitly specifies whether')))).toBe(false);
|
||||
expect(classify(prose(s => s.replace('The plan says', 'Earlier review assessment follows. The plan says')))).toBe(false);
|
||||
for (const tail of ['This finding is withdrawn.', 'This transaction boundary is now specified.', 'Correction: this transaction boundary is "resolved".'])
|
||||
expect(classify(prose(s => s + '\n' + tail))).toBe(false);
|
||||
});
|
||||
test('one current amendment owns order and rollback safety', () => {
|
||||
for (const replacement of ['before commit, inside any DB transaction', 'after commit, inside the DB transaction'])
|
||||
expect(classify(option(0, o => { o.description = o.description!.replace('after commit, outside any DB transaction', replacement); }))).toBe(false);
|
||||
expect(classify(option(0, o => { o.description = o.description!.replace('can never roll back paid status', 'can roll back paid status'); }))).toBe(false);
|
||||
expect(classify(option(0, o => { o.description = o.description!.replace('Lookup and update commit in one transaction;', 'No transactional update is planned;'); }))).toBe(false);
|
||||
expect(classify(option(0, o => { o.label = '3A Write the final report'; }))).toBe(false);
|
||||
for (const prefix of ['Source excerpt: ', 'If approved, '])
|
||||
expect(classify(option(0, o => { o.description = prefix + o.description; }))).toBe(false);
|
||||
for (const tail of ['This amendment is "withdrawn".', 'Correction: do not commit the update before email.'])
|
||||
expect(classify(option(0, o => { o.description += ' ' + tail; }))).toBe(false);
|
||||
});
|
||||
test('new transaction syntax rejects stale and conditional evidence', () => {
|
||||
for (const prefix of ['Assuming approval, ', 'Provided approval, ']) {
|
||||
expect(classify(prose(s => s.replace('ELI10: ', 'ELI10: ' + prefix)))).toBe(false);
|
||||
expect(classify(option(0, o => { o.description = prefix + o.description; }))).toBe(false);
|
||||
expect(classify(option(1, o => { o.description = prefix + o.description; }))).toBe(false);
|
||||
}
|
||||
for (const status of ['superseded', '"superseded"', '"resolved"', 'no longer current', '"no longer current"']) {
|
||||
expect(classify(prose(s => s + '\nThis finding is ' + status + '.'))).toBe(false);
|
||||
expect(classify(option(0, o => { o.description += ' This amendment is ' + status + '.'; }))).toBe(false);
|
||||
expect(classify(option(1, o => { o.description += ' This option is ' + status + '.'; }))).toBe(false);
|
||||
}
|
||||
for (const history of ['> This finding is superseded.', 'Archived note: "This finding is superseded."', 'Archived note: "This finding is no longer current."', '~~~\nThis finding is superseded.\n~~~'])
|
||||
expect(classify(prose(s => s + '\n' + history))).toBe(true);
|
||||
for (const convert of [(s: string) => '> ' + s, (s: string) => '"' + s + '"', (s: string) => '`' + s + '`']) {
|
||||
expect(classify(option(0, o => { o.description = convert(o.description!); }))).toBe(false);
|
||||
expect(classify(option(1, o => { o.description = convert(o.description!); }))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('the opposed option owns the unchanged risk', () => {
|
||||
expect(classify(option(1, o => { o.description = 'The transaction shape is safe and fully specified.'; }))).toBe(false);
|
||||
expect(classify(option(1, o => { o.description = 'Source excerpt: ' + o.description; }))).toBe(false);
|
||||
expect(classify(option(1, o => { o.description += ' This option is withdrawn.'; }))).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -189,18 +189,14 @@ describe('dependency-free CI planner and report execution', () => {
|
||||
for (const tier of ['gate', 'periodic'] as const) {
|
||||
test(`${tier}: host planner preserves the complete manifest and report fails closed`, () => {
|
||||
const sliceCount = tier === 'gate' ? 6 : 7;
|
||||
const dedicatedAutoplanSlice = tier === 'periodic';
|
||||
const reportDir = path.join(fixture, tier);
|
||||
const manifestPath = path.join(reportDir, 'manifest.json');
|
||||
const planned = run([
|
||||
'--emit-plan', manifestPath, '--slices', String(sliceCount),
|
||||
...(dedicatedAutoplanSlice ? ['--autoplan-slice'] : []),
|
||||
], tier);
|
||||
const planned = run(['--emit-plan', manifestPath, '--slices', String(sliceCount)], tier);
|
||||
expect(planned.error).toBeUndefined();
|
||||
expect(planned.status, planned.stderr).toBe(0);
|
||||
const manifest: PaidRunManifest = JSON.parse(fs.readFileSync(manifestPath, 'utf8'));
|
||||
expect(manifest).toEqual(buildRunManifest({
|
||||
tier, sliceCount, dedicatedAutoplanSlice, evalsAll: true, env: { EVALS_ALL: '1' },
|
||||
tier, sliceCount, evalsAll: true, env: { EVALS_ALL: '1' },
|
||||
}));
|
||||
expect(manifest.entries.filter(entry => entry.status === 'planned').length).toBeGreaterThan(0);
|
||||
expect(fs.existsSync(path.join(fixture, 'node_modules'))).toBe(false);
|
||||
|
||||
@@ -1,114 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import calls from './fixtures/design-artifacts-w-calls.json';
|
||||
import { isDesignArtifactGeneration } from './helpers/design-artifact-question';
|
||||
import { nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
const fp = (call: any) => nativePlanCallFingerprint(call as NativePlanQuestionCall, 0, false);
|
||||
const artifacts = [calls[2]!, calls[3]!];
|
||||
|
||||
test('all eight W calls remain visible: five seeded findings, one shell decision and two artifact approvals', () => {
|
||||
const phases = calls.map(call => planCountQuestionPhase(fp(call), true, () => false,
|
||||
undefined, undefined, undefined, isDesignArtifactGeneration));
|
||||
expect(phases.map(p => p.administrative ?? 'review')).toEqual([
|
||||
'review', 'review', 'artifact-generation', 'artifact-generation', 'review', 'review', 'review', 'review',
|
||||
]);
|
||||
for (const call of artifacts) {
|
||||
expect(planCountQuestionPhase(fp(call), false, () => false, () => true,
|
||||
undefined, undefined, isDesignArtifactGeneration)).toEqual({
|
||||
preReview: false, reviewStarted: false, administrative: 'artifact-generation',
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
test('new decisions, missing coverage, altered artifacts, deferrals and quoted examples remain findings', () => {
|
||||
for (const original of artifacts) {
|
||||
for (const mutate of [
|
||||
(c: any) => { c.questions[0].options[0].description += ' Also change the Save behavior.'; },
|
||||
(c: any) => { c.questions[0].options[0].description += ' Drop the error state.'; },
|
||||
(c: any) => { c.questions[0].options[0].description = c.questions[0].options[0].description.replace('No new design decisions', 'Choose new design decisions'); },
|
||||
(c: any) => { c.questions[0].options[1].description += ' The failure contract is still missing.'; },
|
||||
(c: any) => { c.questions[0].options[0].preview = 'Change the save contract'; },
|
||||
(c: any) => { c.answers[c.questions[0].question] = c.questions[0].options[1].label; },
|
||||
(c: any) => { const q = c.questions[0]; const a = c.answers[q.question]; q.question = 'Example: ' + q.question; c.answers = { [q.question]: a }; },
|
||||
(c: any) => { c.questions.push(calls[7]!.questions[0]); },
|
||||
(c: any) => { c.answered = false; }, (c: any) => { c.failed = true; },
|
||||
(c: any) => { delete c.failed; }, (c: any) => { delete c.unansweredQuestionIndices; },
|
||||
(c: any) => { c.unansweredQuestionIndices = [0]; }, (c: any) => { c.answeredAt = 'invalid'; },
|
||||
(c: any) => { c.sessionId = ''; },
|
||||
]) {
|
||||
const call = structuredClone(original); mutate(call);
|
||||
expect(isDesignArtifactGeneration(fp(call))).toBe(false);
|
||||
}
|
||||
const reordered = structuredClone(original); reordered.questions[0]!.options.reverse();
|
||||
expect(isDesignArtifactGeneration(fp(reordered))).toBe(true);
|
||||
expect(isDesignArtifactGeneration({ ...fp(original), signature: 'foreign' })).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
const REPORT = '# Reviewed plan\n\n## GSTACK REVIEW REPORT\n\n| Review | Status | Findings |\n|---|---|---|\n| Design | clean | recorded |\n\nVERDICT: Review complete\n\nNO UNRESOLVED DECISIONS\n';
|
||||
const GATE = 'Exit plan mode?\n\nClaude wants to exit plan mode\n❯ 1. Yes, and switch to default (ask each time) for this session\n 2. No\n';
|
||||
|
||||
test.skipIf(process.platform === 'win32').each(['only-artifacts', 'freshness'] as const)('real fake-PTY artifact %s preserves coverage and fresh-report requirements', async mode => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-artifact-free-'));
|
||||
const fake = path.join(dir, 'fake-claude'), worker = path.join(dir, 'worker.ts');
|
||||
const report = path.join(dir, 'report.md'), output = path.join(dir, 'result.json');
|
||||
const pidFile = path.join(dir, 'pid.json'), inputs = path.join(dir, 'inputs.jsonl');
|
||||
const refreshed = path.join(dir, 'refreshed');
|
||||
const selected = mode === 'only-artifacts' ? artifacts : [calls[0]!, artifacts[0]!];
|
||||
fs.writeFileSync(fake, `#!${process.execPath}\n` + String.raw`
|
||||
import * as fs from 'node:fs'; import * as path from 'node:path';
|
||||
const stat = process.platform === 'linux' ? fs.readFileSync('/proc/self/stat','utf8') : null;
|
||||
fs.writeFileSync(process.env.PID_FILE, JSON.stringify({pid:process.pid,start:stat?.slice(stat.lastIndexOf(')')+2).split(' ')[19]}));
|
||||
let sent=false; process.stdin.setRawMode?.(true);
|
||||
process.stdin.on('data', data => {
|
||||
fs.appendFileSync(process.env.INPUT_FILE,JSON.stringify(data.toString())+'\n');
|
||||
if(sent)return;sent=true;
|
||||
const at=Date.now(),sid='artifact-free';
|
||||
const project=path.join(process.env.CLAUDE_CONFIG_DIR,'projects','owned');fs.mkdirSync(project,{recursive:true});
|
||||
const events=JSON.parse(process.env.CALLS).flatMap((call,i)=>[
|
||||
{cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date(at-100+i*10).toISOString(),message:{role:'assistant',content:[{type:'tool_use',id:call.toolUseId,name:'AskUserQuestion',input:{questions:call.questions}}]}},
|
||||
{cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date(at-99+i*10).toISOString(),toolUseResult:{answers:call.answers},message:{role:'user',content:[{type:'tool_result',tool_use_id:call.toolUseId,content:'Your questions have been answered: '+Object.entries(call.answers).map(([q,a])=>JSON.stringify(q)+'='+JSON.stringify(a)).join(', ')+'. You can now continue with these answers in mind.'}]}}
|
||||
]);
|
||||
events.push({cwd:process.cwd(),sessionId:sid,isSidechain:false,timestamp:new Date(at).toISOString(),message:{role:'assistant',content:[{type:'text',text:'Design review complete.'},{type:'tool_use',id:'exit',name:'ExitPlanMode',input:{}}]}});
|
||||
fs.writeFileSync(path.join(project,sid+'.jsonl'),events.map(e=>JSON.stringify(e)+'\n').join(''));
|
||||
fs.writeFileSync(process.env.REPORT_FILE,process.env.REPORT);
|
||||
if(process.env.MODE==='freshness') {
|
||||
fs.utimesSync(process.env.REPORT_FILE,(at-95)/1000,(at-95)/1000);
|
||||
setTimeout(()=>{fs.writeFileSync(process.env.REPORT_FILE,process.env.REPORT);fs.writeFileSync(process.env.REFRESHED,'yes');},5500);
|
||||
}
|
||||
process.stdout.write(process.env.GATE);
|
||||
});process.stdin.resume();
|
||||
`); fs.chmodSync(fake,0o755);
|
||||
const runner = pathToFileURL(path.join(import.meta.dir,'helpers/claude-pty-runner.ts')).href;
|
||||
const artifactHelper = pathToFileURL(path.join(import.meta.dir,'helpers/design-artifact-question.ts')).href;
|
||||
const env = {PID_FILE:pidFile,INPUT_FILE:inputs,REPORT_FILE:report,REPORT,GATE,CALLS:JSON.stringify(selected),MODE:mode,REFRESHED:refreshed};
|
||||
fs.writeFileSync(worker, `import {runPlanSkillCounting} from ${JSON.stringify(runner)};\nimport {isDesignArtifactGeneration} from ${JSON.stringify(artifactHelper)};\nconst result=await runPlanSkillCounting({skillName:'plan-design-review',slashCommand:'/plan-design-review',followUpPrompt:'# Artifact control',expectedPlanPath:${JSON.stringify(report)},isLastStep0AUQ:()=>false,isFirstReviewAUQ:()=>true,isArtifactGenerationAUQ:isDesignArtifactGeneration,reviewCountCeiling:8,timeoutMs:33000,env:${JSON.stringify(env)}});await Bun.write(${JSON.stringify(output)},JSON.stringify(result));\n`);
|
||||
const child = Bun.spawn([process.execPath,worker],{env:{...process.env,EVALS_HERMETIC:'1',EVALS_RUN_ID:'',BROWSE_TERMINAL_BINARY:fake},stdout:'pipe',stderr:'pipe'});
|
||||
const timer=setTimeout(()=>child.kill('SIGKILL'),35000);
|
||||
try {
|
||||
const [code,out,err]=await Promise.all([child.exited,new Response(child.stdout).text(),new Response(child.stderr).text()]);
|
||||
expect(code,out+err).toBe(0);const result=JSON.parse(fs.readFileSync(output,'utf8'));
|
||||
expect(result.transcript.calls).toHaveLength(2);
|
||||
expect(result.administrativeCount).toBe(mode==='only-artifacts'?2:1);
|
||||
expect(result.reviewCount).toBe(mode==='only-artifacts'?0:1);
|
||||
expect(result.outcome).toBe(mode==='only-artifacts'?'no_review_questions':'plan_ready');
|
||||
if(mode==='freshness') expect(fs.existsSync(refreshed)).toBe(true);
|
||||
expect(fs.readFileSync(inputs,'utf8').trim().split('\n').map(x=>JSON.parse(x))).toEqual(['/plan-design-review\r']);
|
||||
} finally {
|
||||
clearTimeout(timer);child.kill('SIGKILL');
|
||||
if(fs.existsSync(pidFile))try {
|
||||
const p=JSON.parse(fs.readFileSync(pidFile,'utf8'));let owned=false;
|
||||
if(process.platform==='linux') {
|
||||
const s=fs.readFileSync(`/proc/${p.pid}/stat`,'utf8');owned=s.slice(s.lastIndexOf(')')+2).split(' ')[19]===p.start && fs.readFileSync(`/proc/${p.pid}/cmdline`,'utf8').split('\0').includes(fake);
|
||||
} else owned=execFileSync('ps',['-p',String(p.pid),'-o','command='],{encoding:'utf8',timeout:1000}).includes(fake);
|
||||
if(owned)process.kill(p.pid,'SIGKILL');
|
||||
}catch{/* owned fake already gone */}
|
||||
fs.rmSync(dir,{recursive:true,force:true});
|
||||
}
|
||||
},40000);
|
||||
@@ -1,147 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-compact-primary-aw-call.json';
|
||||
import { nativePlanCallFingerprint, designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const fresh = () => structuredClone(captured.call) as NativePlanQuestionCall;
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
const accepted = (call: NativePlanQuestionCall) => isDesignCountFirstReview(fingerprint(call));
|
||||
type Question = NativePlanQuestionCall['questions'][number];
|
||||
function change(edit: (q: Question, c: NativePlanQuestionCall) => void) {
|
||||
const call = fresh(), q = call.questions[0]!;
|
||||
edit(q, call);
|
||||
call.answers = { [q.question]: q.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('compact numbered design decision fields', () => {
|
||||
test('the exact acknowledged primary issue starts review without changing the call', () => {
|
||||
const call = fresh(), before = JSON.stringify(call);
|
||||
expect(accepted(call)).toBe(true);
|
||||
expect(planCountQuestionPhase(fingerprint(call), false, designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup)).toEqual({
|
||||
preReview: false, reviewStarted: true,
|
||||
});
|
||||
expect(JSON.stringify(call)).toBe(before);
|
||||
});
|
||||
|
||||
test('layout, explanatory prose, names, tokens, ordinals and offered answers may vary', () => {
|
||||
expect(accepted(change(q => {
|
||||
q.question = q.question.replaceAll('has no', 'lacks');
|
||||
for (const field of ['Project/branch/task:', 'ELI10:', 'Stakes if we pick wrong:', 'Recommendation:', 'Completeness:', 'Net:']) {
|
||||
q.question = q.question.replace(` ${field}`, `\n${field}`);
|
||||
}
|
||||
}))).toBe(true);
|
||||
expect(accepted(change(q => { q.question = q.question.replace('The user came to do one thing: save. When everything shouts, nothing is heard, and a scanning user can hit Reset by mistake.', 'If a user scans the header, identical styles conceal the intended action.'); }))).toBe(true);
|
||||
expect(accepted(JSON.parse(JSON.stringify(fresh()).replaceAll('Save', 'Publish').replaceAll('Reset', 'Revert').replaceAll('#1d4ed8', '#234abc').replaceAll('white', 'black')))).toBe(true);
|
||||
expect(accepted(change(q => {
|
||||
q.header = 'Issue 9'; q.question = q.question.replace('D2', 'D17').replace('Issue 1', 'Issue 9').replace(/\b1([AB])\b/g, '9$1');
|
||||
q.options.forEach(o => { o.label = o.label.replace(/^1/, '9'); });
|
||||
}))).toBe(true);
|
||||
expect(accepted(change(q => { q.options.reverse(); }))).toBe(true);
|
||||
for (const option of fresh().questions[0]!.options) {
|
||||
const call = fresh(); call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(accepted(call)).toBe(true);
|
||||
}
|
||||
expect(accepted(change(q => {
|
||||
q.question = q.question.replaceAll(', Export', '').replaceAll('four', 'three');
|
||||
q.options.forEach(o => { o.label = o.label.replace('four', 'three'); o.description = o.description?.replace('/Export', ''); });
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('completed native identity and exact offered selection remain mandatory', () => {
|
||||
for (const edit of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.answeredAt = 'invalid'; },
|
||||
(c: NativePlanQuestionCall) => { c.sessionId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.toolUseId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered' }; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
]) { const call = fresh(); edit(call); expect(accepted(call)).toBe(false); }
|
||||
for (const edit of [
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.signature = 'other:call'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.sessionId = 'other'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeQuestionIndex = 1; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.options.reverse(); },
|
||||
]) { const fp = fingerprint(fresh()); edit(fp); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('a numbered setup, mismatched issue or source packet cannot supply a current finding', () => {
|
||||
for (const header of ['Scope', 'Routing', 'Issue 2', 'Outside voices']) expect(accepted(change(q => { q.header = header; }))).toBe(false);
|
||||
for (const field of ['Project/branch/task:', 'ELI10:', 'Stakes if we pick wrong:', 'Recommendation:', 'Completeness:', 'Net:']) {
|
||||
expect(accepted(change(q => { q.question = q.question.replace(field, ''); }))).toBe(false);
|
||||
expect(accepted(change(q => { q.question += ` ${field} Extra.`; }))).toBe(false);
|
||||
}
|
||||
for (const prefix of ['Historical example:\n', 'Source:\n', 'If approved, ', '> ', '```text\n']) {
|
||||
expect(accepted(change(q => { q.question = prefix + q.question + (prefix.startsWith('```') ? '\n```' : ''); }))).toBe(false);
|
||||
}
|
||||
for (const edit of [
|
||||
(q: Question) => { q.question = q.question.replace('Save has no primary-action hierarchy.', 'How should we route the next reviewer?'); },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: The header', 'ELI10: Previously, the header'); },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: The header', 'ELI10: If approved, the header'); },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: The header shows Save, Reset, Cancel, Export as four identical buttons.', 'ELI10: "The header shows Save, Reset, Cancel, Export as four identical buttons."'); },
|
||||
(q: Question) => { q.question = q.question.replace('main, Pass 1', 'Historical example: main, Pass 1'); },
|
||||
(q: Question) => { q.options[0]!.label = q.options[0]!.label.replace('1A', '2A'); },
|
||||
(q: Question) => { q.question = q.question.replace('Recommendation: 1A', 'Recommendation: 2A'); },
|
||||
]) expect(accepted(change(edit))).toBe(false);
|
||||
});
|
||||
|
||||
test('same current actors, style, choice and unresolved opposition must agree', () => {
|
||||
for (const edit of [
|
||||
(q: Question) => { q.question = q.question.replace('as four identical', 'as three identical'); },
|
||||
(q: Question) => { q.question = q.question.replace('shows Save, Reset', 'shows Publish, Reset'); },
|
||||
(q: Question) => { q.question = q.question.replace('shows Save, Reset', 'shows Save, Save'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('Save:', 'Publish:'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('Reset/Cancel/Export', 'Save/Cancel/Export'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('#1d4ed8', '#abcdef'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('white', 'black'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('filled', 'outlined'); },
|
||||
(q: Question) => { q.options[1]!.label = '1B) Keep three equal buttons'; },
|
||||
(q: Question) => { q.options[1]!.description = 'No current gap remains; no fix is needed.'; },
|
||||
(q: Question) => { q.question = q.question.replace('1A) Filled primary', '1B) Filled primary').replace('1B) Keep four', '1A) Keep four'); },
|
||||
(q: Question) => { q.question = q.question.replace('Save becomes the only filled', 'Publish becomes the only filled'); },
|
||||
(q: Question) => { q.question = q.question.replace('become neutral ghost buttons', 'become filled primary buttons'); },
|
||||
]) expect(accepted(change(edit))).toBe(false);
|
||||
for (const index of [0, 1]) for (const prefix of ['Historical example: ', 'If approved, ', 'Do not apply: ', '> ']) {
|
||||
expect(accepted(change(q => { q.options[index]!.description = prefix + q.options[index]!.description; }))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('owned current withdrawals and approval conditions override affirmative earlier prose', () => {
|
||||
for (const target of [-1, 0, 1]) for (const suffix of [
|
||||
'\nThis finding is withdrawn.', '; This finding is "no longer current".', '; This option is \'withdrawn\'.',
|
||||
'\nThis amendment is ‘no longer current’.', '; This style is `withdrawn`.', '\nIssue 1 is resolved.',
|
||||
'\nCorrection: this gap is already resolved.', '\nNo current violation remains.',
|
||||
'\nOnce approved, apply this amendment.', '\nProvided approval, apply this amendment.',
|
||||
'\nDo not apply this amendment.', '\nNever use these tokens.', '\nSave is already the primary action.',
|
||||
'\nThis finding has no current defect.', '\nThis amendment keeps all four buttons identical.',
|
||||
]) expect(accepted(change(q => {
|
||||
if (target < 0) q.question = q.question.replace('Which option?', `${suffix}\nWhich option?`);
|
||||
else q.options[target]!.description += suffix;
|
||||
}))).toBe(false);
|
||||
});
|
||||
|
||||
test('quoted history, foreign issues and behavior conditions cannot withdraw the current decision', () => {
|
||||
for (const target of [-1, 0, 1]) for (const suffix of [
|
||||
' Prior note: "This finding is withdrawn."', '\n> This amendment is withdrawn.',
|
||||
' Earlier review said `This finding is withdrawn.`', '\nIssue 7 is withdrawn.',
|
||||
'\nIf a user scans the header, Save remains easiest to find.',
|
||||
]) expect(accepted(change(q => {
|
||||
if (target < 0) q.question = q.question.replace('Which option?', `${suffix}\nWhich option?`);
|
||||
else q.options[target]!.description += suffix;
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('the small public fixture and focused regression select only Design finding count', () => {
|
||||
for (const dependency of ['test/design-compact-primary-aw.test.ts', 'test/fixtures/design-compact-primary-aw-call.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name)).toEqual(['plan-design-finding-count']);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -2,8 +2,7 @@ import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { designStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCompletionHandoff, pickDesignCountQuestion } from './helpers/design-count-review';
|
||||
import { hasNativePlanTerminal, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/design-handoff-n-calls.json';
|
||||
import capturedQ from './fixtures/design-handoff-q-calls.json';
|
||||
@@ -11,100 +10,7 @@ import capturedQ from './fixtures/design-handoff-q-calls.json';
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const handoff = () => calls().at(-1)!;
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
function pending(call: NativePlanQuestionCall) {
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('scored Design completion and required next gate', () => {
|
||||
test('the complete native sequence retains all eleven substantive approvals and its separate handoff', () => {
|
||||
const input = calls();
|
||||
const original = structuredClone(input);
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of input) {
|
||||
const phase = planCountQuestionPhase(fp(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, undefined, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 0, review: 11, administrative: 1 });
|
||||
expect(counts.review).toBeGreaterThan(7);
|
||||
expect(input.slice(0, -1).every(call => !isDesignCompletionHandoff(fp(call)))).toBe(true);
|
||||
expect(input).toEqual(original);
|
||||
});
|
||||
|
||||
test('only the actual manual action is selected, in either offered order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = pending(handoff());
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBe(reverse ? 1 : 2);
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
expect(pickDesignCountQuestion(fp(call), { ...fp(call), signature: 'foreign:call' })).toBeNull();
|
||||
}
|
||||
const call = pending(handoff());
|
||||
call.questions[0]!.options[1] = { label: 'Run /plan-ceo-review first' };
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBeNull();
|
||||
});
|
||||
|
||||
test('the retry retains eight real approvals and classifies its required-gate recap separately', () => {
|
||||
const input = structuredClone(captured.retry.calls) as NativePlanQuestionCall[];
|
||||
expect(input).toHaveLength(9);
|
||||
expect(input.slice(0, 8).every(call => !isDesignCompletionHandoff(fp(call)))).toBe(true);
|
||||
const call = input.at(-1)!;
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(true);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBe(2);
|
||||
call.questions[0]!.question = call.questions[0]!.question.replace('8 implementation tasks ready.', 'Please add a missing contrast test.');
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBeNull();
|
||||
});
|
||||
|
||||
test('scores, gate wording or a known identity cannot hide unfinished work or a real choice', () => {
|
||||
const mutations: Array<(call: NativePlanQuestionCall) => void> = [
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('complete (', 'complete only after adding contrast ('); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('review complete', 'review is not complete'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('9 decisions', 'one unresolved decision'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('The required', 'One contrast gap remains. The required'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('The required next gate is Eng Review', 'The optional next gate is Eng Review'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('run it now?', 'fix the missing contrast test now?'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('plan-design-review-next-step', 'plan-design-review-contrast'); },
|
||||
call => { call.questions[0]!.question += ' <gstack-qid:plan-design-review-next-step>'; },
|
||||
call => { call.questions[0]!.options[1]!.label = 'Skip — handle manually and add a missing test'; },
|
||||
call => { call.questions[0]!.options[1]!.description = 'Please add a missing contrast test before proceeding.'; },
|
||||
call => { call.questions[0]!.options[1]!.description = 'Proceed to fix the missing contrast test before the next review.'; },
|
||||
call => { call.questions[0]!.options[1]!.description = 'The contrast gap remains unresolved; handle it manually before Eng.'; },
|
||||
call => { call.questions[0]!.options.push({ label: 'Add a new typeface TODO' }); },
|
||||
call => { call.questions.push(calls()[0]!.questions[0]!); },
|
||||
call => { call.questions[0]!.multiSelect = true; },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('failed, partial, missing-native and unoffered answers do not exclude a call', () => {
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.unansweredQuestionIndices = [0]; },
|
||||
(call: NativePlanQuestionCall) => { call.answers = {}; },
|
||||
(call: NativePlanQuestionCall) => { call.answers = { [call.questions[0]!.question]: 'Build another workflow' }; },
|
||||
]) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
}
|
||||
expect(isDesignCompletionHandoff({ ...fp(handoff()), nativeCall: undefined })).toBe(false);
|
||||
});
|
||||
|
||||
test('the captured report predates only handoff; absent native Exit still cannot complete', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-scored-handoff-'));
|
||||
const file = path.join(dir, 'plan.md');
|
||||
@@ -141,97 +47,4 @@ describe('scored Design completion and required next gate', () => {
|
||||
describe('completed Design review with added decisions and an offered manual stop', () => {
|
||||
const qCalls = () => structuredClone(capturedQ.calls) as NativePlanQuestionCall[];
|
||||
const qHandoff = () => qCalls().at(-1)!;
|
||||
|
||||
test('the actual six calls retain five findings and one completed navigation decision', () => {
|
||||
const input = qCalls();
|
||||
const original = structuredClone(input);
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of input) {
|
||||
const phase = planCountQuestionPhase(fp(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, undefined, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 0, review: 5, administrative: 1 });
|
||||
expect(input.slice(0, -1).every(call => !isDesignCompletionHandoff(fp(call)))).toBe(true);
|
||||
expect(input).toEqual(original);
|
||||
});
|
||||
|
||||
test('pending navigation selects only the offered manual stop in its actual order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = pending(qHandoff());
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBe(reverse ? 1 : 3);
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
expect(pickDesignCountQuestion(fp(call), { ...fp(call), signature: 'foreign:call' })).toBeNull();
|
||||
}
|
||||
const call = pending(qHandoff());
|
||||
call.questions[0]!.options.pop();
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBeNull();
|
||||
});
|
||||
|
||||
test('the new spelling cannot hide described repairs, unfinished work or conditional closure', () => {
|
||||
const mutations: Array<(call: NativePlanQuestionCall) => void> = [
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('is complete', 'is not complete'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('What’s next?', 'Should we add the missing contrast test? What’s next?'); },
|
||||
call => { call.questions[0]!.question = call.questions[0]!.question.replace('What’s next?', 'Once the tests pass, all decisions are resolved. What’s next?'); },
|
||||
call => { call.questions[0]!.options[0]!.description = 'Optional next review.'; },
|
||||
call => { call.questions[0]!.options[2]!.label += ' and fix the missing contrast test'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Proceed to fix the missing contrast test before Eng.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Should we add the missing authorization test before Eng?'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'We could fix the missing authorization test before Eng.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'One contrast gap remains unresolved; handle it manually.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'All decisions will be resolved after the tests pass.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Design review complete after the tests pass.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Design review is not complete.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Not all decisions are resolved.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'The review remains incomplete.'; },
|
||||
call => { call.questions[0]!.options[2]!.description = 'Required gate before shipping. We must repair the missing contrast test.'; },
|
||||
call => { call.questions[0]!.options.push({ label: 'Add a typeface TODO' }); },
|
||||
call => { call.questions[0]!.multiSelect = true; },
|
||||
call => { call.questions.push(qCalls()[0]!.questions[0]!); },
|
||||
call => { call.questions[0]!.question += ' <gstack-qid:plan-design-review-next-steps>'; },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = qHandoff();
|
||||
mutate(call);
|
||||
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('only the administrative answer may postdate the actual completed report', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-q-handoff-'));
|
||||
const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(file, capturedQ.report.content);
|
||||
const written = Date.parse(capturedQ.report.successfulUpdateAt) / 1000;
|
||||
fs.utimesSync(file, written, written);
|
||||
const input = qCalls();
|
||||
const transcript = { status: 'ready' as const, calls: input, assistantMessages: [],
|
||||
planReadyRequests: structuredClone(capturedQ.planReadyRequests) };
|
||||
const administrative = new Set(input.filter(c => isDesignCompletionHandoff(fp(c))).map(c => fp(c).signature));
|
||||
const started = Date.parse('2026-09-09T03:25:54Z');
|
||||
expect(Date.parse(input.at(-2)!.answeredAt!)).toBeLessThan(written * 1000);
|
||||
expect(Date.parse(input.at(-1)!.answeredAt!)).toBeGreaterThan(written * 1000);
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(true);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(false);
|
||||
transcript.planReadyRequests[0]!.failed = false;
|
||||
const stale = Date.parse(input.at(-2)!.answeredAt!) / 1000 - 1;
|
||||
fs.utimesSync(file, stale, stale);
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(false);
|
||||
fs.utimesSync(file, written, written);
|
||||
fs.writeFileSync(file, '# Incomplete report\n');
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,121 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { designStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCompletionHandoff, isDesignCountFirstReview, pickDesignCountQuestion } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import actual from './fixtures/design-handoff-u-calls.json';
|
||||
|
||||
const calls = () => structuredClone(actual) as NativePlanQuestionCall[];
|
||||
const handoff = () => calls().at(-1)!;
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
function answer(call: NativePlanQuestionCall) {
|
||||
call.answers = { [call.questions[0]!.question]: call.questions[0]!.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
function pending(call: NativePlanQuestionCall) {
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('Design completed recap before its required review handoff', () => {
|
||||
test('the exact U calls preserve seven issues and classify only the eighth navigation call separately', () => {
|
||||
const input = calls();
|
||||
const before = structuredClone(input);
|
||||
let reviewStarted = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of input) {
|
||||
const phase = planCountQuestionPhase(fp(call), reviewStarted, designStep0Boundary,
|
||||
isDesignCountFirstReview, undefined, isDesignCompletionHandoff);
|
||||
reviewStarted = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 0, review: 7, administrative: 1 });
|
||||
expect(input.slice(0, 7).every(call => !isDesignCompletionHandoff(fp(call)))).toBe(true);
|
||||
expect(input).toEqual(before);
|
||||
});
|
||||
|
||||
test('the pending exact menu chooses its actual manual option in either order', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = pending(handoff());
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickDesignCountQuestion(fp(call), fp(call))).toBe(reverse ? 1 : 2);
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
expect(pickDesignCountQuestion(fp(call), { ...fp(call), signature: 'foreign:request' })).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('completed recap facts vary without changing the native closed-review decision', () => {
|
||||
for (const recap of [
|
||||
'7 decisions resolved, 6 implementation tasks added, 0 deferred.',
|
||||
'All findings resolved. 6 tasks recorded. No deferred issues.',
|
||||
'The design review recorded accessibility and form-layout requirements. 7 issues addressed.',
|
||||
'This review has approved responsive layout constraints. Zero unresolved decisions.',
|
||||
]) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.question = `Design review complete (6/10 → 9/10). ${recap} Engineering Review is the required shipping gate. What next? <gstack-qid:plan-design-review-next-step>`;
|
||||
expect(isDesignCompletionHandoff(fp(answer(call)))).toBe(true);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBe(2);
|
||||
}
|
||||
});
|
||||
|
||||
test('a closed prefix never hides unfinished work, another decision, source claims or new instructions', () => {
|
||||
const invalid = [
|
||||
'One contrast gap remains.', '7 decisions unresolved.', '1 deferred issue.',
|
||||
'The review is not complete.', 'The review will be complete after contrast is fixed.',
|
||||
'The plan claims that all issues are resolved.', 'The design review added a task; configure the missing states.',
|
||||
'The design review added a task. Configure the missing states.',
|
||||
'The design review added a task and then delete the validation.',
|
||||
'The design review added a task — remove the accessibility check.',
|
||||
'The design review added a task. Should we fix its contrast?',
|
||||
'The design review added a task if the user approves it.',
|
||||
'The design review added a task but the contrast is still missing.',
|
||||
];
|
||||
for (const text of invalid) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.question = `Design review complete. ${text} Eng Review is the required shipping gate. What next? <gstack-qid:plan-design-review-next-step>`;
|
||||
expect(isDesignCompletionHandoff(fp(answer(call)))).toBe(false);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('offered action descriptions cannot smuggle new work or conditional closure', () => {
|
||||
for (const extra of [
|
||||
' Configure a new layout.', ' Remove the missing test.', ' Pick the unresolved color.',
|
||||
' Then implement the spinner.', ' The review is incomplete.',
|
||||
' Once contrast is fixed, all decisions are resolved.',
|
||||
' Please fix the contrast before proceeding.',
|
||||
]) {
|
||||
const call = handoff();
|
||||
call.questions[0]!.options[1]!.description += extra;
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
const active = fp(pending(call));
|
||||
expect(pickDesignCountQuestion(active, active)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('native identity, complete offered answers and the exact binary menu remain required', () => {
|
||||
const mutations: Array<(call: NativePlanQuestionCall) => void> = [
|
||||
c => { c.failed = true; }, c => { c.answered = false; },
|
||||
c => { c.unansweredQuestionIndices = [0]; }, c => { delete c.unansweredQuestionIndices; },
|
||||
c => { c.answers = {}; }, c => { c.answers = { [c.questions[0]!.question]: 'repair another issue' }; },
|
||||
c => { c.questions[0]!.header = 'Contrast'; }, c => { c.questions[0]!.multiSelect = true; },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-design-review-next-step', 'plan-ceo-review-next-step'); },
|
||||
c => { c.questions[0]!.question += ' <gstack-qid:plan-design-review-next-step>'; },
|
||||
c => { c.questions[0]!.options.push({ label: 'Run /plan-ceo-review first' }); },
|
||||
c => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[1]!)); },
|
||||
c => { c.questions[0]!.options[1]!.label += ' and fix contrast'; },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('Eng review is the required shipping gate.', 'Eng review is optional.'); },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = handoff(); mutate(call);
|
||||
expect(isDesignCompletionHandoff(fp(call))).toBe(false);
|
||||
}
|
||||
expect(isDesignCompletionHandoff({ ...fp(handoff()), nativeCall: undefined })).toBe(false);
|
||||
expect(isDesignCompletionHandoff({ ...fp(handoff()), signature: 'foreign:request' })).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,150 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { capturePlanCountQuestion, designStep0Boundary, hasNativePlanTerminal, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCompletionHandoff, pickDesignCountQuestion } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/design-handoff-l-calls.json';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const handoff = () => calls().at(-1)!;
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, false);
|
||||
|
||||
function makePending(call: NativePlanQuestionCall) {
|
||||
call.answered = false;
|
||||
delete call.answers;
|
||||
delete call.unansweredQuestionIndices;
|
||||
return call;
|
||||
}
|
||||
|
||||
function activeQuestion(call: NativePlanQuestionCall) {
|
||||
const q = call.questions[0]!;
|
||||
const visible = `☐ ${q.header}\n${q.question}\n` + q.options.map((option, i) =>
|
||||
`${i ? ' ' : '❯'} ${i + 1}. ${option.label}`).join('\n') +
|
||||
'\nEnter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
return capturePlanCountQuestion(visible, new Set(), 0, false, call)!;
|
||||
}
|
||||
|
||||
describe('Design completed handoff without an offered manual action', () => {
|
||||
test('the captured call is administrative, but its missing manual option is never invented', () => {
|
||||
const call = handoff();
|
||||
expect(isDesignCompletionHandoff(fingerprint(call))).toBe(true);
|
||||
expect(pickDesignCountQuestion(fingerprint(call), fingerprint(call))).toBeNull();
|
||||
makePending(call);
|
||||
expect(isDesignCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
expect(pickDesignCountQuestion(fingerprint(call), activeQuestion(call))).toBeNull();
|
||||
});
|
||||
|
||||
test('the full captured sequence retains all ten decisions and still exceeds the seven-call ceiling', () => {
|
||||
const input = calls();
|
||||
const original = structuredClone(input);
|
||||
let started = false;
|
||||
const counts = { setup: 0, review: 0, administrative: 0 };
|
||||
for (const call of input) {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, undefined, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (phase.administrative) counts.administrative++;
|
||||
else if (phase.preReview) counts.setup++;
|
||||
else counts.review++;
|
||||
}
|
||||
expect(counts).toEqual({ setup: 1, review: 10, administrative: 1 });
|
||||
expect(counts.review).toBeGreaterThan(7);
|
||||
expect(input).toEqual(original);
|
||||
});
|
||||
|
||||
test('an explicit absence of outstanding work remains a closed recap', () => {
|
||||
for (const recap of ['No unresolved design decisions.', 'Zero remaining contrast gaps.', 'No gap remains.']) {
|
||||
const call = handoff();
|
||||
const q = call.questions[0]!;
|
||||
q.question = `Design review complete. ${recap} What’s next? <gstack-qid:plan-design-next-steps>`;
|
||||
call.answers = { [q.question]: q.options[0]!.label };
|
||||
expect(isDesignCompletionHandoff(fingerprint(call))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('an actual manual option is selected in either order only with active native identity', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = makePending(handoff());
|
||||
call.questions[0]!.options.push({ label: "E) Skip — I'll handle next steps manually" });
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
expect(pickDesignCountQuestion(fingerprint(call), activeQuestion(call))).toBe(reverse ? 1 : 4);
|
||||
expect(pickDesignCountQuestion(fingerprint(call), { ...fingerprint(call), signature: 'other' })).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('remaining work, mixed actions and unknown identities stay substantive', () => {
|
||||
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('complete (', 'complete only after resolving contrast ('); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('review complete', 'review is not complete'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('7 decisions made', 'one unresolved gap'); },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('plan-design-next-steps', 'plan-design-contrast-finding'); },
|
||||
c => { c.questions[0]!.question += ' <gstack-qid:plan-design-next-steps>'; },
|
||||
c => { c.questions[0]!.question = 'Design review complete. One contrast gap remains unresolved. What’s next? <gstack-qid:plan-design-next-steps>'; },
|
||||
c => { c.questions[0]!.question = 'Design review complete. One contrast gap remains. What’s next? <gstack-qid:plan-design-next-steps>'; },
|
||||
c => { c.questions[0]!.question = 'Design review complete. There is an unresolved contrast gap. What’s next? <gstack-qid:plan-design-next-steps>'; },
|
||||
c => { c.questions[0]!.question = c.questions[0]!.question.replace('What', '<gstack-qid malformed What'); },
|
||||
c => { c.questions[0]!.header = 'Contrast gap'; },
|
||||
c => { c.questions[0]!.options.push({ label: 'Add the missing contrast test' }); },
|
||||
c => { c.questions[0]!.options[0]!.label = 'Run /plan-eng-review and fix contrast'; },
|
||||
c => { c.questions.push(calls()[1]!.questions[0]!); },
|
||||
c => { c.questions[0]!.multiSelect = true; },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
call.answers = Object.fromEntries(call.questions.map(q => [q.question, q.options[0]!.label]));
|
||||
expect(isDesignCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
const phase = planCountQuestionPhase(fingerprint(call), true, designStep0Boundary,
|
||||
isDesignCountFirstReview, undefined, isDesignCompletionHandoff);
|
||||
expect(phase.administrative).toBeUndefined();
|
||||
expect(phase.preReview).toBe(false);
|
||||
const pending = fingerprint(makePending(call));
|
||||
expect(pickDesignCountQuestion(pending, pending)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('failed, partial, unanswered and free-form results cannot exclude a call', () => {
|
||||
const mutations: Array<(c: NativePlanQuestionCall) => void> = [
|
||||
c => { c.failed = true; },
|
||||
c => { c.answered = false; },
|
||||
c => { c.unansweredQuestionIndices = [0]; },
|
||||
c => { c.answers = {}; },
|
||||
c => { c.answers = { [c.questions[0]!.question]: 'First build a new interaction' }; },
|
||||
];
|
||||
for (const mutate of mutations) {
|
||||
const call = handoff();
|
||||
mutate(call);
|
||||
expect(isDesignCompletionHandoff(fingerprint(call))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the actual late handoff does not stale a valid report, while later real work and failed exits still do', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-handoff-report-'));
|
||||
const file = path.join(dir, 'plan.md');
|
||||
try {
|
||||
fs.writeFileSync(file, '# Reviewed plan\n\n## GSTACK REVIEW REPORT\n\n' +
|
||||
'| Review | Status | Findings |\n|---|---|---|\n| Design | complete | resolved |\n\n' +
|
||||
'VERDICT: DESIGN CLEARED — eng review required\n\nNO UNRESOLVED DECISIONS\n');
|
||||
const input = calls();
|
||||
const transcript = { status: 'ready' as const, calls: input, assistantMessages: [],
|
||||
planReadyRequests: structuredClone(captured.planReadyRequests) };
|
||||
const administrative = new Set(input.filter(c => isDesignCompletionHandoff(fingerprint(c)))
|
||||
.map(c => fingerprint(c).signature));
|
||||
const written = Date.parse('2026-09-08T23:19:12.049Z') / 1000;
|
||||
fs.utimesSync(file, written, written);
|
||||
const started = Date.parse('2026-09-08T23:09:43.875Z');
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready')).toBe(false);
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(true);
|
||||
const stale = Date.parse(input[10]!.answeredAt!) / 1000 - 1;
|
||||
fs.utimesSync(file, stale, stale);
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(false);
|
||||
fs.utimesSync(file, written, written);
|
||||
transcript.planReadyRequests[0]!.failed = true;
|
||||
expect(hasNativePlanTerminal(transcript, file, started, 'plan_ready', administrative)).toBe(false);
|
||||
} finally {
|
||||
fs.rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,38 +0,0 @@
|
||||
import {expect,test} from 'bun:test';
|
||||
import captured from './fixtures/design-count-ad-v2.json';
|
||||
import {planCountQuestionPhase,designStep0Boundary,nativePlanCallFingerprint} from './helpers/claude-pty-runner';
|
||||
import {isDesignCountFirstReview,isDesignCountSetup} from './helpers/design-count-review';
|
||||
test('actual completed ordinary design Issue starts review at the finding, without counting later Eng work',()=>{
|
||||
expect(isDesignCountFirstReview(captured.firstFinding)).toBe(true);
|
||||
expect(planCountQuestionPhase(captured.firstFinding,false,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup)).toMatchObject({preReview:false,reviewStarted:true});
|
||||
});
|
||||
|
||||
test('ordinary design finding retains completed native identity and unresolved alternatives',()=>{
|
||||
for(const update of [
|
||||
(f:any)=>{f.nativeCall.answered=false;},(f:any)=>{f.nativeCall.failed=true;},
|
||||
(f:any)=>{f.signature='foreign';},(f:any)=>{f.nativeCall.unansweredQuestionIndices=[0];},
|
||||
(f:any)=>{f.nativeCall.answers={};},(f:any)=>{f.nativeCall.questions[0].header='Issue 2';},
|
||||
(f:any)=>{f.nativeCall.questions[0].multiSelect=true;},
|
||||
(f:any)=>{f.nativeCall.questions[0].question='Example: '+f.nativeCall.questions[0].question;},
|
||||
(f:any)=>{f.options.reverse();},
|
||||
]){const f=structuredClone(captured.firstFinding);update(f);expect(isDesignCountFirstReview(f)).toBe(false);}
|
||||
const f=structuredClone(captured.firstFinding);const q=f.nativeCall.questions[0]!;const old=q.question;
|
||||
q.question=q.question.replace('D4 — ','D38: ');f.nativeCall.answers={[q.question]:f.nativeCall.answers[old]!};
|
||||
expect(isDesignCountFirstReview(f)).toBe(true);
|
||||
});
|
||||
|
||||
test('optional gap tags do not decide the finding boundary',()=>{
|
||||
for(const replacement of [' (Visual Hierarchy)', '']){const f=structuredClone(captured.firstFinding),q=f.nativeCall.questions[0]!,old=q.question;q.question=q.question.replace(' (G1, Visual Hierarchy)',replacement);f.nativeCall.answers={[q.question]:f.nativeCall.answers[old]!};expect(isDesignCountFirstReview(f)).toBe(true);}
|
||||
});
|
||||
import {E2E_TOUCHFILES,selectTests} from './helpers/touchfiles';
|
||||
test('the retained Design calls select the native cadence workflow',()=>{
|
||||
for(const file of ['test/design-count-ad-v2.test.ts','test/fixtures/design-count-ad-v2.json'])
|
||||
expect(selectTests([file],E2E_TOUCHFILES,[]).selected).toEqual(['plan-design-finding-count']);
|
||||
});
|
||||
|
||||
test('an Issue label for participation or next-review routing is still setup',()=>{
|
||||
for(const [title,labels,description] of [
|
||||
['D4 — Issue 1: how should we address optional outside-review participation?', ['Run outside voices','Defer outside voices'],'Applies independent review to the plan.'],
|
||||
['D4 — Issue 1: how should we resolve which review runs next?', ['Run the engineering review','Defer next reviews'],'Closes the required engineering review gate.'],
|
||||
] as const){const c=structuredClone(captured.firstFinding.nativeCall),q=c.questions[0]!;q.question=title;q.options=labels.map((label,i)=>({label:`1${i?'B':'A'}) ${label}`,description}));c.answers={[title]:q.options[0]!.label};expect(isDesignCountFirstReview(nativePlanCallFingerprint(c,0,true))).toBe(false);}
|
||||
});
|
||||
@@ -1,84 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-count-current-pass.json';
|
||||
import { nativePlanCallFingerprint, planCountQuestionPhase, designStep0Boundary } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
const accepts = (call: NativePlanQuestionCall) => isDesignCountFirstReview(fingerprint(call));
|
||||
test('the first current design issue starts review with its actual native answer and no Net summary', () => {
|
||||
const call = calls()[3]!;
|
||||
for (const option of call.questions[0]!.options) {
|
||||
call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(accepts(call)).toBe(true);
|
||||
}
|
||||
});
|
||||
test('all observed substantive calls count, including extra findings and the TODO proposal', () => {
|
||||
const input = calls(), before = JSON.stringify(input);
|
||||
let started = false;
|
||||
const phases = input.map(call => {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(phases.map(p => p.preReview)).toEqual([true, true, true, false, false, false, false, false, false, false, false, false]);
|
||||
// Nine decisions exceed the paid case's existing ceiling of seven.
|
||||
expect(phases.filter(p => !p.preReview && !p.administrative)).toHaveLength(9);
|
||||
expect(JSON.stringify(input)).toBe(before);
|
||||
});
|
||||
const invalid = {
|
||||
'foreign file': (q: any) => { q.question = q.question.replace('of the Account settings plan.', 'of OTHER.md.'); },
|
||||
'quoted owner': (q: any) => { q.question = q.question.replace('Pass 1 (Information Architecture) of the Account settings plan.', '"Pass 1 (Information Architecture) of the Account settings plan."'); },
|
||||
'historical owner': (q: any) => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: Historical example: '); },
|
||||
'setup owner': (q: any) => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: Review setup phase; '); },
|
||||
'quoted assertion': (q: any) => { q.question = q.question.replace(/^ELI10: (.*)$/m, 'ELI10: "$1"'); },
|
||||
'conditional assertion': (q: any) => { q.question = q.question.replace('ELI10: ', 'ELI10: If approved later, '); },
|
||||
'different native identity': (q: any) => { q.header = 'Issue 99'; },
|
||||
'missing remedy': (q: any) => { q.options[0].description = 'We can discuss this later.'; },
|
||||
'foreign opposition without a withdrawal': (q: any) => { q.options[2].description = "Another Issue 99 violates DESIGN.md's stated primary treatment."; },
|
||||
'foreign plan without a filename': (q: any) => { q.question = q.question.replace('Account settings plan', 'unrelated plan'); },
|
||||
'unowned opposition': (q: any) => { q.options[2].description = 'Another issue violates DESIGN.md, this issue is resolved.'; },
|
||||
'missing current opposition': (q: any) => { q.options[2].description = 'This menu remains available.'; },
|
||||
'withdrawn decision': (q: any) => { q.question += '\nD4 is withdrawn.'; },
|
||||
'quoted withdrawn status': (q: any) => { q.question += '\nThis finding is "withdrawn".'; },
|
||||
'foreign recommendation': (q: any) => { q.question = q.question.replace('Recommendation: 1A', 'Recommendation: 99A'); },
|
||||
};
|
||||
for (const [name, mutate] of Object.entries(invalid)) test('count still rejects ' + name, () => {
|
||||
const call = calls()[3]!, q = call.questions[0]!;
|
||||
mutate(q); call.answers = { [q.question]: q.options[0]!.label };
|
||||
expect(accepts(call)).toBe(false);
|
||||
});
|
||||
test('pending, failed, foreign and unoffered native acknowledgments never establish review', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'not offered' }; },
|
||||
]) { const call = calls()[3]!; mutate(call); expect(accepts(call)).toBe(false); }
|
||||
const fp = fingerprint(calls()[3]!); fp.signature = 'foreign:call'; expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
});
|
||||
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { designCountExistingInteractionStates } from './helpers/design-count-fixture';
|
||||
test('both fixture documents define existing error layout and export behavior while preserving all five gaps', () => {
|
||||
const source = readFileSync(join(import.meta.dir, 'skill-e2e-plan-design-finding-count.test.ts'), 'utf8');
|
||||
expect(source).toContain("import { designCountExistingInteractionStates as existingInteractionStates } from './helpers/design-count-fixture';");
|
||||
const start = source.indexOf('const designSystem = ');
|
||||
const end = source.indexOf("describeE2E(", start);
|
||||
expect(start).toBeGreaterThan(0); expect(end).toBeGreaterThan(start);
|
||||
const build = new Function('existingInteractionStates', new Bun.Transpiler({ loader: 'ts' }).transformSync(source.slice(start, end) + '\nreturn { designSystem, plan: planDesign5Findings("/owned/review.md") };'));
|
||||
const { designSystem, plan } = build(designCountExistingInteractionStates);
|
||||
for (const text of [designSystem, plan]) {
|
||||
expect(text).toContain('The existing ErrorSummary mounts in the status/error area below the action\ngroup and above Profile.');
|
||||
expect(text).toContain('Retry wraps below the text as a full-width 44px ghost button');
|
||||
expect(text).toContain('account-settings-YYYY-MM-DD.json');
|
||||
expect(text).toContain('outside the live region');
|
||||
}
|
||||
for (const name of ['Visual Hierarchy', 'Spacing', 'Typography', 'Color', 'Motion']) expect(plan).toContain('## ' + name);
|
||||
expect(plan).toContain('same size, weight, and color');
|
||||
expect(plan).toContain('no consistent vertical rhythm');
|
||||
expect(plan).toContain('14px, 16px, and 18px');
|
||||
});
|
||||
@@ -1,104 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-count-sep20-calls.json';
|
||||
import { designCountExistingInteractionStates } from './helpers/design-count-fixture';
|
||||
import { designStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCompletionHandoff, isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
describe('September 20 design count fixture omissions', () => {
|
||||
test('the failed retry contains eight real decisions, including three unseeded requirements', () => {
|
||||
let started = false;
|
||||
const reviewHeaders: string[] = [];
|
||||
for (const call of structuredClone(captured.calls) as NativePlanQuestionCall[]) {
|
||||
const phase = planCountQuestionPhase(nativePlanCallFingerprint(call, 0, true), started,
|
||||
designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
if (!phase.preReview && !phase.administrative) reviewHeaders.push(call.questions[0]!.header);
|
||||
}
|
||||
expect(reviewHeaders).toEqual(Array.from({ length: 8 }, (_, index) => `Issue ${index + 1}`));
|
||||
expect(captured.provenance.expectedCeiling).toBe(7);
|
||||
for (const header of captured.provenance.unseededHeaders) expect(reviewHeaders).toContain(header);
|
||||
});
|
||||
|
||||
test('the first finding owns its review evidence independently of the earlier mixed setup packet', () => {
|
||||
const first = (structuredClone(captured.calls) as NativePlanQuestionCall[])
|
||||
.find(call => call.questions[0]!.header === 'Issue 1')!;
|
||||
const q = first.questions[0]!;
|
||||
for (const option of q.options) {
|
||||
first.answers = { [q.question]: option.label };
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(first, 0, true))).toBe(true);
|
||||
}
|
||||
for (const opposition of [
|
||||
'Leaves the plan no longer violating DESIGN.md.',
|
||||
'Leaves another plan violating DESIGN.md.',
|
||||
'"Leaves the plan violating DESIGN.md."',
|
||||
'Leaves the plan violating DESIGN.md. This issue is resolved.',
|
||||
]) {
|
||||
const changed = structuredClone(first);
|
||||
changed.questions[0]!.options[2]!.description = opposition;
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(changed, 0, true)), opposition).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('retained contract violations require an affirmative, unconditional alternative', () => {
|
||||
const first = (structuredClone(captured.calls) as NativePlanQuestionCall[])
|
||||
.find(call => call.questions[0]!.header === 'Issue 1')!;
|
||||
const accepts = (description: string) => {
|
||||
const changed = structuredClone(first);
|
||||
changed.questions[0]!.options[2]!.description = description;
|
||||
return isDesignCountFirstReview(nativePlanCallFingerprint(changed, 0, true));
|
||||
};
|
||||
for (const verb of ['Leave', 'Keep']) for (const owner of ['the plan', 'this header', 'the design', 'this page']) {
|
||||
const action = `${verb.toLowerCase()} ${owner} violating DESIGN.md`;
|
||||
const assertion = `${verb}s ${owner} violating DESIGN.md`;
|
||||
for (const positive of [
|
||||
assertion + '.',
|
||||
`✅ No visual change to review. ❌ ${assertion} and users scanning four labels.`,
|
||||
assertion + '. Users still scan the labels. Historical note: "Never ' + action + '."',
|
||||
]) expect(accepts(positive), positive).toBe(true);
|
||||
for (const negative of [
|
||||
`Does not ${action}.`, `Never ${action}.`, `Do not ${action}.`,
|
||||
`Cannot ${action}.`, `Must not ${action}.`, `Should not ${action}.`,
|
||||
`If approved, ${assertion.toLowerCase()}.`,
|
||||
`Assuming approval, ${assertion.toLowerCase()}.`,
|
||||
`${assertion} only if approved later.`, `${assertion} once approval arrives.`,
|
||||
`${assertion} after approval.`, `${assertion} subject to approval.`,
|
||||
`${assertion}; pending approval.`, `${assertion}. This alternative requires approval.`,
|
||||
`${assertion}. Correction: do not ${action}.`,
|
||||
`${assertion}. This option does not ${action}.`,
|
||||
]) expect(accepts(negative), negative).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
const accepted = designCountExistingInteractionStates.join(' ');
|
||||
|
||||
test('the surrounding contract supplies the three missing operation-specific error strings', () => {
|
||||
expect(accepted).toContain('Save: “Couldn’t save your changes. Your edits are still here.”');
|
||||
expect(accepted).toContain('Export: “Couldn’t prepare your export.”');
|
||||
expect(accepted).toContain('Load: “Couldn’t load your settings.”');
|
||||
expect(accepted).toContain('Each uses the existing error icon and its sibling Retry');
|
||||
});
|
||||
|
||||
test('the surrounding contract defines a clean Save without changing its pending or dirty behavior', () => {
|
||||
expect(accepted).toContain('Save stays enabled and focusable while idle, whether clean or dirty.');
|
||||
expect(accepted).toContain('A clean Save is a no-op: no request, validation, pending state, timestamp, status, or focus change.');
|
||||
expect(accepted).toContain('Only a dirty Save sends the existing atomic request.');
|
||||
expect(accepted).toContain('both request buttons use aria-disabled=true');
|
||||
});
|
||||
|
||||
test('the surrounding contract names exports without introducing personal data or a date ambiguity', () => {
|
||||
expect(accepted).toContain('account-settings-YYYY-MM-DD.json');
|
||||
expect(accepted).toContain('the user’s local calendar date');
|
||||
expect(accepted).toContain('no account name or email');
|
||||
expect(accepted).toContain('no account identifiers');
|
||||
expect(accepted).toContain('Repeated same-day exports keep the browser’s normal collision suffix');
|
||||
});
|
||||
|
||||
test('the surrounding contract locates validation errors and responsive retry feedback', () => {
|
||||
expect(accepted).toContain('ErrorSummary mounts in the status/error area below the action group and above Profile');
|
||||
expect(accepted).toContain('focus goes to the first invalid field and the summary is not a second live region');
|
||||
expect(accepted).toContain('error/Retry row is inline above 640px with an 8px gap');
|
||||
expect(accepted).toContain('Retry wraps below the text as a full-width 44px ghost button, outside the live region');
|
||||
expect(accepted).toContain('long errors fit 320px without horizontal scroll');
|
||||
});
|
||||
});
|
||||
@@ -3,11 +3,9 @@ import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import fixture from './fixtures/design-count-native-8525.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff, pickDesignCountQuestion } from './helpers/design-count-review';
|
||||
import { nativePlanCallFingerprint, planCountQuestionPhase, designStep0Boundary, hasNativePlanTerminal } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall, PlanCountTranscript } from './helpers/plan-count-transcript';
|
||||
const calls = () => structuredClone(fixture.transcript.calls) as NativePlanQuestionCall[];
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
|
||||
import { hasNativePlanTerminal } from './helpers/claude-pty-runner';
|
||||
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
|
||||
|
||||
function completion() {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-8525-replay-'));
|
||||
const file = path.join(dir, path.basename(fixture.provenance.planPath));
|
||||
@@ -22,65 +20,10 @@ function completion() {
|
||||
const check = () => hasNativePlanTerminal(transcript, file, startedAt, 'completion_summary');
|
||||
return { dir, file, transcript, final, write, check, cleanup: () => fs.rmSync(dir, {recursive:true, force:true}) };
|
||||
}
|
||||
test('full exact native attempt starts review at Issue 1 and counts six independently acknowledged decisions', () => {
|
||||
const input = calls(); let started = false; const counts = {step0:0,review:0,administrative:0};
|
||||
expect(isDesignCountFirstReview(fp(input[0]!))).toBe(false);
|
||||
expect(isDesignCountFirstReview(fp(input[1]!))).toBe(true);
|
||||
for (const call of input) {
|
||||
const p = planCountQuestionPhase(fp(call), started, designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
counts[p.administrative ? 'administrative' : p.preReview ? 'step0' : 'review']++;
|
||||
started = p.reviewStarted;
|
||||
}
|
||||
expect(counts).toEqual({step0:1,review:6,administrative:0});
|
||||
expect(counts.review).toBeGreaterThanOrEqual(4); expect(counts.review).toBeLessThanOrEqual(7);
|
||||
});
|
||||
|
||||
test('exact native final text and reconstructed read-back-verified report supply completion', () => {
|
||||
const f = completion(); try { expect(f.check()).toBe(true); } finally { f.cleanup(); }
|
||||
});
|
||||
const changedQuestion = (change: (c: NativePlanQuestionCall) => void) => {
|
||||
const c = calls()[1]!; change(c); const q = c.questions[0]!;
|
||||
c.answers = {[q.question]:q.options[0]!.label}; return c;
|
||||
};
|
||||
for (const [name, change] of Object.entries({
|
||||
'unrelated setup header': (c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Routing'; },
|
||||
'wrong native Issue header': (c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Issue 2'; },
|
||||
'wrong offered Issue ids': (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label = '2A: Filled primary Save'; },
|
||||
'multiselect': (c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
'another bundled question': (c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
'missing current source': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('PLAN.md','other.md'); },
|
||||
'quoted current source': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('PLAN.md','"PLAN.md"'); },
|
||||
'multiple source gaps': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('gap G1','gap G1 and gap G2'); },
|
||||
'unowned gap in alternative': (c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.description = 'Leave G2 open; the gap stays open.'; },
|
||||
'no current defect': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('all four header buttons look identical','the header buttons have distinct approved styles'); },
|
||||
'quoted only defect': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace(/ELI10: ([\s\S]*?)\nStakes/, 'ELI10: "$1"\nStakes'); },
|
||||
'historical assessment': (c: NativePlanQuestionCall) => { c.questions[0]!.question = c.questions[0]!.question.replace('ELI10:','ELI10: Historical example:'); },
|
||||
'withdrawn current issue': (c: NativePlanQuestionCall) => { c.questions[0]!.question += '\nThis issue is withdrawn.'; },
|
||||
'quoted withdrawn state': (c: NativePlanQuestionCall) => { c.questions[0]!.question += '\nThis issue is "withdrawn".'; },
|
||||
'no concrete offered remedy': (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = '✅ Follow the design system.'; },
|
||||
'quoted only remedy': (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = '"' + c.questions[0]!.options[0]!.description + '"'; },
|
||||
'no opposed open gap': (c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.description = 'The question remains available for discussion.'; },
|
||||
'quoted question': (c: NativePlanQuestionCall) => { c.questions[0]!.question = '> ' + c.questions[0]!.question.replaceAll('\n','\n> '); },
|
||||
'code example': (c: NativePlanQuestionCall) => { c.questions[0]!.question = '```text\n' + c.questions[0]!.question + '\n```'; },
|
||||
})) test(`named current issue rejects ${name}`, () => expect(isDesignCountFirstReview(fp(changedQuestion(change)))).toBe(false));
|
||||
test('native ownership and actual answer remain required', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered=false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed=true; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices=[0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.answers={[c.questions[0]!.question]:'an unoffered recommendation'}; },
|
||||
]) { const c=calls()[1]!; mutate(c); expect(isDesignCountFirstReview(fp(c))).toBe(false); }
|
||||
const c=calls()[1]!; expect(isDesignCountFirstReview({...fp(c),signature:'foreign:question'})).toBe(false);
|
||||
});
|
||||
test('the source gap and design-defect class are independent of seeded spelling or G-number', () => {
|
||||
const c=changedQuestion(c => { c.questions[0]!.question=c.questions[0]!.question.replaceAll('G1','G22').replaceAll('Save','Submit');
|
||||
c.questions[0]!.options.forEach(o=>{o.label=o.label.replaceAll('Save','Submit');o.description=o.description?.replaceAll('Save','Submit');}); });
|
||||
for (const o of c.questions[0]!.options) { c.answers={[c.questions[0]!.question]:o.label};expect(isDesignCountFirstReview(fp(c))).toBe(true); }
|
||||
const coded=changedQuestion(c=>{c.questions[0]!.question=c.questions[0]!.question.replace('PLAN.md','`PLAN.md`');});
|
||||
expect(isDesignCountFirstReview(fp(coded))).toBe(true);
|
||||
c.answered=false;delete c.answers;delete c.unansweredQuestionIndices;
|
||||
expect(pickDesignCountQuestion(fp(c),fp(c))).toBeNull(); // Existing actor/default answer ownership is unchanged.
|
||||
});
|
||||
test('current typed status accepts presentation, field order and current report prose independently', () => {
|
||||
const f=completion();try {
|
||||
for (const heading of ['## Completion','### Completion summary','## Review complete','## Design review complete','**Review completion:**']) {
|
||||
@@ -139,144 +82,8 @@ test('typed delivery retains source session, answer chronology, fresh file and c
|
||||
const alternate=path.join(f.dir,'alternate.md');fs.writeFileSync(alternate,fixture.report);fs.symlinkSync(alternate,f.file);expect(f.check()).toBe(false);
|
||||
}finally{f.cleanup();}
|
||||
});
|
||||
test('cancelled retry current native Issue is still classified without supplying terminal coverage', () => {
|
||||
const input=structuredClone(fixture.cancelledRetry.calls) as NativePlanQuestionCall[];
|
||||
expect(fixture.cancelledRetry.coverageCredit).toBe(0);
|
||||
expect(input).toHaveLength(2);expect(isDesignCountFirstReview(fp(input[0]!))).toBe(false);
|
||||
expect(isDesignCountFirstReview(fp(input[1]!))).toBe(true);
|
||||
const review=planCountQuestionPhase(fp(input[1]!),false,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff);
|
||||
expect(review.preReview).toBe(false);
|
||||
});
|
||||
test('an unlabelled source gap still needs a current defect, concrete offered repair and its own retained violation', () => {
|
||||
const original=fixture.cancelledRetry.calls[1]! as NativePlanQuestionCall;
|
||||
for (const mutate of [
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('currently look identical','already have distinct correct styles');},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('Pass 1 Information Architecture','planning setup');},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('PLAN.md','other.md');},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('ELI10:','ELI10: Historical example:');},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description='Use the Button component as appropriate.';},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description='This resolves the hierarchy gap completely.';},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description='The plan keeps a documented DESIGN.md violation for G9.';},
|
||||
(q:NativePlanQuestionCall['questions'][number])=>{q.header='Setup';},
|
||||
]) {const c=structuredClone(original);mutate(c.questions[0]!);c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};expect(isDesignCountFirstReview(fp(c))).toBe(false);}
|
||||
});
|
||||
|
||||
|
||||
import phaseEntry77 from './fixtures/design-phase-entry-77.json';
|
||||
function phaseCalls77() { return structuredClone(phaseEntry77.calls) as NativePlanQuestionCall[]; }
|
||||
function phaseSequence77(calls = phaseCalls77()) {
|
||||
let started = false;
|
||||
return calls.map(call => {
|
||||
const f = nativePlanCallFingerprint(call, 0, !started);
|
||||
const phase = planCountQuestionPhase(f, started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return { id: call.toolUseId, ...phase };
|
||||
});
|
||||
}
|
||||
function phaseMutation77(index: number, mutate: (call: NativePlanQuestionCall) => void) {
|
||||
const call = phaseCalls77()[index]!; const before = call.questions[0]!.question;
|
||||
const answer = call.answers![before]!; mutate(call);
|
||||
if (call.questions[0]!.question !== before) call.answers = {[call.questions[0]!.question]: answer};
|
||||
return nativePlanCallFingerprint(call, 0, true);
|
||||
}
|
||||
|
||||
test('actual77 focus ACK opens review, later learnings stays setup, all six real findings count', () => {
|
||||
const phases = phaseSequence77();
|
||||
expect(phases.slice(0, 3).map(p => p.preReview)).toEqual([true, true, true]);
|
||||
expect(phases[1]!.reviewStarted).toBe(true);
|
||||
expect(phases.slice(3).map(p => p.preReview)).toEqual([false, false, false, false, false, false]);
|
||||
expect(phases.filter(p => !p.preReview)).toHaveLength(6);
|
||||
const calls = phaseCalls77();
|
||||
// These remain setup decisions, never substituted for a substantive finding.
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(calls[1]!, 0, true))).toBe(false);
|
||||
expect(isDesignCountSetup(nativePlanCallFingerprint(calls[2]!, 0, false))).toBe(true);
|
||||
});
|
||||
|
||||
test('native focus and learnings classification follows scope actions, not recommendation or order', () => {
|
||||
for (const index of [1,2]) for (const reversed of [false,true]) for (const picked of [0,1]) {
|
||||
const call = phaseCalls77()[index]!; const q=call.questions[0]!;
|
||||
q.options.forEach(o => { o.label=o.label.replace(/\s*\(recommended\)/i,''); });
|
||||
q.options[picked]!.label += ' (recommended)';
|
||||
if(reversed)q.options.reverse();
|
||||
call.answers = {[q.question]:q.options[picked]!.label};
|
||||
const f=nativePlanCallFingerprint(call,0,true);
|
||||
expect(designStep0Boundary(f)).toBe(true);
|
||||
expect(isDesignCountSetup(f)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('equivalent all-seven versus subset focus wording stays a plan-wide setup choice', () => {
|
||||
for(const title of ['Review all 7 design dimensions, or focus on specific areas?', 'Review all 7 dimensions or focus on a subset?', 'Review all 7 design passes, or focus?']) {
|
||||
const f=phaseMutation77(1,c=>{c.questions[0]!.question=c.questions[0]!.question.replace(/^D2[^\n]+/,'D21: '+title);});
|
||||
expect(designStep0Boundary(f)).toBe(true); expect(isDesignCountSetup(f)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
for (const [name, mutate] of Object.entries({
|
||||
'pending': (c: NativePlanQuestionCall) => { c.answered=false; },
|
||||
'failed': (c: NativePlanQuestionCall) => { c.failed=true; },
|
||||
'missing answer time': (c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
'missing session': (c: NativePlanQuestionCall) => { c.sessionId=''; },
|
||||
'missing call ID': (c: NativePlanQuestionCall) => { c.toolUseId=''; },
|
||||
'partial': (c: NativePlanQuestionCall) => { c.unansweredQuestionIndices=[0]; },
|
||||
'unoffered answer': (c: NativePlanQuestionCall) => { c.answers={[c.questions[0]!.question]:'Unrelated answer'}; },
|
||||
'checkbox': (c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect=true; },
|
||||
'mixed packet': (c: NativePlanQuestionCall) => { c.questions.push({header:'Issue',question:'Approve a new layout?',options:[{label:'Approve'},{label:'Defer'}],multiSelect:false}); },
|
||||
'foreign source': (c: NativePlanQuestionCall) => { c.questions[0]!.question=c.questions[0]!.question.replace('of PLAN.md','of OTHER.md'); },
|
||||
'historical source': (c: NativePlanQuestionCall) => { c.questions[0]!.question=c.questions[0]!.question.replace('Project/branch/task:','Project/branch/task: Historical source:'); },
|
||||
'quoted question': (c: NativePlanQuestionCall) => { c.questions[0]!.question='> '+c.questions[0]!.question; },
|
||||
'additional approval': (c: NativePlanQuestionCall) => { c.questions[0]!.question+='\nApprove all findings?'; },
|
||||
'extra option': (c: NativePlanQuestionCall) => { c.questions[0]!.options.push({label:'Approve deployment'}); },
|
||||
'extra option action': (c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label+=' and approve the plan'; c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label}; },
|
||||
})) for(const index of [1,2])test(`native ${index===1?'focus':'learnings'} does not classify ${name} as setup`,()=>{
|
||||
const f=phaseMutation77(index,mutate);
|
||||
expect(designStep0Boundary(f)).toBe(false); expect(isDesignCountSetup(f)).toBe(false);
|
||||
});
|
||||
|
||||
test('narrow-only, duplicated scope, quoted rating and component rating do not open review',()=>{
|
||||
for(const mutate of [
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options[0]!.label='Only the 2 listed gaps';},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options[1]!.label='All 7 dimensions';},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.question=c.questions[0]!.question.replace("ELI10: I've rated this plan", "ELI10: Earlier: I've rated this plan");},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.question=c.questions[0]!.question.replace("rated this plan", "rated this error message");},
|
||||
]) { const f=phaseMutation77(1,mutate);expect(designStep0Boundary(f)).toBe(false);expect(isDesignCountSetup(f)).toBe(false); }
|
||||
const f=nativePlanCallFingerprint(phaseCalls77()[1]!,0,true);
|
||||
for(const changed of [{...f,signature:'foreign:tool'},{...f,nativeQuestionIndex:1},{...f,options:[...f.options].reverse()}]) {
|
||||
expect(designStep0Boundary(changed)).toBe(false);expect(isDesignCountSetup(changed)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
for (const suffix of ['Also approve deployment.', 'Approve all findings.', 'Continue the review and deploy to production.', 'Review while deleting the API.'])
|
||||
for (const location of ['question', 'option'] as const) for (const index of [1, 2])
|
||||
test(`native setup rejects mixed current action in ${location}: ${suffix} (${index})`, () => {
|
||||
const f = phaseMutation77(index, c => {
|
||||
if (location === 'question') c.questions[0]!.question += '\n' + suffix;
|
||||
else c.questions[0]!.options[0]!.description += ' ' + suffix;
|
||||
});
|
||||
expect(designStep0Boundary(f)).toBe(false); expect(isDesignCountSetup(f)).toBe(false);
|
||||
});
|
||||
for (const index of [1, 2]) test(`native setup rejects contradictory duplicate source (${index})`, () => {
|
||||
const f = phaseMutation77(index, c => { c.questions[0]!.question += '\nProject/branch/task: plan-design-review of OTHER.md.'; });
|
||||
expect(designStep0Boundary(f)).toBe(false); expect(isDesignCountSetup(f)).toBe(false);
|
||||
});
|
||||
for (const index of [1, 2]) test(`native setup allows quoted examples and negative consequences without approving them (${index})`, () => {
|
||||
const f = phaseMutation77(index, c => {
|
||||
c.questions[0]!.question += '\nExample of a later finding: "Approve deployment." This scope choice does not approve that action.';
|
||||
c.questions[0]!.options[0]!.description += ' ❌ This does not approve deployment. Example: “Approve all findings.”';
|
||||
});
|
||||
expect(designStep0Boundary(f)).toBe(true); expect(isDesignCountSetup(f)).toBe(true);
|
||||
});
|
||||
|
||||
for (const index of [1, 2]) for (const suffix of ['Also approve the design system.', 'Implement the first dimension.'])
|
||||
test(`native setup rejection cannot fall through to a legacy boundary (${index}): ${suffix}`, () => {
|
||||
const f = phaseMutation77(index, c => { c.questions[0]!.question += '\n' + suffix; });
|
||||
expect(designStep0Boundary(f)).toBe(false); expect(isDesignCountSetup(f)).toBe(false);
|
||||
});
|
||||
|
||||
|
||||
const cf74 = fixture.cf74Retry;
|
||||
function cf74Calls() { return structuredClone(cf74.transcript.calls) as NativePlanQuestionCall[]; }
|
||||
|
||||
function cf74Completion() {
|
||||
const dir=fs.mkdtempSync(path.join(os.tmpdir(),'design-cf74-completion-'));
|
||||
const file=path.join(dir,path.basename(cf74.provenance.planPath));
|
||||
@@ -288,75 +95,10 @@ function cf74Completion() {
|
||||
write();
|
||||
return {dir,file,transcript,final,startedAt,write,check:()=>hasNativePlanTerminal(transcript,file,startedAt,'completion_summary'),cleanup:()=>fs.rmSync(dir,{recursive:true,force:true})};
|
||||
}
|
||||
test('cf74 complete current styling decision starts the seven acknowledged review choices',()=>{
|
||||
const input=cf74Calls();let started=false;const counts={step0:0,review:0,administrative:0};
|
||||
expect(isDesignCountFirstReview(fp(input[0]!))).toBe(true);
|
||||
for(const call of input){const p=planCountQuestionPhase(fp(call),started,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff);counts[p.administrative?'administrative':p.preReview?'step0':'review']++;started=p.reviewStarted;}
|
||||
expect(counts).toEqual({step0:0,review:7,administrative:0});
|
||||
expect(counts.review).toBeGreaterThanOrEqual(4);expect(counts.review).toBeLessThanOrEqual(7);
|
||||
});
|
||||
|
||||
test('cf74 actual completed native report envelope binds the fresh owned Design report',()=>{
|
||||
const f=cf74Completion();try{expect(f.check()).toBe(true);}finally{f.cleanup();}
|
||||
});
|
||||
|
||||
const changeCf74=(change:(q:NativePlanQuestionCall['questions'][number])=>void)=>{
|
||||
const call=cf74Calls()[0]!,q=call.questions[0]!;change(q);call.answers={[q.question]:q.options[0]!.label};return call;
|
||||
};
|
||||
for(const primary of ['Save','Submit'])for(const peerOrder of ['Reset, Cancel, Export','Export, Cancel, Reset'])for(const prefix of ['Matches DESIGN.md exactly','Apply DESIGN.md tokens','Use DESIGN.md'])
|
||||
test(`cf74 complete attributed styling keeps named role ownership: ${primary}/${peerOrder}/${prefix}`,()=>{
|
||||
const call=changeCf74(q=>{q.question=q.question.replaceAll('Save',primary);q.options.forEach(o=>{o.label=o.label.replaceAll('Save',primary);o.description=o.description?.replaceAll('Save',primary);});
|
||||
q.options[0]!.description=q.options[0]!.description!.replace('Matches DESIGN.md exactly',prefix).replace('Reset, Cancel, Export',peerOrder);});
|
||||
for(const chosen of call.questions[0]!.options){call.answers={[call.questions[0]!.question]:chosen.label};expect(isDesignCountFirstReview(fp(call))).toBe(true);}
|
||||
});
|
||||
for(const [name,change] of Object.entries({
|
||||
'foreign current source':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replaceAll('DESIGN.md','OTHER.md');},
|
||||
'quoted source':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replaceAll('DESIGN.md','"DESIGN.md"');},
|
||||
'duplicate source field':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nProject/branch/task: another source.';},
|
||||
'foreign owner':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nThis finding belongs to another project.';},
|
||||
'historical premise':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('ELI10: Right now','ELI10: Historically');},
|
||||
'quoted premise':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace(/ELI10: (.+)/,'ELI10: "$1"');},
|
||||
'single-quoted premise':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace(/ELI10: (.+)/,"ELI10: '$1'");},
|
||||
'quoted entire question':(q:NativePlanQuestionCall['questions'][number])=>{q.question='> '+q.question.replaceAll('\n','\n> ');},
|
||||
'no current equal-weight defect':(q:NativePlanQuestionCall['questions'][number])=>{q.question=q.question.replace('look identical','already have distinct correct styles');},
|
||||
'withdrawn issue':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nThis issue is withdrawn.';},
|
||||
'quoted current withdrawn status':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nThis issue is "withdrawn".';},
|
||||
'single quoted withdrawn status':(q:NativePlanQuestionCall['questions'][number])=>{q.question+="\nThis issue is 'withdrawn'.";},
|
||||
'withdrawn contract':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nThe DESIGN.md contract is no longer current.';},
|
||||
'wrong issue header':(q:NativePlanQuestionCall['questions'][number])=>{q.header='Issue 2';},
|
||||
'setup header':(q:NativePlanQuestionCall['questions'][number])=>{q.header='Focus';},
|
||||
'foreign option IDs':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.label=q.options[0]!.label.replace('1A','2A');},
|
||||
'missing primary styling':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description='✅ Matches DESIGN.md exactly. ❌ Work required.';},
|
||||
'foreign primary styling':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('Save #','Publish #');},
|
||||
'foreign peer styling':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel, Download');},
|
||||
'missing peer':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel');},
|
||||
'duplicate peer':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Reset, Export');},
|
||||
'primary also a ghost':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel, Save');},
|
||||
'quoted remedy':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description='"'+q.options[0]!.description+'"';},
|
||||
'conditional remedy':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description+=' If approved, apply these styles.';},
|
||||
'withdrawn remedy':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description+=' This option is withdrawn.';},
|
||||
'withdrawn quoted remedy status':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description+=' This option is "withdrawn".';},
|
||||
'cancelled styling':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description+=' Do not apply these styles.';},
|
||||
'missing retained violation':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description='This closes the hierarchy gap completely.';},
|
||||
'quoted retained violation':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description='"'+q.options[2]!.description+'"';},
|
||||
'conditional retained violation':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description+=' If approved, leave the gap open.';},
|
||||
'withdrawn retained violation':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description+=' This option is withdrawn.';},
|
||||
'foreign retained violation':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description+=' This deferral belongs to another project.';},
|
||||
}))test(`cf74 current style rejects ${name}`,()=>{expect(isDesignCountFirstReview(fp(changeCf74(change)))).toBe(false);});
|
||||
test('cf74 current styling still requires its own complete native answer and identities',()=>{
|
||||
for(const change of [
|
||||
(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},
|
||||
(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},
|
||||
(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.sessionId='';},
|
||||
(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!));},
|
||||
]){const call=cf74Calls()[0]!;change(call);expect(isDesignCountFirstReview(fp(call))).toBe(false);}
|
||||
const f=fp(cf74Calls()[0]!);expect(isDesignCountFirstReview({...f,signature:'foreign:call'})).toBe(false);
|
||||
});
|
||||
test('cf74 first eight-review failure remains eight with no threshold or TODO exclusion change',()=>{
|
||||
let started=false;const counts={setup:0,review:0,administrative:0};
|
||||
for(const call of cf74.firstFailureCalls as NativePlanQuestionCall[]){const p=planCountQuestionPhase(fp(call),started,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff);counts[p.administrative?'administrative':p.preReview?'setup':'review']++;started=p.reviewStarted;}
|
||||
expect(counts).toEqual({setup:2,review:8,administrative:0});expect(counts.review).toBeGreaterThan(7);
|
||||
});
|
||||
for(const heading of ['## Completion report','### Completion summary','## Completion'])for(const field of ['Plan written:','Plan saved:','Plan written to'])
|
||||
test(`cf74 complete typed delivery: ${heading}/${field}`,()=>{
|
||||
const f=cf74Completion();try{f.final.text=f.final.text.replace('## Completion report',heading).replace('Plan written:',field);expect(f.check()).toBe(true);}finally{f.cleanup();}
|
||||
@@ -393,43 +135,3 @@ test('cf74 typed envelope cannot bypass fresh own Design report and native chron
|
||||
const target=path.join(f.dir,'other.md');fs.writeFileSync(target,cf74.report);fs.symlinkSync(target,f.file);expect(f.check()).toBe(false);
|
||||
}finally{f.cleanup();}
|
||||
});
|
||||
|
||||
for(const [name,change] of Object.entries({
|
||||
'mismatched source color':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('#1d4ed8','#aa0000');},
|
||||
'mismatched source foreground':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description=q.options[0]!.description!.replace('white text','black text');},
|
||||
'unrelated additional approval':(q:NativePlanQuestionCall['questions'][number])=>{q.options[0]!.description+=' Also approve deployment.';},
|
||||
'unrelated extra question action':(q:NativePlanQuestionCall['questions'][number])=>{q.question+='\nThen delete the audit log.';},
|
||||
'retained option actually fixes':(q:NativePlanQuestionCall['questions'][number])=>{q.options[2]!.description+=' This option resolves the hierarchy gap.';},
|
||||
}))test(`cf74 complete role transfer rejects ${name}`,()=>expect(isDesignCountFirstReview(fp(changeCf74(change)))).toBe(false));
|
||||
test('cf74 concrete token identity is source-owned rather than fixed to one palette',()=>{
|
||||
const call=changeCf74(q=>{q.question=q.question.replaceAll('#1d4ed8','#234567').replaceAll('white text','black text');q.options.forEach(o=>{o.description=o.description?.replaceAll('#1d4ed8','#234567').replaceAll('white text','black text');});});
|
||||
expect(isDesignCountFirstReview(fp(call))).toBe(true);
|
||||
});
|
||||
|
||||
for(const field of ['question','option'] as const)for(const action of ['Also implement a webhook handler.','Then replace the database.'])
|
||||
test(`cf74 peer extra work rejects ${field}/${action}`,()=>{
|
||||
const call=changeCf74(q=>{if(field==='question')q.question+='\n'+action;else q.options[0]!.description+=' '+action;});
|
||||
expect(isDesignCountFirstReview(fp(call))).toBe(false);
|
||||
});
|
||||
|
||||
for(const field of ['question','option','opposed'] as const)for(const [prefix,work]of [
|
||||
['Also ','build a webhook handler'],['Then ','migrate the database'],['Please ','configure a new service'],
|
||||
['Next ','install the worker'],['Now ','rewrite the API'],['First ','create an audit endpoint'],
|
||||
['and ','add a billing screen'],['but ','remove the login check'],['while ','launch a second deployment'],
|
||||
] as const)test(`cf74 imperative work class rejects ${field}/${prefix}${work}`,()=>{
|
||||
const call=changeCf74(q=>{const action=prefix+work+'.';if(field==='question')q.question+='\n'+action;else q.options[field==='option'?0:2]!.description+=' '+action;});
|
||||
expect(isDesignCountFirstReview(fp(call))).toBe(false);
|
||||
});
|
||||
for(const field of ['question','option'] as const)for(const text of [
|
||||
'The implementation may replace an existing button variant.',
|
||||
'Replacing the style makes the primary action clearer.',
|
||||
'Do not implement a webhook handler.',
|
||||
'No database replacement belongs to this review.',
|
||||
'Historical note: "Also implement a webhook handler."',
|
||||
"Historical note: 'Then replace the database.'",
|
||||
'Previous example: `Also configure a worker.`',
|
||||
'\n> Also implement a webhook handler.',
|
||||
] as const)test(`cf74 imperative guard preserves explanation/history ${field}/${text}`,()=>{
|
||||
const call=changeCf74(q=>{if(field==='question')q.question+='\n'+text;else q.options[0]!.description+=' '+text;});
|
||||
expect(isDesignCountFirstReview(fp(call))).toBe(true);
|
||||
});
|
||||
@@ -1,323 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-count-native-issue-fields.json';
|
||||
import { designStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
const accepts = (call: NativePlanQuestionCall) => isDesignCountFirstReview(fingerprint(call));
|
||||
type Question = NativePlanQuestionCall['questions'][number];
|
||||
function changed(index: number, edit: (question: Question) => void) {
|
||||
const call = calls()[index]!, question = call.questions[0]!;
|
||||
edit(question);
|
||||
call.answers = { [question.question]: question.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('native numbered design gaps with complete decision fields', () => {
|
||||
test('the exact first four findings each establish review independently', () => {
|
||||
for (const index of [1, 2, 3, 4]) expect(accepts(calls()[index]!)).toBe(true);
|
||||
});
|
||||
|
||||
test('all eight public calls retain one setup and seven review decisions without mutation', () => {
|
||||
const input = calls(), before = JSON.stringify(input);
|
||||
let started = false;
|
||||
const phases = input.map(call => {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(input).toHaveLength(8);
|
||||
expect(captured.assistantMessages).toHaveLength(3);
|
||||
expect(phases.map(phase => phase.preReview)).toEqual([true, false, false, false, false, false, false, false]);
|
||||
expect(phases.filter(phase => phase.administrative)).toHaveLength(0);
|
||||
expect(JSON.stringify(input)).toBe(before);
|
||||
});
|
||||
|
||||
test('a finding keeps its identity across descriptive headers, ordinals and offered answers', () => {
|
||||
for (const index of [1, 2, 3, 4]) {
|
||||
const call = changed(index, question => {
|
||||
question.header = 'Current design requirement';
|
||||
question.question = question.question.replace(/Issue [1-9]\d*/, 'Issue 17')
|
||||
.replace(/\bG[1-9]\d*\b/g, 'G29').replace(/\b[1-9]\d*([ABC])\b/g, '17$1');
|
||||
question.options = question.options.map(option => ({
|
||||
label: option.label.replace(/^[1-9]\d*/, '17'),
|
||||
description: option.description?.replace(/\bG[1-9]\d*\b/g, 'G29'),
|
||||
})).reverse();
|
||||
});
|
||||
for (const option of call.questions[0]!.options) {
|
||||
call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(accepts(call)).toBe(true);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('decision fields tolerate prose layout and equivalent current defect descriptions', () => {
|
||||
const descriptions = [
|
||||
'The header buttons currently share the same visual weight; the primary action is not distinguishable.',
|
||||
'The Save request currently gives no visible feedback while it is pending; users try again.',
|
||||
'The form labels currently mix 14px, 16px and 18px with no consistent role; the hierarchy is unclear.',
|
||||
'The form currently mixes 24px, 32px and 16px section gaps without a spacing rule.',
|
||||
];
|
||||
for (const [offset, assessment] of descriptions.entries()) {
|
||||
const call = changed(offset + 1, question => {
|
||||
question.question = question.question.replace(/^ELI10: .+$/m, `ELI10: ${assessment} DESIGN.md specifies the existing treatment.`)
|
||||
.replace(/\n(?=(?:Stakes if we pick wrong|Recommendation|Completeness|Net):)/g, '\n\n');
|
||||
});
|
||||
expect(accepts(call)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('bare gap IDs, scores, or setup menus cannot replace the current design defect', () => {
|
||||
for (const index of [1, 2, 3, 4]) for (const edit of [
|
||||
(q: Question) => { q.header = 'Focus'; },
|
||||
(q: Question) => { q.header = 'Issue 99'; },
|
||||
(q: Question) => { q.question = q.question.replace(/^D\d+[^\n]+/, 'D2 — Issue 1 (G1): Are we ready to review the design?'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^ELI10: .+$/m, 'ELI10: G1 is a design finding with a score of 6/10.'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^ELI10: .+$/m, 'ELI10: The form already follows every design requirement and has no current defect.'); },
|
||||
(q: Question) => { q.options = [{ label: `${index}A Start review`, description: 'Continue the review.' }, { label: `${index}B Wait`, description: 'Keep the gap open.' }]; },
|
||||
]) expect(accepts(changed(index, edit))).toBe(false);
|
||||
});
|
||||
|
||||
test('source, quoted, conditional, withdrawn and duplicate evidence does not establish review', () => {
|
||||
for (const index of [1, 2, 3, 4]) for (const edit of [
|
||||
(q: Question) => { q.question = `Historical example:\n${q.question}`; },
|
||||
(q: Question) => { q.question = `\`\`\`\n${q.question}\n\`\`\``; },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: ', 'ELI10: If approved, '); },
|
||||
(q: Question) => { q.question = q.question.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'); },
|
||||
(q: Question) => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: Earlier review example: '); },
|
||||
(q: Question) => { q.question += '\nELI10: No current defect exists.'; },
|
||||
(q: Question) => { q.question += '\nCorrection: this finding is withdrawn.'; },
|
||||
(q: Question) => { q.question += '\nCorrection: this gap is already resolved.'; },
|
||||
(q: Question) => { q.options[0]!.description = `If approved later, ${q.options[0]!.description}`; },
|
||||
(q: Question) => { q.options[0]!.description = `> ${q.options[0]!.description}`; },
|
||||
(q: Question) => { q.options[0]!.description += ' This amendment is withdrawn.'; },
|
||||
(q: Question) => { q.options[2]!.description += ' This gap is now closed.'; },
|
||||
]) expect(accepts(changed(index, edit))).toBe(false);
|
||||
});
|
||||
|
||||
test('the offered remedy and retained gap must belong to this decision', () => {
|
||||
for (const index of [1, 2, 3, 4]) for (const edit of [
|
||||
(q: Question) => { q.options[0]!.label = '99A A different issue'; },
|
||||
(q: Question) => { q.options[0]!.description = 'Record a finding after the next review.'; },
|
||||
(q: Question) => { q.options[2]!.description = q.options[2]!.description!.replace(/G\d+/, 'G999'); },
|
||||
(q: Question) => { q.options[2]!.description = 'The gap is resolved; nothing remains open.'; },
|
||||
(q: Question) => { q.options[2]!.label = `${index}C Choose the next workflow`; },
|
||||
(q: Question) => { q.question = q.question.replace('Recommendation:', 'Previous recommendation:'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^Recommendation: [1-9]\d*[A-Z]/m, 'Recommendation: 99A'); },
|
||||
]) expect(accepts(changed(index, edit))).toBe(false);
|
||||
});
|
||||
|
||||
test('owned quoted status scalars still withdraw a decision; quoted history does not', () => {
|
||||
for (const index of [1, 2, 3, 4]) for (const target of ['question', 'remedy', 'deferral']) {
|
||||
const append = (q: Question, text: string) => {
|
||||
if (target === 'question') q.question += text;
|
||||
else q.options[target === 'remedy' ? 0 : 2]!.description += text;
|
||||
};
|
||||
for (const [left, right] of [['"', '"'], ["'", "'"], ['“', '”'], ['‘', '’'], ['`', '`']]) {
|
||||
expect(accepts(changed(index, q => append(q, `\nThis finding is ${left}withdrawn${right}.`)))).toBe(false);
|
||||
}
|
||||
expect(accepts(changed(index, q => append(q, '\nPrior note: "This finding is withdrawn."')))).toBe(true);
|
||||
expect(accepts(changed(index, q => append(q, '\n> This finding is withdrawn.')))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('only a completed, successful native call with its actual selected answer can start review', () => {
|
||||
const changes = [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.sessionId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.toolUseId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'not offered' }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
];
|
||||
for (const index of [1, 2, 3, 4]) {
|
||||
for (const change of changes) { const call = calls()[index]!; change(call); expect(accepts(call)).toBe(false); }
|
||||
for (const change of [
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.signature = 'other:call'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeQuestionIndex = 1; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.options.reverse(); },
|
||||
]) { const fp = fingerprint(calls()[index]!); change(fp); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('dacc95ea current Issue decisions without a G or Pass label', () => {
|
||||
const actual = () => structuredClone(captured.dacc95eaFirstAttempt.calls) as NativePlanQuestionCall[];
|
||||
for (const index of [2, 3, 4, 5, 6]) test(`actual retained Issue ${index - 1} independently starts review`, () => {
|
||||
expect(accepts(actual()[index]!)).toBe(true);
|
||||
});
|
||||
test('actual eight-call phase replay preserves two setup calls and six later decisions', () => {
|
||||
let started = false;
|
||||
const input = actual(), before = JSON.stringify(input);
|
||||
const phases = input.map(call => {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(phases.map(phase => phase.preReview)).toEqual([true, true, false, false, false, false, false, false]);
|
||||
expect(phases.filter(phase => phase.administrative)).toHaveLength(0);
|
||||
expect(JSON.stringify(input)).toBe(before);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('dacc95ea numbered Finding decisions with an owned detailed comparison', () => {
|
||||
const actual = () => structuredClone(captured.dacc95eaRetry.calls) as NativePlanQuestionCall[];
|
||||
for (const index of [3, 4, 5, 6, 7]) test(`actual retained Finding call ${index - 2} independently starts review`, () => {
|
||||
expect(accepts(actual()[index]!)).toBe(true);
|
||||
});
|
||||
test('nine retained retry calls preserve three setup calls and six later decisions', () => {
|
||||
let started = false;
|
||||
const input = actual(), before = JSON.stringify(input);
|
||||
const phases = input.map(call => {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(phases.map(phase => phase.preReview)).toEqual([true, true, true, false, false, false, false, false, false]);
|
||||
expect(JSON.stringify(input)).toBe(before);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('current native design decision boundaries', () => {
|
||||
const specimens = () => [
|
||||
...structuredClone(captured.dacc95eaFirstAttempt.calls).slice(2, 7),
|
||||
...structuredClone(captured.dacc95eaRetry.calls).slice(3, 8),
|
||||
] as NativePlanQuestionCall[];
|
||||
const edit = (input: NativePlanQuestionCall, mutate: (q: Question) => void) => {
|
||||
const call = structuredClone(input), q = call.questions[0]!;
|
||||
mutate(q); call.answers = { [q.question]: q.options[0]!.label }; return call;
|
||||
};
|
||||
for (const [name, mutate] of Object.entries({
|
||||
'whole quoted brief': (q: Question) => { q.question = q.question.split('\n').map(line => '> ' + line).join('\n'); },
|
||||
'whole fenced brief': (q: Question) => { q.question = '\x60\x60\x60md\n' + q.question + '\n\x60\x60\x60'; },
|
||||
'historical preface': (q: Question) => { q.question = 'Historical example:\n' + q.question; },
|
||||
'foreign source': (q: Question) => { q.question = q.question.replaceAll('PLAN.md', 'OTHER.md'); },
|
||||
'quoted source': (q: Question) => { q.question = q.question.replaceAll('PLAN.md', '"PLAN.md"'); },
|
||||
'conditional assessment': (q: Question) => { q.question = q.question.replace('ELI10: ', 'ELI10: If approved later, '); },
|
||||
'quoted assessment': (q: Question) => { q.question = q.question.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'); },
|
||||
'duplicate assessment': (q: Question) => { q.question += '\nELI10: Another assessment.'; },
|
||||
'explicitly closed gap': (q: Question) => { q.question += '\nThis finding is now resolved.'; },
|
||||
'withdrawn current scalar': (q: Question) => { q.question += '\nThis finding is "withdrawn".'; },
|
||||
'setup header': (q: Question) => { q.header = 'Focus'; },
|
||||
'wrong header identity': (q: Question) => { q.header = 'Issue 99'; },
|
||||
'wrong option identity': (q: Question) => { q.options[0]!.label = q.options[0]!.label.replace(/^\d+/, '99'); },
|
||||
'foreign recommendation': (q: Question) => { q.question = q.question.replace(/^Recommendation: \d+[A-Z]/m, 'Recommendation: 99A'); },
|
||||
'withdrawn remedy': (q: Question) => { q.options[0]!.description += '\nThis amendment is withdrawn.'; },
|
||||
'closed deferral': (q: Question) => { q.options.at(-1)!.description += '\nThis gap is now closed.'; },
|
||||
})) test('both captured classes reject ' + name, () => {
|
||||
for (const call of specimens()) expect(accepts(edit(call, mutate))).toBe(false);
|
||||
});
|
||||
test('every offered answer and recommendation-first ordering retains the same owned decision', () => {
|
||||
for (const input of specimens()) {
|
||||
const call = structuredClone(input), q = call.questions[0]!;
|
||||
q.options.reverse();
|
||||
for (const option of q.options) { call.answers = { [q.question]: option.label }; expect(accepts(call)).toBe(true); }
|
||||
}
|
||||
});
|
||||
for (const [name, mutate] of Object.entries({
|
||||
unanswered: (c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
failed: (c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
'pending index': (c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
'missing timestamp': (c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
'unoffered answer': (c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Recommendation A' }; },
|
||||
'multiple questions': (c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
})) test('both captured classes reject native ' + name, () => {
|
||||
for (const call of specimens()) { mutate(call); expect(accepts(call)).toBe(false); }
|
||||
});
|
||||
test('the expanded comparison must keep complete current option ownership', () => {
|
||||
const original = specimens()[5]!;
|
||||
for (const mutate of [
|
||||
(q: Question) => { q.question = q.question.replace(/\nPros \/ cons:[\s\S]*?\nNet:/, '\nNet:'); },
|
||||
(q: Question) => { q.question = q.question.replace(/(\nPros \/ cons:\n)([\s\S]*?)(\nNet:)/, '$1\x60\x60\x60md\n$2\n\x60\x60\x60$3'); },
|
||||
(q: Question) => { q.question = q.question.replace(/(\nPros \/ cons:\n)/, '$1Historical example:\n'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^1A\)/m, '99A)'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^1B\)/m, '1A)'); },
|
||||
(q: Question) => { q.question = q.question.replace(/\n1C\)[\s\S]*?\nNet:/, '\nNet:'); },
|
||||
]) expect(accepts(edit(original, mutate))).toBe(false);
|
||||
});
|
||||
test('only the bound native decision status can withdraw its current finding', () => {
|
||||
for (const original of specimens()) {
|
||||
const title = original.questions[0]!.question.split('\n')[0]!;
|
||||
const owner = /^D[1-9]\d*/.exec(title)?.[0] ?? /Finding [1-9]\d*/.exec(title)![0];
|
||||
for (const status of ['withdrawn', '"withdrawn"', '\x60withdrawn\x60']) {
|
||||
expect(accepts(edit(original, q => { q.question += `\n${owner} is ${status}.`; }))).toBe(false);
|
||||
}
|
||||
expect(accepts(edit(original, q => { q.question += `\nPrior note: "${owner} is withdrawn."`; }))).toBe(true);
|
||||
expect(accepts(edit(original, q => { q.question += `\n> ${owner} is withdrawn.`; }))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('a conforming contrast ratio cannot borrow a low-contrast classification', () => {
|
||||
const original = specimens()[2]!;
|
||||
expect(accepts(edit(original, q => { q.question = q.question.replaceAll('3:1', '4.5:1'); }))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { classifyPlanCountFrame, hasNativePlanTerminal, isQuestionlessNativePlanExit, assertReviewReportAtBottom } from './helpers/claude-pty-runner';
|
||||
import type { PlanCountTranscript } from './helpers/plan-count-transcript';
|
||||
|
||||
test('full first attempt reaches owned completion and passes every unchanged paid callback assertion', () => {
|
||||
const actual = captured.dacc95eaFirstAttempt, ending = actual.completion;
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'design-dacc-completion-'));
|
||||
const file = path.join(dir, path.basename(ending.provenance.file));
|
||||
const transcript: PlanCountTranscript = { status: 'ready', calls: structuredClone(actual.calls) as NativePlanQuestionCall[],
|
||||
assistantMessages: structuredClone(ending.assistantMessages), planReadyRequests: structuredClone(ending.planReadyRequests) };
|
||||
const startedAt = Math.min(...transcript.calls.map(call => Date.parse(call.answeredAt!))) - 1_000;
|
||||
const modifiedAt = Date.parse(ending.provenance.mutations.at(-1)!.at) / 1_000;
|
||||
const write = (body = ending.report) => { fs.writeFileSync(file, body); fs.utimesSync(file, modifiedAt, modifiedAt); };
|
||||
let started = false; const counts = { step0: 0, review: 0, administrative: 0 }, nonReview = new Set<string>();
|
||||
const fingerprints = transcript.calls.map(call => {
|
||||
const fp = fingerprint(call), phase = planCountQuestionPhase(fp, started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff);
|
||||
started = phase.reviewStarted;
|
||||
counts[phase.administrative ? 'administrative' : phase.preReview ? 'step0' : 'review']++;
|
||||
if (phase.preReview || phase.administrative) nonReview.add(fp.signature);
|
||||
return { ...fp, preReview: phase.preReview };
|
||||
});
|
||||
const caller = fs.readFileSync(path.join(import.meta.dir, 'skill-e2e-plan-design-finding-count.test.ts'), 'utf8');
|
||||
const constants = /^const N = .+;\nconst FLOOR = .+;\nconst CEILING = .+;/m.exec(caller)![0];
|
||||
// Bind the actual callback's complete validation block, without importing
|
||||
// the paid registration or changing its assertions, prompt or work limits.
|
||||
const start = caller.indexOf(" if (!['plan_ready', 'completion_summary', 'ceiling_reached'].includes(obs.outcome))");
|
||||
const end = caller.indexOf('\n } finally {', start);
|
||||
expect(start).toBeGreaterThan(0); expect(end).toBeGreaterThan(start);
|
||||
const validate = new Function('fs', 'planPath', 'obs', 'assertReviewReportAtBottom',
|
||||
new Bun.Transpiler({ loader: 'ts' }).transformSync(constants + '\n' + caller.slice(start, end)));
|
||||
try {
|
||||
write();
|
||||
expect(createHash('sha256').update(ending.report).digest('hex')).toBe(ending.reportSha256);
|
||||
expect(counts).toEqual({ step0: 2, review: 6, administrative: 0 });
|
||||
const frame = classifyPlanCountFrame(ending.screen);
|
||||
expect(frame).toBe('plan_ready');
|
||||
expect(hasNativePlanTerminal(transcript, file, startedAt, 'plan_ready')).toBe(true);
|
||||
expect(isQuestionlessNativePlanExit(transcript, file, startedAt, ending.screen, nonReview)).toBe(false);
|
||||
expect(assertReviewReportAtBottom(ending.report).ok).toBe(true);
|
||||
const replayed = { outcome: frame, step0Count: counts.step0, reviewCount: counts.review, fingerprints, elapsedMs: 0, evidence: ending.screen };
|
||||
expect(() => validate(fs, file, replayed, assertReviewReportAtBottom)).not.toThrow();
|
||||
for (const [delta, error] of [
|
||||
[{ outcome: 'no_review_questions' }, 'finding-count FAILED'],
|
||||
[{ reviewCount: 3 }, 'BAND FAIL (below floor)'],
|
||||
[{ reviewCount: 8 }, 'BAND FAIL (above ceiling)'],
|
||||
] as const) expect(() => validate(fs, file, { ...replayed, ...delta }, assertReviewReportAtBottom)).toThrow(error);
|
||||
write(ending.report + '\n## Work after report\n');
|
||||
expect(() => validate(fs, file, replayed, assertReviewReportAtBottom)).toThrow('D19 FAIL');
|
||||
write(); fs.rmSync(file);
|
||||
expect(() => validate(fs, file, replayed, assertReviewReportAtBottom)).toThrow('D19 FAIL');
|
||||
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
|
||||
});
|
||||
@@ -1,135 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { capturePlanCountQuestion, nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import { pickDesignCountOutsideVoices } from './helpers/design-count-outside';
|
||||
import { isDesignCountFirstReview } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
const packet: NativePlanQuestionCall = {
|
||||
"sessionId": "52868766-97d8-4406-8230-4f263be36546",
|
||||
"toolUseId": "toolu_013ghrvY7M8hMrMHox3tnJLn",
|
||||
"questions": [
|
||||
{
|
||||
"question": "D3 (Step 0D) — I've rated this plan 5/10 on design completeness. The three biggest gaps are: (1) the 5 identified implementation gaps describe the problem but not the solution, (2) no explicit state coverage table, (3) no user journey emotional arc. I'll skip mockups and review all 7 dimensions as you requested. Any specific areas to prioritize, or cover all 7 equally? <gstack-qid:plan-design-focus>",
|
||||
"header": "Focus areas",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Cover all 7 equally (Recommended)",
|
||||
"description": "Standard review: all 7 design dimensions get full treatment. Takes longer but produces a complete plan."
|
||||
},
|
||||
{
|
||||
"label": "Focus on the 5 identified gaps first",
|
||||
"description": "Prioritize Pass 5 (Design System Alignment) to close the gap descriptions into actionable specs, then cover remaining passes more quickly."
|
||||
},
|
||||
{
|
||||
"label": "Prioritize accessibility and states",
|
||||
"description": "Focus on Pass 2 (Interaction States) and Pass 6 (Responsive/A11y), since the form has sensitive UX requirements (ARIA, contrast, keyboard)."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"question": "D4 — Want outside design voices before the detailed review? Codex evaluates against OpenAI's design hard rules + litmus checks; a Claude subagent does an independent completeness review. (Requires Codex CLI to be installed.) <gstack-qid:outside-voices-design>",
|
||||
"header": "Outside voices",
|
||||
"multiSelect": false,
|
||||
"options": [
|
||||
{
|
||||
"label": "Yes, run outside design voices",
|
||||
"description": "Launches Codex design critique + Claude subagent completeness review in parallel before the 7 passes. Adds 1–2 minutes."
|
||||
},
|
||||
{
|
||||
"label": "No, proceed without (Recommended)",
|
||||
"description": "Skip outside voices and go straight to the 7 review passes. Faster; sufficient for most plans."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"answered": false,
|
||||
"failed": false
|
||||
};
|
||||
|
||||
function screen(index: number, call = packet) {
|
||||
const q = call.questions[index]!;
|
||||
return '← ☐ Focus areas ☐ Outside voices ✔ Submit →\n│ ' + q.question + '\n' +
|
||||
q.options.map((option, i) => (i === 0 ? '❯' : '') + `${i + 1}. ${option.label}`).join('\n') +
|
||||
'\nEnter to select · Tab/Arrow keys to navigate · Esc to cancel\n';
|
||||
}
|
||||
|
||||
describe('Design count fixture outside-review choice', () => {
|
||||
test('the captured focus tab stays unchanged and only its outside-review tab declines', () => {
|
||||
const seen = new Set<string>();
|
||||
const focus = capturePlanCountQuestion(screen(0), seen, 0, true, packet)!;
|
||||
expect(pickDesignCountOutsideVoices(focus, focus)).toBeNull();
|
||||
const outside = capturePlanCountQuestion(screen(1), seen, 1, true, packet)!;
|
||||
expect(pickDesignCountOutsideVoices(outside, outside)).toBe(2);
|
||||
expect(capturePlanCountQuestion(screen(1), seen, 2, true, packet)).toBeNull();
|
||||
expect(pickDesignCountOutsideVoices(nativePlanCallFingerprint(packet, 0, true), focus)).toBeNull();
|
||||
});
|
||||
|
||||
test('the current opt-in question remains recognizable when native metadata arrives after the answer', () => {
|
||||
const fp = capturePlanCountQuestion(screen(1), new Set(), 0, true)!;
|
||||
expect(fp.nativeCall).toBeUndefined();
|
||||
expect(pickDesignCountOutsideVoices(fp, fp), fp.promptSnippet).toBe(2);
|
||||
const focus = capturePlanCountQuestion(screen(0), new Set(), 0, true)!;
|
||||
expect(pickDesignCountOutsideVoices(focus, focus)).toBeNull();
|
||||
});
|
||||
|
||||
test('single questions and reversed choices still select only the explicit No action', () => {
|
||||
for (const reverse of [false, true]) {
|
||||
const call = structuredClone(packet);
|
||||
call.questions = [call.questions[1]!];
|
||||
if (reverse) call.questions[0]!.options.reverse();
|
||||
const fp = nativePlanCallFingerprint(call, 0, true);
|
||||
expect(pickDesignCountOutsideVoices(fp, fp)).toBe(reverse ? 1 : 2);
|
||||
}
|
||||
});
|
||||
|
||||
test('pending packet metadata without the matching active question cannot steer a choice', () => {
|
||||
const fp = nativePlanCallFingerprint(packet, 0, true);
|
||||
expect(pickDesignCountOutsideVoices(fp, fp)).toBeNull();
|
||||
const outside = capturePlanCountQuestion(screen(1), new Set(), 0, true, packet)!;
|
||||
expect(pickDesignCountOutsideVoices(outside, { ...outside, signature: 'unrelated' })).toBeNull();
|
||||
for (const mutate of [
|
||||
(call: NativePlanQuestionCall) => { call.answered = true; },
|
||||
(call: NativePlanQuestionCall) => { call.failed = true; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[1]!.multiSelect = true; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[1]!.question = 'Should the product ask customers to use outside design voices?'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[1]!.question = call.questions[1]!.question.replace('outside-voices-design', 'design-review-finding'); },
|
||||
(call: NativePlanQuestionCall) => { call.questions[1]!.options[1]!.label = 'No, leave the design defect unfixed'; },
|
||||
(call: NativePlanQuestionCall) => { call.questions[1]!.options.push({ label: 'Change the design now' }); },
|
||||
]) {
|
||||
const call = structuredClone(packet);
|
||||
mutate(call);
|
||||
expect(pickDesignCountOutsideVoices(outside, { ...outside, nativeCall: call })).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('the binary outside-voices variant still selects only its explicit opt-out', () => {
|
||||
for (const id of ['outside-voices-design', 'plan-design-review-outside-voices']) {
|
||||
const call = structuredClone(packet);
|
||||
call.questions = [call.questions[1]!];
|
||||
const q = call.questions[0]!;
|
||||
q.question = `D3 — Want outside voices before the detailed review? <gstack-qid:${id}>`;
|
||||
q.options[0]!.label = 'Yes, run outside voices (recommended)';
|
||||
const native = nativePlanCallFingerprint(call, 0, true);
|
||||
expect(pickDesignCountOutsideVoices(native, native)).toBe(2);
|
||||
const visible = capturePlanCountQuestion(screen(0, call), new Set(), 0, true)!;
|
||||
expect(pickDesignCountOutsideVoices(visible, visible)).toBe(2);
|
||||
q.options[1]!.label = 'No, leave the design defect unfixed';
|
||||
const product = nativePlanCallFingerprint(call, 0, true);
|
||||
expect(pickDesignCountOutsideVoices(product, product)).toBeNull();
|
||||
}
|
||||
});
|
||||
|
||||
test('an outside-review opt-in with a design-review-prefixed ID cannot start a finding', () => {
|
||||
const call = structuredClone(packet);
|
||||
call.questions = [call.questions[1]!];
|
||||
const q = call.questions[0]!;
|
||||
q.question = 'D3 — Want outside voices before the detailed review?\n' +
|
||||
'Project/branch/task: main branch; design review of PLAN.md before the 7 passes. ' +
|
||||
'<gstack-qid:plan-design-review-outside-voices>';
|
||||
call.answered = true;
|
||||
call.unansweredQuestionIndices = [];
|
||||
call.answers = { [q.question]: q.options[0]!.label };
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(call, 0, true))).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,234 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-count-sep21-declared-first-call.json';
|
||||
import headerCaptured from './fixtures/design-count-sep21-header-first-call.json';
|
||||
import { nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview } from './helpers/design-count-review';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
|
||||
const current = () => structuredClone(captured.calls[0]) as NativePlanQuestionCall;
|
||||
type Question = NativePlanQuestionCall['questions'][number];
|
||||
const changed = (change: (q: Question) => void) => {
|
||||
const c = current(), q = c.questions[0]!;
|
||||
change(q);
|
||||
c.answers = {[q.question]: q.options[0]!.label};
|
||||
return nativePlanCallFingerprint(c, 0, true);
|
||||
};
|
||||
|
||||
describe('primary finding facts are independent of presentation', () => {
|
||||
test('the exact native declaration starts review for every offered answer', () => {
|
||||
const c = current(), q = c.questions[0]!;
|
||||
for (const option of q.options) {
|
||||
c.answers = {[q.question]: option.label};
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(c, 0, true))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('title adapters, owned role location, header and separators compose', () => {
|
||||
const titles = [
|
||||
'D2 — Issue 1: Save has no visual primacy in the header action group',
|
||||
'D2 — Issue 1: Make Save the visible primary action',
|
||||
'D2 — Issue 1: How should Save be distinguished from Reset, Cancel, and Export in the header?',
|
||||
];
|
||||
for (const title of titles) for (const header of ['Issue 1', 'Issue 1: Save', 'Issue 1 Save', 'Save primary']) {
|
||||
for (const role of ['label', 'body']) for (const separator of [', ', '; ', '. ']) {
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.header = header;
|
||||
q.question = title + q.question.slice(q.question.indexOf('\n'));
|
||||
if (role === 'body') {
|
||||
q.options[0]!.label = '1A — Apply DESIGN.md token (recommended)';
|
||||
q.options[0]!.description = q.options[0]!.description?.replace('Save #', 'Save filled primary #');
|
||||
}
|
||||
q.options[0]!.description = q.options[0]!.description?.replace('white text, Reset/Cancel/Export', `white text${separator}Export, Reset, Cancel`);
|
||||
q.options.reverse();
|
||||
})), `${title}/${header}/${role}/${separator}`).toBe(true);
|
||||
}
|
||||
}
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.header = 'Issue 3 Publish';
|
||||
q.question = q.question.replace('Issue 1:', 'Issue 3:').replaceAll('Save', 'Publish').replace('Four buttons', '4 buttons');
|
||||
q.options = q.options.map(o => ({label: o.label.replace(/^1/, '3').replaceAll('Save', 'Publish').replace('four', '4'),
|
||||
description: o.description?.replaceAll('Save', 'Publish').replace('#1d4ed8 with white', '#ffee22 with black')}));
|
||||
}))).toBe(true);
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.question = q.question.replace('header action group\n', 'header action group.\n');
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('native identity, current ownership, counts, authority and substantive options remain required', () => {
|
||||
const changes: Array<(q: Question) => void> = [
|
||||
q => {q.header = 'Issue 2';},
|
||||
q => {q.header = 'Issue 1 Publish';},
|
||||
q => {q.question = q.question.replace('Save has no visual primacy', 'Choose the next reviewer');},
|
||||
q => {q.question = q.question.replace('ELI10:', '> ELI10:');},
|
||||
q => {q.question = q.question.replace('ELI10:', 'ELI10: If approved,');},
|
||||
q => {q.question = q.question.replace('Four buttons', 'Three buttons');},
|
||||
q => {q.options[0]!.label = q.options[0]!.label.replace('Save filled primary', 'Publish filled primary');},
|
||||
q => {q.options[0]!.label = '1A — Prepare the review';},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Save #', 'Publish #');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('#1d4ed8', 'blue');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('with white text', '');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Reset/Cancel/Export', 'Reset//Cancel');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Reset/Cancel/Export', 'Reset/Cancel/Save');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('neutral ghost', 'filled primary');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Matches DESIGN.md exactly: ', '');},
|
||||
q => {
|
||||
q.options[0]!.description = q.options[0]!.description?.replace('Matches DESIGN.md exactly: ', '');
|
||||
q.options[1]!.description += ' Matches DESIGN.md exactly.';
|
||||
},
|
||||
q => {q.options[0]!.description += ' ❌ These tokens do not match DESIGN.md.';},
|
||||
q => {q.options[1]!.label = q.options[1]!.label.replace('four', 'three');},
|
||||
q => {q.options[1]!.description = 'The primary action is clear; no gap remains.';},
|
||||
q => {q.options[1]!.description = '> ' + q.options[1]!.description;},
|
||||
q => {q.options[1]!.description = q.options[1]!.description?.replace('Violates DESIGN.md', 'Satisfies DESIGN.md');},
|
||||
q => {q.options[1]!.description += ' This gap is resolved.';},
|
||||
];
|
||||
for (const change of changes) expect(isDesignCountFirstReview(changed(change)), change.toString()).toBe(false);
|
||||
for (const owner of [-1, 0, 1]) for (const suffix of [
|
||||
' This finding is "withdrawn".', ' ❌ Issue 1 is closed.', ' Assuming approval, use this option.',
|
||||
' This finding applies only to another project.',
|
||||
]) expect(isDesignCountFirstReview(changed(q => {
|
||||
if (owner === -1) q.question += suffix;
|
||||
else q.options[owner]!.description += suffix;
|
||||
})), owner + suffix).toBe(false);
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.question += '\n"Issue 1 is closed." Issue 2 is closed.';
|
||||
q.options[0]!.description += ' "This amendment is withdrawn."';
|
||||
}))).toBe(true);
|
||||
for (const owner of [0, 1]) for (const suffix of [
|
||||
'. This option is withdrawn.', '. If approved, apply this option.', '. Do not apply these styles.',
|
||||
]) expect(isDesignCountFirstReview(changed(q => {
|
||||
q.options[owner]!.label += suffix;
|
||||
})), owner + suffix).toBe(false);
|
||||
for (const role of ['Export primary', 'Export filled primary', 'Save ghost']) {
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.options[0]!.label = q.options[0]!.label.replace('others ghost', `${role}, others ghost`);
|
||||
})), role).toBe(false);
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.options[0]!.description += ` ${role}.`;
|
||||
})), role + ' in description').toBe(false);
|
||||
}
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.options[0]!.label += '. These tokens do not match DESIGN.md.';
|
||||
}))).toBe(false);
|
||||
expect(isDesignCountFirstReview(changed(q => {
|
||||
q.options[0]!.label += '. "Export filled primary." "Save ghost." "These tokens do not match DESIGN.md."';
|
||||
q.options[0]!.description += ' "Export filled primary." "Save ghost." "These tokens do not match DESIGN.md."';
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('recognized invalid primary findings cannot fall through to a generic review marker', () => {
|
||||
const c = current(), q = c.questions[0]!;
|
||||
q.question += '\n<gstack-qid:plan-design-review-primary-action>';
|
||||
c.answers = {[q.question]: q.options[0]!.label};
|
||||
const fp = nativePlanCallFingerprint(c, 0, true);
|
||||
// The loose marker is deliberately visible even when the real public
|
||||
// question is too long for a short prompt projection.
|
||||
fp.promptSnippet = 'D2 — Issue 1 <gstack-qid:plan-design-review-primary-action>';
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
expect(isDesignCountFirstReview({...fp, signature: 'foreign:call'})).toBe(false);
|
||||
const multiple = structuredClone(fp);
|
||||
multiple.nativeCall!.questions.push(structuredClone(q));
|
||||
expect(isDesignCountFirstReview(multiple)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('primary facts with identity carried by the native header', () => {
|
||||
const altered = (change: (q: Question) => void = () => {}) => {
|
||||
const c = structuredClone(headerCaptured.calls[0]) as NativePlanQuestionCall;
|
||||
const q = c.questions[0]!;
|
||||
change(q);
|
||||
c.answers = {[q.question]: q.options[0]!.label};
|
||||
return nativePlanCallFingerprint(c, 0, true);
|
||||
};
|
||||
|
||||
test('the exact public question starts review for every answer', () => {
|
||||
const c = structuredClone(headerCaptured.calls[0]) as NativePlanQuestionCall;
|
||||
for (const option of c.questions[0]!.options) {
|
||||
c.answers = {[c.questions[0]!.question]: option.label};
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(c, 0, true))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('identity, actor-list, property separator and authority location vary independently', () => {
|
||||
for (const identity of ['header', 'title', 'both']) for (const list of ['Reset, Cancel, Export', 'Export/Reset/Cancel', 'Cancel, Export and Reset']) {
|
||||
for (const separator of [': ', ' = ', ' ']) for (const authority of ['label', 'body']) {
|
||||
expect(isDesignCountFirstReview(altered(q => {
|
||||
if (identity !== 'header') q.question = q.question.replace('D1 — Should', 'D1 — Issue 1: Should');
|
||||
if (identity === 'title') q.header = 'Issue 1';
|
||||
q.question = q.question.replace('Save, Reset, Cancel and Export', `Save, ${list}`);
|
||||
q.options[0]!.description = q.options[0]!.description?.replace('Save: ', `Save${separator}`)
|
||||
.replace('Reset, Cancel, Export: ', `${list}${separator}`);
|
||||
if (authority === 'body') {
|
||||
q.options[0]!.label = '1A) Filled primary';
|
||||
q.options[0]!.description = q.options[0]!.description?.replace('Uses the exact approved tokens;', 'Matches DESIGN.md exactly;');
|
||||
}
|
||||
q.options.reverse();
|
||||
})), `${identity}/${list}/${separator}/${authority}`).toBe(true);
|
||||
}
|
||||
}
|
||||
expect(isDesignCountFirstReview(altered(q => {
|
||||
q.header = 'Issue 7: Publish';
|
||||
q.question = q.question.replaceAll('Save', 'Publish');
|
||||
q.options = q.options.map(o => ({label: o.label.replace(/^1/, '7').replaceAll('Save', 'Publish'),
|
||||
description: o.description?.replaceAll('Save', 'Publish').replace('#1d4ed8 with white', '#eeeeff with black')}));
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('independent identity and fact fields cannot disagree or borrow evidence', () => {
|
||||
const mutations: Array<(q: Question) => void> = [
|
||||
q => {q.header = 'Issue 2: Save';},
|
||||
q => {q.header = 'Issue 1: Publish';},
|
||||
q => {q.question = q.question.replace('D1 — Should', 'D1 — Issue 2: Should');},
|
||||
q => {q.question = q.question.replace('D1 — Should', 'D1 — Issue 1: Should').replace('Should Save', 'Should Publish');},
|
||||
q => {q.header = 'Issue 1: Save/Publish';},
|
||||
q => {q.question = q.question.replace(/^D1[^\n]+/, 'D1 — Choose the next reviewer for Save primary action');},
|
||||
q => {q.question = q.question.replace(/^D1([^\n]+)/, 'D1 — Historical example:$1');},
|
||||
q => {q.question = q.question.replace(/^D1([^\n]+)/, 'D1 — If approved,$1');},
|
||||
q => {q.question = q.question.replace(/^D1([^\n]+)/, 'D1 — "$1"');},
|
||||
q => {q.question = q.question.replace('Save, Reset, Cancel and Export', 'Save, Reset, Reset and Export');},
|
||||
q => {q.question = q.question.replace('Save, Reset, Cancel and Export', 'Save, Reset and Export');},
|
||||
q => {q.question = q.question.replace('look identical', 'are three identical buttons');},
|
||||
q => {q.question = q.question.replace('Right now Save, Reset, Cancel and Export look identical.', '"Right now Save, Reset, Cancel and Export look identical."');},
|
||||
q => {q.options[0]!.label = '1A) Primary';},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Uses the exact approved tokens', 'Uses unapproved tokens');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Uses the exact approved tokens;', 'Uses the exact approved tokens is false;');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Uses the exact approved tokens;', 'Uses the exact approved tokens from another unrelated design system;');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Save: filled', 'Publish: filled');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Reset, Cancel, Export:', 'Reset, Save, Export:');},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Reset, Cancel, Export:', 'Reset//Export:');},
|
||||
q => {q.options[0]!.label = '1A) Primary'; q.options[1]!.label = '1B) DESIGN.md primary';},
|
||||
q => {q.options[0]!.description = q.options[0]!.description?.replace('Uses the exact approved tokens;', ''); q.options[1]!.description += ' Uses the exact approved tokens.';},
|
||||
q => {q.options[2]!.label = q.options[2]!.label.replace('four', 'three');},
|
||||
q => {q.options[2]!.description = q.options[2]!.description?.replace('Leaves a known DESIGN.md violation and no primary action', 'Resolves the DESIGN.md violation and makes the primary action clear');},
|
||||
];
|
||||
for (const change of mutations) expect(isDesignCountFirstReview(altered(change)), change.toString()).toBe(false);
|
||||
for (const field of ['label', 'description'] as const) for (const statement of [
|
||||
'This option is withdrawn.', 'If approved, apply this option.', 'Do not apply these styles.',
|
||||
'These tokens do not match DESIGN.md.', 'These tokens are not approved.',
|
||||
'Save: ghost.', 'Export: filled primary.',
|
||||
]) expect(isDesignCountFirstReview(altered(q => {q.options[0]![field] += ` ${statement}`;})), `${field}/${statement}`).toBe(false);
|
||||
for (const field of ['label', 'description'] as const) expect(isDesignCountFirstReview(altered(q => {
|
||||
q.options[0]![field] += ' "These tokens do not match DESIGN.md." "Save: ghost."';
|
||||
}))).toBe(true);
|
||||
expect(isDesignCountFirstReview(altered(q => {
|
||||
q.question = q.question.replace(/^D1[^\n]+/, 'D1 — Issue 1: How should Save be distinguished from Reset, Cancel, and Export?')
|
||||
.replace('Save, Reset, Cancel and Export look identical', 'Save, Reset, Cancel and Discard look identical');
|
||||
}))).toBe(false);
|
||||
});
|
||||
|
||||
test('header identity failures stay invalid in the presence of generic review markers', () => {
|
||||
for (const header of ['Issue 1: Save/Publish', 'Design', 'Issue 2: Save', 'Issue 01: Save']) {
|
||||
const fp = altered(q => {
|
||||
q.header = header;
|
||||
q.question += '\n<gstack-qid:plan-design-review-primary-action>';
|
||||
});
|
||||
fp.promptSnippet = 'D1 <gstack-qid:plan-design-review-primary-action>';
|
||||
expect(isDesignCountFirstReview(fp), header).toBe(false);
|
||||
}
|
||||
const quoted = altered(q => {
|
||||
q.question = q.question.replace(/^D1([^\n]+)/, 'D1 — "$1"') + '\n<gstack-qid:plan-design-review-primary-action>';
|
||||
});
|
||||
quoted.promptSnippet = 'D1 <gstack-qid:plan-design-review-primary-action>';
|
||||
expect(isDesignCountFirstReview(quoted)).toBe(false);
|
||||
});
|
||||
});
|
||||
File diff suppressed because it is too large.
Load diff
@@ -1,148 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { execFile } from 'node:child_process';
|
||||
import { promisify } from 'node:util';
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { generateDesignMockup } from '../scripts/resolvers/design';
|
||||
import { HOST_PATHS } from '../scripts/resolvers/types';
|
||||
|
||||
const ROOT = path.resolve(import.meta.dir, '..');
|
||||
|
||||
// The current count driver owns fixture creation; this control materializes
|
||||
// its exact inputs with that same helper and keeps the paid report/band gates.
|
||||
test.each(['success', 'below', 'above', 'missing-report', 'trailing-report', 'timeout', 'throw', 'native-error'])('native Design count registration: %s', async scenario => {
|
||||
const dir = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'design-count-fixture-')));
|
||||
const facts = path.join(dir, 'facts.json');
|
||||
const child = path.join(dir, 'caller.test.ts');
|
||||
try {
|
||||
fs.writeFileSync(child, `
|
||||
import {describe,expect,mock} from 'bun:test';
|
||||
import * as fs from 'node:fs';import * as path from 'node:path';
|
||||
import {execFileSync} from 'node:child_process';
|
||||
import * as runner from ${JSON.stringify(path.join(ROOT,'test/helpers/claude-pty-runner.ts'))};
|
||||
import {createPlanCountFixture} from ${JSON.stringify(path.join(ROOT,'test/helpers/plan-count-fixture.ts'))};
|
||||
const original={...runner},scenario=${JSON.stringify(scenario)};
|
||||
let calls=0;
|
||||
mock.module(${JSON.stringify(path.join(ROOT,'test/helpers/e2e-gate.ts'))},()=>({describeE2ETier:tier=>{expect(tier).toBe('periodic');return describe;}}));
|
||||
mock.module(${JSON.stringify(path.join(ROOT,'test/helpers/claude-pty-runner.ts'))},()=>({...original,
|
||||
runPlanSkillCounting:async opts=>{
|
||||
calls++;const target=opts.expectedPlanPath;
|
||||
fs.writeFileSync(${JSON.stringify(facts)},JSON.stringify({calls,target,validated:false}));
|
||||
expect(opts.cwd).toBeUndefined();
|
||||
expect(opts.followUpPrompt).toContain(target);
|
||||
expect(opts.followUpPrompt).toContain('Text-only review; skip mockups. Review all seven design dimensions.');
|
||||
expect(opts).toMatchObject({skillName:'plan-design-review',slashCommand:'/plan-design-review',reviewCountCeiling:8,
|
||||
timeoutMs:1500000,env:{QUESTION_TUNING:'false',EXPLAIN_LEVEL:'default'}});
|
||||
for(const key of ['isLastStep0AUQ','isFirstReviewAUQ','isSetupAUQ','isCompletionHandoffAUQ','isArtifactGenerationAUQ','pickAUQ'])expect(typeof opts[key]).toBe('function');
|
||||
for(const finding of ['same size, weight, and color as','24px in some places, 32px in others, and 16px',
|
||||
'approximately 3:1 (below WCAG AA)','14px, 16px, and 18px font sizes','2-5 seconds with no loading indicator'])expect(opts.followUpPrompt).toContain(finding);
|
||||
const fixture=createPlanCountFixture(opts.followUpPrompt,{files:opts.fixtureFiles});
|
||||
try {
|
||||
for(const [file,content] of Object.entries({'PLAN.md':opts.followUpPrompt,...opts.fixtureFiles}))
|
||||
expect(execFileSync('git',['show','HEAD:'+file],{cwd:fixture.cwd,encoding:'utf8',timeout:5000})).toBe(content);
|
||||
const design=fs.readFileSync(path.join(fixture.cwd,'DESIGN.md'),'utf8');
|
||||
for(const contract of ['640px maximum width','Save is the only filled primary action','Spacing uses an 8px base',
|
||||
'Typography has two roles','All text must meet WCAG AA contrast','pending-action pattern is an inline spinner'])expect(design).toContain(contract);
|
||||
} finally {fixture.cleanup();}
|
||||
fs.writeFileSync(${JSON.stringify(facts)},JSON.stringify({calls,target,validated:true}));
|
||||
if(scenario==='throw')throw new Error('controlled count observation failure');
|
||||
if(scenario!=='missing-report')fs.writeFileSync(target,'# Reviewed plan\\n\\n## GSTACK REVIEW REPORT\\nVERDICT: APPROVED\\n'+(scenario==='trailing-report'?'\\n## Unreviewed tail\\n':''));
|
||||
return {outcome:scenario==='timeout'?'timeout':scenario==='native-error'?'transcript_unavailable':'plan_ready',
|
||||
reviewCount:scenario==='below'?3:scenario==='above'?8:5,step0Count:2,elapsedMs:1000,fingerprints:[],evidence:'controlled native observation'};
|
||||
},
|
||||
}));
|
||||
await import(${JSON.stringify(path.join(ROOT,'test/skill-e2e-plan-design-finding-count.test.ts'))});
|
||||
`);
|
||||
const result = Bun.spawnSync([process.execPath,'test',child], {
|
||||
cwd:ROOT,timeout:10_000,env:{PATH:process.env.PATH??'',HOME:dir,TMPDIR:dir,TMP:dir,TEMP:dir,GIT_CONFIG_NOSYSTEM:'1',
|
||||
...(process.env.SystemRoot?{SystemRoot:process.env.SystemRoot}:{})},
|
||||
});
|
||||
const output=result.stdout.toString()+result.stderr.toString();
|
||||
expect(result.signalCode??null,output).toBeNull();
|
||||
expect(fs.existsSync(facts),output).toBe(true);
|
||||
const observed=JSON.parse(fs.readFileSync(facts,'utf8'));
|
||||
expect(observed.calls).toBe(1);expect(observed.validated,output).toBe(true);
|
||||
expect(fs.existsSync(path.dirname(observed.target))).toBe(false);
|
||||
expect(result.exitCode,output).toBe(scenario==='success'?0:1);
|
||||
const failure:Record<string,string>={below:'BAND FAIL (below floor)',above:'BAND FAIL (above ceiling)',
|
||||
'missing-report':'D19 FAIL: agent did not produce expected plan file','trailing-report':'trailing ## heading(s) after GSTACK REVIEW REPORT',
|
||||
timeout:'finding-count FAILED: outcome=timeout',throw:'controlled count observation failure','native-error':'finding-count FAILED: outcome=transcript_unavailable'};
|
||||
if(failure[scenario])expect(output).toContain(failure[scenario]);
|
||||
} finally {fs.rmSync(dir,{recursive:true,force:true});}
|
||||
},15_000);
|
||||
|
||||
// Execute the documented setup, not a duplicate implementation of its path choice.
|
||||
// The designer and provider are never invoked; mkdir is the observed side effect.
|
||||
for (const storage of ['configured', 'plugin', 'default']) test(`Design mockup setup honors ${storage} state storage`, async () => {
|
||||
const directory = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'design-output-root-')));
|
||||
const home = path.join(directory, 'operator home');
|
||||
const configured = path.join(directory, 'private state');
|
||||
const plugin = path.join(directory, 'plugin state');
|
||||
const cwd = path.join(directory, 'settings-fixture');
|
||||
try {
|
||||
fs.mkdirSync(path.join(home, '.claude/skills/gstack'), { recursive: true });
|
||||
fs.symlinkSync(path.join(ROOT, 'bin'), path.join(home, '.claude/skills/gstack/bin'), 'dir');
|
||||
fs.mkdirSync(cwd);
|
||||
await promisify(execFile)('git', ['init', '-q', cwd], { timeout: 5000 });
|
||||
const expected = storage === 'configured' ? configured : storage === 'plugin' ? plugin : path.join(home, '.gstack');
|
||||
const env = { PATH: process.env.PATH, HOME: home, USERPROFILE: '', TMPDIR: directory, TMP: directory,
|
||||
...(storage === 'configured' ? { GSTACK_HOME: configured, CLAUDE_PLUGIN_DATA: plugin, CLAUDE_PLUGIN_ROOT: '/plugins/gstack' } : {}),
|
||||
...(storage === 'plugin' ? { CLAUDE_PLUGIN_DATA: plugin, CLAUDE_PLUGIN_ROOT: '/plugins/gstack' } : {}) };
|
||||
const sources = [
|
||||
...['plan-design-review/SKILL.md.tmpl', 'design-shotgun/SKILL.md.tmpl',
|
||||
'design-consultation/sections/proposal-and-preview.md.tmpl', 'design-review/SKILL.md.tmpl']
|
||||
.map(file => fs.readFileSync(path.join(ROOT, file), 'utf8')),
|
||||
generateDesignMockup({ skillName: 'office-hours', tmplPath: '', host: 'claude', paths: HOST_PATHS.claude! }),
|
||||
];
|
||||
let mockupDirectory = '';
|
||||
for (const source of sources) {
|
||||
const block = [...source.matchAll(/```bash\n([\s\S]*?)\n```/g)]
|
||||
.find(match => /(?:_DESIGN_DIR|REPORT_DIR)=/.test(match[1]!))?.[1];
|
||||
expect(block).toBeDefined();
|
||||
const { stdout } = await promisify(execFile)('bash', ['-c', block!.replaceAll('<screen-name>', 'settings-page')], {
|
||||
cwd, env, timeout: 5000,
|
||||
});
|
||||
const output = stdout.match(/^(?:DESIGN_DIR|REPORT_DIR): (.+)$/m)?.[1];
|
||||
expect(output).toBeDefined();
|
||||
expect(path.resolve(path.dirname(output!))).toBe(path.resolve(expected, 'projects', 'settings-fixture', 'designs'));
|
||||
expect(fs.statSync(output!).isDirectory()).toBe(true);
|
||||
if (!mockupDirectory) mockupDirectory = output!;
|
||||
}
|
||||
// Execute the optional ideal-image command with a local stand-in for the
|
||||
// provider binary, observing its exact output argument and created image.
|
||||
const fakeDesign = path.join(home, '.claude/skills/gstack/design/dist/design');
|
||||
fs.mkdirSync(path.dirname(fakeDesign), { recursive: true });
|
||||
fs.writeFileSync(fakeDesign, '#!/bin/sh\nprintf \'%s\n\' "$@" > "$DESIGN_FAKE_ARGS"\nwhile [ "$1" != --output ]; do shift; done\nshift\nprintf fixture > "$1"\n');
|
||||
fs.chmodSync(fakeDesign, 0o755);
|
||||
const argsPath = path.join(directory, 'ideal-args.txt');
|
||||
const idealBlock = [...sources[0]!.matchAll(/```bash\n([\s\S]*?)\n```/g)]
|
||||
.find(match => match[1]!.includes('ideal-<dimension>.png'))?.[1];
|
||||
expect(idealBlock).toBeDefined();
|
||||
const idealResult = await promisify(execFile)('bash', ['-c', idealBlock!.replaceAll('<dimension>', 'hierarchy')], {
|
||||
cwd, env: { ...env, DESIGN_FAKE_ARGS: argsPath }, timeout: 5000,
|
||||
});
|
||||
const idealPath = idealResult.stdout.match(/^IDEAL_IMAGE: (.+)$/m)?.[1];
|
||||
expect(idealPath).toBeDefined();
|
||||
expect(path.resolve(path.dirname(path.dirname(idealPath!))))
|
||||
.toBe(path.resolve(expected, 'projects', 'settings-fixture', 'designs'));
|
||||
expect(fs.readFileSync(argsPath, 'utf8').trim().split('\n')).toEqual([
|
||||
'generate', '--brief', '<description of what 10/10 looks like for this dimension>', '--output', idealPath!,
|
||||
]);
|
||||
expect(fs.readFileSync(idealPath!, 'utf8')).toBe('fixture');
|
||||
for (const file of ['approved.json', 'variant-A.png', 'finalized.html']) fs.writeFileSync(path.join(mockupDirectory, file), 'fixture');
|
||||
const consumer = fs.readFileSync(path.join(ROOT, 'design-html/SKILL.md.tmpl'), 'utf8');
|
||||
let discovered = '';
|
||||
for (const match of consumer.matchAll(/```bash\n([\s\S]*?)\n```/g)) {
|
||||
if (!/_(?:APPROVED|VARIANTS|FINALIZED)=/.test(match[1]!)) continue;
|
||||
const { stdout } = await promisify(execFile)('bash', ['-c', match[1]!], { cwd, env, timeout: 5000 });
|
||||
discovered += stdout;
|
||||
}
|
||||
for (const [label, file] of [['APPROVED', 'approved.json'], ['VARIANTS', 'variant-A.png'], ['FINALIZED', 'finalized.html']]) {
|
||||
expect(discovered).toContain(`${label}: ${mockupDirectory}/${file}`);
|
||||
}
|
||||
// Existing slug-cache behavior is separate from the design artifact namespace.
|
||||
if (storage !== 'default') expect(fs.existsSync(path.join(home, '.gstack/projects'))).toBe(false);
|
||||
expect(fs.readdirSync(cwd)).toEqual(['.git']);
|
||||
} finally { fs.rmSync(directory, { recursive: true, force: true }); }
|
||||
});
|
||||
@@ -1,155 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-first-decision-af.json';
|
||||
import retryCaptured from './fixtures/design-first-decision-af-retry.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, nativePlanCallFingerprint, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
function call(): NativePlanQuestionCall {
|
||||
return structuredClone(captured.nativeCall) as NativePlanQuestionCall;
|
||||
}
|
||||
function answer(c: NativePlanQuestionCall, index = 0) {
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[index]!.label };
|
||||
return nativePlanCallFingerprint(c, 0, true);
|
||||
}
|
||||
|
||||
test('actual completed Make Save decision starts review before later findings', () => {
|
||||
expect(isDesignCountFirstReview(captured)).toBe(true);
|
||||
expect(planCountQuestionPhase(captured, false, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup)).toMatchObject({ preReview: false, reviewStarted: true });
|
||||
});
|
||||
|
||||
test('all offered decisions, including keeping the gap, are review decisions', () => {
|
||||
for (let index = 0; index < 3; index++) {
|
||||
expect(isDesignCountFirstReview(answer(call(), index))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('the actual retry starts at its first finding with a numbered control header', () => {
|
||||
expect(isDesignCountFirstReview(retryCaptured)).toBe(true);
|
||||
for (let i = 0; i < 3; i++) {
|
||||
const c = structuredClone(retryCaptured.nativeCall) as NativePlanQuestionCall;
|
||||
expect(isDesignCountFirstReview(answer(c, i))).toBe(true);
|
||||
}
|
||||
for (const header of ['Issue 2: Save', 'Issue 1: Reset', 'Issue 1.1: Save', 'Issue 1: Save\nMode']) {
|
||||
const c = structuredClone(retryCaptured.nativeCall) as NativePlanQuestionCall;
|
||||
c.questions[0]!.header = header;
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
for (const description of ['Save already complies. Record the completed review.',
|
||||
'Save becomes the single filled primary (#1d4ed8, white text); the report describes the buttons.']) {
|
||||
const c = structuredClone(retryCaptured.nativeCall) as NativePlanQuestionCall;
|
||||
c.questions[0]!.options[0]!.description = description;
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('number-letter option prefixes accept whitespace and existing punctuation', () => {
|
||||
for (const separator of [' ', ') ', '. ']) {
|
||||
const c = call();
|
||||
for (const option of c.questions[0]!.options) option.label = option.label.replace(/^(1[A-C]) /, '$1' + separator);
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a completed native answer remains mandatory', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered answer' }; },
|
||||
(c: NativePlanQuestionCall) => { c.toolUseId = ''; },
|
||||
(c: NativePlanQuestionCall) => { c.sessionId = ''; },
|
||||
]) {
|
||||
const c = call(); mutate(c);
|
||||
expect(isDesignCountFirstReview(nativePlanCallFingerprint(c, 0, true))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('number, menu and event identity cannot be borrowed from another decision', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Issue 2'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = c.questions[0]!.options[0]!.label; },
|
||||
]) {
|
||||
const c = call(); mutate(c);
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
for (const mutate of [
|
||||
(f: typeof captured) => { f.signature = 'foreign'; },
|
||||
(f: typeof captured) => { f.options.reverse(); },
|
||||
]) {
|
||||
const f = structuredClone(captured); mutate(f);
|
||||
expect(isDesignCountFirstReview(f)).toBe(false);
|
||||
}
|
||||
expect(isDesignCountFirstReview({ ...captured, nativeQuestionIndex: 1 })).toBe(false);
|
||||
});
|
||||
|
||||
test('quoted examples and workflow-only Issue titles do not start review', () => {
|
||||
for (const title of [
|
||||
'Example: D1 — Issue 1: Make Save the visible primary action?',
|
||||
'> D1 — Issue 1: Make Save the visible primary action?',
|
||||
'```\nD1 — Issue 1: Make Save the visible primary action?',
|
||||
'D1 — Issue 1: Make outside voices available?',
|
||||
'D1 — Issue 1: Fix which review runs next?',
|
||||
]) {
|
||||
const c = call(); c.questions[0]!.question = title;
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('a source citation or Keep fragment cannot replace opposed design choices', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { for (const o of c.questions[0]!.options) o.description = 'Read DESIGN.md before starting.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = '1Creeps into setup'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.label = '1C Start reviewing'; },
|
||||
]) {
|
||||
const c = call(); mutate(c);
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('a report or reviewer decision about compliant styles is administrative', () => {
|
||||
for (const [title, options] of [
|
||||
['D1 — Issue 1: Make a report about the primary actions?', [
|
||||
{ label: '1A Record the completed review', description: 'Matches DESIGN.md exactly: primary actions already use the approved styles. Write a report describing that existing result.' },
|
||||
{ label: '1B Keep the current review report', description: 'Leave the existing report unchanged. No product or implementation decision remains.' },
|
||||
]],
|
||||
['D1 — Issue 1: Make the typography review the next step?', [
|
||||
{ label: '1A Start the typography reviewer', description: 'Matches DESIGN.md exactly: the existing typography already complies. Ask another reviewer to confirm it.' },
|
||||
{ label: '1B Keep reviewing manually', description: 'Continue the review without another reviewer. No design change is proposed.' },
|
||||
]],
|
||||
] as const) {
|
||||
const c = call();
|
||||
c.questions[0]!.question = title;
|
||||
c.questions[0]!.options = options.map(option => ({ ...option }));
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the alternate primary style remedy must bind the same control and unresolved violation', () => {
|
||||
const renamed = call();
|
||||
renamed.questions[0]!.question = renamed.questions[0]!.question.replaceAll('Save', 'Submit');
|
||||
for (const option of renamed.questions[0]!.options) {
|
||||
option.label = option.label.replaceAll('Save', 'Submit');
|
||||
option.description = option.description?.replaceAll('Save', 'Submit');
|
||||
}
|
||||
expect(isDesignCountFirstReview(answer(renamed))).toBe(true);
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = c.questions[0]!.options[0]!.description!.replace('Save filled', 'Reset filled'); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = 'Matches DESIGN.md exactly: the existing buttons already comply. Record the result.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.description = 'The current buttons already comply. No unresolved design requirement remains.'; },
|
||||
]) {
|
||||
const c = call(); mutate(c);
|
||||
expect(isDesignCountFirstReview(answer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the regression and retained native call select the affected live workflow', () => {
|
||||
for (const file of ['test/design-first-decision-af.test.ts', 'test/fixtures/design-first-decision-af.json', 'test/fixtures/design-first-decision-af-retry.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -1,132 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-first-issue-ai.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
|
||||
const findings = captured.calls.filter(row => row.ordinal >= 3 && row.ordinal <= 7);
|
||||
test.each(findings)('actual completed Design D$ordinal starts review', ({ fingerprint }) => {
|
||||
expect(isDesignCountFirstReview(fingerprint)).toBe(true);
|
||||
expect(planCountQuestionPhase(fingerprint, false, designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup)).toMatchObject({ preReview: false, reviewStarted: true });
|
||||
});
|
||||
|
||||
function first(): any { return structuredClone(findings[0]!.fingerprint); }
|
||||
function question(fp: any, text: string) {
|
||||
const call = fp.nativeCall, old = call.questions[0].question;
|
||||
call.questions[0].question = text;
|
||||
call.answers = { [text]: call.answers[old] };
|
||||
}
|
||||
function options(fp: any, change: (q: any) => void) {
|
||||
const call = fp.nativeCall, q = call.questions[0];
|
||||
change(q);
|
||||
fp.options = q.options.map((o: any, i: number) => ({ index: i + 1, label: o.label }));
|
||||
call.answers = { [q.question]: q.options[0].label };
|
||||
}
|
||||
|
||||
test('actual routing, learnings and future typography TODO do not start a design review', () => {
|
||||
for (const row of captured.calls.filter(row => [1, 2, 9].includes(row.ordinal)))
|
||||
expect(isDesignCountFirstReview(row.fingerprint)).toBe(false);
|
||||
});
|
||||
|
||||
test('any offered answer and menu order can resolve a substantive finding', () => {
|
||||
for (const row of findings) {
|
||||
for (const option of row.fingerprint.nativeCall.questions[0]!.options) {
|
||||
const fp: any = structuredClone(row.fingerprint);
|
||||
fp.nativeCall.answers = { [fp.nativeCall.questions[0].question]: option.label };
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
const fp: any = structuredClone(row.fingerprint);
|
||||
options(fp, q => q.options.reverse());
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('native completion, request identity, current answers and aligned menu are required', () => {
|
||||
for (const change of [
|
||||
(f: any) => { delete f.nativeCall; },
|
||||
(f: any) => { f.nativeCall.answered = false; },
|
||||
(f: any) => { f.nativeCall.failed = true; },
|
||||
(f: any) => { f.nativeCall.sessionId = 'foreign'; },
|
||||
(f: any) => { f.nativeCall.toolUseId = 'stale-request'; },
|
||||
(f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(f: any) => { f.nativeCall.unansweredQuestionIndices = [0]; },
|
||||
(f: any) => { f.nativeCall.answers = {}; },
|
||||
(f: any) => { f.nativeCall.answers[f.nativeCall.questions[0].question] = 'not offered'; },
|
||||
(f: any) => { f.nativeCall.questions[0].question += '\nCorrection: this is a new question.'; },
|
||||
(f: any) => { f.nativeCall.answeredAt = 'invalid'; },
|
||||
(f: any) => { f.nativeCall.questions.push(structuredClone(f.nativeCall.questions[0])); },
|
||||
(f: any) => { f.nativeCall.questions[0].multiSelect = true; },
|
||||
(f: any) => { f.options.reverse(); },
|
||||
(f: any) => { f.nativeCall.questions[0].header = 'Issue 7'; },
|
||||
]) { const fp = first(); change(fp); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('numbered issue and all choice identifiers agree without depending on D numbering', () => {
|
||||
const fp = first(); question(fp, fp.nativeCall.questions[0].question.replace('D3 —', 'D27:'));
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
for (const change of [
|
||||
(q: any) => { q.options[0].label = q.options[0].label.replace('1A:', '2A:'); },
|
||||
(q: any) => { q.options[1].label = q.options[0].label; },
|
||||
]) { const f = first(); options(f, change); expect(isDesignCountFirstReview(f)).toBe(false); }
|
||||
});
|
||||
|
||||
test('a design Issue heading cannot borrow review content for setup, navigation or future work', () => {
|
||||
for (const title of [
|
||||
'Should we run outside design voices now?',
|
||||
'How should we configure design review routing?',
|
||||
'What review should run after the design review?',
|
||||
'Should we record an app-wide typography TODO?',
|
||||
'What type scale will form labels use after a future redesign?',
|
||||
]) {
|
||||
const fp = first(); question(fp, fp.nativeCall.questions[0].question.replace(/Issue 1: [^\n]+/, `Issue 1: ${title}`));
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
}
|
||||
for (const replacement of ['PLAN.md onboarding', 'PLAN.md post-review TODO', 'PLAN.md engineering review']) {
|
||||
const fp = first(); question(fp, fp.nativeCall.questions[0].question.replace('PLAN.md design review', replacement));
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('quoted, hypothetical and withdrawn declarations cannot start the phase', () => {
|
||||
for (const change of [
|
||||
(text: string) => `Example: ${text}`,
|
||||
(text: string) => `\`\`\`text\n${text}\n\`\`\``,
|
||||
(text: string) => text.replace('ELI10: ', 'ELI10: Example only: '),
|
||||
(text: string) => `${text}\nCorrection: that question was hypothetical and is withdrawn.`,
|
||||
(text: string) => text.replace('How should Save', 'If we later proceed, how should Save'),
|
||||
]) { const fp = first(); question(fp, change(fp.nativeCall.questions[0].question)); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('concrete design conformance and an opposed current violation belong to different offered choices', () => {
|
||||
for (const change of [
|
||||
(q: any) => { q.options.forEach((o: any) => { o.description = 'This is an available option.'; }); },
|
||||
(q: any) => { q.options[0].description = '✅ Example only: ' + q.options[0].description; },
|
||||
(q: any) => { q.options[2].description = 'No current design gap remains.'; },
|
||||
(q: any) => { q.options[0].label = '1A: Run primary review (recommended)'; },
|
||||
]) { const fp = first(); options(fp, change); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
test('the new exact public fixture and controls select only the design count workflow', () => {
|
||||
for (const path of ['test/design-first-issue-ai.test.ts', 'test/fixtures/design-first-issue-ai.json'])
|
||||
expect(selectTests([path], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
});
|
||||
|
||||
test('the assessment asserts a current defect, preserving conditional stakes and quoted history', () => {
|
||||
for (const change of [
|
||||
(s: string) => s.replace('ELI10: The header', 'ELI10: Suppose the header'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, "ELI10: '$1'"),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace('\nStakes if we pick wrong:', ' This issue is withdrawn.\nStakes if we pick wrong:'),
|
||||
(s: string) => s.replace('\nStakes if we pick wrong:', ' We have resolved this finding.\nStakes if we pick wrong:'),
|
||||
]) { const fp = first(); question(fp, change(fp.nativeCall.questions[0].question)); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
const conditional = first(); question(conditional, conditional.nativeCall.questions[0].question.replace('Stakes if we pick wrong:', 'Stakes if we pick wrong: If we leave this unchanged,'));
|
||||
expect(isDesignCountFirstReview(conditional)).toBe(true);
|
||||
const history = first(); question(history, history.nativeCall.questions[0].question.replace('\nStakes if we pick wrong:', ' The old report claimed "We have resolved this finding.", but that claim was wrong.\nStakes if we pick wrong:'));
|
||||
expect(isDesignCountFirstReview(history)).toBe(true);
|
||||
});
|
||||
|
||||
test('an explicit no-current-issue assessment cannot borrow the offered fixes', () => {
|
||||
const fp = first();
|
||||
question(fp, fp.nativeCall.questions[0].question.replace(/^ELI10: .+$/m, 'ELI10: The header shows four clearly differentiated buttons. DESIGN.md is fully followed. No current issue remains.'));
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
});
|
||||
@@ -1,231 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-primary-action-aj.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
|
||||
test('the exact completed current primary-action choice starts Design review', () => {
|
||||
expect(isDesignCountFirstReview(captured)).toBe(true);
|
||||
expect(planCountQuestionPhase(captured, false, designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup))
|
||||
.toMatchObject({ preReview: false, reviewStarted: true });
|
||||
});
|
||||
|
||||
function fresh(): any { return structuredClone(captured); }
|
||||
function changeText(fp: any, change: (text: string) => string) {
|
||||
const q = fp.nativeCall.questions[0], answer = fp.nativeCall.answers[q.question];
|
||||
q.question = change(q.question); fp.nativeCall.answers = { [q.question]: answer };
|
||||
}
|
||||
function changeMenu(fp: any, change: (q: any) => void) {
|
||||
const q = fp.nativeCall.questions[0]; change(q);
|
||||
fp.options = q.options.map((o: any, i: number) => ({ index: i + 1, label: o.label }));
|
||||
fp.nativeCall.answers = { [q.question]: q.options[0].label };
|
||||
}
|
||||
|
||||
test('an offered deferral, reordered menu and consistently renamed control remain review decisions', () => {
|
||||
for (const option of captured.nativeCall.questions[0]!.options) {
|
||||
const fp = fresh(); fp.nativeCall.answers = { [fp.nativeCall.questions[0].question]: option.label };
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
const reordered = fresh(); changeMenu(reordered, q => q.options.reverse());
|
||||
expect(isDesignCountFirstReview(reordered)).toBe(true);
|
||||
const renamed = fresh(); changeText(renamed, s => s.replaceAll('Save', 'Submit'));
|
||||
changeMenu(renamed, q => q.options.forEach((o: any) => {
|
||||
o.label = o.label.replaceAll('Save', 'Submit'); o.description = o.description.replaceAll('Save', 'Submit');
|
||||
}));
|
||||
expect(isDesignCountFirstReview(renamed)).toBe(true);
|
||||
const numbered = fresh(); changeText(numbered, s => s.replaceAll('Issue 1', 'Issue 6').replaceAll('1A', '6A').replaceAll('1B', '6B').replaceAll('1C', '6C'));
|
||||
changeMenu(numbered, q => { q.header = 'Issue 6'; q.options.forEach((o: any) => { o.label = o.label.replace(/^1/, '6'); }); });
|
||||
expect(isDesignCountFirstReview(numbered)).toBe(true);
|
||||
});
|
||||
|
||||
test('native completion, owned identity, offered answers and aligned numbering are necessary', () => {
|
||||
for (const change of [
|
||||
(f: any) => { delete f.nativeCall; },
|
||||
(f: any) => { f.nativeCall.answered = false; },
|
||||
(f: any) => { f.nativeCall.failed = true; },
|
||||
(f: any) => { f.nativeCall.sessionId = 'foreign-session'; },
|
||||
(f: any) => { f.nativeCall.toolUseId = 'foreign-request'; },
|
||||
(f: any) => { f.nativeCall.answeredAt = 'invalid'; },
|
||||
(f: any) => { delete f.nativeCall.answeredAt; },
|
||||
(f: any) => { f.nativeCall.unansweredQuestionIndices = [0]; },
|
||||
(f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(f: any) => { f.nativeCall.answers = {}; },
|
||||
(f: any) => { f.nativeCall.answers[f.nativeCall.questions[0].question] = 'unoffered'; },
|
||||
(f: any) => { f.nativeCall.questions[0].question += ' altered'; },
|
||||
(f: any) => { f.nativeCall.questions[0].header = 'Issue 2'; },
|
||||
(f: any) => { f.nativeCall.questions[0].multiSelect = true; },
|
||||
(f: any) => { f.options.reverse(); },
|
||||
(f: any) => { changeMenu(f, q => { q.options[0].label = q.options[0].label.replace('1A', '2A'); }); },
|
||||
]) { const fp = fresh(); change(fp); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('quoted, hypothetical, future and explicitly withdrawn assessments cannot borrow style choices', () => {
|
||||
for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '```text\n' + s + '\n```',
|
||||
(s: string) => s.replace('ELI10: Right now', 'ELI10: Suppose right now'),
|
||||
(s: string) => s.replace('ELI10: Right now', 'ELI10: If approved, right now'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, "ELI10: '$1'"),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: $1 This issue is withdrawn.'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: $1 We have resolved this finding.'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: $1 No current gap remains.'),
|
||||
(s: string) => s + '\nCorrection: this issue is withdrawn.',
|
||||
(s: string) => s.replace('make Save the only filled primary action?', 'make Save the only filled primary action in a future redesign?'),
|
||||
(s: string) => s.replace('make Save the only filled primary action?', 'make the primary reviewer the next step?'),
|
||||
(s: string) => s.replace('ELI10: Right now Save,', 'ELI10: Right now Publish,'),
|
||||
]) { const fp = fresh(); changeText(fp, change); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('the named amendment and unresolved violation belong to distinct current offered choices', () => {
|
||||
for (const change of [
|
||||
(q: any) => { q.options[0].description = q.options[0].description.replace('✅ Save is', '✅ Publish is'); },
|
||||
(q: any) => { q.options[0].description = '✅ Example only: ' + q.options[0].description; },
|
||||
(q: any) => { q.options[0].description = '✅ Save is not the single filled primary action.'; },
|
||||
(q: any) => { q.options[0].description += ' This issue is withdrawn.'; },
|
||||
(q: any) => { q.options[2].description = 'All buttons already comply. No current issue remains.'; },
|
||||
(q: any) => { q.options[2].description = '❌ Hypothetical: Primary-action ambiguity ships; documented DESIGN.md violation remains.'; },
|
||||
(q: any) => { q.options[2].description += ' Correction: this issue is resolved.'; },
|
||||
(q: any) => { q.options[0].description += ' ' + q.options[2].description; q.options[2].description = 'Another compliant option.'; },
|
||||
]) { const fp = fresh(); changeMenu(fp, change); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('conditional stakes and an unrelated quoted historical claim retain the current choice', () => {
|
||||
const fp = fresh(); changeText(fp, s => s.replace('Stakes if we pick wrong:', 'Stakes if we pick wrong: If unchanged,'));
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
const history = fresh(); changeText(history, s => s.replace('\nStakes if we pick wrong:', ' The old report claimed "This issue is resolved.", but that claim was wrong.\nStakes if we pick wrong:'));
|
||||
expect(isDesignCountFirstReview(history)).toBe(true);
|
||||
});
|
||||
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
test('the exact public fixture and controls select only the affected Design count workflow', () => {
|
||||
for (const file of ['test/design-primary-action-aj.test.ts', 'test/fixtures/design-primary-action-aj.json'])
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
});
|
||||
|
||||
import deferredTodo from './fixtures/design-future-todo-aj.json';
|
||||
import { isDesignArtifactGeneration } from './helpers/design-artifact-question';
|
||||
test('the actual completed future-only TODO recording is administrative and retains freshness', () => {
|
||||
expect(isDesignArtifactGeneration(deferredTodo)).toBe(true);
|
||||
expect(planCountQuestionPhase(deferredTodo, true, designStep0Boundary, isDesignCountFirstReview,
|
||||
isDesignCountSetup, undefined, isDesignArtifactGeneration)).toEqual({ preReview: false, reviewStarted: true, administrative: 'artifact-generation' });
|
||||
});
|
||||
|
||||
test('skipping a future TODO is administrative, while building it now remains a review decision', () => {
|
||||
for (const index of [0, 1, 2]) {
|
||||
const fp: any = structuredClone(deferredTodo), q = fp.nativeCall.questions[0];
|
||||
fp.nativeCall.answers = { [q.question]: q.options[index].label };
|
||||
expect(isDesignArtifactGeneration(fp)).toBe(index !== 2);
|
||||
const phase = planCountQuestionPhase(fp, true, designStep0Boundary, isDesignCountFirstReview,
|
||||
isDesignCountSetup, undefined, isDesignArtifactGeneration);
|
||||
expect(phase.preReview).toBe(false);
|
||||
expect(phase.administrative).toBe(index !== 2 ? 'artifact-generation' : undefined);
|
||||
}
|
||||
const fp: any = structuredClone(deferredTodo);
|
||||
expect(planCountQuestionPhase(fp, false, designStep0Boundary, isDesignCountFirstReview,
|
||||
isDesignCountSetup, undefined, isDesignArtifactGeneration).reviewStarted).toBe(false);
|
||||
});
|
||||
|
||||
test('a deferred artifact requires completed native identity, the exact answer and full aligned menu', () => {
|
||||
for (const change of [
|
||||
(f: any) => { delete f.nativeCall; },
|
||||
(f: any) => { f.nativeCall.answered = false; },
|
||||
(f: any) => { f.nativeCall.failed = true; },
|
||||
(f: any) => { f.nativeCall.sessionId = 'foreign'; },
|
||||
(f: any) => { f.nativeCall.answeredAt = 'invalid'; },
|
||||
(f: any) => { f.nativeCall.unansweredQuestionIndices = [0]; },
|
||||
(f: any) => { f.nativeQuestionIndex = 1; },
|
||||
(f: any) => { f.nativeCall.answers = {}; },
|
||||
(f: any) => { f.nativeCall.answers[f.nativeCall.questions[0].question] = 'unoffered'; },
|
||||
(f: any) => { f.options.reverse(); },
|
||||
(f: any) => { f.nativeCall.questions[0].multiSelect = true; },
|
||||
(f: any) => { f.nativeCall.questions.push(structuredClone(f.nativeCall.questions[0])); },
|
||||
(f: any) => { f.nativeCall.questions[0].options.pop(); },
|
||||
]) { const fp: any = structuredClone(deferredTodo); change(fp); expect(isDesignArtifactGeneration(fp)).toBe(false); }
|
||||
});
|
||||
|
||||
test('a deferred TODO cannot conceal current implementation, changed scope or source-only declarations', () => {
|
||||
for (const change of [
|
||||
(s: string) => 'Example: ' + s,
|
||||
(s: string) => '```text\n' + s + '\n```',
|
||||
(s: string) => s.replace('ELI10: DESIGN.md', 'ELI10: Suppose DESIGN.md'),
|
||||
(s: string) => s.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'),
|
||||
(s: string) => s.replace('so nothing changes now.', 'but replace the font now.'),
|
||||
(s: string) => s.replace('in a later design pass?', 'in this update?'),
|
||||
(s: string) => s + '\nCorrection: font replacement is now in scope; implement it now.',
|
||||
]) { const fp: any = structuredClone(deferredTodo); changeText(fp, change); expect(isDesignArtifactGeneration(fp)).toBe(false); }
|
||||
for (const index of [0, 1, 2]) {
|
||||
const fp: any = structuredClone(deferredTodo);
|
||||
fp.nativeCall.questions[0].options[index].description += ' Also fix the current form typography in this PR.';
|
||||
expect(isDesignArtifactGeneration(fp)).toBe(false);
|
||||
}
|
||||
const current = fresh();
|
||||
expect(isDesignArtifactGeneration(current)).toBe(false);
|
||||
expect(isDesignCountFirstReview(current)).toBe(true);
|
||||
});
|
||||
|
||||
test('deferred artifact classification follows offered identities and the current approved font', () => {
|
||||
const reordered: any = structuredClone(deferredTodo);
|
||||
changeMenu(reordered, q => q.options.reverse());
|
||||
reordered.nativeCall.answers = { [reordered.nativeCall.questions[0].question]: 'A Add to TODOS.md (recommended)' };
|
||||
expect(isDesignArtifactGeneration(reordered)).toBe(true);
|
||||
const renamed: any = structuredClone(deferredTodo);
|
||||
changeText(renamed, s => s.replaceAll('system-ui', 'ApprovedSans'));
|
||||
changeMenu(renamed, q => q.options.forEach((o: any) => { o.description = o.description.replaceAll('system-ui', 'ApprovedSans'); }));
|
||||
expect(isDesignArtifactGeneration(renamed)).toBe(true);
|
||||
});
|
||||
|
||||
test('the deferred TODO fixture selects the same affected Design count workflow', () => {
|
||||
expect(selectTests(['test/fixtures/design-future-todo-aj.json'], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
});
|
||||
|
||||
test('deferred TODO scope survives benign explanations, estimates and a concise equivalent proposal', () => {
|
||||
for (const change of [
|
||||
(s: string) => s.replace('Why: a chosen typeface is the cheapest tell that the app was designed rather than assembled. Pros: brand voice across the whole app.', 'Why: a deliberate typeface could make the application recognizable. Pros: a consistent future brand voice.'),
|
||||
(s: string) => s.replace('(for example DM Sans, Instrument Sans, IBM Plex Sans)', '(for example Atkinson Hyperlegible)'),
|
||||
(s: string) => s.replace('Stakes if we pick wrong: either the debt is forgotten, or a note lands in TODOS.md that you consider noise.', 'Stakes if we pick wrong: the future debt may be forgotten, or the backlog may become noisy.').replace('Recommendation: A because the debt is real but explicitly out of scope, and a written TODO costs nothing.', 'Recommendation: A to retain the explicitly out-of-scope debt for later.').replace('Net: keep the typography debt visible vs. drop it.', 'Net: record the deferred typography debt or omit the note.'),
|
||||
(s: string) => s.replace('DESIGN.md and this plan keep system-ui as the app font, and you excluded visual exploration from this update, so nothing changes now.', 'DESIGN.md and this plan retain system-ui as the app font. Visual exploration remains out of scope for this update, so nothing changes now.'),
|
||||
]) { const fp: any = structuredClone(deferredTodo); changeText(fp, change); expect(isDesignArtifactGeneration(fp)).toBe(true); }
|
||||
const estimate: any = structuredClone(deferredTodo);
|
||||
estimate.nativeCall.questions[0].options[0].description = estimate.nativeCall.questions[0].options[0].description.replace('human: ~5min / CC: ~1min to record', 'human: ~10min / CC: ~2min to record');
|
||||
expect(isDesignArtifactGeneration(estimate)).toBe(true);
|
||||
const concise: any = structuredClone(deferredTodo);
|
||||
changeText(concise, s => s.replace('record a deferred TODOS.md item to evaluate a real body typeface', 'add a deferred TODOS.md note to consider an alternate body typeface').replace('in a later design pass?', 'during a future design pass?').replace(/^ELI10: .+$/m,
|
||||
'ELI10: DESIGN.md and the current plan preserve system-ui as the app font. Visual exploration is out of scope for this update, so the current design remains unchanged. This question only records a deferred TODOS.md note for a future /design-consultation; it does not change the current design.'));
|
||||
changeMenu(concise, q => {
|
||||
q.options[0].description = '✅ Records only a TODOS.md note for a future /design-consultation. No design changes in this update; DESIGN.md and system-ui remain unchanged.';
|
||||
q.options[1].description = '✅ No TODO is recorded. No follow-up work.';
|
||||
q.options[2].description = '✅ Replace the font now in this PR.';
|
||||
});
|
||||
expect(isDesignArtifactGeneration(concise)).toBe(true);
|
||||
});
|
||||
|
||||
test('paraphrased facts still require affirmative preservation and reject present work', () => {
|
||||
for (const change of [
|
||||
(s: string) => s.replace('so nothing changes now.', 'so it is false that nothing changes now.'),
|
||||
(s: string) => s.replace('ELI10: DESIGN.md and this plan keep system-ui as the app font', 'ELI10: DESIGN.md and this plan keep Roboto as the app font'),
|
||||
(s: string) => s.replace('Net: keep the typography debt visible vs. drop it.', 'Net: replace the font now.'),
|
||||
(s: string) => s.replace('Net: keep the typography debt visible vs. drop it.', 'Net: this scope is withdrawn.'),
|
||||
]) { const fp: any = structuredClone(deferredTodo); changeText(fp, change); expect(isDesignArtifactGeneration(fp)).toBe(false); }
|
||||
const conditional: any = structuredClone(deferredTodo);
|
||||
conditional.nativeCall.questions[0].options[0].description = conditional.nativeCall.questions[0].options[0].description.replace('Nothing changes in this update;', 'If approved: Nothing changes in this update;');
|
||||
expect(isDesignArtifactGeneration(conditional)).toBe(false);
|
||||
const additional: any = structuredClone(deferredTodo);
|
||||
additional.nativeCall.questions[0].options[0].description += ' Add a 48px button target to this plan.';
|
||||
expect(isDesignArtifactGeneration(additional)).toBe(false);
|
||||
for (const suffix of ['Add a TODOS.md note and make the Save button 48px.', 'Add a TODOS.md note for the future font review and make the Save button 48px.']) {
|
||||
const mixed: any = structuredClone(deferredTodo);
|
||||
mixed.nativeCall.questions[0].options[0].description += ' ' + suffix;
|
||||
expect(isDesignArtifactGeneration(mixed)).toBe(false);
|
||||
}
|
||||
const recordingOnly: any = structuredClone(deferredTodo);
|
||||
recordingOnly.nativeCall.questions[0].options[0].description += ' Add a TODOS.md note for the future font review.';
|
||||
expect(isDesignArtifactGeneration(recordingOnly)).toBe(true);
|
||||
for (const suffix of ['Visual exploration is no longer out of scope.', 'This plan no longer keeps system-ui.']) {
|
||||
const fp: any = structuredClone(deferredTodo);
|
||||
changeText(fp, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: $1 ' + suffix));
|
||||
expect(isDesignArtifactGeneration(fp)).toBe(false);
|
||||
}
|
||||
const archival: any = structuredClone(deferredTodo);
|
||||
changeText(archival, s => s.replace(/^ELI10: (.+)$/m, 'ELI10: $1 Historical note: "Visual exploration is no longer out of scope."'));
|
||||
expect(isDesignArtifactGeneration(archival)).toBe(true);
|
||||
});
|
||||
@@ -1,107 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-primary-assignment-ao.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { AskUserQuestionFingerprint as Fingerprint } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
type Question = NonNullable<Fingerprint['nativeCall']>['questions'][number];
|
||||
function edit(change: (q: Question) => void): Fingerprint {
|
||||
const fp = structuredClone(captured.fingerprints[0]) as Fingerprint;
|
||||
const call = fp.nativeCall!, q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers![q.question]);
|
||||
change(q);
|
||||
call.answers = { [q.question]: q.options[selected]!.label };
|
||||
fp.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return fp;
|
||||
}
|
||||
|
||||
test('the exact style assignment begins review before the following pending-state decision', () => {
|
||||
let started = false;
|
||||
const phases = captured.fingerprints.map(raw => {
|
||||
const phase = planCountQuestionPhase(raw as Fingerprint, started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(phases).toEqual([
|
||||
{ preReview: false, reviewStarted: true },
|
||||
{ preReview: false, reviewStarted: true },
|
||||
]);
|
||||
expect(captured.fingerprints.map(fp => fp.preReview)).toEqual([true, true]);
|
||||
});
|
||||
|
||||
test('assignment whitespace and current deferral phrasing compose', () => {
|
||||
for (const separator of [' = ', '=', ' =']) {
|
||||
for (const action of ['Leave', 'Keep']) {
|
||||
for (const debt of ['debt', 'an open issue']) {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.options[0]!.description = q.options[0]!.description!.replaceAll(' = ', separator);
|
||||
q.options[2]!.description = q.options[2]!.description!.replace('Leave the header', `${action} the header`)
|
||||
.replace('as debt.', `as ${debt}.`);
|
||||
}))).toBe(true);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('the named control and offered answer may change without changing the review phase', () => {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.question = q.question.replaceAll('Save', 'Submit');
|
||||
q.options = q.options.map(o => ({ label: o.label.replaceAll('Save', 'Submit'),
|
||||
description: o.description?.replaceAll('Save', 'Submit') }));
|
||||
}))).toBe(true);
|
||||
for (const selected of [0, 1, 2]) {
|
||||
const fp = edit(() => {}), q = fp.nativeCall!.questions[0]!;
|
||||
fp.nativeCall!.answers = { [q.question]: q.options[selected]!.label };
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
const rejected: Array<[string, (q: Question) => void]> = [
|
||||
['withdrawn contract', q => { q.question += '\nThis DESIGN.md contract is "withdrawn".'; }],
|
||||
['superseded contract', q => { q.question += '\nThis contract is "superseded".'; }],
|
||||
['wrong primary', q => { q.options[0]!.description = q.options[0]!.description!.replace('Save =', 'Reset ='); }],
|
||||
['primary also ghost', q => { q.options[0]!.description = q.options[0]!.description!.replace('Reset/Cancel/Export =', 'Save/Cancel/Export ='); }],
|
||||
['no primary foreground', q => { q.options[0]!.description = q.options[0]!.description!.replace(' with white text', ''); }],
|
||||
['no ghost treatment', q => { q.options[0]!.description = q.options[0]!.description!.replace('neutral ghost Buttons', 'filled Buttons'); }],
|
||||
['no current authority', q => { q.options[0]!.description = q.options[0]!.description!.replace('per DESIGN.md', 'per a draft proposal'); }],
|
||||
['conditional assignment', q => { q.options[0]!.description = 'If approved later: ' + q.options[0]!.description; }],
|
||||
['historical assignment', q => { q.options[0]!.description = 'Historical example: ' + q.options[0]!.description; }],
|
||||
['quoted assignment', q => { q.options[0]!.description = '> ' + q.options[0]!.description; }],
|
||||
['cancelled assignment', q => { q.options[0]!.description += '\nCorrection: do not apply these styles.'; }],
|
||||
['rejected assignment', q => { q.options[0]!.description += '\nThis amendment is "rejected".'; }],
|
||||
['no opposed option', q => { q.options[2]!.label = '1C Configure Export'; }],
|
||||
['no documented violation', q => { q.options[2]!.description = q.options[2]!.description!.replace('Ships a known DESIGN.md violation', 'Satisfies DESIGN.md'); }],
|
||||
['no remaining primary gap', q => { q.options[2]!.description = q.options[2]!.description!.replace('stays undiscoverable', 'becomes obvious'); }],
|
||||
['conditional deferral', q => { q.options[2]!.description = 'If approved later: ' + q.options[2]!.description; }],
|
||||
['historical deferral', q => { q.options[2]!.description = 'Historical example: ' + q.options[2]!.description; }],
|
||||
['quoted deferral', q => { q.options[2]!.description = '> ' + q.options[2]!.description; }],
|
||||
['cancelled deferral', q => { q.options[2]!.description += '\nThis deferral is "cancelled".'; }],
|
||||
['cancelled header instruction', q => { q.options[2]!.description += '\nCorrection: do not leave the header unchanged.'; }],
|
||||
['cancelled keep instruction', q => { q.options[2]!.description += '\nCorrection: do not keep the header unchanged.'; }],
|
||||
['resolved violation', q => { q.options[2]!.description += '\nThis violation is now resolved.'; }],
|
||||
['conditional benefit', q => { q.options[2]!.description = q.options[2]!.description!.replace('✅ Zero implementation', '✅ If approved later: zero implementation'); }],
|
||||
];
|
||||
test.each(rejected)('%s cannot start review', (_, change) => {
|
||||
expect(isDesignCountFirstReview(edit(change))).toBe(false);
|
||||
});
|
||||
|
||||
test('a quoted historical cancellation does not cancel the current deferral', () => {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.options[2]!.description += '\nHistorical note: "Correction: do not leave the header unchanged."';
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('unfinished or foreign native calls cannot start review', () => {
|
||||
const incomplete = edit(() => {}); incomplete.nativeCall!.answered = false;
|
||||
expect(isDesignCountFirstReview(incomplete)).toBe(false);
|
||||
const foreign = edit(() => {}); foreign.signature = 'foreign:tool';
|
||||
expect(isDesignCountFirstReview(foreign)).toBe(false);
|
||||
});
|
||||
|
||||
test('the retry regression maps only to the existing Design workflow owner', () => {
|
||||
for (const file of ['test/design-primary-assignment-ao.test.ts', 'test/fixtures/design-primary-assignment-ao.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -1,131 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-primary-composition-an.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { AskUserQuestionFingerprint as Fingerprint } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
type Question = NonNullable<Fingerprint['nativeCall']>['questions'][number];
|
||||
const original = () => structuredClone(captured.fingerprint) as Fingerprint;
|
||||
function edit(change: (question: Question) => void): Fingerprint {
|
||||
const fp = original(), call = fp.nativeCall!, question = call.questions[0]!;
|
||||
const chosen = question.options.findIndex(option => option.label === call.answers![question.question]);
|
||||
change(question);
|
||||
call.answers = { [question.question]: question.options[chosen]!.label };
|
||||
fp.options = question.options.map((option, index) => ({ index: index + 1, label: option.label }));
|
||||
return fp;
|
||||
}
|
||||
|
||||
test('the exact completed first Issue starts the existing review phase', () => {
|
||||
const fp = original();
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
expect(planCountQuestionPhase(fp, false, designStep0Boundary, isDesignCountFirstReview, isDesignCountSetup))
|
||||
.toEqual({ preReview: false, reviewStarted: true });
|
||||
// The old run's observation is retained; this is a prospective replay.
|
||||
expect(captured.fingerprint.preReview).toBe(true);
|
||||
});
|
||||
|
||||
test('primary qualifiers and authority position compose independently of wording', () => {
|
||||
for (const qualifier of ['single', 'single filled', 'only', 'only filled', 'visible']) {
|
||||
for (const authority of ['Apply DESIGN.md: ', 'Apply DESIGN.md tokens: ', 'suffix']) {
|
||||
const fp = edit(question => {
|
||||
question.question = question.question.replace('single filled primary', `${qualifier} primary`);
|
||||
const style = question.options[0]!.description!.replace('Apply DESIGN.md: ', '');
|
||||
question.options[0]!.description = authority === 'suffix' ? `${style} Exact DESIGN.md.` : authority + style;
|
||||
});
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('an offered alternate answer, renamed control, and numeric retained count preserve the decision', () => {
|
||||
for (const option of original().nativeCall!.questions[0]!.options) {
|
||||
const fp = original(), call = fp.nativeCall!;
|
||||
call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
expect(isDesignCountFirstReview(edit(question => {
|
||||
question.question = question.question.replaceAll('Save', 'Submit');
|
||||
question.options = question.options.map(option => ({
|
||||
label: option.label.replaceAll('Save', 'Submit'),
|
||||
description: option.description?.replaceAll('Save', 'Submit').replace('four header buttons', '4 buttons'),
|
||||
}));
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
const rejected: Array<[string, (question: Question) => void]> = [
|
||||
['foreign issue header', q => { q.header = 'Issue 2'; }],
|
||||
['reviewer setup title', q => { q.question = q.question.replace('Make Save the single filled primary action in the header', 'Run outside design voices'); }],
|
||||
['historical question', q => { q.question = 'Historical example:\n' + q.question; }],
|
||||
['source-framed assessment', q => { q.question = q.question.replace('\nELI10:', '\nSource excerpt:\nELI10:'); }],
|
||||
['quoted assessment', q => { q.question = q.question.replace('\nELI10:', '\n> ELI10:'); }],
|
||||
['conditional assessment', q => { q.question = q.question.replace('ELI10: Right now', 'ELI10: If right now'); }],
|
||||
['unequal current controls', q => { q.question = q.question.replace('look identical', 'do not look identical'); }],
|
||||
['withdrawn finding', q => { q.question += '\nThis issue is withdrawn.'; }],
|
||||
['missing style authority', q => { q.options[0]!.description = q.options[0]!.description!.replace('Apply DESIGN.md: ', ''); }],
|
||||
['wrong named primary', q => { q.options[0]!.description = q.options[0]!.description!.replace('Save filled', 'Reset filled'); }],
|
||||
['missing foreground', q => { q.options[0]!.description = q.options[0]!.description!.replace('/white', ''); }],
|
||||
['missing ghost treatment', q => { q.options[0]!.description = q.options[0]!.description!.replace('neutral ghost buttons', 'filled buttons'); }],
|
||||
['primary also offered as ghost', q => { q.options[0]!.description = q.options[0]!.description!.replace('; Reset', '; Save'); }],
|
||||
['conditional amendment', q => { q.options[0]!.description = 'If approved later: ' + q.options[0]!.description; }],
|
||||
['quoted amendment', q => { q.options[0]!.description = '> ' + q.options[0]!.description; }],
|
||||
['cancelled amendment', q => { q.options[0]!.description += ' Correction: do not apply these styles.'; }],
|
||||
['no opposed choice', q => { q.options[2]!.label = '1C Configure Export'; }],
|
||||
['wrong retained count', q => { q.options[2]!.description = q.options[2]!.description!.replace('four', 'three'); }],
|
||||
['conditional deferral', q => { q.options[2]!.description = 'If approved later: ' + q.options[2]!.description; }],
|
||||
['historical deferral', q => { q.options[2]!.description = 'Historical example: ' + q.options[2]!.description; }],
|
||||
['resolved deferral', q => { q.options[2]!.description += ' The gap is now resolved.'; }],
|
||||
['conditional project metadata', q => { q.question = q.question.replace('Project/branch/task: ', 'Project/branch/task: If approved: '); }],
|
||||
['rejected current issue', q => { q.question += '\nIssue 1 is rejected.'; }],
|
||||
['cancelled current issue', q => { q.question += '\nThis issue is cancelled.'; }],
|
||||
['rejected amendment', q => { q.options[0]!.description += ' This amendment is rejected.'; }],
|
||||
['styles no longer current', q => { q.options[0]!.description += ' Correction: these styles are not current.'; }],
|
||||
['rejected opposed option', q => { q.options[2]!.description += ' This option is rejected.'; }],
|
||||
['cancelled retained buttons', q => { q.options[2]!.description += ' Correction: do not keep all four buttons identical.'; }],
|
||||
['quoted rejected current issue', q => { q.question += '\nThis issue is "rejected".'; }],
|
||||
['quoted cancelled amendment', q => { q.options[0]!.description += ' This amendment is "cancelled".'; }],
|
||||
['quoted styles no longer current', q => { q.options[0]!.description += ' Correction: these styles are "not current".'; }],
|
||||
['quoted rejected opposed option', q => { q.options[2]!.description += ' This option is "rejected".'; }],
|
||||
];
|
||||
test.each(rejected)('%s cannot start review', (_, change) => {
|
||||
expect(isDesignCountFirstReview(edit(change))).toBe(false);
|
||||
});
|
||||
|
||||
test('additional native-field gap prose cannot bypass owned primary facts or rejection guards', () => {
|
||||
for (const suffix of [' Leaves the plan violating DESIGN.md.', ' The gap remains open.']) {
|
||||
expect(isDesignCountFirstReview(edit(q => { q.options[2]!.description += suffix; })), suffix).toBe(true);
|
||||
for (const [name, change] of rejected) {
|
||||
const fp = edit(q => {
|
||||
q.options[2]!.description += suffix;
|
||||
change(q);
|
||||
});
|
||||
expect(isDesignCountFirstReview(fp), name + suffix).toBe(false);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('quoted historical withdrawal does not cancel the current issue', () => {
|
||||
expect(isDesignCountFirstReview(edit(q => { q.question += '\nHistorical note: "This issue is withdrawn."'; }))).toBe(true);
|
||||
});
|
||||
|
||||
test('recognition still requires an owned, completed and aligned native answer', () => {
|
||||
const invalid: Array<(fp: Fingerprint) => void> = [
|
||||
fp => { fp.nativeCall!.answered = false; },
|
||||
fp => { fp.nativeCall!.failed = true; },
|
||||
fp => { fp.signature = 'foreign:tool'; },
|
||||
fp => { fp.nativeQuestionIndex = 1; },
|
||||
fp => { fp.nativeCall!.unansweredQuestionIndices = [0]; },
|
||||
fp => { delete fp.nativeCall!.answeredAt; },
|
||||
fp => { fp.nativeCall!.answers = {}; },
|
||||
fp => { fp.options.reverse(); },
|
||||
];
|
||||
for (const change of invalid) {
|
||||
const fp = original(); change(fp);
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the regression and public fixture select the affected Design workflow', () => {
|
||||
for (const path of ['test/design-primary-composition-an.test.ts', 'test/fixtures/design-primary-composition-an.json'])
|
||||
expect(selectTests([path], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
});
|
||||
@@ -1,101 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import fixture from './fixtures/design-primary-contract-ak.json';
|
||||
import { isDesignCountFirstReview } from './helpers/design-count-review';
|
||||
import type { AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
|
||||
type FP = AskUserQuestionFingerprint;
|
||||
type Question = NonNullable<FP['nativeCall']>['questions'][number];
|
||||
const original = () => structuredClone(fixture.fingerprint) as unknown as FP;
|
||||
function edit(change: (q: Question, fp: FP) => void): FP {
|
||||
const fp = original(), call = fp.nativeCall!, q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers![q.question]);
|
||||
change(q, fp);
|
||||
call.answers = { [q.question]: q.options[selected]!.label };
|
||||
fp.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return fp;
|
||||
}
|
||||
function replaceText(q: Question, from: string, to: string): void {
|
||||
q.question = q.question.replaceAll(from, to);
|
||||
q.options = q.options.map(o => ({
|
||||
...o, label: o.label.replaceAll(from, to),
|
||||
description: o.description?.replaceAll(from, to),
|
||||
}));
|
||||
}
|
||||
|
||||
describe('current primary-action contract from a completed native issue', () => {
|
||||
test('the owned Issue 1 is a review decision despite colon-numbered options', () => {
|
||||
expect(isDesignCountFirstReview(original())).toBe(true);
|
||||
});
|
||||
test.each([
|
||||
['renamed primary control', (q: Question) => replaceText(q, 'Save', 'Submit')],
|
||||
['other prescribed color', (q: Question) => replaceText(q, '#1d4ed8', '#234abc')],
|
||||
['numeric button count', (q: Question) => replaceText(q, 'are four identical buttons', 'are 4 identical buttons')],
|
||||
['explicit all button count', (q: Question) => replaceText(q, 'are four identical buttons', 'are all four identical buttons')],
|
||||
['uncounted current equality', (q: Question) => replaceText(q, 'are four identical buttons', 'are identical buttons')],
|
||||
['parenthesized option separators', (q: Question) => {
|
||||
q.options = q.options.map(o => ({ ...o, label: o.label.replace(/^1([ABC]):/, '1$1)') }));
|
||||
}],
|
||||
['current pro/con decline', (q: Question) => { q.options[2]!.description = '✅ No implementation work now. ✅ No visual retesting. ❌ ' + q.options[2]!.description; }],
|
||||
['same contract with explicit button noun', (q: Question) => {
|
||||
q.options[0]!.description = q.options[0]!.description!.replace('neutral ghost.', 'neutral ghost buttons.');
|
||||
}],
|
||||
])('%s preserves the actual contract', (_, change) => {
|
||||
expect(isDesignCountFirstReview(edit(change))).toBe(true);
|
||||
});
|
||||
const negatives: Array<[string, (q: Question, fp: FP) => void]> = [
|
||||
['failed call', (_, fp) => { fp.nativeCall!.failed = true; }],
|
||||
['unanswered call', (_, fp) => { fp.nativeCall!.answered = false; }],
|
||||
['pending question', (_, fp) => { fp.nativeCall!.unansweredQuestionIndices = [0]; }],
|
||||
['missing successful answer time', (_, fp) => { delete fp.nativeCall!.answeredAt; }],
|
||||
['invalid answer time', (_, fp) => { fp.nativeCall!.answeredAt = 'unknown'; }],
|
||||
['unowned signature', (_, fp) => { fp.signature = 'other:call'; }],
|
||||
['missing session', (_, fp) => { fp.nativeCall!.sessionId = ''; }],
|
||||
['multiple questions', (q, fp) => { fp.nativeCall!.questions.push(structuredClone(q)); }],
|
||||
['multiple selections', q => { q.multiSelect = true; }],
|
||||
['competing issue header', q => { q.header = 'Issue 2'; }],
|
||||
['competing control header', q => { q.header = 'Issue 1: Cancel'; }],
|
||||
['competing option identity', q => { q.options[0]!.label = q.options[0]!.label.replace('1A:', '2A:'); }],
|
||||
['duplicate options', q => { q.options[1]!.label = q.options[0]!.label; }],
|
||||
['historical assessment', q => replaceText(q, 'ELI10: Right now', 'ELI10: Previously')],
|
||||
['quoted assessment', q => replaceText(q, 'ELI10: Right now', 'ELI10: "Right now')],
|
||||
['conditional assessment', q => replaceText(q, 'ELI10: Right now', 'ELI10: If right now')],
|
||||
['negated equality', q => replaceText(q, 'are four identical buttons', 'are not identical buttons')],
|
||||
['other equal controls', q => replaceText(q, 'Right now Save, Reset', 'Right now Undo, Reset')],
|
||||
['no current assessment', q => { q.question = q.question.replace(/^ELI10:.*\n/m, ''); }],
|
||||
['hypothetical issue', q => { q.question += '\nThis issue is hypothetical.'; }],
|
||||
['withdrawn issue', q => { q.question += '\nIssue 1 has been withdrawn.'; }],
|
||||
['resolved issue', q => { q.question += '\nNo current gap remains.'; }],
|
||||
['amendment only quotes source', q => { q.options[0]!.description = '> ' + q.options[0]!.description; }],
|
||||
['conditional amendment', q => { q.options[0]!.description = 'If approved later, ' + q.options[0]!.description; }],
|
||||
['negated amendment', q => { q.options[0]!.description = 'Do not ' + q.options[0]!.description; }],
|
||||
['other primary amendment', q => { q.options[0]!.description = q.options[0]!.description!.replace('Save #', 'Reset #'); }],
|
||||
['administrative record action', q => { q.options[0]!.description = 'Record the current review in the plan file.'; }],
|
||||
['wrong design authority', q => { q.options[0]!.description = q.options[0]!.description!.replace('DESIGN.md', 'an archived example'); }],
|
||||
['no fill prescribed', q => { q.options[0]!.description = q.options[0]!.description!.replace('filled with', 'outlined with'); }],
|
||||
['no ghost secondary controls', q => { q.options[0]!.description = q.options[0]!.description!.replace('neutral ghost', 'identical filled'); }],
|
||||
['no opposed choice', q => { q.options[2]!.label = '1C: Export settings instead'; }],
|
||||
['opposed choice does not retain the gap', q => { q.options[2]!.description = 'The gap is already fixed; file the report.'; }],
|
||||
['opposed choice withdraws finding', q => { q.options[2]!.description += ' Issue 1 is withdrawn.'; }],
|
||||
['fenced assessment', q => { q.question = q.question.replace(/^(ELI10:.*)$/m, '```text\n$1\n```'); }],
|
||||
['competing assessments', q => { q.question += '\nELI10: Save is already the unique primary action; all secondary controls are ghosts.'; }],
|
||||
['assessment relabelled as history', q => { q.question = q.question.replace(/^(ELI10:.*)$/m, '$1 Correction: the identical-buttons sentence is a historical example, not the current UI.'); }],
|
||||
['primary also styled as secondary', q => { q.options[0]!.description = q.options[0]!.description!.replace('Reset, Cancel, Export neutral ghost.', 'Save, Reset, Cancel, Export neutral ghost.'); }],
|
||||
['later style cancellation', q => { q.options[0]!.description += ' Correction: do not apply these tokens; Save remains identical to the other buttons.'; }],
|
||||
['historical icon-prefixed decline', q => { q.options[2]!.description = 'Historical source excerpt: ❌ ' + q.options[2]!.description; }],
|
||||
['conditional icon-prefixed decline', q => { q.options[2]!.description = 'If approved later: ❌ ' + q.options[2]!.description; }],
|
||||
['historical pro/con decline', q => { q.options[2]!.description = '✅ Historical example: no implementation work. ❌ ' + q.options[2]!.description; }],
|
||||
['conditional pro/con decline', q => { q.options[2]!.description = '✅ If approved later: no implementation work. ❌ ' + q.options[2]!.description; }],
|
||||
['opposed gap later resolved', q => { q.options[2]!.description += ' Correction: this gap is already resolved; no style change is required.'; }],
|
||||
];
|
||||
test.each(negatives)('%s is not current completed review evidence', (_, change) => {
|
||||
expect(isDesignCountFirstReview(edit(change))).toBe(false);
|
||||
});
|
||||
test('an unknown answer or mismatched rendered menu cannot supply completion', () => {
|
||||
const answer = original();
|
||||
answer.nativeCall!.answers = { [answer.nativeCall!.questions[0]!.question]: 'not offered' };
|
||||
expect(isDesignCountFirstReview(answer)).toBe(false);
|
||||
const menu = original();
|
||||
menu.options[0]!.label = 'different visible choice';
|
||||
expect(isDesignCountFirstReview(menu)).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,62 +0,0 @@
|
||||
import {describe, expect, test} from 'bun:test';
|
||||
import fixture from './fixtures/design-primary-decision-al.json';
|
||||
import {isDesignCountFirstReview} from './helpers/design-count-review';
|
||||
import type {AskUserQuestionFingerprint} from './helpers/claude-pty-runner';
|
||||
type FP=AskUserQuestionFingerprint;
|
||||
type Q=NonNullable<FP['nativeCall']>['questions'][number];
|
||||
const original=()=>structuredClone(fixture.fingerprint) as unknown as FP;
|
||||
function edit(change:(q:Q,fp:FP)=>void):FP {
|
||||
const fp=original(),c=fp.nativeCall!,q=c.questions[0]!,chosen=q.options.findIndex(o=>o.label===c.answers![q.question]);
|
||||
change(q,fp);c.answers={[q.question]:q.options[chosen]!.label};
|
||||
fp.options=q.options.map((o,i)=>({index:i+1,label:o.label}));return fp;
|
||||
}
|
||||
describe('answered primary-action decision with compact style choices',()=>{
|
||||
test('recognizes the exact current native issue independently of its interrogative title',()=>{
|
||||
expect(isDesignCountFirstReview(original())).toBe(true);
|
||||
});
|
||||
const positive:Array<[string,(q:Q,fp:FP)=>void]>=[
|
||||
['different named primary',q=>{q.question=q.question.replaceAll('Save','Submit');q.options=q.options.map(o=>({...o,label:o.label.replaceAll('Save','Submit'),description:o.description?.replaceAll('Save','Submit')}));}],
|
||||
['explicit fill role and foreground',q=>{q.options[0]!.description=q.options[0]!.description!.replace('filled #1d4ed8/white','filled primary #234abc with black text').replace('neutral ghost.','neutral ghost buttons.');}],
|
||||
['actions without header qualification',q=>{q.question=q.question.replace('the header actions','actions');}],
|
||||
['imperative title with the compact style',q=>{q.question=q.question.replace('How should the header actions establish that Save is the primary action','Make Save the visible primary action');}],
|
||||
['interrogative title with expanded style',q=>{q.options[0]!.description='Apply DESIGN.md tokens: Save #1d4ed8 filled with white text; Reset, Cancel, Export neutral ghost.';}],
|
||||
['existing explicit open-gap deferral',q=>{q.options[2]!.description='Decline the fix; gap stays documented and lowers the score.';}],
|
||||
['numeric control count in deferral',q=>{q.options[2]!.description=q.options[2]!.description!.replace('four','4');}],
|
||||
['quoted historical note does not withdraw current amendment',q=>{q.options[0]!.description+=' Prior note: "This amendment is withdrawn."';}],
|
||||
];
|
||||
test.each(positive)('%s retains the same owned decision',(_,change)=>expect(isDesignCountFirstReview(edit(change))).toBe(true));
|
||||
const negative:Array<[string,(q:Q,fp:FP)=>void]>=[
|
||||
['equality qualified as archived only',q=>{q.question=q.question.replace('look identical.','look identical only in the archived screenshot. Today they are distinct.');}],
|
||||
['amendment relabelled as historical',q=>{q.options[0]!.description+=' This is a historical example, not the current amendment.';}],
|
||||
['amendment explicitly withdrawn',q=>{q.options[0]!.description+=' This amendment is withdrawn.';}],
|
||||
['deferral relabelled as historical',q=>{q.options[2]!.description+=' This is a historical example, not the current deferral.';}],
|
||||
['failed native call',(_,fp)=>{fp.nativeCall!.failed=true;}],
|
||||
['unanswered native call',(_,fp)=>{fp.nativeCall!.answered=false;}],
|
||||
['unbound signature',(_,fp)=>{fp.signature='other:call';}],
|
||||
['missing completion time',(_,fp)=>{delete fp.nativeCall!.answeredAt;}],
|
||||
['competing issue number',q=>{q.header='Issue 2';}],
|
||||
['competing option identity',q=>{q.options[0]!.label='2A Primary + ghost';}],
|
||||
['source-framed question',q=>{q.question='Historical example:\n'+q.question;}],
|
||||
['historical premise',q=>{q.question=q.question.replace('ELI10: Right now','ELI10: Previously');}],
|
||||
['quoted premise',q=>{q.question=q.question.replace('ELI10: Right now','> ELI10: Right now');}],
|
||||
['conditional premise',q=>{q.question=q.question.replace('ELI10: Right now','ELI10: If right now');}],
|
||||
['negated equality',q=>{q.question=q.question.replace('look identical','do not look identical');}],
|
||||
['competing premise',q=>{q.question+='\nELI10: No current hierarchy gap exists.';}],
|
||||
['resolved finding',q=>{q.question+='\nCorrection: this gap is already resolved.';}],
|
||||
['other primary in remedy',q=>{q.options[0]!.description=q.options[0]!.description!.replace('Save filled','Reset filled');}],
|
||||
['no prescribed fill',q=>{q.options[0]!.description=q.options[0]!.description!.replace('filled','outlined');}],
|
||||
['no prescribed foreground',q=>{q.options[0]!.description=q.options[0]!.description!.replace('/white','/unknown');}],
|
||||
['primary also a ghost',q=>{q.options[0]!.description=q.options[0]!.description!.replace('; Reset','; Save, Reset');}],
|
||||
['quoted amendment',q=>{q.options[0]!.description='> '+q.options[0]!.description;}],
|
||||
['conditional amendment',q=>{q.options[0]!.description='If approved later, '+q.options[0]!.description;}],
|
||||
['negated amendment',q=>{q.options[0]!.description='Do not apply: '+q.options[0]!.description;}],
|
||||
['withdrawn amendment',q=>{q.options[0]!.description+=' Correction: do not apply these tokens.';}],
|
||||
['wrong design authority',q=>{q.options[0]!.description=q.options[0]!.description!.replace('Exact DESIGN.md.','Archived example.');}],
|
||||
['no opposed choice',q=>{q.options[2]!.label='1C Export preferences';}],
|
||||
['defer does not retain equality',q=>{q.options[2]!.description=q.options[2]!.description!.replace('identical','distinct');}],
|
||||
['historical deferral',q=>{q.options[2]!.description='Historical source excerpt: '+q.options[2]!.description;}],
|
||||
['conditional deferral',q=>{q.options[2]!.description='If accepted later: '+q.options[2]!.description;}],
|
||||
['deferral closes gap',q=>{q.options[2]!.description+=' Correction: the violation is now closed.';}],
|
||||
];
|
||||
test.each(negative)('%s is not completed current-review evidence',(_,change)=>expect(isDesignCountFirstReview(edit(change))).toBe(false));
|
||||
});
|
||||
@@ -1,150 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-primary-emphasis-av-calls.json';
|
||||
import { nativePlanCallFingerprint, designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff } from './helpers/design-count-review';
|
||||
import { isDesignArtifactGeneration } from './helpers/design-artifact-question';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles-data';
|
||||
|
||||
const batches = [captured.calls, captured.retry.calls] as NativePlanQuestionCall[][];
|
||||
const indices = [1, 2];
|
||||
const fresh = (index: number) => structuredClone(batches[index]![indices[index]!]!);
|
||||
const fingerprint = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
const accepted = (call: NativePlanQuestionCall) => isDesignCountFirstReview(fingerprint(call));
|
||||
type Question = NativePlanQuestionCall['questions'][number];
|
||||
function change(index: number, edit: (question: Question, call: NativePlanQuestionCall) => void): NativePlanQuestionCall {
|
||||
const call = fresh(index), question = call.questions[0]!;
|
||||
edit(question, call);
|
||||
call.answers = { [question.question]: question.options[0]!.label };
|
||||
return call;
|
||||
}
|
||||
|
||||
describe('current primary emphasis and annotated header signal decisions', () => {
|
||||
test('exact public first and retry calls enter review at their first real issue', () => {
|
||||
for (const [index, batch] of batches.entries()) {
|
||||
const before = JSON.stringify(batch);
|
||||
let started = false;
|
||||
const phases = batch.map(call => {
|
||||
const phase = planCountQuestionPhase(fingerprint(call), started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup, isDesignCompletionHandoff, isDesignArtifactGeneration);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(batch.map(accepted)).toEqual(batch.map((_, i) => i === indices[index]));
|
||||
expect(phases.map(phase => phase.preReview)).toEqual(batch.map((_, i) => i < indices[index]!));
|
||||
expect(phases.filter(phase => phase.administrative)).toHaveLength(0);
|
||||
// Five actual issues satisfy the original floor without relying on the
|
||||
// retry's later TODO question, which is outside this first-entry fix.
|
||||
expect(phases.slice(indices[index], indices[index]! + 5).filter(phase => !phase.preReview)).toHaveLength(5);
|
||||
expect(JSON.stringify(batch)).toBe(before);
|
||||
}
|
||||
});
|
||||
|
||||
test('consistent actors, palette, finding ordinal and offered selections preserve meaning', () => {
|
||||
for (const index of [0, 1]) {
|
||||
const renamed = JSON.parse(JSON.stringify(fresh(index)).replaceAll('Save', 'Submit').replaceAll('Reset', 'Revert').replaceAll('#1d4ed8', '#234abc'));
|
||||
expect(accepted(renamed)).toBe(true);
|
||||
const ordinal = change(index, q => {
|
||||
q.header = q.header.replace('Issue 1', 'Issue 9');
|
||||
q.question = q.question.replace('Issue 1', 'Issue 9').replace(/\b1([ABC])\b/g, '9$1');
|
||||
q.options.forEach(option => { option.label = option.label.replace(/^1/, '9'); });
|
||||
});
|
||||
expect(accepted(ordinal)).toBe(true);
|
||||
for (const option of fresh(index).questions[0]!.options) {
|
||||
const call = fresh(index);
|
||||
call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(accepted(call)).toBe(true);
|
||||
}
|
||||
expect(accepted(change(index, q => q.options.reverse()))).toBe(true);
|
||||
expect(accepted(change(index, q => {
|
||||
q.question = q.question.replace(/D[23] —/, 'D17 —');
|
||||
}))).toBe(true);
|
||||
}
|
||||
for (const header of ['Primary CTA', 'Header hierarchy', 'Issue 1', 'Issue 1: Save']) {
|
||||
expect(accepted(change(1, q => { q.header = header; }))).toBe(true);
|
||||
}
|
||||
expect(accepted(change(1, q => { q.question = q.question.replace('(G1)', '(G19)'); }))).toBe(true);
|
||||
});
|
||||
|
||||
test('unacknowledged, failed, foreign and mismatched native identities do not start review', () => {
|
||||
const mutations = [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.answeredAt; },
|
||||
(c: NativePlanQuestionCall) => { c.answeredAt = 'invalid'; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered answer' }; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
];
|
||||
for (const index of [0, 1]) {
|
||||
for (const mutation of mutations) { const call = fresh(index); mutation(call); expect(accepted(call)).toBe(false); }
|
||||
for (const mutation of [
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.signature = 'foreign:call'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.sessionId = 'foreign'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeCall!.toolUseId = 'foreign'; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.nativeQuestionIndex = 1; },
|
||||
(fp: ReturnType<typeof fingerprint>) => { fp.options.reverse(); },
|
||||
]) { const fp = fingerprint(fresh(index)); mutation(fp); expect(isDesignCountFirstReview(fp)).toBe(false); }
|
||||
}
|
||||
});
|
||||
|
||||
test('descriptive headers cannot override conflicting ordinals or become setup navigation', () => {
|
||||
for (const index of [0, 1]) {
|
||||
for (const header of ['Issue 2', 'Issue 2: Save', 'Scope', 'Routing', 'Learnings', 'Outside voices', 'Next steps']) {
|
||||
expect(accepted(change(index, q => { q.header = header; }))).toBe(false);
|
||||
}
|
||||
expect(accepted(change(index, q => { q.options[0]!.label = q.options[0]!.label.replace('1A', '2A'); }))).toBe(false);
|
||||
expect(accepted(change(index, q => { q.question = q.question.replace('Save primary emphasis', 'the reviewer primary emphasis').replace('that Save is', 'that the reviewer is'); }))).toBe(false);
|
||||
expect(accepted(change(index, q => { q.question = q.question.replace('primary emphasis', 'review readiness').replace('primary action?', 'next reviewer?'); }))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('current equal-weight premise cannot come from a quote, source or future condition', () => {
|
||||
for (const index of [0, 1]) for (const edit of [
|
||||
(q: Question) => { q.question = 'Historical example:\n' + q.question; },
|
||||
(q: Question) => { q.question = '```text\n' + q.question + '\n```'; },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: Right now', 'ELI10: Previously'); },
|
||||
(q: Question) => { q.question = q.question.replace('ELI10: Right now', 'ELI10: If approved, right now'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^ELI10: (.+)$/m, '> ELI10: $1'); },
|
||||
(q: Question) => { q.question = q.question.replace(/^ELI10: (.+)$/m, 'ELI10: "$1"'); },
|
||||
(q: Question) => { q.question = q.question.replace(/(?:all )?look (?:the same|identical)/, 'do not look identical'); },
|
||||
(q: Question) => { q.question += '\nELI10: No current gap remains.'; },
|
||||
(q: Question) => { q.question += '\nCorrection: this gap is already resolved.'; },
|
||||
(q: Question) => { q.question = q.question.replace('Right now Save,', 'Right now Publish,'); },
|
||||
]) expect(accepted(change(index, edit))).toBe(false);
|
||||
});
|
||||
|
||||
test('the current named correction and distinct unresolved choice must both be present', () => {
|
||||
for (const index of [0, 1]) for (const edit of [
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('✅ Save', '✅ Publish'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace('; Reset', '; Save, Reset'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace(/filled/g, 'outlined'); },
|
||||
(q: Question) => { q.options[0]!.description = q.options[0]!.description!.replace(/DESIGN\.md/g, 'ARCHIVED.md'); },
|
||||
(q: Question) => { q.options[2]!.label = '1C Choose another workflow'; },
|
||||
(q: Question) => { q.options[2]!.description = 'All buttons already comply; no violation remains.'; },
|
||||
(q: Question) => { q.options[2]!.description += '\nCorrection: the violation is now closed.'; },
|
||||
]) expect(accepted(change(index, edit))).toBe(false);
|
||||
for (const index of [0, 1]) for (const option of [0, 2]) for (const prefix of ['Historical source excerpt: ', 'If approved later: ', '> ', 'Do not apply: ']) {
|
||||
expect(accepted(change(index, q => { q.options[option]!.description = prefix + q.options[option]!.description; }))).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('owned current withdrawal overrides earlier assertions while quoted history does not', () => {
|
||||
for (const index of [0, 1]) for (const target of [-1, 0, 2]) {
|
||||
for (const suffix of ['\nThis finding is withdrawn.', '\nThis finding is "no longer current".', '\nThis finding is \'withdrawn\'.', '\nThis finding is ‘no longer current’.', '\nThis finding is `no longer current`.', '\nAssessment complete; This finding is withdrawn.', '\nAssessment complete; This finding is \'no longer current\'.', '\nCorrection: this gap is already resolved.', '\nProvided approval, apply this amendment.', '\nOnce approved, apply this amendment.']) {
|
||||
expect(accepted(change(index, q => { if (target < 0) q.question += suffix; else q.options[target]!.description += suffix; }))).toBe(false);
|
||||
}
|
||||
for (const suffix of [' Prior note: "This finding is withdrawn."', '\n> This amendment is withdrawn.', ' Earlier review said `This finding is withdrawn.`', '\nIf a user scans the header, Save remains easiest to find.']) {
|
||||
expect(accepted(change(index, q => { if (target < 0) q.question += suffix; else q.options[target]!.description += suffix; }))).toBe(true);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('new public fixture and regression tests select the Design finding-count workflow only', () => {
|
||||
for (const dependency of ['test/design-primary-emphasis-av.test.ts', 'test/fixtures/design-primary-emphasis-av-calls.json']) {
|
||||
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes(dependency)).map(([name]) => name)).toEqual(['plan-design-finding-count']);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,138 +0,0 @@
|
||||
import {describe,expect,test} from 'bun:test';
|
||||
import {nativePlanCallFingerprint,planCountQuestionPhase,designStep0Boundary} from './helpers/claude-pty-runner';
|
||||
import {isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff} from './helpers/design-count-review';
|
||||
import {isDesignArtifactGeneration} from './helpers/design-artifact-question';
|
||||
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/design-primary-group-as-calls.json';
|
||||
|
||||
const calls=()=>structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const first=()=>calls()[1]!;
|
||||
const fp=(c:NativePlanQuestionCall)=>nativePlanCallFingerprint(c,0,true);
|
||||
const reanswer=(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:c.questions[0]!.options[0]!.label};return c;};
|
||||
const accepted=(c:NativePlanQuestionCall)=>isDesignCountFirstReview(fp(c));
|
||||
|
||||
describe('Design primary action named by the Issue header',()=>{
|
||||
test('exact eight native calls retain the three setup questions in one call and the six issues plus TODO',()=>{
|
||||
const input=calls(), before=JSON.stringify(input);
|
||||
expect(input.map(c=>c.questions.length)).toEqual([3,1,1,1,1,1,1,1]);
|
||||
expect(input.flatMap(c=>c.questions)).toHaveLength(10);
|
||||
let started=false; const phases=input.map(call=>{
|
||||
const phase=planCountQuestionPhase(fp(call),started,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff,isDesignArtifactGeneration);
|
||||
started=phase.reviewStarted;return phase;
|
||||
});
|
||||
expect(phases.map(p=>p.preReview)).toEqual([true,false,false,false,false,false,false,false]);
|
||||
expect(phases.filter(p=>p.administrative)).toHaveLength(0);
|
||||
expect(input.map(accepted)).toEqual([false,true,false,false,false,false,false,false]);
|
||||
expect(phases.filter(p=>!p.preReview)).toHaveLength(7);
|
||||
expect(phases.filter(p=>p.preReview)).toHaveLength(1);
|
||||
expect(JSON.stringify(input)).toBe(before);
|
||||
});
|
||||
test('finding annotations, control names, palette and option order do not supply or restrict identity',()=>{
|
||||
for(const annotation of [' (F1)',' (F27)','']){
|
||||
const c=first(),q=c.questions[0]!;
|
||||
q.question=q.question.replace(' (F1)',annotation);
|
||||
expect(accepted(reanswer(c))).toBe(true);
|
||||
}
|
||||
const c=JSON.parse(JSON.stringify(first()).replaceAll('Save','Publish').replaceAll('#1d4ed8','#123abc')) as NativePlanQuestionCall;
|
||||
const q=c.questions[0]!;
|
||||
q.header='Issue 8: Publish';
|
||||
q.question=q.question.replace('D4 — Issue 1 (F1)','D31 — Issue 8 (F12)').replace(/\b1([ABC])\b/g,'8$1');
|
||||
for(const o of q.options)o.label=o.label.replace(/^1/,'8');
|
||||
q.options.reverse();
|
||||
for(const o of q.options){c.answers={[q.question]:o.label};expect(accepted(c)).toBe(true);}
|
||||
q.options.reverse();q.options=q.options.filter(o=>!o.label.startsWith('8B:'));expect(accepted(reanswer(c))).toBe(true);
|
||||
});
|
||||
test('the separately completed retry retains its existing two setup and six review calls',()=>{
|
||||
const input=structuredClone(captured.retryCalls) as NativePlanQuestionCall[];
|
||||
let started=false;const phases=input.map(call=>{
|
||||
const phase=planCountQuestionPhase(fp(call),started,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff,isDesignArtifactGeneration);
|
||||
started=phase.reviewStarted;return phase;
|
||||
});
|
||||
expect(input).toHaveLength(8);
|
||||
expect(phases.map(p=>p.preReview)).toEqual([true,true,false,false,false,false,false,false]);
|
||||
expect(phases.filter(p=>p.administrative)).toHaveLength(0);
|
||||
});
|
||||
test('primary and ghost controls, documented count, issue identity and offered choice identities remain bound',()=>{
|
||||
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
|
||||
c=>{c.questions[0]!.header='Issue 1';},
|
||||
c=>{c.questions[0]!.header='Issue 1: Export';},
|
||||
c=>{c.questions[0]!.header='Issue 2: Save';},
|
||||
c=>{c.questions[0]!.question=c.questions[0]!.question.replace('other three','other two');},
|
||||
c=>{c.questions[0]!.question=c.questions[0]!.question.replace('Reset, Cancel and Export','Reset, Reset and Export');},
|
||||
c=>{c.questions[0]!.question=c.questions[0]!.question.replace('Reset, Cancel and Export','Reset, Save and Export');},
|
||||
c=>{c.questions[0]!.options[0]!.description=c.questions[0]!.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel, Delete');},
|
||||
c=>{c.questions[0]!.options[0]!.description=c.questions[0]!.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel, Save');},
|
||||
c=>{c.questions[0]!.options[0]!.description=c.questions[0]!.options[0]!.description!.replace('Reset, Cancel, Export','Reset, Cancel, Export, Export');},
|
||||
c=>{c.questions[0]!.options[2]!.label='1C: Keep all three identical';},
|
||||
c=>{c.questions[0]!.options[2]!.label='2C: Keep all four identical';},
|
||||
];
|
||||
for(const change of changes){const c=first();change(c);expect(accepted(reanswer(c))).toBe(false);}
|
||||
});
|
||||
test('readiness, focus, navigation, source-only questions and naked F labels cannot begin review',()=>{
|
||||
for(const title of [
|
||||
'D4 — Issue 1 (F1): Ready to review the header action group?',
|
||||
'D4 — Issue 1 (F1): Which design source should the reviewer use?',
|
||||
'D4 — Issue 1 (F1): Fix the primary action?',
|
||||
'D4 — Issue 1 (F1): How should the header action group establish the primary action? Ready?',
|
||||
'Example: D4 — Issue 1 (F1): How should the header action group establish the primary action?',
|
||||
'> D4 — Issue 1 (F1): How should the header action group establish the primary action?',
|
||||
]){const c=first(),q=c.questions[0]!;q.question=title+'\n'+q.question.split('\n').slice(1).join('\n');expect(accepted(reanswer(c))).toBe(false);}
|
||||
for(const header of ['Focus','Routing','Next steps','Outside voices']){const c=first();c.questions[0]!.header=header;expect(accepted(c)).toBe(false);}
|
||||
const c=first();c.questions[0]!.options=[{label:'Start the review'},{label:'Wait'}];expect(accepted(reanswer(c))).toBe(false);
|
||||
});
|
||||
test('current gap and contract cannot be replaced by quoted, historical or conditional material',()=>{
|
||||
for(const prefix of ['Historical example: ','Hypothetical example: ','Quoted assessment: ','Source example: ','If approved, ','When approved, ','Unless rejected, ','Assuming approval, ','Provided approval, ']){
|
||||
const c=first();c.questions[0]!.question=c.questions[0]!.question.replace('ELI10: ','ELI10: '+prefix);expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
for(const transform of [(s:string)=>'"'+s+'"',(s:string)=>'> '+s,(s:string)=>' '+s,(s:string)=>'```\n'+s+'\n```']){
|
||||
const c=first(),q=c.questions[0]!;q.question=q.question.split('\n').map(line=>line.startsWith('ELI10:')?transform(line):line).join('\n');expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
for(const suffix of [' This finding is no longer current.',' This finding is "no longer current".',' This requirement is withdrawn.',' This contract is "withdrawn".',' This gap is now resolved.',' This issue is superseded.']){
|
||||
const c=first();c.questions[0]!.question+=suffix;expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('each offered amendment and deferral must remain current and unconditional',()=>{
|
||||
for(const index of [0,2])for(const prefix of ['Historical example: ','Source example: ','Assuming approval, ','Provided approval, ','✅ Assuming approval, ','✅ Provided approval, ']){
|
||||
const c=first(),o=c.questions[0]!.options[index]!;o.description=prefix+o.description;expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
for(const index of [0,2])for(const suffix of [' This finding is no longer current.',' This amendment is "withdrawn".',' This deferral is rejected.',' This choice is superseded.',' This gap is closed.',' This contract is withdrawn.',' This requirement is "no longer current".',' Assuming approval, this is proposed only.',' Provided approval, this will become current.']){
|
||||
const c=first();c.questions[0]!.options[index]!.description+=suffix;expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
for(const suffix of [' These tokens are withdrawn.',' These styles are "no longer current".',' Do not apply these tokens.']){
|
||||
const c=first();c.questions[0]!.options[0]!.description+=suffix;expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
const c=first();c.questions[0]!.options[2]!.description+=' Do not keep all four buttons identical.';expect(accepted(reanswer(c))).toBe(false);
|
||||
});
|
||||
test('quoted past statuses do not erase the current finding, style or opposed choice',()=>{
|
||||
for(const target of [-1,0,2])for(const history of [' The prior review said "This finding is no longer current."'," The prior review said 'This finding is withdrawn.'",' The prior review said ‘This finding is no longer current.’',' The prior review said "Estimate (human: ~1h / CC: ~5min) This finding is no longer current."',' The prior review said `This finding is no longer current.`',' The earlier decision was `no longer current`.','\n> This amendment is withdrawn.']){
|
||||
const c=first();if(target<0)c.questions[0]!.question+=history;else c.questions[0]!.options[target]!.description+=history;
|
||||
expect(accepted(reanswer(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('current status scalars retain their subjects across quote styles and semicolon boundaries',()=>{
|
||||
for(const target of [-1,0,2])for(const subject of ['finding','amendment','contract'])for(const status of ['withdrawn','no longer current'])for(const quote of ['',"'","‘",'"','“','`'])for(const boundary of [' ','; ']){
|
||||
const closing=quote==='‘'?'’':quote==='“'?'”':quote;
|
||||
const suffix=boundary+'This '+subject+' is '+quote+status+closing+'.';
|
||||
const c=first();if(target<0)c.questions[0]!.question+=suffix;else c.questions[0]!.options[target]!.description+=suffix;
|
||||
expect(accepted(reanswer(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('completed native ownership, answer membership, one question and exact option indices are required',()=>{
|
||||
const changes:Array<(c:NativePlanQuestionCall)=>void>=[
|
||||
c=>{c.answered=false;},c=>{delete (c as Partial<NativePlanQuestionCall>).answered;},
|
||||
c=>{c.failed=true;},c=>{delete c.failed;},c=>{c.sessionId='';},c=>{c.toolUseId='';},
|
||||
c=>{c.answers={};},c=>{c.answers={[c.questions[0]!.question]:'not offered'};},
|
||||
c=>{delete c.answeredAt;},c=>{c.answeredAt='invalid';},c=>{delete c.unansweredQuestionIndices;},c=>{c.unansweredQuestionIndices=[0];},
|
||||
c=>{c.questions[0]!.multiSelect=true;},c=>{c.questions.push(structuredClone(c.questions[0]!));},
|
||||
c=>{c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!));},
|
||||
];
|
||||
for(const change of changes){const c=first();change(c);expect(accepted(c)).toBe(false);}
|
||||
for(const mutate of [
|
||||
(f:ReturnType<typeof fp>)=>{f.signature='foreign';},
|
||||
(f:ReturnType<typeof fp>)=>{f.nativeQuestionIndex=1;},
|
||||
(f:ReturnType<typeof fp>)=>{f.options=[];},
|
||||
(f:ReturnType<typeof fp>)=>{f.options[0]!.index=2;},
|
||||
(f:ReturnType<typeof fp>)=>{f.options[0]!.label='unrelated';},
|
||||
]){const f=fp(first());mutate(f);expect(isDesignCountFirstReview(f)).toBe(false);}
|
||||
});
|
||||
});
|
||||
@@ -1,63 +0,0 @@
|
||||
import {describe,expect,test} from 'bun:test';
|
||||
import {nativePlanCallFingerprint,planCountQuestionPhase,designStep0Boundary} from './helpers/claude-pty-runner';
|
||||
import {isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff} from './helpers/design-count-review';
|
||||
import {isDesignArtifactGeneration} from './helpers/design-artifact-question';
|
||||
import type {NativePlanQuestionCall} from './helpers/plan-count-transcript';
|
||||
import actual from './fixtures/design-primary-header-aq.json';
|
||||
|
||||
const calls=()=>structuredClone(actual.calls) as NativePlanQuestionCall[];
|
||||
const first=()=>calls()[2]!;
|
||||
const fp=(call=first())=>nativePlanCallFingerprint(call,244227,true);
|
||||
const classify=(call=first())=>isDesignCountFirstReview(fp(call));
|
||||
function mutate(fn:(call:NativePlanQuestionCall)=>void){const c=first();fn(c);return c;}
|
||||
function text(change:(s:string)=>string){return mutate(c=>{const q=c.questions[0]!,answer=c.answers![q.question]!;q.question=change(q.question);c.answers={[q.question]:answer};});}
|
||||
|
||||
describe('AQ current primary-header amendment starts Design review',()=>{
|
||||
test('exact owned four-call prefix starts on Issue 1 with all question bytes unchanged',()=>{
|
||||
let started=false;const phases=calls().map(c=>{const p=planCountQuestionPhase(fp(c),started,designStep0Boundary,isDesignCountFirstReview,isDesignCountSetup,isDesignCompletionHandoff,isDesignArtifactGeneration);started=p.reviewStarted;return p;});
|
||||
expect(phases.map(p=>p.preReview)).toEqual([true,true,false,false]);
|
||||
expect(phases.every(p=>!p.administrative)).toBe(true);
|
||||
expect(classify()).toBe(true);expect(isDesignCountSetup(fp())).toBe(false);expect(isDesignCompletionHandoff(fp())).toBe(false);
|
||||
});
|
||||
test('consistent control, palette, decision and issue identities can vary',()=>{
|
||||
const c=first(),q=c.questions[0]!;q.question=q.question.replaceAll('Save','Submit').replaceAll('#1d4ed8','#123abc').replace('D3 — Issue 1:','D9 — Issue 4:').replaceAll('1A','4A').replaceAll('1B','4B').replaceAll('1C','4C');q.header='Issue 4';
|
||||
for(const o of q.options){o.label=o.label.replaceAll('Save','Submit').replace(/^1/,'4');o.description=o.description?.replaceAll('Save','Submit').replaceAll('#1d4ed8','#123abc');}
|
||||
q.options.reverse();for(const o of q.options){c.answers={[q.question]:o.label};expect(classify(c)).toBe(true);}
|
||||
});
|
||||
test('only and single primary-header descriptions retain the same current action',()=>{
|
||||
for(const title of ['make Save the single visually primary header action?','make Save the only primary header action?','make Save the single primary header action?','Make Save the only visually primary action in the header?'])expect(classify(text(s=>s.replace('make Save the only visually primary header action?',title)))).toBe(true);
|
||||
});
|
||||
test('source, historical, conditional and noncurrent assessments cannot start review',()=>{
|
||||
for(const prefix of ['Source excerpt: ','Earlier review assessment: ','If approved, ','For historical context, ','Hypothetical example: '])expect(classify(text(s=>s.replace('ELI10: ','ELI10: '+prefix)))).toBe(false);
|
||||
for(const heading of ['Source excerpt:','Earlier review assessment:','If approved later:'])expect(classify(text(s=>s.replace('ELI10:',heading+'\nELI10:')))).toBe(false);
|
||||
for(const suffix of [' This finding is withdrawn.',' This amendment is "closed".',' This remedy is a historical example, not the current option.',' Correction: this finding is not current.',' This issue is superseded.',' This issue is \"superseded\".'])expect(classify(text(s=>s+suffix))).toBe(false);
|
||||
});
|
||||
test('current context and assessment owners must be unique',()=>{
|
||||
for(const insertion of ['Project/branch/task: other, another project with an archived design.','ELI10: Right now Save, Reset, Cancel and Export look identical.'])expect(classify(text(s=>s.replace('ELI10:',insertion+'\nELI10:')))).toBe(false);
|
||||
expect(classify(text(s=>s.replace(/^Project\/branch\/task:.*\n/m,'')))).toBe(false);
|
||||
for(const frame of ['If approved,','Provided approval,','Assuming approval,','Earlier review assessment:'])expect(classify(text(s=>s.replace('Project/branch/task: main','Project/branch/task: '+frame+' main')))).toBe(false);
|
||||
});
|
||||
test('native identity, completion, selected answer and original displayed options stay required',()=>{
|
||||
for(const change of [
|
||||
(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},
|
||||
(c:NativePlanQuestionCall)=>{delete c.answeredAt;},(c:NativePlanQuestionCall)=>{c.answeredAt='invalid';},(c:NativePlanQuestionCall)=>{c.sessionId='';},(c:NativePlanQuestionCall)=>{c.toolUseId='';},
|
||||
(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered'};},
|
||||
(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));}
|
||||
])expect(classify(mutate(change))).toBe(false);
|
||||
for(const f of [{...fp(),signature:'foreign:call'},{...fp(),nativeQuestionIndex:1},{...fp(),nativeCall:undefined},{...fp(),options:[...fp().options].reverse()}])expect(isDesignCountFirstReview(f)).toBe(false);
|
||||
});
|
||||
test('explicit issue identities and setup-only action labels cannot grant review',()=>{
|
||||
for(const c of [mutate(c=>{c.questions[0]!.header='Issue 2';}),mutate(c=>{c.questions[0]!.header='Routing';}),text(s=>s.replace('Issue 1:','Issue 01:')),text(s=>s.replace('D3 —','D03 —')),text(s=>s.replace('make Save the only visually primary header action?','start reviewing the header?')),mutate(c=>{c.questions[0]!.options[0]!.label='Start review';c.answers={[c.questions[0]!.question]:'Start review'};})])expect(classify(c)).toBe(false);
|
||||
});
|
||||
test('the original gap, exact named remedy, and an opposed retained violation are all required',()=>{
|
||||
expect(classify(text(s=>s.replace('look identical','no longer look identical')))).toBe(false);
|
||||
for(const body of ['Source excerpt: Apply DESIGN.md: Save #1d4ed8 filled white text; Reset, Cancel, Export neutral ghost buttons.','If approved, Apply DESIGN.md: Save #1d4ed8 filled white text; Reset, Cancel, Export neutral ghost buttons.','Apply DESIGN.md: Publish #1d4ed8 filled white text; Reset, Cancel, Export neutral ghost buttons.','Apply DESIGN.md: Save filled; Reset, Cancel, Export ghost.'])expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description=body;}))).toBe(false);
|
||||
for(const suffix of [' This amendment is withdrawn.',' This amendment is "closed".',' This remedy is a historical example, not the current option.',' Do not apply these tokens.',' This option is superseded.',' This option is \"superseded\".'])expect(classify(mutate(c=>{c.questions[0]!.options[0]!.description+=suffix;}))).toBe(false);
|
||||
for(const body of ['The design is accepted.','Source excerpt: Decline the fix; document the violation as accepted.','If approved, decline the fix; document the violation as accepted.','Decline the fix; document the violation as accepted. This deferral is withdrawn.','Decline the fix; document the violation as accepted. This deferral is superseded.','Decline the fix; document the violation as accepted. This deferral is \"superseded\".'])expect(classify(mutate(c=>{c.questions[0]!.options[2]!.description=body;}))).toBe(false);
|
||||
expect(classify(mutate(c=>{c.questions[0]!.options[2]!.label='Proceed with review';}))).toBe(false);
|
||||
});
|
||||
test('a wholly quoted archival note does not withdraw the current owned decision',()=>{
|
||||
expect(classify(text(s=>s+'\n"Earlier review assessment: This finding is withdrawn."'))).toBe(true);
|
||||
});
|
||||
});
|
||||
@@ -1,125 +0,0 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import captured from './fixtures/design-primary-treatment-ao.json';
|
||||
import { isDesignCountFirstReview, isDesignCountSetup } from './helpers/design-count-review';
|
||||
import { designStep0Boundary, planCountQuestionPhase } from './helpers/claude-pty-runner';
|
||||
import type { AskUserQuestionFingerprint as Fingerprint } from './helpers/claude-pty-runner';
|
||||
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
|
||||
|
||||
type Question = NonNullable<Fingerprint['nativeCall']>['questions'][number];
|
||||
const original = () => structuredClone(captured.fingerprints[0]) as Fingerprint;
|
||||
function edit(change: (q: Question) => void): Fingerprint {
|
||||
const fp = original(), call = fp.nativeCall!, q = call.questions[0]!;
|
||||
const selected = q.options.findIndex(o => o.label === call.answers![q.question]);
|
||||
change(q);
|
||||
call.answers = { [q.question]: q.options[selected]!.label };
|
||||
fp.options = q.options.map((o, i) => ({ index: i + 1, label: o.label }));
|
||||
return fp;
|
||||
}
|
||||
|
||||
test('the exact first primary-treatment decision starts review and retains the following decision', () => {
|
||||
expect(isDesignCountFirstReview(original())).toBe(true);
|
||||
let started = false;
|
||||
const phases = captured.fingerprints.map(raw => {
|
||||
const phase = planCountQuestionPhase(raw as Fingerprint, started, designStep0Boundary,
|
||||
isDesignCountFirstReview, isDesignCountSetup);
|
||||
started = phase.reviewStarted;
|
||||
return phase;
|
||||
});
|
||||
expect(phases).toEqual([
|
||||
{ preReview: false, reviewStarted: true },
|
||||
{ preReview: false, reviewStarted: true },
|
||||
]);
|
||||
expect(captured.fingerprints.map(fp => fp.preReview)).toEqual([true, true]);
|
||||
});
|
||||
|
||||
test('equivalent primary qualifiers, singular filled treatment and authority compose', () => {
|
||||
for (const qualifier of ['visible', 'visually', 'single filled']) {
|
||||
for (const treatment of ['one filled button', 'single filled primary', 'one filled primary button']) {
|
||||
for (const authority of ['per DESIGN.md', 'exactly as DESIGN.md specifies']) {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.question = q.question.replace('visually primary', `${qualifier} primary`);
|
||||
q.options[0]!.description = q.options[0]!.description!
|
||||
.replace('one filled button', treatment).replace('per DESIGN.md', authority);
|
||||
}))).toBe(true);
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('a different control or an opposed answer preserves the current decision', () => {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.question = q.question.replaceAll('Save', 'Submit');
|
||||
q.options = q.options.map(o => ({ label: o.label.replaceAll('Save', 'Submit'),
|
||||
description: o.description?.replaceAll('Save', 'Submit') }));
|
||||
}))).toBe(true);
|
||||
for (const option of original().nativeCall!.questions[0]!.options) {
|
||||
const fp = original(), call = fp.nativeCall!;
|
||||
call.answers = { [call.questions[0]!.question]: option.label };
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
const rejected: Array<[string, (q: Question) => void]> = [
|
||||
['foreign header', q => { q.header = 'Issue 2'; }],
|
||||
['workflow title', q => { q.question = q.question.replace('Make Save the visually primary action in the header', 'Run outside design voices'); }],
|
||||
['historical question', q => { q.question = 'Historical example:\n' + q.question; }],
|
||||
['source assessment', q => { q.question = q.question.replace('\nELI10:', '\nSource excerpt:\nELI10:'); }],
|
||||
['quoted assessment', q => { q.question = q.question.replace('\nELI10:', '\n> ELI10:'); }],
|
||||
['conditional assessment', q => { q.question = q.question.replace('ELI10: Right now', 'ELI10: If right now'); }],
|
||||
['no current equal-weight gap', q => { q.question = q.question.replace('all look identical', 'do not look identical'); }],
|
||||
['withdrawn contract', q => { q.question += '\nThis DESIGN.md contract is withdrawn.'; }],
|
||||
['quoted withdrawn contract', q => { q.question += '\nThis DESIGN.md contract is "withdrawn".'; }],
|
||||
['superseded requirement', q => { q.question += '\nThis requirement is superseded.'; }],
|
||||
['quoted superseded requirement', q => { q.question += '\nThis requirement is \"superseded\".'; }],
|
||||
['rejected contract', q => { q.question += '\nThis DESIGN.md contract is rejected.'; }],
|
||||
['quoted cancelled contract', q => { q.question += '\nThis DESIGN.md contract is \"cancelled\".'; }],
|
||||
['withdrawn issue', q => { q.question += '\nThis issue is withdrawn.'; }],
|
||||
['quoted rejected issue', q => { q.question += '\nThis issue is "rejected".'; }],
|
||||
['wrong named primary', q => { q.options[0]!.description = q.options[0]!.description!.replace('Save becomes', 'Reset becomes'); }],
|
||||
['primary also ghost', q => { q.options[0]!.description = q.options[0]!.description!.replace('; Reset,', '; Save,'); }],
|
||||
['missing foreground', q => { q.options[0]!.description = q.options[0]!.description!.replace(', white text', ''); }],
|
||||
['missing ghost treatment', q => { q.options[0]!.description = q.options[0]!.description!.replace('neutral ghost buttons', 'filled buttons'); }],
|
||||
['missing style authority', q => { q.options[0]!.description = q.options[0]!.description!.replace('per DESIGN.md', 'per a future proposal'); }],
|
||||
['conditional amendment', q => { q.options[0]!.description = 'If approved later: ' + q.options[0]!.description; }],
|
||||
['quoted amendment', q => { q.options[0]!.description = '> ' + q.options[0]!.description; }],
|
||||
['cancelled amendment', q => { q.options[0]!.description += '\nCorrection: do not apply these styles.'; }],
|
||||
['quoted rejected amendment', q => { q.options[0]!.description += '\nThis amendment is "rejected".'; }],
|
||||
['no opposed choice', q => { q.options[2]!.label = '1C Configure Export'; }],
|
||||
['opposed gap closed', q => { q.options[2]!.description = q.options[2]!.description!.replace('gap stays open', 'gap is closed'); }],
|
||||
['historical opposed choice', q => { q.options[2]!.description = 'Historical example: ' + q.options[2]!.description; }],
|
||||
['conditional opposed choice', q => { q.options[2]!.description = 'If approved later: ' + q.options[2]!.description; }],
|
||||
['quoted opposed choice', q => { q.options[2]!.description = '> ' + q.options[2]!.description; }],
|
||||
['resolved gap', q => { q.options[2]!.description += '\nThe gap is now resolved.'; }],
|
||||
['quoted rejected opposed choice', q => { q.options[2]!.description += '\nThis option is "rejected".'; }],
|
||||
];
|
||||
test.each(rejected)('%s does not start review', (_, change) => {
|
||||
expect(isDesignCountFirstReview(edit(change))).toBe(false);
|
||||
});
|
||||
|
||||
test('a wholly quoted historical cancellation does not withdraw this requirement', () => {
|
||||
expect(isDesignCountFirstReview(edit(q => {
|
||||
q.question += '\nHistorical note: "This DESIGN.md contract is withdrawn."';
|
||||
}))).toBe(true);
|
||||
});
|
||||
|
||||
test('recognition requires the same completed native identity and offered answer', () => {
|
||||
for (const change of [
|
||||
(fp: Fingerprint) => { fp.nativeCall!.answered = false; },
|
||||
(fp: Fingerprint) => { fp.nativeCall!.failed = true; },
|
||||
(fp: Fingerprint) => { fp.signature = 'foreign:tool'; },
|
||||
(fp: Fingerprint) => { fp.nativeQuestionIndex = 1; },
|
||||
(fp: Fingerprint) => { fp.nativeCall!.unansweredQuestionIndices = [0]; },
|
||||
(fp: Fingerprint) => { delete fp.nativeCall!.answeredAt; },
|
||||
(fp: Fingerprint) => { fp.nativeCall!.answers = {}; },
|
||||
(fp: Fingerprint) => { fp.options.reverse(); },
|
||||
]) {
|
||||
const fp = original(); change(fp);
|
||||
expect(isDesignCountFirstReview(fp)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('both source regressions select the existing Design workflow owner', () => {
|
||||
for (const file of ['test/design-primary-treatment-ao.test.ts', 'test/fixtures/design-primary-treatment-ao.json']) {
|
||||
expect(selectTests([file], E2E_TOUCHFILES, []).selected).toEqual(['plan-design-finding-count']);
|
||||
}
|
||||
});
|
||||
@@ -1,122 +0,0 @@
|
||||
import {expect, test} from 'bun:test';
|
||||
import fixture from './fixtures/design-variant-choice-am.json';
|
||||
import retry from './fixtures/design-variant-choice-am-retry.json';
|
||||
import {isDesignCountFirstReview} from './helpers/design-count-review';
|
||||
import type {AskUserQuestionFingerprint as FP} from './helpers/claude-pty-runner';
|
||||
type Q=NonNullable<FP['nativeCall']>['questions'][number];
|
||||
const original=()=>structuredClone(fixture.fingerprint) as unknown as FP;
|
||||
function edit(change:(q:Q,fp:FP)=>void):FP {
|
||||
const fp=original(),c=fp.nativeCall!,q=c.questions[0]!,chosen=q.options.findIndex(o=>o.label===c.answers![q.question]);
|
||||
change(q,fp);c.answers={[q.question]:q.options[chosen]!.label};fp.options=q.options.map((o,i)=>({index:i+1,label:o.label}));return fp;
|
||||
}
|
||||
test('exact completed primary choice binds the current token contract to existing component variants',()=>expect(isDesignCountFirstReview(original())).toBe(true));
|
||||
const yes:Array<[string,(q:Q,fp:FP)=>void]>=[
|
||||
['renamed primary',q=>{q.question=q.question.replaceAll('Save','Submit');q.options=q.options.map(o=>({...o,label:o.label.replaceAll('Save','Submit'),description:o.description?.replaceAll('Save','Submit')}));}],
|
||||
['another prescribed color and foreground',q=>{q.question=q.question.replace('#1d4ed8 with white text, about 6.7:1 contrast','#ffcc22 with black text');}],
|
||||
['unqualified action position',q=>{q.question=q.question.replace('single primary action in the header?','single primary action?');}],
|
||||
['one benefit sufficient',q=>{q.options[0]!.description=q.options[0]!.description!.split('\n').slice(1).join('\n');}],
|
||||
['an existing open-gap deferral',q=>{q.options[2]!.description='Leaves a documented DESIGN.md violation in place.';}],
|
||||
['quoted historical note does not cancel current choice',q=>{q.options[0]!.description+=' Prior note: "This amendment is withdrawn."';}],
|
||||
];
|
||||
test.each(yes)('%s preserves current owned review',(_,change)=>expect(isDesignCountFirstReview(edit(change))).toBe(true));
|
||||
const no:Array<[string,(q:Q,fp:FP)=>void]>=[
|
||||
['proposed token contract only',q=>{q.question=q.question.replace('DESIGN.md already says','A proposed example follows. DESIGN.md already says');}],
|
||||
['withdrawn token requirement',q=>{q.question=q.question.replace('This is Design Principle 2:','Correction: this DESIGN.md requirement is withdrawn. This is Design Principle 2:');}],
|
||||
['superseded token contract',q=>{q.question=q.question.replace('This is Design Principle 2:','That token contract is no longer current. This is Design Principle 2:');}],
|
||||
['variant amendment cancelled directly',q=>{q.options[0]!.description+=' Correction: do not use the primary and ghost variants.';}],
|
||||
['variant amendment contradicts its remedy',q=>{q.options[0]!.description+=' The current amendment keeps all four buttons identical.';}],
|
||||
['failed call',(_,f)=>{f.nativeCall!.failed=true;}],
|
||||
['unanswered call',(_,f)=>{f.nativeCall!.answered=false;}],
|
||||
['unbound call',(_,f)=>{f.signature='other:call';}],
|
||||
['missing completion time',(_,f)=>{delete f.nativeCall!.answeredAt;}],
|
||||
['wrong question index',(_,f)=>{f.nativeQuestionIndex=1;}],
|
||||
['unanswered member',(_,f)=>{f.nativeCall!.unansweredQuestionIndices=[0];}],
|
||||
['wrong header identity',q=>{q.header='Issue 2';}],
|
||||
['wrong choice identity',q=>{q.options[0]!.label=q.options[0]!.label.replace('1A','2A');}],
|
||||
['setup framing',q=>{q.header='Routing';}],
|
||||
['source-framed question',q=>{q.question='Historical example:\n'+q.question;}],
|
||||
['historical assessment',q=>{q.question=q.question.replace('ELI10: Right now','ELI10: Previously');}],
|
||||
['conditional assessment',q=>{q.question=q.question.replace('ELI10: Right now','ELI10: If right now');}],
|
||||
['quoted assessment',q=>{q.question=q.question.replace('ELI10:','> ELI10:');}],
|
||||
['negated equality',q=>{q.question=q.question.replace('all look identical','do not look identical');}],
|
||||
['archived-only equality',q=>{q.question=q.question.replace('all look identical.','all look identical only in an archived screenshot.');}],
|
||||
['absent token contract',q=>{q.question=q.question.replace('DESIGN.md already says','Archived notes say');}],
|
||||
['wrong primary contract',q=>{q.question=q.question.replace('says Save is','says Export is');}],
|
||||
['no fill contract',q=>{q.question=q.question.replace('only filled button','outlined button');}],
|
||||
['no foreground contract',q=>{q.question=q.question.replace('with white text','with unknown text');}],
|
||||
['no ghost contract',q=>{q.question=q.question.replace('neutral ghost buttons','also filled buttons');}],
|
||||
['quoted token contract',q=>{q.question=q.question.replace('DESIGN.md already says','"DESIGN.md already says').replace('neutral ghost buttons.','neutral ghost buttons."');}],
|
||||
['conditional token contract',q=>{q.question=q.question.replace('DESIGN.md already says','If DESIGN.md already says');}],
|
||||
['resolved current gap',q=>{q.question+='\nCorrection: the gap is already resolved.';}],
|
||||
['wrong proposed primary',q=>{q.options[0]!.label=q.options[0]!.label.replace('Filled Save','Filled Reset');}],
|
||||
['wrong proposed ghost role',q=>{q.options[0]!.label=q.options[0]!.label.replace('ghost others','filled others');}],
|
||||
['no variant authority',q=>{q.options[0]!.description=q.options[0]!.description!.replace('from DESIGN.md','from an archived example');}],
|
||||
['no variant amendment',q=>{q.options[0]!.description=q.options[0]!.description!.replace('Uses the existing','Mentions the existing');}],
|
||||
['quoted variant amendment',q=>{q.options[0]!.description=q.options[0]!.description!.replace('✅ Uses','> ✅ Uses');}],
|
||||
['conditional variant amendment',q=>{q.options[0]!.description='If approved later:\n'+q.options[0]!.description;}],
|
||||
['conditional benefit prefix',q=>{q.options[0]!.description=q.options[0]!.description!.replace('✅ Save reads','✅ If Save reads');}],
|
||||
['historical variant amendment',q=>{q.options[0]!.description+=' This is a historical example, not the current amendment.';}],
|
||||
['withdrawn amendment',q=>{q.options[0]!.description+=' This amendment is withdrawn.';}],
|
||||
['cancelled style',q=>{q.options[0]!.description+=' Correction: do not apply these styles.';}],
|
||||
['no opposed choice',q=>{q.options[2]!.label='1C Export preferences';}],
|
||||
['deferral no longer retains violation',q=>{q.options[2]!.description=q.options[2]!.description!.replace('Violates DESIGN.md','Matches DESIGN.md');}],
|
||||
['historical deferral',q=>{q.options[2]!.description='Historical source excerpt:\n'+q.options[2]!.description;}],
|
||||
['conditional deferral',q=>{q.options[2]!.description='If accepted later:\n'+q.options[2]!.description;}],
|
||||
['closed deferral',q=>{q.options[2]!.description+=' Correction: the violation is now closed.';}],
|
||||
];
|
||||
test.each(no)('%s is not current completed review evidence',(_,change)=>expect(isDesignCountFirstReview(edit(change))).toBe(false));
|
||||
|
||||
function retryEdit(change:(q:Q,fp:FP)=>void):FP {
|
||||
const fp=structuredClone(retry.fingerprint) as unknown as FP,c=fp.nativeCall!,q=c.questions[0]!;
|
||||
const chosen=q.options.findIndex(o=>o.label===c.answers![q.question]);
|
||||
change(q,fp);c.answers={[q.question]:q.options[chosen]!.label};
|
||||
fp.options=q.options.map((o,i)=>({index:i+1,label:o.label}));return fp;
|
||||
}
|
||||
test('retry first decision supplies concrete tokens in the offered label and DESIGN.md authority in its description',()=>{
|
||||
expect(isDesignCountFirstReview(retryEdit(()=>{}))).toBe(true);
|
||||
expect(retry.provenance.historicalOutcome).toBe('no_review_questions');
|
||||
});
|
||||
test('current labelled token choice permits any offered alternate and harmless historical quotes',()=>{
|
||||
for(const index of [0,1,2]){
|
||||
const fp=retryEdit(q=>{q.question+='\nArchived note: "This requirement was withdrawn."';});
|
||||
const c=fp.nativeCall!,q=c.questions[0]!;c.answers={[q.question]:q.options[index]!.label};
|
||||
expect(isDesignCountFirstReview(fp)).toBe(true);
|
||||
}
|
||||
expect(isDesignCountFirstReview(retryEdit(q=>{
|
||||
q.question=q.question.replaceAll('Save','Submit');
|
||||
q.options=q.options.map(o=>({...o,label:o.label.replaceAll('Save','Submit'),description:o.description?.replaceAll('Save','Submit')}));
|
||||
}))).toBe(true);
|
||||
});
|
||||
const retryNo:Array<[string,(q:Q,fp:FP)=>void]>=[
|
||||
['wrong offered secondary count',q=>{q.options[0]!.description=q.options[0]!.description!.replace('three neutral ghosts','two neutral ghosts');}],
|
||||
['wrong stated secondary count',q=>{q.question=q.question.replace('other three','other five');}],
|
||||
['consistent but wrong secondary counts',q=>{q.question=q.question.replace('other three','other five');q.options[0]!.description=q.options[0]!.description!.replace('three neutral ghosts','five neutral ghosts');}],
|
||||
['token authority withdrawn',q=>{q.options[0]!.description+=' Correction: these tokens do not match DESIGN.md.';}],
|
||||
['unanswered',(_,f)=>{f.nativeCall!.answered=false;}],
|
||||
['failed',(_,f)=>{f.nativeCall!.failed=true;}],
|
||||
['wrong owner',(_,f)=>{f.signature='foreign:use';}],
|
||||
['missing completion time',(_,f)=>{delete f.nativeCall!.answeredAt;}],
|
||||
['unanswered member',(_,f)=>{f.nativeCall!.unansweredQuestionIndices=[0];}],
|
||||
['wrong issue',q=>{q.header='Issue 2';}],
|
||||
['wrong option issue',q=>{q.options[0]!.label=q.options[0]!.label.replace('1A','2A');}],
|
||||
['non-design pass',q=>{q.question=q.question.replace('Visual Hierarchy','Routing');}],
|
||||
['invalid pass',q=>{q.question=q.question.replace('Pass 1,','Pass 9,');}],
|
||||
['source assessment',q=>{q.question=q.question.replace('\nELI10:','\nSource excerpt:\nELI10:');}],
|
||||
['proposed contract',q=>{q.question=q.question.replace('DESIGN.md already says','A proposed example follows. DESIGN.md already says');}],
|
||||
['withdrawn contract',q=>{q.question+=' Correction: this DESIGN.md requirement is withdrawn.';}],
|
||||
['superseded contract',q=>{q.question+=' That token contract is no longer current.';}],
|
||||
['wrong contract control',q=>{q.question=q.question.replace('says Save is','says Reset is');}],
|
||||
['not an exclusive primary',q=>{q.question=q.question.replace('only filled primary','outlined');}],
|
||||
['not ghost secondaries',q=>{q.question=q.question.replace('neutral ghost buttons','filled buttons');}],
|
||||
['wrong labelled control',q=>{q.options[0]!.label=q.options[0]!.label.replace('Save filled','Reset filled');}],
|
||||
['missing concrete color',q=>{q.options[0]!.label=q.options[0]!.label.replace('#1d4ed8','blue');}],
|
||||
['missing foreground',q=>{q.options[0]!.label=q.options[0]!.label.replace('/white','');}],
|
||||
['missing style authority',q=>{q.options[0]!.description=q.options[0]!.description!.replace('Matches DESIGN.md exactly','Matches a historical example');}],
|
||||
['quoted remedy',q=>{q.options[0]!.description='> '+q.options[0]!.description;}],
|
||||
['conditional remedy',q=>{q.options[0]!.description='If approved: '+q.options[0]!.description;}],
|
||||
['cancelled variants',q=>{q.options[0]!.description+=' Correction: do not use the primary and ghost variants.';}],
|
||||
['contradictory remedy',q=>{q.options[0]!.description+=' The current amendment keeps all four buttons identical.';}],
|
||||
['resolved deferral',q=>{q.options[2]!.description+=' This violation is now resolved.';}],
|
||||
['no remaining violation',q=>{q.options[2]!.description=q.options[2]!.description!.replace('Documented DESIGN.md violation ships','No documented violation ships');}],
|
||||
];
|
||||
test.each(retryNo)('retry %s is not positive review evidence',(_,change)=>expect(isDesignCountFirstReview(retryEdit(change))).toBe(false));
|
||||
@@ -1,116 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import { isDevexReviewIssue } from './helpers/devex-count-fixture';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import captured from './fixtures/devex-ac-first-attempt-calls.json';
|
||||
|
||||
const calls = () => structuredClone(captured) as NativePlanQuestionCall[];
|
||||
const classify = (call: NativePlanQuestionCall) => isDevexReviewIssue(nativePlanCallFingerprint(call, 0, true));
|
||||
function changeQuestion(call: NativePlanQuestionCall, transform: (text: string) => string): void {
|
||||
const q = call.questions[0]!;
|
||||
const answer = call.answers![q.question]!;
|
||||
q.question = transform(q.question);
|
||||
call.answers = { [q.question]: answer };
|
||||
}
|
||||
|
||||
describe('AC DX accounting preserves all accepted obligations', () => {
|
||||
test('D4 confirms accuracy, D12 approves a real repair, and the failed attempt still contains eight issues', () => {
|
||||
const actual = calls().map(classify);
|
||||
expect(actual).toEqual([false, false, false, false, true, true, true, true, true, true, false, true, true]);
|
||||
expect(actual.filter(Boolean)).toHaveLength(8);
|
||||
expect(actual.filter(Boolean).length).toBeGreaterThan(7);
|
||||
});
|
||||
|
||||
test('all three accuracy/correction choices and their order remain observational', () => {
|
||||
for (const option of calls()[3]!.questions[0]!.options) {
|
||||
const c = calls()[3]!;
|
||||
c.answers = { [c.questions[0]!.question]: option.label };
|
||||
c.questions[0]!.options.reverse();
|
||||
changeQuestion(c, text => text.replaceAll('EvalKit', 'RenderKit').replaceAll('ML engineer', 'backend developer'));
|
||||
expect(classify(c)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('the structured frame cannot hide a request in any of its sections', () => {
|
||||
const obligations = [
|
||||
'Should we remove the CI gate?', 'Remove the CI gate.',
|
||||
'I recommend packaging the missing example. Do you approve?',
|
||||
'I approve removing the CI gate; please apply that change.',
|
||||
'I see the missing example. Please update the README.',
|
||||
'I see the missing example. The plan must include it.',
|
||||
'I see the CI gate. Ship a local escape hatch.',
|
||||
'I look at the README. Provide a working command.',
|
||||
];
|
||||
for (const extra of obligations) for (const where of ['headline', 'preamble', 'body', 'closing']) {
|
||||
const c = calls()[3]!;
|
||||
changeQuestion(c, text => {
|
||||
if (where === 'headline') return text.replace('today?', `today? ${extra}`);
|
||||
if (where === 'preamble') return text.replace('\n\nNARRATIVE', ` ${extra}\n\nNARRATIVE`);
|
||||
if (where === 'body') return text.replace('I open the README.', `I open the README. ${extra}`);
|
||||
return text + ` ${extra}`;
|
||||
});
|
||||
expect(classify(c), `${where}: ${extra}`).toBe(true);
|
||||
}
|
||||
for (const extra of [', remove the CI gate', ' and ship a local escape hatch', '; the plan must include a keyless path']) {
|
||||
const c = calls()[3]!;
|
||||
changeQuestion(c, text => text.replace('I open the README.', `I open the README${extra}.`));
|
||||
expect(classify(c)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('each full option description, title, and frame boundary is required', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description += ' Remove the CI gate.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Please update the README.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.description += ' The plan must package the example.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' and ship now'; },
|
||||
(c: NativePlanQuestionCall) => { delete c.questions[0]!.options[0]!.description; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({label: 'Fix the CI gate', description: 'Approve the repair.'}); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'CI gate'; },
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => text.replace('NARRATIVE (', 'PROPOSAL (')),
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => text.replace('Recommendation: A because every step', 'Recommendation: A because we should fix every step')),
|
||||
]) {
|
||||
const c = calls()[3]!; mutate(c);
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
expect(classify(c)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a second answered issue stays substantive while a pending issue contributes no coverage', () => {
|
||||
const c = calls()[3]!; const issue = calls()[4]!;
|
||||
c.questions.push(...issue.questions); Object.assign(c.answers!, issue.answers);
|
||||
expect(classify(c)).toBe(true);
|
||||
delete c.answers![issue.questions[0]!.question]; c.unansweredQuestionIndices = [1];
|
||||
expect(classify(c)).toBe(false);
|
||||
});
|
||||
|
||||
test('the accepted keyless-demo obligation is independent of option position and score', () => {
|
||||
const c = calls()[11]!;
|
||||
c.questions[0]!.options.reverse();
|
||||
changeQuestion(c, text => text.replace('3/10 today', '5/10 today').replaceAll('EVALKIT_API_KEY', 'RENDERKIT_API_KEY'));
|
||||
expect(classify(c)).toBe(true);
|
||||
});
|
||||
|
||||
test('pending, failed, unbound, unselected and quoted keyless-demo proposals supply no accepted-obligation credit', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.header = 'Review mode'; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[1]!.label }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description = 'Confirm that the demo already works without a key.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!)); },
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => '> ' + text),
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => text.replace('should the golden path', 'should not the golden path')),
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => text.replace('reads install, set', 'does not read install, set')),
|
||||
(c: NativePlanQuestionCall) => changeQuestion(c, text => text + ' <gstack-qid:review-mode>'),
|
||||
]) { const c = calls()[11]!; mutate(c); expect(classify(c)).toBe(false); }
|
||||
const fp = nativePlanCallFingerprint(calls()[11]!, 0, true);
|
||||
expect(isDevexReviewIssue({ ...fp, signature: 'foreign:call' })).toBe(false);
|
||||
expect(isDevexReviewIssue({ ...fp, options: [] })).toBe(false);
|
||||
expect(isDevexReviewIssue({ ...fp, nativeCall: undefined })).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,680 +0,0 @@
|
||||
import capturedZ from './fixtures/devex-count-z-calls.json';
|
||||
import capturedURetry from './fixtures/devex-count-u-retry-calls.json';
|
||||
import capturedY from './fixtures/devex-count-y-calls.json';
|
||||
import capturedV from './fixtures/devex-empathy-v-calls.json';
|
||||
import capturedU from './fixtures/devex-count-u-calls.json';
|
||||
|
||||
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import type { AskUserQuestionFingerprint } from './helpers/claude-pty-runner';
|
||||
import capturedL from './fixtures/devex-review-l-calls.json';
|
||||
import capturedN from './fixtures/devex-review-n-calls.json';
|
||||
import capturedT from './fixtures/devex-review-t-calls.json';
|
||||
import { nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import {
|
||||
DEVEX_COUNT_FILES,
|
||||
planDevexCountFixture,
|
||||
isDevexReviewIssue,
|
||||
devexReviewModePick,
|
||||
} from './helpers/devex-count-fixture';
|
||||
|
||||
let nextCall = 0;
|
||||
|
||||
describe('Y agreed TTHW versus retained CI block decision', () => {
|
||||
const captured = () => structuredClone(capturedY[0]!) as NativePlanQuestionCall;
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
|
||||
const change = (c: NativePlanQuestionCall, transform: (text: string) => string) => {
|
||||
const q = c.questions[0]!; const answer = c.answers![q.question]!;
|
||||
q.question = transform(q.question); c.answers = {[q.question]:answer}; return c;
|
||||
};
|
||||
test('all five exact completed native calls carry independent issues', () => {
|
||||
expect(capturedY.map(c => isDevexReviewIssue(fp(structuredClone(c) as NativePlanQuestionCall)))).toEqual([true,true,true,true,true]);
|
||||
expect(capturedY[0]!.answers[capturedY[0]!.questions[0]!.question]).toBe('Demo-only CI bypass (Recommended)');
|
||||
});
|
||||
test('numeric contradiction and selected remedy are independent of literal minutes and option order', () => {
|
||||
const c = change(captured(), text => text.replace('<2 min','<3.5 min').replace('5-min','4-minute').replace('devex-d1-tthw-contradiction','plan-devex-review-timing-conflict'));
|
||||
c.questions[0]!.options.reverse();
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
for (const index of [0,1,2]) {
|
||||
const alternative = captured(); alternative.answers = {[alternative.questions[0]!.question]:alternative.questions[0]!.options[index]!.label};
|
||||
expect(isDevexReviewIssue(fp(alternative))).toBe(true);
|
||||
}
|
||||
for (const [from,to] of [['5-min','1-min'],['<2 min','<0 min'],['5-min','0-min']])
|
||||
expect(isDevexReviewIssue(fp(change(captured(), text => text.replace(from!,to!))))).toBe(false);
|
||||
});
|
||||
test('the complete affirmative statement excludes setup, negation, examples and conditional timings', () => {
|
||||
for (const transform of [
|
||||
(s:string) => s.replace('is mathematically impossible','is not mathematically impossible'),
|
||||
(s:string) => s.replace('is mathematically impossible','is achievable'),
|
||||
(s:string) => s.replace('The agreed','If the agreed'),
|
||||
(s:string) => s.replace('The agreed','Example: The agreed'),
|
||||
(s:string) => '> '+s,
|
||||
(s:string) => '```text\n'+s+'\n```',
|
||||
(s:string) => s.replace('5-min CI block.', '5-min CI block only if optional simulation is enabled.'),
|
||||
(s:string) => s.replace('Which resolution belongs in the plan?', 'Which review mode should we use?'),
|
||||
(s:string) => s.replace('Which resolution belongs in the plan?', 'Should we begin the review?'),
|
||||
(s:string) => s.replace('devex-d1-tthw-contradiction','devex-review-mode'),
|
||||
(s:string) => s.replace('devex-d1-tthw-contradiction','foreign-tthw-contradiction'),
|
||||
(s:string) => s.replace('Which resolution belongs in the plan?', 'Which resolution belongs in the plan? Also approve deployment.'),
|
||||
]) expect(isDevexReviewIssue(fp(change(captured(),transform)))).toBe(false);
|
||||
const c=captured();c.questions[0]!.header='TTHW target';expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
});
|
||||
test('only complete current native answers to offered remedies enter the new arm', () => {
|
||||
for (const mutate of [
|
||||
(c:NativePlanQuestionCall)=>{c.answered=false;},
|
||||
(c:NativePlanQuestionCall)=>{c.failed=true;},
|
||||
(c:NativePlanQuestionCall)=>{delete c.failed;},
|
||||
(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},
|
||||
(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},
|
||||
(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!));},
|
||||
(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered remedy'};},
|
||||
(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:c.questions[0]!.options[3]!.label};},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options[0]!.label='Confirm benchmark';c.answers={[c.questions[0]!.question]:'Confirm benchmark'};},
|
||||
]) { const c=captured();mutate(c);expect(isDevexReviewIssue(fp(c))).toBe(false); }
|
||||
expect(isDevexReviewIssue({...fp(captured()),signature:'foreign:call'})).toBe(false);
|
||||
expect(isDevexReviewIssue({...fp(captured()),nativeCall:undefined})).toBe(false);
|
||||
expect(isDevexReviewIssue({...fp(captured()),options:[...fp(captured()).options].reverse()})).toBe(false);
|
||||
expect(isDevexReviewIssue({...fp(captured()),options:[]})).toBe(false);
|
||||
});
|
||||
});
|
||||
function call(question: string, labels = ['Add to plan', 'Defer']): AskUserQuestionFingerprint {
|
||||
const toolUseId = `tool-${++nextCall}`;
|
||||
return {
|
||||
signature: `session:${toolUseId}`, promptSnippet: question,
|
||||
options: labels.map((label, i) => ({ index: i + 1, label })),
|
||||
observedAtMs: 0, preReview: true,
|
||||
nativeCall: {
|
||||
sessionId: 'session', toolUseId, answered: true,
|
||||
answers: { [question]: labels[0]! },
|
||||
questions: [{ header: 'DX decision', question, options: labels.map(label => ({ label })) }],
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
describe('empathy accuracy is setup, not approval of the quoted findings', () => {
|
||||
const actual = () => structuredClone(capturedV) as NativePlanQuestionCall[];
|
||||
const fp = (native: NativePlanQuestionCall) => nativePlanCallFingerprint(native, 0, true);
|
||||
const mutateQuestion = (native: NativePlanQuestionCall, transform: (question: string) => string) => {
|
||||
const q = native.questions[0]!;
|
||||
const selected = native.answers![q.question]!;
|
||||
q.question = transform(q.question);
|
||||
native.answers = { [q.question]: selected };
|
||||
};
|
||||
test('the exact six answered V calls are one confirmation and five issue decisions', () => {
|
||||
expect(actual().map(c => isDevexReviewIssue(fp(c)))).toEqual([false, true, true, true, true, true]);
|
||||
});
|
||||
test('accuracy-only menus survive reordering, product names, headers and absent IDs', () => {
|
||||
for (const header of ['Empathy narrative', 'Empathy trace', 'Narrative']) {
|
||||
const c = actual()[0]!;
|
||||
c.questions[0]!.header = header;
|
||||
c.questions[0]!.options.reverse();
|
||||
mutateQuestion(c, q => q.replaceAll('EvalKit', 'AnotherSDK').replace('Python ML engineer', 'TypeScript backend developer').replace(/ <gstack-qid:[^>]+>/, ''));
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('correcting the trace still does not approve a remedy', () => {
|
||||
for (const option of actual()[0]!.questions[0]!.options) {
|
||||
const c = actual()[0]!;
|
||||
c.answers = { [c.questions[0]!.question]: option.label };
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
}
|
||||
});
|
||||
test('a remedy option or an instruction in an accuracy description is substantive', () => {
|
||||
for (const edit of [
|
||||
(c: NativePlanQuestionCall) => c.questions[0]!.options.push({label:'Package the missing example', description:'Approve this repair.'}),
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description += ' Repair the missing example.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description = 'Correct the package and its missing example.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' and fix the missing example'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[2]!.description = 'The actual flow differs. Remove the CI gate.'; },
|
||||
]) {
|
||||
const c = actual()[0]!; edit(c);
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
}
|
||||
});
|
||||
test('additional approval questions and unquoted obligations are not confirmation', () => {
|
||||
for (const extra of [
|
||||
' Should we package the missing example?',
|
||||
' Repair the missing example.',
|
||||
' Proceeding also approves the CI bypass.',
|
||||
]) {
|
||||
const c = actual()[0]!;
|
||||
mutateQuestion(c, q => q.replace('Does this match reality? Where am I wrong?', 'Does this match reality? Where am I wrong?'+extra));
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
}
|
||||
const c = actual()[0]!;
|
||||
mutateQuestion(c, q => q.replace('The persona:', 'Repair the missing example. The persona:'));
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
const grant = actual()[0]!;
|
||||
mutateQuestion(grant, q => q.replace('The persona:', 'Grant access to every account. The persona:'));
|
||||
expect(isDevexReviewIssue(fp(grant))).toBe(true);
|
||||
for (const change of [
|
||||
(q: string) => q.replace('the EvalKit getting-started reality', 'the current state and approve packaging the missing quickstart as future reality'),
|
||||
(q: string) => q.replace('The persona: Python ML engineer', 'The persona: Python ML engineer — now package the missing example for this release, a Python ML engineer'),
|
||||
]) { const c = actual()[0]!; mutateQuestion(c,change); expect(isDevexReviewIssue(fp(c))).toBe(true); }
|
||||
});
|
||||
test('a second answered issue tab still counts one issue-bearing call', () => {
|
||||
const c = actual()[0]!; const issue = actual()[1]!;
|
||||
c.questions.push(...issue.questions);
|
||||
Object.assign(c.answers!, issue.answers);
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
delete c.answers![issue.questions[0]!.question];
|
||||
c.unansweredQuestionIndices = [1];
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
const issues = [
|
||||
['CI gate', 'Journey Stage: HELLO WORLD. The mandatory five-minute CI gate blocks the first local evaluation. Remove the gate or make it optional for local runs?'],
|
||||
['Argument order', 'run_eval(dataset, evaluator) and run_batch(evaluator, dataset) reverse the positional order. Should we standardize these signatures or require keyword arguments?'],
|
||||
['Authentication error', 'An invalid API key raises AuthError("request failed"), with no explanation or recovery guidance. How should we replace this opaque error?'],
|
||||
['Packaged example', 'The quickstart tells developers to run examples/first_eval.py, but it is absent from the published package. Include the example or fix the documented command?'],
|
||||
['Breaking rename', 'Version 2 removes Client.evaluate and replaces it with Client.run without a migration guide or deprecation warning. Add a compatibility alias or a migration path?'],
|
||||
] as const;
|
||||
|
||||
describe('DevEx substantive finding coverage', () => {
|
||||
test('the actual untagged empathy confirmation does not borrow a finding from its recap', () => {
|
||||
const actual = structuredClone(capturedL.calls[0]!);
|
||||
const fp = call(actual.questions[0]!.question);
|
||||
fp.nativeCall = actual;
|
||||
expect(isDevexReviewIssue(fp)).toBe(false);
|
||||
// An empathy-derived remedy decision is still substantive. The exclusion
|
||||
// requires the confirmation question, not merely a familiar header.
|
||||
const question = issues[0][1];
|
||||
actual.questions[0]!.question = question;
|
||||
actual.answers = { [question]: actual.questions[0]!.options[0]!.label };
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
test('an empathy-shaped question with remedy choices stays substantive', () => {
|
||||
for (const mixed of [false, true]) {
|
||||
const actual = structuredClone(capturedL.calls[0]!);
|
||||
const fp = call(actual.questions[0]!.question);
|
||||
fp.nativeCall = actual;
|
||||
if (mixed) actual.questions[0]!.options.push({label:'Package the missing example',description:'Fix the quickstart now'});
|
||||
else actual.questions[0]!.options = [{label:'Package the missing example',description:'Fix the quickstart now'}, {label:'Leave the example absent',description:'Defer the fix'}];
|
||||
actual.answers = { [actual.questions[0]!.question]: actual.questions[0]!.options[0]!.label };
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
}
|
||||
});
|
||||
test('a correction label cannot hide an instruction to fix the package', () => {
|
||||
const actual = structuredClone(capturedL.calls[0]!);
|
||||
actual.questions[0]!.options[1]!.label = 'Partially — the example is absent; package it now';
|
||||
const fp = call(actual.questions[0]!.question);
|
||||
fp.nativeCall = actual;
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
test('the observed mandatory confirmations alone contribute zero findings', () => {
|
||||
const confirmations = [
|
||||
'No design doc found. Run /office-hours first? <gstack-qid:plan-devex-review-office-hours-preflight>',
|
||||
'Who is your primary target developer? <gstack-qid:plan-devex-review-persona>',
|
||||
'Does the empathy narrative match reality? <gstack-qid:plan-devex-review-empathy-check>',
|
||||
// A concrete defect in a benchmark recap does not make the target
|
||||
// confirmation itself a resolution decision for that defect.
|
||||
'Remove the mandatory CI wait before first eval to reach the agreed benchmark. Which tier do you confirm? <gstack-qid:plan-devex-review-tthw-tier>',
|
||||
'What should the magical first-eval moment look like? <gstack-qid:plan-devex-review-magical-moment>',
|
||||
'How deep should this DX review go? <gstack-qid:plan-devex-review-mode>',
|
||||
'Confusion report reviewed. Which items should be addressed? <gstack-qid:plan-devex-review-confusion-report>',
|
||||
'Which onboarding setup should run next? <gstack-qid:future-setup-choice>',
|
||||
];
|
||||
expect(confirmations.map(question => call(question)).filter(isDevexReviewIssue)).toEqual([]);
|
||||
});
|
||||
|
||||
test.each(issues)('%s is a finding in investigation or scoring', (_name, question) => {
|
||||
const fp = call(question);
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
fp.preReview = false;
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
|
||||
test('full native question evidence survives a short diagnostic snippet', () => {
|
||||
const fp = call('Context from the SDK audit. '.repeat(20) + issues[3][1]);
|
||||
fp.promptSnippet = fp.promptSnippet.slice(0, 240);
|
||||
expect(fp.promptSnippet).not.toContain('examples/first_eval.py');
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
|
||||
test('a real argument-order decision does not need a particular resolution verb', () => {
|
||||
const fp = call('Which argument order should run_eval and run_batch use?', [
|
||||
'Dataset first in both functions', 'Evaluator first in both functions',
|
||||
]);
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
|
||||
test.each([
|
||||
'Design doc', 'Target persona', 'Narrative check', 'TTHW target',
|
||||
'Magic delivery', 'Review mode', 'Fix scope',
|
||||
])('observed administrative header %s cannot borrow a defect from its recap', header => {
|
||||
const fp = call(`${issues[0][1]} This is the context for our confirmation.`);
|
||||
fp.nativeCall!.questions[0]!.header = header;
|
||||
expect(isDevexReviewIssue(fp)).toBe(false);
|
||||
});
|
||||
|
||||
test('a CI issue stays substantive when it references persona and TTHW evidence', () => {
|
||||
const fp = call('The target persona confirmed our TTHW target. The mandatory CI gate blocks the first eval. Which local bypass should the SDK support?');
|
||||
fp.nativeCall!.questions[0]!.header = 'CI gate fix';
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
});
|
||||
|
||||
test('one call batching the defects does not become five finding decisions', () => {
|
||||
const distinct = issues.map(([, question]) => call(question));
|
||||
expect(distinct.filter(isDevexReviewIssue)).toHaveLength(5);
|
||||
const batched = call('Review these issues together.');
|
||||
batched.nativeCall!.questions = distinct.flatMap(fp => fp.nativeCall!.questions);
|
||||
batched.nativeCall!.answers = Object.assign({}, ...distinct.map(fp => fp.nativeCall!.answers));
|
||||
expect([batched].filter(isDevexReviewIssue)).toHaveLength(1);
|
||||
});
|
||||
|
||||
test('an unanswered issue tab cannot turn an administrative answer into coverage', () => {
|
||||
const admin = call('How deep should this DX review go? <gstack-qid:plan-devex-review-mode>');
|
||||
const issue = call(issues[0][1]);
|
||||
admin.nativeCall!.questions.push(issue.nativeCall!.questions[0]!);
|
||||
admin.nativeCall!.unansweredQuestionIndices = [1];
|
||||
expect(isDevexReviewIssue(admin)).toBe(false);
|
||||
Object.assign(admin.nativeCall!.answers!, issue.nativeCall!.answers);
|
||||
admin.nativeCall!.unansweredQuestionIndices = [];
|
||||
expect(isDevexReviewIssue(admin)).toBe(true);
|
||||
});
|
||||
|
||||
test.each([
|
||||
'Which files should I review? <gstack-qid:unknown-administrative-choice>',
|
||||
'I noted the mandatory CI gate before first eval. Can we continue the setup?',
|
||||
'Should I add a developer community Slack channel?',
|
||||
'Should the plan reference run_eval and run_batch?',
|
||||
'The package includes examples/first_eval.py. Shall I read it?',
|
||||
'Authentication errors already include a cause and a fix. Ready to continue?',
|
||||
])('unknown or unsupported prompts do not count: %s', question => {
|
||||
expect(isDevexReviewIssue(call(question))).toBe(false);
|
||||
});
|
||||
|
||||
test('a generic question cannot borrow issue evidence from its option labels', () => {
|
||||
expect(isDevexReviewIssue(call('What should I inspect next?', [issues[0][1], issues[1][1]]))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('DevEx count review-mode selection', () => {
|
||||
const modeQuestion = 'D6 — How deep should this DX review go? <gstack-qid:plan-devex-review-mode>';
|
||||
|
||||
test('selects POLISH from the actual menu that previously chose EXPANSION', () => {
|
||||
expect(devexReviewModePick(call(modeQuestion, [
|
||||
'DX EXPANSION (Recommended)', 'DX POLISH', 'DX TRIAGE',
|
||||
]))).toBe(2);
|
||||
});
|
||||
|
||||
test('retains the observed POLISH index after menu reordering', () => {
|
||||
expect(devexReviewModePick(call(modeQuestion, [
|
||||
'DX TRIAGE', 'DX EXPANSION', 'DX POLISH (Recommended)',
|
||||
]))).toBe(3);
|
||||
});
|
||||
|
||||
test('recognizes the same mode question without a question ID', () => {
|
||||
expect(devexReviewModePick(call('HowdeepshouldthisDXreviewgo?', [
|
||||
'DXEXPANSION(Recommended)', 'DXPOLISH', 'DXTRIAGE',
|
||||
]))).toBe(2);
|
||||
});
|
||||
|
||||
test('unrelated questions cannot select a mode from quoted labels', () => {
|
||||
expect(devexReviewModePick(call('Which documentation example should be included?', [
|
||||
'DX EXPANSION', 'DX POLISH', 'DX TRIAGE',
|
||||
]))).toBeNull();
|
||||
expect(devexReviewModePick(call(issues[0][1]))).toBeNull();
|
||||
});
|
||||
|
||||
test('missing or ambiguous mode menus keep the existing choice policy', () => {
|
||||
expect(devexReviewModePick(call(modeQuestion, ['DX EXPANSION', 'DX TRIAGE']))).toBeNull();
|
||||
expect(devexReviewModePick(call(modeQuestion, [
|
||||
'DX EXPANSION', 'DX POLISH', 'DX POLISH', 'DX TRIAGE',
|
||||
]))).toBeNull();
|
||||
expect(devexReviewModePick(call(modeQuestion, [
|
||||
'DX EXPANSION │ DX POLISH', 'Example │ DX POLISH', 'DX TRIAGE',
|
||||
]))).toBeNull();
|
||||
});
|
||||
|
||||
test('a multi-question call is not treated as a single mode menu', () => {
|
||||
const fp = call(modeQuestion, ['DX EXPANSION', 'DX POLISH', 'DX TRIAGE']);
|
||||
fp.nativeCall!.questions.push(call(issues[0][1]).nativeCall!.questions[0]!);
|
||||
expect(devexReviewModePick(fp)).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('DevEx calibrated fixture instructions', () => {
|
||||
test('keeps the reviewed artifact path without telling the model an expected count', () => {
|
||||
const plan = planDevexCountFixture('/tmp/owned-plan.md');
|
||||
expect(plan).toContain('write your plan-mode plan to /tmp/owned-plan.md');
|
||||
const suppliedContext = [plan, ...Object.values(DEVEX_COUNT_FILES)].join('\n');
|
||||
expect(suppliedContext).not.toMatch(/(?:exactly|at least|at most)\s+(?:five|5)|(?:five|5)[- ]findings?|4[-–]7|reviewCount|CEILING|FLOOR/i);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('native first-local-run CI decisions', () => {
|
||||
const question =
|
||||
'D3 \u2014 Journey Stage: FIRST RESULT \u2014 5-minute CI gate makes the <2min TTHW target mathematically unreachable\n\nELI10: On every first local run, the SDK blocks for 5 minutes waiting for a remote CI check (docs/current-contracts.md). There is no skip flag. The TTHW study measured EvalKit at 6 minutes total (docs/benchmarks.md). The agreed target is under 2 minutes. With a mandatory 5-minute wait baked in, you cannot reach that target \u2014 the CI gate alone exceeds it. Competitors: A=2min, B=4min, C=3min. EvalKit currently loses on TTHW.\n\nStakes if we pick wrong: If the target stays <2min but the gate stays too, the benchmark is aspirational theatre. If the gate stays and the target is adjusted, the competitive position is weaker.\n\nRecommendation: A \u2014 add a local skip path. The CI gate adds real value in production CI, but blocking local first-runs is the wrong tradeoff for an SDK that wants sub-2min TTHW.\nNote: options differ in kind, not coverage \u2014 no completeness score.\n\n<gstack-qid:plan-devex-review-ci-gate>';
|
||||
test('the answered first-local-run CI gate is substantive, including its plural variant', () => {
|
||||
for (const text of [
|
||||
question,
|
||||
question.replace('first local run', 'first local runs'),
|
||||
]) {
|
||||
const fp = call(text);
|
||||
fp.nativeCall!.questions[0]!.header = 'CI gate TTHW';
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
}
|
||||
});
|
||||
test('an unanswered CI tab and an administrative recap never create coverage', () => {
|
||||
const fp = call('Does the empathy narrative match reality?');
|
||||
fp.nativeCall!.questions[0]!.header = 'Empathy check';
|
||||
fp.nativeCall!.questions.push({
|
||||
header: 'CI gate TTHW',
|
||||
question,
|
||||
options: [{ label: 'Skip CI' }, { label: 'Keep CI' }],
|
||||
});
|
||||
fp.nativeCall!.unansweredQuestionIndices = [1];
|
||||
expect(isDevexReviewIssue(fp)).toBe(false);
|
||||
fp.nativeCall!.answers![question] = 'Skip CI';
|
||||
fp.nativeCall!.unansweredQuestionIndices = [];
|
||||
expect(isDevexReviewIssue(fp)).toBe(true);
|
||||
const recap = call(question);
|
||||
recap.nativeCall!.questions[0]!.header = 'Empathy check';
|
||||
expect(isDevexReviewIssue(recap)).toBe(false);
|
||||
expect(
|
||||
isDevexReviewIssue(
|
||||
call(
|
||||
'The production CI gate waits five minutes. Change the release check?',
|
||||
),
|
||||
),
|
||||
).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('native developer-trace accuracy confirmation', () => {
|
||||
const actualCalls = () => structuredClone(capturedN.calls) as NativePlanQuestionCall[];
|
||||
const actual = () => actualCalls()[0]!;
|
||||
const fp = (native: NativePlanQuestionCall) => nativePlanCallFingerprint(native, 0, true);
|
||||
|
||||
test('the captured developer narrative confirms evidence and retains all five actual issue decisions', () => {
|
||||
const input = actualCalls();
|
||||
const before = structuredClone(input);
|
||||
expect(isDevexReviewIssue(fp(input[0]!))).toBe(false);
|
||||
expect(input.filter(native => isDevexReviewIssue(fp(native)))).toHaveLength(5);
|
||||
expect(input.slice(1).every(native => isDevexReviewIssue(fp(native)))).toBe(true);
|
||||
expect(input).toEqual(before);
|
||||
});
|
||||
|
||||
test('accuracy labels cannot hide remedy choices or a substantive repair question', () => {
|
||||
for (const mutate of [
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = issues[3][1]; },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.options[0]!.label = 'Package the missing example now'; },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.options[1]!.description = 'Package the missing example now.'; },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question += ' Should I package the missing examples/first_eval.py to fix this quickstart?'; },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = native.questions[0]!.question.replace('Does this match the actual experience?', 'Should I package the missing examples/first_eval.py to fix this quickstart? Does this match the actual experience?'); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = native.questions[0]!.question.replace('Does this match the actual experience?', 'Do you want me to package the missing examples/first_eval.py to fix this quickstart? Does this match the actual experience?'); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = native.questions[0]!.question.replace('Does this match the actual experience?', 'Would you like the missing examples/first_eval.py packaged? Does this match the actual experience?'); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = native.questions[0]!.question.replace('Does this match the actual experience?', 'Approve packaging the missing examples/first_eval.py? Does this match the actual experience?'); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question = native.questions[0]!.question.replace('Does this match the actual experience?', 'Please package the missing examples/first_eval.py. Does this match the actual experience?'); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.options[0]!.description = 'Proceed to package the missing examples/first_eval.py so the quickstart works.'; },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.options.push({ label: 'Fix the API argument order' }); },
|
||||
(native: NativePlanQuestionCall) => { native.questions[0]!.question += ' <gstack-qid:plan-devex-example-fix>'; },
|
||||
]) {
|
||||
const native = actual();
|
||||
mutate(native);
|
||||
native.answers = { [native.questions[0]!.question]: native.questions[0]!.options[0]!.label };
|
||||
expect(isDevexReviewIssue(fp(native))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('an answered issue beside the narrative still counts the native call once', () => {
|
||||
const native = actual();
|
||||
const issue = actualCalls()[1]!;
|
||||
native.questions.push(issue.questions[0]!);
|
||||
native.unansweredQuestionIndices = [1];
|
||||
expect(isDevexReviewIssue(fp(native))).toBe(false);
|
||||
Object.assign(native.answers!, issue.answers);
|
||||
native.unansweredQuestionIndices = [];
|
||||
expect([native].filter(value => isDevexReviewIssue(fp(value)))).toHaveLength(1);
|
||||
native.answered = false;
|
||||
expect(isDevexReviewIssue(fp(native))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('T native documentation follow-up decisions', () => {
|
||||
const calls = () => structuredClone(capturedT.calls) as NativePlanQuestionCall[];
|
||||
const fp = (native: NativePlanQuestionCall) => nativePlanCallFingerprint(native, 0, true);
|
||||
const changeQuestion = (native: NativePlanQuestionCall, transform: (s: string) => string) => {
|
||||
const q = native.questions[0]!;
|
||||
const answer = native.answers![q.question]!;
|
||||
q.question = transform(q.question);
|
||||
native.answers = { [q.question]: answer };
|
||||
return native;
|
||||
};
|
||||
|
||||
test('the complete captured census keeps empathy setup and seven distinct issue calls', () => {
|
||||
const actual = calls(); const before = structuredClone(actual);
|
||||
expect(actual.map(c => isDevexReviewIssue(fp(c)))).toEqual([false, true, true, true, true, true, true, true]);
|
||||
expect(actual).toEqual(before);
|
||||
});
|
||||
|
||||
for (const index of [6, 7]) {
|
||||
test(`follow-up ${index} requires complete native offered-answer identity`, () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'Foreign answer' }; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!)); },
|
||||
]) { const c = calls()[index]!; mutate(c); expect(isDevexReviewIssue(fp(c))).toBe(false); }
|
||||
const foreign = fp(calls()[index]!); foreign.signature = 'foreign:call'; expect(isDevexReviewIssue(foreign)).toBe(false);
|
||||
const screen = fp(calls()[index]!); delete screen.nativeCall; expect(isDevexReviewIssue(screen)).toBe(false);
|
||||
});
|
||||
|
||||
test(`follow-up ${index} cannot borrow an unselected remedy or a setup identity`, () => {
|
||||
const skipped = calls()[index]!; const q = skipped.questions[0]!;
|
||||
skipped.answers = { [q.question]: q.options.at(-1)!.label };
|
||||
expect(isDevexReviewIssue(fp(skipped))).toBe(false);
|
||||
for (const header of ['Empathy check', 'Review mode', 'Next steps']) {
|
||||
const c = calls()[index]!; c.questions[0]!.header = header; expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
}
|
||||
for (const replacement of ['<gstack-qid:devex-mode>', '<gstack-qid:devex-next-steps>', '<gstack-qid:foreign>']) {
|
||||
const c = changeQuestion(calls()[index]!, s => s.replace(/<gstack-qid:[^>]+>/, replacement));
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
}
|
||||
const duplicate = changeQuestion(calls()[index]!, s => s + ' <gstack-qid:devex-extra>');
|
||||
expect(isDevexReviewIssue(fp(duplicate))).toBe(false);
|
||||
});
|
||||
}
|
||||
|
||||
test('a resolved documentation gap, quoted example or removed follow-up obligation earns no new credit', () => {
|
||||
for (const transform of [
|
||||
(s: string) => s.replace('but never says where to get one', 'and already says where to get one'),
|
||||
(s: string) => s.replace('Documentation — README', 'Documentation — It is false that README'),
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```text\n' + s + '\n```',
|
||||
]) expect(isDevexReviewIssue(fp(changeQuestion(calls()[6]!, transform)))).toBe(false);
|
||||
for (const transform of [
|
||||
(s: string) => s.replace('**What:** Add', '**What:** Do not add'),
|
||||
(s: string) => s.replace('additional examples/ files', 'the already-approved quickstart file'),
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```text\n' + s + '\n```',
|
||||
]) expect(isDevexReviewIssue(fp(changeQuestion(calls()[7]!, transform)))).toBe(false);
|
||||
});
|
||||
|
||||
test('option reordering preserves the exact selected remedy and each native call counts once', () => {
|
||||
for (const c of calls().slice(6)) {
|
||||
c.questions[0]!.options.reverse();
|
||||
expect([c].filter(c => isDevexReviewIssue(fp(c)))).toHaveLength(1);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('U completed first-pass contract decisions', () => {
|
||||
const calls = () => structuredClone(capturedU.calls) as NativePlanQuestionCall[];
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
const change = (call: NativePlanQuestionCall, transform: (s: string) => string) => {
|
||||
const q = call.questions[0]!; const answer = call.answers![q.question]!;
|
||||
q.question = transform(q.question); call.answers = {[q.question]: answer}; return call;
|
||||
};
|
||||
test('all five actual seed decisions count once, without mutating evidence', () => {
|
||||
const actual = calls(); const before = structuredClone(actual);
|
||||
expect(actual.map(call => isDevexReviewIssue(fp(call)))).toEqual([true, true, true, true, true]);
|
||||
expect(actual).toEqual(before);
|
||||
});
|
||||
for (const index of [0, 1]) {
|
||||
test(`decision ${index + 1} requires complete native identity and an offered answer`, () => {
|
||||
expect(isDevexReviewIssue(fp(calls()[index]!))).toBe(true);
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {}; },
|
||||
(c: NativePlanQuestionCall) => { c.answers = {[c.questions[0]!.question]: 'Foreign answer'}; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!)); },
|
||||
]) { const c = calls()[index]!; mutate(c); expect(isDevexReviewIssue(fp(c))).toBe(false); }
|
||||
const foreign = fp(calls()[index]!); foreign.signature = 'foreign:tool'; expect(isDevexReviewIssue(foreign)).toBe(false);
|
||||
const ui = fp(calls()[index]!); delete ui.nativeCall; expect(isDevexReviewIssue(ui)).toBe(false);
|
||||
});
|
||||
test(`decision ${index + 1} cannot borrow issue words for setup or quoted examples`, () => {
|
||||
for (const transform of [
|
||||
(s: string) => '> ' + s,
|
||||
(s: string) => '```text\n' + s + '\n```',
|
||||
(s: string) => s.replace(/Pass 1 \(Getting Started\):/, 'Pass 1 (Getting Started): It is false that'),
|
||||
(s: string) => s.replace(/<gstack-qid:[^>]+>/, '<gstack-qid:plan-devex-review-mode>'),
|
||||
(s: string) => s + ' <gstack-qid:another>',
|
||||
]) expect(isDevexReviewIssue(fp(change(calls()[index]!, transform)))).toBe(false);
|
||||
const c = calls()[index]!; c.questions[0]!.header = 'Review mode'; expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
});
|
||||
test(`decision ${index + 1} keeps a distinct accepted or deferred decision independent of option order`, () => {
|
||||
const c = calls()[index]!; const q = c.questions[0]!;
|
||||
q.options.reverse(); expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
c.answers = {[q.question]: q.options[0]!.label}; expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
});
|
||||
}
|
||||
test('resolved or negated first-run contracts and pure navigation do not count', () => {
|
||||
for (const transform of [
|
||||
(s: string) => s.replace("doesn't ship", 'already ships'),
|
||||
(s: string) => s.replace('quickstart points to', 'quickstart no longer points to'),
|
||||
(s: string) => s.replace('Should we fix the quickstart path in the plan?', 'Should we begin the review?'),
|
||||
(s: string) => s.replace('Should we fix', 'Should we not fix'),
|
||||
]) expect(isDevexReviewIssue(fp(change(calls()[0]!, transform)))).toBe(false);
|
||||
for (const transform of [
|
||||
(s: string) => s.replace('makes that unreachable', 'makes that reachable'),
|
||||
(s: string) => s.replace('makes that unreachable', 'does not make that unreachable'),
|
||||
(s: string) => s.replace('The plan retains the gate.', 'The plan already skips the gate.'),
|
||||
(s: string) => s.replace('How should this plan handle the contradiction?', 'Should we begin the review?'),
|
||||
]) expect(isDevexReviewIssue(fp(change(calls()[1]!, transform)))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
describe('U demo timing decision after completed measurements', () => {
|
||||
const captured = () => structuredClone(capturedURetry[0]!) as NativePlanQuestionCall;
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, true);
|
||||
function replace(c: NativePlanQuestionCall, from: string, to: string) {
|
||||
const q = c.questions[0]!; const old = q.question; q.question = old.replace(from, to);
|
||||
if (c.answers) c.answers = { [q.question]: c.answers[old]! };
|
||||
return c;
|
||||
}
|
||||
test('all five actual completed calls are independent issue decisions', () => {
|
||||
const calls = structuredClone(capturedURetry) as NativePlanQuestionCall[];
|
||||
expect(calls.map(c => isDevexReviewIssue(fp(c)))).toEqual([true, true, true, true, true]);
|
||||
expect(calls).toEqual(capturedURetry);
|
||||
const c = captured(); c.questions[0]!.options.reverse();
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true); // Deferring the repair is still this decision.
|
||||
});
|
||||
test('timings are compared instead of pinning the observed minutes', () => {
|
||||
let c = captured();
|
||||
for (const [from,to] of [['<2 min','<3 min'],['under 2 minutes','under 3 minutes'],['blocks for 5 minutes','blocks for 4 minutes'],['measured TTHW of 6 minutes','measured TTHW of 5 minutes']]) c=replace(c,from!,to!);
|
||||
expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
for (const [from,to] of [['blocks for 5 minutes','blocks for 1 minutes'],['measured TTHW of 6 minutes','measured TTHW of 4 minutes'],['under 2 minutes','under 9 minutes']])
|
||||
expect(isDevexReviewIssue(fp(replace(captured(),from!,to!)))).toBe(false);
|
||||
});
|
||||
test('setup, negated, quoted and merely hypothetical timing claims remain outside the new arm', () => {
|
||||
for (const [from,to] of [
|
||||
['should it bypass the mandatory CI check to reach the <2 min TTHW target?', 'which TTHW target should we confirm?'],
|
||||
['ELI10: The agreed onboarding target is under 2 minutes', 'Example: ELI10: The agreed onboarding target is under 2 minutes'],
|
||||
['ELI10: The agreed onboarding target is under 2 minutes', '> ELI10: The agreed onboarding target is under 2 minutes'],
|
||||
['ELI10: The agreed onboarding target is under 2 minutes', '```text\nELI10: The agreed onboarding target is under 2 minutes'],
|
||||
['Today `python -m evalkit.demo` blocks', 'Today `python -m evalkit.demo` no longer blocks'],
|
||||
['Today `python -m evalkit.demo` blocks', 'It is false that `python -m evalkit.demo` blocks'],
|
||||
['Today `python -m evalkit.demo` blocks', 'If `python -m evalkit.demo` blocks'],
|
||||
['giving a measured TTHW of 6 minutes', 'giving a measured TTHW of 6 minutes only if the optional slow simulation is enabled'],
|
||||
['giving a measured TTHW of 6 minutes', 'giving a measured TTHW of 6 minutes only in a hypothetical example'],
|
||||
['devex-demo-ci-bypass', 'plan-devex-review-tthw-tier'],
|
||||
]) expect(isDevexReviewIssue(fp(replace(captured(),from!,to!)))).toBe(false);
|
||||
const c = captured(); c.questions[0]!.header = 'TTHW target'; expect(isDevexReviewIssue(fp(c))).toBe(false);
|
||||
});
|
||||
test('the new measured branch requires one complete matched native decision', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.failed = true; },
|
||||
(c: NativePlanQuestionCall) => { delete c.failed; },
|
||||
(c: NativePlanQuestionCall) => { c.answered = false; },
|
||||
(c: NativePlanQuestionCall) => { delete c.unansweredQuestionIndices; },
|
||||
(c: NativePlanQuestionCall) => { c.unansweredQuestionIndices = [0]; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.multiSelect = true; },
|
||||
(c: NativePlanQuestionCall) => { c.questions.push(structuredClone(c.questions[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!)); },
|
||||
(c: NativePlanQuestionCall) => { c.answers = { [c.questions[0]!.question]: 'unoffered answer' }; },
|
||||
]) { const c=captured(); mutate(c); expect(isDevexReviewIssue(fp(c))).toBe(false); }
|
||||
expect(isDevexReviewIssue({...fp(captured()), signature:'foreign:call'})).toBe(false);
|
||||
expect(isDevexReviewIssue({...fp(captured()), nativeCall:undefined})).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Z written migration guide as an additional accepted obligation', () => {
|
||||
const call = () => structuredClone(capturedZ[7]!) as NativePlanQuestionCall;
|
||||
const fp = (c: NativePlanQuestionCall) => nativePlanCallFingerprint(c, 0, false);
|
||||
const change = (c: NativePlanQuestionCall, from: string, to: string) => {
|
||||
const q=c.questions[0]!;const answer=c.answers![q.question]!;q.question=q.question.replaceAll(from,to);c.answers={[q.question]:answer};return c;
|
||||
};
|
||||
test('all eight real calls retain empathy plus seven distinct issue decisions', () => {
|
||||
const calls=structuredClone(capturedZ) as NativePlanQuestionCall[];
|
||||
expect(calls.map(c=>isDevexReviewIssue(fp(c)))).toEqual([false,true,true,true,true,true,true,true]);expect(calls).toEqual(capturedZ);
|
||||
const c=call();c.questions[0]!.options.reverse();expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
});
|
||||
test('version, decision and task numbers do not determine finding credit', () => {
|
||||
const c=call();for(const [from,to] of [['D8','D17'],['TODO-2','TODO-9'],['todo2-migration','todo9-migration'],['v1','v3'],['v2','v4'],['T4','T11'],['P2','P1']]) {
|
||||
change(c,from!,to!);const q=c.questions[0]!;q.header=q.header.replaceAll(from!,to!);q.options.forEach(o=>{o.description=o.description?.replaceAll(from!,to!);});
|
||||
}expect(isDevexReviewIssue(fp(c))).toBe(true);
|
||||
});
|
||||
test('setup, quoted, hypothetical and already satisfied claims confer no new acceptance', () => {
|
||||
for(const [from,to] of [
|
||||
['TODO: should the plan include','TODO: should the review confirm'],
|
||||
['But there is currently no written migration guide in docs/.','The written migration guide already exists in docs/.'],
|
||||
['But there is currently no written migration guide in docs/.','But there is currently no written migration guide in docs/ only in this hypothetical example.'],
|
||||
['The deprecation shim (T4) handles','If the deprecation shim (T4) handles'],
|
||||
['The deprecation shim (T4) handles','> The deprecation shim (T4) handles'],
|
||||
['The deprecation shim (T4) handles','```text\nThe deprecation shim (T4) handles'],
|
||||
['A one-page migration guide covers:','The already-approved migration guide covers:'],
|
||||
['Without it, developers','This is only an example. Without it, developers'],
|
||||
['<gstack-qid:plan-devex-review-todo2-migration-guide>','<gstack-qid:plan-devex-review-mode>'],
|
||||
['<gstack-qid:plan-devex-review-todo2-migration-guide>','<gstack-qid:foreign>'],
|
||||
])expect(isDevexReviewIssue(fp(change(call(),from!,to!)))).toBe(false);
|
||||
for(const header of ['Review mode','Empathy check','Next steps']){const c=call();c.questions[0]!.header=header;expect(isDevexReviewIssue(fp(c))).toBe(false);}
|
||||
expect(isDevexReviewIssue(fp(change(call(),'<gstack-qid:plan-devex-review-todo2-migration-guide>','<gstack-qid:plan-devex-review-todo2-migration-guide> <gstack-qid:extra>')))).toBe(false);
|
||||
});
|
||||
test('one exact successful native call must select the offered written-guide task', () => {
|
||||
for(const mutate of [
|
||||
(c:NativePlanQuestionCall)=>{c.answered=false;},(c:NativePlanQuestionCall)=>{c.failed=true;},(c:NativePlanQuestionCall)=>{delete c.failed;},
|
||||
(c:NativePlanQuestionCall)=>{delete c.unansweredQuestionIndices;},(c:NativePlanQuestionCall)=>{c.unansweredQuestionIndices=[0];},
|
||||
(c:NativePlanQuestionCall)=>{c.answers={};},(c:NativePlanQuestionCall)=>{c.answers={[c.questions[0]!.question]:'unoffered answer'};},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.multiSelect=true;},(c:NativePlanQuestionCall)=>{c.questions.push(structuredClone(c.questions[0]!));},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options.push(structuredClone(c.questions[0]!.options[0]!));},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options[0]!.description+=' Also remove authentication.';},
|
||||
(c:NativePlanQuestionCall)=>{c.questions[0]!.options[0]!.description=undefined;},
|
||||
]){const c=call();mutate(c);expect(isDevexReviewIssue(fp(c))).toBe(false);}
|
||||
for(const index of [1,2]){const c=call();const q=c.questions[0]!;c.answers={[q.question]:q.options[index]!.label};expect(isDevexReviewIssue(fp(c))).toBe(false);}
|
||||
expect(isDevexReviewIssue({...fp(call()),signature:'foreign:call'})).toBe(false);expect(isDevexReviewIssue({...fp(call()),nativeCall:undefined})).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,93 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import { isDevexReviewIssue } from './helpers/devex-count-fixture';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import recorded from './fixtures/devex-empathy-ab-calls.json';
|
||||
|
||||
const calls = () => structuredClone(recorded) as NativePlanQuestionCall[];
|
||||
const classify = (call: NativePlanQuestionCall) => isDevexReviewIssue(nativePlanCallFingerprint(call, 0, true));
|
||||
function changeQuestion(call: NativePlanQuestionCall, transform: (text: string) => string): void {
|
||||
const q = call.questions[0]!;
|
||||
const answer = call.answers![q.question]!;
|
||||
q.question = transform(q.question);
|
||||
call.answers = { [q.question]: answer };
|
||||
}
|
||||
|
||||
describe('DX delimited empathy accuracy confirmation', () => {
|
||||
test('the seven completed AB calls are two setup confirmations and five issue decisions', () => {
|
||||
expect(calls().map(classify)).toEqual([false, false, true, true, true, true, true]);
|
||||
});
|
||||
|
||||
test('accuracy and correction choices do not approve the defects described in the trace', () => {
|
||||
for (const answer of calls()[1]!.questions[0]!.options.map(o => o.label)) {
|
||||
const c = calls()[1]!;
|
||||
c.answers = { [c.questions[0]!.question]: answer };
|
||||
c.questions[0]!.options.reverse();
|
||||
changeQuestion(c, text => text.replaceAll('EvalKit', 'RenderKit').replace('Python ML engineer', 'TypeScript frontend developer'));
|
||||
expect(classify(c)).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
test('extra obligations outside the delimited journey are still substantive', () => {
|
||||
for (const transform of [
|
||||
(s: string) => s.replace('Does this match reality?', 'Does this match reality? Also package the missing example.'),
|
||||
(s: string) => s.replace("Here's what I think", "Package the missing example. Here's what I think"),
|
||||
(s: string) => s.replace('Does this match reality?', 'Should we fix the missing example? Does this match reality?'),
|
||||
(s: string) => s.replace('your actual developer experience?', 'your actual developer experience and approve packaging the example?'),
|
||||
(s: string) => s.replace(/\n\n---\n\nDoes this match reality\?$/, '\n\n---\n\nRemove the CI gate.\n\nDoes this match reality?'),
|
||||
(s: string) => s.replace(/\n\n---\n\nDoes this match reality\?$/, '\n\nDoes this match reality?'),
|
||||
]) { const c = calls()[1]!; changeQuestion(c, transform); expect(classify(c)).toBe(true); }
|
||||
});
|
||||
|
||||
test('delimiters cannot hide remedy paragraphs or appended decision clauses', () => {
|
||||
for (const extra of [
|
||||
'Should we remove the CI gate?',
|
||||
'Remove the CI gate.',
|
||||
'I recommend packaging the missing example. Do you approve?',
|
||||
'I approve removing the CI gate; please apply that change.',
|
||||
'I look at the package. Should we add the missing example?',
|
||||
'I run the demo; remove the CI gate.',
|
||||
'I got results. We should package the missing example.',
|
||||
'I found the CI gate. Please disable it.',
|
||||
'I see the missing example. I decide to package it.',
|
||||
'I got results. We will remove the CI gate.',
|
||||
"I check the package. Let's add the missing example.",
|
||||
'I see the missing example. Please update the README.',
|
||||
'I see the missing example. The plan must include it.',
|
||||
'I see the CI gate. Ship a local escape hatch.',
|
||||
'I see the CI gate; Update the documentation.',
|
||||
'I look at the README. Provide a working command.',
|
||||
]) {
|
||||
for (const separator of ['\n\n', ' ']) {
|
||||
const c = calls()[1]!;
|
||||
changeQuestion(c, text => text.replace(/\n\n---\n\nDoes this match reality\?$/,
|
||||
separator + extra + '\n\n---\n\nDoes this match reality?'));
|
||||
expect(classify(c)).toBe(true);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('a remedy inside an option cannot borrow an accuracy label', () => {
|
||||
for (const mutate of [
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.description += ' Remove the CI gate.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.description += ' Package the missing example.'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options.push({ label: 'Package the missing example', description: 'Fix the documented quickstart.' }); },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[0]!.label += ' and remove the CI gate'; },
|
||||
(c: NativePlanQuestionCall) => { c.questions[0]!.options[1]!.label = 'Partially wrong — fix the missing example'; },
|
||||
]) {
|
||||
const c = calls()[1]!; mutate(c);
|
||||
c.answers = { [c.questions[0]!.question]: c.questions[0]!.options[0]!.label };
|
||||
expect(classify(c)).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a second answered finding remains one substantive native call', () => {
|
||||
const c = calls()[1]!; const issue = calls()[2]!;
|
||||
c.questions.push(...issue.questions);
|
||||
Object.assign(c.answers!, issue.answers);
|
||||
expect(classify(c)).toBe(true);
|
||||
delete c.answers![issue.questions[0]!.question];
|
||||
c.unansweredQuestionIndices = [1];
|
||||
expect(classify(c)).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -5,8 +5,6 @@ import * as path from 'node:path';
|
||||
import { runGeneration } from '../scripts/gen-skill-docs';
|
||||
import { ALL_HOST_NAMES, getHostConfig } from '../hosts';
|
||||
|
||||
const ROOT = path.resolve(import.meta.dir, '..');
|
||||
|
||||
test('every host exposes the DX per-call rule before the pre-review audit and Step 0', async () => {
|
||||
const outputRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'devex-rule-free-'));
|
||||
try {
|
||||
@@ -128,387 +126,3 @@ test('every host exposes the DX per-call rule before the pre-review audit and St
|
||||
}
|
||||
} finally { fs.rmSync(outputRoot, { recursive: true, force: true }); }
|
||||
}, 20_000);
|
||||
|
||||
// Import the actual paid registration in a child with only its process boundary
|
||||
// mocked. Coverage and report validation remain the production predicates.
|
||||
function runDxRegistration(scenario: string) {
|
||||
const directory = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'devex-registration-free-')));
|
||||
const script = path.join(directory, 'registration.test.ts');
|
||||
const facts = path.join(directory, 'facts.json');
|
||||
fs.writeFileSync(script, `
|
||||
import { describe, expect, mock } from 'bun:test';
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { DEVEX_COUNT_FILES, planDevexCountFixture, isDevexReviewIssue, devexReviewModePick }
|
||||
from ${JSON.stringify(path.join(ROOT, 'test/helpers/devex-count-fixture.ts'))};
|
||||
import captured from ${JSON.stringify(path.join(ROOT, 'test/fixtures/devex-seed-coverage-ad-v3.json'))};
|
||||
const { assertReviewReportAtBottom: actualReport, devexStep0Boundary } =
|
||||
await import(${JSON.stringify(path.join(ROOT, 'test/helpers/claude-pty-runner.ts'))});
|
||||
const scenario = ${JSON.stringify(scenario)};
|
||||
const factsPath = ${JSON.stringify(facts)};
|
||||
const finalPlan = '# Reviewed DX plan\\n\\nThe five seeded gaps each have a recorded decision.\\n\\n## GSTACK REVIEW REPORT\\n\\nDX review complete.\\n';
|
||||
const facts = { checked: false, runnerCalls: 0, reportCalls: 0, judgeCalls: 0, planPath: '', finalPlan: '' };
|
||||
const save = () => fs.writeFileSync(factsPath, JSON.stringify(facts));
|
||||
mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/e2e-gate.ts'))}, () => ({
|
||||
describeE2ETier: tier => { expect(tier).toBe('periodic'); return describe; },
|
||||
}));
|
||||
mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/claude-pty-runner.ts'))}, () => ({
|
||||
devexStep0Boundary,
|
||||
assertReviewReportAtBottom: content => {
|
||||
facts.reportCalls++;
|
||||
facts.finalPlan = content;
|
||||
save();
|
||||
expect(content).toBe(fs.readFileSync(facts.planPath, 'utf8'));
|
||||
return actualReport(content);
|
||||
},
|
||||
runPlanSkillCounting: async opts => {
|
||||
facts.runnerCalls++;
|
||||
facts.planPath = opts.expectedPlanPath;
|
||||
save();
|
||||
expect(path.dirname(path.dirname(opts.expectedPlanPath))).toBe(${JSON.stringify(directory)});
|
||||
expect(fs.existsSync(path.dirname(opts.expectedPlanPath))).toBe(true);
|
||||
// The runner owns the seeded Git project. The caller owns only its report.
|
||||
expect(opts.cwd).toBeUndefined();
|
||||
expect(opts.skillName).toBe('plan-devex-review');
|
||||
expect(opts.slashCommand).toBe('/plan-devex-review');
|
||||
expect(opts.timeoutMs).toBe(1500000);
|
||||
expect(opts.reviewCountCeiling).toBe(Infinity);
|
||||
expect(opts.pickAUQ).toBe(devexReviewModePick);
|
||||
expect(opts.isReviewAUQ).toBe(isDevexReviewIssue);
|
||||
expect(opts.isLastStep0AUQ).toBe(devexStep0Boundary);
|
||||
expect(opts.env).toEqual({ QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' });
|
||||
expect(opts.followUpPrompt).toBe(planDevexCountFixture(opts.expectedPlanPath) +
|
||||
'\\nFinish this DX review; I will handle subsequent reviews manually.');
|
||||
expect(opts.fixtureFiles).toEqual(DEVEX_COUNT_FILES);
|
||||
expect(opts.followUpPrompt).toContain('Use DX POLISH');
|
||||
// These are the five unresolved contracts in the current native fixture.
|
||||
expect(opts.fixtureFiles['docs/current-contracts.md']).toContain('There is no skip flag or offline first-run path.');
|
||||
expect(opts.fixtureFiles['docs/package-contents.txt']).toContain('that file is absent');
|
||||
expect(opts.fixtureFiles['docs/api.md']).toContain('run_eval(dataset, evaluator)');
|
||||
expect(opts.fixtureFiles['docs/api.md']).toContain('run_batch(evaluator, dataset)');
|
||||
expect(opts.fixtureFiles['docs/api.md']).toContain('AuthError("request failed")');
|
||||
expect(opts.fixtureFiles['docs/api.md']).toContain('removes the old name immediately');
|
||||
facts.checked = true;
|
||||
save();
|
||||
if (scenario === 'throw') throw new Error('controlled DX runner failure');
|
||||
// Replay public question/reply evidence only. Its historical run did not
|
||||
// complete; the terminal/report below are controlled caller-boundary inputs.
|
||||
const transcript = { status: 'ready', calls: structuredClone(captured.attempts[0].calls), assistantMessages: [] };
|
||||
if (scenario.startsWith('missing-seed-')) transcript.calls.splice(Number(scenario.slice(-1)), 1);
|
||||
if (scenario === 'missing-native') transcript.status = 'missing';
|
||||
if (scenario === 'repeated-seed') transcript.calls = Array.from({length: 5}, (_, i) =>
|
||||
({ ...structuredClone(transcript.calls[0]), toolUseId: 'repeated-' + i }));
|
||||
if (scenario === 'batched') {
|
||||
const call = structuredClone(transcript.calls[0]);
|
||||
call.questions = transcript.calls.flatMap(item => item.questions);
|
||||
call.answers = Object.fromEntries(transcript.calls.flatMap(item => Object.entries(item.answers)));
|
||||
transcript.calls = [call];
|
||||
}
|
||||
if (scenario === 'pending') transcript.calls[0].answered = false;
|
||||
if (scenario !== 'missing-report') fs.writeFileSync(opts.expectedPlanPath,
|
||||
finalPlan + (scenario === 'trailing-report' ? '\\n## Unexpected follow-up\\n' : ''));
|
||||
return { outcome: scenario === 'timeout' ? 'timeout' : scenario === 'summary' ? 'completion_summary' : 'plan_ready',
|
||||
transcript, fingerprints: [], step0Count: 2, reviewCount: scenario.startsWith('missing-seed-') ? 100 : 5,
|
||||
elapsedMs: 100, evidence: 'controlled DX observation' };
|
||||
},
|
||||
}));
|
||||
mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/plan-review-decisions.ts'))}, () => ({
|
||||
evaluatePlanReviewDecisions: () => { facts.judgeCalls++; save(); throw new Error('unexpected paid judge'); },
|
||||
}));
|
||||
await import(${JSON.stringify(path.join(ROOT, 'test/skill-e2e-plan-devex-finding-count.test.ts'))});
|
||||
`);
|
||||
try {
|
||||
const child = Bun.spawnSync([process.execPath, 'test', script], {
|
||||
cwd: ROOT, timeout: 10_000,
|
||||
env: { PATH: process.env.PATH ?? '', HOME: directory, TMPDIR: directory, TMP: directory, TEMP: directory,
|
||||
GIT_CONFIG_NOSYSTEM: '1', ...(process.env.SystemRoot ? { SystemRoot: process.env.SystemRoot } : {}) },
|
||||
});
|
||||
const output = child.stdout.toString() + child.stderr.toString();
|
||||
expect(child.signalCode ?? null, output).toBeNull();
|
||||
expect(fs.existsSync(facts), output).toBe(true);
|
||||
const observed = JSON.parse(fs.readFileSync(facts, 'utf8'));
|
||||
expect(observed.checked, output).toBe(true);
|
||||
expect(observed.runnerCalls, output).toBe(1);
|
||||
expect(observed.judgeCalls, output).toBe(0);
|
||||
expect(fs.existsSync(path.dirname(observed.planPath)), 'actual paid finally must remove its owned report directory').toBe(false);
|
||||
return { output, exitCode: child.exitCode, observed };
|
||||
} finally { fs.rmSync(directory, { recursive: true, force: true }); }
|
||||
}
|
||||
|
||||
test('the actual DX registration supplies its complete native fixture and preserves runner failure', () => {
|
||||
const result = runDxRegistration('throw');
|
||||
expect(result.exitCode, result.output).toBe(1);
|
||||
expect(result.output).toContain('controlled DX runner failure');
|
||||
expect(result.observed.reportCalls).toBe(0);
|
||||
});
|
||||
|
||||
for (const outcome of ['success', 'summary']) test(`DX registration accepts completed seed decisions and the owned final report: ${outcome}`, () => {
|
||||
const result = runDxRegistration(outcome);
|
||||
expect(result.exitCode, result.output).toBe(0);
|
||||
expect(result.observed.reportCalls).toBe(1);
|
||||
expect(result.observed.finalPlan).toContain('## GSTACK REVIEW REPORT');
|
||||
});
|
||||
|
||||
for (const scenario of [
|
||||
...Array.from({ length: 5 }, (_, index) => 'missing-seed-' + index),
|
||||
'missing-native', 'repeated-seed', 'batched', 'pending',
|
||||
]) test(`DX registration requires complete distinct native coverage: ${scenario}`, () => {
|
||||
const result = runDxRegistration(scenario);
|
||||
expect(result.exitCode, result.output).toBe(1);
|
||||
expect(result.output).toContain('SEEDED COVERAGE FAIL');
|
||||
expect(result.observed.reportCalls).toBe(0);
|
||||
});
|
||||
|
||||
for (const scenario of ['missing-report', 'trailing-report', 'timeout']) test(`DX registration rejects incomplete delivery: ${scenario}`, () => {
|
||||
const result = runDxRegistration(scenario);
|
||||
expect(result.exitCode, result.output).toBe(1);
|
||||
expect(result.output).toContain(scenario === 'timeout' ? 'outcome=timeout' : 'D19 FAIL');
|
||||
expect(result.observed.reportCalls).toBe(scenario === 'trailing-report' ? 1 : 0);
|
||||
});
|
||||
|
||||
test('materialized DX references have working local links without inventing completed launch work', () => {
|
||||
const fixture = path.join(ROOT, 'test/fixtures/devex-existing-sdk');
|
||||
const files = ['README.md', 'docs/getting-started.md', 'docs/feedback.md', 'docs/reference-v1.md'];
|
||||
for (const file of files) {
|
||||
const body = fs.readFileSync(path.join(fixture, file), 'utf8');
|
||||
for (const [, target] of body.matchAll(/\[[^\]]+\]\(([^)]+)\)/g)) {
|
||||
const [relative, anchor] = target!.split('#');
|
||||
const destination = path.resolve(path.dirname(path.join(fixture, file)), relative || path.basename(file));
|
||||
expect(destination.startsWith(fixture + path.sep)).toBe(true);
|
||||
const linked = fs.readFileSync(destination, 'utf8');
|
||||
if (anchor) {
|
||||
const headings = [...linked.matchAll(/^#+ (.+)$/gm)].map(match => match[1]!.toLowerCase()
|
||||
.replace(/[^\w\s-]/g, '').replace(/\s/g, '-'));
|
||||
expect(headings, target).toContain(anchor);
|
||||
}
|
||||
}
|
||||
}
|
||||
const readme = fs.readFileSync(path.join(fixture, 'README.md'), 'utf8');
|
||||
expect(readme).toContain('no selected primary developer persona or peer-DX study');
|
||||
expect(readme).toContain('No first-run duration has\nbeen measured');
|
||||
expect(readme).toContain('There is no skip');
|
||||
expect(readme).toContain('no interactive demo or designed aha sequence');
|
||||
expect(readme).toContain('one ordinary passing case; it has no staged regression');
|
||||
const reference = fs.readFileSync(path.join(fixture, 'docs/reference-v1.md'), 'utf8');
|
||||
const guide = fs.readFileSync(path.join(fixture, 'docs/getting-started.md'), 'utf8');
|
||||
const errorLink = /^Reference: (docs\/[^#]+)#([^\s]+)$/m.exec(guide);
|
||||
expect(errorLink).not.toBeNull();
|
||||
expect(fs.readFileSync(path.join(fixture, errorLink![1]!), 'utf8')).toBe(reference);
|
||||
const errorHeadings = [...reference.matchAll(/^### (.+)$/gm)].map(match => match[1]!.toLowerCase().replace(/\s/g, '-'));
|
||||
expect(errorHeadings).toContain(errorLink![2]!);
|
||||
|
||||
expect(reference).toContain('cannot interrupt arbitrary application code or cap requests made by a separate');
|
||||
expect(reference).toContain('those calls have not been executed against\nthe SDK here');
|
||||
expect(reference).toContain('Fixture checks execute the local application files and explicit\ncontract doubles');
|
||||
});
|
||||
|
||||
test('materialized DX error examples identify their cause, bound and reachable code reference', () => {
|
||||
const reference = fs.readFileSync(path.join(ROOT, 'test/fixtures/devex-existing-sdk/docs/reference-v1.md'), 'utf8');
|
||||
const expected = [
|
||||
{ heading: '### SDK E002', count: 1, causes: ['MetricTypeError'], values: ['cases[0]'] },
|
||||
{ heading: '### SDK E003', count: 2, causes: ['DeadlineExceeded', 'ManagedProviderCostLimit'],
|
||||
values: ['deadline_seconds=20', 'max_cost_usd=0.25'] },
|
||||
];
|
||||
for (const spec of expected) {
|
||||
const start = reference.indexOf(spec.heading);
|
||||
const next = reference.indexOf('\n##', start + spec.heading.length);
|
||||
const section = reference.slice(start, next < 0 ? undefined : next);
|
||||
const blocks = [...section.matchAll(/```text\n([\s\S]*?)\n```/g)].map(match => match[1]!);
|
||||
expect(blocks, spec.heading).toHaveLength(spec.count);
|
||||
for (const [index, block] of blocks.entries()) {
|
||||
const code = spec.heading.replace('### SDK ', 'SDK_');
|
||||
expect(block.split('\n')[0]).toStartWith(code + ':');
|
||||
expect(block).toContain('Cause: ' + spec.causes[index]);
|
||||
expect(block).toContain(spec.values[index]!);
|
||||
expect(block).toMatch(/^Next: .+/m);
|
||||
const anchor = spec.heading.replace('### ', '').toLowerCase().replace(/ /g, '-');
|
||||
expect(block).toContain('Reference: docs/reference-v1.md#' + anchor);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
// Materialize the documented files in a temp directory. These controls test
|
||||
// application examples against an explicit contract stub, never the absent SDK.
|
||||
const dxDocs = () => Object.fromEntries(['README.md', 'docs/getting-started.md', 'docs/reference-v1.md']
|
||||
.map(file => [file, fs.readFileSync(path.join(ROOT, 'test/fixtures/devex-existing-sdk', file), 'utf8')]));
|
||||
function dxBlock(body: string, after: string, language: string): string {
|
||||
const offset = body.indexOf(after);
|
||||
expect(offset, `missing documented example: ${after}`).toBeGreaterThanOrEqual(0);
|
||||
const match = new RegExp('```' + language + '\\n([\\s\\S]*?)\\n```').exec(body.slice(offset));
|
||||
expect(match, `missing ${language} block after ${after}`).not.toBeNull();
|
||||
return match![1]!;
|
||||
}
|
||||
function runDxDocumentationControl(code: string, payload: unknown) {
|
||||
const directory = fs.mkdtempSync(path.join(os.tmpdir(), 'devex-doc-control-'));
|
||||
try {
|
||||
const script = path.join(directory, 'control.py');
|
||||
fs.writeFileSync(script, code);
|
||||
const input = path.join(directory, 'input.json');
|
||||
fs.writeFileSync(input, JSON.stringify(payload));
|
||||
// Like bin/gstack-config, support both Python command names. Windows
|
||||
// installs normally expose python.exe; avoid preferring its python3 Store alias.
|
||||
const python = (process.platform === 'win32' ? ['python', 'python3'] : ['python3', 'python'])
|
||||
.map(command => Bun.which(command)).find((command): command is string => command !== null);
|
||||
if (!python) throw new Error('Python 3 is required for the DX documentation controls');
|
||||
const child = Bun.spawnSync([python, script, input], { cwd: directory, timeout: 10_000,
|
||||
stdin: 'ignore', stdout: 'pipe', stderr: 'pipe' });
|
||||
expect(child.signalCode ?? null, child.stderr.toString()).toBeNull();
|
||||
expect(child.exitCode, child.stderr.toString()).toBe(0);
|
||||
return child.stdout.toString();
|
||||
} finally { fs.rmSync(directory, { recursive: true, force: true }); }
|
||||
}
|
||||
|
||||
test('materialized DX success blocks print the documented structured fields without assuming SDK repr', () => {
|
||||
const docs = dxDocs();
|
||||
const first = dxBlock(docs['README.md']!, '## Quick start', 'python');
|
||||
expect(first).toBe(dxBlock(docs['docs/getting-started.md']!, '## Neutral first evaluation', 'python'));
|
||||
const examples = [
|
||||
{ code: first, expected: dxBlock(docs['README.md']!, 'Shown application output', 'text') },
|
||||
{ code: dxBlock(docs['docs/getting-started.md']!, '## Neutral first evaluation', 'python'),
|
||||
expected: dxBlock(docs['docs/getting-started.md']!, 'Shown application output', 'text') },
|
||||
{ code: dxBlock(docs['docs/getting-started.md']!, '## Caller-owned metric for free text', 'python'),
|
||||
expected: dxBlock(docs['docs/getting-started.md']!, 'Shown free-text application output', 'text') },
|
||||
];
|
||||
for (const { code } of examples) { expect(code).not.toContain('print(result)'); expect(code).toContain('result.cases'); }
|
||||
const output = runDxDocumentationControl(String.raw`
|
||||
import contextlib, io, json, sys, types
|
||||
with open(sys.argv[1], encoding='utf-8') as source:
|
||||
examples = json.load(source)
|
||||
# Deliberate assumed-contract double: not an implementation of eval-sdk.
|
||||
def evaluate(target, cases, metric):
|
||||
result = []
|
||||
for case in cases:
|
||||
actual = target(case['inputs'])
|
||||
result.append(types.SimpleNamespace(actual=actual, expected=case['expected'], score=metric(actual, case['expected'])))
|
||||
return types.SimpleNamespace(cases=result)
|
||||
stub = types.ModuleType('eval_sdk'); stub.evaluate = evaluate; sys.modules['eval_sdk'] = stub
|
||||
for example in examples:
|
||||
output = io.StringIO(); namespace = {}
|
||||
with contextlib.redirect_stdout(output): exec(example['code'], namespace)
|
||||
assert output.getvalue().strip() == example['expected']
|
||||
assert json.loads(output.getvalue())[0]['score'] == 1.0
|
||||
# Preserve the caller-owned metric's mismatching-prose behavior separately.
|
||||
assert namespace['text_metric']('red', 'green') == 0.0
|
||||
print('three documented outputs match the contract stub; no SDK executed')
|
||||
`, examples);
|
||||
expect(output).toContain('three documented outputs match the contract stub; no SDK executed');
|
||||
});
|
||||
|
||||
test('materialized DX application client bounds actual local process timeouts, retries and reservations', () => {
|
||||
const guide = dxDocs()['docs/getting-started.md']!;
|
||||
const client = dxBlock(guide, 'Save as `bounded_client.py`', 'python');
|
||||
const transport = dxBlock(guide, 'Save as `fixture_transport.py`', 'python');
|
||||
const usage = dxBlock(guide, 'Use the application client in the callable', 'python');
|
||||
expect(guide).toContain('verified upper bound');
|
||||
expect(guide).toContain('not refunded');
|
||||
expect(guide).toContain('does not prove that a remote provider cancelled');
|
||||
const output = runDxDocumentationControl(String.raw`
|
||||
import json, pathlib, subprocess, sys, time, types
|
||||
with open(sys.argv[1], encoding='utf-8') as source:
|
||||
payload = json.load(source)
|
||||
pathlib.Path('bounded_client.py').write_text(payload['client'])
|
||||
pathlib.Path('fixture_transport.py').write_text(payload['transport'])
|
||||
from bounded_client import BoundedClient
|
||||
# Observe the real handles; subprocess.run still owns timeout/kill/wait.
|
||||
original_popen = subprocess.Popen
|
||||
children = []
|
||||
def capture_popen(*args, **kwargs):
|
||||
child = original_popen(*args, **kwargs)
|
||||
children.append(child)
|
||||
return child
|
||||
subprocess.Popen = capture_popen
|
||||
command = [sys.executable, 'fixture_transport.py']
|
||||
client = BoundedClient(command, timeout_seconds=1, max_attempts=2, total_cents=4, attempt_cents=2)
|
||||
assert client({'enabled': True}) == {'ready': True}
|
||||
assert client.reserved_cents == 2
|
||||
assert client({'enabled': False}) == {'ready': False}
|
||||
assert client.reserved_cents == 4
|
||||
try: client({'enabled': True}); raise AssertionError('budget exceeded')
|
||||
except RuntimeError as e: assert 'spending limit' in str(e)
|
||||
assert client.reserved_cents == 4
|
||||
# Calls sharing this application client also share one reservation ceiling.
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
client = BoundedClient(command, timeout_seconds=1, max_attempts=2, total_cents=4, attempt_cents=2)
|
||||
def concurrent_call(_):
|
||||
try: return client({'enabled': True})
|
||||
except RuntimeError as e:
|
||||
assert 'spending limit' in str(e); return None
|
||||
with ThreadPoolExecutor(max_workers=4) as pool: outputs = list(pool.map(concurrent_call, range(4)))
|
||||
assert outputs.count({'ready': True}) == 2 and outputs.count(None) == 2
|
||||
assert client.reserved_cents == 4
|
||||
# Real child failure/retry and real child timeout: no network or SDK involved.
|
||||
pathlib.Path('controlled_transport.py').write_text('''import sys, time
|
||||
mode = sys.argv[1]
|
||||
with open('attempts', 'a') as f: f.write('attempt\\n')
|
||||
if mode == 'stall': time.sleep(30)
|
||||
if mode == 'fail': sys.exit(75)
|
||||
''')
|
||||
for mode in ('fail', 'stall'):
|
||||
first_child = len(children)
|
||||
pathlib.Path('attempts').unlink(missing_ok=True)
|
||||
client = BoundedClient([sys.executable, 'controlled_transport.py', mode], timeout_seconds=0.2,
|
||||
max_attempts=2, total_cents=6, attempt_cents=2)
|
||||
started = time.monotonic()
|
||||
try: client({'enabled': True}); raise AssertionError('failed transport succeeded')
|
||||
except (RuntimeError, subprocess.TimeoutExpired): pass
|
||||
assert time.monotonic() - started < 3
|
||||
attempts = pathlib.Path('attempts').read_text().splitlines()
|
||||
owned_children = children[first_child:]
|
||||
assert len(attempts) == len(owned_children) == 2 and client.reserved_cents == 4
|
||||
for child in owned_children:
|
||||
assert child.poll() is not None, 'transport process leaked'
|
||||
assert child.wait(timeout=0) == child.returncode
|
||||
assert child.returncode != 0
|
||||
if mode == 'fail': assert child.returncode == 75
|
||||
# Insufficient reservation prevents even the retry from starting.
|
||||
pathlib.Path('attempts').unlink()
|
||||
client = BoundedClient([sys.executable, 'controlled_transport.py', 'fail'], timeout_seconds=1,
|
||||
max_attempts=2, total_cents=2, attempt_cents=2)
|
||||
try: client({'enabled': True}); raise AssertionError('budget exceeded')
|
||||
except RuntimeError as e: assert 'spending limit' in str(e)
|
||||
assert len(pathlib.Path('attempts').read_text().splitlines()) == 1
|
||||
assert client.reserved_cents == 2
|
||||
# The full shown usage sends independent limits to the SDK contract double.
|
||||
seen = []
|
||||
def evaluate(target, cases, metric, **options):
|
||||
seen.append(options)
|
||||
assert target(cases[0]['inputs']) == cases[0]['expected']
|
||||
return types.SimpleNamespace(cases=[])
|
||||
stub = types.ModuleType('eval_sdk'); stub.evaluate = evaluate; sys.modules['eval_sdk'] = stub
|
||||
exec(payload['usage'], {})
|
||||
assert seen == [{'deadline_seconds': 20, 'max_cost_usd': 0.25}]
|
||||
print('local timeout/retry/reservation bounds verified; no SDK/provider call')
|
||||
`, { client, transport, usage });
|
||||
expect(output).toContain('local timeout/retry/reservation bounds verified; no SDK/provider call');
|
||||
});
|
||||
|
||||
test('materialized DX CLI cases and import targets match the exact shown invocation in an offline contract double', () => {
|
||||
const reference = dxDocs()['docs/reference-v1.md']!;
|
||||
const cli = reference.slice(reference.indexOf('## CLI'), reference.indexOf('## Errors'));
|
||||
const payload = { app: dxBlock(cli, 'Save as `app.py`', 'python'), cases: dxBlock(cli, 'Save as `cases.json`', 'json'),
|
||||
command: dxBlock(cli, 'Run with the assumed SDK', 'bash') };
|
||||
expect(JSON.parse(payload.cases)).toEqual([{ inputs: { enabled: true }, expected: { ready: true } }]);
|
||||
const output = runDxDocumentationControl(String.raw`
|
||||
import argparse, importlib, json, pathlib, shlex, sys
|
||||
with open(sys.argv[1], encoding='utf-8') as source:
|
||||
payload = json.load(source)
|
||||
pathlib.Path('app.py').write_text(payload['app']); pathlib.Path('cases.json').write_text(payload['cases'])
|
||||
# Parse the documented command as an explicit contract double, not the absent CLI.
|
||||
args = shlex.split(payload['command']); assert args[:2] == ['eval-sdk', 'run']
|
||||
parser = argparse.ArgumentParser()
|
||||
for flag in ('target', 'cases', 'metric', 'deadline-seconds', 'max-cost-usd'): parser.add_argument('--' + flag, required=True)
|
||||
parser.add_argument('--no-input', action='store_true')
|
||||
options = parser.parse_args(args[2:])
|
||||
assert options.no_input and options.deadline_seconds == '20' and options.max_cost_usd == '0.25'
|
||||
def resolve(value):
|
||||
module, name = value.split(':'); return getattr(importlib.import_module(module), name)
|
||||
target, metric = resolve(options.target), resolve(options.metric)
|
||||
cases = json.loads(pathlib.Path(options.cases).read_text())
|
||||
assert isinstance(cases, list) and len(cases) == 1
|
||||
for case in cases:
|
||||
assert set(case) == {'inputs', 'expected'}
|
||||
assert metric(target(case['inputs']), case['expected']) == 1.0
|
||||
print('shown cases file, CLI arguments and import targets agree; no SDK executed')
|
||||
`, payload);
|
||||
expect(output).toContain('shown cases file, CLI arguments and import targets agree; no SDK executed');
|
||||
});
|
||||
@@ -1,77 +0,0 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import captured from './fixtures/devex-review-o-calls.json';
|
||||
import retry from './fixtures/devex-output-o-retry-call.json';
|
||||
import { isDevexReviewIssue } from './helpers/devex-count-fixture';
|
||||
import { nativePlanCallFingerprint } from './helpers/claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './helpers/plan-count-transcript';
|
||||
import { E2E_TOUCHFILES } from './helpers/touchfiles';
|
||||
|
||||
const calls = () => structuredClone(captured.calls) as NativePlanQuestionCall[];
|
||||
const fp = (call: NativePlanQuestionCall) => nativePlanCallFingerprint(call, 0, true);
|
||||
describe('documented expected-output gaps are substantive DX decisions', () => {
|
||||
test('all eleven actual O calls retain two setup and nine substantive decisions', () => {
|
||||
const original=calls();
|
||||
expect(original.map(call=>isDevexReviewIssue(fp(call)))).toEqual([false,false,true,true,true,true,true,true,true,true,true]);
|
||||
expect(original).toEqual(calls());
|
||||
// The later migration note and demo exemption are additional offered
|
||||
// changes, not retroactively included in earlier selected options.
|
||||
expect(original[3]!.answers![original[3]!.questions[0]!.question]).toContain('EVALKIT_SKIP_CI_CHECK');
|
||||
expect(original[10]!.questions[0]!.header).toBe('TODO: Demo CI exemption');
|
||||
});
|
||||
|
||||
test('the actual README output decision is counted before or after the review boundary', () => {
|
||||
const fingerprint=fp(calls()[7]!);fingerprint.promptSnippet='Short diagnostic text';
|
||||
for(const preReview of [true,false])expect(isDevexReviewIssue({...fingerprint,preReview})).toBe(true);
|
||||
});
|
||||
|
||||
test('equivalent output-documentation gaps do not depend on an issue number', () => {
|
||||
for(const question of [
|
||||
'The quickstart has no expected output, so developers cannot recognize a successful run.',
|
||||
'Expected output is absent from the documentation. Add an example of a successful command?',
|
||||
'The README does not show the output to expect. Should we document the success signal?',
|
||||
]) {
|
||||
const call=calls()[7]!;call.questions[0]!.header='Documentation gap';call.questions[0]!.question=question;
|
||||
call.answers={[question]:call.questions[0]!.options[0]!.label};expect(isDevexReviewIssue(fp(call))).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('missing answers, confirmation-only options and references to working output do not count', () => {
|
||||
const actual=calls()[7]!;
|
||||
for(const mutate of [
|
||||
(call:NativePlanQuestionCall)=>{call.answered=false;},
|
||||
(call:NativePlanQuestionCall)=>{call.answers={};},
|
||||
(call:NativePlanQuestionCall)=>{call.questions[0]!.header='Empathy check';},
|
||||
(call:NativePlanQuestionCall)=>{call.questions[0]!.options=[{label:'Read the documentation'},{label:'Continue the review'}];},
|
||||
(call:NativePlanQuestionCall)=>{const q=call.questions[0]!;q.question='The README already documents the expected output. Which file should I inspect next?';call.answers={[q.question]:q.options[0]!.label};},
|
||||
(call:NativePlanQuestionCall)=>{const q=call.questions[0]!;q.question='Which documentation should I inspect next?';call.answers={[q.question]:q.options[0]!.label};},
|
||||
]) {const call=structuredClone(actual);mutate(call);expect(isDevexReviewIssue(fp(call))).toBe(false);}
|
||||
const partial=calls()[1]!;partial.questions.push(actual.questions[0]!);partial.unansweredQuestionIndices=[1];
|
||||
expect(isDevexReviewIssue(fp(partial))).toBe(false);
|
||||
});
|
||||
|
||||
test('the actual retry sample-demo output proposal is the same documentation gap', () => {
|
||||
const call=structuredClone(retry.call) as NativePlanQuestionCall;
|
||||
expect(isDevexReviewIssue(fp(call))).toBe(true);
|
||||
expect(call.questions[0]!.options[0]!.label).toContain('Add to plan: include sample demo output in README');
|
||||
for (const question of [
|
||||
'The README shows no example output, so success is unspecified.',
|
||||
'Sample demo output is missing from the quickstart documentation.',
|
||||
]) {const next=structuredClone(call);next.questions[0]!.question=question;next.answers={[question]:next.questions[0]!.options[0]!.label};expect(isDevexReviewIssue(fp(next))).toBe(true);}
|
||||
for (const question of [
|
||||
'The README already shows sample demo output. Which documentation should I read next?',
|
||||
'Should we inspect example output in the README?',
|
||||
'README expected output is not missing.',
|
||||
'README already shows expected output; the missing item is a changelog.',
|
||||
'No expected output is missing from README.',
|
||||
'The README has no missing expected output. Should we show another example?',
|
||||
'No sample demo output is missing from README. Should we show another example?',
|
||||
]) {const next=structuredClone(call);next.questions[0]!.question=question;next.answers={[question]:next.questions[0]!.options[0]!.label};expect(isDevexReviewIssue(fp(next))).toBe(false);}
|
||||
call.questions[0]!.options=[{label:'Read the README'},{label:'Continue the review'}];
|
||||
expect(isDevexReviewIssue(fp(call))).toBe(false);
|
||||
});
|
||||
|
||||
test('the captured documentation regression remains a paid dependency', () => {
|
||||
for(const file of ['test/devex-output-o.test.ts','test/fixtures/devex-review-o-calls.json','test/fixtures/devex-output-o-retry-call.json'])
|
||||
expect(E2E_TOUCHFILES['plan-devex-finding-count']).toContain(file);
|
||||
});
|
||||
});
|
||||
Loaded 100 of 317 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user