mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 02:16:56 +02:00
* refactor(resolvers): split review.ts into MECE resolver modules (pure move) Move every function from scripts/resolvers/review.ts, unchanged, into: - review-dashboard.ts: review dashboard, plan-file review report - plan-gates.ts: approval check, exit-plan-mode gate, plan-file discovery, plan-completion audit/gate (ship + review), plan verification exec - spec-review.ts: both spec review loops, benefits-from, anti-shortcut clause - outside-voice-steps.ts: Codex second opinion, adversarial step, Codex plan review, Codex doc review, disabled-outside record - review-scope.ts: scope drift, cross-review dedup, shared-code reuse review.ts is deleted; index.ts imports the new modules. gen-skill-docs output is byte-identical for every host (--host all). Test imports and source-path references are re-pointed; the two source-text report/gate tests in gen-skill-docs.test.ts become behavioral renders across every consuming skill and host. All 46 touchfile entries that named review.ts now name all five modules, guarded by a recorded selection golden. * test(browse): black-box auth matrix for every server route and both surfaces Drives buildFetchHandler fetchLocal/fetchTunnel with no token, wrong token, root token, scoped token and the SSE cookie for all 33 routes, plus unmatched paths and wrong methods. Denials assert today's exact status, body and content type; allowed credentials assert the handler was reached. Written against the unchanged if-chain server so the W3 route-table refactor must keep it green. * refactor(shard-engine): move scripts/test-strict-output.ts to scripts/lib/shard-engine.ts The shared shard engine grows from the existing strict-output module (runShardChild, killProcessGroup, signal forwarding, strict classifier). scripts/test-strict-output.ts stays as a re-export so existing importers, mock.module paths and the strict-output/run-shard-child tests are unchanged. The engine inherits the global touchfile entry; the free runner's CLI-routing fixture copies the new module. * refactor(resolvers): decompose the three >150-line review resolvers (output-neutral) Split generateAdversarialStep, generateCodexPlanReview and generatePlanCompletionAuditInner into per-section helpers whose template literals are copied verbatim, so every function in the new modules is at or under 150 lines. gen-skill-docs output is byte-identical for every host (--host all, compared against96764e80with a fixed --link-root). * refactor(resolvers): one outside-voice failure policy (deliberate prose unification) outsideVoiceFailurePolicy(ctx, opts) in outside-voice.ts now renders the auth / timeout / empty-response bullets for all four call sites that hand-typed them (Codex second opinion, adversarial step, Codex plan review, design outside voices). Options are explicit per site (timeoutMinutes, onTimeout, stderrOnEmpty, fallback, escape) with no defaults. Deliberate generated-prose changes (every host): - office-hours: 'Fall back to <native> subagent.' becomes 'Fall back to the <native> subagent below.' - plan-devex-review: the plain 'Auth failure (stderr contains ...)' bullets become the canonical bold bullets; auth also triggers on 'API key'; 'auth failed' becomes 'authentication failed'. - review/ship adversarial: 'exceeded 9 minutes and was terminated' becomes 'timed out after 9 minutes and was terminated'; the timeout is still MISSING COVERAGE. - design outside voices: unchanged. Adds ratchet (d) (test/outside-voice-failure-policy.test.ts) with a reasoned allowlist for /codex's own CLI errors, the MISSING COVERAGE retention test, refreshed codex/factory ship goldens, and outside-voice.ts in every touchfile entry of review.ts and design.ts (selection golden extended). * test(pty): fake PTY session driver with an injectable clock through the runner launch seam The three plan-skill runners take an optional PtyDriver (launch, now, monotonic, sleep); omitted, they use the real launcher and clocks exactly as before. test/helpers/pty/fake-session.ts feeds scripted frames through that seam, and claude-pty-runner.runners.unit.test.ts runs observation, counting and floor for success, deadline timeout, permission prompt and plan-ready outcomes with no CLI or real timers. These cases must stay green unchanged through the W4 split and the runPtySession extraction. Touchfiles: every entry that lists claude-pty-runner.ts or pty-screen.ts now also lists test/helpers/pty/**. * refactor(shard-engine): run both lanes on the shared engine; lane policy injected Engine (scripts/lib/shard-engine.ts) gains the W2 primitives: per-shard tmp/Chromium sandbox + async cleanup backstop, log-path allocation and full-stream log capture, one duration-seed reader/writer with a lane predicate, LanePolicy (seed predicate + zero-execution verdict), strictShardStatus, and the shared CLI flag loop. runShardChild takes an optional companion (signal/settle) and waits a bounded 250ms to reap a wall-killed child. Free lane stops spawning shards itself: runFreeShard uses runShardChild with trackShardBrowser as the companion (win32 path unchanged: no process group, no negative-pid kill). Its sync state-dir removal stays lane policy. Paid lane uses the sandbox, log, seed, verdict and flag primitives; the hollow-shard guard applies PAID_LANE_POLICY. Lane outcomes are unchanged (free keeps >= 0 seeds and file-count zero-exec rule; paid keeps > 0 seeds, warning under selection and passed-empty under EVALS_ALL). paid-free-boundary's closure assertion now names the engine module, where the strict classifier lives. * test(shard-engine): engine unit tests, fixture-corpus equivalence, per-lane CLI parity - test/shard-engine.test.ts: failing/unhandled/module-load output fails both lanes, per-lane zero-execution and seed rules, whole-group kill on a wall timeout (both lanes), mocked-win32 path with no negative-pid kill, companion settle order, log capture, sandbox isolation, flag loop. - test/shard-engine-equivalence.test.ts + test/fixtures/shard-equivalence: seven outcome fixtures plus one real shard, run through both lanes and compared with classifications recorded from the base runners (96764e80). - test/shard-cli-parity.test.ts + test/fixtures/shard-cli-parity: flag set, defaults, validation errors and the Unknown argument error per lane match the base runners. * refactor(shard-engine): decompose runFreeShard and runPaidShard to <= 150 lines Output-neutral extraction under the fixture-corpus equivalence and runner tests: captureFreeStream, explainFreeVerdict and logFreeRecovery (free); paidShardCommand, settleShardSpool, settleBootstrapRetention and printLogTail (paid). The bootstrap scope-creation block that bootstrap-retention.test.ts evaluates stays verbatim. * refactor(pty): split claude-pty-runner.ts into test/helpers/pty/* behind a barrel Pure move: every line of the former 5,047-line runner lands verbatim in one module (four private helpers gain `export` for cross-module use): binary, screen (absorbs test/helpers/pty-screen.ts, which now re-exports it), launch, session (PtyDriver), judge, classify, auq, plan-native, boundaries, runners/{observation,counting,floor}. claude-pty-runner.ts re-exports the original public surface by name; pty/ modules import siblings directly. Tests that read the runner's source text: - rewritten as behavioral: the unit test's model-pin tripwire (fake CLI argv: fallback chain, --model before extraArgs, hermetic --strict-mcp-config), pty-skill-seeding-wiring (runners through the fake driver; launcher through a fake CLI reporting CLAUDE_CONFIG_DIR). The "three wrappers forward model" grep is replaced by the runners' fake-driver launch assertions. - pty-screen-session / pty-screen-supervision: stop copying runner source; they mock.module the real pty/screen.ts (and the fixture cleanup) instead. - re-pointed to the owning module (they execute a sliced runner body with injected boundaries; no seam exists for those boundaries yet): eng-seeded-completion-ai, plan-floor-permission, plan-create-prepublication, plan-count-completion; hermetic-wiring's source guard now reads pty/launch.ts and scans every pty/ module for raw process.env spreads. - plan-count-timeout and pty-output-wake mock the viewport at pty/screen.ts. * test(ratchet-c): enforcing module/function size ratchet and moved-code touchfile coverage Ratchet (c) ships enforcing: test/helpers/module-size.ts counts file and top-level function lengths by brace matching over masked source (strings, comments, regex literals and template text masked; ${} expressions kept), covering function declarations, arrow functions assigned to consts and route-table handler properties, with no parser dependency. Its self-test uses template literals and code-fence braces copied from scripts/resolvers/review.ts and design.ts. test/fixtures/module-size-ratchet.json binds scripts/lib/shard-engine.ts (<= 800 lines, <= 150 per function) and records the residual runner sizes (free 2352, paid 1921) as non-growth caps; allowlist entries are keyed on file plus matched text and need a reason. Failure output lists file:line, the rule, Fix: and the allowlist path. touchfiles.test.ts gains the moved-code superset check over test/fixtures/touchfile-move-goldens/ (W2 golden recorded at96764e80: test-strict-output.ts and test-paid-shards.ts global, test-free-shards.ts none). * refactor(browse): declared route table replaces the buildFetchHandler if-chain The ~1,300-line if-chain in buildFetchHandler becomes a route table: each entry declares method, path, auth kind and surfaces, and one auth gate in browse/src/routes/table.ts returns the per-kind denial (root-bearer, scoped, root-or-sse-cookie: 401 Unauthorized; root-token: 403 Root token required; extension-origin: 403 Forbidden). Unmatched requests take the declared fallthrough (root-bearer check, then plain-text 404). Handlers move to browse/src/routes/{core,pairing,pty,tokens,tunnel,activity,commands,files, inspector}.ts and receive a RouteContext with auth checks as functions instead of closing over factory locals. Dispatch order is unchanged: tunnel filter, beforeRoute overlay, gate, handler. TUNNEL_PATHS stays a literal in server.ts. Behavior-preserving: the black-box auth matrix from the previous commit passes unchanged. /memory and /inspector/events are declared root-bearer because the blanket check always ran before their SSE-cookie branch. Source-text route tests are rewritten as behavioral tests through buildFetchHandler or a route's real handler with a stub RouteContext (browse/test/route-test-harness.ts). Checks with no runtime seam are re-pointed to the route modules: Surface type, /inspector/events SSE helper, sanitizeReplacer imports, /pty-inject-scan sidecar-client import, and the ngrok config lookup and startTunnel wiring that stay in server.ts. * test(browse): stubbed-handler auth matrix and route inventory for the route table Every ROUTES entry runs through the real dispatcher and gate with stub handlers on each declared surface and six credentials; denials assert the exact status and body each auth kind returned at96764e8, admitted credentials assert the handler ran (with the gate's TokenInfo for scoped routes). Also pins the reviewed route inventory (method, path, auth kind, surfaces), that every entry declares auth and surfaces, that the table's tunnel paths equal the TUNNEL_PATHS literal with GET /connect admitted, the unmatched fallthrough, and that the root token is rejected on every tunnel route through buildFetchHandler. * test(browse): ratchet (b) keeps route dispatch inside the route table Scans browse/src/server.ts and browse/src/routes/*.ts for pathname comparisons; only the table matcher and the tunnel-surface filter are allowed, listed with reasons in browse/test/fixtures/route-dispatch-allowlist.json (keyed on file plus line text). Also checks every entry declares auth and surfaces and that gstack registers no beforeRoute overlay itself. Self-tests plant a violation and assert the file:line, Fix: and allowlist path in the message, that a shifted line stays allowlisted, and that a reasonless entry is rejected. * test: touchfile superset check for modules moved out of browse/src/server.ts Records the paid evals selected by touching browse/src/server.ts at96764e80(17 E2E, 1 LLM judge) and asserts every browse/src/routes/*.ts module selects a superset. The test reads every golden in test/fixtures/moved-module-selection/ so other moved-code goldens can sit beside it. * test(shard-engine): give non-timeout corpus fixtures CI headroom; keep the POSIX golden off the Windows lane Only the wall-timeout fixture keeps a 3s wall; the rest get 60s so a loaded host cannot turn a pass into a timeout. Base and branch runners still agree on every classification under the new walls. The Windows exclusion entry moves the free runner's ratchet (c) residual cap to 2356 lines. * refactor(pty): one runPtySession loop drives observation, counting and floor test/helpers/pty/session.ts owns launch -> start -> (poll -> tick)* -> timeout and the failure contract the three runners each hand-rolled: the run's own error wins over capture and close errors, close always runs, owned fixture cleanup runs last (also when launch fails). Each runner now supplies a PtySessionPlan: its boot/command step, poll cadence (2s observation/floor sleep; counting's output wake + 250ms coalesce), tick policy (permission handling, native identity, terminal rules stay per runner because they differ) and capture hooks. The runner bodies are decomposed into top-level steps so no function exceeds 150 lines; behavior is unchanged and the fake-driver cases from the first W4 commit pass unmodified. The counting capture step and the native completion-summary predicate are now named functions (countingCapture, isNativeCompletionSummary), so plan-create-prepublication and plan-count-completion call them directly instead of executing sliced source. The two harnesses that still execute a sliced runner body with injected boundaries (eng-seeded-completion-ai, plan-floor-permission) pass the PtyDriver seam instead of overriding Date/Bun.sleep. * test(ratchet-c): register route modules, review resolver modules and server.ts residual cap * refactor(pty): decompose launchClaudePty and engNumberedFindingAUQ under 150 lines launchClaudePty (349 lines) becomes launch preparation (args, hermetic child env, owned state roots), recorder creation, spawn, the trust-dialog watcher, close, and the session handle over one PtyProcess state object. The failure order is unchanged: abort the viewport, dispose any recorders created so far, dispose the viewport, rethrow. The --model / --strict-mcp-config ordering and seedSkills wiring stay pinned by the behavioral fake-CLI tests. engNumberedFindingAUQ (345 lines) keeps its guards and dispatch; each self-contained issue family (declared cache, library retry hooks, cache owner, injected singleton, shared writers, injected export) moves verbatim into its own function. Every pty/ module is now <= 800 lines and every top-level function <= 150 lines. * test(pty): split claude-pty-runner.unit.test.ts along the pty/ module seams The 188 unit tests move verbatim into claude-pty-runner.{screen,classify, auq,launch,plan-native,boundaries}.unit.test.ts (test names unchanged; each file imports only what it uses from the barrel). The five files that no longer read a SKILL.md template join the test-of-test ratchet baseline with a reason. * test(touchfiles): moved PTY modules keep their paid-eval selection test/fixtures/touchfile-selection/w4-pty.json records, at96764e8, the paid evals selected by touching test/helpers/claude-pty-runner.ts (20) and test/helpers/pty-screen.ts (20). touchfiles.test.ts now asserts every .ts file under test/helpers/pty/ (and pty/screen.ts for both sources) selects a superset, reading every golden in that directory so later moves can add one; a planted-violation case pins the report and its Fix line. * fix(browse): unexchanged pair setup keys no longer authenticate bearer requests validateToken accepted a gsk_setup_ key as a bearer on /command, /batch and /file (found while building the W3 auth matrix). A setup key now only authenticates the /connect exchange. * W1: one state-root owner (lib/state-root.ts + bin/gstack-state-root.sh), gstack-paths --explain and fail-stop, parity tests * W1: guarded migration of every executable state-root site; uninstall deletes only ~/.gstack Bins, careful/freeze hooks, setup, upgrade migrations, browse/src, design, ios-qa daemon, lib and scripts resolve the state root through bin/gstack-state-root.sh (bash) or lib/state-root.ts (TS). Bins source the twin and stop with a reinstall message when it is missing; hooks source it and never spawn gstack-paths. browse/src/config.ts and lib/cso/state.ts delegate to resolveStateRoot. Analytics writers and readers move together so the usage log stays one file. gstack-uninstall deletes state only at ~/.gstack, refuses (exit 2) when it resolves to /, $HOME or an ancestor, the checkout or the git root, and leaves any other resolved root in place with the removal command. Fixtures that copy single bins now copy the twin. * W1: privacy keys and trust-policy deny tiers merge across state roots; gstack-config reporting; test hermeticity readConfigKey / gstack_read_config_key return the most restrictive telemetry, memorable_recall, codex_reviews and update_check across the resolved root and ~/.gstack; other keys read the resolved root only. gstack-config set reports an overriding root with the exact override command, list shows the winning root and a root-variable disagreement line. gstack-gbrain-repo-policy get merges deny/read-only tiers. gstack-egress reads through readConfigKey. test-setup.ts strips inherited GSTACK_STATE_ROOT/GSTACK_STATE_DIR and redirects the legacy root. * W1: shared hook logging helper (hosts/claude/hooks/hook-log.ts) One hook-errors.log writer: root from resolveStateRoot, 0600 on every append, opt-in rate limit used only by memorable-user-prompt. The five hooks route through it. * W1: docs/state-root.md and README troubleshooting pointer Precedence table, a real --explain example, the move-your-state recipe, merged privacy keys, the uninstall rule, the resolver-failure fix, and the plugin-mode note (evidence gate: no official plugin distribution). * W1b: template and resolver prose resolve state through guarded gstack-paths; ratchet (a) Every gstack-paths eval in templates and resolvers carries the fail-stop guard; executable ~/.gstack paths in bash blocks (context recovery preamble, eureka log, analytics, project artifacts, upgrade snooze, setup-gbrain lock, retro snapshots, ship consent marker) use $GSTACK_STATE_ROOT, and the writer prose that pairs with them points at the printed PROJECT_DIR / RETRO_FILE. ship drops export GSTACK_STATE_ROOT. SKILL.md regenerated (claude + codex), ship goldens re-pinned, parity and context-budget caps raised to the measured sizes with notes. test/state-root-ratchet.test.ts enforces the rule with a reasoned allowlist; W1 touchfile entries plus a superset golden. * refactor: apply W1 state-root edits in W2/W3/W5-owned files; one moved-code touchfile golden for all workstreams * test: fold the moved-code touchfile golden into touchfiles.test.ts; fix integration fixture closure and caps * v1.91.11.0: CHANGELOG, TODOS, docs and conventions for the refactor wave * test: re-measure plan-ceo/design-consultation caps and ship goldens after the guarded plan-discovery and spec-review blocks; add the state-root twin to the workflow-boundaries fixture * fix(windows): migrations resolve their directory with either path separator; state-root parity compares under the HOME Git Bash actually sees * fix(review,ship): state plan-check timing after smoke expiry and test_stub Skip semantics (review workflow judge clarity) * test(qa-eval): webhook fix eval asks for the fix loop's post-repair probes; eight-scenario coverage stays in the report-only case and the harness recheck * test(qa-eval): re-pin the webhook prompt contract to the fix-loop stage; R29 coverage omissions stay bound by the report-only case * fix(review,ship): plan checks publish a checkpoint before each probe; only the smoke expiry stop is skipped * fix(qa): carry #2999's checkpoint receipt link, report-template line and full-revision placeholder (identical hunks) * test(qa-callers): disable git auto maintenance in the caller fixture Git 2.47+ runs auto maintenance detached after commit; on the CI runner's git 2.55 it rewrote .git/objects fan-out directories while the write observer was running, which surfaced as unauthorized mutations. Same gc.auto=0 / maintenance.auto=false guard the shared-libs fixture already uses. * test(plan-mode-no-op): require prose evidence for the prose-fallback members so a spinner-frame judge verdict cannot end the run as asked * test(ship-docsync): carry #2999's seeded-attempt docsync harness (identical files) The doc-sync fault cases replayed attempt 1 before reaching their gate and ran out of their 285s budget. The fixture now seeds attempt 1 and the parent starts at the gate under test. Taken byte-identical from origin/capy/audit-fix-wave (fb526898,e6ac813d,6ce10ff7,d0c53577,77cce3be). Local: stale-before, recovery and late-result 6/6 PASS (97-164s); the whole file 12/12 PASS.
568 lines
38 KiB
TypeScript
568 lines
38 KiB
TypeScript
import { expect, test } from 'bun:test';
|
||
import { readFileSync } from 'node:fs';
|
||
import { join } from 'node:path';
|
||
import { generateReviewArmy } from '../scripts/resolvers/review-army';
|
||
import { generateCrossReviewDedup, generateSharedCodeReuse, generateScopeDrift } from '../scripts/resolvers/review-scope';
|
||
import { generatePlanCompletionAuditReview, generatePlanCompletionAuditShip } from '../scripts/resolvers/plan-gates';
|
||
import { generateQAExploratory, generateQAReview } from '../scripts/resolvers/qa';
|
||
import { generateConfidenceCalibration } from '../scripts/resolvers/confidence';
|
||
import { HOST_PATHS, type TemplateContext } from '../scripts/resolvers/types';
|
||
|
||
const root = join(import.meta.dir, '..');
|
||
const skill = readFileSync(join(root, 'review/SKILL.md.tmpl'), 'utf8');
|
||
const adversarial = readFileSync(join(root, 'review/sections/adversarial.md.tmpl'), 'utf8');
|
||
|
||
test('review audits deliverables before deferring behavioral plan checks to the QA preflight', () => {
|
||
const ctx: TemplateContext = { skillName: 'review', tmplPath: '', host: 'claude', paths: HOST_PATHS.claude };
|
||
const audit = generatePlanCompletionAuditReview(ctx).replace(/\s+/g, ' ');
|
||
for (const contract of [
|
||
'Separate static audit evidence from behavioral checks',
|
||
'retain the exact command, expected outcome and source for Step 4.7',
|
||
'They remain pending execution, never DONE from a diff',
|
||
'A mixed item contributes to both lists',
|
||
'Zero audited deliverables do not waive these checks',
|
||
'Keep external-state and human-only checks under the existing audit rules',
|
||
'If only behavioral checks remain, report zero audited deliverables and retain their pending Step 4.7 list',
|
||
'For each audited deliverable, run the verification dispatch',
|
||
'File-existence checks and verified read-only content validators are static audit checks, not behavioral probes',
|
||
'Inspect the validator and its hooks before running it; verify read-only effects and access to the target',
|
||
'leave the item UNVERIFIABLE and defer the command to Step 4.7',
|
||
'Do not start applications, exercise APIs or mutate state during this audit',
|
||
'If found and verified safe above, invoke it',
|
||
]) expect(audit).toContain(contract);
|
||
expect(audit.indexOf('Inspect the validator and its hooks')).toBeLessThan(audit.indexOf('If found and verified safe above, invoke it'));
|
||
expect(audit).not.toContain('For each extracted plan item, run the verification dispatch');
|
||
const qa = generateQAReview(ctx);
|
||
expect(qa).toContain('Then run required plan checks, even after smoke expires');
|
||
expect(qa).toContain('Report clean/completed only when all required checks pass on current inputs');
|
||
});
|
||
|
||
test('review prior-Skip matching includes adversarial and Greptile findings without relaxing eligibility', () => {
|
||
const dedup = generateCrossReviewDedup({ skillName: 'review', tmplPath: '', host: 'claude', paths: HOST_PATHS.claude });
|
||
expect(dedup).toContain('For every combined finding, including core, specialist, exploratory QA, adversarial and valid actionable Greptile findings, check:');
|
||
expect(dedup).toContain('Suppress only when all conditions hold: the user skipped the same unchanged finding');
|
||
expect(dedup).toContain('same advisory/defect kind');
|
||
expect(dedup).toContain('Never use a skipped advisory to suppress a real defect');
|
||
expect(dedup).toContain('Only suppress `skipped` findings — never `fixed` or `auto-fixed`');
|
||
expect(skill).toContain('Run Step 5.0 severity/prior-skip dedup on all');
|
||
});
|
||
|
||
test('Review audit and prior-Skip clarifications do not route Ship through Review steps', () => {
|
||
for (const host of Object.keys(HOST_PATHS) as TemplateContext['host'][]) {
|
||
const ctx: TemplateContext = { skillName: 'ship', tmplPath: '', host, paths: HOST_PATHS[host] };
|
||
const audit = generatePlanCompletionAuditShip(ctx);
|
||
expect(audit).toContain('Step 8.1/9');
|
||
expect(audit).not.toContain('Step 4.7');
|
||
expect(audit).not.toContain('Separate static audit evidence from behavioral checks');
|
||
const dedup = generateCrossReviewDedup(ctx);
|
||
expect(dedup).toContain('Step 9.3: Cross-review finding dedup');
|
||
expect(dedup).not.toContain('Step 5.0');
|
||
expect(dedup).not.toContain('For every combined finding');
|
||
}
|
||
});
|
||
|
||
test('review collects every source before its single parent fix phase', () => {
|
||
const markers = [
|
||
'## Step 4: Critical pass', '### TODOS cross-reference',
|
||
'### Documentation staleness check', '{{SECTION:review-army}}',
|
||
'{{QA_REVIEW}}', '{{SECTION:adversarial}}',
|
||
'## Step 5: Fix-First Review', '{{CROSS_REVIEW_DEDUP}}',
|
||
'### Step 5a:', '### Step 5b:', '### Step 5c:', '### Step 5d:',
|
||
'## Step 5.8: Persist Eng Review result',
|
||
];
|
||
const positions = markers.map(marker => skill.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
expect(skill.match(/## Step 5: Fix-First Review/g)).toHaveLength(1);
|
||
expect(skill.replace(/\s+/g, ' ')).toContain('Do not edit reviewed source until Step 5');
|
||
expect(skill.replace(/\s+/g, ' ')).toContain('every dispatched reader has returned or is confirmed stopped');
|
||
});
|
||
|
||
test('review settles adversarial attempts before fixing and has one full-pass back edge', () => {
|
||
const generated = readFileSync(join(root, 'review/sections/adversarial.md'), 'utf8');
|
||
const decision = skill.slice(skill.indexOf('## Step 5.8: Persist Eng Review result')).replace(/\s+/g, ' ');
|
||
expect(generated).toContain('## Step 4.8: Adversarial review');
|
||
expect(generated).toContain("queued for the parent's Fix-First handling at Step 5; do not edit during Step 4.8");
|
||
expect(generated.replace(/\s+/g, ' ')).toContain('before the parent applies queued fixes');
|
||
expect(generated).toContain('Return all findings and structured-review decisions to Step 5');
|
||
expect(decision).toContain('A pass covers Steps 3–5, including all reviewers before fixes');
|
||
expect(decision).toContain('Below 3, repeat Steps 3–5 with a new REVIEW_START');
|
||
expect(decision).not.toContain('Route Step');
|
||
expect(decision).not.toContain('Steps 5.0–5d');
|
||
expect(decision).toContain('without a clean summary or a fourth pass');
|
||
});
|
||
|
||
test('review small-diff and failed-reader paths retain QA and the required adversarial pass', () => {
|
||
const army = generateReviewArmy({ skillName: 'review', tmplPath: 'review/SKILL.md.tmpl',
|
||
host: 'claude', paths: HOST_PATHS.claude });
|
||
expect(army).toContain("Continue to Step 4.6 with the core findings and an empty specialist list, then the parent's Exploratory QA step and Step 4.8 (adversarial review), then Step 5");
|
||
expect(army).not.toContain('design-lite');
|
||
expect(army).toContain('Missing dispatched coverage remains incomplete, never completed or clean');
|
||
expect(army).toContain('Continue independent Step 4.7 QA and Step 4.8 adversarial review');
|
||
expect(army).not.toContain("Exploratory QA step, then continue to Step 5.");
|
||
expect(army).not.toContain('If the Red Team subagent fails or times out, skip silently');
|
||
});
|
||
|
||
test('review defines QA confidence, severity, impact selection and numeric version comparison', () => {
|
||
const flat = skill.replace(/\s+/g, ' ');
|
||
expect(flat).toContain('For QA findings, assign confidence (1–10) from replay/code evidence');
|
||
expect(flat).toContain("retain Step 4.7's severity, not a severity inferred from confidence");
|
||
expect(flat).toContain('A probe is affected when its entrypoint, dependencies, contract or replay inputs change');
|
||
expect(flat).toContain('If impact is uncertain, rerun it');
|
||
expect(flat).toContain('Compare dotted version components as integers from left to right');
|
||
});
|
||
|
||
test('review emits scope check after the plan audit and before the checklist', () => {
|
||
const markers = [
|
||
'{{SCOPE_DRIFT}}', '{{SECTION:plan-completion}}',
|
||
'## Step 2: Read the checklist',
|
||
];
|
||
const positions = markers.map(marker => skill.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
const audit = readFileSync(join(root, 'review/sections/plan-completion.md'), 'utf8').replace(/\s+/g, ' ');
|
||
expect(audit).toContain('When continuing after the audit (no HIGH-impact gate, or option B/C), emit the single final Scope Check');
|
||
expect(audit).toContain("Step 1.5's provisional notes and this plan context");
|
||
expect(audit).toContain('Emit Step 1.5\'s Scope Check once without plan fields');
|
||
expect(skill).not.toContain('Finish Step 1.5 here');
|
||
});
|
||
|
||
test('review composes confidence-tagged findings into one final report with explicit incomplete coverage', () => {
|
||
const flat = skill.replace(/\s+/g, ' ');
|
||
expect(flat).toContain('Use CRITICAL/INFORMATIONAL labels in the finding format');
|
||
expect(flat).toContain("Step 5.8 combines these finding lines with the checklist's action groups");
|
||
const report = flat.slice(flat.indexOf('### Report the final review'), flat.indexOf('{{LEARNINGS_LOG}}'));
|
||
expect(report).toContain('Emit one final report, merging all reviewers rather than concatenating their reports');
|
||
expect(report).toContain('counts final unresolved non-advisory defects');
|
||
expect(report).toContain('State INCOMPLETE if `COMPLETED` is false, even when N=0');
|
||
expect(report).toContain("Use the checklist's action groups with confidence-tagged finding lines");
|
||
expect(report).toContain('Keep fixed, skipped and advisory items separate from unresolved defects; retain their dispositions');
|
||
expect(report).toContain("Append Step 4.7's single `## Exploratory QA and Verification Results` section");
|
||
expect(report).toContain('Neither coverage gaps nor advice are defects');
|
||
});
|
||
|
||
test('small-diff persistence uses an empty specialist map without manufacturing skipped coverage', () => {
|
||
const army = generateReviewArmy({ skillName: 'review', tmplPath: 'review/SKILL.md.tmpl',
|
||
host: 'claude', paths: HOST_PATHS.claude });
|
||
expect(army).toContain('For DIFF_LINES < 50, keep `specialists: {}`; do not manufacture per-specialist scope records');
|
||
expect(army).toContain('Otherwise record each considered specialist');
|
||
expect(skill).toContain("Use Step 4.6's `specialists` object unchanged, including its empty small-diff map");
|
||
const ship = readFileSync(join(root, 'ship/sections/review-army.md.tmpl'), 'utf8');
|
||
expect(ship).toContain('`specialists`: `{}` for a small-diff skip');
|
||
expect(skill).toContain('every required Step 4.7 probe passes');
|
||
expect(ship).toContain('all required probes pass');
|
||
});
|
||
|
||
test('caller QA runs charter and setup after resource loading and has a severity for unmatched functional failures', () => {
|
||
for (const skillName of ['review', 'ship']) {
|
||
const body = generateQAReview({ skillName, tmplPath: '', host: 'claude', paths: HOST_PATHS.claude });
|
||
const preparation = body.indexOf(skillName === 'review'
|
||
? '**1. Set the charter and isolation.**'
|
||
: 'Run the shared preflight;');
|
||
const probes = body.indexOf('**3. Run smoke and plan checks.**');
|
||
expect(preparation).toBeGreaterThan(-1);
|
||
if (skillName === 'review') {
|
||
const readiness = body.indexOf('**2. Check readiness and list required checks.**');
|
||
expect(readiness).toBeGreaterThan(preparation);
|
||
expect(probes).toBeGreaterThan(readiness);
|
||
expect(body.slice(preparation, readiness).replace(/\s+/g, ' ')).toContain('complete the shared isolation/permission preflight before setup');
|
||
} else expect(preparation).toBeGreaterThan(body.indexOf('**2. List required checks.**'));
|
||
expect(probes).toBeGreaterThan(preparation);
|
||
expect(body.replace(/\s+/g, ' ')).toContain('unmatched functional failures are `functional-contract`, `CRITICAL`');
|
||
expect(body).toContain('Setup/permission blockers are not defects');
|
||
expect(body).toContain('Test creation needs user approval');
|
||
expect(body).toContain('Return verified defects to Fix-First');
|
||
expect(body).not.toContain('for parent approval');
|
||
}
|
||
});
|
||
|
||
test('review prepares context and deduplicates before classifying findings', () => {
|
||
const positions = [
|
||
'## Step 3.5: Slop scan', '## Step 3.6: Gather review context',
|
||
'{{LEARNINGS_SEARCH}}', '{{ASIDE_RESEARCH}}', '## Step 4: Critical pass',
|
||
'## Step 5: Fix-First Review', '{{CROSS_REVIEW_DEDUP}}',
|
||
'**Keep decisions through fix cycles:**', '### Step 5a: Classify each finding',
|
||
].map(marker => skill.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
expect(skill).toContain('findings before Step 5a classification');
|
||
});
|
||
|
||
test('review owns the complete persistence contract after the adversarial read', () => {
|
||
const step = skill.slice(skill.indexOf('## Step 5.8: Persist Eng Review result'));
|
||
expect(skill.indexOf('{{SECTION:adversarial}}')).toBeLessThan(skill.indexOf('## Step 5.8: Persist Eng Review result'));
|
||
expect(adversarial).not.toContain('### Before persisting Eng Review (Step 5.8)');
|
||
for (const contract of [
|
||
'repeat Steps 3–5', 'at most 3 fix cycles', 'final zero-edit pass, reconcile',
|
||
'by structural identity', 'advisory/defect kind',
|
||
'original `evidence_paths`/`helper_target`', 'without requiring deleted pre-extraction blocks',
|
||
'Current findings determine recurring defects and unresolved counts', 'earlier fixes do not suppress them',
|
||
'snapshot_covered_paths',
|
||
'raw bytes equal the bound snapshot blobs', 'prior-cycle, supplied or prior-record coverage',
|
||
'REVIEW_START', 'COMPLETED', 'CONVERGED', 'CYCLES', 'Step 4.7',
|
||
'native Step 4.8 adversarial pass', 'means false, as does a failed native review',
|
||
'optional outside', 'their own incomplete records when unavailable', 'named-risk',
|
||
'zero counts', '`completed:false`', '`specialists`', '`findings`',
|
||
'verified exploratory QA findings',
|
||
'approved **and completed**', 'explicit Skip', 'sharedLibsFingerprint',
|
||
'`review_binding`', 'validated captured branch',
|
||
]) expect(step.toLowerCase().replace(/\s+/g, ' ')).toContain(contract.toLowerCase());
|
||
expect(step.indexOf('### 1. Re-review after edits'))
|
||
.toBeLessThan(step.indexOf('### 2. Fill the record'));
|
||
expect(step.indexOf('### 2. Fill the record'))
|
||
.toBeLessThan(step.indexOf('~/.claude/skills/gstack/bin/gstack-review-log'));
|
||
expect(step).toContain('`quality_score` is Step 4.6\'s specialist score');
|
||
expect(step).toContain('unresolved non-advisory core defects still count');
|
||
expect(step.indexOf('Pre-Landing Review: N issues (X critical, Y informational)'))
|
||
.toBeGreaterThan(step.indexOf('~/.claude/skills/gstack/bin/gstack-review-log'));
|
||
expect(step).toContain('`## Exploratory QA and Verification Results`');
|
||
});
|
||
|
||
test('review distinguishes required native coverage from optional outside coverage', () => {
|
||
const section = readFileSync(join(root, 'review/sections/adversarial.md'), 'utf8');
|
||
expect(section).toContain('Only this optional outside adversarial pass is non-blocking');
|
||
expect(section).not.toContain('All errors are non-blocking');
|
||
expect(section).toContain('The native pass is required for Step 5.8 completion');
|
||
expect(skill).toContain('Core findings use the confidence gates below');
|
||
expect(skill).toContain('Step 4.6 applies its specialist gates');
|
||
});
|
||
|
||
test('review identifies probe selection, report assets and the detected diff base', () => {
|
||
const generated = generateQAReview({ skillName: 'review', tmplPath: 'review/SKILL.md.tmpl',
|
||
host: 'claude', paths: HOST_PATHS.claude });
|
||
const checklist = readFileSync(join(root, 'review/checklist.md'), 'utf8');
|
||
expect(generated).toContain('Smoke: 5 minutes/12 probes, one success and the riskiest changed failure/edge');
|
||
expect(generated).toContain('Required even for small diffs or missing plans/servers');
|
||
expect(generated.replace(/\s+/g, ' ')).toContain("Use checklist severity; unmatched functional failures are `functional-contract`, `CRITICAL`");
|
||
expect(generated).toContain('Setup/permission blockers are not defects');
|
||
expect(generated).toContain('Test creation needs user approval');
|
||
expect(generated).toContain("Read QA\'s `templates/functional-report-template.md`. Title it");
|
||
expect(generated).toContain('Link every checkpoint');
|
||
expect(generated).toContain('No second report');
|
||
expect(checklist).toContain('merge-base diff from the caller');
|
||
expect(checklist).not.toContain('git diff origin/main');
|
||
});
|
||
|
||
test('caller QA defines execution, evidence ownership and report adaptation before handoff', () => {
|
||
for (const skillName of ['review', 'ship']) {
|
||
const generated = generateQAReview({ skillName, tmplPath: `${skillName}/SKILL.md.tmpl`,
|
||
host: 'claude', paths: HOST_PATHS.claude }).replace(/\s+/g, ' ');
|
||
for (const contract of [
|
||
'Only the parent runs report-only discovery',
|
||
'Never overwrite another run',
|
||
'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit',
|
||
'using the same procedure but no smoke guard; never reset the clock',
|
||
'Read agent/user updates and await results without batching them with reporting/logging',
|
||
'Compare each probe\'s recorded source, tests, contracts, commands and fixtures (or input fingerprint) with current inputs, even without updates',
|
||
'Re-review changed or uncertain coverage',
|
||
"Use checklist severity",
|
||
]) expect(generated).toContain(contract);
|
||
const shared = generateQAExploratory({ skillName: 'qa', tmplPath: '', host: 'claude', paths: HOST_PATHS.claude });
|
||
for (const contract of ['First demonstrate success: output AND durable effects',
|
||
'Wait for successful checkpoint publication before dispatch',
|
||
'Replay the exact failing command/request from the same initial fixture state']) {
|
||
expect(shared).toContain(contract);
|
||
}
|
||
if (skillName === 'review') {
|
||
expect(generated).toContain('Title it `## Exploratory QA and Verification Results`');
|
||
expect(generated).toContain('keep metadata/outcome tables');
|
||
expect(generated).toContain('demote other headings one level');
|
||
expect(generated).toContain('include it here under `### Browser results`');
|
||
expect(generated).toContain('other headings demoted two levels');
|
||
expect(generated).toContain('Keep browser/functional scores and outcomes separate');
|
||
expect(generated).toContain('save browser baseline/evidence normally');
|
||
expect(generated).toContain('Prepare one provisional QA section');
|
||
expect(generated).toContain('Update affected outcomes/checkpoint links through repairs/revalidation');
|
||
expect(generated).toContain('Continue to Step 4.8 even if blocked');
|
||
expect(generated).toContain('Step 5.8 appends this section once after final findings and decides completion');
|
||
} else {
|
||
expect(generated).toContain('PR section `## Exploratory QA`');
|
||
expect(generated).toContain('fields as subsections');
|
||
}
|
||
}
|
||
});
|
||
|
||
test('review section index follows the actual pre-fix execution order', () => {
|
||
const manifest = JSON.parse(readFileSync(join(root, 'review/sections/manifest.json'), 'utf8'));
|
||
expect(manifest.sections.map((section: { id: string }) => section.id)).toEqual([
|
||
'plan-completion', 'review-army', 'adversarial', 'shared-code-reuse',
|
||
]);
|
||
});
|
||
|
||
test('review finalization ownership: initialize invocation state and capture the core token before reading', () => {
|
||
const start = skill.slice(skill.indexOf('## Step 3: Get the diff'), skill.indexOf('## Step 3.4:'));
|
||
expect(start).toContain('one invocation action list and CYCLES=0');
|
||
expect(start).toContain('Keep both through re-reviews');
|
||
expect(skill.match(/CYCLES=0/g)).toHaveLength(1);
|
||
expect(start).toContain('gstack-review-log --start review\ngit diff "$DIFF_BASE"');
|
||
expect(start).toContain('Save the printed REVIEW_START for this core candidate before reading its diff');
|
||
expect(start).toContain('Each re-review captures a new token before reading, never at log time');
|
||
expect(start.replace(/\s+/g, ' ')).toContain('Earlier core tokens remain unused');
|
||
expect(start).toContain('Native/outside reviewer attempts own separate PASS_START tokens, not REVIEW_START');
|
||
expect(start).toContain('Step 5.8 finishes only the final core token');
|
||
});
|
||
|
||
test('review finalization ownership: late findings use Fix-First before the bounded parent transition', () => {
|
||
const flat = skill.replace(/\s+/g, ' ');
|
||
const markers = [
|
||
'{{SECTION:adversarial}}',
|
||
'## Step 5: Fix-First Review',
|
||
'Structured approval does not waive advisory/test_stub ASK gates',
|
||
'## Step 5.8: Persist Eng Review result',
|
||
'Edited: increment CYCLES once',
|
||
'No edits: fill the record below',
|
||
'### 2. Fill the record',
|
||
'--finish REVIEW_START',
|
||
];
|
||
const positions = markers.map(marker => flat.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
expect(flat).toContain('Structured approval does not waive advisory/test_stub ASK gates');
|
||
expect(flat).toContain('A pass covers Steps 3–5, including all reviewers before fixes');
|
||
expect(flat).toContain('at most 3 fix cycles');
|
||
expect(flat).toContain('Below 3, repeat Steps 3–5 with a new REVIEW_START');
|
||
expect(flat).toContain('At 3, persist `converged:false` and remaining findings');
|
||
const limit = flat.slice(flat.indexOf('At 3,'), flat.indexOf('No edits:'));
|
||
expect(limit).toContain('by filling and saving the record below');
|
||
expect(limit).toContain('Report nonconvergence and coverage gaps, then STOP this invocation');
|
||
expect(limit).toContain('without a clean summary or a fourth pass');
|
||
});
|
||
|
||
test('review finalization ownership: affected QA reuse cannot replace a full review or erase decisions', () => {
|
||
const step = skill.slice(skill.indexOf('## Step 5.8: Persist Eng Review result'));
|
||
const flat = step.replace(/\s+/g, ' ');
|
||
expect(flat).toContain('On a repeat, execute Steps 3–5 in order');
|
||
expect(flat).toContain('rerun affected probes after source, test, contract, command or fixture changes');
|
||
expect(flat).toContain("At Step 4.7, reuse only this invocation's unchanged-input QA evidence");
|
||
expect(flat).toContain('Reusing a probe never skips a review step');
|
||
expect(flat).toContain("final zero-edit pass, reconcile this invocation's actions with current findings");
|
||
expect(flat).toContain('original `evidence_paths`/`helper_target`');
|
||
expect(flat).toContain('Current findings determine recurring defects and unresolved counts; earlier fixes do not suppress them');
|
||
expect(flat).toContain('Re-read its final-snapshot supporting source and reconfirm the decision');
|
||
expect(flat).toContain('otherwise report its history without a reusable skip');
|
||
expect(flat).toContain('The logger computes `snapshot_covered_paths` from eligible paths whose raw bytes equal the bound snapshot blobs');
|
||
expect(flat).toContain('Never carry prior-cycle, supplied or prior-record coverage forward or build this proof yourself');
|
||
expect(flat).toContain('Fixed advice needs no skip coverage');
|
||
expect(flat).toContain('include this invocation\'s revalidated decisions');
|
||
});
|
||
|
||
test('review finalization ownership: required native completion and optional outside records stay separate', () => {
|
||
const step = skill.slice(skill.indexOf('### 2. Fill the record'));
|
||
const flat = step.replace(/\s+/g, ' ');
|
||
expect(flat).toContain('native Step 4.8 adversarial pass finish');
|
||
expect(flat).toContain('every required Step 4.7 probe passes');
|
||
expect(flat).toContain('Any failed, blocked, inconclusive or not-run required probe means false, as does a failed native review');
|
||
expect(flat).toContain('`/ship` named-risk acceptance cannot complete `/review`');
|
||
expect(flat).toContain('The required in-host adversarial result controls native completion');
|
||
expect(flat).toContain('Optional outside attempts keep their own incomplete records when unavailable');
|
||
expect(flat).toContain('cannot substitute for the native result, or vice versa');
|
||
expect(flat).toContain('structured-review gate still applies');
|
||
expect(flat).toContain('zero counts and `completed:false`');
|
||
expect(flat).toContain('`CONVERGED`: true only for a completed zero-edit pass');
|
||
});
|
||
|
||
test('review finalization ownership: finish only the final core token without log-time capture', () => {
|
||
const step = skill.slice(skill.indexOf('## Step 5.8: Persist Eng Review result'));
|
||
const flat = step.replace(/\s+/g, ' ');
|
||
expect(step.match(/--finish REVIEW_START/g)).toHaveLength(1);
|
||
expect(step).not.toContain('--start review');
|
||
expect(step).not.toContain('--finish PASS_START');
|
||
expect(flat).toContain('Never invent a binding or replace REVIEW_START at log time');
|
||
expect(flat).toContain('finish only the final core token');
|
||
expect(step).toContain('"completed":COMPLETED,"converged":CONVERGED,"cycles":CYCLES');
|
||
});
|
||
|
||
test('review finalization ownership: the plan audit retains its high-impact gate before the final scope check', () => {
|
||
const plan = readFileSync(join(root, 'review/sections/plan-completion.md.tmpl'), 'utf8');
|
||
expect(plan).toContain('INFORMATIONAL except for the HIGH-impact discrepancy question below');
|
||
expect(plan).toContain('resolve that gate before the final Scope Check');
|
||
expect(plan).not.toContain('never blocks the review');
|
||
expect(plan).toContain('{{PLAN_COMPLETION_AUDIT_REVIEW}}');
|
||
const audit = readFileSync(join(root, 'review/sections/plan-completion.md'), 'utf8');
|
||
const gate = audit.indexOf('**HIGH-impact discrepancies** trigger AskUserQuestion');
|
||
expect(gate).toBeGreaterThan(-1);
|
||
expect(gate).toBeLessThan(audit.indexOf('When continuing after the audit (no HIGH-impact gate, or option B/C)'));
|
||
expect(audit).toContain('then it gates via AskUserQuestion');
|
||
expect(audit).toContain('A ends this invocation before code review or implementation');
|
||
expect(audit).toContain('after implementation, start a fresh /review');
|
||
expect(audit).toContain('B queues the approved TODO changes for Step 5, not this read-only audit');
|
||
expect(audit).toContain('B/C continue to the final Scope Check and Step 2');
|
||
expect(audit).toContain('None of these choices authorizes shipping or waives required verification');
|
||
expect(skill).toContain("including Step 1.5's approved TODO changes");
|
||
});
|
||
|
||
test('review scope notes remain provisional until the plan section emits the only final scope check', () => {
|
||
const scope = generateScopeDrift({ skillName: 'review', tmplPath: 'review/SKILL.md.tmpl',
|
||
host: 'claude', paths: HOST_PATHS.claude }).replace(/\s+/g, ' ');
|
||
expect(scope).toContain('Keep these notes provisional. Next, execute the plan-completion section');
|
||
expect(scope).toContain('it resolves the HIGH-impact decision and emits the single final Scope Check before Step 2');
|
||
expect(scope).not.toContain('Scope Check: [CLEAN');
|
||
expect(scope).not.toContain('available plan-audit results');
|
||
});
|
||
|
||
test('review confidence uses its severity labels without an undefined P0 exception', () => {
|
||
const ctx: TemplateContext = { skillName: 'review', tmplPath: 'review/SKILL.md.tmpl',
|
||
host: 'claude', paths: HOST_PATHS.claude };
|
||
const confidence = generateConfidenceCalibration(ctx);
|
||
expect(confidence).toContain('Only report a suspected release-blocking catastrophe');
|
||
expect(confidence).toContain('widespread data loss, total outage or system-wide compromise');
|
||
expect(confidence).toContain('label it CRITICAL and explicitly speculative');
|
||
expect(confidence).toContain('[CRITICAL]');
|
||
expect(confidence).toContain('[CRITICAL|INFORMATIONAL]');
|
||
expect(confidence).not.toMatch(/\bP[012]\b/);
|
||
expect(confidence).toContain('If you cannot quote the motivating line(s), the finding is unverified');
|
||
const flat = confidence.replace(/\s+/g, ' ');
|
||
for (const rule of [
|
||
'| 9-10 | Specific code verifies a concrete bug or exploit. | Show normally |',
|
||
'| 7-8 | High-confidence pattern match; very likely correct. | Show normally |',
|
||
'| 5-6 | Moderate; could be a false positive. | Show with caveat:',
|
||
'Medium confidence, verify this is actually an issue',
|
||
'| 3-4 | Suspicious but may be fine. | Suppress from main report. Include in appendix only. |',
|
||
'| 1-2 | Speculation. | Only report a suspected release-blocking catastrophe',
|
||
'Quote the specific code line', 'file:line and verbatim text',
|
||
'For a missing field, quote its class definition; for a nullable value, its initialization; for a race, both sides',
|
||
'Force its confidence to 4-5: use 4 for appendix-only reporting, or 5 only when the finding belongs in the main report with the medium-confidence caveat below',
|
||
'Never invent speculative confidence 7+',
|
||
'read and quote their generating metaclass, descriptor, ORM Meta block, migration, decorator or schema',
|
||
'Missing literal names in the class body or grep results do not prove absence',
|
||
'If the user confirms a reported finding scored < 7 is real, log the corrected pattern as a learning',
|
||
]) expect(flat).toContain(rule);
|
||
expect(confidence.indexOf('Pre-emit verification gate')).toBeLessThan(confidence.indexOf('| Score |'));
|
||
expect(confidence).not.toContain('FP classes the gate kills');
|
||
expect(confidence).not.toContain('1539-framework-aware-review.md');
|
||
expect(generateConfidenceCalibration({ ...ctx, skillName: 'ship' })).toContain('Only report if severity would be P0');
|
||
});
|
||
|
||
test('review names the lifecycle and record owners before using their persistence rules', () => {
|
||
const start = skill.slice(skill.indexOf('## Step 3: Get the diff'), skill.indexOf('## Step 3.4:')).replace(/\s+/g, ' ');
|
||
expect(start).toContain('An invocation is this /review run; a pass reviews one candidate before any fixes');
|
||
expect(start).toContain('REVIEW_START / PASS_START | Opaque start receipts from the logger');
|
||
expect(start).toContain('a matching key alone never proves a prior Skip is reusable');
|
||
expect(start).toContain("`review_binding` | The logger's proof tying a finished review to its captured candidate");
|
||
expect(start).toContain('`snapshot_covered_paths` | Supporting advice files the logger proved byte-identical to that candidate');
|
||
expect(start).toContain('Used by the prior-Skip checker, never supplied by the reviewer');
|
||
});
|
||
|
||
test('review invocation-local advice reuse keeps raw-source and changed-decision gates', () => {
|
||
const decisions = skill.slice(skill.indexOf('**Keep decisions through fix cycles:**'),
|
||
skill.indexOf('### Step 5a:')).replace(/\s+/g, ' ');
|
||
for (const contract of [
|
||
'Immediately save completed AUTO-FIX/fix and explicit Skip actions',
|
||
'keeping defects separate from advice',
|
||
"retain the helper's fingerprint, `advisory`, `evidence_paths` and `helper_target`",
|
||
're-read every supporting caller and helper destination',
|
||
'including secondary callers and transformed/indirect paths',
|
||
'Compare their raw source with the decision evidence',
|
||
'Unrelated auto-fixes do not reopen unchanged identity, contract and tradeoffs',
|
||
'Material proposal, behavior, migration or risk changes require a new question',
|
||
"cannot suppress new/recurring defects or replace Step 5.0's prior-review checker",
|
||
]) expect(decisions).toContain(contract);
|
||
});
|
||
|
||
for (const skillName of ['review', 'ship']) {
|
||
const ctx: TemplateContext = { skillName, tmplPath: `${skillName}/SKILL.md.tmpl`,
|
||
host: 'claude', paths: HOST_PATHS.claude };
|
||
const army = generateReviewArmy(ctx);
|
||
const flat = army.replace(/\s+/g, ' ');
|
||
|
||
test(`${skillName} clarity: terminal failure permits independent work but never certifies coverage`, () => {
|
||
expect(flat).toContain('Confirm that each task has finished or is stopped');
|
||
expect(flat).toContain('A timeout alone does not prove termination');
|
||
expect(flat).toContain("If a reader or writer is still active, wait; if its state is unknown, inspect its task/process status");
|
||
expect(flat).toContain("If you cannot confirm it stopped, use the parent's Fix-First stop path without edits");
|
||
expect(flat).toContain('Continue independent evidence collection after a terminal failure');
|
||
expect(flat).toContain('Missing dispatched coverage remains incomplete, never completed or clean');
|
||
expect(flat).not.toContain('Specialists are additive — partial results are better than no results');
|
||
const redTeam = flat.slice(flat.indexOf('### Red Team dispatch'));
|
||
expect(redTeam).toContain('confirm it stopped and record its review as incomplete, just as for other specialists');
|
||
expect(redTeam).toContain('original specialist outputs and rerun stages 1–7');
|
||
});
|
||
|
||
test(`${skillName} clarity: ordered specialist merge separates validation from scoring and provenance`, () => {
|
||
const markers = ['#### 1. Parse outputs', '#### 2. Validate severity',
|
||
'#### 3. Identify and merge', '#### 4. Apply specialist confidence gates',
|
||
'#### 5. Score and present specialists', '#### 6. Save specialist activity',
|
||
'#### 7. Hand off to Fix-First'];
|
||
const positions = markers.map(marker => army.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
const validation = flat.slice(flat.indexOf('#### 2.'), flat.indexOf('#### 3.'));
|
||
expect(validation).toContain('core and specialist findings');
|
||
expect(validation).toContain('remove `advisory` and retain its `CRITICAL` severity');
|
||
expect(validation).toContain('Never downgrade severity');
|
||
expect(validation).toContain('Valid INFORMATIONAL advisories remain advisory');
|
||
const merge = flat.slice(flat.indexOf('#### 3.'), flat.indexOf('#### 4.'));
|
||
expect(merge.indexOf('Partition defects and advisories')).toBeLessThan(merge.indexOf('grouping by fingerprint'));
|
||
for (const gate of ['sharedLibsFingerprint', 'literal JSON on stdin', 'never trust a supplied hash',
|
||
'Missing/malformed metadata cannot deduplicate', 'highest confidence', '+1 (cap at 10)',
|
||
'distinct specialists', 'all source names', 'Core findings never earn a specialist confidence boost']) {
|
||
expect(merge).toContain(gate);
|
||
}
|
||
for (const gate of ['Confidence 7+', 'Confidence 5-6', 'Confidence 3-4', 'Confidence 1-2']) expect(army).toContain(gate);
|
||
const scoring = flat.slice(flat.indexOf('#### 5.'), flat.indexOf('#### 6.'));
|
||
expect(scoring).toContain('Only specialist findings enter this header and `quality_score`; core findings do not');
|
||
expect(scoring).toContain('NON-advisory');
|
||
expect(scoring).toContain('quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))');
|
||
expect(scoring).toContain('unresolved-defect totals');
|
||
expect(flat).toContain('Advisory findings COUNT in the stats `findings` field');
|
||
expect(flat).toContain('Count only findings that specialist actually returned');
|
||
expect(flat).toContain('core-only advice must not create a specialist dispatch or finding');
|
||
expect(flat).toContain('ASK-only');
|
||
});
|
||
|
||
test(`${skillName} clarity: shared-code reuse gives executable decisions and retains checker safeguards`, () => {
|
||
const reuse = generateSharedCodeReuse(ctx).replace(/\s+/g, ' ');
|
||
const markers = ['1. **Read the evidence.**', '2. **Run the checker.**',
|
||
'3. **Act on its result.**', '4. **Persist through the logger.**'];
|
||
const positions = markers.map(marker => reuse.indexOf(marker));
|
||
expect(positions.every(position => position >= 0)).toBe(true);
|
||
expect(positions).toEqual([...positions].sort((a, b) => a - b));
|
||
for (const gate of ['first-party authored provenance', 'all supporting callers and the helper destination',
|
||
'--check-shared-libs REVIEW_START', "<<'GSTACK_SHARED_LIBS_REUSE_JSON'", 'literal JSON on stdin',
|
||
'`reusable: true`', 'False, command failure or unreadable output', 'never suppression',
|
||
'Do not supply your own snapshot, prior record or coverage', 'sharedLibsFingerprint',
|
||
'without consuming/replacing it', 'actual repo, raw branch and current snapshot',
|
||
'completed/converged', 'verified binding', 'explicit Skip', 'logger-versioned `snapshot_covered_paths`',
|
||
'older unversioned coverage', 'Sanitized branch names are not identity', 'canReuseSharedLibsAdvisory',
|
||
'byte-for-byte with its blob', 'assume-unchanged, skip-worktree', 'sparse index', 'symlinks/ancestors',
|
||
'submodules', 'ignored/outside or unreadable', 'active/unknown Git filters', 'encodings and line conversion',
|
||
'fsmonitor and optional locks', 'never uses external diff/textconv', 'Unknown evidence fails closed',
|
||
'logger recomputes final coverage', 'Real defects retain normal Fix-First handling independently']) {
|
||
expect(reuse).toContain(gate);
|
||
}
|
||
});
|
||
}
|
||
|
||
test('review clarity: settlement gates edits separately from incomplete required coverage', () => {
|
||
const fix = skill.slice(skill.indexOf('## Step 5: Fix-First Review'), skill.indexOf('{{CROSS_REVIEW_DEDUP}}')).replace(/\s+/g, ' ');
|
||
expect(fix).toContain('every dispatched reader has returned or is confirmed stopped');
|
||
expect(fix).toContain('active or unknown reader/writer');
|
||
expect(fix).toContain('persist incomplete at Step 5.8 and STOP without edits');
|
||
expect(fix).toContain('Terminal failure does not block fixes from independent evidence');
|
||
expect(fix).toContain('Missing required output still makes the pass incomplete');
|
||
});
|
||
|
||
test('review clarity: Greptile reply choices never substitute for Fix-First approval', () => {
|
||
const fix = skill.slice(skill.indexOf('## Step 5: Fix-First Review'), skill.indexOf('{{CROSS_REVIEW_DEDUP}}'));
|
||
expect(fix).toContain('VALID & ACTIONABLE Greptile findings');
|
||
const greptile = skill.slice(skill.indexOf('### Greptile comment resolution'), skill.indexOf('## Step 5.8:'));
|
||
const flat = greptile.replace(/\s+/g, ' ');
|
||
expect(flat).toContain('Step 5c alone supplies A) Fix / B) Skip');
|
||
expect(flat).not.toContain('A: Fix it now, B: Acknowledge, C: False positive');
|
||
expect(flat).toContain('reply decisions, not code approval');
|
||
expect(flat).toContain('B) Propose a code change');
|
||
expect(flat).toContain('return to Steps 5c–5d with an ASK proposal');
|
||
expect(flat).toContain('Show the exact change and any `test_stub`; wait for approval before editing');
|
||
expect(flat).toContain('no new fix permission');
|
||
});
|
||
|
||
test('ship review clarity: parent settlement gate precedes classification and cannot waive coverage', () => {
|
||
const ship = readFileSync(join(root, 'ship/sections/review-army.md.tmpl'), 'utf8');
|
||
const gate = ship.slice(ship.indexOf('## Step 9.4:'), ship.indexOf('1. **Classify')).replace(/\s+/g, ' ');
|
||
expect(gate).toContain("Before edits, inspect every dispatched reader/writer's handle");
|
||
expect(gate).toContain('Wait for return or confirm termination');
|
||
expect(gate).toContain('otherwise log incomplete through items 5–6 and STOP without edits');
|
||
expect(gate).toContain('After terminal failure, independent evidence may support fixes');
|
||
expect(gate).toContain('missing dispatched output still blocks continuation, even with a QA exception');
|
||
});
|