mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 18:36:54 +02:00
* refactor(resolvers): split review.ts into MECE resolver modules (pure move) Move every function from scripts/resolvers/review.ts, unchanged, into: - review-dashboard.ts: review dashboard, plan-file review report - plan-gates.ts: approval check, exit-plan-mode gate, plan-file discovery, plan-completion audit/gate (ship + review), plan verification exec - spec-review.ts: both spec review loops, benefits-from, anti-shortcut clause - outside-voice-steps.ts: Codex second opinion, adversarial step, Codex plan review, Codex doc review, disabled-outside record - review-scope.ts: scope drift, cross-review dedup, shared-code reuse review.ts is deleted; index.ts imports the new modules. gen-skill-docs output is byte-identical for every host (--host all). Test imports and source-path references are re-pointed; the two source-text report/gate tests in gen-skill-docs.test.ts become behavioral renders across every consuming skill and host. All 46 touchfile entries that named review.ts now name all five modules, guarded by a recorded selection golden. * test(browse): black-box auth matrix for every server route and both surfaces Drives buildFetchHandler fetchLocal/fetchTunnel with no token, wrong token, root token, scoped token and the SSE cookie for all 33 routes, plus unmatched paths and wrong methods. Denials assert today's exact status, body and content type; allowed credentials assert the handler was reached. Written against the unchanged if-chain server so the W3 route-table refactor must keep it green. * refactor(shard-engine): move scripts/test-strict-output.ts to scripts/lib/shard-engine.ts The shared shard engine grows from the existing strict-output module (runShardChild, killProcessGroup, signal forwarding, strict classifier). scripts/test-strict-output.ts stays as a re-export so existing importers, mock.module paths and the strict-output/run-shard-child tests are unchanged. The engine inherits the global touchfile entry; the free runner's CLI-routing fixture copies the new module. * refactor(resolvers): decompose the three >150-line review resolvers (output-neutral) Split generateAdversarialStep, generateCodexPlanReview and generatePlanCompletionAuditInner into per-section helpers whose template literals are copied verbatim, so every function in the new modules is at or under 150 lines. gen-skill-docs output is byte-identical for every host (--host all, compared against96764e80with a fixed --link-root). * refactor(resolvers): one outside-voice failure policy (deliberate prose unification) outsideVoiceFailurePolicy(ctx, opts) in outside-voice.ts now renders the auth / timeout / empty-response bullets for all four call sites that hand-typed them (Codex second opinion, adversarial step, Codex plan review, design outside voices). Options are explicit per site (timeoutMinutes, onTimeout, stderrOnEmpty, fallback, escape) with no defaults. Deliberate generated-prose changes (every host): - office-hours: 'Fall back to <native> subagent.' becomes 'Fall back to the <native> subagent below.' - plan-devex-review: the plain 'Auth failure (stderr contains ...)' bullets become the canonical bold bullets; auth also triggers on 'API key'; 'auth failed' becomes 'authentication failed'. - review/ship adversarial: 'exceeded 9 minutes and was terminated' becomes 'timed out after 9 minutes and was terminated'; the timeout is still MISSING COVERAGE. - design outside voices: unchanged. Adds ratchet (d) (test/outside-voice-failure-policy.test.ts) with a reasoned allowlist for /codex's own CLI errors, the MISSING COVERAGE retention test, refreshed codex/factory ship goldens, and outside-voice.ts in every touchfile entry of review.ts and design.ts (selection golden extended). * test(pty): fake PTY session driver with an injectable clock through the runner launch seam The three plan-skill runners take an optional PtyDriver (launch, now, monotonic, sleep); omitted, they use the real launcher and clocks exactly as before. test/helpers/pty/fake-session.ts feeds scripted frames through that seam, and claude-pty-runner.runners.unit.test.ts runs observation, counting and floor for success, deadline timeout, permission prompt and plan-ready outcomes with no CLI or real timers. These cases must stay green unchanged through the W4 split and the runPtySession extraction. Touchfiles: every entry that lists claude-pty-runner.ts or pty-screen.ts now also lists test/helpers/pty/**. * refactor(shard-engine): run both lanes on the shared engine; lane policy injected Engine (scripts/lib/shard-engine.ts) gains the W2 primitives: per-shard tmp/Chromium sandbox + async cleanup backstop, log-path allocation and full-stream log capture, one duration-seed reader/writer with a lane predicate, LanePolicy (seed predicate + zero-execution verdict), strictShardStatus, and the shared CLI flag loop. runShardChild takes an optional companion (signal/settle) and waits a bounded 250ms to reap a wall-killed child. Free lane stops spawning shards itself: runFreeShard uses runShardChild with trackShardBrowser as the companion (win32 path unchanged: no process group, no negative-pid kill). Its sync state-dir removal stays lane policy. Paid lane uses the sandbox, log, seed, verdict and flag primitives; the hollow-shard guard applies PAID_LANE_POLICY. Lane outcomes are unchanged (free keeps >= 0 seeds and file-count zero-exec rule; paid keeps > 0 seeds, warning under selection and passed-empty under EVALS_ALL). paid-free-boundary's closure assertion now names the engine module, where the strict classifier lives. * test(shard-engine): engine unit tests, fixture-corpus equivalence, per-lane CLI parity - test/shard-engine.test.ts: failing/unhandled/module-load output fails both lanes, per-lane zero-execution and seed rules, whole-group kill on a wall timeout (both lanes), mocked-win32 path with no negative-pid kill, companion settle order, log capture, sandbox isolation, flag loop. - test/shard-engine-equivalence.test.ts + test/fixtures/shard-equivalence: seven outcome fixtures plus one real shard, run through both lanes and compared with classifications recorded from the base runners (96764e80). - test/shard-cli-parity.test.ts + test/fixtures/shard-cli-parity: flag set, defaults, validation errors and the Unknown argument error per lane match the base runners. * refactor(shard-engine): decompose runFreeShard and runPaidShard to <= 150 lines Output-neutral extraction under the fixture-corpus equivalence and runner tests: captureFreeStream, explainFreeVerdict and logFreeRecovery (free); paidShardCommand, settleShardSpool, settleBootstrapRetention and printLogTail (paid). The bootstrap scope-creation block that bootstrap-retention.test.ts evaluates stays verbatim. * refactor(pty): split claude-pty-runner.ts into test/helpers/pty/* behind a barrel Pure move: every line of the former 5,047-line runner lands verbatim in one module (four private helpers gain `export` for cross-module use): binary, screen (absorbs test/helpers/pty-screen.ts, which now re-exports it), launch, session (PtyDriver), judge, classify, auq, plan-native, boundaries, runners/{observation,counting,floor}. claude-pty-runner.ts re-exports the original public surface by name; pty/ modules import siblings directly. Tests that read the runner's source text: - rewritten as behavioral: the unit test's model-pin tripwire (fake CLI argv: fallback chain, --model before extraArgs, hermetic --strict-mcp-config), pty-skill-seeding-wiring (runners through the fake driver; launcher through a fake CLI reporting CLAUDE_CONFIG_DIR). The "three wrappers forward model" grep is replaced by the runners' fake-driver launch assertions. - pty-screen-session / pty-screen-supervision: stop copying runner source; they mock.module the real pty/screen.ts (and the fixture cleanup) instead. - re-pointed to the owning module (they execute a sliced runner body with injected boundaries; no seam exists for those boundaries yet): eng-seeded-completion-ai, plan-floor-permission, plan-create-prepublication, plan-count-completion; hermetic-wiring's source guard now reads pty/launch.ts and scans every pty/ module for raw process.env spreads. - plan-count-timeout and pty-output-wake mock the viewport at pty/screen.ts. * test(ratchet-c): enforcing module/function size ratchet and moved-code touchfile coverage Ratchet (c) ships enforcing: test/helpers/module-size.ts counts file and top-level function lengths by brace matching over masked source (strings, comments, regex literals and template text masked; ${} expressions kept), covering function declarations, arrow functions assigned to consts and route-table handler properties, with no parser dependency. Its self-test uses template literals and code-fence braces copied from scripts/resolvers/review.ts and design.ts. test/fixtures/module-size-ratchet.json binds scripts/lib/shard-engine.ts (<= 800 lines, <= 150 per function) and records the residual runner sizes (free 2352, paid 1921) as non-growth caps; allowlist entries are keyed on file plus matched text and need a reason. Failure output lists file:line, the rule, Fix: and the allowlist path. touchfiles.test.ts gains the moved-code superset check over test/fixtures/touchfile-move-goldens/ (W2 golden recorded at96764e80: test-strict-output.ts and test-paid-shards.ts global, test-free-shards.ts none). * refactor(browse): declared route table replaces the buildFetchHandler if-chain The ~1,300-line if-chain in buildFetchHandler becomes a route table: each entry declares method, path, auth kind and surfaces, and one auth gate in browse/src/routes/table.ts returns the per-kind denial (root-bearer, scoped, root-or-sse-cookie: 401 Unauthorized; root-token: 403 Root token required; extension-origin: 403 Forbidden). Unmatched requests take the declared fallthrough (root-bearer check, then plain-text 404). Handlers move to browse/src/routes/{core,pairing,pty,tokens,tunnel,activity,commands,files, inspector}.ts and receive a RouteContext with auth checks as functions instead of closing over factory locals. Dispatch order is unchanged: tunnel filter, beforeRoute overlay, gate, handler. TUNNEL_PATHS stays a literal in server.ts. Behavior-preserving: the black-box auth matrix from the previous commit passes unchanged. /memory and /inspector/events are declared root-bearer because the blanket check always ran before their SSE-cookie branch. Source-text route tests are rewritten as behavioral tests through buildFetchHandler or a route's real handler with a stub RouteContext (browse/test/route-test-harness.ts). Checks with no runtime seam are re-pointed to the route modules: Surface type, /inspector/events SSE helper, sanitizeReplacer imports, /pty-inject-scan sidecar-client import, and the ngrok config lookup and startTunnel wiring that stay in server.ts. * test(browse): stubbed-handler auth matrix and route inventory for the route table Every ROUTES entry runs through the real dispatcher and gate with stub handlers on each declared surface and six credentials; denials assert the exact status and body each auth kind returned at96764e8, admitted credentials assert the handler ran (with the gate's TokenInfo for scoped routes). Also pins the reviewed route inventory (method, path, auth kind, surfaces), that every entry declares auth and surfaces, that the table's tunnel paths equal the TUNNEL_PATHS literal with GET /connect admitted, the unmatched fallthrough, and that the root token is rejected on every tunnel route through buildFetchHandler. * test(browse): ratchet (b) keeps route dispatch inside the route table Scans browse/src/server.ts and browse/src/routes/*.ts for pathname comparisons; only the table matcher and the tunnel-surface filter are allowed, listed with reasons in browse/test/fixtures/route-dispatch-allowlist.json (keyed on file plus line text). Also checks every entry declares auth and surfaces and that gstack registers no beforeRoute overlay itself. Self-tests plant a violation and assert the file:line, Fix: and allowlist path in the message, that a shifted line stays allowlisted, and that a reasonless entry is rejected. * test: touchfile superset check for modules moved out of browse/src/server.ts Records the paid evals selected by touching browse/src/server.ts at96764e80(17 E2E, 1 LLM judge) and asserts every browse/src/routes/*.ts module selects a superset. The test reads every golden in test/fixtures/moved-module-selection/ so other moved-code goldens can sit beside it. * test(shard-engine): give non-timeout corpus fixtures CI headroom; keep the POSIX golden off the Windows lane Only the wall-timeout fixture keeps a 3s wall; the rest get 60s so a loaded host cannot turn a pass into a timeout. Base and branch runners still agree on every classification under the new walls. The Windows exclusion entry moves the free runner's ratchet (c) residual cap to 2356 lines. * refactor(pty): one runPtySession loop drives observation, counting and floor test/helpers/pty/session.ts owns launch -> start -> (poll -> tick)* -> timeout and the failure contract the three runners each hand-rolled: the run's own error wins over capture and close errors, close always runs, owned fixture cleanup runs last (also when launch fails). Each runner now supplies a PtySessionPlan: its boot/command step, poll cadence (2s observation/floor sleep; counting's output wake + 250ms coalesce), tick policy (permission handling, native identity, terminal rules stay per runner because they differ) and capture hooks. The runner bodies are decomposed into top-level steps so no function exceeds 150 lines; behavior is unchanged and the fake-driver cases from the first W4 commit pass unmodified. The counting capture step and the native completion-summary predicate are now named functions (countingCapture, isNativeCompletionSummary), so plan-create-prepublication and plan-count-completion call them directly instead of executing sliced source. The two harnesses that still execute a sliced runner body with injected boundaries (eng-seeded-completion-ai, plan-floor-permission) pass the PtyDriver seam instead of overriding Date/Bun.sleep. * test(ratchet-c): register route modules, review resolver modules and server.ts residual cap * refactor(pty): decompose launchClaudePty and engNumberedFindingAUQ under 150 lines launchClaudePty (349 lines) becomes launch preparation (args, hermetic child env, owned state roots), recorder creation, spawn, the trust-dialog watcher, close, and the session handle over one PtyProcess state object. The failure order is unchanged: abort the viewport, dispose any recorders created so far, dispose the viewport, rethrow. The --model / --strict-mcp-config ordering and seedSkills wiring stay pinned by the behavioral fake-CLI tests. engNumberedFindingAUQ (345 lines) keeps its guards and dispatch; each self-contained issue family (declared cache, library retry hooks, cache owner, injected singleton, shared writers, injected export) moves verbatim into its own function. Every pty/ module is now <= 800 lines and every top-level function <= 150 lines. * test(pty): split claude-pty-runner.unit.test.ts along the pty/ module seams The 188 unit tests move verbatim into claude-pty-runner.{screen,classify, auq,launch,plan-native,boundaries}.unit.test.ts (test names unchanged; each file imports only what it uses from the barrel). The five files that no longer read a SKILL.md template join the test-of-test ratchet baseline with a reason. * test(touchfiles): moved PTY modules keep their paid-eval selection test/fixtures/touchfile-selection/w4-pty.json records, at96764e8, the paid evals selected by touching test/helpers/claude-pty-runner.ts (20) and test/helpers/pty-screen.ts (20). touchfiles.test.ts now asserts every .ts file under test/helpers/pty/ (and pty/screen.ts for both sources) selects a superset, reading every golden in that directory so later moves can add one; a planted-violation case pins the report and its Fix line. * fix(browse): unexchanged pair setup keys no longer authenticate bearer requests validateToken accepted a gsk_setup_ key as a bearer on /command, /batch and /file (found while building the W3 auth matrix). A setup key now only authenticates the /connect exchange. * W1: one state-root owner (lib/state-root.ts + bin/gstack-state-root.sh), gstack-paths --explain and fail-stop, parity tests * W1: guarded migration of every executable state-root site; uninstall deletes only ~/.gstack Bins, careful/freeze hooks, setup, upgrade migrations, browse/src, design, ios-qa daemon, lib and scripts resolve the state root through bin/gstack-state-root.sh (bash) or lib/state-root.ts (TS). Bins source the twin and stop with a reinstall message when it is missing; hooks source it and never spawn gstack-paths. browse/src/config.ts and lib/cso/state.ts delegate to resolveStateRoot. Analytics writers and readers move together so the usage log stays one file. gstack-uninstall deletes state only at ~/.gstack, refuses (exit 2) when it resolves to /, $HOME or an ancestor, the checkout or the git root, and leaves any other resolved root in place with the removal command. Fixtures that copy single bins now copy the twin. * W1: privacy keys and trust-policy deny tiers merge across state roots; gstack-config reporting; test hermeticity readConfigKey / gstack_read_config_key return the most restrictive telemetry, memorable_recall, codex_reviews and update_check across the resolved root and ~/.gstack; other keys read the resolved root only. gstack-config set reports an overriding root with the exact override command, list shows the winning root and a root-variable disagreement line. gstack-gbrain-repo-policy get merges deny/read-only tiers. gstack-egress reads through readConfigKey. test-setup.ts strips inherited GSTACK_STATE_ROOT/GSTACK_STATE_DIR and redirects the legacy root. * W1: shared hook logging helper (hosts/claude/hooks/hook-log.ts) One hook-errors.log writer: root from resolveStateRoot, 0600 on every append, opt-in rate limit used only by memorable-user-prompt. The five hooks route through it. * W1: docs/state-root.md and README troubleshooting pointer Precedence table, a real --explain example, the move-your-state recipe, merged privacy keys, the uninstall rule, the resolver-failure fix, and the plugin-mode note (evidence gate: no official plugin distribution). * W1b: template and resolver prose resolve state through guarded gstack-paths; ratchet (a) Every gstack-paths eval in templates and resolvers carries the fail-stop guard; executable ~/.gstack paths in bash blocks (context recovery preamble, eureka log, analytics, project artifacts, upgrade snooze, setup-gbrain lock, retro snapshots, ship consent marker) use $GSTACK_STATE_ROOT, and the writer prose that pairs with them points at the printed PROJECT_DIR / RETRO_FILE. ship drops export GSTACK_STATE_ROOT. SKILL.md regenerated (claude + codex), ship goldens re-pinned, parity and context-budget caps raised to the measured sizes with notes. test/state-root-ratchet.test.ts enforces the rule with a reasoned allowlist; W1 touchfile entries plus a superset golden. * refactor: apply W1 state-root edits in W2/W3/W5-owned files; one moved-code touchfile golden for all workstreams * test: fold the moved-code touchfile golden into touchfiles.test.ts; fix integration fixture closure and caps * v1.91.11.0: CHANGELOG, TODOS, docs and conventions for the refactor wave * test: re-measure plan-ceo/design-consultation caps and ship goldens after the guarded plan-discovery and spec-review blocks; add the state-root twin to the workflow-boundaries fixture * fix(windows): migrations resolve their directory with either path separator; state-root parity compares under the HOME Git Bash actually sees * fix(review,ship): state plan-check timing after smoke expiry and test_stub Skip semantics (review workflow judge clarity) * test(qa-eval): webhook fix eval asks for the fix loop's post-repair probes; eight-scenario coverage stays in the report-only case and the harness recheck * test(qa-eval): re-pin the webhook prompt contract to the fix-loop stage; R29 coverage omissions stay bound by the report-only case * fix(review,ship): plan checks publish a checkpoint before each probe; only the smoke expiry stop is skipped * fix(qa): carry #2999's checkpoint receipt link, report-template line and full-revision placeholder (identical hunks) * test(qa-callers): disable git auto maintenance in the caller fixture Git 2.47+ runs auto maintenance detached after commit; on the CI runner's git 2.55 it rewrote .git/objects fan-out directories while the write observer was running, which surfaced as unauthorized mutations. Same gc.auto=0 / maintenance.auto=false guard the shared-libs fixture already uses. * test(plan-mode-no-op): require prose evidence for the prose-fallback members so a spinner-frame judge verdict cannot end the run as asked * test(ship-docsync): carry #2999's seeded-attempt docsync harness (identical files) The doc-sync fault cases replayed attempt 1 before reaching their gate and ran out of their 285s budget. The fixture now seeds attempt 1 and the parent starts at the gate under test. Taken byte-identical from origin/capy/audit-fix-wave (fb526898,e6ac813d,6ce10ff7,d0c53577,77cce3be). Local: stale-before, recovery and late-result 6/6 PASS (97-164s); the whole file 12/12 PASS.
730 lines
55 KiB
TypeScript
730 lines
55 KiB
TypeScript
import * as fs from 'node:fs';
|
|
import * as path from 'node:path';
|
|
import * as os from 'node:os';
|
|
import { createHash } from 'node:crypto';
|
|
import { spawnSync } from 'node:child_process';
|
|
import { isDeepStrictEqual } from 'node:util';
|
|
import { readQaDeadline } from '../../lib/qa-deadline';
|
|
import { readQaCaptureRecord } from '../../lib/qa-evidence';
|
|
import type { SkillTestResult } from './session-runner';
|
|
import { runSkillTest, SESSION_DRAIN_GRACE_MS } from './session-runner';
|
|
import { CAPTURE_MS } from './eval-budgets';
|
|
import { refreshHermeticSkillRuntime } from './hermetic-skill-runtime';
|
|
import { seedHermeticGstackHome } from './hermetic-env';
|
|
import { observeQAWrites, type QAWriteObservation } from './qa-functional-observer';
|
|
import { nativeCalls, readQACheckpointFiles, validateQACheckpoints } from './qa-checkpoint-evidence';
|
|
import { ownedPath } from './qa-functional-fixture';
|
|
import { qaEvidenceCommand, qaNativeCapture, qaProducerReceipt, type QaEvidenceContext } from './qa-evidence-producer';
|
|
import { qaCaptureArtifacts } from './qa-functional-evidence';
|
|
|
|
export type QaCaller = 'review' | 'ship';
|
|
export const QA_CALLER_ROOT = path.resolve(import.meta.dir, '../..');
|
|
|
|
export function callerExcerpt(source: string, start: string, end: string): string {
|
|
const first = source.indexOf(start);
|
|
const last = source.indexOf(end, first + start.length);
|
|
if (first < 0 || last < 0 || last <= first || source.indexOf(start, first + start.length) >= 0) {
|
|
throw new Error(`Missing or ambiguous caller excerpt boundary: ${start} -> ${end}`);
|
|
}
|
|
return source.slice(first, last);
|
|
}
|
|
|
|
export function qaCallerInstructions(caller: QaCaller, root = QA_CALLER_ROOT): string {
|
|
const read = (file: string) => fs.readFileSync(path.join(root, file), 'utf8');
|
|
const source = read(`${caller}/SKILL.md`);
|
|
const excerpt = caller === 'review'
|
|
? callerExcerpt(source, '## Step 4: Critical pass (core review)', '## Step 5.8: Persist Eng Review result')
|
|
: callerExcerpt(source,
|
|
'> **STOP.** Before auditing plan completion, verification, and scope drift (Step 8),',
|
|
'> **STOP.** Before addressing Greptile review comments');
|
|
const review = caller === 'review' ? excerpt : read('ship/sections/review-army.md');
|
|
if (!review.includes(`### Step ${caller === 'review' ? '4.7' : '9.2.1'}: Exploratory QA (before Fix-First)`) || !review.includes('sections/exploratory.md')) {
|
|
throw new Error(`Generated /${caller} parent is missing the exploratory QA integration`);
|
|
}
|
|
const army = read(`${caller}/sections/review-army.md`);
|
|
if (!army.toLowerCase().includes('exploratory') || !army.includes('50')) {
|
|
throw new Error(`Generated /${caller} specialist bypass is missing its exploration handoff`);
|
|
}
|
|
for (const id of ['scope', 'exploratory', 'system-functional']) {
|
|
const body = read(`qa/sections/${id}.md`);
|
|
if (body.trim().length < 200 || /\{\{[A-Z_]+/.test(body)) {
|
|
throw new Error(`Incomplete generated QA resource: ${id}`);
|
|
}
|
|
}
|
|
if (caller === 'ship') {
|
|
const plan = read('ship/sections/plan-completion.md');
|
|
if (!plan.includes('## Step 8.1: Plan Verification') || plan.includes('### 3. Invoke /qa-only inline')) {
|
|
throw new Error('Generated plan verification has not been integrated');
|
|
}
|
|
}
|
|
return excerpt;
|
|
}
|
|
|
|
export interface CallerTool {
|
|
id: string;
|
|
parent: string | null;
|
|
name: string;
|
|
input: Record<string, unknown>;
|
|
output: string;
|
|
failed: boolean;
|
|
index: number;
|
|
resultIndex: number;
|
|
messageId?: string;
|
|
sessionId?: string;
|
|
handoffContent?: string;
|
|
}
|
|
|
|
const UNCHANGED_READ = 'Wasted call — file unchanged since your last Read. Refer to that earlier tool_result instead.';
|
|
|
|
export function callerTools(transcript: unknown[]): CallerTool[] {
|
|
const failures: string[] = [];
|
|
const calls = nativeCalls(transcript, failures);
|
|
if (failures.length) throw new Error(failures.join('; '));
|
|
const tools: CallerTool[] = [];
|
|
for (const call of calls) {
|
|
const start = transcript[call.start] as any;
|
|
const event = transcript[call.end] as any;
|
|
const tool: CallerTool = { id: call.id, parent: call.parent, name: call.name, input: call.input,
|
|
output: call.output, failed: call.failed, index: call.start, resultIndex: call.end,
|
|
messageId: start.message.id, sessionId: start.session_id };
|
|
tools.push(tool);
|
|
const native = event.tool_use_result;
|
|
if (!tool.failed && tool.name === 'Read' && typeof tool.input.file_path === 'string'
|
|
&& tool.input.file_path.endsWith('/HANDOFF.md') && Object.keys(tool.input).length === 1
|
|
&& typeof tool.sessionId === 'string' && event.session_id === tool.sessionId
|
|
&& event.message.content.length === 1 && native?.file?.filePath === tool.input.file_path) {
|
|
if (native.type === 'text' && typeof native.file.content === 'string') {
|
|
const lines = native.file.content.split('\n');
|
|
if (native.file.startLine === 1 && native.file.numLines === lines.length && native.file.totalLines === lines.length
|
|
&& tool.output === lines.map((line: string, i: number) => `${i + 1}\t${line}`).join('\n')) {
|
|
tool.handoffContent = native.file.content;
|
|
}
|
|
} else if (native.type === 'file_unchanged' && tool.output === UNCHANGED_READ) {
|
|
const prior = tools.findLast(read => read.name === 'Read' && read.parent === tool.parent && read.sessionId === tool.sessionId
|
|
&& read.input.file_path === tool.input.file_path
|
|
&& read.resultIndex >= 0 && read.resultIndex < tool.index && (!tool.messageId || read.messageId !== tool.messageId));
|
|
tool.handoffContent = prior?.handoffContent;
|
|
}
|
|
}
|
|
}
|
|
return tools;
|
|
}
|
|
|
|
export interface CallerProbe {
|
|
id: string;
|
|
charter: string;
|
|
input: string;
|
|
snapshot: string;
|
|
status: 'pass' | 'fail' | 'blocked' | 'inconclusive';
|
|
stdout: string;
|
|
stderr: string;
|
|
exit: number | null;
|
|
}
|
|
|
|
export interface CallerReceipt {
|
|
status: 'pass' | 'fail' | 'blocked' | 'inconclusive';
|
|
probes: string[];
|
|
remaining: string[];
|
|
}
|
|
|
|
const literalCallerArgument = /(?:[^\s'"\\;&|<>`$(){}*?\[\]~#]+|'[^'\r\n]*'|"[^"\\$`\r\n]*")/.source;
|
|
const literalCallerCLI = new RegExp(`^bun (?:scripts/probe\\.ts|cli\\.ts)(?: ${literalCallerArgument})?$`);
|
|
const literalCallerProbe = new RegExp(`^bun scripts/probe\\.ts(?: ${literalCallerArgument})?$`);
|
|
const literalDeadlineCommand = new RegExp(`^bun (${literalCallerArgument}) (start|status|run) (${literalCallerArgument})(?: (.*))?$`);
|
|
const literalCallerListing = new RegExp(`^ls(?: -[la]+)?(?: --)?(?: ${literalCallerArgument})*$`);
|
|
const callerDiffMode = '(--stat|--numstat|--name-only|--name-status)';
|
|
const literalCallerDiff = new RegExp(`^git diff(?: ${callerDiffMode})?(?: (origin/main|"\\$DIFF_BASE"))?(?: ${callerDiffMode})?(?: --(?: ${literalCallerArgument})+)?$`);
|
|
const callerDiffPreface = 'DIFF_BASE=$(git merge-base origin/main HEAD) && ';
|
|
|
|
function literalCallerDiffAllowed(text: string): boolean {
|
|
const recomputedBase = text.startsWith(callerDiffPreface);
|
|
const diff = literalCallerDiff.exec(recomputedBase ? text.slice(callerDiffPreface.length) : text);
|
|
return !!diff && !(diff[1] && diff[3]) && recomputedBase === (diff[2] === '"$DIFF_BASE"');
|
|
}
|
|
|
|
function literalCallerListingAllowed(text: string): boolean {
|
|
if (!literalCallerListing.test(text)) return false;
|
|
const operands = text.match(new RegExp(literalCallerArgument, 'g'))!.slice(1);
|
|
if (operands[0]?.match(/^-[la]+$/)) operands.shift();
|
|
return operands[0] === '--' || operands.every(operand => !operand.replace(/^['"]/, '').startsWith('-'));
|
|
}
|
|
|
|
export interface CallerDeadlineContext {
|
|
runtime: string;
|
|
fixtureRoot: string;
|
|
}
|
|
|
|
function callerEvidenceContext(context?: CallerDeadlineContext): QaEvidenceContext | undefined {
|
|
return context ? { cwd: context.fixtureRoot, reportRoot: path.join(context.fixtureRoot, 'reports'), executable: path.join(context.runtime, 'bin/gstack-qa-evidence') } : undefined;
|
|
}
|
|
|
|
function callerEvidenceCommand(command: string, context?: CallerDeadlineContext) {
|
|
const parsed = qaEvidenceCommand(command, callerEvidenceContext(context));
|
|
if (!context || !parsed) return;
|
|
try {
|
|
if (!path.isAbsolute(context.fixtureRoot) || path.resolve(context.fixtureRoot) !== context.fixtureRoot
|
|
|| context.runtime !== path.join(path.dirname(context.fixtureRoot), 'host/runtime')
|
|
|| fs.realpathSync(context.runtime) !== context.runtime
|
|
|| fs.realpathSync(path.join(context.runtime, 'bin/gstack-qa-evidence')) !== path.join(fs.realpathSync(QA_CALLER_ROOT), 'bin/gstack-qa-evidence')
|
|
|| !fs.lstatSync(ownedPath(context.fixtureRoot, 'reports')).isDirectory()) return;
|
|
if (parsed.action !== 'capture') return parsed;
|
|
if (!literalCallerProbe.test(parsed.nativeCommand!)) return;
|
|
if (parsed.deadline === path.join(context.fixtureRoot, 'reports/deadline.json') || parsed.timeoutMs === 10000) return parsed;
|
|
} catch {}
|
|
}
|
|
|
|
function callerDeadlineCommand(command: string, context?: CallerDeadlineContext) {
|
|
if (!context) return;
|
|
const capture = callerEvidenceCommand(command, context);
|
|
if (capture?.action === 'capture' && capture.deadline) return { action: 'run', stateFile: capture.deadline };
|
|
const match = literalDeadlineCommand.exec(command.trim());
|
|
if (!match) return;
|
|
const literal = (value: string) => /^['"]/.test(value) ? value.slice(1, -1) : value;
|
|
const helper = path.join(context.runtime, 'bin/gstack-qa-deadline');
|
|
const stateFile = path.join(context.fixtureRoot, 'reports/deadline.json');
|
|
if (literal(match[1]) !== helper || literal(match[3]) !== stateFile) return;
|
|
try {
|
|
if (!path.isAbsolute(context.runtime) || path.resolve(context.runtime) !== context.runtime
|
|
|| !path.isAbsolute(context.fixtureRoot) || path.resolve(context.fixtureRoot) !== context.fixtureRoot
|
|
|| context.runtime !== path.join(path.dirname(context.fixtureRoot), 'host/runtime')
|
|
|| fs.realpathSync(context.runtime) !== context.runtime
|
|
|| ownedPath(context.fixtureRoot, 'reports/deadline.json') !== stateFile
|
|
|| !fs.lstatSync(path.dirname(stateFile)).isDirectory()
|
|
|| !fs.statSync(helper).isFile()
|
|
|| fs.realpathSync(helper) !== path.join(fs.realpathSync(QA_CALLER_ROOT), 'bin/gstack-qa-deadline')) return;
|
|
} catch { return; }
|
|
const action = match[2];
|
|
const rest = match[4];
|
|
if (action === 'status' && rest === undefined) return { action, stateFile };
|
|
if (action === 'run' && rest?.startsWith('-- ') && literalCallerProbe.test(rest.slice(3))) {
|
|
return { action, stateFile };
|
|
}
|
|
if (action !== 'start' || rest === undefined) return;
|
|
const start = new RegExp(`^(0|[1-9]\\d*)(\\.\\d{1,3})?(?: (${literalCallerArgument}))?$`).exec(rest);
|
|
if (!start) return;
|
|
const seconds = Number(start[1] + (start[2] ?? ''));
|
|
if (!(seconds > 0 && seconds <= 300)) return;
|
|
const earlier = start[3] === undefined ? undefined : literal(start[3]);
|
|
if (earlier !== undefined) {
|
|
const time = Date.parse(earlier);
|
|
if (!Number.isFinite(time) || ![new Date(time).toISOString(), new Date(time).toISOString().replace('.000Z', 'Z')].includes(earlier)) return;
|
|
}
|
|
return { action, stateFile };
|
|
}
|
|
|
|
export function qaCallerCommandAllowed(command: string, workflowCommands: string[] = [], deadline?: CallerDeadlineContext): boolean {
|
|
const text = command.trim();
|
|
if (callerEvidenceCommand(text, deadline)) return true;
|
|
if (callerDeadlineCommand(text, deadline)) return true;
|
|
if (/\bgstack-qa-(?:deadline|evidence)\b/.test(text) && !literalCallerCLI.test(text)
|
|
&& !literalCallerDiffAllowed(text) && !literalCallerListingAllowed(text)) return false;
|
|
if (workflowCommands.includes(text)) return true;
|
|
if (literalCallerCLI.test(text)) return true;
|
|
if (literalCallerDiffAllowed(text)) return true;
|
|
if (literalCallerListingAllowed(text)) return true;
|
|
let normalizedRecord = text;
|
|
const substitutedFields: string[] = [];
|
|
for (const [field, command] of [['timestamp', 'date -u +%Y-%m-%dT%H:%M:%SZ'], ['commit', 'git rev-parse --short HEAD']]) {
|
|
const fragment = `"${field}":"'"$(${command})"'"`;
|
|
if (!normalizedRecord.includes(fragment)) continue;
|
|
if (normalizedRecord.split(fragment).length !== 2) return false;
|
|
normalizedRecord = normalizedRecord.replace(fragment, `"${field}":"native-${field}"`);
|
|
substitutedFields.push(field);
|
|
}
|
|
const reviewRecord = normalizedRecord.match(/^\/?[\w./-]+\/bin\/gstack-review-log '([^'\r\n]+)'(?: --finish ([a-zA-Z0-9._:-]+))?$/);
|
|
if (reviewRecord) {
|
|
try {
|
|
const record = JSON.parse(reviewRecord[1]);
|
|
if (substitutedFields.some(field => record[field] !== `native-${field}`)) return false;
|
|
if (record?.skill === 'review') return !!reviewRecord[2];
|
|
return record?.skill === 'adversarial-review'
|
|
&& (!!reviewRecord[2] || record.completed === false && record.converged === false);
|
|
} catch { return false; }
|
|
}
|
|
if (/[\n\r;&|<>`$\\(){}]/.test(text)) return false;
|
|
return /^(?:pwd|ls(?: -la)?|bun --version|date -u \+%Y-%m-%dT%H:%M:%SZ|bun (?:run test|test(?: cli\.test\.ts)?))$/.test(text)
|
|
|| /^git (?:status --(?:short|porcelain)|branch --show-current|rev-parse (?:--short )?HEAD|merge-base origin\/main HEAD|ls-files(?: --others --exclude-standard)?)$/.test(text)
|
|
|| /^\/?[\w./-]+\/bin\/gstack-review-log --start (?:review|adversarial-review)$/.test(text)
|
|
|| /^\/?[\w./-]+\/bin\/gstack-(?:review-read|specialist-stats)$/.test(text);
|
|
}
|
|
|
|
export function validateCallerEvidence(input: {
|
|
caller: QaCaller;
|
|
result: Pick<SkillTestResult, 'transcript' | 'exitReason'>;
|
|
probes: CallerProbe[];
|
|
receipt: CallerReceipt;
|
|
currentSnapshot: string;
|
|
requiredCharters: string[];
|
|
mutations: string[];
|
|
observerComplete: boolean;
|
|
workflowCommands?: string[];
|
|
fixtureRoot?: string;
|
|
runtime?: string;
|
|
requireGuardedSmoke?: boolean;
|
|
requireCapturedEvidence?: boolean;
|
|
reportRoot: string;
|
|
checkpointFiles: Record<string, string>;
|
|
reportMarkdown: string;
|
|
}): string[] {
|
|
const errors: string[] = [];
|
|
const deadline = input.fixtureRoot && input.runtime ? { fixtureRoot: input.fixtureRoot, runtime: input.runtime } : undefined;
|
|
const producerContext = callerEvidenceContext(deadline);
|
|
if (deadline && input.reportRoot !== path.join(deadline.fixtureRoot, 'reports')) errors.push('deadline report root differs from the owned caller report root');
|
|
if (input.result.exitReason !== 'success') errors.push(`session did not complete: ${input.result.exitReason}`);
|
|
if (!input.observerComplete) errors.push('observer incomplete');
|
|
if (input.mutations.length) errors.push(...input.mutations.map(file => `unauthorized mutation: ${file}`));
|
|
let tools: CallerTool[];
|
|
try { tools = callerTools(input.result.transcript); } catch (error) {
|
|
return [...errors, (error as Error).message];
|
|
}
|
|
const reads = tools.filter(tool => tool.name === 'Read' && !tool.failed && tool.output.trim());
|
|
const captureOf = (tool: CallerTool) => qaNativeCapture({ ...tool, start: tool.index, end: tool.resultIndex }, producerContext);
|
|
const readOf = (suffix: string) => reads.find(tool => String(tool.input.file_path ?? '').endsWith(suffix));
|
|
for (const resource of ['/qa/sections/exploratory.md', '/qa/sections/system-functional.md']) {
|
|
if (!readOf(resource)) errors.push(`missing executed resource read: ${resource}`);
|
|
}
|
|
const parentRead = readOf(`/caller-${input.caller}.md`);
|
|
if (!parentRead) errors.push('missing parent entrypoint read');
|
|
if (input.caller === 'ship' && !readOf('/ship/sections/review-army.md')) errors.push('missing ship Step 9 read');
|
|
for (const tool of tools) {
|
|
const file = String(tool.input.file_path ?? '');
|
|
const guarded = tool.name === 'Bash' ? callerDeadlineCommand(String(tool.input.command), deadline) : undefined;
|
|
if (input.fixtureRoot && ['Write', 'Edit', 'MultiEdit'].includes(tool.name)) {
|
|
try {
|
|
const relative = path.relative(input.fixtureRoot, ownedPath(input.fixtureRoot, file));
|
|
if (!/^(?:reports|\.qa-state)\//.test(relative)) errors.push('write outside the declared report/fixture interface');
|
|
if (relative === 'reports/deadline.json' || /^reports\/\.qa-deadline-/.test(relative)) errors.push('actor attempted to replace reserved deadline state');
|
|
if (input.requireCapturedEvidence && /^reports\/exploration-\d{3}\.json$/.test(relative)) errors.push('actor transcribed or overwrote helper-owned checkpoint');
|
|
} catch { errors.push('write outside the declared report/fixture interface'); }
|
|
}
|
|
if (tool.name === 'Bash' && !qaCallerCommandAllowed(String(tool.input.command), input.workflowCommands, deadline)) {
|
|
errors.push('command outside declared caller observation interface');
|
|
}
|
|
if (tool.name === 'Read' && /\/(?:browse|devex-review)\/SKILL\.md$|\/qa\/sections\/(?:browser-[^/]+|qa-patterns)\.md$/.test(file)) {
|
|
errors.push(`unexpected browser/DX load: ${file}`);
|
|
}
|
|
if (tool.name === 'Skill' && /^(?:gstack-)?(?:qa|qa-only|review|ship)$/.test(String(tool.input.skill))) {
|
|
errors.push(`recursive full skill: ${tool.input.skill}`);
|
|
}
|
|
if (tool.name === 'Read' && /\/(?:qa|qa-only|review|ship)\/SKILL\.md$/.test(file)) {
|
|
errors.push(`recursive full skill read: ${file}`);
|
|
}
|
|
if (tool.name === 'Bash' && guarded?.action !== 'run' && !literalCallerCLI.test(String(tool.input.command).trim()) && !literalCallerDiffAllowed(String(tool.input.command).trim()) && !literalCallerListingAllowed(String(tool.input.command).trim()) && /\bgit\s+(?:(?:-C|-c)\s+\S+\s+)*(?:add|commit|push|stash|reset|checkout|restore|merge|rebase|cherry-pick)(?=\s|[;&|<>]|$)|\bgh\s+pr\s+(?:create|merge)\b/.test(String(tool.input.command))) {
|
|
errors.push('unauthorized git/publication action');
|
|
}
|
|
if (tool.parent && ['Write', 'Edit', 'MultiEdit'].includes(tool.name) && !/\/(?:reports|evidence|state)\//.test(file)) {
|
|
errors.push('discovery child attempted a product/test edit');
|
|
}
|
|
}
|
|
const seen = new Set<string>();
|
|
let guardedDiagnostics = 0;
|
|
const checkpointProbes: Array<{ command: string; observed: CallerProbe; index: number }> = [];
|
|
for (const probe of input.probes) {
|
|
const key = JSON.stringify([probe.charter, probe.input, probe.snapshot]);
|
|
if (seen.has(key) && probe.status === 'pass') errors.push(`duplicate unchanged passing probe: ${probe.id}`);
|
|
seen.add(key);
|
|
const tool = tools.find(tool => tool.name === 'Bash'
|
|
&& (/^(?:\s*cd\s+[^\n&;]+\s*&&)?\s*(?:bun|[\w./-]+\/bun)\s+(?:run\s+)?(?:scripts\/probe\.ts|'scripts\/probe\.ts'|"scripts\/probe\.ts")(?:\s|$)/.test(String(tool.input.command))
|
|
|| callerDeadlineCommand(String(tool.input.command), deadline)?.action === 'run'
|
|
|| callerEvidenceCommand(String(tool.input.command), deadline)?.action === 'capture')
|
|
&& (callerEvidenceCommand(String(tool.input.command), deadline)?.action === 'capture'
|
|
? isDeepStrictEqual(captureOf(tool)?.captured.observed, probe)
|
|
: tool.output.split('\n').some(line => {
|
|
try { return JSON.stringify(JSON.parse(line)) === JSON.stringify(probe); } catch { return false; }
|
|
})));
|
|
if (!tool) errors.push(`probe missing native command/result: ${probe.id}`);
|
|
else {
|
|
checkpointProbes.push({ command: String(tool.input.command), observed: probe, index: tool.index });
|
|
if (parentRead && (tool.index <= parentRead.resultIndex || tool.messageId && tool.messageId === parentRead.messageId)) errors.push(`probe preceded parent entrypoint: ${probe.id}`);
|
|
for (const id of ['exploratory', 'system-functional']) {
|
|
const resource = readOf(`/qa/sections/${id}.md`);
|
|
if (resource && (tool.index <= resource.resultIndex || tool.messageId && tool.messageId === resource.messageId)) errors.push(`probe preceded resource read: ${id}`);
|
|
}
|
|
if (input.requireGuardedSmoke) {
|
|
const command = String(tool.input.command);
|
|
const guarded = callerDeadlineCommand(command, deadline);
|
|
const captured = captureOf(tool);
|
|
const requiredPlan = probe.charter === 'plan:nine' && probe.input === '9' && input.requiredCharters.includes('plan:nine')
|
|
&& /^bun scripts\/probe\.ts (?:9|'9'|"9")$/.test(captured?.command.nativeCommand ?? command.trim());
|
|
if (guarded?.action !== 'run') {
|
|
if (!requiredPlan) errors.push(`smoke probe missing trusted deadline run: ${probe.id}`);
|
|
} else {
|
|
let valid = false;
|
|
try {
|
|
const receipts = captured?.captured.receipt.timing ?? tool.output.split('\n').filter(line => line.startsWith('QA_DEADLINE '))
|
|
.map(line => JSON.parse(line.slice('QA_DEADLINE '.length)));
|
|
const [started, finished] = receipts;
|
|
const state = readQaDeadline(guarded.stateFile);
|
|
const start = Date.parse(started.observedAt), end = Date.parse(finished.observedAt);
|
|
const limit = Date.parse(state.deadlineAt);
|
|
valid = receipts.length === 2 && state.budgetMs <= 300_000
|
|
&& Number.isFinite(start) && Number.isFinite(end)
|
|
&& new Date(start).toISOString() === started.observedAt && new Date(end).toISOString() === finished.observedAt
|
|
&& start >= Date.parse(state.startedAt) && start < limit && end >= start && end < limit
|
|
&& Number.isInteger(probe.exit) && probe.exit !== null && probe.exit >= 0 && probe.exit <= 255
|
|
&& tool.failed === (probe.exit !== 0)
|
|
&& isDeepStrictEqual(started, { guard: 'qa-deadline', event: 'started', ...state, observedAt: started.observedAt, remainingMs: limit - start, expired: false })
|
|
&& isDeepStrictEqual(finished, { guard: 'qa-deadline', event: 'finished', observedAt: finished.observedAt, deadlineAt: state.deadlineAt, timedOut: false, exitCode: probe.exit });
|
|
} catch {}
|
|
if (valid) guardedDiagnostics++;
|
|
else errors.push(`smoke probe missing consistent deadline receipts: ${probe.id}`);
|
|
}
|
|
}
|
|
}
|
|
if (probe.status === 'pass' && probe.exit === null) errors.push(`unfinished probe reported pass: ${probe.id}`);
|
|
}
|
|
if (input.requireGuardedSmoke && !guardedDiagnostics) errors.push('no authenticated guarded diagnostic executed');
|
|
checkpointProbes.sort((a, b) => a.index - b.index);
|
|
if (input.requireCapturedEvidence && checkpointProbes.some(probe => qaEvidenceCommand(probe.command, producerContext)?.action !== 'capture')) errors.push('diagnostic probe bypassed production capture');
|
|
const expiredTargets = tools.flatMap(tool => {
|
|
if (tool.name !== 'Bash' || !tool.failed) return [];
|
|
const command = String(tool.input.command);
|
|
const guarded = callerDeadlineCommand(command, deadline);
|
|
if (guarded?.action !== 'run') return [];
|
|
try {
|
|
const completion = qaProducerReceipt({ ...tool, start: tool.index, end: tool.resultIndex }, producerContext, 'incomplete');
|
|
let diagnostics: any[];
|
|
if (completion?.command.action === 'capture' && producerContext) {
|
|
const capture = readQaCaptureRecord(input.reportRoot, completion.command.id!, completion.receipt.sha256);
|
|
if (capture.receipt.status !== 'incomplete' || capture.receipt.exitCode !== 124 || completion.receipt.exitCode !== 124
|
|
|| capture.receipt.signal !== null || capture.receipt.cwd !== producerContext.cwd || capture.receipt.deadline !== guarded.stateFile
|
|
|| !isDeepStrictEqual(capture.receipt.argv, completion.command.argv) || capture.stdout.length || capture.stderr.length) return [];
|
|
diagnostics = capture.receipt.timing;
|
|
} else diagnostics = tool.output.split('\n').filter(line => line.startsWith('QA_DEADLINE ')).map(line => JSON.parse(line.slice('QA_DEADLINE '.length)));
|
|
if (diagnostics.length !== 1) return [];
|
|
const receipt = diagnostics[0];
|
|
const state = readQaDeadline(guarded.stateFile);
|
|
const observed = Date.parse(receipt.observedAt);
|
|
if (state.budgetMs > 300_000 || !Number.isFinite(observed) || new Date(observed).toISOString() !== receipt.observedAt
|
|
|| observed < Date.parse(state.startedAt) || observed < Date.parse(state.deadlineAt)
|
|
|| !isDeepStrictEqual(receipt, { guard: 'qa-deadline', event: 'expired', ...state, observedAt: receipt.observedAt, remainingMs: 0, expired: true })
|
|
|| tool.output.split('\n').some(line => { try { JSON.parse(line); return true; } catch { return false; } })) return [];
|
|
return [{ command, output: tool.output }];
|
|
} catch { return []; }
|
|
});
|
|
errors.push(...validateQACheckpoints({
|
|
transcript: input.result.transcript, reportRoot: input.reportRoot,
|
|
producer: producerContext,
|
|
probes: checkpointProbes, requiredProbes: checkpointProbes.slice(1),
|
|
additionalTargets: expiredTargets,
|
|
files: input.checkpointFiles, reportMarkdown: input.reportMarkdown,
|
|
}));
|
|
const selected = input.receipt.probes.map(id => input.probes.find(probe => probe.id === id));
|
|
if (selected.some(probe => !probe)) errors.push('receipt references an unobserved probe');
|
|
const currentProbes = selected.filter((probe): probe is CallerProbe => !!probe && probe.snapshot === input.currentSnapshot);
|
|
const boundary = currentProbes.find(probe => probe.charter === 'plan:nine' && probe.input === '9'
|
|
&& probe.exit === 0 && probe.stdout === '18\n' && probe.stderr === '' && probe.status === 'pass');
|
|
const happyProof = currentProbes.find(probe => probe.charter === 'happy' && probe.status === 'pass') ?? boundary;
|
|
for (const charter of input.requiredCharters) {
|
|
const current = currentProbes.filter(probe => probe.charter === charter
|
|
|| charter === 'happy' && probe === boundary
|
|
|| charter === 'adverse' && probe === boundary && (!input.requiredCharters.includes('happy') || probe.id !== happyProof?.id));
|
|
if (!current.length && !input.receipt.remaining.length) errors.push(`missing current charter: ${charter}`);
|
|
if (input.receipt.status === 'pass' && !current.some(probe => probe.status === 'pass')) {
|
|
errors.push(`false green for charter: ${charter}`);
|
|
}
|
|
}
|
|
const handoffs = reads.filter(tool => String(tool.input.file_path ?? '').endsWith('/HANDOFF.md'));
|
|
for (const tool of tools) {
|
|
if (tool.name !== 'Bash' || !/gstack-review-log['"]?\s/.test(String(tool.input.command))
|
|
|| !/"completed"\s*:\s*true/.test(String(tool.input.command))) continue;
|
|
if (handoffs.some(read => (read.resultIndex >= tool.index || tool.messageId && tool.messageId === read.messageId)
|
|
&& !handoffs.some(prior => prior.input.file_path === read.input.file_path && prior.parent === read.parent && prior.sessionId === read.sessionId
|
|
&& (read.output !== UNCHANGED_READ && prior.output === read.output
|
|
|| read.handoffContent !== undefined && prior.handoffContent === read.handoffContent)
|
|
&& prior.resultIndex < tool.index && (!tool.messageId || tool.messageId !== prior.messageId)))) {
|
|
errors.push('review completion preceded handoff freshness decision');
|
|
}
|
|
}
|
|
if (input.receipt.status === 'pass' && (input.receipt.remaining.length || selected.some(probe => probe?.status !== 'pass'))) {
|
|
errors.push('blocked, failing or incomplete coverage reported green');
|
|
}
|
|
return errors;
|
|
}
|
|
|
|
export function callerSnapshot(files: Record<string, string>): string {
|
|
return createHash('sha256').update(JSON.stringify(Object.entries(files).sort(([a], [b]) => a.localeCompare(b)))).digest('hex');
|
|
}
|
|
|
|
export const QA_CALLER_CASES = [
|
|
'review-exploratory-small-cli',
|
|
'ship-exploratory-small-cli',
|
|
'ship-exploratory-unavailable',
|
|
'ship-exploratory-plan-checks',
|
|
'ship-exploratory-late-input',
|
|
] as const;
|
|
|
|
export type QaCallerCase = typeof QA_CALLER_CASES[number];
|
|
export const QA_CALLER_TEST_MS = CAPTURE_MS + SESSION_DRAIN_GRACE_MS + 10_000;
|
|
const productFiles = ['scale.ts', 'cli.ts', 'cli.test.ts', 'README.md', 'package.json'];
|
|
|
|
export interface QaCallerFixture {
|
|
root: string;
|
|
cwd: string;
|
|
state: string;
|
|
runtime: string;
|
|
config: string;
|
|
instructions: string;
|
|
caller: QaCaller;
|
|
caseId: QaCallerCase;
|
|
reviewStart?: string;
|
|
gitEnvironment: Record<'GIT_OBJECT_DIRECTORY' | 'GIT_ALTERNATE_OBJECT_DIRECTORIES', string>;
|
|
journal: string;
|
|
mutationEvents: string[];
|
|
observerErrors: string[];
|
|
observation?: QAWriteObservation;
|
|
workflowCommands: string[];
|
|
lateApplied: boolean;
|
|
snapshot(): string;
|
|
probes(): CallerProbe[];
|
|
observe(): Promise<void>;
|
|
close(): Promise<void>;
|
|
}
|
|
|
|
export function createQaCallerFixture(caseId: QaCallerCase, options: { instructions?: string; installRuntime?: boolean } = {}): QaCallerFixture {
|
|
const caller: QaCaller = caseId.startsWith('review-') ? 'review' : 'ship';
|
|
const instructions = options.instructions ?? qaCallerInstructions(caller);
|
|
const root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), 'qc-'));
|
|
const cwd = path.join(root, 'product');
|
|
const state = path.join(root, 'state');
|
|
for (const dir of [cwd, state, path.join(cwd, 'scripts'), path.join(cwd, 'reports'), path.join(cwd, '.qa-state')]) fs.mkdirSync(dir);
|
|
const runtimeParent = path.join(root, 'host');
|
|
fs.mkdirSync(runtimeParent);
|
|
const config = options.installRuntime === false ? '' : refreshHermeticSkillRuntime(QA_CALLER_ROOT, runtimeParent);
|
|
const runtime = path.join(runtimeParent, 'runtime');
|
|
const journal = path.join(root, 'probes.jsonl');
|
|
const fixtureInput = path.join(state, 'fixture.json');
|
|
fs.writeFileSync(journal, '', { mode: 0o600 });
|
|
seedHermeticGstackHome(state);
|
|
fs.appendFileSync(path.join(state, 'config.yaml'), 'cross_project_learnings: false\n');
|
|
fs.writeFileSync(fixtureInput, '{"locale":"C"}\n');
|
|
const base = 'export function scale(value: string) {\n const n = Number(value);\n if (!/^\\d+$/.test(value) || n > 9) throw new Error("integer required: 0..9");\n return 2 * n;\n}\n';
|
|
fs.writeFileSync(path.join(cwd, 'scale.ts'), base);
|
|
fs.writeFileSync(path.join(cwd, 'cli.ts'), 'import { scale } from "./scale";\ntry { console.log(scale(process.argv[2] ?? "")); } catch (error) { console.error((error as Error).message); process.exit(2); }\n');
|
|
if (caseId === 'ship-exploratory-unavailable') {
|
|
fs.writeFileSync(path.join(cwd, 'cli.ts'), 'import "./vendor/native-engine.ts";\n' + fs.readFileSync(path.join(cwd, 'cli.ts'), 'utf8'));
|
|
}
|
|
fs.writeFileSync(path.join(cwd, 'cli.test.ts'), 'import { test, expect } from "bun:test";\nimport { scale } from "./scale";\ntest("double a positive number", () => expect(scale("3")).toBe(6));\n');
|
|
fs.writeFileSync(path.join(cwd, 'package.json'), JSON.stringify({ name: 'caller-scale', private: true, scripts: { test: 'bun test cli.test.ts' } }, null, 2) + '\n');
|
|
fs.writeFileSync(path.join(cwd, 'README.md'), `# Scale CLI\n\nThe CLI doubles an integer from 0 through 9 inclusive. Success exits 0 with the doubled integer and a newline on stdout, and empty stderr. Other inputs exit 2, with empty stdout and exactly \"integer required: 0..9\\n\" on stderr. It has no persistent business state or network effects.\n\nRun \`bun cli.ts <integer>\`. The repository's diagnostic client \`bun scripts/probe.ts <integer>\` launches that same CLI, records its separate streams and exit status, and prints a JSON receipt. The client preserves the command's exit status. Every diagnostic receipt field is synthetic, nonsecret evidence; snapshot identifies the owned source and fixture inputs. Its journal is owned temporary evidence outside the source tree. It is safe to use with arbitrary local input. Native tests: \`bun run test\`. No browser or service is used.\n`);
|
|
if (caseId === 'ship-exploratory-unavailable') {
|
|
fs.appendFileSync(path.join(cwd, 'README.md'), '\nThis target additionally requires the generated native-engine module at vendor/native-engine.ts. The module has not been built in this environment; no dependency generation, installation or substitute engine is authorized. Pure scale unit tests do not exercise that CLI prerequisite.\n');
|
|
}
|
|
const probeSource = `import * as fs from 'node:fs';
|
|
import { createHash, randomUUID } from 'node:crypto';
|
|
import { spawnSync } from 'node:child_process';
|
|
const files = ${JSON.stringify(productFiles)};
|
|
const snapshot = createHash('sha256').update(JSON.stringify([...files.map(file => [file, fs.readFileSync(file, 'utf8')]), ['fixture.json', fs.readFileSync(${JSON.stringify(fixtureInput)}, 'utf8')]].sort(([a], [b]) => a.localeCompare(b)))).digest('hex');
|
|
const input = process.argv[2] ?? '';
|
|
const n = Number(input), valid = /^\\d+$/.test(input) && n <= 9;
|
|
const fixture = JSON.parse(fs.readFileSync(${JSON.stringify(fixtureInput)}, 'utf8'));
|
|
const result = spawnSync(${JSON.stringify(process.execPath)}, ['cli.ts', input], { cwd: process.cwd(), env: { ...process.env, LC_ALL: fixture.locale }, encoding: 'utf8', timeout: 2000 });
|
|
const stdout = result.stdout ?? '', stderr = result.stderr ?? result.error?.message ?? '';
|
|
const exit = result.status ?? null;
|
|
const expected = valid ? { exit: 0, stdout: String(n * 2) + '\\n', stderr: '' } : { exit: 2, stdout: '', stderr: 'integer required: 0..9\\n' };
|
|
const missingPrerequisite = ${caseId === 'ship-exploratory-unavailable'} && !fs.existsSync('vendor/native-engine.ts');
|
|
const status = result.error || missingPrerequisite && exit !== 0 ? 'blocked' : exit === expected.exit && stdout === expected.stdout && stderr === expected.stderr ? 'pass' : 'fail';
|
|
const charter = input === '9' ? 'plan:nine' : valid && n > 0 ? 'happy' : 'adverse';
|
|
const probe = { id: 'probe-' + randomUUID(), charter, input, snapshot, status, stdout, stderr, exit };
|
|
fs.appendFileSync(${JSON.stringify(journal)}, JSON.stringify(probe) + '\\n');
|
|
console.log(JSON.stringify(probe));
|
|
process.exit(exit ?? 127);
|
|
`;
|
|
fs.writeFileSync(path.join(cwd, 'scripts/probe.ts'), probeSource);
|
|
fs.writeFileSync(path.join(cwd, 'HANDOFF.md'), 'No concurrent input update.\n');
|
|
if (caseId === 'ship-exploratory-plan-checks') {
|
|
fs.writeFileSync(path.join(cwd, 'PLAN.md'), '# Scale change\n\n## Verification\n\nThe upper boundary is a required release check: run `bun scripts/probe.ts 9`; require exit 0, stdout `18\\n`, and empty stderr. Ordinary positive input is not a substitute for this check.\n');
|
|
}
|
|
const run = (...args: string[]) => {
|
|
const result = spawnSync('git', args, { cwd, encoding: 'utf8', timeout: 5000 });
|
|
if (result.status !== 0) throw new Error(`Caller fixture git ${args[0]} failed: ${result.stderr || result.error?.message}`);
|
|
return result.stdout.trim();
|
|
};
|
|
try {
|
|
run('init', '-b', 'main');
|
|
run('config', 'user.name', 'QA Caller Fixture');
|
|
run('config', 'user.email', 'qa-caller-fixture@gstack.test');
|
|
run('config', 'commit.gpgsign', 'false');
|
|
run('config', 'gc.auto', '0');
|
|
run('config', 'maintenance.auto', 'false');
|
|
fs.writeFileSync(path.join(cwd, '.gitignore'), 'reports/\n.gstack/\n.qa-state/\ncaller-*.md\n');
|
|
run('add', '.');
|
|
run('commit', '-m', 'Seed caller fixture');
|
|
run('update-ref', 'refs/remotes/origin/main', 'HEAD');
|
|
run('checkout', '-b', 'caller-change');
|
|
fs.writeFileSync(path.join(cwd, 'scale.ts'), caseId === 'review-exploratory-small-cli'
|
|
? base.replace('if (!/^', 'if (!n || !/^')
|
|
: base.replace('return 2 * n;', 'return n + n;'));
|
|
const numstat = run('diff', '--numstat', 'origin/main').split('\n').filter(Boolean);
|
|
if (numstat.reduce((sum, line) => sum + line.split('\t').slice(0, 2).reduce((n, value) => n + Number(value), 0), 0) >= 50) {
|
|
throw new Error('Caller fixture no longer exercises the small-diff bypass');
|
|
}
|
|
if (fs.realpathSync(cwd).startsWith(fs.realpathSync(QA_CALLER_ROOT) + path.sep)) throw new Error('Product fixture is inside source checkout');
|
|
if (run('rev-parse', '--show-toplevel') !== fs.realpathSync(cwd)) throw new Error('Wrong fixture repository root');
|
|
const instructionFile = path.join(cwd, `caller-${caller}.md`);
|
|
fs.writeFileSync(instructionFile, instructions.replace(/~\/\.claude\/skills\/gstack\b/g, runtime));
|
|
const gitEnvironment = {
|
|
GIT_OBJECT_DIRECTORY: path.join(state, 'git-objects'),
|
|
GIT_ALTERNATE_OBJECT_DIRECTORIES: '',
|
|
};
|
|
fs.cpSync(fs.realpathSync(path.join(cwd, '.git/objects')), gitEnvironment.GIT_OBJECT_DIRECTORY, { recursive: true });
|
|
let reviewStart: string | undefined;
|
|
if (caller === 'review') {
|
|
const start = spawnSync('bash', [path.join(QA_CALLER_ROOT, 'bin/gstack-review-log'), '--start', 'review'], {
|
|
cwd, env: { ...process.env, ...gitEnvironment, GSTACK_HOME: state, GSTACK_STATE_ROOT: state },
|
|
encoding: 'utf8', timeout: 5000,
|
|
});
|
|
if (start.status !== 0 || !/^[0-9a-f-]{36}$/.test(start.stdout.trim())) {
|
|
throw new Error(`Caller review start failed: ${start.stderr || start.error?.message}`);
|
|
}
|
|
reviewStart = start.stdout.trim();
|
|
}
|
|
const mutationEvents: string[] = [], observerErrors: string[] = [];
|
|
let observer: Awaited<ReturnType<typeof observeQAWrites>> | undefined;
|
|
const sources = [instructions, ...[
|
|
`${caller}/sections/review-army.md`, 'ship/sections/plan-completion.md',
|
|
'qa/sections/scope.md', 'qa/sections/exploratory.md', 'qa/sections/system-functional.md',
|
|
...(caller === 'review' ? ['review/sections/adversarial.md'] : []),
|
|
].map(file => fs.readFileSync(path.join(QA_CALLER_ROOT, file), 'utf8'))];
|
|
const workflowCommands = sources.flatMap(source => [...source.matchAll(/```bash\n([\s\S]*?)\n```/g)]
|
|
.map(match => match[1].replaceAll('~/.claude/skills/gstack', runtime).replaceAll('<base>', 'main').trim()));
|
|
if (caller === 'review') {
|
|
const native = fs.readFileSync(path.join(QA_CALLER_ROOT, 'review/sections/adversarial.md'), 'utf8');
|
|
for (const match of native.matchAll(/`([^`\n]+)`/g)) {
|
|
const command = match[1].replaceAll('<base>', 'main');
|
|
if (command.startsWith('DIFF_BASE=$(git merge-base origin/main HEAD) && git diff')) workflowCommands.push(command);
|
|
if (command.startsWith('git diff "$DIFF_BASE"')) {
|
|
workflowCommands.push(command, `DIFF_BASE=$(git merge-base origin/main HEAD) && ${command}`);
|
|
}
|
|
}
|
|
}
|
|
const snapshot = () => callerSnapshot({ ...Object.fromEntries(productFiles.map(file => [file, fs.readFileSync(path.join(cwd, file), 'utf8')])), 'fixture.json': fs.readFileSync(fixtureInput, 'utf8') });
|
|
const probes = () => fs.readFileSync(journal, 'utf8').split('\n').filter(Boolean).map(line => JSON.parse(line) as CallerProbe);
|
|
const fixture: QaCallerFixture = {
|
|
root, cwd, state, runtime, config, caller, caseId, instructions, journal, mutationEvents, observerErrors, workflowCommands, reviewStart, gitEnvironment,
|
|
lateApplied: false, snapshot, probes,
|
|
observe: async () => {
|
|
if (observer || fixture.observation) throw new Error('Caller observation cannot restart mid-capture');
|
|
observer = await observeQAWrites(cwd, { reportDirectory: 'reports', evidenceProducer: true });
|
|
},
|
|
close: async () => {
|
|
coordinator?.close();
|
|
if (!observer) return;
|
|
fixture.observation = observer.stop();
|
|
observer = undefined;
|
|
observerErrors.push(...fixture.observation.failures);
|
|
if (!fixture.observation.complete) observerErrors.push('incomplete caller write observation');
|
|
for (const file of new Set([...fixture.observation.events.map(event => event.path), ...fixture.observation.changed])) {
|
|
if (!/^(?:reports|\.qa-state)(?:\/|$)/.test(file)) mutationEvents.push(file);
|
|
}
|
|
},
|
|
};
|
|
const coordinator = caseId === 'ship-exploratory-late-input' ? fs.watch(journal, () => {
|
|
if (fixture.lateApplied || probes().length < 2) return;
|
|
fixture.lateApplied = true;
|
|
fs.writeFileSync(fixtureInput, '{"locale":"POSIX"}\n');
|
|
fs.writeFileSync(path.join(cwd, 'reports/HANDOFF.md'), `The fixture coordinator changed the selected native process locale from C to POSIX in ${fixtureInput}.\n`);
|
|
}) : undefined;
|
|
coordinator?.on('error', error => observerErrors.push(error.message));
|
|
return fixture;
|
|
} catch (error) {
|
|
fs.rmSync(root, { recursive: true, force: true });
|
|
throw error;
|
|
}
|
|
}
|
|
|
|
export function qaCallerSessionOptions(fixture: QaCallerFixture, runId: string): Parameters<typeof runSkillTest>[0] {
|
|
return {
|
|
prompt: `Load gstack's /${fixture.caller} supplied parent phase from caller-${fixture.caller}.md and resume it on the selected working-tree diff against origin/main. This excerpt comes from ${fixture.runtime}/${fixture.caller}/SKILL.md; resolve installed-relative references there, not from the excerpt file or product directory. That path identifies the asset base, not another entrypoint: do not read or invoke the full parent SKILL.md or rerun its preamble. Earlier preamble/branch/base setup is complete; use the existing local origin/main ref without fetch. Earlier-phase asset locators are ${fixture.runtime}/review/checklist.md and ${fixture.runtime}/qa/templates/functional-report-template.md. Read those files directly when referenced; recursive Glob does not follow the installed asset symlinks. Cross-project learnings are configured off in this owned fixture. ${fixture.reviewStart ? `The actual review-start helper already captured REVIEW_START=${fixture.reviewStart} for this unchanged core pass; retain that token. ` : ''} This fixture evaluates only the supplied parent phase, not later publication stages. Read README.md for the project contract and commands. Use diagnostic-client commands such as \`bun scripts/probe.ts <literal>\` for exploratory discoveries and their checkpoint evidence. A required \`bun run test\` is separate suite verification: report it as verification, never as a diagnostic observation or checkpoint anchor/target. Use the production evidence helper to publish each diagnostic checkpoint as \`reports/exploration-NNN.json\`, not inside a nested directory; do not transcribe its observed payload. ${fixture.caseId === 'ship-exploratory-plan-checks' ? 'The previously discovered plan is PLAN.md.' : 'No plan file was found.'} There is no remote service and no release publication is authorized. There is no interactive approver; do not invent answers or permission. Keep normal parent decision gates. Before every completion report or bookkeeping log, read HANDOFF.md and reports/HANDOFF.md if present for any concurrent collaborator update, await the results, and compare evidence with current inputs. A gstack-review-log completed:true record is a completion, not preliminary bookkeeping; a later handoff read cannot validate an earlier completion.\n\nDeadline bookkeeping additionally permits \`bun ${fixture.runtime}/bin/gstack-qa-deadline start ${fixture.cwd}/reports/deadline.json SECONDS [EARLIER_UTC]\`, \`bun ${fixture.runtime}/bin/gstack-qa-deadline status ${fixture.cwd}/reports/deadline.json\`, and \`bun ${fixture.runtime}/bin/gstack-qa-deadline run ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These are closed literal forms: SECONDS must be positive and at most 300, EARLIER_UTC is the optional caller absolute deadline: use the section clock's Hard deadline UTC, never its Runner entry UTC, reserve-start time or a clock-read time. The child is only the existing diagnostic client with zero or one literal argument. Resolve these exact helper and state paths; do not use variables, another helper, another state file, nested wrappers, scripts, operators or substitutions. Only this helper may create or change reports/deadline.json and its .qa-deadline- temporary files; never use Write/Edit/MultiEdit on those paths. Record the full outer run command in checkpoints and evidence; keep the unchanged child JSON as observed, separate from prefixed guard diagnostics. A completed expired guard-run is not a probe or a pass: retain its unused checkpoint, report not-run coverage and do not restart the deadline. Keep the 12-probe smoke limit. Required suites and explicit plan checks are outside the bounded smoke budget, not permission to reset it.\n\nFunctional evidence uses the same production helper and existing diagnostic client: \`bun ${fixture.runtime}/bin/gstack-qa-evidence capture ${fixture.cwd}/reports NNN --public --deadline ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These diagnostic receipts are declared public/synthetic, so --public is approved; a fresh three-digit ID is required each time. Explicit plan probes outside the smoke budget may replace --deadline with --timeout-ms 10000; this does not reset or bypass the smoke deadline. Publish causal intent with \`bun ${fixture.runtime}/bin/gstack-qa-evidence checkpoint ${fixture.cwd}/reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command'\`; quote arguments literally. For complex quoting, Write only capture, observationCommand, hypothesis and nextCommand to reports/intent.json; publish with the same helper: checkpoint REPORT_ROOT NNN intent.json. Materialize is supported when evidence.json is required. Sources stay inside reports. Decide to execute the next probe before publishing its checkpoint, then await successful publication and dispatch that exact probe. If you defer an optional idea or stop exploration, do not publish a checkpoint for it; describe it as not run in Markdown. An unused checkpoint requires an actual authenticated expired-capture result; nearing the deadline or choosing to stop is not enough. A complete capture can truthfully retain a nonzero domain exit (including expected rejection or unavailable dependency); it is not automatically a pass. Wrapper-failed, interrupted or incomplete captures are not observations or passes; an authenticated expired capture can retain an unused checkpoint as not-run evidence. No new native child command or shell authority is granted.\n\nThe supported Bash interface is one literal documented native command or one exact generated workflow shell block with main substituted for <base>; even read-only commands must not be chained except for the exact DIFF_BASE preface below. Native CLI commands accept no argument or one literal argument: an unquoted shell-safe word, single-quoted text, or double-quoted text without expansion or escapes; no multiline arguments or shell composition. Inventory commands are pwd, bun --version, and ls with an optional combined -l/-a flag, optional -- separator, and literal path operands using the same quoting rules; operands beginning with - require --. Listing never permits other options, glob expansion, substitution, redirection or composition. The supported Git forms are git status --short, git status --porcelain, git branch --show-current, git rev-parse HEAD, git rev-parse --short HEAD, git merge-base origin/main HEAD, git diff (optional origin/main base and at most one output mode: --stat, --numstat, --name-only, or --name-status, before or after the base; optional literal pathspec arguments after --), git ls-files, and git ls-files --others --exclude-standard. Diff pathspecs use the same literal argument syntax as the CLI; quote globs so Git, not the shell, interprets them. The only variable-base form is DIFF_BASE=$(git merge-base origin/main HEAD) && git diff with exactly the double-quoted "$DIFF_BASE" base, the same optional output mode and literal pathspecs. That exact preface recomputes the base in the same command; other assignments and variable expansions are forbidden; the only additional command substitutions are the two installed bookkeeping fields described below. No other Git options, configuration overrides, --output, --no-index, --ext-diff, --textconv, external helpers or mutations are authorized. Exact generated workflow blocks remain allowed. Review bookkeeping permits the installed gstack-review-log start commands and single-quoted JSON records for review/adversarial-review (the installed quote-spliced double-quoted $(date -u +%Y-%m-%dT%H:%M:%SZ) is permitted only as the complete top-level timestamp value, and $(git rev-parse --short HEAD) only as the complete top-level commit value; no other substitutions, field placements or shell evaluation), gstack-review-read, gstack-specialist-stats, plus date -u +%Y-%m-%dT%H:%M:%SZ; use an original start token for a started attempt and completed:false/converged:false without a token for an unstarted one. Bookkeeping has an owned state directory and private Git object store; the real product, index, refs and config must stay unchanged. Outside providers are unavailable in this isolated fixture; the native reviewer is still required. Other interpreters, shell composition outside those generated blocks or the exact DIFF_BASE preface, global setup and external operations are not authorized. Use Read/Glob/Grep for discovery instead of cat/head/tail; use native Write/Edit tools for any authorized file output. Pass these same command and write boundaries to any child; generated shell fragments remain verbatim apart from the declared substitutions.\n\nWrite the phase report to reports/review.md. Also save reports/receipt.json with this machine-readable result shape: {\"status\":\"pass|fail|blocked|inconclusive\",\"probes\":[\"observed diagnostic receipt ids\"],\"remaining\":[\"names of incomplete contracts\"]}. The status is the overall supplied phase gate, not whether some probes passed. The probes array contains the exact id values from the captured child JSON (probe-...), never the helper's three-digit capture IDs or QA_EVIDENCE.id. Capture IDs select stored observations for checkpoint/materialize; child diagnostic IDs identify the actual probes in this final receipt. Read the child JSON's id for each probe used as evidence; do not derive it from a filename or execution counter. remaining names unresolved required contracts or gates in this phase. Pass requires no remaining required contracts or gates. Optional unavailable providers and later stages outside this excerpt are not required remainder. This format does not establish that any work succeeded. Stop at the end of this phase, or at its first unresolved approval gate.`,
|
|
appendSystemPrompt: `Caller execution scheduling (fixture contract):
|
|
This session has at most 25 assistant turns, including required verification and final artifacts. The command boundary applies to each Bash call, not to the number of independent tool calls in an assistant turn.
|
|
After required clock and approval prerequisites settle, issue independent source Reads and read-only discovery together as separate native tool calls once their paths and inputs are known. Wait for their results before decisions that depend on them.
|
|
The completion reserve is for required verification, affected-input revalidation and artifacts, not an earlier deadline. Keep completing required work within the actual remaining deadline; reserve entry alone is not a reason to stop. Use each native probe's snapshot to distinguish current from superseded evidence before deciding which checks still need revalidation.
|
|
Never group diagnostic probes, checkpoint publication with its next probe, or any action with the clock/status/approval result it needs. Shell composition remains forbidden outside the declared forms. Preserve every required Read, probe, verification, freshness check, approval and report field; the turn limit does not authorize skipping work or reporting incomplete work as passed.`,
|
|
workingDirectory: fixture.cwd,
|
|
timeout: CAPTURE_MS,
|
|
completionReserveMs: CAPTURE_MS / 4,
|
|
maxTurns: 25,
|
|
allowedTools: ['Read', 'Grep', 'Glob', 'Bash', 'Write', 'Edit', 'Agent', 'Skill', 'AskUserQuestion'],
|
|
testName: fixture.caseId,
|
|
runId,
|
|
publicStreamDiagnostics: true,
|
|
env: { CLAUDE_CONFIG_DIR: fixture.config, GSTACK_HOME: fixture.state, GSTACK_STATE_ROOT: fixture.state, GSTACK_HEADLESS: '1', GIT_OPTIONAL_LOCKS: '0', ...fixture.gitEnvironment },
|
|
};
|
|
}
|
|
|
|
export async function runQaCaller(fixture: QaCallerFixture, runId: string, runner = runSkillTest): Promise<SkillTestResult> {
|
|
const options = qaCallerSessionOptions(fixture, runId);
|
|
if (!options.env?.CLAUDE_CONFIG_DIR) throw new Error('Live caller capture requires the hermetic skill runtime');
|
|
return runner(options);
|
|
}
|
|
|
|
export function readCallerReceipt(fixture: QaCallerFixture): CallerReceipt {
|
|
const receipt = JSON.parse(fs.readFileSync(path.join(fixture.cwd, 'reports/receipt.json'), 'utf8'));
|
|
if (!receipt || !['pass', 'fail', 'blocked', 'inconclusive'].includes(receipt.status)
|
|
|| !Array.isArray(receipt.probes) || !receipt.probes.every((id: unknown) => typeof id === 'string')
|
|
|| !Array.isArray(receipt.remaining) || !receipt.remaining.every((name: unknown) => typeof name === 'string')) {
|
|
throw new Error('Malformed caller receipt');
|
|
}
|
|
return receipt;
|
|
}
|
|
|
|
export function retainQaCallerEvidence(fixture: QaCallerFixture, dir: string, result: SkillTestResult | undefined): void {
|
|
fs.mkdirSync(dir, { recursive: true, mode: 0o700 });
|
|
if (!fs.lstatSync(dir).isDirectory() || fs.realpathSync(dir) !== path.resolve(dir)) throw new Error('Evidence directory must be a real owned directory');
|
|
fs.chmodSync(dir, 0o700);
|
|
const handoffReads = new Set<string>();
|
|
const publicEvents = result?.transcript.flatMap(event => {
|
|
if (!['assistant', 'user'].includes(event.type)) return [];
|
|
const content = (event.message?.content ?? []).filter((block: any) => ['tool_use', 'tool_result'].includes(block.type));
|
|
for (const block of content) {
|
|
if (event.type === 'assistant' && block.type === 'tool_use' && block.name === 'Read'
|
|
&& typeof block.input?.file_path === 'string' && block.input.file_path.endsWith('/HANDOFF.md')) {
|
|
handoffReads.add(JSON.stringify([event.parent_tool_use_id ?? null, block.id]));
|
|
}
|
|
}
|
|
const handoffResult = event.type === 'user' && content.length === 1 && content[0].type === 'tool_result'
|
|
&& (handoffReads.has(JSON.stringify([event.parent_tool_use_id ?? null, content[0].tool_use_id]))
|
|
|| /\/\.qa-evidence\/\d{3}\/observation\.json$/.test(event.tool_use_result?.file?.filePath ?? ''));
|
|
const metadata = handoffResult ? event.tool_use_result
|
|
: typeof event.tool_use_result?.interrupted === 'boolean' ? { interrupted: event.tool_use_result.interrupted } : undefined;
|
|
return content.length ? [{ type: event.type, parent_tool_use_id: event.parent_tool_use_id ?? null,
|
|
session_id: event.session_id, message: { id: event.message?.id, content },
|
|
...(metadata ? { tool_use_result: metadata } : {}) }] : [];
|
|
}) ?? [];
|
|
const retain = (file: string, content: string) => {
|
|
const destination = path.join(dir, file);
|
|
if (fs.existsSync(destination)) throw new Error('Evidence capture must not overwrite an earlier attempt');
|
|
fs.writeFileSync(destination, content, { mode: 0o600, flag: 'wx' });
|
|
};
|
|
let captures: unknown;
|
|
let captureFailure: unknown;
|
|
try { captures = qaCaptureArtifacts(path.join(fixture.cwd, 'reports')); }
|
|
catch (error) { captureFailure = error; captures = { error: String(error) }; }
|
|
for (const [file, content] of Object.entries({
|
|
'native-events.json': JSON.stringify(publicEvents, null, 2),
|
|
'native-probes.jsonl': fs.readFileSync(fixture.journal, 'utf8'),
|
|
'captures.json': JSON.stringify(captures, null, 2),
|
|
'observer.json': JSON.stringify({ observation: fixture.observation, events: fixture.mutationEvents, errors: fixture.observerErrors, lateApplied: fixture.lateApplied, snapshot: fixture.snapshot(), exitReason: result?.exitReason ?? 'capture failed' }, null, 2),
|
|
'consumed-parent.md': fixture.instructions,
|
|
'report.md': fs.existsSync(path.join(fixture.cwd, 'reports/review.md')) ? fs.readFileSync(path.join(fixture.cwd, 'reports/review.md'), 'utf8') : 'No report produced.\n',
|
|
'receipt.json': fs.existsSync(path.join(fixture.cwd, 'reports/receipt.json')) ? fs.readFileSync(path.join(fixture.cwd, 'reports/receipt.json'), 'utf8') : 'null\n',
|
|
})) {
|
|
retain(file, content);
|
|
}
|
|
try {
|
|
const deadline = path.join(fixture.cwd, 'reports/deadline.json');
|
|
if (fs.lstatSync(deadline, { throwIfNoEntry: false })) retain('deadline.json', JSON.stringify(readQaDeadline(deadline)) + '\n');
|
|
} catch (error) {
|
|
retain('deadline-capture-error.txt', `Deadline capture failed: ${(error as Error).message}\n`);
|
|
}
|
|
let checkpoints: Record<string, string>;
|
|
try { checkpoints = readQACheckpointFiles(path.join(fixture.cwd, 'reports')); }
|
|
catch (error) {
|
|
retain('checkpoint-capture-error.txt', `Checkpoint capture failed: ${(error as Error).message}\n`);
|
|
return;
|
|
}
|
|
for (const [file, content] of Object.entries(checkpoints)) retain(file, content);
|
|
if (captureFailure) throw captureFailure;
|
|
}
|