mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-28 23:52:28 +02:00
v1.91.2.0 fix: consolidate gstack reliability wave (#2959)
* fix(memory-ingest): --scan-secrets scans the rendered page and fails closed --scan-secrets ran gitleaks on the raw transcript .jsonl, then imported a page rendered from it. gitleaks' assignment rules don't match across a JSON-escaped quote (KEY=\"v\" on disk), so a secret the rendered page shows as KEY="v" was imported unflagged. And the gate skipped a file only on scanner "gitleaks" with findings, so a scan that errored (non-zero exit, 16MB maxBuffer overflow on a file with many findings, unparseable report) or could not run (gitleaks missing, slow-probe cooldown) imported the file unscanned. Scan the rendered page body, the exact bytes writeStaged() writes, via a new secretScanText() helper, and skip the file whenever the scan did not complete. Skipped files stay out of the state file, so the next run retries them. Reword the helper warnings and setup-gbrain/memory.md, which described the fail-open as intended. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(test): reconcile Bun failure markers and footer counts * fix(sync-gbrain): verify source-scoped reads without mutation * fix(test): recognize grounded TTHW target choices structurally * fix(aside): make the readiness probe work under zsh and report why it failed The probe built its deadline into `_T` and expanded it unquoted, so `$_T aside repl …` only worked in a shell that word-splits. zsh does not: it looked for a command literally named "gtimeout 30", the probe answered ASIDE_NOT_RUNNING with Aside installed and ready, and every browsing skill fell back to the bundled Chromium in silence. zsh is the macOS default and Aside is macOS-only, so on a stock Mac the probe could never report READY. The deadline becomes a function, `_gs_d`. It receives the command as "$@", already split, so sh, bash and zsh all behave the same, and the gtimeout → timeout → perl alarm chain is unchanged. A 4th arm runs the call unbounded when none of the three is present, which is what the empty `_T` did before. Not `eval`: it re-parses the string, so the parens and `;` of the perl arm become syntax and that arm dies in bash *and* zsh — on a stock Mac, the arm that actually runs. On failure the probe now prints the CLI's reason after ASIDE_NOT_RUNNING:, the shape gstack-render already uses: the first line that starts with a capital letter, i.e. the CLI's own sentence or Node's `Error:` line below its loader frame. "Not running" covers states with different fixes — no window open for the profile, a NODE_OPTIONS preload that kills the CLI — and a bare verdict sent all of them to "open the Aside app". The BROWSER SETUP prose quotes that reason before asking the user to open the app. The text pin asserted the broken invocation verbatim, so it now pins the function and asserts neither `$_T aside repl` nor an eval form comes back. A second test executes the rendered probe in sh, bash and zsh on each of the four deadline arms with stubbed binaries on a narrowed PATH, plus two failing CLIs: one that prints its own sentence, one that crashes like Node with the useful line below the frame. The deadline function costs zero bytes against the lines it replaces; the reason costs 53 per copy of the probe (44 where the reworded BROWSER SETUP line gives 9 back). That moves four guards by the measured amount: plan-devex-review's skeleton cap to 68,550 (measured 68,544), plan-ceo-review's skeleton cap to 80,150 (measured 80,111) and union ratio to 1.081 (measured 1.0803), and plan-eng-review's union ratio to 1.151 (measured 1.1504). Fixes #2842, #2941. * Clarify engineering review startup and decision flow * Fix Windows readiness fixture PATH and command shim * fix(test): recognize grounded TTHW target choices structurally * Clarify engineering review startup and decision flow * fix(test): restrict QA-only fixture tools to its no-Edit contract * v1.90.0.0 fix(sync-gbrain): guard readiness verdicts and refresh metadata * fix(browse): validate canonical upload targets * fix(gbrain): classify structured PGLite busy response * fix(browse): preserve native extension runtime APIs * Fix displayless browser handoff ownership * Accept unique installed autoplan methodology aliases * fix(skills): preserve positional literals during installation * fix(browse): checksum installer contents through stdin * fix(test): normalize Windows checksum fixture paths * test: emulate unavailable shasum in Windows checksum fixture * fix(investigate): preserve owned freeze lifecycle * fix(review): preserve N+1 retry and Red Team completion * fix: bound Aside readiness and preserve safe fallback * test: exercise setup and Chromium on native ARM * fix: preserve install ownership and ARM browser selection * Fix gbrain ingest scan boundaries and seed observation * Refresh managed ship hooks and supervise expanded paid census * Reject resumed gbrain pages excluded by current policy * Recover zombie agent locks safely and enable CI Python venv * Repair paid actor declarations and Aside pitch assertions * Bump consolidated wave to next free minor release * Clarify CEO review admin choices and option tradeoffs * Preserve CEO mode handoff anchors in clarified workflow * Make Windows portability fixtures use shell-native paths * Restore ARM Bun alias and clarify ship review gates * Refresh ship workflow golden snapshots * Fix Windows DX documentation controls without piped stdin * Decode Codex child pipes without Bun's encoded-stream stall * Bound DX pre-review audit before product questions * Clarify trusted review-start read in paid revalidation * Bump consolidated wave to next free minor release * Clarify CEO review admin choices and option tradeoffs * Preserve CEO mode handoff anchors in clarified workflow * Make Windows portability fixtures use shell-native paths * Restore ARM Bun alias and clarify ship review gates * Refresh ship workflow golden snapshots * Fix Windows DX documentation controls without piped stdin * Decode Codex child pipes without Bun's encoded-stream stall * Bound DX pre-review audit before product questions * Clarify trusted review-start read in paid revalidation * Reconcile new main planning flow and paid judge census * fix: reconcile rebased planning and source-bound validation * test: pin cookie workflow judge to scored Sonnet model * fix: keep terminal agent boot out of module imports * fix: preserve pending-question uncertainty in engineering review * fix: stabilize Windows reliability-wave fixtures * fix: clarify design consultation research workflow * fix: preserve independent design consultation inputs * fix: resolve design taste scope and browser research guidance * fix: make consultation opt-in preflight unambiguous * test: await native Edge owner readiness or terminal result --------- Co-authored-by: Bruce Krysiak <brucek@alum.mit.edu> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Antonio Vitalic <antoninte99@gmail.com>
This commit is contained in:
co-authored by
Bruce Krysiak
Claude Opus 5.5
Antonio Vitalic
parent
2a113ae7e6
commit
01593aa67c
@@ -181,12 +181,12 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
|
||||
// Fork port wave 2 (#703): the repo-doc-preference block in the design
|
||||
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
|
||||
maxSkeletonBytes: 80_100, // + depth-specific output and 0H/0I feasibility boundary clarity; measured 80,073.
|
||||
maxSkeletonBytes: 80_150, // + depth-specific output and 0H/0I feasibility boundary clarity + the Aside probe's failure reason; measured 80,111.
|
||||
minUnionBytes: 123_600, // token-reduction Phases 1-2 (v1.69.x branch): preamble bash -> bin/gstack-skill-start, onboarding -> gated emission; measured union 137,346
|
||||
mustContain: ['SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'HOLD SCOPE', 'SCOPE REDUCTION'],
|
||||
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
||||
// prose replacing the smaller opt-in question) lands this ~5.2% over baseline.
|
||||
maxSizeRatio: 1.08,
|
||||
maxSizeRatio: 1.081, // + the Aside probe's failure reason; measured 1.0803
|
||||
},
|
||||
'plan-eng-review': {
|
||||
skill: 'plan-eng-review',
|
||||
@@ -218,7 +218,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
|
||||
// hardening: host-anchored mode signal, precedence, passing-mention
|
||||
// guards) and the plan-mode preamble reword land the union at 1.092.
|
||||
maxSizeRatio: 1.15, // + clarity rules for saved decisions/setup gates; measured 1.146
|
||||
maxSizeRatio: 1.151, // + clarity rules for saved decisions/setup gates + the Aside probe's failure reason; measured 1.1504
|
||||
},
|
||||
'plan-design-review': {
|
||||
skill: 'plan-design-review',
|
||||
@@ -264,7 +264,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
|
||||
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
|
||||
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
|
||||
maxSkeletonBytes: 68_500, // + v2.0 {{ASIDE_RESEARCH}} (Aside first, WebSearch fallback); measured 67_129
|
||||
maxSkeletonBytes: 68_550, // + v2.0 {{ASIDE_RESEARCH}} (Aside first, WebSearch fallback) + the Aside probe's failure reason; measured 68_544
|
||||
minUnionBytes: 99_700, // token-reduction Phases 1-2 (v1.69.x branch); measured union 110,833
|
||||
mustContain: ['developer experience', 'Getting Started'],
|
||||
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
||||
|
||||
@@ -4054,7 +4054,10 @@ export async function launchClaudePty(
|
||||
let childEnv = hermeticChildEnv(opts.env);
|
||||
// The opted-in viewport emulates xterm; placeholder styles are required to
|
||||
// distinguish an empty suggestion from text the user has actually entered.
|
||||
if (opts.observeScreen) childEnv.TERM = 'xterm-256color';
|
||||
if (opts.observeScreen) {
|
||||
childEnv.TERM = 'xterm-256color';
|
||||
childEnv.FORCE_COLOR = '1';
|
||||
}
|
||||
let hermeticSkillStateRoot: string | undefined;
|
||||
if (opts.seedSkills && hermetic && !opts.env?.CLAUDE_CONFIG_DIR) {
|
||||
childEnv.CLAUDE_CONFIG_DIR = hermeticSkillsConfigDir();
|
||||
|
||||
@@ -104,6 +104,7 @@ export interface CodexEvalOptions {
|
||||
name: string;
|
||||
suite: string;
|
||||
budgetMs: number;
|
||||
drainGraceMs?: number;
|
||||
run: (signal: AbortSignal) => Promise<CodexResult>;
|
||||
validate: (result: CodexResult) => void | Promise<void>;
|
||||
record: (entry: EvalTestEntry) => void;
|
||||
@@ -114,7 +115,9 @@ export interface CodexEvalOptions {
|
||||
|
||||
export async function runRecordedCodexEval(opts: CodexEvalOptions): Promise<CodexResult> {
|
||||
const started = Date.now();
|
||||
const deadlineAt = started + opts.budgetMs + CODEX_DRAIN_GRACE_MS;
|
||||
const drainGraceMs = opts.drainGraceMs ?? CODEX_DRAIN_GRACE_MS;
|
||||
const timeoutMs = opts.budgetMs + drainGraceMs;
|
||||
const deadlineAt = started + timeoutMs;
|
||||
const controller = new AbortController();
|
||||
let result: CodexResult | undefined;
|
||||
let stage: 'runner' | 'validation' = 'runner';
|
||||
@@ -122,7 +125,7 @@ export async function runRecordedCodexEval(opts: CodexEvalOptions): Promise<Code
|
||||
let failure: unknown;
|
||||
let deadline: ReturnType<typeof setTimeout>;
|
||||
|
||||
const timeoutError = () => new CodexEvalTimeout(`Codex eval exceeded ${opts.budgetMs}ms plus ${CODEX_DRAIN_GRACE_MS}ms drain grace`);
|
||||
const timeoutError = () => new CodexEvalTimeout(`Codex eval exceeded ${opts.budgetMs}ms plus ${drainGraceMs}ms drain grace`);
|
||||
const checkDeadline = () => {
|
||||
if (!controller.signal.aborted && Date.now() >= deadlineAt) controller.abort(timeoutError());
|
||||
controller.signal.throwIfAborted();
|
||||
@@ -149,7 +152,7 @@ export async function runRecordedCodexEval(opts: CodexEvalOptions): Promise<Code
|
||||
const error = timeoutError();
|
||||
controller.abort(error);
|
||||
reject(error);
|
||||
}, opts.budgetMs + CODEX_DRAIN_GRACE_MS);
|
||||
}, timeoutMs);
|
||||
});
|
||||
return await Promise.race([work(), timedOut]);
|
||||
} catch (error) {
|
||||
|
||||
@@ -16,6 +16,7 @@ import * as fs from 'fs';
|
||||
import * as path from 'path';
|
||||
import * as os from 'os';
|
||||
import { spawn } from 'child_process';
|
||||
import { StringDecoder } from 'node:string_decoder';
|
||||
import { hermeticChildEnv } from './hermetic-env';
|
||||
import { extractSkillSections } from './skill-fixture';
|
||||
import { killProcessGroup } from '../../scripts/test-strict-output';
|
||||
@@ -286,6 +287,8 @@ export async function runCodexSkill(opts: {
|
||||
let stdoutEnded = false;
|
||||
let stderrEnded = false;
|
||||
let finalized = false;
|
||||
const stdoutDecoder = new StringDecoder('utf8');
|
||||
const stderrDecoder = new StringDecoder('utf8');
|
||||
let workTimer: ReturnType<typeof setTimeout> | undefined;
|
||||
let drainTimer: ReturnType<typeof setTimeout> | undefined;
|
||||
let finish!: () => void;
|
||||
@@ -337,14 +340,14 @@ export async function runCodexSkill(opts: {
|
||||
// drained. A destroyed pipe can emit 'close' without either EOF or error.
|
||||
const onStdoutDone = () => { if (!finalized) { stdoutDone = true; maybeFinish(); } };
|
||||
const onStderrDone = () => { if (!finalized) { stderrDone = true; maybeFinish(); } };
|
||||
const onStdoutEnd = () => { if (!finalized) { stdoutEnded = true; onStdoutDone(); } };
|
||||
const onStderrEnd = () => { if (!finalized) { stderrEnded = true; onStderrDone(); } };
|
||||
const onStdoutEnd = () => { if (!finalized) { stdoutBuffer += stdoutDecoder.end(); stdoutEnded = true; onStdoutDone(); } };
|
||||
const onStderrEnd = () => { if (!finalized) { stderr += stderrDecoder.end(); stderrEnded = true; onStderrDone(); } };
|
||||
const onStreamError = (stream: 'stdout' | 'stderr', error: Error) => {
|
||||
if (!finalized) streamError ??= { stream, error };
|
||||
};
|
||||
const onStdout = (chunk: string) => {
|
||||
const onStdout = (chunk: Buffer) => {
|
||||
if (finalized) return;
|
||||
stdoutBuffer += chunk;
|
||||
stdoutBuffer += stdoutDecoder.write(chunk);
|
||||
const lines = stdoutBuffer.split('\n');
|
||||
stdoutBuffer = lines.pop() || '';
|
||||
for (const line of lines) {
|
||||
@@ -364,12 +367,10 @@ export async function runCodexSkill(opts: {
|
||||
} catch { /* malformed JSONL is ignored by parseCodexJSONL too */ }
|
||||
}
|
||||
};
|
||||
const onStderr = (chunk: string) => { if (!finalized) stderr += chunk; };
|
||||
const onStderr = (chunk: Buffer) => { if (!finalized) stderr += stderrDecoder.write(chunk); };
|
||||
|
||||
proc.on('exit', onExit);
|
||||
proc.on('error', onSpawnError);
|
||||
proc.stdout!.setEncoding('utf8');
|
||||
proc.stderr!.setEncoding('utf8');
|
||||
proc.stdout!.on('data', onStdout);
|
||||
proc.stderr!.on('data', onStderr);
|
||||
proc.stdout!.on('end', onStdoutEnd).on('close', onStdoutDone).on('error', error => onStreamError('stdout', error));
|
||||
|
||||
@@ -4,6 +4,7 @@ import { join } from 'node:path';
|
||||
import { buildWorkflowJudgePrompt, type WorkflowJudgeFile, type WorkflowJudgeInput } from './workflow-judge-input';
|
||||
|
||||
export const COOKIE_WORKFLOW_JUDGE = {
|
||||
model: 'claude-sonnet-4-6',
|
||||
judgeContext: 'a fallback-browser cookie import workflow',
|
||||
judgeGoal: 'how to select an authorized source browser, profile, and domain without guessing an account; configure optional authentication verification before mutation; obtain explicit consent for precisely scoped storage reset; distinguish copied cookies from positive sign-in evidence; and recover within the documented platform and privacy boundaries',
|
||||
thresholds: { clarity: 4, completeness: 3, actionability: 4 },
|
||||
|
||||
@@ -105,9 +105,9 @@ export const STRICT_RETRY_CASE_BUDGETS = [...FINDING_RETRY_BUDGETS, AUQ_CONSISTE
|
||||
export const FILE_RETRY_BUDGETS = [
|
||||
...STRICT_RETRY_CASE_BUDGETS,
|
||||
...[
|
||||
// Fourteen workflow judges include their 10s recording grace; the other
|
||||
// eleven judges retain 120s. Supervise all 25 and the existing one retry.
|
||||
{ file: 'test/skill-llm-eval.test.ts', attemptMs: 15 * (JUDGE_MS + 10_000) + 11 * JUDGE_MS, retries: 1 },
|
||||
// Sixteen workflow judges include their 10s recording grace; the other
|
||||
// eleven judges retain 120s. Supervise all 27 and the existing one retry.
|
||||
{ file: 'test/skill-llm-eval.test.ts', attemptMs: 16 * (JUDGE_MS + 10_000) + 11 * JUDGE_MS, retries: 1 },
|
||||
{ file: 'test/codex-e2e-plan-format.test.ts', attemptMs: 4 * (CAPTURE_LONG_MS + 10_000), retries: 1 },
|
||||
{ file: 'test/skill-e2e-auq-matrix.test.ts', attemptMs: 6 * CAPTURE_MS, retries: 1 },
|
||||
{ file: 'test/skill-e2e-plan-format.test.ts', attemptMs: 4 * (CAPTURE_MS + 10_000), retries: 1 },
|
||||
|
||||
@@ -1,15 +1,27 @@
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { resolve } from 'node:path';
|
||||
import type { EvalTestEntry } from './eval-store';
|
||||
import { buildCookieWorkflowJudgeInput } from './cookie-workflow-judge-input';
|
||||
import { COOKIE_MANUAL_REVIEW_FILE } from './cookie-workflow-manual-review';
|
||||
|
||||
export function approvedCookieWorkflowSource(source: string): string {
|
||||
return source
|
||||
.replace('sha256sum < "$tmpfile" | awk \'{print $(1)}\'', 'sha256sum "$tmpfile" | awk \'{print $1}\'')
|
||||
.replace('shasum -a 256 < "$tmpfile" | awk \'{print $(1)}\'', 'shasum -a 256 "$tmpfile" | awk \'{print $1}\'');
|
||||
}
|
||||
|
||||
export function manualReviewFixture(root = resolve(import.meta.dir, '../..')): EvalTestEntry {
|
||||
const approval = JSON.parse(readFileSync(resolve(root, COOKIE_MANUAL_REVIEW_FILE), 'utf8'));
|
||||
const prompt = approvedCookieWorkflowSource(buildCookieWorkflowJudgeInput(root).prompt);
|
||||
if (createHash('sha256').update(prompt).digest('hex') !== approval.prompt_sha256
|
||||
|| Buffer.byteLength(prompt) !== approval.prompt_bytes) {
|
||||
throw new Error('Historical cookie approval fixture no longer reconstructs the exact approved prompt');
|
||||
}
|
||||
return {
|
||||
name: 'setup-browser-cookies/SKILL.md workflow', suite: 'Cookie setup workflow quality', tier: 'llm-judge',
|
||||
passed: false, execution: 'executed', exit_reason: 'provider_refusal', attempt: 1, duration_ms: 1, cost_usd: 0,
|
||||
model: approval.model, prompt: buildCookieWorkflowJudgeInput(root).prompt,
|
||||
model: approval.model, prompt,
|
||||
error: 'Synthetic provider refusal fixture, not live model evidence',
|
||||
manual_review: { approval, refusal: { stop_reason: 'refusal', response_id: 'msg_synthetic_fixture',
|
||||
request_id: 'req_synthetic_fixture', model: approval.model, input_tokens: 1, output_tokens: 0, text_blocks: 0 } },
|
||||
|
||||
@@ -163,13 +163,14 @@ function deterministicPlanFloorFinding(input: PlanFloorReview): PlanFloorAssessm
|
||||
const q = input.candidate.question;
|
||||
const combined = `${q.header}\n${q.question}`.replace(/\s+/g, ' ');
|
||||
const lower = combined.toLowerCase();
|
||||
const journeyContext = lower.replace(/\btime[- ]to[- ]first[- ]call\b/g, '');
|
||||
const hasTthwTargetConcept =
|
||||
/\b(?:tthw|time-to-first-call|time to first call|time-to-hello-world|time to hello world)\b/.test(lower) ||
|
||||
(/\b(?:yardstick|score against|bar i compare|target is recorded)\b/.test(lower) &&
|
||||
/\b(?:under-?10|2-5|min|minutes|clock)\b/.test(lower));
|
||||
const isDevexTthwTarget =
|
||||
hasTthwTargetConcept &&
|
||||
/\b(?:quickstart|first-call journey|sdk quickstart|onboarding flow|8-step onboarding|gap report)\b/.test(lower) &&
|
||||
/\b(?:quickstart|onboarding flow|8-step onboarding|gap report|first(?:[- ](?:sdk|api|successful))*[- ]call)\b/.test(journeyContext) &&
|
||||
/\b(?:email|key|wait|unattended)\b/.test(lower) &&
|
||||
q.options.some(o => /(?:under|<)\s*10\s*min|measured wait|competitive|champion|current trajectory|copy-pasteable first call|key turnaround/i.test(`${o.label}\n${o.description}`));
|
||||
if (!isDevexTthwTarget) return null;
|
||||
@@ -178,17 +179,13 @@ function deterministicPlanFloorFinding(input: PlanFloorReview): PlanFloorAssessm
|
||||
'Step 7: register an API key by emailing the team.',
|
||||
'No quickstart command, no hosted sandbox, no copy-pasteable curl example.',
|
||||
].find(text => input.seed.includes(text));
|
||||
const questionQuote = q.question.match(/Which (?:time-to-first-call|TTHW|Time-to-Hello-World) target should this quickstart (?:aim for|be measured against|be held to)\?/i)?.[0]
|
||||
?? q.question.match(/Which Time-to-Hello-World target fits this first-call journey\?/i)?.[0]
|
||||
?? q.question.match(/Which time-to-first-call target should this review (?:hold the plan to|aim the plan at)\?/i)?.[0]
|
||||
?? q.question.match(/Which yardstick should the gap report score against\?/i)?.[0];
|
||||
const optionIndex = q.options.findIndex(o => /<\s*10\s*min|measured wait|competitive|champion|current trajectory|copy-pasteable first call|key turnaround/i.test(`${o.label}\n${o.description}`));
|
||||
const firstLine = q.question.split(/\r?\n/)[0]!.trim();
|
||||
const brief = firstLine.replace(/^D\d+(?:\s*\(re-ask\))?\s*[—–:-]\s*/i, '');
|
||||
const currentQuestion = /^D\d+\s*\(re-ask\)/i.test(firstLine) ? brief.split(/\.\s+/).at(-1)! : brief;
|
||||
const questionQuote = currentQuestion.match(/^(?:Which|What) (?:(?:time[- ]to[- ]first[- ]call|TTHW|time[- ]to[- ]hello[- ]world) target|yardstick) (?:should|fits)\b[^?]*\?$/i)?.[0];
|
||||
const optionIndex = q.options.findIndex(o => /^(?:(?:under|<)\s*10\s*min\b|(?:champion|competitive|current trajectory)(?=$|\s*[(,]))/i.test(o.label.replace(/^[A-D][).:]\s*/, '').trim()));
|
||||
const option = optionIndex >= 0 ? q.options[optionIndex] : undefined;
|
||||
const optionQuote = option && /<\s*10\s*min/i.test(option.label) ? option.label
|
||||
: option && /competitive|champion|current trajectory/i.test(option.label) ? option.label
|
||||
: option?.description.match(/[^.]*?(?:under|<)\s*10\s*min[^.]*\./i)?.[0]
|
||||
?? option?.description.match(/[^.]*copy-pasteable first call[^.]*\./i)?.[0]
|
||||
?? option?.description.match(/[^.]*measured wait[^.]*\./i)?.[0];
|
||||
const optionQuote = option?.label;
|
||||
if (!seedQuote || !questionQuote || optionIndex < 0 || !optionQuote) return null;
|
||||
|
||||
return validatePlanFloorAssessment(input, {
|
||||
|
||||
@@ -730,7 +730,7 @@ export function reviewRevalidationPrompt(f: SharedLibsFixture, instructions: str
|
||||
|
||||
Revalidation fixture execution contract:
|
||||
- The runtime allows ${SHARED_INTERACTIVE_MAX_TURNS} assistant turns. Batch independent required source reads, Git configuration/attribute checks, and snapshot checks within each phase. Preserve every required evidence check and dependency: capture the real start token before reading the diff, and complete final evidence verification before persistence.
|
||||
- The trusted start-record location is ${startRecord}. Replace <REVIEW_START> with the token actually returned by --start; read and verify that record. Use the supplied helper interfaces; discovering helper CLI options is outside this replay.
|
||||
- The trusted start-record location is ${startRecord}. Replace <REVIEW_START> with the token actually returned by --start. Read that token's record in a separate, successful Read tool call or a single cat command before continuing. Verify its repo, branch, working tree and start time. Do not combine the record read with --start, the diff or other diagnostic commands whose failure could invalidate the read; if the read fails, retry it before proceeding. Use the supplied helper interfaces; discovering helper CLI options is outside this replay.
|
||||
- After final verification, combine successful --finish persistence and one complete, untruncated read-back through gstack-review-read in the same tool invocation. Read back only after persistence succeeds, inspect the full current record and binding, then return the final review summary in conversation.
|
||||
- Failed persistence or verification remains a failure. Late source changes still require the workflow's normal re-review; never skip checks, questions, or convergence rules to finish within the bound.`;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,200 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { createHash, randomUUID } from 'node:crypto';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import type { CanUseTool, HookCallback, SDKMessage } from '@anthropic-ai/claude-agent-sdk';
|
||||
import type { AgentSdkResult, QueryProvider } from './agent-sdk-runner';
|
||||
import type { EvalTestEntry } from './eval-store';
|
||||
import { CAPTURE_MS } from './eval-budgets';
|
||||
import { readWorkflowExcerpt } from './workflow-excerpt';
|
||||
|
||||
export type ShipHookCase = 'ship-managed-hook-refresh' | 'ship-unmanaged-hook-consent' | 'ship-local-hook-preservation';
|
||||
const ROOT = path.resolve(import.meta.dir, '../..');
|
||||
const quote = (value: string) => `'${value.replaceAll("'", "'\"'\"'")}'`;
|
||||
const read = (file: string) => fs.existsSync(file) ? fs.readFileSync(file, 'utf8') : '';
|
||||
const oldHook = '#!/bin/sh\n# gstack-redact pre-push (managed)\n_input="$(cat)"\nprintf "%s" "$_input" | "$(git rev-parse --git-path hooks/pre-push.local)" "$@"\n';
|
||||
const unmanagedHook = '#!/bin/sh\n# gstack-redact pre-push (managed) extra\nexit 42\n';
|
||||
const localPolicy = '#!/bin/sh\ncat > "$HOME/received"\nexit 37\n';
|
||||
const refs = 'refs/heads/a aaaa refs/heads/a bbbb\nrefs/heads/b cccc refs/heads/b dddd\n';
|
||||
|
||||
export function createShipHookFixture(id: ShipHookCase, root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), 'shook-'))) {
|
||||
fs.mkdirSync(root, { recursive: true });
|
||||
fs.chmodSync(root, 0o700);
|
||||
const repo = path.join(root, 'project');
|
||||
const home = path.join(root, 'home');
|
||||
const state = path.join(root, 'state');
|
||||
const installed = path.join(home, '.claude/skills/gstack/bin');
|
||||
const receipts = path.join(root, 'receipts');
|
||||
const protectedFiles = new Map<string, string>();
|
||||
const write = (file: string, text: string, executable = false) => {
|
||||
fs.mkdirSync(path.dirname(file), { recursive: true });
|
||||
const target = fs.lstatSync(file, { throwIfNoEntry: false }) ? fs.realpathSync(file) : path.join(fs.realpathSync(path.dirname(file)), path.basename(file));
|
||||
const relative = path.relative(fs.realpathSync(root), target);
|
||||
if (relative.startsWith('..') || path.isAbsolute(relative)) throw new Error('fixture write escaped its temporary root');
|
||||
fs.writeFileSync(file, text, { mode: executable ? 0o755 : 0o600 });
|
||||
protectedFiles.set(file, text);
|
||||
};
|
||||
fs.mkdirSync(repo);
|
||||
fs.mkdirSync(state);
|
||||
const env = { HOME: home, GSTACK_HOME: state, GSTACK_STATE_ROOT: state, CLAUDE_PLUGIN_DATA: '',
|
||||
CLAUDE_CONFIG_DIR: path.join(root, 'claude-config'), GIT_CONFIG_GLOBAL: '/dev/null', GIT_CONFIG_SYSTEM: '/dev/null',
|
||||
GIT_CONFIG_NOSYSTEM: '1', GIT_CONFIG_COUNT: '0',
|
||||
PATH: `${path.dirname(process.execPath)}:${process.env.PATH ?? '/usr/bin:/bin'}` };
|
||||
const initialized = spawnSync('git', ['init', '-q', repo], { env, encoding: 'utf8', timeout: 10000 });
|
||||
if (initialized.status !== 0) throw new Error(initialized.stderr);
|
||||
const hook = path.join(repo, '.git/hooks/pre-push');
|
||||
write(hook, id === 'ship-managed-hook-refresh' ? oldHook : unmanagedHook, true);
|
||||
if (id !== 'ship-unmanaged-hook-consent') write(hook + '.local', localPolicy, true);
|
||||
if (id === 'ship-managed-hook-refresh') protectedFiles.delete(hook);
|
||||
write(path.join(state, 'config.yaml'), 'redact_prepush_hook: true\n');
|
||||
write(path.join(installed, 'gstack-config'), `#!/bin/sh\nprintf 'CONFIG:%s\\n' "$*" >> ${quote(receipts)}\nexec bash ${quote(path.join(ROOT, 'bin/gstack-config'))} "$@"\n`, true);
|
||||
write(path.join(installed, 'gstack-redact'), `#!/bin/sh\nprintf 'INSTALL:%s\\n' "$*" >> ${quote(receipts)}\nexec ${quote(process.execPath)} ${quote(path.join(ROOT, 'bin/gstack-redact'))} "$@"\n`, true);
|
||||
const workflow = readWorkflowExcerpt('ship/SKILL.md', '**Credential pre-push guard', '**Idempotency check');
|
||||
write(path.join(repo, 'workflow.md'), workflow);
|
||||
const guard = workflow.match(/```bash\n([\s\S]*?)```/)![1].trim();
|
||||
const readable = new Set([path.join(repo, 'workflow.md'), hook, hook + '.local']);
|
||||
const commands = new Set([guard, 'git config --get core.hooksPath', 'git rev-parse --git-path hooks/pre-push', 'git rev-parse --git-path hooks/pre-push.local']);
|
||||
const executions: Array<{ tool: string; input: Record<string, unknown>; allowed: boolean }> = [];
|
||||
const interactions: Array<{ disposition: string; input: unknown }> = [];
|
||||
const prompt = `Load gstack's /ship credential pre-push guard from workflow.md and run only that workflow excerpt. Inspect the actual repository, apply its rules, report the outcome, and stop before pushing.
|
||||
|
||||
Declared fixture interactions: Read workflow.md and the local hook files; Bash for the supplied guard block verbatim and these read-only commands: git config --get core.hooksPath, git rev-parse --git-path hooks/pre-push, git rev-parse --git-path hooks/pre-push.local. Run each listed inspection command in its own Bash call; do not chain commands or append an echo/status probe. A missing core.hooksPath setting may return exit 1 without requiring another command. One native AskUserQuestion about hook modification is supported if the workflow calls for it; the owner declines modifications. Use Read rather than shell commands to inspect files. No other command, file edit, direct receipt access, directory change, background command, commit, push, config change, installation phase or unrelated /ship phase is supported. HOME and GSTACK_HOME already belong to this fixture; do not change them.`;
|
||||
const invalid = (tool: string, input: Record<string, unknown>) => {
|
||||
if (tool === 'Read') return typeof input.file_path !== 'string' || !readable.has(path.resolve(repo, input.file_path)) ? 'Read outside declared fixture paths' : undefined;
|
||||
if (tool === 'Bash') return typeof input.command !== 'string' || input.run_in_background || !commands.has(input.command.trim()) ? 'Bash outside declared fixture commands' : undefined;
|
||||
if (tool === 'AskUserQuestion') return interactions.some(event => event.disposition === 'declined') ? 'Duplicate owner question' : undefined;
|
||||
return 'Tool outside declared fixture interactions';
|
||||
};
|
||||
const preToolUse: HookCallback = async input => {
|
||||
if (input.hook_event_name !== 'PreToolUse') throw new Error('unexpected native hook event');
|
||||
const toolInput = input.tool_input as Record<string, unknown>;
|
||||
const reason = invalid(input.tool_name, toolInput);
|
||||
executions.push({ tool: input.tool_name, input: toolInput, allowed: !reason });
|
||||
return { hookSpecificOutput: { hookEventName: 'PreToolUse', permissionDecision: reason ? 'deny' : input.tool_name === 'AskUserQuestion' ? 'ask' : 'allow',
|
||||
...(reason ? { permissionDecisionReason: reason } : {}),
|
||||
...(input.tool_name === 'Bash' && !reason ? { updatedInput: { command: toolInput.command, timeout: 10000, run_in_background: false } } : {}) } };
|
||||
};
|
||||
const canUseTool: CanUseTool = async (tool, input) => {
|
||||
const reason = invalid(tool, input);
|
||||
if (reason) {
|
||||
interactions.push({ disposition: 'unsupported', input });
|
||||
return { behavior: 'deny', message: reason };
|
||||
}
|
||||
if (tool !== 'AskUserQuestion') return { behavior: 'allow', updatedInput: input };
|
||||
const questions = input.questions as Array<{ question: string; multiSelect?: boolean; options: Array<{ label: string; description: string }> }>;
|
||||
const question = questions?.[0];
|
||||
const decline = question?.options?.find(option => /^(?:[A-D][).]\s*)?(no\b|decline\b|do not\b|skip\b|leave\b|keep\b)/i.test(option.label));
|
||||
if (questions?.length !== 1 || question.multiSelect || !/hook|guard|chain|credential/i.test(question.question) || !decline) {
|
||||
interactions.push({ disposition: 'unsupported', input });
|
||||
return { behavior: 'deny', message: 'Only one hook modification question with a decline option is supported.' };
|
||||
}
|
||||
interactions.push({ disposition: 'declined', input });
|
||||
return { behavior: 'allow', updatedInput: { ...input, answers: { [question.question]: decline.label } } };
|
||||
};
|
||||
const snapshot = () => ({ receipts: read(receipts), hook: read(hook), localHook: read(hook + '.local'), localExists: fs.existsSync(hook + '.local'),
|
||||
executions, interactions, changedProtectedFiles: [...protectedFiles].filter(([file, bytes]) => read(file) !== bytes).map(([file]) => path.relative(root, file)),
|
||||
workflowSha256: createHash('sha256').update(workflow).digest('hex') });
|
||||
return { id, root, repo, home, env, hook, receipts, guard, prompt, preToolUse, canUseTool, snapshot };
|
||||
}
|
||||
|
||||
export function shipHookFailures(fixture: ReturnType<typeof createShipHookFixture>, result: Pick<AgentSdkResult, 'exitReason' | 'output'>) {
|
||||
const evidence = fixture.snapshot();
|
||||
const failures: string[] = [];
|
||||
const check = (ok: boolean, message: string) => { if (!ok) failures.push(message); };
|
||||
check(result.exitReason === 'success', `actor ended: ${result.exitReason}`);
|
||||
check(evidence.changedProtectedFiles.length === 0, 'protected policy or fixture files changed');
|
||||
check(!evidence.executions.some(event => !event.allowed) && !evidence.interactions.some(event => event.disposition === 'unsupported'), 'undeclared interaction');
|
||||
check(evidence.executions.some(event => event.tool === 'Read' && event.allowed && path.resolve(fixture.repo, event.input.file_path as string) === path.join(fixture.repo, 'workflow.md')), 'workflow was not read through native hook');
|
||||
check(evidence.executions.some(event => event.tool === 'Bash' && event.allowed && (event.input.command as string).trim() === fixture.guard), 'guard was not executed through native hook');
|
||||
check(evidence.receipts.includes('CONFIG:get redact_prepush_hook\n'), 'actual config helper was not called');
|
||||
const installs = evidence.receipts.match(/^INSTALL:install-prepush-hook$/gm)?.length ?? 0;
|
||||
let callback: { status: number | null; received: string } | undefined;
|
||||
if (fixture.id === 'ship-managed-hook-refresh') {
|
||||
check(installs === 1, 'expected exactly one actual installer call');
|
||||
check(evidence.hook !== oldHook && evidence.hook.includes('_input="$(cat; printf x)"'), 'managed wrapper was not refreshed');
|
||||
check(evidence.localHook === localPolicy, 'local policy changed');
|
||||
check(evidence.interactions.length === 0, 'managed refresh unnecessarily requested consent');
|
||||
const invoked = spawnSync('bash', [fixture.hook, 'origin', 'synthetic'], { cwd: fixture.repo, env: fixture.env, input: refs, encoding: 'utf8', timeout: 10000 });
|
||||
callback = { status: invoked.status, received: read(path.join(fixture.home, 'received')) };
|
||||
check(callback.status === 37 && callback.received === refs, 'refreshed wrapper lost local policy status or complete stdin');
|
||||
check(/refresh|updat|install|current/i.test(result.output), 'refresh outcome was not acknowledged');
|
||||
} else {
|
||||
check(installs === 0, 'installer invoked without applicable consent');
|
||||
check(evidence.hook === unmanagedHook, 'unmanaged policy changed');
|
||||
if (fixture.id === 'ship-unmanaged-hook-consent') {
|
||||
check(evidence.interactions.filter(event => event.disposition === 'declined').length === 1, 'unmanaged-hook consent was not requested');
|
||||
check(!evidence.localExists, 'declined install created a local policy');
|
||||
check(/declin|unchanged|not install|not modif|preserv|left|leave/i.test(result.output), 'declined modification was not acknowledged');
|
||||
} else {
|
||||
check(evidence.localHook === localPolicy, 'existing local policy changed');
|
||||
check(evidence.interactions.length === 0, 'existing local policy requires manual integration, not consent to overwrite');
|
||||
check(/manual/i.test(result.output), 'manual integration was not reported');
|
||||
}
|
||||
}
|
||||
return { failures, evidence: { ...evidence, callback } };
|
||||
}
|
||||
|
||||
export async function runShipHookActor(id: ShipHookCase, record: (entry: EvalTestEntry) => void, injectedQuery?: QueryProvider, artifactDirectory?: string) {
|
||||
const artifacts = artifactDirectory ?? process.env.GSTACK_EVAL_DIR;
|
||||
if (!artifacts) throw new Error('ship hook actor requires an explicit GSTACK_EVAL_DIR for durable evidence');
|
||||
fs.mkdirSync(artifacts, { recursive: true, mode: 0o700 });
|
||||
const { query } = await import('@anthropic-ai/claude-agent-sdk');
|
||||
const { runAgentSdkTest, resolveClaudeBinary } = await import('./agent-sdk-runner');
|
||||
let fixture = createShipHookFixture(id);
|
||||
const started = Date.now();
|
||||
const attempts: Array<{ events: SDKMessage[]; evidence?: ReturnType<typeof fixture.snapshot> }> = [];
|
||||
let result: AgentSdkResult | undefined;
|
||||
let finalEvidence: unknown;
|
||||
let failure: unknown;
|
||||
let passed = false;
|
||||
const evidenceFile = path.join(artifacts, `${id}-${randomUUID()}.json`);
|
||||
try {
|
||||
const binary = injectedQuery ? undefined : resolveClaudeBinary();
|
||||
if (!injectedQuery && !binary) throw new Error('native ship hook actor requires the pinned Claude CLI');
|
||||
result = await runAgentSdkTest({ systemPrompt: { type: 'preset', preset: 'claude_code' }, userPrompt: fixture.prompt,
|
||||
workingDirectory: fixture.repo, env: fixture.env, maxTurns: 12, signal: AbortSignal.timeout(CAPTURE_MS - 20000),
|
||||
allowedTools: ['Read', 'Bash', 'AskUserQuestion'], settingSources: [], testName: id,
|
||||
pathToClaudeCodeExecutable: binary ?? undefined,
|
||||
canUseTool: (...args) => fixture.canUseTool(...args),
|
||||
onRetry: () => {
|
||||
attempts.at(-1)!.evidence = fixture.snapshot();
|
||||
const root = fixture.root;
|
||||
fs.rmSync(root, { recursive: true, force: true });
|
||||
fixture = createShipHookFixture(id, root);
|
||||
},
|
||||
queryProvider: options => {
|
||||
const attempt: typeof attempts[number] = { events: [] };
|
||||
attempts.push(attempt);
|
||||
const stream = (injectedQuery ?? query)({ ...options, options: { ...options.options, allowedTools: [],
|
||||
hooks: { PreToolUse: [{ hooks: [(...args) => fixture.preToolUse(...args)] }] } } });
|
||||
return new Proxy(stream, { get(target, property) {
|
||||
if (property === Symbol.asyncIterator) return async function* () {
|
||||
try { for await (const event of target) { attempt.events.push(event); yield event; } }
|
||||
finally { attempt.evidence = fixture.snapshot(); }
|
||||
};
|
||||
const value = Reflect.get(target, property, target);
|
||||
return typeof value === 'function' ? value.bind(target) : value;
|
||||
} });
|
||||
},
|
||||
});
|
||||
const verdict = shipHookFailures(fixture, result);
|
||||
finalEvidence = verdict.evidence;
|
||||
if (verdict.failures.length) throw new Error(verdict.failures.join('; '));
|
||||
passed = true;
|
||||
} catch (error) {
|
||||
failure = error;
|
||||
throw error;
|
||||
} finally {
|
||||
try {
|
||||
const output = JSON.stringify({ assistant: result?.output, error: failure === undefined ? undefined : String(failure),
|
||||
evidence: finalEvidence ?? fixture.snapshot(), attempts });
|
||||
fs.writeFileSync(evidenceFile, output + '\n', { mode: 0o600 });
|
||||
record({ name: id, suite: 'ship-hook-boundary', tier: 'e2e', passed, duration_ms: Date.now() - started,
|
||||
cost_usd: result?.costUsd ?? 0, transcript: result?.events, prompt: fixture.prompt, turns_used: result?.turnsUsed,
|
||||
model: result?.model, output, error: failure === undefined ? undefined : String(failure),
|
||||
exit_reason: passed ? 'success' : result?.exitReason === 'success' ? 'assertion_failed' : result?.exitReason ?? 'runner_error' });
|
||||
} finally { fs.rmSync(fixture.root, { recursive: true, force: true }); }
|
||||
}
|
||||
return evidenceFile;
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
|
||||
const root = path.resolve(import.meta.dir, '../..');
|
||||
|
||||
export function createReadinessFixture(kind: 'ready' | 'unknown') {
|
||||
const workDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-'));
|
||||
const home = path.join(workDir, '.fixture-home');
|
||||
const bin = path.join(workDir, '.fixture-bin');
|
||||
fs.mkdirSync(home); fs.mkdirSync(bin);
|
||||
const init = spawnSync('git', ['init', '--quiet'], { cwd: workDir, timeout: 10_000 });
|
||||
if (init.status !== 0) throw new Error('readiness fixture git init failed');
|
||||
fs.writeFileSync(path.join(workDir, '.gbrain-source'), 'client-fixture\n');
|
||||
const stateDir = path.join(home, '.gstack');
|
||||
fs.mkdirSync(stateDir);
|
||||
fs.writeFileSync(path.join(stateDir, '.gbrain-sync-state.json'), JSON.stringify({
|
||||
schema_version: 1, last_writer: 'gstack-gbrain-sync', last_stages: [{
|
||||
name: 'code', ran: true, ok: true,
|
||||
detail: { status: 'ok', source_id: 'client-fixture', source_path: workDir },
|
||||
}],
|
||||
}, null, 2));
|
||||
const guidance = '<!-- gstack-gbrain-search-guidance:start -->\nExisting search guidance\n<!-- gstack-gbrain-search-guidance:end -->';
|
||||
fs.writeFileSync(path.join(workDir, 'CLAUDE.md'), kind === 'unknown' ? `# Fixture\n${guidance}\n` : '# Fixture\n');
|
||||
const skill = fs.readFileSync(path.join(root, 'sync-gbrain/SKILL.md'), 'utf8');
|
||||
const start = skill.indexOf('## Step 4: Refresh');
|
||||
const end = skill.indexOf('## Concurrency note', start);
|
||||
if (start < 0 || end < 0) throw new Error('sync-gbrain Step 4/5 fixture anchors missing');
|
||||
fs.writeFileSync(path.join(workDir, 'readiness.md'), skill.slice(start, end)
|
||||
.replaceAll('~/.claude/skills/gstack/bin/gstack-gbrain-read-capability.ts', path.join(root, 'bin/gstack-gbrain-read-capability.ts')));
|
||||
const log = path.join(home, 'gbrain-calls');
|
||||
fs.writeFileSync(path.join(bin, 'gbrain'), `#!/usr/bin/env bun
|
||||
import { appendFileSync } from 'node:fs';
|
||||
const args = process.argv.slice(2).join(' ');
|
||||
appendFileSync(${JSON.stringify(log)}, args + '\\n');
|
||||
if (args === 'sources list --json') console.log(${JSON.stringify(JSON.stringify({ sources: [{ id: 'client-fixture', local_path: workDir }] }))});
|
||||
else if (args === 'list --source client-fixture --limit 1') console.log('code/fixture/readme\\tcode\\t2026-09-24\\tReadme');
|
||||
else if (args === 'get code/fixture/readme --source client-fixture --json') ${kind === 'ready'
|
||||
? `console.log(${JSON.stringify(JSON.stringify({ source_id: 'client-fixture', slug: 'code/fixture/readme', content: '# Readme' }))});`
|
||||
: "{ console.error('temporary read failure'); process.exit(2); }"}
|
||||
else { console.error('unsupported operation'); process.exit(3); }
|
||||
`);
|
||||
fs.chmodSync(path.join(bin, 'gbrain'), 0o755);
|
||||
if (process.platform === 'win32') {
|
||||
fs.writeFileSync(path.join(bin, 'gbrain.cmd'), `@echo off\r\n"${process.execPath}" "%~dp0gbrain" %*\r\n`);
|
||||
}
|
||||
const pathKey = Object.keys(process.env).find(key => key.toLowerCase() === 'path') ?? 'PATH';
|
||||
const pin = fs.readFileSync(path.join(workDir, '.gbrain-source'), 'utf8');
|
||||
const state = fs.readFileSync(path.join(stateDir, '.gbrain-sync-state.json'), 'utf8');
|
||||
return {
|
||||
workDir,
|
||||
env: { HOME: home, GSTACK_HOME: stateDir, [pathKey]: `${bin}${path.delimiter}${process.env[pathKey] ?? ''}` },
|
||||
guidance,
|
||||
calls: () => fs.existsSync(log) ? fs.readFileSync(log, 'utf8').trim().split('\n') : [],
|
||||
content: () => fs.readFileSync(path.join(workDir, 'CLAUDE.md'), 'utf8'),
|
||||
sourceIntact: () => fs.readFileSync(path.join(workDir, '.gbrain-source'), 'utf8') === pin
|
||||
&& fs.readFileSync(path.join(stateDir, '.gbrain-sync-state.json'), 'utf8') === state
|
||||
&& !fs.existsSync(path.join(workDir, 'code')),
|
||||
cleanup: () => fs.rmSync(workDir, { recursive: true, force: true }),
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
export function readinessVerdictProblems(kind: 'ready' | 'unknown', output: string): string[] {
|
||||
const problems: string[] = [];
|
||||
if (kind === 'ready') {
|
||||
const capability = [...output.matchAll(/^\s*Capability[\s.:-]+(OK|FIX|WARN|ERR)\b[^\r\n]*/gim)];
|
||||
if (capability.length !== 1 || capability[0]![1]?.toUpperCase() !== 'OK' || !/\bsource-scoped page read verified\b/i.test(capability[0]![0]))
|
||||
problems.push('ready result lacks verified source-scoped Capability OK');
|
||||
const overallStatuses = [...output.matchAll(/\b(?:gbrain\s+status|verdict)\s*:\s*(GREEN|YELLOW|RED)\b/gi)].map((match) => match[1].toUpperCase());
|
||||
if (overallStatuses.includes('GREEN'))
|
||||
problems.push('ready result claims GREEN with unavailable rows');
|
||||
if (overallStatuses.includes('RED'))
|
||||
problems.push('ready result has conflicting overall verdict');
|
||||
if (!overallStatuses.includes('YELLOW'))
|
||||
problems.push('ready result lacks YELLOW overall verdict');
|
||||
} else {
|
||||
if (!/unknown|unverified|retry|could not verify/i.test(output)) problems.push('unknown status not reported');
|
||||
if (!/\bCapability\s*[.: ]+\s*WARN\b|\b(?:gbrain\s+status|verdict)\s*:\s*YELLOW\b/i.test(output))
|
||||
problems.push('unknown result lacks WARN/YELLOW verdict');
|
||||
if (/\b(?:gbrain\s+status|verdict)\s*:\s*GREEN\b|\bCapability\s*[.: ]+\s*OK\b/i.test(output))
|
||||
problems.push('unknown result claims GREEN or capability OK');
|
||||
}
|
||||
for (const claim of output.matchAll(/\b(?:semantic search|writes?|write readiness|write availability)[^.!?\n]{0,60}\b(?:ready|verified|proven|confirmed|working)\b/gi)) {
|
||||
if (!/\b(?:not|never|without|unknown|unverified)\b/i.test(claim[0]))
|
||||
problems.push('read-only check claims semantic search or write readiness');
|
||||
}
|
||||
return problems;
|
||||
}
|
||||
@@ -21,6 +21,9 @@
|
||||
* Each test lists the file patterns that, if changed, require the test to run.
|
||||
*/
|
||||
export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
'investigate-owned-completion': ['investigate/**', 'freeze/**', 'guard/**', 'unfreeze/**', 'careful/bin/hook-extract.sh', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/workflow-boundaries-fixture.ts', 'test/workflow-boundaries-fixture.test.ts', 'test/skill-e2e-investigate-owned-completion.test.ts'],
|
||||
'investigate-owned-abort': ['investigate/**', 'freeze/**', 'guard/**', 'unfreeze/**', 'careful/bin/hook-extract.sh', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/workflow-boundaries-fixture.ts', 'test/workflow-boundaries-fixture.test.ts', 'test/skill-e2e-investigate-owned-termination.test.ts'],
|
||||
'investigate-owned-ending-error': ['investigate/**', 'freeze/**', 'guard/**', 'unfreeze/**', 'careful/bin/hook-extract.sh', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/workflow-boundaries-fixture.ts', 'test/workflow-boundaries-fixture.test.ts', 'test/skill-e2e-investigate-owned-termination.test.ts'],
|
||||
'shared-libs-review-path-eligibility': ['review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/shared-libs-fixture.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/shared-libs-revalidation-prompt.test.ts', 'test/shared-libs-source-reads.test.ts', 'test/fixtures/shared-libs-resolved-reads-public.json'],
|
||||
'shared-libs-review-index-flags': ['review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/shared-libs-fixture.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/shared-libs-revalidation-prompt.test.ts', 'test/fixtures/shared-libs-paths-max-turns-public.json', 'test/shared-libs-source-reads.test.ts', 'test/fixtures/shared-libs-resolved-reads-public.json'],
|
||||
'shared-libs-review-prior-coverage': ['review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/shared-libs-fixture.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/shared-libs-revalidation-prompt.test.ts', 'test/shared-libs-source-reads.test.ts', 'test/fixtures/shared-libs-resolved-reads-public.json'],
|
||||
@@ -80,7 +83,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
'qa-b8-checkout': ['test/session-runner-stream-lifecycle.test.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval-checkout.html', 'test/fixtures/qa-eval-checkout-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts',
|
||||
'scripts/resolvers/testing.ts'
|
||||
],
|
||||
'qa-only-no-fix': ['test/session-runner-stream-lifecycle.test.ts', 'qa-only/**', 'qa/templates/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts'],
|
||||
'qa-only-no-fix': ['test/qa-only-capability.test.ts', 'test/session-runner-stream-lifecycle.test.ts', 'qa-only/**', 'qa/templates/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts'],
|
||||
'qa-fix-loop': ['test/session-runner-stream-lifecycle.test.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts',
|
||||
'test/qa-fix-loop-fixture.test.ts', 'scripts/resolvers/testing.ts'
|
||||
],
|
||||
@@ -102,7 +105,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
|
||||
// Review Army (specialist dispatch)
|
||||
'review-army-migration-safety': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts'],
|
||||
'review-army-perf-n-plus-one': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts'],
|
||||
'review-army-perf-n-plus-one': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts', 'test/review-army-budget.test.ts', 'test/review-n-plus-one-contract.test.ts', 'test/fixtures/review-n-plus-one-dispatch.json'],
|
||||
'review-army-delivery-audit': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review.ts', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts'],
|
||||
'review-army-quality-score': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts'],
|
||||
'review-army-json-findings': ['test/session-runner-stream-lifecycle.test.ts', 'review/**', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts'],
|
||||
@@ -945,6 +948,18 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
],
|
||||
|
||||
// Ship
|
||||
'ship-managed-hook-refresh': ['ship/**', 'bin/gstack-redact', 'bin/gstack-config', 'scripts/gen-skill-docs.ts',
|
||||
'test/helpers/ship-hook-actor.ts', 'test/ship-hook-actor.test.ts', 'test/ship-hook-refresh.test.ts',
|
||||
'test/helpers/workflow-excerpt.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-ship-hook-refresh.test.ts',
|
||||
'test/paid-pr-profile.test.ts'],
|
||||
'ship-unmanaged-hook-consent': ['ship/**', 'bin/gstack-redact', 'bin/gstack-config', 'scripts/gen-skill-docs.ts',
|
||||
'test/helpers/ship-hook-actor.ts', 'test/ship-hook-actor.test.ts', 'test/ship-hook-refresh.test.ts',
|
||||
'test/helpers/workflow-excerpt.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-ship-hook-consent.test.ts',
|
||||
'test/paid-pr-profile.test.ts'],
|
||||
'ship-local-hook-preservation': ['ship/**', 'bin/gstack-redact', 'bin/gstack-config', 'scripts/gen-skill-docs.ts',
|
||||
'test/helpers/ship-hook-actor.ts', 'test/ship-hook-actor.test.ts', 'test/ship-hook-refresh.test.ts',
|
||||
'test/helpers/workflow-excerpt.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-ship-hook-consent.test.ts',
|
||||
'test/paid-pr-profile.test.ts'],
|
||||
'ship-base-branch': ['test/session-runner-stream-lifecycle.test.ts', 'ship/**', 'bin/gstack-repo-mode', 'test/skill-e2e-review-attribution.test.ts',
|
||||
'scripts/resolvers/testing.ts'
|
||||
],
|
||||
@@ -1356,6 +1371,8 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
'scripts/resolvers/gbrain.ts',
|
||||
'test/skill-e2e-gbrain-roundtrip-local.test.ts',
|
||||
],
|
||||
'sync-gbrain-read-ready': ['sync-gbrain/SKILL.md.tmpl', 'sync-gbrain/SKILL.md', 'bin/gstack-gbrain-read-capability.ts', 'lib/gbrain-exec.ts', 'test/helpers/sync-gbrain-readiness-fixture.ts', 'test/helpers/sync-gbrain-readiness-verdict.ts', 'test/skill-e2e-sync-gbrain-readiness.test.ts'],
|
||||
'sync-gbrain-read-unknown': ['sync-gbrain/SKILL.md.tmpl', 'sync-gbrain/SKILL.md', 'bin/gstack-gbrain-read-capability.ts', 'lib/gbrain-exec.ts', 'test/helpers/sync-gbrain-readiness-fixture.ts', 'test/helpers/sync-gbrain-readiness-verdict.ts', 'test/skill-e2e-sync-gbrain-readiness.test.ts'],
|
||||
|
||||
// WS2 arm benchmark — with-skill vs without-skill agentic arms scored on
|
||||
// the git diff left behind (research instrument, never a release gate).
|
||||
@@ -1416,6 +1433,9 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
* Must have exactly the same keys as E2E_TOUCHFILES.
|
||||
*/
|
||||
export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
|
||||
'investigate-owned-completion': 'gate',
|
||||
'investigate-owned-abort': 'gate',
|
||||
'investigate-owned-ending-error': 'gate',
|
||||
'shared-libs-review-path-eligibility': 'gate',
|
||||
'shared-libs-review-index-flags': 'gate',
|
||||
'shared-libs-review-prior-coverage': 'gate',
|
||||
@@ -1491,6 +1511,8 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
|
||||
// GBrain CLI round-trip — periodic per Voyage embedding cost (~$0.001/run)
|
||||
// and external-API-dependency (skips cleanly if VOYAGE_API_KEY unset).
|
||||
'gbrain-roundtrip-local': 'periodic',
|
||||
'sync-gbrain-read-ready': 'periodic',
|
||||
'sync-gbrain-read-unknown': 'periodic',
|
||||
'office-hours-forcing-energy': 'periodic', // D2a demotion 2026-08: posture score, periodic-grade signal (sibling precedent at office-hours-tone)
|
||||
// 'office-hours-builder-wildness' retiered to periodic in v1.32 contributor
|
||||
// wave: this is an LLM-judge creativity score (axis_a ≥4 on a "wildness"
|
||||
@@ -1633,6 +1655,9 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
|
||||
// Ship — gate (end-to-end ship path)
|
||||
'ship-base-branch': 'gate',
|
||||
'ship-local-workflow': 'gate',
|
||||
'ship-managed-hook-refresh': 'gate',
|
||||
'ship-unmanaged-hook-consent': 'gate',
|
||||
'ship-local-hook-preservation': 'gate',
|
||||
'ship-coverage-audit': 'gate',
|
||||
'ship-triage': 'gate',
|
||||
'ship-docsync': 'gate',
|
||||
@@ -1831,6 +1856,7 @@ export const LLM_JUDGE_TOUCHFILES: Record<string, string[]> = {
|
||||
'retro/SKILL.md instructions': ['retro/sections/**', 'retro/SKILL.md', 'retro/SKILL.md.tmpl', 'test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-input.ts', 'test/helpers/workflow-judge-cache.ts', 'test/workflow-judge-cache.test.ts', 'scripts/eval-input-cache.ts', 'test/eval-input-cache.test.ts', 'test/workflow-judge-input.test.ts', 'test/helpers/workflow-excerpt.ts'],
|
||||
'qa-only/SKILL.md workflow': ['qa-only/SKILL.md', 'qa-only/SKILL.md.tmpl', 'test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-input.ts', 'test/helpers/workflow-judge-cache.ts', 'test/workflow-judge-cache.test.ts', 'scripts/eval-input-cache.ts', 'test/eval-input-cache.test.ts', 'test/workflow-judge-input.test.ts', 'test/helpers/workflow-excerpt.ts'],
|
||||
'gstack-upgrade/SKILL.md upgrade flow': ['gstack-upgrade/SKILL.md', 'gstack-upgrade/SKILL.md.tmpl', 'test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-input.ts', 'test/helpers/workflow-judge-cache.ts', 'test/workflow-judge-cache.test.ts', 'scripts/eval-input-cache.ts', 'test/eval-input-cache.test.ts', 'test/workflow-judge-input.test.ts', 'test/helpers/workflow-excerpt.ts'],
|
||||
'sync-gbrain/SKILL.md read-only readiness': ['sync-gbrain/SKILL.md', 'sync-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-read-capability.ts', 'test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-input.ts', 'test/helpers/workflow-judge-cache.ts', 'test/workflow-judge-cache.test.ts', 'scripts/eval-input-cache.ts', 'test/eval-input-cache.test.ts', 'test/workflow-judge-input.test.ts'],
|
||||
|
||||
// Voice directive
|
||||
'voice directive tone': ['scripts/resolvers/preamble.ts', 'review/SKILL.md', 'review/SKILL.md.tmpl', 'scripts/gen-skill-docs.ts', 'test/skill-llm-eval.test.ts'],
|
||||
|
||||
@@ -0,0 +1,267 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import type { CanUseTool, HookCallback, SDKMessage } from '@anthropic-ai/claude-agent-sdk';
|
||||
import type { AgentSdkResult, QueryProvider } from './agent-sdk-runner';
|
||||
import type { EvalTestEntry } from './eval-store';
|
||||
import { CAPTURE_MS } from './eval-budgets';
|
||||
|
||||
export type BoundaryCase = 'investigate-owned-completion' | 'investigate-owned-abort' |
|
||||
'investigate-owned-ending-error';
|
||||
|
||||
const ROOT = path.resolve(import.meta.dir, '../..');
|
||||
const quote = (value: string) => `'${value.replaceAll("'", "'\"'\"'")}'`;
|
||||
const read = (file: string) => fs.existsSync(file) ? fs.readFileSync(file, 'utf8') : '';
|
||||
|
||||
export function createBoundaryFixture(id: BoundaryCase, root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), 'gbound-'))) {
|
||||
fs.mkdirSync(root, { recursive: true });
|
||||
root = fs.realpathSync(root);
|
||||
fs.chmodSync(root, 0o700);
|
||||
const repo = path.join(root, 'project');
|
||||
const home = path.join(root, 'home');
|
||||
const state = path.join(root, 'state');
|
||||
const installed = path.join(home, '.claude/skills/gstack');
|
||||
const receipts = path.join(root, 'receipts');
|
||||
const boundary = path.join(state, 'freeze-dir.txt');
|
||||
const source = path.join(repo, 'src/value.js');
|
||||
const interactions: Array<{ tool: string; disposition: string; input: unknown; owner?: string }> = [];
|
||||
const executions: Array<{ tool: string; input: Record<string, unknown>; allowed: boolean }> = [];
|
||||
const commands = new Map<string, string>();
|
||||
const readable = new Set([path.join(repo, 'workflow.md'), source]);
|
||||
const protectedFiles = new Map<string, string>();
|
||||
const write = (file: string, text: string, executable = false) => {
|
||||
let ancestor = path.dirname(file);
|
||||
while (!fs.lstatSync(ancestor, { throwIfNoEntry: false })) ancestor = path.dirname(ancestor);
|
||||
const resolved = fs.realpathSync(ancestor);
|
||||
if (resolved !== root && !resolved.startsWith(root + path.sep)) throw new Error('fixture write escapes its isolated root');
|
||||
fs.mkdirSync(path.dirname(file), { recursive: true });
|
||||
const parent = fs.realpathSync(path.dirname(file));
|
||||
if (!parent.startsWith(root + path.sep) ||
|
||||
(fs.lstatSync(file, { throwIfNoEntry: false }) && !fs.realpathSync(file).startsWith(root + path.sep))) {
|
||||
throw new Error('fixture write escapes its isolated root');
|
||||
}
|
||||
fs.writeFileSync(file, text, { mode: executable ? 0o755 : 0o600 });
|
||||
protectedFiles.set(file, text);
|
||||
};
|
||||
fs.mkdirSync(repo);
|
||||
fs.mkdirSync(state);
|
||||
const env = {
|
||||
HOME: home, GSTACK_HOME: state, CLAUDE_PLUGIN_DATA: '', CLAUDE_CONFIG_DIR: path.join(root, 'claude-config'),
|
||||
GIT_CONFIG_GLOBAL: '/dev/null', GIT_CONFIG_SYSTEM: '/dev/null', GIT_CONFIG_NOSYSTEM: '1',
|
||||
PATH: `${path.join(root, 'bin')}:${path.dirname(process.execPath)}:${process.env.PATH ?? '/usr/bin:/bin'}`,
|
||||
};
|
||||
const generated = read(path.join(ROOT, 'investigate/SKILL.md'));
|
||||
const scope = generated.match(/^## Scope Lock\n[\s\S]*?(?=^## )/m)?.[0];
|
||||
if (!scope) throw new Error('generated /investigate Scope Lock missing');
|
||||
write(path.join(repo, 'workflow.md'), scope);
|
||||
const blocks = [...scope.matchAll(/```bash\n([\s\S]*?)```/g)].map(match => match[1].trim());
|
||||
if (blocks.length !== 3) throw new Error('Scope Lock must have availability, acquisition and release blocks');
|
||||
commands.set(blocks[0], 'availability');
|
||||
commands.set(blocks[1].replace('<detected-directory>', 'src'), 'acquire');
|
||||
commands.set('bash ./verify.sh', 'verify');
|
||||
const releaseTemplate = blocks[2];
|
||||
for (const file of ['careful/bin/hook-extract.sh', 'freeze/bin/check-freeze.sh']) {
|
||||
write(path.join(installed, file), read(path.join(ROOT, file)), true);
|
||||
}
|
||||
write(path.join(installed, 'freeze/bin/freeze-state-real.sh'), read(path.join(ROOT, 'freeze/bin/freeze-state.sh')), true);
|
||||
write(path.join(installed, 'freeze/bin/freeze-state.sh'), `#!/bin/bash
|
||||
printf 'ACTION:%s:%s\\n' "$1" "\${2:-}" >> ${quote(receipts)}
|
||||
_out=$(bash ${quote(path.join(installed, 'freeze/bin/freeze-state-real.sh'))} "$@")
|
||||
_rc=$?
|
||||
printf '%s\\nRESULT:%s\\n' "$_out" "$_rc" >> ${quote(receipts)}
|
||||
printf '%s\\n' "$_out"
|
||||
exit "$_rc"
|
||||
`, true);
|
||||
write(source, 'export function value() { return 1; }\n');
|
||||
protectedFiles.delete(source);
|
||||
const shellSource = process.platform === 'win32'
|
||||
? spawnSync('bash', ['-c', 'cd "$(cygpath -u "$1")" && pwd -P', '_', path.join(repo, 'src')],
|
||||
{ encoding: 'utf8', timeout: 10_000 })
|
||||
: null;
|
||||
if (shellSource && shellSource.status !== 0) throw new Error(`Cannot resolve Bash fixture source: ${shellSource.stderr}`);
|
||||
const physicalSource = shellSource ? shellSource.stdout.trim() : path.join(repo, 'src');
|
||||
write(path.join(repo, 'verify.sh'), `#!/bin/bash
|
||||
printf 'VERIFY\\n' >> ${quote(receipts)}
|
||||
${id === 'investigate-owned-ending-error'
|
||||
? "printf 'FIXTURE_ENDING_ERROR: required verifier is unavailable; end this investigation without retry.\\n' >&2\n_rc=69"
|
||||
: `${quote(process.execPath)} -e 'import { value } from "./src/value.js"; if (value() !== 2) process.exit(1); console.log("FIXTURE_VERIFIED")'\n_rc=$?`}
|
||||
printf 'VERIFY_STATUS:%s\\n' "$_rc" >> ${quote(receipts)}
|
||||
exit "$_rc"
|
||||
`, true);
|
||||
const prompt = `Load gstack's /investigate Scope Lock workflow from workflow.md. This is a bounded
|
||||
fixture: the root cause is established, value() returns 1 but must return 2. Apply
|
||||
the scope workflow. After acquisition, ask the owner via AskUserQuestion whether
|
||||
to continue or abort. If continuing, use Edit to change src/value.js, then run
|
||||
bash ./verify.sh. Do not run other investigation phases.
|
||||
|
||||
Supported interactions: Read workflow.md and src/value.js; Bash for the supplied
|
||||
scope blocks and verifier; Edit only src/value.js, changing only the return value
|
||||
from 1 to 2; one continue/abort question. Use each entire Bash block verbatim,
|
||||
including its comment-only lines. Substitute
|
||||
src for <detected-directory>, and your returned token for <retained-owner-token>.
|
||||
The only verifier command is exactly bash ./verify.sh. Never combine commands,
|
||||
change directories or run a command in the background.
|
||||
No installation, other file
|
||||
edits, direct state-file access, or deleting state outside the supplied helper.
|
||||
The owner may explicitly abort at that question. Honor that answer.
|
||||
A FIXTURE_ENDING_ERROR verifier result ends the investigation, not a new repair.
|
||||
Finish with the actual outcome and scope disposition. HOME and GSTACK_HOME are
|
||||
already isolated to this fixture. Do not change them.`;
|
||||
const invalidTool = (tool: string, input: Record<string, unknown>): string | undefined => {
|
||||
const continued = interactions.some(event => event.disposition === 'continue-investigation');
|
||||
const owner = read(receipts).match(/^FREEZE_OWNER=([a-f0-9]{32})$/m)?.[1];
|
||||
const owned = !!owner && read(boundary) === `${physicalSource}\ngstack-freeze-v1:${owner}\n`;
|
||||
if (tool === 'Read') {
|
||||
if (typeof input.file_path !== 'string' || !readable.has(path.resolve(repo, input.file_path))) return 'Read is limited to the workflow and declared source paths.';
|
||||
} else if (tool === 'Bash') {
|
||||
if (typeof input.command !== 'string' || input.run_in_background) return 'Only foreground fixture commands are supported.';
|
||||
const command = input.command.trim();
|
||||
if (owner && command === releaseTemplate.replace('<retained-owner-token>', owner)) return;
|
||||
const action = commands.get(command);
|
||||
if (!action) return 'Use one of the declared Bash blocks verbatim; state reads and receipt writes are forbidden.';
|
||||
if (action === 'acquire' && owner) return 'This run already acquired its boundary.';
|
||||
if (action === 'verify' && (!continued || !owned)) return 'Verification requires continuation and the active owned boundary.';
|
||||
} else if (tool === 'Edit') {
|
||||
if (typeof input.file_path !== 'string' || path.resolve(repo, input.file_path) !== source || !continued || !owned) return 'Only the continued investigation under its owned boundary may edit its source file.';
|
||||
const current = read(source);
|
||||
if (typeof input.old_string !== 'string' || !input.old_string || typeof input.new_string !== 'string' ||
|
||||
current !== 'export function value() { return 1; }\n' ||
|
||||
current.replace(input.old_string, input.new_string) !== 'export function value() { return 2; }\n') return 'The only permitted edit changes the existing return value from 1 to 2.';
|
||||
} else if (tool === 'AskUserQuestion') {
|
||||
if (interactions.some(event => event.tool === 'AskUserQuestion')) return 'Only one declared owner question is supported.';
|
||||
} else return 'This tool is outside the declared fixture interactions.';
|
||||
};
|
||||
const preToolUse: HookCallback = async input => {
|
||||
if (input.hook_event_name !== 'PreToolUse') throw new Error('unexpected fixture hook');
|
||||
const toolInput = input.tool_input as Record<string, unknown>;
|
||||
const reason = invalidTool(input.tool_name, toolInput);
|
||||
executions.push({ tool: input.tool_name, input: toolInput, allowed: !reason });
|
||||
if (reason) interactions.push({ tool: input.tool_name, disposition: 'unsupported-tool', input: toolInput });
|
||||
return { hookSpecificOutput: { hookEventName: 'PreToolUse', permissionDecision: reason ? 'deny' : input.tool_name === 'AskUserQuestion' ? 'ask' : 'allow',
|
||||
...(reason ? { permissionDecisionReason: reason } : {}),
|
||||
...(input.tool_name === 'Bash' && !reason ? { updatedInput: { command: toolInput.command, timeout: 10000, run_in_background: false } } : {}) } };
|
||||
};
|
||||
const canUseTool: CanUseTool = async (tool, input) => {
|
||||
const invalid = invalidTool(tool, input);
|
||||
if (invalid) {
|
||||
interactions.push({ tool, disposition: 'unsupported-tool', input });
|
||||
return { behavior: 'deny', message: invalid };
|
||||
}
|
||||
if (tool === 'AskUserQuestion') {
|
||||
const owner = read(boundary).match(/gstack-freeze-v1:([a-f0-9]{32})/)?.[1];
|
||||
const questions = input.questions as Array<{ question: string; options: Array<{ label: string }> }>;
|
||||
const question = questions?.[0];
|
||||
const abort = id === 'investigate-owned-abort';
|
||||
const choice = question?.options?.find(option => (abort ? /\b(abort|stop|cancel)\b/i : /\b(continue|proceed)\b/i).test(option.label));
|
||||
if (!owner || questions?.length !== 1 || !choice || !/continue|proceed|abort/i.test(question.question)) {
|
||||
interactions.push({ tool, disposition: 'unsupported-question', input, owner });
|
||||
return { behavior: 'deny', message: 'Ask only whether to continue or abort, after scope acquisition.' };
|
||||
}
|
||||
interactions.push({ tool, disposition: abort ? 'explicit-abort' : 'continue-investigation', input, owner });
|
||||
return { behavior: 'allow', updatedInput: { ...input, answers: { [question.question]: choice.label } } };
|
||||
}
|
||||
return { behavior: 'allow', updatedInput: input };
|
||||
};
|
||||
const snapshot = () => ({ receipts: read(receipts), boundary: read(boundary), source: read(source), interactions, executions,
|
||||
changedProtectedFiles: [...protectedFiles].filter(([file, bytes]) => read(file) !== bytes).map(([file]) => path.relative(root, file)) });
|
||||
return { id, root, repo, env, prompt, source, receipts, boundary, installed, physicalSource, canUseTool, preToolUse, snapshot };
|
||||
}
|
||||
|
||||
export function boundaryFailures(fixture: ReturnType<typeof createBoundaryFixture>, result: Pick<AgentSdkResult, 'exitReason' | 'events' | 'assistantTurns'>) {
|
||||
const evidence = fixture.snapshot();
|
||||
const session = result.events.find(event => event.type === 'system' && event.subtype === 'init')?.session_id;
|
||||
const assistant = result.assistantTurns.filter(event => event.session_id === session && event.parent_tool_use_id === null && event.message.role === 'assistant');
|
||||
const finalId = assistant.at(-1)?.message.id;
|
||||
const finalText = session && finalId ? assistant.filter(event => event.message.id === finalId)
|
||||
.flatMap(event => event.message.content).filter(block => block.type === 'text').map(block => block.text).join('\n') : '';
|
||||
const failures: string[] = [];
|
||||
const check = (ok: boolean, message: string) => { if (!ok) failures.push(message); };
|
||||
check(result.exitReason === 'success', `actor ended: ${result.exitReason}`);
|
||||
check(evidence.changedProtectedFiles.length === 0, 'protected fixture files changed');
|
||||
check(!evidence.interactions.some(event => /invalid|unsupported/.test(event.disposition)), 'undeclared interaction');
|
||||
check(evidence.executions.some(event => event.tool === 'Bash' && event.allowed), 'registered native hook saw no Bash execution');
|
||||
const owners = [...evidence.receipts.matchAll(/^FREEZE_OWNER=([a-f0-9]{32})$/gm)].map(match => match[1]);
|
||||
check(owners.length === 1, 'actor must acquire exactly one run-owned boundary');
|
||||
check(evidence.receipts.includes(`FREEZE_DIR=${fixture.physicalSource}\n`), 'actor did not acquire the affected module');
|
||||
check(evidence.receipts.includes(`ACTION:release:${owners[0]}\n`), 'actor did not release its acquired owner token');
|
||||
check(evidence.receipts.includes('FREEZE_RELEASED:'), 'helper did not confirm owned cleanup');
|
||||
check(!evidence.boundary && !fs.existsSync(fixture.boundary), 'owned boundary remains');
|
||||
if (fixture.id === 'investigate-owned-abort') {
|
||||
check(evidence.interactions.some(event => event.disposition === 'explicit-abort' && event.owner === owners[0]), 'explicit abort was not delivered after acquisition');
|
||||
check(evidence.source === 'export function value() { return 1; }\n', 'edit occurred despite explicit abort');
|
||||
check(!evidence.receipts.includes('VERIFY\n'), 'verifier ran after abort');
|
||||
check(/abort|stop|cancel/i.test(finalText), 'actor did not acknowledge abort');
|
||||
} else {
|
||||
check(evidence.executions.filter(event => event.tool === 'Edit' && event.allowed).length === 1, 'expected one allowed native Edit');
|
||||
check(evidence.interactions.some(event => event.disposition === 'continue-investigation' && event.owner === owners[0]), 'continuation was not authorized under the acquired owner');
|
||||
check(evidence.source.includes('return 2'), 'requested correction is missing');
|
||||
check(evidence.receipts.match(/^VERIFY$/gm)?.length === 1, 'verifier did not run exactly once');
|
||||
if (fixture.id === 'investigate-owned-ending-error') {
|
||||
check(evidence.receipts.includes('VERIFY_STATUS:69\n'), 'ending verifier error was not delivered');
|
||||
check(/error|unavailable|cannot|could not|unable/i.test(finalText), 'ending error not acknowledged');
|
||||
} else check(evidence.receipts.includes('VERIFY_STATUS:0\n'), 'verification did not succeed');
|
||||
}
|
||||
return failures;
|
||||
}
|
||||
|
||||
export async function runBoundaryActor(id: BoundaryCase, record: (entry: EvalTestEntry) => void, provider?: QueryProvider) {
|
||||
const started = Date.now();
|
||||
let fixture: ReturnType<typeof createBoundaryFixture> | undefined;
|
||||
let result: AgentSdkResult | undefined;
|
||||
let passed = false;
|
||||
let error: unknown;
|
||||
const attempts: Array<{ events: SDKMessage[]; evidence?: ReturnType<ReturnType<typeof createBoundaryFixture>['snapshot']> }> = [];
|
||||
try {
|
||||
fixture = createBoundaryFixture(id);
|
||||
const { runAgentSdkTest, resolveClaudeBinary } = await import('./agent-sdk-runner');
|
||||
const executable = resolveClaudeBinary();
|
||||
if (!provider && !executable) throw new Error('F9 actor requires the pinned native Claude CLI');
|
||||
const query = provider ?? (await import('@anthropic-ai/claude-agent-sdk')).query;
|
||||
result = await runAgentSdkTest({
|
||||
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
||||
userPrompt: fixture.prompt, workingDirectory: fixture.repo, env: fixture.env,
|
||||
pathToClaudeCodeExecutable: executable ?? undefined,
|
||||
allowedTools: ['Read', 'Bash', 'Edit', 'AskUserQuestion'], settingSources: [],
|
||||
canUseTool: (...args) => fixture!.canUseTool(...args), maxTurns: 12, signal: AbortSignal.timeout(CAPTURE_MS - 20000),
|
||||
testName: id,
|
||||
onRetry: () => {
|
||||
attempts.at(-1)!.evidence = fixture!.snapshot();
|
||||
const root = fixture!.root;
|
||||
fs.rmSync(root, { recursive: true, force: true });
|
||||
fixture = createBoundaryFixture(id, root);
|
||||
},
|
||||
queryProvider: options => {
|
||||
const attempt: typeof attempts[number] = { events: [] };
|
||||
attempts.push(attempt);
|
||||
const stream = query({ ...options, options: { ...options.options,
|
||||
allowedTools: [], hooks: { PreToolUse: [{ hooks: [(...args) => fixture!.preToolUse(...args)] }] },
|
||||
} });
|
||||
return new Proxy(stream, { get(target, property) {
|
||||
if (property === Symbol.asyncIterator) return async function* () {
|
||||
try { for await (const event of target) { attempt.events.push(event); yield event; } }
|
||||
finally { attempt.evidence = fixture!.snapshot(); }
|
||||
};
|
||||
const value = Reflect.get(target, property, target);
|
||||
return typeof value === 'function' ? value.bind(target) : value;
|
||||
} });
|
||||
},
|
||||
});
|
||||
const failures = boundaryFailures(fixture, result);
|
||||
if (failures.length) throw new Error(failures.join('; '));
|
||||
passed = true;
|
||||
} catch (failure) {
|
||||
error = failure;
|
||||
throw failure;
|
||||
} finally {
|
||||
try {
|
||||
record({ name: id, suite: 'workflow-boundaries', tier: 'e2e', passed,
|
||||
duration_ms: Date.now() - started, cost_usd: result?.costUsd ?? 0,
|
||||
transcript: result?.events, prompt: fixture?.prompt, turns_used: result?.turnsUsed, model: result?.model,
|
||||
output: JSON.stringify({ assistant: result?.output, error: error === undefined ? undefined : String(error), evidence: fixture?.snapshot(), attempts }),
|
||||
exit_reason: passed ? 'success' : result?.exitReason === 'success' ? 'assertion_failed' : result?.exitReason ?? 'runner_error' });
|
||||
} finally {
|
||||
if (fixture) fs.rmSync(fixture.root, { recursive: true, force: true });
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -13,7 +13,7 @@ import { buildEvalInputIdentity, lookupEvalInputCache, storeEvalInputCache,
|
||||
type Thresholds = { clarity: number; completeness: number; actionability: number };
|
||||
export interface WorkflowCacheOptions {
|
||||
root: string; testName: string; skillPath: string; startMarker: string; endMarker: string | null;
|
||||
judgeContext: string; judgeGoal: string; thresholds: Thresholds; prompt: string; attempt: number;
|
||||
judgeContext: string; judgeGoal: string; model?: string; thresholds: Thresholds; prompt: string; attempt: number;
|
||||
env?: NodeJS.ProcessEnv;
|
||||
}
|
||||
export interface WorkflowJudgeReuse {
|
||||
@@ -101,7 +101,7 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): {
|
||||
parameters: { rootPackage, thresholds: opts.thresholds, max_tokens: DEFAULT_JUDGE_MAX_TOKENS, temperature: null, budget_ms: JUDGE_MS,
|
||||
request: 'messages.create/user', retries: 1 },
|
||||
runtime: { image: env.EVALS_CACHE_RUNTIME_ID!, bun: Bun.version, node: process.versions.node,
|
||||
platform: process.platform, arch: process.arch, judge: resolveEvalModel('judge', undefined, env),
|
||||
platform: process.platform, arch: process.arch, judge: resolveEvalModel('judge', opts.model, env),
|
||||
anthropic_base_url: env.ANTHROPIC_BASE_URL ?? 'https://api.anthropic.com',
|
||||
anthropic_log: env.ANTHROPIC_LOG ?? null,
|
||||
proxies: Object.fromEntries(['HTTP_PROXY', 'HTTPS_PROXY', 'ALL_PROXY', 'NO_PROXY', 'http_proxy', 'https_proxy', 'all_proxy', 'no_proxy']
|
||||
|
||||
Reference in New Issue
Block a user