mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 09:55:29 +02:00
v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
This commit is contained in:
co-authored by
OpenAI Codex
parent
71f6048e8a
commit
9f81911136
@@ -20,6 +20,14 @@ import { runSkillTest, type SkillTestResult } from './session-runner';
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..', '..');
|
||||
|
||||
/**
|
||||
* Existing long section-loader work budget (v1.71): complete workflows can
|
||||
* load their sections quickly, then need 300–450s to generate the full report.
|
||||
* Keep 120s of the CAPTURE_LONG_MS outer budget for setup and reporting.
|
||||
* Ordinary captureSectionReads callers retain the 300s default below.
|
||||
*/
|
||||
export const LONG_SECTION_CAPTURE_MS = 480_000;
|
||||
|
||||
/** The 7 decision-brief format elements graded on the captured AUQ text. */
|
||||
export const AUQ_FORMAT_ELEMENTS: Array<{ field: string; re: RegExp }> = [
|
||||
{ field: 'ELI10:', re: /ELI10\s*:/i },
|
||||
@@ -217,6 +225,22 @@ This is a capture test, not an interactive session. Skip any system-audit / envi
|
||||
* resolves to, so Read/Grep/Glob/Write is all the agent needs (no Bash → it cannot
|
||||
* `find /` its way out, nor run git/gh mutations).
|
||||
*/
|
||||
export function hasDisabledOutsideReview(output: string): boolean {
|
||||
const headings = [...output.matchAll(/^## GSTACK REVIEW REPORT[ \t]*\r?$/gm)];
|
||||
const heading = headings.at(-1);
|
||||
if (!heading) return false;
|
||||
const section = output.slice(heading.index! + heading[0].length).split(/^##[ \t]+/m, 1)[0];
|
||||
const plain = (cell: string) => cell.replace(/[*_`]/g, '').trim().replace(/\s+/g, ' ').toLowerCase();
|
||||
for (const line of section.split('\n')) {
|
||||
if (!line.trimStart().startsWith('|')) continue;
|
||||
const cells = line.split('|').map(plain);
|
||||
if (cells[1] === 'outside review') {
|
||||
return /^disabled(?:$|\s|[(:—–-])/.test(cells[5] ?? '');
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
export async function captureSectionReads(opts: {
|
||||
planDir: string;
|
||||
skillName: string;
|
||||
@@ -230,9 +254,25 @@ export async function captureSectionReads(opts: {
|
||||
model?: string;
|
||||
maxTurns?: number;
|
||||
timeout?: number;
|
||||
/** Measure native section loading with the documented extra-review opt-out. */
|
||||
nativeReviewOnly?: boolean;
|
||||
}): Promise<{ readSections: Set<string>; reportProduced: boolean; toolCalls: SkillTestResult['toolCalls']; output: string }> {
|
||||
const outFile = path.join(opts.planDir, opts.reportFile ?? 'REPORT.md');
|
||||
const skillPath = path.join(opts.planDir, opts.skillName, 'SKILL.md');
|
||||
// Outside-review dispatch has separate behavioral coverage. Native-only
|
||||
// captures use the real supported control in state owned by this call;
|
||||
// never mutate the operator's or another capture's gstack configuration.
|
||||
// Keep the model-facing config path relative to the fixture's working directory.
|
||||
const stateDir = opts.nativeReviewOnly
|
||||
? fs.mkdtempSync(path.join(path.resolve(opts.planDir), '.gstack-section-state-')) : null;
|
||||
const nativeReviewRule = stateDir
|
||||
? `\n- Read ${path.relative(path.resolve(opts.planDir), path.join(stateDir, 'config.yaml'))}, the isolated gstack configuration for this capture. It sets codex_reviews: disabled. Follow that documented control: skip the entire extra outside-review step, including its native fallback, and report outside coverage as disabled. Complete all native review sections and the full required report.`
|
||||
: '';
|
||||
// Preserve full method execution while avoiding a second written walkthrough
|
||||
// of decisions already represented in the amended plan and required outputs.
|
||||
const planReviewWritingRule = opts.skillName === 'plan-ceo-review'
|
||||
? `\n- Write a concise, complete decision record: preserve original requirements and accepted plan amendments. Record each finding once with concrete evidence, the selected remedy, residual risks, and verification. Give all 11 sections an explicit outcome (including no issues or justified skips); retain the complete required registries, applicable diagrams, tasks, completion summary, and exact GSTACK REVIEW REPORT table. Cross-reference those records instead of repeating findings, option deliberations, diagrams, or registries in each section. Use compact outcome entries and short table cells; execute the review checklists without copying their questions or narrating every check into the artifact. Brevity must preserve every finding, accepted requirement, required field, and required diagram in its specified format. Do not expand the artifact into full implementation or test code unless that code is needed to specify an accepted plan change. This is a writing rule only: execute the full review, perform every required lazy-file Read, and complete all required artifacts before returning.`
|
||||
: '';
|
||||
const prompt = `You are running an automated skill-execution test. No human is present, so AskUserQuestion is unavailable. The ONLY skill file you may read is this absolute path: ${skillPath}. Do NOT Glob/find/search for any other SKILL.md anywhere — especially nothing under ~/.claude or /Users.
|
||||
|
||||
Read ${skillPath} and EXECUTE its workflow for this scenario:
|
||||
@@ -244,18 +284,28 @@ Rules for this run:
|
||||
- At any decision point that would call AskUserQuestion, silently pick the skill's recommended option and continue. Do NOT stop to ask.
|
||||
- This skill's body has been carved into on-demand sections/. When the skill gives a STOP-Read directive (for example "Read \`.../sections/<file>\` and execute it in full"), you MUST actually Read that sections/ file with the Read tool BEFORE doing the work it covers. Do not work from memory.
|
||||
- Do NOT run git, gh, commit, push, or any mutating command.
|
||||
- When the workflow is complete, write the skill's final output (the full review report / ship plan, including any required report table) to ${outFile}.`;
|
||||
- When the workflow is complete, write the skill's final output (the full review report / ship plan, including any required report table) to ${outFile}.${nativeReviewRule}${planReviewWritingRule}
|
||||
- After all required writes are complete, return a brief completion message and STOP. Do not reproduce the full report in the final response.`;
|
||||
|
||||
const result = await runSkillTest({
|
||||
prompt,
|
||||
workingDirectory: opts.planDir,
|
||||
allowedTools: ['Read', 'Grep', 'Glob', 'Write'],
|
||||
maxTurns: opts.maxTurns ?? 25,
|
||||
timeout: opts.timeout ?? 300_000,
|
||||
testName: opts.testName,
|
||||
runId: opts.runId,
|
||||
model: resolveEvalModel('capture', opts.model),
|
||||
});
|
||||
let result: SkillTestResult;
|
||||
try {
|
||||
if (stateDir) fs.writeFileSync(path.join(stateDir, 'config.yaml'), 'codex_reviews: disabled\n');
|
||||
result = await runSkillTest({
|
||||
prompt,
|
||||
workingDirectory: opts.planDir,
|
||||
allowedTools: ['Read', 'Grep', 'Glob', 'Write'],
|
||||
tools: ['Read', 'Grep', 'Glob', 'Write'],
|
||||
publicStreamDiagnostics: true,
|
||||
maxTurns: opts.maxTurns ?? 25,
|
||||
timeout: opts.timeout ?? 300_000,
|
||||
testName: opts.testName,
|
||||
runId: opts.runId,
|
||||
model: resolveEvalModel('capture', opts.model),
|
||||
...(stateDir ? { env: { GSTACK_HOME: stateDir, GSTACK_STATE_ROOT: stateDir } } : {}),
|
||||
});
|
||||
} finally {
|
||||
if (stateDir) fs.rmSync(stateDir, { recursive: true, force: true });
|
||||
}
|
||||
|
||||
const readSections = new Set<string>();
|
||||
for (const c of result.toolCalls) {
|
||||
@@ -267,7 +317,8 @@ Rules for this run:
|
||||
|
||||
let output = '';
|
||||
try { output = fs.readFileSync(outFile, 'utf-8'); } catch { output = result.output ?? ''; }
|
||||
const reportProduced = opts.reportMarker ? opts.reportMarker.test(output) : output.trim().length > 0;
|
||||
const reportProduced = result.exitReason === 'success'
|
||||
&& (opts.reportMarker ? opts.reportMarker.test(output) : output.trim().length > 0);
|
||||
|
||||
return { readSections, reportProduced, toolCalls: result.toolCalls, output };
|
||||
}
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
/** Bounded, body-free evidence from one owned PreToolUse Edit request. */
|
||||
import * as fs from 'node:fs';
|
||||
import { createHash } from 'node:crypto';
|
||||
const MAX_BYTES = 1024 * 1024, MAX_LINES = 512;
|
||||
// Covers ordinary 120-column crops, including combining scalars, without
|
||||
// storing any request text. Overflow is explicit and supplies no crop authority.
|
||||
const MAX_SUFFIX_SCALARS = 256, MAX_SUFFIX_HASHES = 8192;
|
||||
type ClippedAdditions = {version: 1; status: 'overflow'} | {version: 1; status: 'complete'; startLine: number;
|
||||
lines: Array<{line: number; lineHash: string; nextLineHash: string; suffixHashes: string[]}>};
|
||||
const suffixHash = (line: number, lineHash: string, nextLineHash: string, suffix: string) =>
|
||||
sha(JSON.stringify([line, lineHash, nextLineHash, suffix]));
|
||||
const sha = (bytes: string | Buffer) => createHash('sha256').update(bytes).digest('hex');
|
||||
export const autoplanEditLineHash = (line: string) => sha(line.replace(/\s/g, ''));
|
||||
export interface AutoplanEditDigest {
|
||||
version: 1;
|
||||
beforeSHA256: string;
|
||||
requestSHA256: string;
|
||||
oldLineHashes: string[];
|
||||
newLineHashes: string[];
|
||||
clippedAdditions?: ClippedAdditions;
|
||||
}
|
||||
export function validAutoplanEditDigest(value: unknown): value is AutoplanEditDigest {
|
||||
if (!value || typeof value !== 'object' || Array.isArray(value)) return false;
|
||||
const v = value as Record<string, unknown>;
|
||||
const hashes = (a: unknown): a is string[] => Array.isArray(a) && a.length > 0 && a.length <= MAX_LINES &&
|
||||
Array.from(a).every(h => typeof h === 'string' && /^[a-f0-9]{64}$/.test(h));
|
||||
const base = Object.keys(v).length === (v.clippedAdditions === undefined ? 5 : 6) && v.version === 1 && typeof v.beforeSHA256 === 'string' &&
|
||||
/^[a-f0-9]{64}$/.test(v.beforeSHA256) && typeof v.requestSHA256 === 'string' &&
|
||||
/^[a-f0-9]{64}$/.test(v.requestSHA256) && hashes(v.oldLineHashes) && hashes(v.newLineHashes) &&
|
||||
v.oldLineHashes.length + v.newLineHashes.length <= MAX_LINES;
|
||||
if (!base) return false;
|
||||
if (v.clippedAdditions === undefined) return true;
|
||||
const c = v.clippedAdditions as Record<string, any>;
|
||||
if (!c || typeof c !== 'object' || Array.isArray(c) || c.version !== 1) return false;
|
||||
if (c.status === 'overflow') return Object.keys(c).length === 2;
|
||||
if (c.status !== 'complete' || Object.keys(c).length !== 4 || !Number.isSafeInteger(c.startLine) || c.startLine < 1 ||
|
||||
!Array.isArray(c.lines) || !c.lines.length || c.lines.length > MAX_LINES) return false;
|
||||
let count = 0, previousLine = c.startLine - 1;
|
||||
return Array.from(c.lines).every((row: any) => {
|
||||
if (!row || typeof row !== 'object' || Array.isArray(row) || Object.keys(row).length !== 4 ||
|
||||
!Number.isSafeInteger(row.line) || row.line <= previousLine || row.line < c.startLine ||
|
||||
row.line >= c.startLine + (v.newLineHashes as string[]).length ||
|
||||
row.lineHash !== (v.newLineHashes as string[])[row.line - c.startLine] ||
|
||||
typeof row.nextLineHash !== 'string' || !/^[a-f0-9]{64}$/.test(row.nextLineHash) ||
|
||||
(row.line - c.startLine + 1 < (v.newLineHashes as string[]).length &&
|
||||
row.nextLineHash !== (v.newLineHashes as string[])[row.line - c.startLine + 1]) ||
|
||||
!hashes(row.suffixHashes) || row.suffixHashes.length > MAX_SUFFIX_SCALARS) return false;
|
||||
previousLine = row.line; count += row.suffixHashes.length;
|
||||
return count <= MAX_SUFFIX_HASHES;
|
||||
});
|
||||
}
|
||||
/** The caller must validate the owned path before this capped, no-follow read. */
|
||||
export function readAutoplanDigestFile(file: string): Buffer | undefined {
|
||||
let fd: number | undefined;
|
||||
try {
|
||||
fd = fs.openSync(file, fs.constants.O_RDONLY | (fs.constants.O_NOFOLLOW ?? 0));
|
||||
const initial = fs.fstatSync(fd);
|
||||
if (!initial.isFile() || initial.size > MAX_BYTES) return undefined;
|
||||
const buffer = Buffer.alloc(MAX_BYTES + 1); let length = 0;
|
||||
while (length < buffer.length) {
|
||||
const n = fs.readSync(fd, buffer, length, buffer.length - length, null);
|
||||
if (!n) break; length += n;
|
||||
}
|
||||
const final = fs.fstatSync(fd);
|
||||
if (length > MAX_BYTES || length !== initial.size || final.size !== initial.size ||
|
||||
final.mtimeMs !== initial.mtimeMs || final.ctimeMs !== initial.ctimeMs) return undefined;
|
||||
return buffer.subarray(0, length);
|
||||
} catch { return undefined; }
|
||||
finally { if (fd !== undefined) fs.closeSync(fd); }
|
||||
}
|
||||
export function createAutoplanEditDigest(file: string, removed: string, added: string, clipped = true): AutoplanEditDigest | undefined {
|
||||
const oldLines = removed.split('\n'), newLines = added.split('\n');
|
||||
if (!removed || removed === added || oldLines.length + newLines.length > MAX_LINES) return undefined;
|
||||
const bytes = readAutoplanDigestFile(file); if (!bytes) return undefined;
|
||||
const before = bytes.toString('utf8'), at = before.indexOf(removed);
|
||||
// Require one exact old substring. Only complete request lines can match a
|
||||
// displayed row; partial edge lines cannot establish added-line authority.
|
||||
if (!Buffer.from(before).equals(bytes) || at < 0 || at !== before.lastIndexOf(removed)) return undefined;
|
||||
const digest: AutoplanEditDigest = { version: 1, beforeSHA256: sha(bytes), requestSHA256: sha(JSON.stringify([removed, added])),
|
||||
oldLineHashes: oldLines.map(autoplanEditLineHash), newLineHashes: newLines.map(autoplanEditLineHash) };
|
||||
const end = at + removed.length;
|
||||
if (!clipped || (at > 0 && before[at - 1] !== '\n') || (end < before.length && before[end] !== '\n')) return digest;
|
||||
if (Buffer.byteLength(added) > MAX_BYTES) { digest.clippedAdditions = {version: 1, status: 'overflow'}; return digest; }
|
||||
const startLine = before.slice(0, at).split('\n').length;
|
||||
const afterLine = end < before.length ? before.slice(end + 1).split('\n', 1)[0] : undefined;
|
||||
const originals = new Set(before.split('\n').map(autoplanEditLineHash));
|
||||
const candidates = newLines.map((text, index) => ({index, scalars: Array.from(text.replace(/\s/g, ''))}))
|
||||
.filter(row => row.scalars.length && !originals.has(digest.newLineHashes[row.index]!) &&
|
||||
!digest.oldLineHashes.includes(digest.newLineHashes[row.index]!) && (row.index + 1 < newLines.length || afterLine !== undefined));
|
||||
const count = candidates.reduce((sum, row) => sum + Math.min(row.scalars.length, MAX_SUFFIX_SCALARS), 0);
|
||||
if (count > MAX_SUFFIX_HASHES) digest.clippedAdditions = {version: 1, status: 'overflow'};
|
||||
else if (count) digest.clippedAdditions = {version: 1, status: 'complete', startLine, lines: candidates.map(row => {
|
||||
const line = startLine + row.index, lineHash = digest.newLineHashes[row.index]!;
|
||||
const nextLineHash = digest.newLineHashes[row.index + 1] ?? autoplanEditLineHash(afterLine!);
|
||||
return {line, lineHash, nextLineHash, suffixHashes: Array.from({length: Math.min(row.scalars.length, MAX_SUFFIX_SCALARS)},
|
||||
(_, index) => suffixHash(line, lineHash, nextLineHash, row.scalars.slice(-index - 1).join('')))};
|
||||
})};
|
||||
return digest;
|
||||
}
|
||||
|
||||
/** Require complete numbered rows and the marker column of this exact diff. */
|
||||
export function matchesAutoplanDigestRows(rows: string[], before: Buffer, digest: AutoplanEditDigest,
|
||||
allowReplacementReset = false): boolean {
|
||||
if (!validAutoplanEditDigest(digest) || sha(before) !== digest.beforeSHA256) return false;
|
||||
const fullRow = /^( {0,3})([1-9]\d*) ([+ -])(.*)$/;
|
||||
const first = rows.findIndex(row => fullRow.test(row));
|
||||
const leading = first > 0 ? rows.slice(0, first) : [];
|
||||
const chunks: Array<{kind: string; text: string; line: number}> = [];
|
||||
let column: number | undefined, previousLine = 0;
|
||||
let removedStart: number | undefined, replacementReset = false;
|
||||
for (const row of leading.length ? rows.slice(first) : rows) {
|
||||
const full = fullRow.exec(row);
|
||||
if (full) {
|
||||
const line = Number(full[2]), markerColumn = full[1]!.length + full[2]!.length + 1;
|
||||
const kind = full[3]!, previousKind = chunks.at(-1)?.kind;
|
||||
// The complete native replacement panel numbers the old block first,
|
||||
// then restarts additions at that block's first line. Only the caller
|
||||
// that owns this full panel opts in; all row hashes still must match.
|
||||
const reset = allowReplacementReset && !replacementReset && previousKind === '-' && kind === '+' && line === removedStart;
|
||||
if (!Number.isSafeInteger(line) || (line < previousLine && !reset) ||
|
||||
(column !== undefined && column !== markerColumn)) return false;
|
||||
if (allowReplacementReset) {
|
||||
if ((kind === '-' || kind === '+') && kind === previousKind && line !== previousLine + 1) return false;
|
||||
if (kind === '-' && previousKind !== '-') {
|
||||
if (removedStart !== undefined || chunks.some(c => c.kind === '+')) return false;
|
||||
removedStart = line;
|
||||
}
|
||||
if (previousKind === '-' && kind !== '-' && !reset) return false;
|
||||
if (kind === '+' && previousKind !== '+' && removedStart !== undefined && !reset) return false;
|
||||
if (reset) replacementReset = true;
|
||||
}
|
||||
column = markerColumn; previousLine = line;
|
||||
chunks.push({kind, text: full[4]!, line});
|
||||
} else {
|
||||
if (column === undefined || !row.startsWith(' '.repeat(column))) return false;
|
||||
const last = chunks.at(-1), kind = row[column];
|
||||
if (!last || kind !== last.kind) return false;
|
||||
last.text += row.slice(column + 1);
|
||||
}
|
||||
}
|
||||
const originals = new Set(before.toString('utf8').split('\n').map(autoplanEditLineHash));
|
||||
if (allowReplacementReset) {
|
||||
const oldRows = chunks.filter(c => c.kind !== '+').map(c => autoplanEditLineHash(c.text));
|
||||
const newRows = chunks.filter(c => c.kind !== '-').map(c => autoplanEditLineHash(c.text));
|
||||
const starts = (rows: string[], hashes: string[]) => rows.flatMap((_, index) =>
|
||||
hashes.every((hash, offset) => rows[index + offset] === hash) ? [index] : []);
|
||||
const oldStarts = starts(oldRows, digest.oldLineHashes), newStarts = starts(newRows, digest.newLineHashes);
|
||||
if (oldStarts.length !== 1 || newStarts.length !== 1 || oldStarts[0] !== newStarts[0]) return false;
|
||||
const originalLines = before.toString('utf8').split('\n').map(autoplanEditLineHash);
|
||||
const delta = digest.newLineHashes.length - digest.oldLineHashes.length;
|
||||
let oldLine = chunks[0]!.line, newLine = oldLine, added = false;
|
||||
for (const row of chunks) {
|
||||
if (row.kind !== '+') {
|
||||
if (row.line !== oldLine + (added ? delta : 0) || autoplanEditLineHash(row.text) !== originalLines[oldLine - 1]) return false;
|
||||
oldLine++;
|
||||
}
|
||||
if (row.kind !== '-' && row.line !== newLine++) return false;
|
||||
if (row.kind === '+') added = true;
|
||||
}
|
||||
}
|
||||
let authenticatedClip = false;
|
||||
if (leading.length) {
|
||||
const c = digest.clippedAdditions, next = chunks[0];
|
||||
if (!c || c.status !== 'complete' || column === undefined || !next || chunks.length < 2 ||
|
||||
leading.length > MAX_SUFFIX_SCALARS || leading.some(row =>
|
||||
!row.startsWith(' '.repeat(column)) || row[column] !== '+')) return false;
|
||||
const fragment = leading.map(row => row.slice(column + 1)).join('').replace(/\s/g, '');
|
||||
const length = Array.from(fragment).length, record = c.lines.find(row => row.line === next.line - 1);
|
||||
if (!record || !length || length > MAX_SUFFIX_SCALARS || record.nextLineHash !== autoplanEditLineHash(next.text) ||
|
||||
originals.has(record.lineHash) || digest.oldLineHashes.includes(record.lineHash) ||
|
||||
record.suffixHashes[length - 1] !== suffixHash(record.line, record.lineHash, record.nextLineHash, fragment)) return false;
|
||||
const beforeLines = before.toString('utf8').split('\n');
|
||||
if (chunks.some((row, index) => {
|
||||
if (row.line !== next.line + index || row.kind === '-') return true;
|
||||
const relative = row.line - c.startLine, hash = autoplanEditLineHash(row.text);
|
||||
if (relative < digest.newLineHashes.length) return digest.newLineHashes[relative] !== hash;
|
||||
const originalIndex = row.line - (digest.newLineHashes.length - digest.oldLineHashes.length) - 1;
|
||||
return row.kind !== ' ' || originalIndex < 0 || originalIndex >= beforeLines.length ||
|
||||
autoplanEditLineHash(beforeLines[originalIndex]!) !== hash;
|
||||
})) return false;
|
||||
authenticatedClip = true;
|
||||
}
|
||||
return chunks.length >= 2 && (authenticatedClip || chunks.some(c => c.kind === '+' && /\S/.test(c.text) &&
|
||||
!originals.has(autoplanEditLineHash(c.text)) && !digest.oldLineHashes.includes(autoplanEditLineHash(c.text)))) &&
|
||||
chunks.every(c => c.kind === '+' ? digest.newLineHashes.includes(autoplanEditLineHash(c.text)) :
|
||||
originals.has(autoplanEditLineHash(c.text)) && (c.kind !== '-' || digest.oldLineHashes.includes(autoplanEditLineHash(c.text))));
|
||||
}
|
||||
@@ -0,0 +1,594 @@
|
||||
/** One-time input for a cropped native Edit of an already-owned review artifact. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { createHash } from 'node:crypto';
|
||||
import { validAutoplanEditDigest, readAutoplanDigestFile, matchesAutoplanDigestRows, createAutoplanEditDigest, autoplanEditLineHash } from './autoplan-artifact-digest';
|
||||
import type { PendingAutoplanArtifact } from './autoplan-artifact-recorder';
|
||||
import type { NativePublicToolEvent } from './plan-count-transcript';
|
||||
|
||||
interface ArtifactPermissionContext {
|
||||
cwd: string;
|
||||
/** Set only by the launcher that created HOME/.gstack, never from ambient env. */
|
||||
ownedStateRoot?: string;
|
||||
/** Optional native plans root supplied by the same isolated launcher. */
|
||||
ownedNativePlansRoot?: string;
|
||||
commandStartedAt: number;
|
||||
now?: number;
|
||||
transcriptStatus: string;
|
||||
publicTools: NativePublicToolEvent[];
|
||||
}
|
||||
|
||||
const MAX_BYTES = 1024 * 1024;
|
||||
const compact = (text: string) => text.replace(/\s/g, '');
|
||||
|
||||
/** A completed write distinguishes a new same-looking file confirmation. */
|
||||
export function autoplanPermissionProgressKey(viewport: string, events: readonly NativePublicToolEvent[]): string | undefined {
|
||||
const file = /^ {0,3}Do you want to (?:create|overwrite|edit) ([^\n?]+)\? *$/m.exec(viewport)?.[1];
|
||||
if (!file || new Set(events.map(event => event.sessionId)).size !== 1) return;
|
||||
const menu = compact(viewport);
|
||||
for (let i = events.length - 1; i >= 0; i--) {
|
||||
const result = events[i]!;
|
||||
if (result.kind !== 'result' || result.isError !== false) continue;
|
||||
const uses = events.slice(0, i).filter(event => event.kind === 'use' &&
|
||||
event.sessionId === result.sessionId && event.toolUseId === result.toolUseId);
|
||||
if (uses.length !== 1) continue;
|
||||
const use = uses[0]!, target = use.input?.file_path;
|
||||
if (!['Write', 'Edit'].includes(use.name ?? '') || typeof target !== 'string' ||
|
||||
!path.isAbsolute(target) || path.basename(target) !== file ||
|
||||
!menu.includes(`alwaysallowaccessto${compact(path.dirname(target))}forthissession`) ||
|
||||
!Number.isFinite(Date.parse(use.timestamp)) || Date.parse(result.timestamp) < Date.parse(use.timestamp) ||
|
||||
!Number.isFinite(Date.parse(result.timestamp))) continue;
|
||||
return `${result.sessionId}:${result.toolUseId}`;
|
||||
}
|
||||
}
|
||||
|
||||
/** The native header may remain above the diff; both displayed paths must bind. */
|
||||
function ownedEditDiffRows(rows: string[], file: string, ownedStateRoot?: string): string[] | null {
|
||||
const header = rows.findIndex(row => /^[●⏺] Update\(/.test(row));
|
||||
if (header > 0 && rows.slice(0,header).some(row => row.trim())) {
|
||||
// A completed native tool's diff may remain above the active edit panel.
|
||||
// Only its indented diff output is ignored; competing panels or prose are
|
||||
// not evidence for the current request and cannot be used as a prefix.
|
||||
const prefix = rows.slice(0, header), fullRow = /^ {6}([1-9]\d*) ([+ -])/;
|
||||
const first = prefix.map(row => fullRow.exec(row)).find(Boolean);
|
||||
if (!first) return null;
|
||||
const markerColumn = 6 + first[1]!.length + 1;
|
||||
let kind: string | undefined;
|
||||
for (const row of prefix) {
|
||||
if (!row.trim()) continue;
|
||||
const full = fullRow.exec(row);
|
||||
if (full) {
|
||||
if (!Number.isSafeInteger(Number(full[1])) || 6 + full[1]!.length + 1 !== markerColumn) return null;
|
||||
kind = full[2];
|
||||
} else {
|
||||
const wrappedKind = row[markerColumn];
|
||||
if (!row.startsWith(' '.repeat(markerColumn)) || !['+', ' ', '-'].includes(wrappedKind ?? '') ||
|
||||
(kind !== undefined && wrappedKind !== kind)) return null;
|
||||
kind = wrappedKind;
|
||||
}
|
||||
}
|
||||
rows = rows.slice(header);
|
||||
}
|
||||
// A redraw can repeat the same native tool title above one current panel.
|
||||
// Those homogeneous titles supply no authority: the full panel below must
|
||||
// still bind its path, current request, content, and exact one-time menu.
|
||||
const repeated: string[] = [];
|
||||
let panelAt = 0;
|
||||
for (; panelAt < rows.length; panelAt++) {
|
||||
if (!rows[panelAt]!.trim()) continue;
|
||||
const title = /^[●⏺] Update\(([^\n]+)\)$/.exec(rows[panelAt]!);
|
||||
if (!title) break;
|
||||
repeated.push(title[1]!);
|
||||
}
|
||||
if (repeated.length > 1 && ownedStateRoot && /^[─╌]{8,}$/.test(rows[panelAt] ?? '') &&
|
||||
rows[panelAt + 1]?.trim() === 'Edit file') {
|
||||
const relative = path.relative(ownedStateRoot, file).split(path.sep).join('/');
|
||||
const alias = path.basename(ownedStateRoot) === '.gstack' ? `~/.gstack/${relative}` : undefined;
|
||||
if (repeated.some(title => title !== repeated[0]) || (repeated[0] !== file && repeated[0] !== alias)) return null;
|
||||
rows = rows.slice(panelAt);
|
||||
}
|
||||
if (rows.filter(row => /^[●⏺] Update\(/.test(row)).length > 1) return null;
|
||||
// A viewport can start at the native Edit panel after its tool title has
|
||||
// scrolled away. The remaining displayed path must still bind the complete
|
||||
// owned project/artifact path; the menu and current request are checked below.
|
||||
if (/^[─╌]{8,}$/.test(rows[0] ?? '') && rows[1]?.trim() === 'Edit file') {
|
||||
if (!ownedStateRoot || !/^[─╌]{8,}$/.test(rows[3] ?? '')) return null;
|
||||
const relative = path.relative(ownedStateRoot, file).split(path.sep).join('/');
|
||||
const alias = path.basename(ownedStateRoot) === '.gstack' ? `~/.gstack/${relative}` : undefined;
|
||||
const displayed = rows[2]?.trim() ?? '';
|
||||
if (displayed !== file && displayed !== alias) {
|
||||
const suffix = displayed.startsWith('…') ? displayed.slice(1).split(path.sep).join('/') : '';
|
||||
if ((suffix !== relative && !suffix.endsWith('/' + relative)) || !file.split(path.sep).join('/').endsWith(suffix)) return null;
|
||||
}
|
||||
return rows.slice(4);
|
||||
}
|
||||
const update = /^[●⏺] Update\(([^\n]+)\)$/.exec(rows[0] ?? '');
|
||||
if (!update) return rows; // Existing cropped-only row guards still apply.
|
||||
if (!ownedStateRoot || rows[1]?.trim() !== '' || !/^[─╌]{8,}$/.test(rows[2] ?? '') ||
|
||||
rows[3]?.trim() !== 'Edit file' || !/^[─╌]{8,}$/.test(rows[5] ?? '')) return null;
|
||||
const relative = path.relative(ownedStateRoot,file).split(path.sep).join('/');
|
||||
const alias = path.basename(ownedStateRoot) === '.gstack' ? `~/.gstack/${relative}` : undefined;
|
||||
if (update[1] !== file && update[1] !== alias) return null;
|
||||
const displayed = rows[4]?.trim() ?? '';
|
||||
if (displayed !== file && displayed !== alias) {
|
||||
const suffix = displayed.startsWith('…') ? displayed.slice(1).split(path.sep).join('/') : '';
|
||||
// A truncated prefix must still retain the complete owned project/artifact
|
||||
// path. A basename or sibling-project suffix cannot bind this request.
|
||||
if ((suffix !== relative && !suffix.endsWith('/'+relative)) || !file.split(path.sep).join('/').endsWith(suffix)) return null;
|
||||
}
|
||||
return rows.slice(6);
|
||||
}
|
||||
|
||||
export function ownedAutoplanArtifact(file: string, context: Pick<ArtifactPermissionContext, 'cwd' | 'ownedStateRoot'>): boolean {
|
||||
if (!context.ownedStateRoot || !path.isAbsolute(file) || path.resolve(file) !== file) return false;
|
||||
const project = path.join(context.ownedStateRoot, 'projects', path.basename(context.cwd));
|
||||
const relative = path.relative(project, file).split(path.sep).join('/');
|
||||
// Current CEO plan and both explicit Eng test-plan layouts. Design/DX amend
|
||||
// ACTIVE_PLAN; they have no separate state-root plan directory. Do not admit
|
||||
// restore points, methodology snapshots, task logs, config, or mockups.
|
||||
if (!/^ceo-plans\/\d{4}-\d{2}-\d{2}-[a-z0-9][a-z0-9-]*\.md$/.test(relative) &&
|
||||
!/^[a-zA-Z0-9][a-zA-Z0-9_.-]*-test-plan-\d{8}-\d{6}\.md$/.test(relative)) return false;
|
||||
try {
|
||||
// System temp parents may be aliases (/var -> /private/var on macOS).
|
||||
// Canonicalize above the owned root; no symlink at or below it is admitted.
|
||||
const expected = path.join(fs.realpathSync(context.ownedStateRoot), path.relative(context.ownedStateRoot, file));
|
||||
return fs.lstatSync(context.ownedStateRoot).isDirectory() &&
|
||||
fs.realpathSync(file) === expected && fs.lstatSync(file).isFile() && fs.statSync(file).size <= MAX_BYTES;
|
||||
} catch { return false; }
|
||||
}
|
||||
|
||||
/** Require the whole current cropped diff, exact menu, and requested edit text. */
|
||||
function matchesCroppedEdit(viewport: string, file: string, before: string, removed: string, after: string,
|
||||
ownedStateRoot?: string): boolean {
|
||||
if (viewport.length > MAX_BYTES) return false;
|
||||
const text = viewport.replace(/\r\n?/g, '\n');
|
||||
const menu = /^ {0,3}Do you want to make this edit to ([^\n?]+)\? *\n {0,3}❯ *1\. Yes *\n {0,3}2\. Yes, and switch to accept edits \(auto-approve file edits and common file commands\) for this session(?: \(shift\+tab\))? *\n {0,3}3\. No *\n\s*Esc to cancel [·•] Tab to amend\s*$/m.exec(text);
|
||||
if (!menu || menu.index + menu[0].length !== text.length || menu[1] !== path.basename(file)) return false;
|
||||
const rows = text.slice(0, menu.index).trimEnd().split('\n');
|
||||
if (!/^[╌─]{8,}$/.test(rows.pop() ?? '')) return false;
|
||||
const diffRows = ownedEditDiffRows(rows,file,ownedStateRoot);
|
||||
if (!diffRows) return false;
|
||||
const chunks: Array<{ kind: string; text: string }> = [];
|
||||
let markerColumn: number | undefined;
|
||||
for (const row of diffRows) {
|
||||
const numbered = /^( {0,3})([1-9]\d*) ([+ -])(.*)$/.exec(row);
|
||||
if (numbered) {
|
||||
const line = Number(numbered[2]), column = numbered[1]!.length + numbered[2]!.length + 1;
|
||||
// Match the digest parser: every row owns one marker column, regardless
|
||||
// of line-number width; continuations retain that column and diff kind.
|
||||
if (!Number.isSafeInteger(line) || (markerColumn !== undefined && markerColumn !== column)) return false;
|
||||
markerColumn = column;
|
||||
chunks.push({ kind: numbered[3]!, text: numbered[4]! });
|
||||
} else {
|
||||
const last = chunks.at(-1);
|
||||
if (markerColumn === undefined || !row.startsWith(' '.repeat(markerColumn)) ||
|
||||
!last || row[markerColumn] !== last.kind) return false;
|
||||
last.text += row.slice(markerColumn + 1);
|
||||
}
|
||||
}
|
||||
// The crop itself cannot contain an example introduction, quote, unrelated
|
||||
// prompt, or arbitrary diff: each row must occur in this exact pending edit.
|
||||
const originals = before.split('\n').map(compact);
|
||||
const deletions = removed.split('\n').map(compact);
|
||||
const replacements = after.split('\n').map(compact);
|
||||
const changed = chunks.some(chunk => chunk.kind !== ' ' && compact(chunk.text));
|
||||
const matches = (oldRows: string[], newRows: string[]) => chunks.every(chunk =>
|
||||
(chunk.kind === '+' ? newRows : chunk.kind === '-' ? oldRows : originals).includes(compact(chunk.text)));
|
||||
if (changed && matches(deletions, replacements)) return true;
|
||||
// Edit arguments can start or end inside a line while the native preview
|
||||
// displays the whole line. Reconstruct only those unchanged edge bytes
|
||||
// from the unique current old substring; no viewport text supplies them.
|
||||
const at = before.indexOf(removed), end = at + removed.length;
|
||||
if (!changed || !removed || at < 0 || at !== before.lastIndexOf(removed) || after.length > MAX_BYTES) return false;
|
||||
const prefix = before.slice(before.slice(0, at).lastIndexOf('\n') + 1, at);
|
||||
const newline = before.indexOf('\n', end);
|
||||
const suffix = before.slice(end, newline < 0 ? before.length : newline);
|
||||
if (!prefix && !suffix) return false;
|
||||
const oldRows = (prefix + removed + suffix).split('\n').map(compact);
|
||||
const newRows = (prefix + after + suffix).split('\n').map(compact);
|
||||
return matches(oldRows, newRows) && chunks.some(chunk =>
|
||||
chunk.kind === '+' ? !oldRows.includes(compact(chunk.text)) :
|
||||
chunk.kind === '-' && !newRows.includes(compact(chunk.text)));
|
||||
}
|
||||
|
||||
export function autoplanArtifactPermissionInput(
|
||||
viewport: string, context: ArtifactPermissionContext, seen: ReadonlySet<string>,
|
||||
): { input: '1\r'; signature: string; file: string } | null {
|
||||
const now = context.now ?? Date.now();
|
||||
if (context.transcriptStatus !== 'ready' || !Number.isFinite(context.commandStartedAt) ||
|
||||
context.commandStartedAt > now || context.publicTools.length > 10_000 ||
|
||||
context.publicTools.some(event => !Number.isFinite(Date.parse(event.timestamp)))) return null;
|
||||
const events = context.publicTools.filter(event => Date.parse(event.timestamp) >= context.commandStartedAt);
|
||||
if (!events.length || events.some(event => !event.sessionId || !event.toolUseId ||
|
||||
!Number.isFinite(Date.parse(event.timestamp)) || Date.parse(event.timestamp) > now) ||
|
||||
new Set(events.map(event => event.sessionId)).size !== 1) return null;
|
||||
// Bind the latest file mutation, which must be the sole unresolved Write/Edit.
|
||||
// Claude may publish a queued Bash while its current Edit permission is open;
|
||||
// that unrelated request supplies no file permission authority.
|
||||
const edit = events.filter(event => event.kind === 'use' &&
|
||||
(event.name === 'Write' || event.name === 'Edit')).at(-1);
|
||||
if (!edit || edit.kind !== 'use' || edit.name !== 'Edit' || typeof edit.input?.file_path !== 'string' ||
|
||||
typeof edit.input.old_string !== 'string' || !edit.input.old_string ||
|
||||
typeof edit.input.new_string !== 'string' || edit.input.new_string === edit.input.old_string ||
|
||||
(edit.input.replace_all !== undefined && edit.input.replace_all !== false)) return null;
|
||||
const signature = `${edit.sessionId}:${edit.toolUseId}`;
|
||||
if (seen.has(signature) || !ownedAutoplanArtifact(edit.input.file_path, context)) return null;
|
||||
const uses = new Map<string, NativePublicToolEvent>();
|
||||
const results = new Map<string, NativePublicToolEvent>();
|
||||
let previousTime = context.commandStartedAt;
|
||||
for (const event of events) {
|
||||
const time = Date.parse(event.timestamp);
|
||||
if (time < previousTime) return null;
|
||||
previousTime = time;
|
||||
const map = event.kind === 'use' ? uses : results;
|
||||
if (map.has(event.toolUseId)) return null;
|
||||
map.set(event.toolUseId, event);
|
||||
if (event.kind === 'result' && !uses.has(event.toolUseId)) return null;
|
||||
}
|
||||
const writes = [...uses.values()].filter(event => event.name === 'Write' || event.name === 'Edit');
|
||||
if (results.has(edit.toolUseId) || writes.filter(event => !results.has(event.toolUseId)).length !== 1) return null;
|
||||
if (!writes.some(event => event.toolUseId !== edit.toolUseId &&
|
||||
event.input?.file_path === edit.input!.file_path && results.has(event.toolUseId) &&
|
||||
results.get(event.toolUseId)!.isError === false)) return null;
|
||||
try {
|
||||
const before = fs.readFileSync(edit.input.file_path, 'utf8');
|
||||
if (!before.includes(edit.input.old_string) ||
|
||||
!matchesCroppedEdit(viewport, edit.input.file_path, before, edit.input.old_string, edit.input.new_string,
|
||||
context.ownedStateRoot)) return null;
|
||||
return { input: '1\r', signature, file: edit.input.file_path };
|
||||
} catch { return null; }
|
||||
}
|
||||
|
||||
/** A previously granted viewport cannot establish a newer unpublished request. */
|
||||
export const autoplanArtifactMenuKey = (viewport: string) =>
|
||||
`menu:${createHash('sha256').update(viewport.replace(/\r\n?/g, '\n')).digest('hex')}`;
|
||||
|
||||
/** Queued native-plan edits remain unstarted; this never grants their input. */
|
||||
function ownedQueuedNativePlan(file: unknown, root: string | undefined, pendingTime: number): file is string {
|
||||
if (!root || typeof file !== 'string' || !path.isAbsolute(root) || path.resolve(root) !== root ||
|
||||
!path.isAbsolute(file) || path.resolve(file) !== file || path.dirname(file) !== root ||
|
||||
!/^[a-z][a-z0-9-]*\.md$/.test(path.basename(file))) return false;
|
||||
try {
|
||||
return fs.lstatSync(root).isDirectory() && fs.lstatSync(file).isFile() &&
|
||||
fs.realpathSync(file) === path.join(fs.realpathSync(root), path.basename(file)) && fs.statSync(file).size <= MAX_BYTES &&
|
||||
Math.floor(fs.statSync(file).mtimeMs) <= pendingTime;
|
||||
} catch { return false; }
|
||||
}
|
||||
|
||||
/** Bind completed snapshot output and queued native-plan redraw labels before
|
||||
* the full current panel. Arbitrary prose, examples and competing panels stay. */
|
||||
function queuedPlanViewport(viewport: string, queuedPlans: number, current: NativePublicToolEvent,
|
||||
events: NativePublicToolEvent[]): string {
|
||||
if (!queuedPlans) return viewport;
|
||||
const lines = viewport.replace(/\r\n?/g, '\n').split('\n');
|
||||
const title = lines.findIndex(line => /^[●⏺] Update\(/.test(line));
|
||||
if (title < 0 || lines.filter(line => /^[●⏺] Update\(/.test(line)).length !== 1) return viewport;
|
||||
const panel = lines.findIndex((line, i) => i > title && /^[─╌]{8,}$/.test(line));
|
||||
const redraws = lines.slice(title + 1, panel).filter(line => line.trim());
|
||||
if (panel < 0 || redraws.length !== queuedPlans || redraws.some(line => !/^[●⏺] Updated plan$/.test(line))) return viewport;
|
||||
const prefix = lines.slice(0, title);
|
||||
while (prefix.at(-1)?.trim() === '') prefix.pop();
|
||||
if (prefix.some(line => line.trim())) {
|
||||
const prior = events.slice(0, events.indexOf(current));
|
||||
const result = prior.filter(event => event.kind === 'result').at(-1);
|
||||
const use = result && prior.find(event => event.kind === 'use' && event.toolUseId === result.toolUseId);
|
||||
const command = /^ {8}"([^"\n]+)…\)$/.exec(prefix[0] ?? '');
|
||||
const outputEnd = prefix.length - 2;
|
||||
if (!command || !use || use.name !== 'Bash' || use.messageId !== current.messageId ||
|
||||
use.requestId !== current.requestId || typeof use.input?.command !== 'string' ||
|
||||
!use.input.command.includes(command[1]!) || result?.isError !== false || typeof result.content !== 'string' ||
|
||||
!/^ {2}⎿[ \u00a0]+\{$/.test(prefix[1] ?? '') ||
|
||||
!/^ {5}… \+[1-9]\d* lines \(ctrl\+o to expand\)$/.test(prefix[outputEnd] ?? '') ||
|
||||
!/^ {2}⎿[ \u00a0]+Allowed by auto mode classifier$/.test(prefix.at(-1) ?? '') || outputEnd < 3) return viewport;
|
||||
const displayed = ['{', ...prefix.slice(2, outputEnd).map(line => /^ {5}( {2}\S.*)$/.exec(line)?.[1])];
|
||||
if (displayed.some(line => line === undefined) ||
|
||||
result.content.split('\n').slice(0, displayed.length).join('\n') !== displayed.join('\n')) return viewport;
|
||||
}
|
||||
return [lines[title], '', ...lines.slice(panel)].join('\n');
|
||||
}
|
||||
|
||||
/** Native batch redraws are display-only: bind their titles, waiting command,
|
||||
* and one clipped context row to public events before removing the prefix. */
|
||||
function queuedArtifactViewport(viewport: string, current: NativePublicToolEvent,
|
||||
events: NativePublicToolEvent[], queued: ReadonlySet<string>, file: string,
|
||||
pendingTime: number, ownedStateRoot?: string): string {
|
||||
if (!queued.size || !ownedStateRoot) return viewport;
|
||||
const successors = events.filter(e => e.kind === 'use' && queued.has(e.toolUseId));
|
||||
if (successors.some(e => e.input?.file_path !== file)) return viewport;
|
||||
const lines = viewport.replace(/\r\n?/g, '\n').split('\n');
|
||||
const firstTitle = lines.findIndex(line => /^[●⏺] Update\(/.test(line));
|
||||
const panel = lines.findIndex((line, i) => i > firstTitle && /^[─╌]{8,}$/.test(line));
|
||||
if (firstTitle < 1 || panel < 0) return viewport;
|
||||
const prefix = lines.slice(0, firstTitle).filter(line => line.trim());
|
||||
const clipped = prefix.length === 1 && /^ {10}(\S.{15,})$/.exec(prefix[0]!);
|
||||
if (!clipped) return viewport;
|
||||
const relative = path.relative(ownedStateRoot, file).split(path.sep).join('/');
|
||||
const alias = path.basename(ownedStateRoot) === '.gstack' ? `~/.gstack/${relative}` : undefined;
|
||||
const rows = lines.slice(firstTitle, panel).filter(line => line.trim());
|
||||
const titles = rows.slice(0, successors.length + 1);
|
||||
if (titles.length !== successors.length + 1 || titles.some(row => {
|
||||
const title = /^[●⏺] Update\(([^\n]+)\)$/.exec(row);
|
||||
return !title || (title[1] !== file && title[1] !== alias);
|
||||
})) return viewport;
|
||||
const bash = rows.slice(titles.length);
|
||||
if (bash.length < 2 || !/^ {2}⎿[ \u00a0]+Waiting…$/.test(bash.at(-1)!)) return viewport;
|
||||
const parts = bash.slice(0, -1).map((row, i) =>
|
||||
(i === 0 ? /^[●⏺] Bash\((.+)$/ : /^ {6}(.+)$/).exec(row)?.[1]);
|
||||
if (parts.some(part => part === undefined)) return viewport;
|
||||
const rendered = parts.join('');
|
||||
if (!rendered.endsWith('…)')) return viewport;
|
||||
const commandPrefix = compact(rendered.slice(0, -2));
|
||||
const waiting = events.filter(e => e.kind === 'use' && e.name === 'Bash' &&
|
||||
e.messageId === current.messageId && e.requestId === current.requestId &&
|
||||
events.indexOf(e) > Math.max(...successors.map(s => events.indexOf(s))) &&
|
||||
!events.some(result => result.kind === 'result' && result.toolUseId === e.toolUseId) &&
|
||||
typeof e.input?.command === 'string' && compact(e.input.command).startsWith(commandPrefix));
|
||||
if (commandPrefix.length < 32 || waiting.length !== 1) return viewport;
|
||||
const completed = events.filter(e => e.kind === 'result' && e.isError === false &&
|
||||
Date.parse(e.timestamp) <= pendingTime).map(result => ({result, use:events.find(e =>
|
||||
e.kind === 'use' && e.toolUseId === result.toolUseId)})).filter(({use}) =>
|
||||
use?.name === 'Edit' && use.input?.file_path === file &&
|
||||
use.messageId === current.messageId && use.requestId === current.requestId).at(-1);
|
||||
const replacement = completed?.use?.input?.new_string;
|
||||
if (typeof replacement !== 'string' || !replacement) return viewport;
|
||||
const before = fs.readFileSync(file, 'utf8'), at = before.indexOf(replacement);
|
||||
if (at < 0 || before.indexOf(replacement, at + 1) !== -1) return viewport;
|
||||
// Native diffs display at most three unchanged context lines after an edit.
|
||||
// The cropped row must be a suffix of one of those current, unchanged lines.
|
||||
const lineEnd = before.indexOf('\n', at + replacement.length);
|
||||
const context = lineEnd < 0 ? [] : before.slice(lineEnd + 1).split('\n').slice(0, 3);
|
||||
if (!context.some(line => compact(line).endsWith(compact(clipped[1]!)))) return viewport;
|
||||
return lines.slice(panel).join('\n');
|
||||
}
|
||||
|
||||
/** A cropped command caption is display only. Bind the complete wrapped command
|
||||
* to one unstarted successor in the current published batch before discarding it. */
|
||||
function queuedCommandViewport(viewport: string, current: NativePublicToolEvent,
|
||||
events: NativePublicToolEvent[], queued: ReadonlySet<string>, hookSeenIds: readonly string[],
|
||||
viewportCapturedAt: number): string {
|
||||
if (!queued.size) return viewport;
|
||||
const text = viewport.replace(/\r\n?/g, '\n');
|
||||
const panels = [...text.matchAll(/^[─╌]{8,}\n {0,3}Edit file[ \t]*\n/gm)];
|
||||
if (panels.length !== 1 || panels[0]!.index === 0) return viewport;
|
||||
const rows = text.slice(0, panels[0]!.index).split('\n');
|
||||
while (rows.at(-1)?.trim() === '') rows.pop();
|
||||
const parts = rows.map((row, i) =>
|
||||
(i === 0 ? /^ {2}⎿[ \u00a0]+\$ (\S.*)$/ : /^ {5}(\S.*)$/).exec(row)?.[1]);
|
||||
if (!parts.length || parts.some(part => part === undefined)) return viewport;
|
||||
const currentIndex = events.indexOf(current);
|
||||
const waiting = events.filter(e => e.kind === 'use' && e.name === 'Bash' &&
|
||||
e.messageId === current.messageId && e.requestId === current.requestId && events.indexOf(e) > currentIndex &&
|
||||
!events.some(result => result.kind === 'result' && result.toolUseId === e.toolUseId));
|
||||
const command = waiting[0], input = command?.input?.command;
|
||||
if (waiting.length !== 1 || !command || typeof input !== 'string' || !input || input.length > MAX_BYTES ||
|
||||
/[\x00-\x1f\x7f]/.test(input) || hookSeenIds.includes(command.toolUseId) ||
|
||||
Date.parse(command.timestamp) > viewportCapturedAt ||
|
||||
events.some(e => e.kind === 'use' && queued.has(e.toolUseId) && events.indexOf(e) >= events.indexOf(command))) return viewport;
|
||||
// Preserve every displayed character, including spaces inside quoted arguments.
|
||||
// Only whitespace omitted at a renderer soft-wrap boundary may be skipped.
|
||||
let remaining = input;
|
||||
for (let i = 0; i < parts.length; i++) {
|
||||
if (!remaining.startsWith(parts[i]!)) return viewport;
|
||||
remaining = remaining.slice(parts[i]!.length);
|
||||
if (i < parts.length - 1) remaining = remaining.replace(/^[ \t]+/, '');
|
||||
}
|
||||
return remaining === '' ? text.slice(panels[0]!.index) : viewport;
|
||||
}
|
||||
|
||||
/** A native command description can remain above an unpublished Edit panel.
|
||||
* Its text supplies no command identity, completion, or approval authority.
|
||||
* Only the digest-bound pending path may discard this one display prefix. */
|
||||
function pendingCommandDisplayViewport(viewport: string, file: string, ownedStateRoot?: string): string {
|
||||
const text = viewport.replace(/\r\n?/g, '\n');
|
||||
const panels = [...text.matchAll(/^[─╌]{8,}\n {0,3}Edit file[ \t]*\n/gm)];
|
||||
if (panels.length !== 1 || panels[0]!.index === 0) return viewport;
|
||||
const prefix = text.slice(0, panels[0]!.index).split('\n').filter(line => line.trim());
|
||||
// An unpublished batch can leave the current Update title, plan redraws,
|
||||
// and a queued Bash card above the panel. These cards grant no authority:
|
||||
// only the one current Edit's owned path and complete digest below do so.
|
||||
const update = /^[●⏺] Update\(([^\n]+)\)$/.exec(prefix[0] ?? '');
|
||||
if (update && ownedStateRoot) {
|
||||
const relative = path.relative(ownedStateRoot, file).split(path.sep).join('/');
|
||||
const alias = path.basename(ownedStateRoot) === '.gstack' ? `~/.gstack/${relative}` : undefined;
|
||||
const bash = prefix.findIndex(row => /^[●⏺] Bash\(/.test(row));
|
||||
const command = prefix.slice(bash, -1);
|
||||
if ((update[1] === file || update[1] === alias) && bash > 1 &&
|
||||
prefix.slice(1, bash).every(row => /^[●⏺] Updated plan$/.test(row)) &&
|
||||
/^ {2}⎿[ \u00a0]+Waiting…$/.test(prefix.at(-1) ?? '') && command.length > 0 &&
|
||||
command.every((row, i) => (i === 0 ? /^[●⏺] Bash\(\S.*$/ : /^ {6}\S.*$/).test(row)) &&
|
||||
command.at(-1)!.endsWith('…)') &&
|
||||
!command.slice(1).some(row => /^ {6}[●⏺❯☐□>]|^ {6}(?:`{3,}|~{3,})/.test(row)) &&
|
||||
!/(?:Do you want|Would you like|Bash command[^\n]*permission|requested permissions?|allow all edits|always allow access|Esc to cancel|Edit file)/i.test(command.join('\n'))) {
|
||||
return text.slice(panels[0]!.index);
|
||||
}
|
||||
return viewport;
|
||||
}
|
||||
const title = /^[●⏺] ([^\n]+)$/.exec(prefix[0] ?? '')?.[1];
|
||||
if (!title || /^(?:["'`“‘]|(?:source|example|quoted|history|historical|hypothetical|previous|earlier)\b)/i.test(title) ||
|
||||
!/^ {2}⎿[ \u00a0]+\$ \S.*$/.test(prefix[1] ?? '') ||
|
||||
prefix.slice(2).some(line => !/^ {5}\S.*$/.test(line)) ||
|
||||
prefix.slice(1).some(line => /^[ \t]*[●⏺❯☐□>]|^[ \t]*(?:`{3,}|~{3,})/.test(line)) ||
|
||||
/(?:Do you want|Would you like|Bash command[^\n]*permission|requested permissions?|allow all edits|always allow access|Esc to cancel|Edit file)/i.test(prefix.join('\n'))) return viewport;
|
||||
return text.slice(panels[0]!.index);
|
||||
}
|
||||
|
||||
/** Same native-history gate for an unpublished pending Edit, independent of its display. */
|
||||
export function hasPendingAutoplanArtifactHistory(
|
||||
context: ArtifactPermissionContext & { pending?: PendingAutoplanArtifact },
|
||||
): boolean {
|
||||
const p = context.pending, now = context.now ?? Date.now();
|
||||
if (!p || p.source !== 'pre_tool_use' || p.tool !== 'Edit' || context.transcriptStatus !== 'ready' ||
|
||||
!Number.isFinite(now) || !Number.isFinite(context.commandStartedAt) ||
|
||||
!/^[A-Za-z0-9_-]{1,160}$/.test(p.sessionId) || !/^[A-Za-z0-9_-]{1,160}$/.test(p.toolUseId) ||
|
||||
!ownedAutoplanArtifact(p.file, context) || context.publicTools.length > 10_000) return false;
|
||||
const pendingTime = Date.parse(p.timestamp);
|
||||
if (!Number.isFinite(pendingTime) || pendingTime < context.commandStartedAt || pendingTime > now) return false;
|
||||
const events = context.publicTools.filter(e => Date.parse(e.timestamp) >= context.commandStartedAt);
|
||||
if (!events.length || context.publicTools.some(e => !Number.isFinite(Date.parse(e.timestamp)))) return false;
|
||||
const uses = new Map<string, NativePublicToolEvent>(), results = new Map<string, NativePublicToolEvent>();
|
||||
let last = context.commandStartedAt;
|
||||
for (const event of events) {
|
||||
const time = Date.parse(event.timestamp);
|
||||
if (event.sessionId !== p.sessionId || !event.toolUseId || event.toolUseId === p.toolUseId || time < last || time > now) return false;
|
||||
last = time;
|
||||
const map = event.kind === 'use' ? uses : results;
|
||||
if (map.has(event.toolUseId) || (event.kind === 'result' && !uses.has(event.toolUseId))) return false;
|
||||
map.set(event.toolUseId, event);
|
||||
}
|
||||
const mutations = [...uses.values()].filter(e => e.name === 'Write' || e.name === 'Edit');
|
||||
// Hook metadata cannot replace a published request/result or an unresolved
|
||||
// mutation. Public successful same-file history remains mandatory.
|
||||
if (mutations.some(e => !results.has(e.toolUseId) || Date.parse(e.timestamp) > pendingTime ||
|
||||
Date.parse(results.get(e.toolUseId)!.timestamp) > pendingTime) ||
|
||||
!mutations.some(e => e.input?.file_path === p.file && results.get(e.toolUseId)?.isError === false &&
|
||||
Date.parse(results.get(e.toolUseId)!.timestamp) <= pendingTime)) return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Metadata-only fallback. Added rows are display evidence, never request content. */
|
||||
export function pendingAutoplanArtifactPermissionInput(viewport: string,
|
||||
context: ArtifactPermissionContext & { pending?: PendingAutoplanArtifact; viewportCapturedAt: number },
|
||||
seen: ReadonlySet<string>,
|
||||
): { input: '1\r'; signature: string; file: string } | null {
|
||||
const p = context.pending, now = context.now ?? Date.now();
|
||||
if (p?.editDigest !== undefined && !validAutoplanEditDigest(p.editDigest)) return null;
|
||||
if (!p || !Number.isFinite(now) || context.transcriptStatus !== 'ready' || !Number.isFinite(context.commandStartedAt) ||
|
||||
!Number.isFinite(context.viewportCapturedAt) || context.viewportCapturedAt > now ||
|
||||
context.commandStartedAt > context.viewportCapturedAt || viewport.length > MAX_BYTES ||
|
||||
p.source !== 'pre_tool_use' || p.tool !== 'Edit' || typeof p.file !== 'string' ||
|
||||
!/^[A-Za-z0-9_-]{1,160}$/.test(p.sessionId) || !/^[A-Za-z0-9_-]{1,160}$/.test(p.toolUseId) ||
|
||||
!ownedAutoplanArtifact(p.file, context)) return null;
|
||||
const pendingTime = Date.parse(p.timestamp), signature = `${p.sessionId}:${p.toolUseId}`;
|
||||
if (!Number.isFinite(pendingTime) || pendingTime < context.commandStartedAt || pendingTime > context.viewportCapturedAt ||
|
||||
seen.has(signature) || seen.has(autoplanArtifactMenuKey(viewport)) || context.publicTools.length > 10_000) return null;
|
||||
if (!hasPendingAutoplanArtifactHistory(context)) return null;
|
||||
try {
|
||||
const currentViewport = p.editDigest ? pendingCommandDisplayViewport(viewport, p.file, context.ownedStateRoot) : viewport;
|
||||
const text = currentViewport.replace(/\r\n?/g, '\n');
|
||||
const menu = /^ {0,3}Do you want to make this edit to ([^\n?]+)\? *\n {0,3}❯ *1\. Yes *\n {0,3}2\. Yes, and switch to accept edits \(auto-approve file edits and common file commands\) for this session(?: \(shift\+tab\))? *\n {0,3}3\. No *\n\s*Esc to cancel [·•] Tab to amend\s*$/m.exec(text);
|
||||
if (!menu || menu.index + menu[0].length !== text.length || menu[1] !== path.basename(p.file)) return null;
|
||||
const rows = text.slice(0, menu.index).trimEnd().split('\n');
|
||||
if (!/^[╌─]{8,}$/.test(rows.pop() ?? '')) return null;
|
||||
const diffRows = ownedEditDiffRows(rows,p.file,context.ownedStateRoot);
|
||||
if (!diffRows) return null;
|
||||
if (Math.floor(fs.statSync(p.file).mtimeMs) > pendingTime) return null;
|
||||
if (p.editDigest) {
|
||||
const before = readAutoplanDigestFile(p.file);
|
||||
if (!before || createHash('sha256').update(before).digest('hex') !== p.editDigest.beforeSHA256) return null;
|
||||
if (matchesAutoplanDigestRows(diffRows,before,p.editDigest,currentViewport !== viewport)) return {input:'1\r', signature, file:p.file};
|
||||
if (currentViewport !== viewport) return null; // The new prefix path requires the exact digest, including additions.
|
||||
// Legacy deletion crops below must still honor the recorded request digest.
|
||||
}
|
||||
const originals = fs.readFileSync(p.file, 'utf8').split('\n').map(compact);
|
||||
const chunks: Array<{kind:string; text:string; partial?:boolean}> = [];
|
||||
const fullRow = /^( {0,3})([1-9]\d*) ([+ -])(.*)$/;
|
||||
const first = diffRows.map(row => fullRow.exec(row)).find(Boolean);
|
||||
if (!first) return null;
|
||||
const markerColumn = first[1]!.length + first[2]!.length + 1;
|
||||
let numbered = 0;
|
||||
for (const row of diffRows) {
|
||||
const full = fullRow.exec(row);
|
||||
if (full) {
|
||||
if (!Number.isSafeInteger(Number(full[2])) || full[1]!.length + full[2]!.length + 1 !== markerColumn) return null;
|
||||
numbered++; chunks.push({kind:full[3]!, text:full[4]!});
|
||||
} else {
|
||||
const kind = row[markerColumn], text = row.slice(markerColumn + 1);
|
||||
if (!row.startsWith(' '.repeat(markerColumn)) || !['+', ' ', '-'].includes(kind ?? '')) return null;
|
||||
if (!chunks.length) chunks.push({kind:kind!, text, partial:true});
|
||||
else {
|
||||
const previous = chunks.at(-1)!;
|
||||
if (previous.kind !== kind) return null;
|
||||
previous.text += text;
|
||||
}
|
||||
}
|
||||
}
|
||||
// A leading cropped deletion/context fragment must be an actual suffix.
|
||||
// Complete removed/context rows must occur in the current owned file.
|
||||
if (numbered < 2 || !chunks.some(c => c.kind === '-' && compact(c.text)) ||
|
||||
chunks.some(c => c.kind !== '+' && !originals.some(line => c.partial
|
||||
? line.endsWith(compact(c.text)) : line === compact(c.text)))) return null;
|
||||
const digest = p.editDigest;
|
||||
if (digest && chunks.some(chunk => {
|
||||
const hash = autoplanEditLineHash(chunk.text);
|
||||
if (chunk.kind === '+') return chunk.partial || !digest.newLineHashes.includes(hash);
|
||||
if (chunk.kind !== '-') return false; // Current-file context was checked above.
|
||||
return chunk.partial
|
||||
? !originals.some(line => line.endsWith(compact(chunk.text)) && digest.oldLineHashes.includes(autoplanEditLineHash(line)))
|
||||
: !digest.oldLineHashes.includes(hash);
|
||||
})) return null;
|
||||
return {input:'1\r', signature, file:p.file};
|
||||
} catch { return null; }
|
||||
}
|
||||
|
||||
/** A native hook identifies the executing request within a published tool batch. */
|
||||
export function publishedAutoplanArtifactPermissionInput(viewport: string,
|
||||
context: ArtifactPermissionContext & { pending?: PendingAutoplanArtifact; viewportCapturedAt: number },
|
||||
seen: ReadonlySet<string>,
|
||||
): { input: '1\r'; signature: string; file: string } | null {
|
||||
const p=context.pending, now=context.now??Date.now();
|
||||
if (!p || p.source!=='pre_tool_use' || p.tool!=='Edit' || !validAutoplanEditDigest(p.editDigest) ||
|
||||
context.transcriptStatus!=='ready' || !Number.isFinite(now) || !Number.isFinite(context.commandStartedAt) ||
|
||||
!Number.isFinite(context.viewportCapturedAt) || context.commandStartedAt>context.viewportCapturedAt ||
|
||||
context.viewportCapturedAt>now || context.publicTools.length>10_000 || viewport.length>MAX_BYTES ||
|
||||
!/^[A-Za-z0-9_-]{1,160}$/.test(p.sessionId) || !/^[A-Za-z0-9_-]{1,160}$/.test(p.toolUseId) ||
|
||||
!Array.isArray(p.hookSeenIds) || !p.hookSeenIds.length || p.hookSeenIds.length>128 ||
|
||||
p.hookSeenIds.some(id=>typeof id!=='string'||!/^[A-Za-z0-9_-]{1,160}$/.test(id)) ||
|
||||
new Set(p.hookSeenIds).size!==p.hookSeenIds.length || !p.hookSeenIds.includes(p.toolUseId) ||
|
||||
seen.has(`${p.sessionId}:${p.toolUseId}`) || seen.has(autoplanArtifactMenuKey(viewport)) ||
|
||||
!ownedAutoplanArtifact(p.file,context)) return null;
|
||||
const pendingTime=Date.parse(p.timestamp);
|
||||
if (!Number.isFinite(pendingTime) || pendingTime<context.commandStartedAt || pendingTime>context.viewportCapturedAt ||
|
||||
context.publicTools.some(e=>!Number.isFinite(Date.parse(e.timestamp)))) return null;
|
||||
const events=context.publicTools.filter(e=>Date.parse(e.timestamp)>=context.commandStartedAt);
|
||||
const uses=new Map<string,NativePublicToolEvent>(), results=new Map<string,NativePublicToolEvent>();
|
||||
let last=context.commandStartedAt;
|
||||
for (const event of events) {
|
||||
const time=Date.parse(event.timestamp), map=event.kind==='use'?uses:results;
|
||||
if (event.sessionId!==p.sessionId || !event.toolUseId || time<last || time>now || map.has(event.toolUseId) ||
|
||||
(event.kind==='result'&&!uses.has(event.toolUseId))) return null;
|
||||
last=time;map.set(event.toolUseId,event);
|
||||
}
|
||||
const current=uses.get(p.toolUseId), input=current?.input;
|
||||
if (!current || current.name!=='Edit' || results.has(p.toolUseId) || Date.parse(current.timestamp)>pendingTime ||
|
||||
!/^msg_[A-Za-z0-9_-]{1,160}$/.test(current.messageId??'') || !/^req_[A-Za-z0-9_-]{1,160}$/.test(current.requestId??'') ||
|
||||
input?.file_path!==p.file || typeof input.old_string!=='string' || !input.old_string ||
|
||||
typeof input.new_string!=='string' || input.new_string===input.old_string ||
|
||||
(input.replace_all!==undefined&&input.replace_all!==false)) return null;
|
||||
const queued=new Set<string>();
|
||||
let queuedPlans = 0;
|
||||
for (const mutation of [...uses.values()].filter(e=>e.name==='Edit'||e.name==='Write')) {
|
||||
const result=results.get(mutation.toolUseId);
|
||||
if (result && (Date.parse(mutation.timestamp)>pendingTime || Date.parse(result.timestamp)>pendingTime)) return null;
|
||||
if (mutation.toolUseId===p.toolUseId || result) continue;
|
||||
// Later publications are queued only when this exact batch owns them and
|
||||
// the recorder has not started them. They never supply current authority.
|
||||
const nativePlan = mutation.input?.file_path !== p.file &&
|
||||
ownedQueuedNativePlan(mutation.input?.file_path, context.ownedNativePlansRoot, pendingTime);
|
||||
if (mutation.name!=='Edit' || (mutation.input?.file_path!==p.file && !nativePlan) ||
|
||||
typeof mutation.input.old_string!=='string' || !mutation.input.old_string ||
|
||||
typeof mutation.input.new_string!=='string' || mutation.input.old_string===mutation.input.new_string ||
|
||||
(mutation.input.replace_all!==undefined && mutation.input.replace_all!==false) ||
|
||||
mutation.messageId!==current.messageId || mutation.requestId!==current.requestId ||
|
||||
Date.parse(mutation.timestamp)>context.viewportCapturedAt ||
|
||||
events.indexOf(mutation)<=events.indexOf(current) || p.hookSeenIds.includes(mutation.toolUseId)) return null;
|
||||
if (nativePlan) {
|
||||
if (![...uses.values()].some(previous => (previous.name === 'Write' || previous.name === 'Edit') &&
|
||||
previous.input?.file_path === mutation.input!.file_path &&
|
||||
results.get(previous.toolUseId)?.isError === false &&
|
||||
Date.parse(results.get(previous.toolUseId)!.timestamp) <= pendingTime)) return null;
|
||||
queuedPlans++;
|
||||
}
|
||||
queued.add(mutation.toolUseId);
|
||||
}
|
||||
try {
|
||||
if (Math.floor(fs.statSync(p.file).mtimeMs)>pendingTime) return null;
|
||||
const actual=createAutoplanEditDigest(p.file,input.old_string,input.new_string), expected=p.editDigest!;
|
||||
if (!actual || actual.beforeSHA256!==expected.beforeSHA256 || actual.requestSHA256!==expected.requestSHA256 ||
|
||||
JSON.stringify(actual.oldLineHashes)!==JSON.stringify(expected.oldLineHashes) ||
|
||||
JSON.stringify(actual.newLineHashes)!==JSON.stringify(expected.newLineHashes)) return null;
|
||||
const commandViewport = queuedCommandViewport(viewport,current,events,queued,p.hookSeenIds,context.viewportCapturedAt);
|
||||
const rendered = queuedArtifactViewport(commandViewport,current,events,queued,p.file,pendingTime,context.ownedStateRoot);
|
||||
return autoplanArtifactPermissionInput(queuedPlanViewport(rendered,queuedPlans,current,events),{...context,
|
||||
publicTools:events.filter(e=>!queued.has(e.toolUseId))},seen);
|
||||
} catch { return null; }
|
||||
}
|
||||
@@ -0,0 +1,235 @@
|
||||
/** Opt-in pending Edit metadata for launcher-owned Autoplan plan artifacts. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { ownedAutoplanArtifact, hasPendingAutoplanArtifactHistory } from './autoplan-artifact-permission';
|
||||
import { createAutoplanEditDigest, validAutoplanEditDigest, type AutoplanEditDigest } from './autoplan-artifact-digest';
|
||||
import { readPlanCountTranscript, type NativePublicToolEvent } from './plan-count-transcript';
|
||||
|
||||
// Suffix commitments are capped at 8192 hashes; old metadata/input limits stay unchanged.
|
||||
const MAX_BYTES = 1024 * 1024, MAX_INPUT = 4 * 1024 * 1024, MAX_IDS = 128;
|
||||
const reasons = ['invalid_event','record_error','lock_conflict','concurrent_pending','conflicting_replay',
|
||||
'record_overflow','input_overflow','stdin_timeout','hook_error','approval_withheld'] as const;
|
||||
type Reason = typeof reasons[number];
|
||||
const object = (v: unknown): v is Record<string, any> => v !== null && typeof v === 'object' && !Array.isArray(v);
|
||||
const id = (v: unknown): v is string => typeof v === 'string' && /^[A-Za-z0-9_-]{1,160}$/.test(v);
|
||||
const keys = (v: Record<string, any>, allowed: string[]) => Object.keys(v).every(k => allowed.includes(k));
|
||||
const quote = (v: string) => `'${(process.platform === 'win32' ? v.replaceAll('\\','/') : v).replaceAll("'", "'\\''")}'`;
|
||||
export interface PendingAutoplanArtifact {
|
||||
source:'pre_tool_use'; sessionId:string; toolUseId:string; tool:'Edit'; file:string; timestamp:string; editDigest?:AutoplanEditDigest;
|
||||
/** Validated recorder tombstones; exposed only by the published-current opt-in. */
|
||||
hookSeenIds?:string[];
|
||||
}
|
||||
interface Pending extends PendingAutoplanArtifact { transcriptPath:string }
|
||||
interface State { version:1; cwd:string; config:string; stateRoot:string; sessionId?:string; approvalStartedAt?:number; seenIds:string[]; pending:Pending|null }
|
||||
|
||||
function scopedTranscript(file: unknown, config: string, session: string): file is string {
|
||||
if (typeof file !== 'string' || !path.isAbsolute(file)) return false;
|
||||
const rel = path.relative(path.join(config, 'projects'), file).split(path.sep);
|
||||
if (rel.length !== 2 || !rel[0] || rel[0] === '..' || rel[0] === '.' || rel[1] !== `${session}.jsonl`) return false;
|
||||
try { return fs.lstatSync(file).isFile() && !fs.lstatSync(path.join(config,'projects')).isSymbolicLink() &&
|
||||
!fs.lstatSync(path.dirname(file)).isSymbolicLink() &&
|
||||
fs.realpathSync(file) === path.join(fs.realpathSync(config),'projects',...rel); } catch { return false; }
|
||||
}
|
||||
function readState(file:string, cwd:string, config:string, stateRoot:string): State {
|
||||
const stat = fs.lstatSync(file);
|
||||
if (!stat.isFile() || stat.size > MAX_BYTES) throw Error('record');
|
||||
const s = JSON.parse(fs.readFileSync(file,'utf8'));
|
||||
if (!object(s) || !keys(s,['version','cwd','config','stateRoot','sessionId','approvalStartedAt','seenIds','pending']) || s.version !== 1 ||
|
||||
(s.approvalStartedAt !== undefined && (!Number.isSafeInteger(s.approvalStartedAt) || s.approvalStartedAt <= 0)) ||
|
||||
s.cwd !== cwd || s.config !== config || s.stateRoot !== stateRoot || (s.sessionId !== undefined && !id(s.sessionId)) ||
|
||||
!Array.isArray(s.seenIds) || s.seenIds.length > MAX_IDS || !s.seenIds.every(id) || new Set(s.seenIds).size !== s.seenIds.length ||
|
||||
(s.sessionId === undefined && (s.seenIds.length || s.pending !== null))) throw Error('record');
|
||||
const p = s.pending;
|
||||
if (p !== null && (!object(p) || !keys(p,['source','sessionId','toolUseId','tool','file','timestamp','transcriptPath','editDigest']) ||
|
||||
p.source !== 'pre_tool_use' || p.tool !== 'Edit' || p.sessionId !== s.sessionId || !id(p.toolUseId) || !s.seenIds.includes(p.toolUseId) ||
|
||||
(p.editDigest!==undefined && !validAutoplanEditDigest(p.editDigest)) ||
|
||||
!scopedTranscript(p.transcriptPath,config,p.sessionId) || !Number.isFinite(Date.parse(p.timestamp)) ||
|
||||
!ownedAutoplanArtifact(p.file,{cwd,ownedStateRoot:stateRoot}))) throw Error('record');
|
||||
return s as State;
|
||||
}
|
||||
function poison(file:string, reason:Reason) {
|
||||
try { fs.writeFileSync(file+'.invalid',JSON.stringify({reason})+'\n',{mode:0o600,flag:'wx'}); } catch { /* already invalid or closed */ }
|
||||
}
|
||||
export function createAutoplanArtifactRecorder(cwd:string, config:string, stateRoot:string, approveEdits=false) {
|
||||
if (![cwd,config,stateRoot].every(path.isAbsolute) || !fs.lstatSync(stateRoot).isDirectory()) throw Error('owned runtime required');
|
||||
const createdAt = Date.now();
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(),'gstack-autoplan-artifact-')), file = path.join(dir,'state.json');
|
||||
fs.writeFileSync(file,JSON.stringify({version:1,cwd,config,stateRoot,seenIds:[],pending:null})+'\n',{mode:0o600});
|
||||
const command = [process.execPath,import.meta.path,'--record',file,cwd,config,stateRoot,...(approveEdits ? ['--approve-edits'] : [])].map(quote).join(' ');
|
||||
// Foreign/current file mutations invalidate concurrent owned identity. Only
|
||||
// allowlisted Edit requests can become pending; Write supplies no authority.
|
||||
const hook = {matcher:'^(Write|Edit)$',hooks:[{type:'command',command,timeout:5}]};
|
||||
return {file,hooks:{PreToolUse:[hook],PostToolUse:[hook],PostToolUseFailure:[hook]},
|
||||
// Only the owning caller activates approval, once, at the actual slash-command start.
|
||||
startEditApproval: approveEdits ? (startedAt:number) => {
|
||||
let lock:number|undefined;
|
||||
const temporary = `${file}.start.tmp`;
|
||||
try {
|
||||
if (!Number.isSafeInteger(startedAt) || startedAt < createdAt || startedAt > Date.now() ||
|
||||
fs.existsSync(file+'.invalid')) throw Error('invalid approval start');
|
||||
lock=fs.openSync(file+'.lock','wx',0o600);
|
||||
const state=readState(file,cwd,config,stateRoot);
|
||||
if (state.approvalStartedAt !== undefined || state.pending) throw Error('approval already started or pending');
|
||||
state.approvalStartedAt=startedAt;
|
||||
fs.writeFileSync(temporary,JSON.stringify(state)+'\n',{mode:0o600,flag:'wx'});fs.renameSync(temporary,file);
|
||||
} catch (error) { poison(file,'record_error'); throw error; }
|
||||
finally {
|
||||
try { fs.rmSync(temporary,{force:true}); } catch {}
|
||||
if (lock!==undefined) { fs.closeSync(lock);fs.unlinkSync(file+'.lock'); }
|
||||
}
|
||||
} : undefined,
|
||||
dispose:()=>fs.rmSync(dir,{recursive:true,force:true})};
|
||||
}
|
||||
|
||||
/** Reuses the pending UI path's history gate; current publication supplies no success. */
|
||||
function canApproveEdit(e:Record<string,any>, pending:Pending, state:State):boolean {
|
||||
if (state.approvalStartedAt === undefined || !validAutoplanEditDigest(pending.editDigest)) return false;
|
||||
const tools:NativePublicToolEvent[]=[];
|
||||
const transcript=readPlanCountTranscript(state.config,state.cwd,event=>tools.push(event),pending.transcriptPath);
|
||||
const current=tools.filter(event=>event.toolUseId===pending.toolUseId);
|
||||
if (current.length) {
|
||||
const use=current[0]!, input=use.input;
|
||||
// A published request must be the same one that invoked this hook. No
|
||||
// completed or conflicting identity can be removed from the history gate.
|
||||
if (current.length!==1 || use.kind!=='use' || use.name!=='Edit' || use.sessionId!==pending.sessionId ||
|
||||
!/^msg_[A-Za-z0-9_-]{1,160}$/.test(use.messageId??'') || !/^req_[A-Za-z0-9_-]{1,160}$/.test(use.requestId??'') ||
|
||||
!Number.isFinite(Date.parse(use.timestamp)) || Date.parse(use.timestamp)<state.approvalStartedAt ||
|
||||
Date.parse(use.timestamp)>Date.parse(pending.timestamp) || input?.file_path!==pending.file ||
|
||||
input.old_string!==e.tool_input.old_string || input.new_string!==e.tool_input.new_string ||
|
||||
(input.replace_all!==undefined && input.replace_all!==false)) return false;
|
||||
}
|
||||
if (!hasPendingAutoplanArtifactHistory({cwd:state.cwd,ownedStateRoot:state.stateRoot,
|
||||
commandStartedAt:state.approvalStartedAt,now:Date.now(),transcriptStatus:transcript.status,
|
||||
publicTools:tools.filter(event=>event.toolUseId!==pending.toolUseId),pending})) return false;
|
||||
// Re-read after the journal checks. Missing/duplicate old_string, changed
|
||||
// bytes, or a file modified after this hook's timestamp cannot gain approval.
|
||||
const fresh=createAutoplanEditDigest(pending.file,e.tool_input.old_string,e.tool_input.new_string);
|
||||
return !!fresh && JSON.stringify(fresh)===JSON.stringify(pending.editDigest) &&
|
||||
Math.floor(fs.statSync(pending.file).mtimeMs)<=Date.parse(pending.timestamp);
|
||||
}
|
||||
|
||||
/** Persists metadata/digests only. Opted-in approval never supplies tool success or phase credit. */
|
||||
export function recordAutoplanArtifact(input:string, file:string, cwd:string, config:string, stateRoot:string, approveEdits=false) {
|
||||
let approved=false;
|
||||
let lock:number|undefined, reason:Reason='invalid_event';
|
||||
const temporary = `${file}.${process.pid}.tmp`;
|
||||
try {
|
||||
if (Buffer.byteLength(input)>MAX_INPUT) { reason='input_overflow'; throw Error('input'); }
|
||||
const e = JSON.parse(input);
|
||||
if (object(e) && (e.agent_id !== undefined || (typeof e.cwd === 'string' && e.cwd !== cwd))) return;
|
||||
if (fs.existsSync(file+'.invalid')) return;
|
||||
reason='lock_conflict'; lock=fs.openSync(file+'.lock','wx',0o600);
|
||||
reason='record_error'; const old=readState(file,cwd,config,stateRoot);
|
||||
reason='invalid_event';
|
||||
if (!object(e) || e.cwd!==cwd || !['PreToolUse','PostToolUse','PostToolUseFailure'].includes(e.hook_event_name) ||
|
||||
!['Write','Edit'].includes(e.tool_name) || !id(e.session_id) || !id(e.tool_use_id) ||
|
||||
!scopedTranscript(e.transcript_path,config,e.session_id) || !object(e.tool_input) ||
|
||||
typeof e.tool_input.file_path!=='string') throw Error('event');
|
||||
if (old.sessionId!==undefined && old.sessionId!==e.session_id) return;
|
||||
const previous=old.pending;
|
||||
// An unseen completion reveals an unobserved concurrent mutation. It cannot
|
||||
// leave an older request current; a known old completion is harmless.
|
||||
if (previous && previous.toolUseId!==e.tool_use_id && !old.seenIds.includes(e.tool_use_id)) {
|
||||
reason='concurrent_pending'; throw Error('unobserved concurrent mutation');
|
||||
}
|
||||
if (previous?.toolUseId===e.tool_use_id && (e.tool_name!=='Edit' || previous.file!==e.tool_input.file_path || previous.transcriptPath!==e.transcript_path)) {
|
||||
reason='conflicting_replay'; throw Error('identity changed');
|
||||
}
|
||||
const eligible=e.tool_name==='Edit' && ownedAutoplanArtifact(e.tool_input.file_path,{cwd,ownedStateRoot:stateRoot});
|
||||
if (!eligible) {
|
||||
if (previous && e.hook_event_name==='PreToolUse') { reason='concurrent_pending'; throw Error('foreign pending'); }
|
||||
return;
|
||||
}
|
||||
const state:State={...old,sessionId:e.session_id};
|
||||
if (e.hook_event_name==='PreToolUse') {
|
||||
if (old.seenIds.includes(e.tool_use_id)) {
|
||||
if (previous?.toolUseId===e.tool_use_id && previous.editDigest) {
|
||||
reason='conflicting_replay';
|
||||
const digest=typeof e.tool_input.old_string==='string' && typeof e.tool_input.new_string==='string' &&
|
||||
(e.tool_input.replace_all===undefined || e.tool_input.replace_all===false)
|
||||
? createAutoplanEditDigest(e.tool_input.file_path,e.tool_input.old_string,e.tool_input.new_string, previous.editDigest.clippedAdditions !== undefined) : undefined;
|
||||
if (!digest || JSON.stringify(digest)!==JSON.stringify(previous.editDigest)) throw Error('request changed');
|
||||
}
|
||||
return;
|
||||
}
|
||||
if (previous) { reason='concurrent_pending'; throw Error('concurrent'); }
|
||||
// Match the published-path edit scope; inspect but never retain content.
|
||||
if (typeof e.tool_input.old_string!=='string' || !e.tool_input.old_string || typeof e.tool_input.new_string!=='string' ||
|
||||
e.tool_input.new_string===e.tool_input.old_string ||
|
||||
(e.tool_input.replace_all!==undefined && e.tool_input.replace_all!==false)) throw Error('edit scope');
|
||||
state.pending={source:'pre_tool_use',sessionId:e.session_id,toolUseId:e.tool_use_id,tool:'Edit',file:e.tool_input.file_path,
|
||||
timestamp:new Date().toISOString(),transcriptPath:e.transcript_path};
|
||||
const digest=createAutoplanEditDigest(e.tool_input.file_path,e.tool_input.old_string,e.tool_input.new_string);
|
||||
if (digest) state.pending.editDigest=digest;
|
||||
} else if (previous?.toolUseId===e.tool_use_id) state.pending=null;
|
||||
// Success/failure hooks clear pending and tombstone the ID, but never
|
||||
// supply successful history or phase coverage. Those still require JSONL.
|
||||
if (!state.seenIds.includes(e.tool_use_id)) state.seenIds=[...state.seenIds,e.tool_use_id];
|
||||
if (approveEdits && e.hook_event_name==='PreToolUse' && state.pending) approved=canApproveEdit(e,state.pending,state);
|
||||
const serialized=JSON.stringify(state)+'\n'; reason='record_overflow';
|
||||
if (state.seenIds.length>MAX_IDS || Buffer.byteLength(serialized)>MAX_BYTES) throw Error('record overflow');
|
||||
reason='record_error'; fs.writeFileSync(temporary,serialized,{mode:0o600,flag:'wx'}); fs.renameSync(temporary,file);
|
||||
// Persist the exact failed pending identity for diagnostics, but never let
|
||||
// the opted-in caller retry a rejected Edit through terminal navigation.
|
||||
if (approveEdits && state.approvalStartedAt!==undefined && e.hook_event_name==='PreToolUse' && !approved)
|
||||
poison(file,'approval_withheld');
|
||||
} catch { approved=false; poison(file,reason); }
|
||||
finally {
|
||||
try { fs.rmSync(temporary,{force:true}); } catch {}
|
||||
if (lock!==undefined) { try { fs.closeSync(lock); fs.unlinkSync(file+'.lock'); } catch { poison(file,'record_error'); } }
|
||||
}
|
||||
return approved && !fs.existsSync(file+'.invalid');
|
||||
}
|
||||
|
||||
export function autoplanArtifactRecorderStatus(file:string|undefined,cwd:string,config:string|null,stateRoot:string|undefined):
|
||||
{status:'disabled'|'missing'|'busy'|'idle'|'pending'|'invalid';reason?:string} {
|
||||
if (!file || !config || !stateRoot) return {status:'disabled'};
|
||||
try {
|
||||
if (fs.existsSync(file+'.invalid')) {
|
||||
const stat=fs.lstatSync(file+'.invalid');
|
||||
if (!stat.isFile() || stat.size>1024) return {status:'invalid',reason:'record_error'};
|
||||
const r=JSON.parse(fs.readFileSync(file+'.invalid','utf8'))?.reason;
|
||||
return {status:'invalid',reason:reasons.includes(r) ? r : 'unknown'};
|
||||
}
|
||||
if (fs.existsSync(file+'.lock')) return {status:'busy'};
|
||||
if (!fs.existsSync(file)) return {status:'missing'};
|
||||
return {status:readState(file,cwd,config,stateRoot).pending ? 'pending' : 'idle'};
|
||||
} catch { return {status:'invalid',reason:'record_error'}; }
|
||||
}
|
||||
/** Native approval owns pending Edits; UI navigation resumes only when idle. */
|
||||
export function autoplanArtifactApprovalBoundary(status:ReturnType<typeof autoplanArtifactRecorderStatus>):'clear'|'pending'|'failed' {
|
||||
if (status.status==='idle') return 'clear';
|
||||
if (status.status==='pending' || status.status==='busy') return 'pending';
|
||||
return 'failed';
|
||||
}
|
||||
export function readPendingAutoplanArtifact(file:string|undefined,cwd:string,config:string|null,stateRoot:string|undefined,
|
||||
startedAt:number, publicTools:readonly NativePublicToolEvent[], now=Date.now(), allowPublished=false): PendingAutoplanArtifact|undefined {
|
||||
if (!file || !config || !stateRoot || !Number.isFinite(startedAt) || !Number.isFinite(now) ||
|
||||
autoplanArtifactRecorderStatus(file,cwd,config,stateRoot).status!=='pending') return undefined;
|
||||
try {
|
||||
const state=readState(file,cwd,config,stateRoot), p=state.pending!;
|
||||
const time=Date.parse(p.timestamp), sessions=new Set(publicTools.map(e=>e.sessionId));
|
||||
if (sessions.size!==1 || !sessions.has(p.sessionId) || time<startedAt || time>now ||
|
||||
(!allowPublished && publicTools.some(e=>e.toolUseId===p.toolUseId))) return undefined;
|
||||
const {transcriptPath:_,...pending}=p;
|
||||
return allowPublished ? {...pending,hookSeenIds:[...state.seenIds]} : pending;
|
||||
} catch { return undefined; }
|
||||
}
|
||||
if (import.meta.main && process.argv[2]==='--record') {
|
||||
const [file,cwd,config,stateRoot,approval]=process.argv.slice(3);
|
||||
if (file && cwd && config && stateRoot) {
|
||||
const timer=setTimeout(()=>{poison(file,'stdin_timeout');process.exit(0);},4000);
|
||||
try {
|
||||
const chunks:Uint8Array[]=[]; let size=0;
|
||||
for await (const chunk of Bun.stdin.stream()) {
|
||||
size+=chunk.byteLength;
|
||||
if (size>MAX_INPUT) { poison(file,'input_overflow');process.exit(0); }
|
||||
chunks.push(chunk);
|
||||
}
|
||||
const approved=recordAutoplanArtifact(Buffer.concat(chunks).toString('utf8'),file,cwd,config,stateRoot,approval==='--approve-edits');
|
||||
if (approved) process.stdout.write(JSON.stringify({hookSpecificOutput:{hookEventName:'PreToolUse',permissionDecision:'allow'}})+'\n');
|
||||
} catch { poison(file,'hook_error'); }
|
||||
finally { clearTimeout(timer); }
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,144 @@
|
||||
/** Exact parent tool-delivery audit; neither preparation nor a claimed range is a Read. */
|
||||
import { createHash } from 'node:crypto';
|
||||
import { lstatSync, readFileSync, realpathSync } from 'node:fs';
|
||||
import { basename, dirname, join, sep } from 'node:path';
|
||||
import type { NativePublicToolEvent } from './plan-count-transcript';
|
||||
|
||||
export interface MethodologyReadBinding {
|
||||
phase: string;
|
||||
path: string;
|
||||
sha256: string;
|
||||
bytes: number;
|
||||
content: string;
|
||||
lines: number;
|
||||
}
|
||||
export interface AutoplanMethodReadAudit {
|
||||
phase: string;
|
||||
sessionId: string;
|
||||
dispatchToolUseId: string;
|
||||
at: string;
|
||||
passed: boolean;
|
||||
methodologyPath?: string;
|
||||
sha256?: string;
|
||||
ranges: Array<{ startLine: number; endLine: number; toolUseId: string }>;
|
||||
missing: Array<{ startLine: number; endLine: number }>;
|
||||
error?: string;
|
||||
}
|
||||
const hash = (text: string | Buffer) => createHash('sha256').update(text).digest('hex');
|
||||
const object = (x: unknown): x is Record<string, any> => x !== null && typeof x === 'object' && !Array.isArray(x);
|
||||
const identity = (x: unknown): x is string => typeof x === 'string' && x.trim().length > 0;
|
||||
const integer = (x: unknown): x is number => Number.isSafeInteger(x) && (x as number) > 0;
|
||||
const phaseOf = (prompt: string) => /^You are the independent (CEO|DESIGN|DX|ENG) reviewer for this phase\.\n/.exec(prompt)?.[1]?.toLowerCase();
|
||||
|
||||
/** Paths come from the actual dispatch but may only resolve inside this fixture's owned roots. */
|
||||
export function loadAutoplanMethodologyBinding(prompt: string, ownedRoots: string[]): MethodologyReadBinding {
|
||||
const phase = phaseOf(prompt);
|
||||
const pathLine = /^Read file: ("[^\n]+")$/m.exec(prompt)?.[1];
|
||||
if (!phase || !pathLine) throw new Error('Unrecognized native phase dispatch');
|
||||
const nativePath: unknown = JSON.parse(pathLine);
|
||||
const roots = ownedRoots.map(root => realpathSync(root));
|
||||
const read = (path: unknown): Buffer => {
|
||||
if (typeof path !== 'string' || realpathSync(path) !== path ||
|
||||
!roots.some(root => path.startsWith(root + sep))) throw new Error('Artifact outside owned fixture roots or aliased');
|
||||
const stat = lstatSync(path);
|
||||
if (!stat.isFile() || stat.size > 32 * 1024 * 1024 ||
|
||||
(process.platform !== 'win32' && (stat.mode & 0o777) !== 0o444)) throw new Error('Artifact is not immutable bounded regular data');
|
||||
return readFileSync(path);
|
||||
};
|
||||
if (typeof nativePath !== 'string' || basename(nativePath) !== 'native-prompt.md' ||
|
||||
!basename(dirname(nativePath)).startsWith(`autoplan-${phase}-`)) throw new Error('Foreign phase native input');
|
||||
const native = read(nativePath);
|
||||
const snapshot = JSON.parse(read(join(dirname(nativePath), 'snapshot.json')).toString('utf8'));
|
||||
if (snapshot.schemaVersion !== 2 || snapshot.phase !== phase || snapshot.nativePromptPath !== nativePath ||
|
||||
snapshot.nativeDispatchPrompt !== prompt || snapshot.nativePromptSha256 !== hash(native) ||
|
||||
snapshot.nativePromptBytes !== native.length || snapshot.nativePromptLines !== native.toString('utf8').split('\n').length ||
|
||||
!object(snapshot.methodology)) throw new Error('Dispatch does not match immutable snapshot');
|
||||
const identity = snapshot.methodology;
|
||||
if (typeof identity.methodologyPath !== 'string' || basename(identity.methodologyPath) !== 'methodology.md' ||
|
||||
!basename(dirname(identity.methodologyPath)).startsWith(`autoplan-${phase}-`)) throw new Error('Foreign methodology identity');
|
||||
const content = read(identity.methodologyPath);
|
||||
const manifestBytes = read(join(dirname(identity.methodologyPath), 'methodology.json'));
|
||||
const manifest = JSON.parse(manifestBytes.toString('utf8'));
|
||||
if (hash(manifestBytes) !== identity.manifestSha256 || manifest.phase !== phase ||
|
||||
manifest.methodologyPath !== identity.methodologyPath || hash(content) !== identity.sha256 ||
|
||||
manifest.sha256 !== identity.sha256 || content.length !== identity.bytes || manifest.bytes !== identity.bytes ||
|
||||
content.toString('utf8').split('\n').length !== identity.lines || manifest.lines !== identity.lines) {
|
||||
throw new Error('Methodology bytes do not match dispatched snapshot');
|
||||
}
|
||||
return { phase, path: identity.methodologyPath, content: content.toString('utf8'),
|
||||
sha256: identity.sha256, bytes: identity.bytes, lines: identity.lines };
|
||||
}
|
||||
|
||||
/**
|
||||
* The caller supplies only public records from its existing owned parent transcript reader.
|
||||
* Bind session + tool ID, actual request/result order, exact path and every delivered line.
|
||||
* A child Read, error, self-report, future result or plausible hash alone supplies no coverage.
|
||||
*/
|
||||
export function auditAutoplanMethodReads(
|
||||
events: NativePublicToolEvent[],
|
||||
bindingFor: (prompt: string) => MethodologyReadBinding,
|
||||
): AutoplanMethodReadAudit[] {
|
||||
const audits: AutoplanMethodReadAudit[] = [];
|
||||
for (let dispatchIndex = 0; dispatchIndex < events.length; dispatchIndex++) {
|
||||
const dispatch = events[dispatchIndex]!;
|
||||
if (dispatch.kind !== 'use' || dispatch.name !== 'Agent' || typeof dispatch.input?.prompt !== 'string') continue;
|
||||
const phase = phaseOf(dispatch.input.prompt);
|
||||
if (!phase) continue;
|
||||
const audit: AutoplanMethodReadAudit = { phase, sessionId: dispatch.sessionId,
|
||||
dispatchToolUseId: dispatch.toolUseId, at: dispatch.timestamp, passed: false, ranges: [], missing: [] };
|
||||
audits.push(audit);
|
||||
try {
|
||||
const dispatchTime = Date.parse(dispatch.timestamp);
|
||||
if (!identity(dispatch.sessionId) || !identity(dispatch.toolUseId) || !Number.isFinite(dispatchTime)) {
|
||||
throw new Error('Invalid native dispatch identity or timestamp');
|
||||
}
|
||||
const binding = bindingFor(dispatch.input.prompt);
|
||||
if (binding.phase !== phase || !integer(binding.lines) || binding.lines > 1_000_000 ||
|
||||
hash(binding.content) !== binding.sha256 || Buffer.byteLength(binding.content) !== binding.bytes ||
|
||||
binding.content.split('\n').length !== binding.lines) throw new Error('Invalid methodology binding');
|
||||
audit.methodologyPath = binding.path; audit.sha256 = binding.sha256;
|
||||
const lines = binding.content.split('\n');
|
||||
const covered = new Set<number>();
|
||||
const requests = new Map<string, NativePublicToolEvent>();
|
||||
const results = new Map<string, string>();
|
||||
for (const event of events.slice(0, dispatchIndex)) {
|
||||
if (event.sessionId !== dispatch.sessionId || !identity(event.toolUseId) ||
|
||||
!Number.isFinite(Date.parse(event.timestamp)) || Date.parse(event.timestamp) > dispatchTime) continue;
|
||||
if (event.kind === 'use') {
|
||||
const old = requests.get(event.toolUseId);
|
||||
if (old && JSON.stringify(old) !== JSON.stringify(event)) throw new Error('Conflicting native tool identity');
|
||||
requests.set(event.toolUseId, event); continue;
|
||||
}
|
||||
const request = requests.get(event.toolUseId);
|
||||
if (!request || request.name !== 'Read' || request.input?.file_path !== binding.path ||
|
||||
Date.parse(request.timestamp) > Date.parse(event.timestamp)) continue;
|
||||
const signature = JSON.stringify({ file: event.file, isError: event.isError });
|
||||
const previous = results.get(event.toolUseId);
|
||||
if (previous !== undefined && previous !== signature) throw new Error('Conflicting native Read results');
|
||||
results.set(event.toolUseId, signature);
|
||||
const file = event.file;
|
||||
if (event.isError || !object(file) || file.filePath !== binding.path || typeof file.content !== 'string' ||
|
||||
!integer(file.startLine) || !integer(file.numLines) || file.totalLines !== binding.lines ||
|
||||
file.startLine + file.numLines - 1 > binding.lines) continue;
|
||||
const offset = request.input?.offset ?? 1;
|
||||
const limit = request.input?.limit;
|
||||
if (offset !== file.startLine || (limit !== undefined && (!integer(limit) || file.numLines > limit))) continue;
|
||||
const expected = lines.slice(file.startLine - 1, file.startLine - 1 + file.numLines).join('\n');
|
||||
if (file.content !== expected) continue;
|
||||
if (previous !== undefined) continue;
|
||||
const end = file.startLine + file.numLines - 1;
|
||||
audit.ranges.push({ startLine: file.startLine, endLine: end, toolUseId: event.toolUseId });
|
||||
for (let line = file.startLine; line <= end; line++) covered.add(line);
|
||||
}
|
||||
for (let line = 1; line <= binding.lines; line++) {
|
||||
if (covered.has(line)) continue;
|
||||
const startLine = line;
|
||||
while (line < binding.lines && !covered.has(line + 1)) line++;
|
||||
audit.missing.push({ startLine, endLine: line });
|
||||
}
|
||||
audit.passed = audit.missing.length === 0;
|
||||
if (!audit.passed) audit.error = 'Incomplete successful parent methodology Read content before native dispatch';
|
||||
} catch (error) { audit.error = String(error); }
|
||||
}
|
||||
return audits;
|
||||
}
|
||||
@@ -0,0 +1,122 @@
|
||||
import type { PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
export interface AutoplanPhaseHit {
|
||||
phase: number;
|
||||
ts: number;
|
||||
}
|
||||
|
||||
function phaseDeclaration(text: string): RegExpExecArray | null {
|
||||
const declaration = String.raw`Phase[ \t]+(1|2(?:\.5)?|3)(?:[ \t]+\(([^()]*)\))?[ \t]+(?:is[ \t]+)?(?:complete(?:d)?|done|finished|wrapped[ \t]+up)`;
|
||||
const plain = text.replace(new RegExp(String.raw`^\*\*(${declaration}[.:]?)\*\*`, 'i'), '$1');
|
||||
if (/\bEmit\s+phase-transition\s+summary\s*:/i.test(plain)) return null;
|
||||
let match = new RegExp(String.raw`^${declaration}(?:[.:](?:[ \t]+.*)?|)$`, 'i').exec(plain);
|
||||
if (!match) {
|
||||
// A dash or "with" can introduce the results of an actual completion.
|
||||
// Keep the recap affirmative; source, future and withdrawn claims cannot
|
||||
// supply the missing phase declaration.
|
||||
const recap = new RegExp(String.raw`^${declaration}(?:[ \t]*[—–][ \t]*|[ \t]+(?<withResult>with)[ \t]+)(.+)$`, 'i').exec(plain);
|
||||
const tail = recap?.[4]?.trim();
|
||||
if (tail && !/^["“'‘>]|\?|\b(?:if|unless|when|once|pending|maybe|perhaps|would|could|will|source|example|sample|quote(?:d)?|historical|earlier|previous(?:ly)?|template|not|no|never|superseded|provided|rejected|incomplete|unfinished|withdrawn|retracted|cancelled|canceled)\b/i.test(tail)) match = recap;
|
||||
if (match?.groups?.withResult && /\bhypothetic(?:al|ally)\b/i.test(tail!)) match = null;
|
||||
}
|
||||
// Optional phase names are metadata, and must agree with the phase number.
|
||||
const names: Record<string, RegExp> = {
|
||||
'1': /^CEO(?:[ \t]+review)?$/i,
|
||||
'2': /^design(?:[ \t]+review)?$/i,
|
||||
'2.5': /^DX(?:[ \t]+review)?$/i,
|
||||
'3': /^eng(?:ineering)?(?:[ \t]+review)?$/i,
|
||||
};
|
||||
if (match?.[2] !== undefined && !names[match[1]!]!.test(match[2])) return null;
|
||||
return match;
|
||||
}
|
||||
|
||||
/** An explicit current withdrawal in the same announcement cancels a new with-result claim. */
|
||||
function withResultWithdrawn(lines: string[], index: number, phase: string): boolean {
|
||||
let source = false;
|
||||
let fence: {char: string; length: number} | undefined;
|
||||
const owner = new RegExp(String.raw`^(?:(?:correction|current status)[ \t]*:[ \t]*)?(?:this[ \t]+(?:phase|completion|declaration|announcement)|Phase[ \t]+${phase.replace('.', '\\.')})(?:[ \t]+(?:completion|declaration|status))?[ \t]*(?:(?:is|was|has been)[ \t]+|:[ \t]*)(.+)$`, 'i');
|
||||
for (const line of lines.slice(index + 1)) {
|
||||
if (/^(?: {4}|\t)/.test(line)) continue;
|
||||
const text = line.trim().replace(/\*\*/g, '');
|
||||
const delimiter = /^(`{3,}|~{3,})/.exec(text)?.[1];
|
||||
if (delimiter) {
|
||||
if (!fence) fence = {char: delimiter[0]!, length: delimiter.length};
|
||||
else if (delimiter[0] === fence.char && delimiter.length >= fence.length && !text.slice(delimiter.length).trim()) fence = undefined;
|
||||
continue;
|
||||
}
|
||||
if (fence || /^[>"“'‘]/.test(text)) continue;
|
||||
if (/\b(?:example|sample|quote(?:d)?|source|historical|earlier|previous|archived|template)\b.*[::]\s*$/i.test(text)) { source = true; continue; }
|
||||
if (/^(?:current (?:status|review|phase)|correction)\b/i.test(text.replace(/^#{1,6}[ \t]+/, ''))) source = false;
|
||||
if (source) continue;
|
||||
if (phaseDeclaration(text)) break;
|
||||
const status = owner.exec(text)?.[1]?.replace(/["“”'‘’`]/g, '');
|
||||
if (status && /^(?:withdrawn|retracted|cancelled|canceled|superseded|rejected|incomplete|unfinished|not (?:complete(?:d)?|current)|no longer (?:complete(?:d)?|current))\b/i.test(status)) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* Observe actual assistant announcements from this fixture's native transcript.
|
||||
* The terminal renders Markdown bold as ANSI, and its Read output can contain
|
||||
* the same source markers. Neither rendered styling nor tool output is evidence
|
||||
* that a phase completed. Native timestamps also preserve order when several
|
||||
* completed messages arrive between two polls.
|
||||
*/
|
||||
export function autoplanPhaseCompletions(
|
||||
transcript: PlanCountTranscript,
|
||||
commandStartedAt: number,
|
||||
): AutoplanPhaseHit[] {
|
||||
if (transcript.status !== 'ready') return [];
|
||||
const hits: AutoplanPhaseHit[] = [];
|
||||
const messages = [...transcript.assistantMessages]
|
||||
.filter(message => Number.isFinite(Date.parse(message.timestamp)) && Date.parse(message.timestamp) >= commandStartedAt)
|
||||
.sort((a, b) => Date.parse(a.timestamp) - Date.parse(b.timestamp));
|
||||
|
||||
for (const message of messages) {
|
||||
let fence: { char: string; length: number } | undefined;
|
||||
let previousLine = '';
|
||||
const lines = message.text.split(/\r?\n/);
|
||||
for (let index = 0; index < lines.length; index++) {
|
||||
const line = lines[index]!;
|
||||
// Four spaces/a tab creates an indented Markdown code block. Preserve
|
||||
// that distinction before trimming the declaration's whitespace.
|
||||
if (/^(?: {4}|\t)/.test(line)) continue;
|
||||
const text = line.trim();
|
||||
const delimiter = /^(`{3,}|~{3,})/.exec(text)?.[1];
|
||||
if (delimiter) {
|
||||
if (!fence) fence = { char: delimiter[0]!, length: delimiter.length };
|
||||
else if (delimiter[0] === fence.char && delimiter.length >= fence.length &&
|
||||
!text.slice(delimiter.length).trim()) fence = undefined;
|
||||
previousLine = text;
|
||||
continue;
|
||||
}
|
||||
if (fence) continue;
|
||||
// Accept a plain/bold declaration, never a heading, quoted source,
|
||||
// table cell, checklist, or a sentence promising future completion.
|
||||
let match = phaseDeclaration(text);
|
||||
if (/^>\s/.test(text)) {
|
||||
// The skill's transition summary itself is a blockquote. Accept a
|
||||
// filled-in single-phase summary with concrete consensus counts;
|
||||
// a bare quotation or the template's [N]/[X/Y] examples cannot pass.
|
||||
const block: string[] = [];
|
||||
for (let cursor = index; cursor < lines.length && /^ {0,3}>/.test(lines[cursor]!); cursor++) {
|
||||
block.push(lines[cursor]!.replace(/^ {0,3}>\s?/, ''));
|
||||
}
|
||||
const concreteSummary = block.filter(value => phaseDeclaration(value)).length === 1 &&
|
||||
block.some(value => /^Consensus:\s*\d+\s*\/\s*\d+\b/i.test(value)) &&
|
||||
!/\[[^\]]*\]|\{\{/.test(block.join('\n'));
|
||||
if (concreteSummary) match = phaseDeclaration(text.replace(/^>\s?/, ''));
|
||||
}
|
||||
const introducedExample = /\b(?:example|sample|quote(?:d)?|source|template|instruction|marker|expected\s+(?:output|announcement))\b.*[::]\s*$/i.test(previousLine) ||
|
||||
(!!match?.groups?.withResult && /\b(?:example|sample|quote(?:d)?|source|historical|earlier|previous|archived|hypothetical|template)\b.*[::]\s*$/i.test(previousLine.replace(/\*\*/g, '')));
|
||||
// Keep an example introduction across all its marker/quoted lines,
|
||||
// rather than allowing its second marker to look like real completion.
|
||||
if (text && !(introducedExample && (match || text.startsWith('>')))) previousLine = text;
|
||||
if (!match || introducedExample) continue;
|
||||
if (match.groups?.withResult && withResultWithdrawn(lines, index, match[1]!)) continue;
|
||||
const phase = Number(match[1]);
|
||||
if (!hits.some(hit => hit.phase === phase)) hits.push({ phase, ts: Date.parse(message.timestamp) });
|
||||
}
|
||||
}
|
||||
return hits;
|
||||
}
|
||||
@@ -0,0 +1,48 @@
|
||||
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
|
||||
// Real project routing, matching the section offered by gstack-skill-start.
|
||||
const ROUTING = `## Skill routing
|
||||
|
||||
When the user's request matches an available skill, invoke it via the Skill tool. When in doubt, invoke the skill.
|
||||
|
||||
Key routing rules:
|
||||
- Product ideas/brainstorming → invoke /office-hours
|
||||
- Strategy/scope → invoke /plan-ceo-review
|
||||
- Architecture → invoke /plan-eng-review
|
||||
- Design system/plan review → invoke /design-consultation or /plan-design-review
|
||||
- Full review pipeline → invoke /autoplan
|
||||
- Bugs/errors → invoke /investigate
|
||||
- QA/testing site behavior → invoke /qa or /qa-only
|
||||
- Code review/diff check → invoke /review
|
||||
- Visual polish → invoke /design-review
|
||||
- Ship/deploy/PR → invoke /ship or /land-and-deploy
|
||||
- Save progress → invoke /context-save
|
||||
- Resume context → invoke /context-restore
|
||||
- Author a backlog-ready spec/issue → invoke /spec
|
||||
`;
|
||||
|
||||
/** Only the chain-ordering fixture starts with these onboarding prerequisites. */
|
||||
export function seedAutoplanOnboarding(cwd: string): void {
|
||||
const plan = readFileSync(join(cwd, '.claude/plans/ui-heavy-feature.md'), 'utf8');
|
||||
const section = (title: string): string => {
|
||||
const heading = `## ${title}\n`;
|
||||
const starts = [...plan.matchAll(/^## .*\n/gm)].filter(match => match[0] === heading);
|
||||
if (starts.length !== 1) throw new Error(`Expected one ${title} section in the chain plan`);
|
||||
const start = starts[0].index!;
|
||||
const next = plan.indexOf('\n## ', start + heading.length);
|
||||
const text = plan.slice(start, next < 0 ? undefined : next + 1);
|
||||
if (!text.slice(heading.length).trim()) throw new Error(`Empty ${title} section in the chain plan`);
|
||||
return text;
|
||||
};
|
||||
// Copy background only. The plan retains all new work and review decisions.
|
||||
const brief = section('Context') + section('Existing product and application contracts');
|
||||
const routingFile = join(cwd, 'CLAUDE.md');
|
||||
const designDir = join(cwd, 'docs/designs');
|
||||
if (existsSync(routingFile) || existsSync(join(cwd, 'DESIGN.md')) || existsSync(designDir)) {
|
||||
throw new Error('Autoplan onboarding seed requires a fresh chain fixture');
|
||||
}
|
||||
mkdirSync(designDir, { recursive: true });
|
||||
writeFileSync(join(designDir, 'dashboard-context.md'), brief, { flag: 'wx' });
|
||||
writeFileSync(routingFile, ROUTING, { flag: 'wx' });
|
||||
}
|
||||
@@ -0,0 +1,472 @@
|
||||
import { capturePlanCountQuestion, parseNumberedOptions, planCountPrerequisitePick, planCountQuestionInput, planCountSubmissionInput, type AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
|
||||
import type { NativePlanQuestion, NativePlanQuestionCall, NativePublicToolEvent, PlanCountTranscript } from './plan-count-transcript';
|
||||
import type { readPendingQuestion } from './plan-count-pending-question';
|
||||
|
||||
/** A copied native panel is not actionable after prose or inside a code example. */
|
||||
function activeSetupPanel(visible: string): boolean {
|
||||
const lines = visible.replace(/\r+\n?/g, '\n').trimEnd().split('\n');
|
||||
if (!/^Enter[\t ]+to[\t ]+select[\t ]*·[\t ]*↑\/↓[\t ]+to[\t ]+navigate[\t ]*·[\t ]*Esc[\t ]+to[\t ]+cancel$/i.test(lines.at(-1)?.trim() ?? '')) return false;
|
||||
const header = lines.findLastIndex(line => /^ {0,3}[☐□][^\n]+$/.test(line));
|
||||
if (header < 0) return false;
|
||||
let fence: { char: string; length: number } | undefined;
|
||||
for (const line of lines.slice(0, header)) {
|
||||
const match = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (!match) continue;
|
||||
if (!fence) fence = { char: match[1]![0]!, length: match[1]!.length };
|
||||
else if (match[1]![0] === fence.char && match[1]!.length >= fence.length && !match[2]!.trim()) fence = undefined;
|
||||
}
|
||||
if (fence) return false;
|
||||
const introduction = lines.slice(0, header).findLast(line => !/^[\t ─━-]*$/.test(line)) ?? '';
|
||||
return !/\b(?:example|sample|quot(?:e[sd]?|ed)|template|source)\b.*\b(?:panel|menu|choices?|prompt|question|below|following)\b/i.test(introduction);
|
||||
}
|
||||
|
||||
export type AutoplanSetupDecision =
|
||||
| { kind: 'input'; input: string; signatures: string[] }
|
||||
| { kind: 'waiting' | 'unrelated' }
|
||||
| { kind: 'unsupported_setup'; setup: 'routing' | 'prerequisite'; prompt: string;
|
||||
options: Array<{ index: number; label: string }>; identitySource: 'native-bound' | 'current-native-panel' };
|
||||
|
||||
/** A long boxed routing question can retain its title after only the header scrolls away. */
|
||||
function clippedRoutingTitle(visible: string, question: AskUserQuestionFingerprint): string | null {
|
||||
const lines = visible.replace(/\r+\n?/g, '\n').trimEnd().split('\n');
|
||||
const title = /^ {0,3}[│┃][\t ]*(.+)$/.exec(lines[0] ?? '')?.[1];
|
||||
if (!title || !/^(?:D\s*\d+\s*[—–:-]\s*)?Add\s+(?:gstack\s+)?skill\s+routing\s+rules\s+to\s+CLAUDE\.md\?\s*<gstack-qid:routing-injection>$/i.test(title)) return null;
|
||||
const cursors = lines.flatMap((line, index) => /❯\s*[1-9]\./.test(line) ? [index] : []);
|
||||
if (cursors.length !== 1 || !/^ {0,3}❯\s*1\./.test(lines[cursors[0]!]!)) return null;
|
||||
const before = lines.slice(0, cursors[0]);
|
||||
if (before.some(line => line.trim() && !/^ {0,3}[│┃](?:[\t ]|$)/.test(line)) ||
|
||||
before.some(line => /^ {0,3}[│┃][\t ]*(?:`{3,}|~{3,}|>)/.test(line)) ||
|
||||
(before.join('\n').match(/<gstack-qid/gi)?.length ?? 0) !== 1 ||
|
||||
/[☐□☒]|←[^\n]*Submit|(?:^|\n)[^\n]*[1-9]\.\s*\[[ ✓✔xX]\]/.test(visible)) return null;
|
||||
if (!/^Enter[\t ]+to[\t ]+select[\t ]*·[\t ]*↑\/↓[\t ]+to[\t ]+navigate[\t ]*·[\t ]*Esc[\t ]+to[\t ]+cancel$/i.test(lines.at(-1)?.trim() ?? '')) return null;
|
||||
const rows = lines.slice(cursors[0]).flatMap(line => {
|
||||
const match = /^ {0,3}(?:❯\s*)?([1-9])\.[\t ]*(\S.*?)\s*$/.exec(line);
|
||||
return match ? [{ index: Number(match[1]), label: match[2]! }] : [];
|
||||
});
|
||||
if (rows.length !== 4 || rows.some((row, index) => row.index !== index + 1) ||
|
||||
rows[2]!.label !== 'Type something.' || rows[3]!.label !== 'Chat about this' ||
|
||||
JSON.stringify(rows) !== JSON.stringify(question.options)) return null;
|
||||
return title;
|
||||
}
|
||||
|
||||
/** Identify a complete current setup panel; absence or stale/partial metadata is insufficient. */
|
||||
function completeSetupOptions(visible: string, pending?: NativePlanQuestionCall): Array<{ index: number; label: string }> | null {
|
||||
if (!activeSetupPanel(visible)) return null;
|
||||
const lines = visible.replace(/\r+\n?/g, '\n').trimEnd().split('\n');
|
||||
const header = lines.findLastIndex(line => /^ {0,3}[☐□][^\n]+$/.test(line));
|
||||
const introduction = lines.slice(0, header).findLast(line => !/^[\t ─━-]*$/.test(line)) ?? '';
|
||||
if (/\b(?:example|sample|quot(?:e[sd]?|ed)|template|source)\b[^\n]*:\s*$/i.test(introduction)) return null;
|
||||
const panel = lines.slice(header);
|
||||
if (/[←→☒]|✔\s*Submit/.test(panel[0]!) ||
|
||||
panel.some(line => /(?:^|\s)[1-9]\.\s*\[[ ✓✔xX]\]/.test(line))) return null;
|
||||
if (panel.filter(line => /^ {0,3}❯\s*1\./.test(line)).length !== 1 ||
|
||||
panel.filter(line => /❯\s*[1-9]\./.test(line)).length !== 1) return null;
|
||||
const rows = panel.flatMap(line => {
|
||||
const match = /^ {0,3}(?:❯\s*)?([1-9])\.[\t ]*(\S.*?)\s*$/.exec(line);
|
||||
return match ? [{ index: Number(match[1]), label: match[2]! }] : [];
|
||||
});
|
||||
const compact = (text: string) => text.replace(/\s+/g, '').toLowerCase();
|
||||
if (rows.length < 4 || rows.some((row, index) => row.index !== index + 1) ||
|
||||
compact(rows.at(-2)!.label) !== 'typesomething.' || compact(rows.at(-1)!.label) !== 'chataboutthis') return null;
|
||||
const options = rows.slice(0, -2);
|
||||
if (pending) {
|
||||
const native = pending.questions[0];
|
||||
const cursor = panel.findIndex(line => /^ {0,3}❯\s*1\./.test(line));
|
||||
if (pending.answered || pending.failed || pending.questions.length !== 1 || native?.multiSelect ||
|
||||
compact(panel[0]!.replace(/^ {0,3}[☐□]/, '')) !== compact(native!.header) ||
|
||||
!compact(panel.slice(1, cursor).join(' ')).includes(compact(native!.question)) ||
|
||||
options.length !== native!.options.length || options.some((row, index) => compact(row.label) !== compact(native!.options[index]!.label))) return null;
|
||||
}
|
||||
return options;
|
||||
}
|
||||
|
||||
function unsupportedSetup(visible: string, question: AskUserQuestionFingerprint,
|
||||
setup: 'routing' | 'prerequisite', pending?: NativePlanQuestionCall): AutoplanSetupDecision {
|
||||
const options = completeSetupOptions(visible, pending);
|
||||
return options ? { kind: 'unsupported_setup', setup, prompt: question.promptSnippet, options,
|
||||
identitySource: pending ? 'native-bound' : 'current-native-panel' } : { kind: 'waiting' };
|
||||
}
|
||||
|
||||
/** Pure routing policy; native packet validation still requires every displayed identity. */
|
||||
function routingSetupActions(question: AskUserQuestionFingerprint, allowTemporarySkip: boolean, knownSetupOffer = false) {
|
||||
const primary = question.promptSnippet.replace(/^(?:Routing\s*rules|CLAUDE\.md)\s*/i, '').split('?', 1)[0]!;
|
||||
const prompt = primary.replace(/\s+/g, '');
|
||||
const options = question.options.map(option => {
|
||||
const label = option.label.split(/[│┌\r\n]/, 1)[0]!.trim();
|
||||
// A native label may repeat its menu letter. Remove one corresponding
|
||||
// marker only for action matching; keep the original display identity.
|
||||
const marker = /^([A-Z])\)[\t ]+/i.exec(label);
|
||||
const action = marker && marker[1]!.toUpperCase().charCodeAt(0) - 64 === option.index
|
||||
? label.slice(marker[0].length) : label;
|
||||
return { index: option.index, title: action.replace(/\s+/g, '') };
|
||||
});
|
||||
const add = options.filter(option => /^Add(?:routingrules(?:toCLAUDE\.md)?|toCLAUDE\.md)(?:\(Recommended\))?$/i.test(option.title));
|
||||
// Match the declined setup action, not every English label separately:
|
||||
// No thanks/Skip may stand alone or opt into manual invocation. A manual
|
||||
// migration, deletion, or unrelated workflow is not the opposed action.
|
||||
// "Only" limits the same manual action; it does not add a second action.
|
||||
// Use one whole-label grammar with and without a courtesy/Skip prefix.
|
||||
const manualAction = /^(?:manual(?:invocation|skills)?|(?:I['’]ll)?invoke(?:skills)?manually)(?:[-–—]?only)?$/i;
|
||||
// A temporary Skip is the same opposed setup action only on an intact
|
||||
// two-choice panel. Its description may corroborate manual invocation;
|
||||
// the routing premise and unique Add action below establish its scope.
|
||||
const decline = options.filter(option => {
|
||||
const title = option.title.replace(/\(Recommended\)$/i, '');
|
||||
if (/^Skipfornow$/i.test(title)) return allowTemporarySkip;
|
||||
// The action can stand alone or follow a short courtesy ('No thanks').
|
||||
// Cursor redraws can damage that courtesy while leaving 'invoke skills
|
||||
// manually' intact. Match the complete action, not the spelling of No;
|
||||
// arbitrary preceding instructions and extra trailing actions still fail.
|
||||
const manual = title.replace(/^[a-z]{0,3}thanks[,—–-]/i, '');
|
||||
if (manualAction.test(manual)) return true;
|
||||
const prefix = /^(?:Nothanks|Skip)(?:[,—–-])?/i.exec(title);
|
||||
if (!prefix) return false;
|
||||
// 'No thanks' can be followed by the same explicit Skip action. Strip
|
||||
// that decline verb before checking any optional manual-invocation text.
|
||||
const action = title.slice(prefix[0].length).replace(/^skip(?:[,—–-])?/i, '');
|
||||
return action === '' || manualAction.test(action);
|
||||
});
|
||||
const routingId = /<gstack-qid:routing-injection>/i.test(question.promptSnippet);
|
||||
const routingPremise = /gstack/i.test(prompt) && /CLAUDE\.md/i.test(prompt) && /skillroutingrules/i.test(prompt);
|
||||
// A qid can replace the longer premise, but cannot override a question
|
||||
// about a different target. The Add action and question must agree on
|
||||
// project setup rather than a product routing or taste decision.
|
||||
const claudeTarget = /CLAUDE\.md/i.test(prompt);
|
||||
const quotedPremise = /\b(?:plan|spec|document)\s+(?:quotes?|cites?|references?)\b/i.test(primary);
|
||||
if (!claudeTarget || quotedPremise || (!knownSetupOffer && !routingId && !routingPremise)) return null;
|
||||
return { add, decline };
|
||||
}
|
||||
|
||||
/** A direct setup offer can carry a decision brief without changing its actions. */
|
||||
function contextualPacketSetup(question: NativePlanQuestion, pending: NativePlanQuestionCall) {
|
||||
if (pending.answered !== false || pending.failed !== false || !pending.sessionId || !pending.toolUseId ||
|
||||
question.options.length !== 2 || question.options.some(option => !option.description?.trim())) return null;
|
||||
const ids = [...question.question.matchAll(/<gstack-qid:[a-z0-9-]+>/gi)];
|
||||
const text = question.question.replace(/<gstack-qid:[a-z0-9-]+>/gi, '').trim();
|
||||
const split = /^([^?]+\?)([\s\S]+)$/.exec(text);
|
||||
if (!split || !split[2]!.trim() || split[2]!.includes('?')) return null;
|
||||
// Strip one presentation label, then match the complete substantive offer.
|
||||
// A quoted/conditional/adjacent offer cannot borrow another tab's actions.
|
||||
const offer = split[1]!.replace(/^D[1-9]\d*\s*[—–:-]\s*/, '').replace(/\s+/g, ' ').trim();
|
||||
const context = [split[2]!.trim(), ...question.options.map(option => option.description!.trim())].join('\n');
|
||||
if (/(?:^|\n|[.!]\s+)(?:Also|Then)\b|\bonly\s+(?:if|after)\b|\bunless\b|\bprovided\s+that\b/i.test(context) ||
|
||||
/\b(?:must|need\s+to|have\s+to)\s+(?:run|complete|finish)\s*\/office-hours\b/i.test(context) ||
|
||||
/(?:\/office-hours|design\s+doc(?:ument)?)\s+(?:is\s+)?(?:required|mandatory)\b/i.test(context) ||
|
||||
/\breview\s+is\s+(?:forbidden|blocked)\b|\b(?:skip|bypass|omit)\s+(?:(?:the|this|full|standard|entire|CEO|design|DX|engineering)\s+)*review\b/i.test(context)) return null;
|
||||
const descriptions = question.options.map(option => option.description!.trim().replace(/^[✅❌]\s*/, ''));
|
||||
const fp: AskUserQuestionFingerprint = { signature: '', observedAtMs: 0, preReview: true,
|
||||
promptSnippet: `${question.header} ${question.question}`,
|
||||
options: question.options.map((option, index) => ({ index: index + 1, label: option.label })) };
|
||||
if (/^Routing(?: rules)?$/i.test(question.header.trim()) &&
|
||||
/^Add\s+(?:gstack\s+)?(?:skill\s+)?routing\s+rules\s+to\s+CLAUDE\.md\?$/i.test(offer) &&
|
||||
(!ids.length || ids.length === 1 && ids[0]![0].toLowerCase() === '<gstack-qid:routing-injection>')) {
|
||||
const actions = routingSetupActions(fp, true, true);
|
||||
if (!actions || actions.add.length !== 1 || actions.decline.length !== 1 ||
|
||||
actions.add[0]!.index === actions.decline[0]!.index) return null;
|
||||
const add = actions.add[0]!.index, decline = actions.decline[0]!.index;
|
||||
if (/^(?:Do not|Don't|Never|Skip|Decline)\s+(?:add(?:ing)?\s+)?routing\s+rules\b/i.test(descriptions[add - 1]!) ||
|
||||
/^Add\s+(?:skill\s+)?routing\s+rules\b/i.test(descriptions[decline - 1]!)) return null;
|
||||
return { kind: 'routing', pick: add };
|
||||
}
|
||||
if (!/^(?:Design doc|Prerequisites?(?: doc)?)$/i.test(question.header.trim()) || ids.length ||
|
||||
!/^Run\s*\/office-hours\s+(?:now|first)(?:\s+for\s+(?:a|the)\s+design\s+doc(?:ument)?)?\?$/i.test(offer)) return null;
|
||||
const run = question.options.findIndex(option => /^Run\s*\/office-hours\s*(?:now|first)(?:\s*\(recommended\))?$/i.test(option.label));
|
||||
const skip = question.options.findIndex(option => /^Skip\s*[—–-]\s*(?:proceed\s+with\s+)?standard\s+review(?:\s*\(recommended\))?$/i.test(option.label));
|
||||
if (run < 0 || skip < 0 || run === skip ||
|
||||
/^(?:Skip|Don't|Do not)\b/i.test(descriptions[run]!) ||
|
||||
/^Run\s*\/office-hours\b/i.test(descriptions[skip]!)) return null;
|
||||
// Corroborate the selected label with its short action clause; the rest
|
||||
// of the description may explain tradeoffs without changing that action.
|
||||
const sentence = /^([^.!?]+)([.!?]|$)/.exec(descriptions[skip]!);
|
||||
const action = sentence?.[1]?.trim() ?? '';
|
||||
if (sentence?.[2] === '?' || !/^(?:Review\s+(?:starts?|begins?)\s+(?:immediately|now)|(?:Start|Begin)\s+(?:the\s+)?(?:standard\s+)?review\s+(?:immediately|now)|Proceed\s+(?:directly\s+)?with\s+(?:the\s+)?standard\s+review)(?:\s+(?:using|with|on)\s+[^.!?]+)?$/i.test(action) ||
|
||||
/\b(?:after|when|once|until|if|unless|provided|not|never)\b|n['’]t\b/i.test(action)) return null;
|
||||
return { kind: 'prerequisite', pick: skip + 1 };
|
||||
}
|
||||
|
||||
/** Numbered setup wording may vary; newly admitted forms still consume every description. */
|
||||
function numberedPacketSetup(question: NativePlanQuestion, pending: NativePlanQuestionCall) {
|
||||
if (pending.answered !== false || pending.failed !== false || !pending.sessionId || !pending.toolUseId || question.options.length !== 2) return null;
|
||||
const compact = (value: string | undefined) => (value ?? '').trim().replace(/\s+/g, ' ');
|
||||
const ids = [...question.question.matchAll(/<gstack-qid:[a-z0-9-]+>/gi)];
|
||||
const text = compact(question.question.replace(/<gstack-qid:[a-z0-9-]+>/gi, ''))
|
||||
.replace(/^D[1-9]\d*\s*[—–:-]\s*/, '');
|
||||
const fp: AskUserQuestionFingerprint = { signature: '', observedAtMs: 0, preReview: true,
|
||||
promptSnippet: `${question.header} ${question.question}`,
|
||||
options: question.options.map((option, index) => ({ index: index + 1, label: option.label })) };
|
||||
const routing = routingSetupActions(fp, true);
|
||||
if (ids.length === 1 && ids[0]![0].toLowerCase() === '<gstack-qid:routing-injection>' &&
|
||||
/^Add\s+(?:gstack\s+)?skill\s+routing\s+rules\s+to\s+CLAUDE\.md\?$/i.test(text) &&
|
||||
routing?.add.length === 1 && routing.decline.length === 1 && routing.add[0]!.index !== routing.decline[0]!.index) {
|
||||
const add = question.options[routing.add[0]!.index - 1]!;
|
||||
const decline = question.options[routing.decline[0]!.index - 1]!;
|
||||
if (/^Appends a skill routing section to CLAUDE\.md so future sessions automatically invoke the right skill(?: \(e\.g\. \/autoplan for reviews, \/ship for deploys\))? without you needing to type the command each time\. One-time setup per project\.$/i.test(compact(add.description)) &&
|
||||
/^No change to CLAUDE\.md\. You['’]ll continue calling skills yourself with \/skill-name as you do now\.$/i.test(compact(decline.description))) {
|
||||
return { kind: 'routing', pick: routing.add[0]!.index };
|
||||
}
|
||||
}
|
||||
if (ids.length || !/^No design doc (?:found|exists) for (?:this|the) (?:branch|project)\. Run \/office-hours (?:first|now)(?: to sharpen (?:the|this) review input)?\?$/i.test(text)) return null;
|
||||
const run = question.options.findIndex(option => /^Run\s*\/office-hours\s*(?:now|first)(?:\s*\(recommended\))?$/i.test(option.label));
|
||||
const skip = question.options.findIndex(option => /^Skip\s*[,—–-]\s*proceed\s+with\s+(?:standard\s+)?review(?:\s*\(recommended\))?$/i.test(option.label));
|
||||
if (run < 0 || skip < 0 || run === skip ||
|
||||
!/^Start the full CEO (?:→|->) Design (?:→|->) DX (?:→|->) Eng review pipeline now using the plan as-is\. (?:Recommended when the plan context is already rich enough\. ?)?(?:\(Recommended\))?$/i.test(compact(question.options[skip]!.description)) ||
|
||||
!/^Produces a structured problem statement, premise challenge, and explored alternatives before the review\. (?:Takes ~?\d+(?:[–-]\d+)? min\. )?Gives the review sharper, better-grounded input\.$/i.test(compact(question.options[run]!.description))) return null;
|
||||
return { kind: 'prerequisite', pick: skip + 1 };
|
||||
}
|
||||
|
||||
/** Answer only the known pair of setup offers, using the actual native active tab. */
|
||||
function setupPacketDecision(visible: string, seen: ReadonlySet<string>, pending: NativePlanQuestionCall): AutoplanSetupDecision {
|
||||
const waiting: AutoplanSetupDecision = { kind: 'waiting' };
|
||||
// Validate all questions before touching any tab: a setup question cannot
|
||||
// lend its policy to an adjacent finding, taste decision or checkbox.
|
||||
if (pending.questions.length !== 2) return waiting;
|
||||
const policies = pending.questions.map(question => {
|
||||
if (question.options.length !== 2) return null;
|
||||
const ids = [...question.question.matchAll(/<gstack-qid:[a-z0-9-]+>/gi)];
|
||||
if (ids.length > 1 || (question.question.match(/<gstack-qid/gi)?.length ?? 0) !== ids.length) return null;
|
||||
const offerText = question.question.replace(/<gstack-qid:[a-z0-9-]+>/gi, '').trim();
|
||||
const fp: AskUserQuestionFingerprint = { signature: '', observedAtMs: 0, preReview: true,
|
||||
promptSnippet: `${question.header} ${question.question}`,
|
||||
options: question.options.map((option, index) => ({ index: index + 1, label: option.label })) };
|
||||
const routing = routingSetupActions(fp, true);
|
||||
// A packet must contain only setup. Scope the entire question, including
|
||||
// any premise, rather than borrowing the first question mark's identity
|
||||
// while a later sentence asks for an unrelated approval.
|
||||
const routingOffer = /^(?:gstack\s+works\s+best\s+when\s+(?:your|this|the)\s+project['’]s\s+CLAUDE\.md\s+includes\s+skill\s+routing\s+rules\.\s*)?(?:(?:Should|Can)\s+(?:I|gstack)\s+|Would\s+you\s+like\s+(?:me|gstack)\s+to\s+)?Add\s+(?:(?:gstack\s+)?skill\s+routing\s+rules\s+to\s+(?:this\s+project['’]s\s+)?CLAUDE\.md|(?:skill\s+)?routing\s+rules(?:\s+to\s+CLAUDE\.md)?|them)(?:\s+now)?\?$/i.test(offerText);
|
||||
if (routingOffer && routing?.add.length === 1 && routing.decline.length === 1 && routing.add[0]!.index !== routing.decline[0]!.index) {
|
||||
return { kind: 'routing', pick: routing.add[0]!.index };
|
||||
}
|
||||
const prerequisite = planCountPrerequisitePick(fp);
|
||||
const run = question.options.filter(option => /^Run\s*\/office-hours\s*(?:now|first)(?:\s*\(recommended\))?$/i.test(option.label));
|
||||
// Require an actual prerequisite offer, not a product question that
|
||||
// happens to mention the absence of an office-hours design document.
|
||||
const offer = /^No\s+design\s+doc\s+(?:found|exists)(?:\s+for\s+(?:this|the)\s+(?:branch|project))?\.\s*(?:\/office-hours\s+(?:produces|creates|provides)\s+(?:a\s+)?(?:structured\s+)?(?:design\s+doc(?:ument)?|problem\s+statement)(?:,?\s+(?:and\s+)?(?:premise\s+challenge|(?:explored\s+)?alternatives))*(?:\s*[—–-]\s*(?:sharper|better)\s+input\s+for\s+(?:the|this)\s+review)?\.\s*)?(?:Want\s+to\s+|Would\s+you\s+like\s+to\s+)?Run\s+(?:it|\/office-hours)\s+(?:now|first)(?:\s+or\s+proceed\s+with\s+standard\s+review)?\s*\?$/i.test(offerText);
|
||||
return prerequisite !== null && run.length === 1 && offer ? { kind: 'prerequisite', pick: prerequisite }
|
||||
: numberedPacketSetup(question, pending) ?? contextualPacketSetup(question, pending);
|
||||
});
|
||||
if (policies.some(policy => !policy) || new Set(policies.map(policy => policy!.kind)).size !== 2) return waiting;
|
||||
|
||||
const text = visible.replace(/\r+\n?/g, '\n').trimEnd();
|
||||
const bars = [...text.matchAll(/^ {0,3}←([^\n]*[☐☒][^\n]*)✔\s*Submit\s*→[\t ]*$/gm)];
|
||||
const footer = /Enter[\t ]+to[\t ]+select[\t ]*·[\t ]*Tab\/Arrow[\t ]+keys[\t ]+to[\t ]+navigate[\t ]*·[\t ]*Esc[\t ]+to[\t ]+cancel$/i;
|
||||
if (bars.length !== 1 || !footer.test(text)) return waiting;
|
||||
const bar = bars[0]!;
|
||||
const tabs = [...bar[1]!.matchAll(/([☐☒])\s*([^☐☒]+)/g)];
|
||||
const compact = (value: string) => value.replace(/\s+/g, '');
|
||||
const introduction = text.slice(0, bar.index).split('\n').findLast(line => !/^[\t ─━-]*$/.test(line)) ?? '';
|
||||
if (/\b(?:example|sample|quot(?:e[sd]?|ed)|template|source)\b[^\n]*:\s*$/i.test(introduction) ||
|
||||
compact(bar[1]!) !== tabs.map(tab => compact(tab[0])).join('') ||
|
||||
tabs.length !== pending.questions.length || tabs.some((tab, index) => compact(tab[2]!) !== compact(pending.questions[index]!.header)) ||
|
||||
/(?:^|\n)[^\n]*[1-9]\.\s*\[[ ✓✔xX]\]/.test(text)) return waiting;
|
||||
// Project only this actual pane's decoration for the existing full-panel
|
||||
// validator. The question, labels and native identity remain unchanged.
|
||||
const project = (header: string) => (text.slice(0, bar.index) + '☐ ' + header +
|
||||
text.slice(bar.index + bar[0].length).replace(footer, 'Enter to select · ↑/↓ to navigate · Esc to cancel'))
|
||||
.replace(/(^|\n)[\t ]*[│┃][\t ]?/g, '$1');
|
||||
if (!activeSetupPanel(project('Setup packet'))) return waiting;
|
||||
const packetKey = 'autoplan-setup-packet:' + JSON.stringify({sessionId:pending.sessionId,toolUseId:pending.toolUseId,questions:pending.questions});
|
||||
const choiceKey = (index: number) => `${packetKey}:choice:${index}:${policies[index]!.pick}`;
|
||||
const submitKey = packetKey + ':submit';
|
||||
if (planCountSubmissionInput(text) === '\r') {
|
||||
// Checked tabs are corroboration. Only choices this caller actually
|
||||
// sent for this same native packet can authorize its final submission.
|
||||
const options = parseNumberedOptions(text);
|
||||
if (!/Ready\s+to\s+submit\s+your\s+answers\?\s*❯\s*1\./.test(text) ||
|
||||
options.length !== 2 || options[0]?.index !== 1 || options[0]?.label !== 'Submit answers' ||
|
||||
options[1]?.index !== 2 || options[1]?.label !== 'Cancel' || seen.has(submitKey) ||
|
||||
tabs.some((tab, index) => tab[1] !== '☒' || !seen.has(choiceKey(index)))) return waiting;
|
||||
return { kind: 'input', input: '\r', signatures: [submitKey] };
|
||||
}
|
||||
|
||||
const captured = new Set(seen);
|
||||
const fp = capturePlanCountQuestion(text, captured, 0, true, pending);
|
||||
const index = fp?.nativeQuestionIndex;
|
||||
if (!fp || fp.nativeCall !== pending || index === undefined || tabs[index]?.[1] !== '☐' || seen.has(choiceKey(index))) return waiting;
|
||||
const question = pending.questions[index]!;
|
||||
if (!completeSetupOptions(project(question.header), { ...pending, questions: [question] })) return waiting;
|
||||
return { kind: 'input', input: planCountQuestionInput(text, fp, policies[index]!.pick),
|
||||
signatures: [...captured].filter(signature => !seen.has(signature)).concat(choiceKey(index)) };
|
||||
}
|
||||
|
||||
/** Pure classification: only the caller that sends input commits returned identities. */
|
||||
export function autoplanSetupDecision(visible: string, seen: ReadonlySet<string>, pending?: NativePlanQuestionCall): AutoplanSetupDecision {
|
||||
if (pending && (pending.answered || pending.failed || !pending.questions.length || pending.questions.some(question => question.multiSelect))) return { kind: 'waiting' };
|
||||
if (pending && pending.questions.length > 1) return setupPacketDecision(visible, seen, pending);
|
||||
// A visible packet without its complete native metadata cannot prove that
|
||||
// its other tabs are setup. Wait for persistence instead of guessing.
|
||||
if (/←[^\r\n]*[☐☒][^\r\n]*✔\s*Submit\s*→|Enter\s*to\s*select\s*·\s*Tab\/Arrow\s*keys\s*to\s*navigate/.test(visible)) return { kind: 'waiting' };
|
||||
// Box borders are terminal decoration, not part of an untagged native
|
||||
// question's wrapped text. Keep its full content for identity matching.
|
||||
const display = visible.replace(/(^|[\r\n])[\t ]*[│┃][\t ]?/g, '$1');
|
||||
// A rejected/mismatched menu has received no input. Keep its identities
|
||||
// available when the actual native call arrives after the visible prompt.
|
||||
const captured = new Set(seen);
|
||||
const question = capturePlanCountQuestion(display, captured, 0, true, pending);
|
||||
if (!question) return { kind: 'waiting' };
|
||||
// Recover only a still-visible direct routing title from this complete
|
||||
// boxed native panel. A present native call keeps its existing binding.
|
||||
const clippedTitle = !pending ? clippedRoutingTitle(visible, question) : null;
|
||||
if (clippedTitle) question.promptSnippet = clippedTitle;
|
||||
const answered = (input: string | null): AutoplanSetupDecision => input === null
|
||||
? { kind: 'waiting' }
|
||||
: { kind: 'input', input, signatures: [...captured].filter(signature => !seen.has(signature)) };
|
||||
|
||||
const prerequisite = planCountPrerequisitePick(question);
|
||||
const prerequisitePrompt = /\/office-hours/i.test(question.promptSnippet) &&
|
||||
/(?:no\s*design\s*doc|produce\s*a\s*design\s*doc)/i.test(question.promptSnippet);
|
||||
if (prerequisitePrompt) {
|
||||
// The canonical offer has two opposed choices. A known native call must
|
||||
// match this one-question panel; no mixed packet or substantive choice
|
||||
// may borrow its skip. Before JSONL flushes, require the complete native
|
||||
// single-select UI and the existing prerequisite premise/action guard.
|
||||
const choices = question.options.filter(option =>
|
||||
!/^(?:Type something\.|Chat about this)$/.test(option.label));
|
||||
if (choices.length !== 2 || (pending &&
|
||||
(!question.nativeCall || pending.questions.length !== 1 || pending.questions[0]?.multiSelect))) return { kind: 'waiting' };
|
||||
if (prerequisite === null) {
|
||||
// Mentioning a missing design doc or /office-hours in a product/taste
|
||||
// question does not make it setup. Independently identify an offer to
|
||||
// run that prerequisite even when its opposed skip is unsupported.
|
||||
const run = choices.filter(({ label }) => /^Run\s*\/office-hours\s*(?:now|first)(?:\s*\(recommended\))?$/i.test(label));
|
||||
// UI fingerprints abbreviate long questions; inspect the current
|
||||
// question before its cursor, with full-panel validation below.
|
||||
const beforeOptions = display.slice(0, display.search(/❯\s*1\./));
|
||||
const offer = /\brun\s*\/office-hours\s*(?:now|first)\s*\?(?:\s*<gstack-qid:[^>]+>)?\s*$/i.test(beforeOptions);
|
||||
return run.length === 1 && offer ? unsupportedSetup(display, question, 'prerequisite', pending) : { kind: 'unrelated' };
|
||||
}
|
||||
const input = planCountQuestionInput(display, question, prerequisite);
|
||||
if (!/^[1-9]\d*$/.test(input ?? '') || !activeSetupPanel(display)) return { kind: 'waiting' };
|
||||
return answered(input);
|
||||
}
|
||||
|
||||
// The model rephrases the setup question's closing sentence. Its routing
|
||||
// identity/premise and two opposed setup actions establish what is being
|
||||
// asked; an exact "Add them now?" sentence is not a stable interface.
|
||||
// Keep the actual question/premise separate from its header and later ELI10
|
||||
// prose, which may mention CLAUDE.md even on an unrelated question.
|
||||
const temporarySkipPanel = question.options.some(option => /^Skip\s*for\s*now(?:\s*\(Recommended\))?$/i.test(option.label))
|
||||
? completeSetupOptions(display, pending) : null;
|
||||
const actions = routingSetupActions(question, temporarySkipPanel?.length === 2);
|
||||
if (!actions) return { kind: 'unrelated' };
|
||||
// Lettered labels are a new action presentation, not permission to use
|
||||
// the older damaged-option fallback. Match the complete original menu.
|
||||
if (question.options.some(option => /^[A-Z]\)[\t ]+/i.test(option.label.trim())) &&
|
||||
completeSetupOptions(display, pending)?.length !== 2) return { kind: 'waiting' };
|
||||
const { add, decline } = actions;
|
||||
if (pending && (!question.nativeCall || pending.questions.length !== 1 || pending.questions[0]?.multiSelect)) return { kind: 'waiting' };
|
||||
// An intact Add-to-CLAUDE.md action identifies this setup offer even if
|
||||
// its opposed decline is unsupported. A qid or premise alone must not
|
||||
// turn substantive/ambiguous choices into an early setup failure.
|
||||
if (add.length !== 1) return { kind: 'waiting' };
|
||||
if (decline.length !== 1 || add[0]!.index === decline[0]!.index) return unsupportedSetup(display, question, 'routing', pending);
|
||||
// Newly admitted shorthand still needs the complete two-choice setup
|
||||
// panel; a qid alone cannot lend it stale or mismatched native metadata.
|
||||
if (/manualskills/i.test(decline[0]!.title) && completeSetupOptions(display, pending)?.length !== 2) return { kind: 'waiting' };
|
||||
// The verified clipped panel has the same native numeric shortcut. Do
|
||||
// not queue Enter behind it when the single-select header is offscreen.
|
||||
return answered(clippedTitle ? String(add[0]!.index) : planCountQuestionInput(display, question, add[0]!.index));
|
||||
}
|
||||
|
||||
/** Compatibility wrapper: preserve the existing input-only API. */
|
||||
export function autoplanRoutingSetupInput(visible: string, seen: Set<string>, pending?: NativePlanQuestionCall): string | null {
|
||||
const decision = autoplanSetupDecision(visible, seen, pending);
|
||||
if (decision.kind !== 'input') return null;
|
||||
for (const signature of decision.signatures) seen.add(signature);
|
||||
return decision.input;
|
||||
}
|
||||
|
||||
/** A long final gate may scroll its header away. This proves a wait, never an answer. */
|
||||
function croppedFinalApprovalPanel(visible: string, call: NativePlanQuestionCall): boolean {
|
||||
const question = call.questions[0]!;
|
||||
if (!/^Final (?:approval )?gate$/i.test(question.header) ||
|
||||
!/^(?:D\s*\d+\s*[—–:-]\s*)?Final approval(?: gate)?\s*:[^\n]*\?$/i.test(question.question.split('\n')[0]!) ||
|
||||
question.options.length < 2 || question.options.length > 7) return false;
|
||||
const lines = visible.replace(/\r+\n?/g, '\n').trimEnd().split('\n');
|
||||
const footer = 'Enter to select · ↑/↓ to navigate · Esc to cancel';
|
||||
if (lines.at(-1)?.trim() !== footer ||
|
||||
/[☐□☒]|✔\s*Submit|←|(?:^|\n)[^\n]*[1-9]\.\s*\[[ ✓✔xX]\]/.test(visible)) return false;
|
||||
const rowPattern = /^ {0,3}(❯\s*)?([1-9])\.[\t ]+(\S.*?)\s*$/;
|
||||
const rows = lines.flatMap((line, at) => {
|
||||
const match = rowPattern.exec(line);
|
||||
return match ? [{ at, cursor: !!match[1], index: Number(match[2]), label: match[3]! }] : [];
|
||||
});
|
||||
if (rows.length !== question.options.length + 2 || rows.some((row, i) => row.index !== i + 1) ||
|
||||
!rows[0]!.cursor || rows.slice(1).some(row => row.cursor) ||
|
||||
rows.at(-2)!.label !== 'Type something.' || rows.at(-1)!.label !== 'Chat about this') return false;
|
||||
const before = lines.slice(0, rows[0]!.at).filter(line => line.trim());
|
||||
if (!before.length || before.some(line => !/^ {0,3}[│┃][\t ]/.test(line))) return false;
|
||||
const excerptLines = before.map(line => line.replace(/^ {0,3}[│┃][\t ]?/, ''));
|
||||
if (excerptLines.some(line => /^\s*(?:`{3,}|~{3,}|>)/.test(line))) return false;
|
||||
const normalize = (text: string) => text.replace(/\s+/g, ' ').trim();
|
||||
// The renderer can truncate the last displayed line as well as crop the top.
|
||||
// Require one contiguous owned excerpt, including at least one complete native
|
||||
// line. Never assemble disconnected words or borrow a different question.
|
||||
const excerpt = normalize(excerptLines.join(' ')).replace(/…$/, '');
|
||||
const native = normalize(question.question), at = native.indexOf(excerpt);
|
||||
if (at <= 0 || native.indexOf(excerpt, at + 1) !== -1 ||
|
||||
!question.question.split('\n').slice(1).some(line => normalize(line) && excerpt.includes(normalize(line)))) return false;
|
||||
for (const [i, option] of question.options.entries()) {
|
||||
if (!option.label.trim() || normalize(rows[i]!.label) !== normalize(option.label)) return false;
|
||||
const description = lines.slice(rows[i]!.at + 1, rows[i + 1]!.at);
|
||||
if (description.some(line => line.trim() && !/^ {4,}\S/.test(line)) ||
|
||||
normalize(description.join(' ')) !== normalize(option.description ?? '')) return false;
|
||||
}
|
||||
// Only native footer decoration may follow the two utility rows.
|
||||
const decoration = (line: string) => /^[\t ─━-]*$/.test(line);
|
||||
return lines.slice(rows.at(-2)!.at + 1, rows.at(-1)!.at).every(decoration) &&
|
||||
lines.slice(rows.at(-1)!.at + 1, -1).every(decoration);
|
||||
}
|
||||
|
||||
/** Identify a remaining native human wait. The caller must treat it as failure, never phase credit. */
|
||||
export function autoplanBlockingQuestionBoundary(visible: string, context: {
|
||||
commandStartedAt: number; viewportCapturedAt: number;
|
||||
/** Both projections must come from the same owned readPlanCountTranscript poll. */
|
||||
transcript: PlanCountTranscript; publicTools: NativePublicToolEvent[];
|
||||
/** Only the current return from the owned, post-command readPendingQuestion. */
|
||||
pending?: ReturnType<typeof readPendingQuestion>;
|
||||
}): { sessionId: string; toolUseId: string; source: 'native' | 'pre_tool_use' } | null {
|
||||
const {transcript, commandStartedAt, viewportCapturedAt, publicTools} = context;
|
||||
if (transcript.status !== 'ready' || !Number.isFinite(commandStartedAt) ||
|
||||
!Number.isFinite(viewportCapturedAt) || commandStartedAt < 0 || viewportCapturedAt < commandStartedAt) return null;
|
||||
const sessions = new Set([...transcript.calls.map(call => call.sessionId),
|
||||
...transcript.assistantMessages.map(message => message.sessionId)]);
|
||||
if (sessions.size !== 1 || ![...sessions][0]) return null;
|
||||
const unanswered = transcript.calls.filter(call => !call.answered && !call.failed);
|
||||
if (unanswered.length > 1) return null;
|
||||
const call = unanswered[0] ?? context.pending;
|
||||
if (!call || !sessions.has(call.sessionId) || !call.toolUseId || call.answered !== false ||
|
||||
call.failed !== false || call.questions.length !== 1 || call.questions[0]!.multiSelect) return null;
|
||||
const identity = (questions: unknown): string | null => Array.isArray(questions) && questions.every(q =>
|
||||
q && typeof q.header === 'string' && typeof q.question === 'string' && Array.isArray(q.options) &&
|
||||
q.options.every((o: any) => o && typeof o.label === 'string' &&
|
||||
(o.description === undefined || typeof o.description === 'string')))
|
||||
? JSON.stringify(questions.map(q => [q.header,q.question,q.multiSelect ?? false,
|
||||
q.options.map((o: any) => [o.label,o.description ?? ''])])) : null;
|
||||
if (identity(call.questions) === null) return null;
|
||||
const uses = publicTools.filter(event => event.sessionId === call.sessionId && event.toolUseId === call.toolUseId && event.kind === 'use');
|
||||
if (publicTools.some(event => event.sessionId === call.sessionId && event.toolUseId === call.toolUseId && event.kind === 'result')) return null;
|
||||
let source: 'native' | 'pre_tool_use';
|
||||
if (unanswered.length) {
|
||||
const use = uses[0], at = Date.parse(use?.timestamp ?? '');
|
||||
if (uses.length !== 1 || use?.name !== 'AskUserQuestion' || !Number.isFinite(at) ||
|
||||
at < commandStartedAt || at > viewportCapturedAt || !Array.isArray(use.input?.questions) ||
|
||||
identity(use.input.questions as NativePlanQuestion[]) !== identity(call.questions)) return null;
|
||||
source = 'native';
|
||||
} else {
|
||||
// The reader has already checked cwd/config, parent session, timestamp,
|
||||
// recorder poison/lock state and absence of a published result or call.
|
||||
if (context.pending?.source !== 'pre_tool_use' || uses.length ||
|
||||
transcript.calls.some(row => row.sessionId === call.sessionId && row.toolUseId === call.toolUseId)) return null;
|
||||
source = 'pre_tool_use';
|
||||
}
|
||||
const display = visible.replace(/(^|[\r\n])[\t ]*[│┃][\t ]?/g, '$1');
|
||||
const headerAt = display.search(/^ {0,3}[☐□]/m);
|
||||
if (headerAt < 0) {
|
||||
// Crop recovery is limited to an already published native use. The pending
|
||||
// hook route still requires its original complete current panel.
|
||||
if (source !== 'native' || !croppedFinalApprovalPanel(visible, call)) return null;
|
||||
} else if (/^(?:Source|Example|Quoted|Historical|Template)\b[^\n]*:/im.test(display.slice(0, headerAt)) ||
|
||||
!completeSetupOptions(display, call)) return null;
|
||||
return {sessionId:call.sessionId,toolUseId:call.toolUseId,source};
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
import path from 'node:path';
|
||||
|
||||
/** Rebase recorded Unix paths without inserting raw backslashes into JSON. */
|
||||
export function capturedPathRebaser(replacements: Array<[string, string]>) {
|
||||
const text = (value: string) => replacements.reduce((value, [before, after]) =>
|
||||
value.replaceAll(before, after.split(path.sep).join('/')), value);
|
||||
// Translate separators without resolving traversal, redundant separators, or
|
||||
// relative spelling that an ownership rejection control must still observe.
|
||||
const file = (value: string) => text(value).split('/').join(path.sep);
|
||||
const pathKeys = new Set(['cwd', 'config', 'stateRoot', 'file', 'file_path', 'transcriptPath']);
|
||||
const json = <T>(value: T): T => JSON.parse(JSON.stringify(value), (key, value) =>
|
||||
typeof value === 'string' ? (pathKeys.has(key) ? file(value) : text(value)) : value);
|
||||
return { text, file, json };
|
||||
}
|
||||
@@ -163,7 +163,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// wave's headline capability) grows the union to 1.195x. Deliberate:
|
||||
// the section is on-demand (loads only for Apple store targets), so
|
||||
// per-invocation cost for non-iOS ships is one manifest line.
|
||||
maxSizeRatio: 1.22,
|
||||
maxSizeRatio: 1.28, // Harness-aware dispatch adds validated commands and per-pass provenance (~1.25x).
|
||||
},
|
||||
'plan-ceo-review': {
|
||||
skill: 'plan-ceo-review',
|
||||
@@ -181,7 +181,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
|
||||
// Fork port wave 2 (#703): the repo-doc-preference block in the design
|
||||
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
|
||||
maxSkeletonBytes: 79_000, // + v2.0 {{ASIDE_RESEARCH}} (Aside first, WebSearch fallback); measured 77_657
|
||||
maxSkeletonBytes: 79_744, // Exact 744-byte anti-shortcut move: main 79,739 + section 75,559 = unchanged 155,298-byte union; retains 5-byte slack.
|
||||
minUnionBytes: 123_600, // token-reduction Phases 1-2 (v1.69.x branch): preamble bash -> bin/gstack-skill-start, onboarding -> gated emission; measured union 137,346
|
||||
mustContain: ['SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'HOLD SCOPE', 'SCOPE REDUCTION'],
|
||||
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
||||
@@ -207,7 +207,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
|
||||
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
|
||||
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
|
||||
maxSkeletonBytes: 56_500, // + v2.0 {{ASIDE_RESEARCH}} (Aside first, WebSearch fallback); measured 55_457
|
||||
maxSkeletonBytes: 57_200, // Eng per-issue approval exit check, including regression-test authority; measured 57,113 bytes.
|
||||
minUnionBytes: 99_800, // token-reduction Phases 1-2 (v1.69.x branch); measured union 110,910
|
||||
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
|
||||
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
|
||||
@@ -240,7 +240,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// tier-2+ skeleton (measured 89,184). Main's v1.64.0.0 adds ~340 B more
|
||||
// (telemetry --error-message/--failed-step preamble prose, PR #769).
|
||||
// Budget covers the sum of both waves.
|
||||
maxSkeletonBytes: 73_800, // + v1.78 AUQ objectivity + v1.79 foreground-dispatch sweep (merged); measured 73_398
|
||||
maxSkeletonBytes: 79_500, // Harness-aware outside voice: validated dispatch and provenance.
|
||||
minUnionBytes: 99_200, // token-reduction Phases 1-2 (v1.69.x branch); measured union 110,293
|
||||
mustContain: ['design', 'visual'],
|
||||
maxSizeRatio: 1.12, // D1 1.104 + main's ~0.008
|
||||
@@ -295,7 +295,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// the #538 opt-out + D1 evidence directive — ratio 1.104 measured.
|
||||
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
|
||||
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
|
||||
maxSkeletonBytes: 76_800, // + v2.0 {{ASIDE_RESEARCH}} (Aside first, WebSearch fallback); measured 75_804
|
||||
maxSkeletonBytes: 87_500, // Office-hours + sketch outside voices include host guards and completion checks.
|
||||
minUnionBytes: 115_800, // Phase 4 wave 4; measured union 118,175
|
||||
mustContain: ['design doc', 'problem statement'],
|
||||
maxSizeRatio: 1.12,
|
||||
@@ -347,7 +347,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
|
||||
// v1.64.1.0: shared-preamble prose from the two parallel v1.64 waves lands
|
||||
// the skeleton at 69,022 B; +~1 KB headroom.
|
||||
maxSkeletonBytes: 67_500, // + v1.82 open DESIGN.md format check ({{DESIGN_MD_CHECK}} in Phase 0); measured 67_014
|
||||
maxSkeletonBytes: 72_500, // Outside review + upstream DESIGN.md format check; merged render 72,175 bytes.
|
||||
minUnionBytes: 65_000, // token-reduction Phases 1-2 (v1.69.x branch): preamble bash -> bin/gstack-skill-start, onboarding -> gated emission; measured union 72,252
|
||||
mustContain: ['Typography', 'Color', 'Aesthetic Direction'],
|
||||
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback ~2KB +
|
||||
@@ -440,7 +440,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
'## Filesystem Boundary',
|
||||
'Synthesis recommendation (REQUIRED)',
|
||||
'Recommendation: <action> because',
|
||||
'UNDER_CODEX',
|
||||
'If the runtime guard reports a harness mismatch, stop.',
|
||||
],
|
||||
mustPrecedeStop: ['## Step 1: Detect mode', '## Filesystem Boundary'],
|
||||
mustMoveToSection: [
|
||||
@@ -495,7 +495,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
'## Sequential Execution — MANDATORY',
|
||||
'## Decision Classification',
|
||||
'## Filesystem Boundary — Codex Prompts',
|
||||
'## Phase 0.5: Codex auth + version preflight',
|
||||
'## Phase 0.5: Outside reviewer preflight',
|
||||
'## Pre-Gate Verification',
|
||||
'## Phase 2: Design Review (conditional — skip if no UI scope)',
|
||||
'## Phase 2.5: DX Review (conditional — skip if no developer-facing scope)',
|
||||
@@ -504,7 +504,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
mustPrecedeStop: ['## The 6 Decision Principles', '## Sequential Execution — MANDATORY', '## Decision Classification'],
|
||||
mustMoveToSection: [
|
||||
'CEO DUAL VOICES — CONSENSUS TABLE:',
|
||||
'CODEX SAYS (design — UX challenge)',
|
||||
'Codex SAYS (design — UX challenge)',
|
||||
'ENG DUAL VOICES — CONSENSUS TABLE:',
|
||||
'DX DUAL VOICES — CONSENSUS TABLE:',
|
||||
'## Implementation Tasks aggregator',
|
||||
@@ -513,9 +513,10 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
},
|
||||
behavioral: 'external',
|
||||
externalTest: 'test/skill-e2e-autoplan-chain.test.ts', // phase-complete markers live ONLY in sections — its assertions ARE section-read proof
|
||||
maxSkeletonBytes: 65_100, // + v1.78 AUQ objectivity + #2745 broken-install preflight arm + outside-voice honest labeling; measured 64_668
|
||||
maxSkeletonBytes: 70_000, // Phase-specific outside coverage, native fallback, and harness guard.
|
||||
minUnionBytes: 85_000, // measured union 86,926
|
||||
mustContain: ['6 Decision Principles', 'TASTE DECISION', 'USER CHALLENGE', 'consensus', 'Restore Point'],
|
||||
maxSizeRatio: 1.12, // Four validated outside invocations replace raw CLI calls; phases keep independent coverage.
|
||||
},
|
||||
spec: {
|
||||
skill: 'spec',
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
import { pickCeoCompletionHandoff } from './ceo-completion-handoff';
|
||||
import { findCeoModeOption } from './ceo-mode-option';
|
||||
|
||||
/** Follow the offered recommendation only in the native pre-review approach menu. */
|
||||
export function pickCeoRecommendedApproach(fp: AskUserQuestionFingerprint): number | null {
|
||||
const call = fp.nativeCall;
|
||||
if (!fp.preReview || !call || call.answered !== false || call.failed !== false ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` || call.questions.length !== 1 ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return null;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2 || !/^Approach$/i.test(q.header.trim())) return null;
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:\s*([a-z0-9-]+)\s*>/gi)];
|
||||
if (ids.length !== 1 || (q.question.match(/<gstack-qid/gi)?.length ?? 0) !== 1) return null;
|
||||
const qid = ids[0]![1]!.toLowerCase();
|
||||
const decision = /^D\s*([1-9]\d*)\s*[—–:-]\s*/i.exec(q.question.trim());
|
||||
const approachId = /^plan-ceo(?:-review)?-approach(?:-selection)?(?:-d([1-9]\d*))?$/.exec(qid);
|
||||
// Native routing ids may carry the explicit decision number. An id for a
|
||||
// different decision cannot borrow this question's recommendation policy.
|
||||
const planApproachId = approachId !== null &&
|
||||
(!approachId[1] || approachId[1] === decision?.[1]);
|
||||
const question = q.question.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '');
|
||||
// Native approach menus vary their routing id and use/follow wording. Match
|
||||
// the selector's structure, retaining the explicit recommendation below.
|
||||
const planApproach = planApproachId &&
|
||||
/^Which implementation approach should this plan (?:use|follow)\?(?:\s|$)/i.test(question);
|
||||
const testApproach = qid === 'plan-ceo-review-impl-approach' &&
|
||||
/^Which implementation approach for the [a-z_$][\w$]*(?:\.[a-z_$][\w$]*)*\(\) tests\?(?:\s|$)/i.test(question);
|
||||
// A named component can be the subject instead of "this plan". Consume the
|
||||
// complete direct question; the routing id alone cannot authorize a choice.
|
||||
const directQuestion = question.replace(/\s*<gstack-qid:[^>]+>\s*$/i, '').trim();
|
||||
const component = /^Which implementation approach for (?:the|this) ((?:[a-z_$][\w$.-]*\s+){0,5})(?:handler|endpoint|service|module|component|adapter|client|worker|pipeline|integration)\?$/i.exec(directQuestion);
|
||||
const componentApproach = planApproachId &&
|
||||
component !== null && !/\b(?:and|or|then)\b/i.test(component[1]!);
|
||||
if (!planApproach && !testApproach && !componentApproach) return null;
|
||||
if (fp.options.length !== q.options.length || !fp.options.every((option, i) =>
|
||||
option.index === i + 1 && option.label === q.options[i]!.label)) return null;
|
||||
const labels = q.options.map(option => option.label.trim());
|
||||
if (new Set(labels).size !== labels.length) return null;
|
||||
const recommended = labels.map((label, i) => ({ label, index: i + 1 })).filter(({ label }) =>
|
||||
/\s\(Recommended\)\s*$/i.test(label) &&
|
||||
(label.match(/\brecommended\b/gi)?.length ?? 0) === 1 &&
|
||||
!/\b(?:not|never)\s*\(recommended\)/i.test(label));
|
||||
return recommended.length === 1 ? recommended[0]!.index : null;
|
||||
}
|
||||
|
||||
/** Fixed-count fixtures review existing defects without opting into expansions. */
|
||||
function pickCeoCountMode(fp: AskUserQuestionFingerprint): number | null {
|
||||
const call = fp.nativeCall;
|
||||
if (!fp.preReview || !call || call.answered !== false || call.failed !== false ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` || call.questions.length !== 1 ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return null;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || !/^(?:Review )?mode$/i.test(q.header.trim()) || q.options.length !== 4 ||
|
||||
fp.options.length !== 4 || !fp.options.every((option, i) =>
|
||||
option.index === i + 1 && option.label === q.options[i]!.label)) return null;
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:\s*([a-z0-9-]+)\s*>/gi)];
|
||||
if (ids.length !== 1 || (q.question.match(/<gstack-qid/gi)?.length ?? 0) !== 1 ||
|
||||
!/^(?:plan-)?ceo-(?:review-)?mode(?:-selection)?$/i.test(ids[0]![1]!)) return null;
|
||||
const question = q.question.trim().replace(/^D\s*[1-9]\d*\s*[—–:-]\s*/i, '')
|
||||
.replace(/\s*<gstack-qid:[^>]+>\s*$/i, '');
|
||||
if (!/^Which review mode should I (?:use|apply)(?: for this (?:test coverage )?plan)?\?$/i.test(question)) return null;
|
||||
try {
|
||||
// Require all four modes exactly once. Reuse the mode-routing parser so
|
||||
// displayed option order and side-panel text cannot pick a different mode.
|
||||
const positions = (['HOLD SCOPE', 'SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'SCOPE REDUCTION'] as const)
|
||||
.map(mode => findCeoModeOption(fp.options, mode));
|
||||
return positions.every(position => position !== null) && new Set(positions).size === 4
|
||||
? positions[0]! : null;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/** Preserve the existing manual handoff and all other caller/default choices. */
|
||||
export function pickCeoCountQuestion(
|
||||
fp: AskUserQuestionFingerprint,
|
||||
activeCapture: AskUserQuestionFingerprint = fp,
|
||||
): number | null {
|
||||
return pickCeoCountMode(activeCapture) ?? pickCeoRecommendedApproach(activeCapture) ?? pickCeoCompletionHandoff(fp, activeCapture);
|
||||
}
|
||||
@@ -0,0 +1,477 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
|
||||
/** A native choice may carry the CEO-specific recap beside a generic completion question. */
|
||||
function closedCeoRecap(description: string): boolean {
|
||||
const clause = /(?:^|[.!?]\s+)((?:The\s+)?CEO\s+review\b[^.!?]{0,240})(?=[.!?]|$)/i.exec(description)?.[1];
|
||||
if (!clause || /\b(?:if|unless|until|once|when|after|not|never)\b|n['’]t\b/i.test(clause)) return false;
|
||||
return /\b(?:all(?:\s+(?:gaps?|issues?|findings?))?(?:\s+(?:are|were))?\s+resolved|(?:no|0)\s+unresolved\s+(?:decisions|gaps|issues|findings))(?=\s*\)?\s*(?:;|$))/i.test(clause);
|
||||
}
|
||||
|
||||
/** Past-tense resolution can close a native next-review recap without the word "complete". */
|
||||
function resolvedCeoRecap(description: string): boolean {
|
||||
const clause = /(?:^|[.!?]\s+)((?:(?:This|The)\s+)?CEO\s+review\s+resolved\s+[^.!?;]{1,180}\b(?:bugs|gaps|issues|findings))(?=\s*(?:[.!?;]|$))/i.exec(description)?.[1];
|
||||
return Boolean(clause && !/\b(?:if|unless|until|once|when|after|not|never|some|most|partially|only|of|but|several|few)\b|n['’]t\b/i.test(clause));
|
||||
}
|
||||
|
||||
/** A closed-review declaration plus one direct navigation query, even when its recap follows it. */
|
||||
function closedReviewNavigation(declaration: string, context: string): boolean {
|
||||
const question = declaration.replace(/<gstack-qid:[^>]+>/gi, '');
|
||||
return /^CEO review (?:is )?(?:complete|done|cleared|clean)[.!](?:\s|$)/i.test(question) &&
|
||||
/(?:^|[.!]\s+)What(?:['’]s)? next\?(?:\s|$)/i.test(question) && closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** The metadata recap is native question text, not a new substantive choice. */
|
||||
function isMetadataNavigationQuestion(declaration: string): boolean {
|
||||
return /^What(?:['’]s|\s+is)\s+(?:the\s+)?next(?:\s+(?:step|review))?\s+after\s+(?:this|the)\s+CEO\s+review\?\s*$/i.test(declaration.trim().split('\n')[0]!);
|
||||
}
|
||||
|
||||
function metadataClosedReviewNavigation(declaration: string, context: string): boolean {
|
||||
const question = declaration.replace(/<gstack-qid:[^>]+>/gi, '').trim();
|
||||
return isMetadataNavigationQuestion(question) &&
|
||||
/^[ \t]{0,3}ELI10:\s*(?:The\s+)?CEO\s+review\s+(?:is\s+)?(?:complete|cleared|clean|done(?:\s+and\s+clear(?:ed)?)?)[.!](?:\s|$)/im.test(question) &&
|
||||
!/`{3}|~{3}|(?:^|[.!?]\s+)[ \t]*>|\b(?:example|quoted source)\s*:/im.test(context) &&
|
||||
!/\b(?:incomplete|unfinished)\b|\b(?:review|decisions|findings|issues|gaps)\b[^.!?\n]{0,60}\b(?:not|never)\b|\b(?:isn['’]t|aren['’]t|wasn['’]t|weren['’]t)\b/i.test(context) &&
|
||||
!/(?:^|[.!?;:]\s+|\b(?:proceed to|continue to|should|must|will|need to|can|could|would|may|might)\s+)(?:(?:please|first|then|also)\s+)*(?:add|fix|repair|implement|resolve|decide)\b/im.test(context) &&
|
||||
closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** Scope/risk explanations can contain "if" and "not" without reopening CEO work. */
|
||||
function explainedMetadataNavigation(declaration: string, descriptions: string[]): boolean {
|
||||
const question = declaration.replace(/<gstack-qid:[^>]+>/gi, '').trim();
|
||||
if (!isMetadataNavigationQuestion(question)) return false;
|
||||
const sentences = [question.split('\n').slice(1).join('\n'), ...descriptions]
|
||||
.flatMap(text => text.trim().split(/[.!](?:\s+|$)/).map(sentence => sentence.trim()).filter(Boolean));
|
||||
const resolved = /^(?:\d+|one|two|three|four|five|six|seven|eight|nine|ten) assertion spec gaps were caught and resolved$/i;
|
||||
const noDesignScope = /^No UI scope was detected, so a design review is not needed$/i;
|
||||
const stakes = /^Stakes if we pick wrong: skipping the eng review means shipping without an architecture \+ code quality pass$/i;
|
||||
// Validate every whole sentence before discounting the two inert phrases.
|
||||
// Additional repair, conditional closure, or a different review denial is
|
||||
// substantive even when it follows a valid metadata heading or recap.
|
||||
if (sentences.filter(sentence => resolved.test(sentence)).length !== 1 ||
|
||||
!sentences.every(sentence => resolved.test(sentence) || noDesignScope.test(sentence) || stakes.test(sentence) ||
|
||||
/^ELI10: The CEO review is (?:done|complete|cleared|clean)$/i.test(sentence) ||
|
||||
/^The plan is now ready for the Eng Review, which is the required gate before shipping$/i.test(sentence) ||
|
||||
/^For test code this is lower risk than production code, but the eng review also validates that the test infrastructure is used correctly$/i.test(sentence) ||
|
||||
/^Recommendation: [A-Z] because eng review is the required shipping gate, and this plan is ready for it$/i.test(sentence) ||
|
||||
/^Required gate$/i.test(sentence) ||
|
||||
/^Validates architecture, test infrastructure usage, code quality, and that the \d+-test plan will be implementable without hidden issues$/i.test(sentence) ||
|
||||
/^Proceed to implementation without the eng review$/i.test(sentence) ||
|
||||
/^Lower confidence that the test infrastructure is wired correctly, but acceptable for low-risk test coverage work$/i.test(sentence))) return false;
|
||||
const normalized = [question.split('\n')[0]!, ...sentences
|
||||
.filter(sentence => !noDesignScope.test(sentence))
|
||||
.map(sentence => stakes.test(sentence) ? sentence.replace(/^Stakes if we pick wrong:/i, 'Stakes:') : sentence)]
|
||||
.join('\n');
|
||||
return metadataClosedReviewNavigation(question, normalized);
|
||||
}
|
||||
|
||||
/** A direct Eng/manual choice can put its unconditional CEO recap in a native description. */
|
||||
function describedEngNavigation(question: string, descriptions: string[], context: string): boolean {
|
||||
if (!/^Run\s+\/plan-eng-review\s+(?:next|now)\s*\((?:the\s+)?required(?:\s+shipping)?\s+gate\),?\s+or\s+handle\s+reviews\s+manually\?$/i.test(question)) return false;
|
||||
const recap = /^(?:The\s+)?CEO\s+review\s+is\s+(?:clear|complete|cleared|clean|done)(?:\s+but\s+eng\s+review\s+is\s+the\s+(?:required\s+)?shipping\s+gate)?$/i;
|
||||
const topics = String.raw`(?:test isolation|factory patterns|test coverage|architecture|dependencies)`;
|
||||
const reviewExplanation = new RegExp(String.raw`^(?:Validates|Checks|Reviews)\s+${topics}(?:,\s+${topics})*(?:,?\s+and\s+(?:${topics}|confirms no hidden dependencies))?$`, 'i');
|
||||
const sentences = descriptions.flatMap(description => description.trim().split(/[.!](?:\s+|$)/).map(sentence => sentence.trim()).filter(Boolean));
|
||||
const closedRecap = sentences.some(sentence => recap.test(sentence));
|
||||
// Every sentence must explain this closed handoff. Arbitrary prose after
|
||||
// a valid recap could add work (including verbs no blacklist anticipates).
|
||||
if (!sentences.every(sentence => recap.test(sentence) || reviewExplanation.test(sentence) ||
|
||||
/^Required(?:\s+shipping)?\s+gate\s+before\s+(?:shipping|merging|implementation)$/i.test(sentence) ||
|
||||
/^(?:You['’]ll|You will)\s+need\s+to\s+run\s+\/plan-eng-review\s+(?:separately\s+)?before\s+(?:merging|shipping)$/i.test(sentence) ||
|
||||
/^Run\s+\/plan-eng-review\s+(?:next|now|before\s+(?:merging|shipping)|after\s+implementation\s+and\s+before\s+shipping)$/i.test(sentence) ||
|
||||
/^(?:Fast|Quick|Short)\s+(?:run|review)\s+expected\s+given\s+(?:zero|no|0)\s+CEO\s+findings$/i.test(sentence))) return false;
|
||||
if (!closedRecap || /`{3}|~{3}|(?:^|\n)[ \t]*>|\b(?:example|quoted source)\s*:/im.test(context) ||
|
||||
/\b(?:incomplete|unfinished|not|never)\b|n['’]t\b/i.test(context)) return false;
|
||||
// "Clear" is also a closure claim here; a future condition cannot supply it.
|
||||
const clearClosure = String.raw`(?:(?:the\s+)?CEO|the)\s+review\s+(?:(?:is|was|becomes?|became|(?:will|would|can|could|may|might)\s+(?:be|become))\s+)?clear`;
|
||||
if (new RegExp(String.raw`\b(?:once|when|after)\b[^.!?]{0,180}\b${clearClosure}\b|\b${clearClosure}\b[^.!?]{0,100}\b(?:once|when|after)\b`, 'i').test(context)) return false;
|
||||
return !/(?:^|[.!?;:]\s+|\b(?:proceed to|continue to|should|must|will|need to|can|could|would|may|might)\s+)(?:(?:please|first|then|also)\s+)*(?:add|fix|repair|implement|resolve|decide)\b/im.test(context) &&
|
||||
closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** The next sentence may name the required gate with "it" after the Eng query. */
|
||||
function pronounEngGate(question: string, descriptions: string[], context: string): boolean {
|
||||
if (!/^(?:The\s+)?CEO\s+review\s+is\s+(?:complete|cleared|clean|done)[.!]\s+Run\s+\/plan-eng-review\s+next\?\s+It(?:['’]s|\s+is)\s+the\s+required(?:\s+shipping)?\s+gate\.$/i.test(question)) return false;
|
||||
const topics = String.raw`(?:architecture|security|test quality|performance)`;
|
||||
const covers = new RegExp(String.raw`^Covers\s+${topics}(?:,\s+${topics})*(?:,?\s+and\s+${topics})?$`, 'i');
|
||||
const sentences = descriptions.flatMap(description => description.trim().split(/[.!](?:\s+|$)/)
|
||||
.map(sentence => sentence.trim()).filter(Boolean));
|
||||
return sentences.every(sentence => covers.test(sentence) ||
|
||||
/^Required\s+gate\s+before\s+shipping$/i.test(sentence) ||
|
||||
/^This\s+CEO\s+review\s+found\s+no\s+architecture\s+concerns,\s+so\s+eng\s+review\s+should\s+be\s+fast$/i.test(sentence) ||
|
||||
/^Proceed\s+without\s+the\s+eng\s+review\s+gate$/i.test(sentence) ||
|
||||
/^You\s+own\s+ensuring\s+correctness\s+before\s+shipping$/i.test(sentence)) &&
|
||||
closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** A next-review question may explain completed CEO work only in its choices. */
|
||||
function describedPostReviewNavigation(question: string, descriptions: string[], context: string): boolean {
|
||||
if (!/^What(?:['’]s|\s+is)\s+the\s+next\s+review\s+step\s+after\s+(?:this|the)\s+CEO\s+review\?$/i.test(question) ||
|
||||
/`{3}|~{3}|(?:^|\n)[ \t]*>|\b(?:example|quoted source)\s*:/im.test(context)) return false;
|
||||
const closed = /^(?:The|This) CEO review resolved all findings, but the eng review validates the approach at a lower implementation level$/i;
|
||||
const topics = String.raw`(?:[\w-]+ integration|parameterized queries|async [\w-]+ queue)`;
|
||||
const changedApproach = new RegExp(String.raw`^This CEO review changed the implementation approach \(Approach [A-Z]: ${topics}(?:, ${topics})*\) [—–-] a fresh eng review should validate the new approach before implementation begins$`, 'i');
|
||||
const sentences = descriptions.flatMap(description => description.trim().split(/[.!](?:\s+|$)/)
|
||||
.map(sentence => sentence.trim()).filter(Boolean));
|
||||
// Whole sentences keep extra work out of the recap, including actions that
|
||||
// an imperative-verb blacklist would miss. Only the next review is offered.
|
||||
return sentences.filter(sentence => closed.test(sentence)).length === 1 &&
|
||||
sentences.every(sentence => closed.test(sentence) || changedApproach.test(sentence) ||
|
||||
/^Eng review is the required shipping gate$/i.test(sentence) ||
|
||||
/^It covers architecture details, code quality, and test verification$/i.test(sentence) ||
|
||||
/^Proceed to implementation without the eng review gate$/i.test(sentence) ||
|
||||
/^Skipping is not recommended for a handler that processes payment webhooks$/i.test(sentence)) &&
|
||||
closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** A resolved-gap count may qualify completion before the required next gate. */
|
||||
function countedCeoNavigation(question: string, descriptions: string[]): boolean {
|
||||
if (!/^(?:The )?CEO review is complete \(0 critical gaps, [1-9]\d* (?:spec )?gaps resolved\)\. Eng Review is the required shipping gate\. What['’]s next\?$/i.test(question)) return false;
|
||||
const topics = String.raw`(?:architecture|code quality|tests|performance)`;
|
||||
const covers = new RegExp(String.raw`^Covers ${topics}(?:, ${topics})*(?:,? and ${topics})?$`, 'i');
|
||||
return descriptions.flatMap(description => description.trim().split(/[.!](?:\s+|$)/)
|
||||
.map(sentence => sentence.trim()).filter(Boolean)).every(sentence => covers.test(sentence) ||
|
||||
/^Required gate before shipping$/i.test(sentence) ||
|
||||
/^This is a test-only plan so eng review should be fast$/i.test(sentence) ||
|
||||
/^You manage the eng review yourself$/i.test(sentence) ||
|
||||
/^The dashboard will show NOT CLEARED until it runs$/i.test(sentence));
|
||||
}
|
||||
|
||||
/** An unconditional CLEAR recap followed by one direct required-Eng query. */
|
||||
function clearRequiredEngNavigation(question: string, descriptions: string[], context: string): boolean {
|
||||
if (!/^(?:The\s+)?CEO\s+review\s+is\s+CLEAR\.\s+Eng\s+review\s+is\s+the\s+required\s+shipping\s+gate\s+[—–-]\s+run\s+it\s+next\?$/i.test(question) ||
|
||||
descriptions.some(description => !description.trim())) return false;
|
||||
const topics = String.raw`(?:architecture|code quality|tests|performance)`;
|
||||
const topicsReview = new RegExp(String.raw`^${topics}(?:,\s+${topics})*(?:,?\s+and\s+${topics})?\s+review$`, 'i');
|
||||
const resolved = /^This\s+CEO\s+review\s+held\s+scope\s+and\s+resolved\s+[1-9]\d*\s+assertion\s+gaps\s+[—–-]\s+eng\s+review\s+verifies\s+the\s+test\s+structure\s+is\s+sound$/i;
|
||||
const sentences = descriptions.flatMap(description => description.trim().split(/[.!](?:\s+|$)/)
|
||||
.map(sentence => sentence.trim()).filter(Boolean));
|
||||
// CLEAR is accepted only with this complete navigation grammar. Do not add
|
||||
// it to the permissive legacy completion regex or discard appended prose.
|
||||
return sentences.filter(sentence => resolved.test(sentence)).length === 1 &&
|
||||
sentences.every(sentence => resolved.test(sentence) || topicsReview.test(sentence) ||
|
||||
/^Required\s+gate\s+before\s+shipping$/i.test(sentence) ||
|
||||
/^You\s+manage\s+the\s+review\s+pipeline\s+yourself$/i.test(sentence) ||
|
||||
/^Note:\s+eng\s+review\s+is\s+required\s+to\s+CLEAR\s+for\s+\/ship$/i.test(sentence)) &&
|
||||
closedNavigationContext(context);
|
||||
}
|
||||
|
||||
/** A bare next-workflow choice is administration, never proof of completed review. */
|
||||
function bareEngNavigation(question: string, descriptions: string[]): boolean {
|
||||
if (!/^run \/plan-eng-review\?$/i.test(question) || descriptions.some(s => !s.trim())) return false;
|
||||
const sentences = descriptions.flatMap(s => s.trim().split(/\n+|[.!](?:\s+|$)/))
|
||||
.map(s => s.trim().replace(/^\[[+-]\]\s*/, '')).filter(Boolean);
|
||||
const approved = /^Proceed directly to implementation with the approved changes from this CEO review$/i;
|
||||
const gate = /^Eng Review is the required shipping gate$/i;
|
||||
// Consume the complete offered context. Past findings and already-approved
|
||||
// changes are recaps; an added remedy or unfinished-review choice is not.
|
||||
return sentences.some(s => approved.test(s)) && sentences.some(s => gate.test(s)) &&
|
||||
sentences.every(s => approved.test(s) || gate.test(s) ||
|
||||
/^It covers architecture depth, code quality, test gaps, and performance [—–-] complementing what this CEO review found$/i.test(s) ||
|
||||
/^Since this CEO review expanded the plan \(added [a-z0-9_ +/-]{1,120} requirements\), a fresh eng review is especially valuable$/i.test(s) ||
|
||||
/^Required before shipping; catches implementation issues the plan-level review cannot$/i.test(s) ||
|
||||
/^This CEO review found critical issues \([a-z0-9_ +/-]{1,80}\) [—–-] eng review will verify the fix approach is architecturally sound$/i.test(s) ||
|
||||
/^Adds another review session before implementation starts$/i.test(s) ||
|
||||
/^Faster path to implementation$/i.test(s) ||
|
||||
/^Eng review is the required shipping gate [—–-] skipping it means less confidence before enabling the feature flag$/i.test(s));
|
||||
}
|
||||
|
||||
/** A completed CEO review may distinguish the still-unrun Eng shipping gate. */
|
||||
function unrunEngNavigation(fp: AskUserQuestionFingerprint, question: string): number | null {
|
||||
const call = fp.nativeCall!;
|
||||
const q = call.questions[0]!;
|
||||
if (!call.sessionId || !call.toolUseId || call.failed !== false ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0) ||
|
||||
q.options.length !== 2 || fp.options.length !== 2 ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!/^Next review$/i.test(q.header.trim()) ||
|
||||
!/^CEO Review is CLEAR\. Eng Review is the required shipping gate and (?:hasn['’]t|has not) run yet\. What(?:['’]s| is) next\?$/i.test(question)) return null;
|
||||
if (call.answered === false) {
|
||||
if (call.answers !== undefined || call.answeredAt !== undefined ||
|
||||
(call.unansweredQuestionIndices !== undefined &&
|
||||
(call.unansweredQuestionIndices.length !== 1 || call.unansweredQuestionIndices[0] !== 0))) return null;
|
||||
} else if (call.answered !== true || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length) return null;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).]\s*/i, '').replace(/\s*\(recommended\)\s*$/i, '').trim());
|
||||
const run = labels.findIndex(s => /^Run \/plan-eng-review(?: next| now)?$/i.test(s));
|
||||
const manual = labels.findIndex(s => /^Skip\s*[—–-]\s*I['’]ll handle reviews manually$/i.test(s));
|
||||
if (run < 0 || manual < 0 || run === manual) return null;
|
||||
const topics = String.raw`(?:architecture|code quality|test design|performance|deployment)`;
|
||||
const runDescription = new RegExp(String.raw`^${topics}(?:,\s+${topics})*(?:,?\s+and\s+${topics})?\s+review\.\s+The required gate before shipping\.\s+Run this before implementation begins to catch any structural issues in how the tests are wired up\.$`, 'i');
|
||||
// Consume each complete description in its own offered role. The temporal
|
||||
// qualification is about the next review, not an unfinished CEO decision.
|
||||
if (!runDescription.test(q.options[run]!.description?.trim() ?? '') ||
|
||||
!/^Proceed to implementation directly\.\s+You can run \/plan-eng-review later if needed\.\s+Eng Review is required before shipping but not before starting implementation\.$/i.test(q.options[manual]!.description?.trim() ?? '')) return null;
|
||||
return manual + 1;
|
||||
}
|
||||
|
||||
/** A completed review can explain the cost of skipping its next required gate. */
|
||||
function explainedRequiredEngNavigation(fp: AskUserQuestionFingerprint, question: string): number | null {
|
||||
const call = fp.nativeCall!, q = call.questions[0]!;
|
||||
if (!call.sessionId || !call.toolUseId || call.failed !== false ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0) ||
|
||||
q.header.trim() !== 'Next step' || q.options.length !== 2 || fp.options.length !== 2 ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label)) return null;
|
||||
if (call.answered === false) {
|
||||
if (call.answers !== undefined || call.answeredAt !== undefined ||
|
||||
(call.unansweredQuestionIndices !== undefined &&
|
||||
(call.unansweredQuestionIndices.length !== 1 || call.unansweredQuestionIndices[0] !== 0))) return null;
|
||||
} else if (call.answered !== true || !Array.isArray(call.unansweredQuestionIndices) ||
|
||||
call.unansweredQuestionIndices.length || Object.keys(call.answers ?? {}).length !== 1) return null;
|
||||
const compact = (s: string | undefined) => (s ?? '').replace(/\s+/g, ' ').trim();
|
||||
const match = /^What(?:['’]s| is) next after this CEO review\? ELI10: The CEO review is done and the plan is CLEARED\. But Eng Review is the required shipping gate [—–-] it covers architecture, test plan rigor, and implementation correctness in more depth\. Running it next locks in the plan before implementation starts\. Stakes if we pick wrong: Skipping eng review means the plan goes to implementation without a required gate check [—–-] leaving architecture and test-correctness gaps unverified\. Recommendation: ([A-Z]) because the dashboard shows Eng Review at 0 runs [—–-] required gate, not yet cleared\. Note: options differ in kind, not coverage [—–-] no completeness score\.$/.exec(compact(question));
|
||||
if (!match) return null;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).]\s*/, '').replace(/\s*\(Recommended\)$/, ''));
|
||||
const run = labels.indexOf('Run /plan-eng-review next'), manual = labels.indexOf('Skip — handle reviews manually');
|
||||
if (run < 0 || manual < 0 || run === manual || !q.options[run]!.label.startsWith(`${match[1]}) `)) return null;
|
||||
// All question prose and each role-specific description must be closed
|
||||
// navigation. The risk explanation is not a new CEO repair decision.
|
||||
if (!/^Required shipping gate\. Covers implementation correctness, test plan rigor, and any architecture concerns\. Takes ~[1-9]\d* minutes\.$/.test(compact(q.options[run]!.description)) ||
|
||||
compact(q.options[manual]!.description) !== 'Proceed to implementation without the eng review gate. CEO review findings still apply.') return null;
|
||||
return manual + 1;
|
||||
}
|
||||
|
||||
/** Shared closed-review guards; next-review sequencing is still navigation. */
|
||||
function closedNavigationContext(context: string): boolean {
|
||||
const unfinished = context.replace(/\b(?:no|0)\s+unresolved\s+(?:decisions|gaps|issues|findings)\b/gi, '');
|
||||
// Conditional closure of this review is unfinished work. Sequencing the
|
||||
// next review after implementation does not reopen the completed CEO review.
|
||||
const stateVerb = String.raw`(?:is|are|was|were|becomes?|became|(?:will|would|can|could|may|might)\s+(?:be|become))`;
|
||||
const closure = String.raw`(?:(?:all\s+)?(?:decisions|gaps|issues|findings)\s+(?:${stateVerb}\s+)?resolved|(?:the\s+)?CEO\s+review\s+(?:${stateVerb}\s+)?(?:complete|done|cleared|clean)|the\s+review\s+(?:${stateVerb}\s+)?(?:complete|done|cleared|clean))`;
|
||||
const conditionalClosure = new RegExp(String.raw`\b(?:once|when|after)\b[^.!?]{0,180}\b${closure}\b|\b${closure}\b[^.!?]{0,100}\b(?:once|when|after)\b`, 'i');
|
||||
return (context.match(/\?/g)?.length ?? 0) === 1 &&
|
||||
!/\b(?:unresolved|outstanding|remaining|pending|if|unless|until)\b|\b(?:gap|issue|finding|decision)s?\s+(?:still\s+)?remains?\b|\bstill\s+open\b/i.test(unfinished) &&
|
||||
!/\bnot\s+(?:all|no|0)\b/i.test(context) &&
|
||||
!conditionalClosure.test(context) &&
|
||||
!/(?:^|[.!?;]\s+|\b(?:proceed to|continue to|should|must|will|need to|can|could|would)\s+)(?:(?:please|first|then|also)\s+)*(?:add|fix|implement|resolve|decide)\b/im.test(context);
|
||||
}
|
||||
|
||||
/** A pure next-review menu remains navigation when question tuning is off. */
|
||||
function sequencedReviewNavigation(fp: AskUserQuestionFingerprint): number | null {
|
||||
const call = fp.nativeCall!, q = call.questions[0]!;
|
||||
if (call.failed !== false || !call.sessionId || !call.toolUseId || q.header.trim() !== 'Next review' ||
|
||||
q.options.length !== 2 || fp.options.length !== 2 || q.multiSelect ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return null;
|
||||
if (call.answered === false) {
|
||||
if (call.answers !== undefined || call.answeredAt !== undefined ||
|
||||
(call.unansweredQuestionIndices !== undefined &&
|
||||
(call.unansweredQuestionIndices.length !== 1 || call.unansweredQuestionIndices[0] !== 0))) return null;
|
||||
} else if (call.answered !== true || !Array.isArray(call.unansweredQuestionIndices) ||
|
||||
call.unansweredQuestionIndices.length || Object.keys(call.answers ?? {}).length !== 1) return null;
|
||||
const lines = q.question.trim().split('\n').map(line => line.trim()).filter(Boolean);
|
||||
if (!/^D[1-9]\d* [—–-] Which review runs next\?$/.test(lines[0] ?? '')) return null;
|
||||
// Consume the entire brief, including the displayed option explanations.
|
||||
// Only the next gate is open; adding a new remedy anywhere rejects this arm.
|
||||
const grammar = [
|
||||
/^Project\/branch\/task: [\w.-]+ on [\w./-]+; CEO review of [\w./-]+ is complete and clean \(HOLD SCOPE, 0 critical gaps, [1-9]\d* P1 tasks\)\.$/,
|
||||
/^ELI10: gstack chains reviews\. The CEO review just settled scope and strategy\. The engineering review is the required gate before shipping: it checks architecture, test design, and code quality in detail\. skip_eng_review is false, so it is still required\. No UI scope was detected, so the design review does not apply here\.$/,
|
||||
/^Stakes if we pick wrong: skipping eng review leaves the ship gate NOT CLEARED; the plan is small, so the eng review should be quick\.$/,
|
||||
/^Recommendation: A because eng review is the required gate and the plan now has exact assertions worth a second structured pass on test design\.$/,
|
||||
/^Note: options differ in kind, not coverage [—–-] no completeness score\.$/,
|
||||
/^A\) Run \/plan-eng-review next \(recommended\)$/,
|
||||
/^✅ Clears the required shipping gate on a plan that is small and already decided$/,
|
||||
/^✅ Gives the three tasks a test-design pass focused on the assertion mechanics \(mock implementation, sleeper record shape\)$/,
|
||||
/^❌ One more review session before implementation starts \(human ~[1-9]\d* min \/ CC ~[1-9]\d* min\)$/,
|
||||
/^B\) Skip, handle reviews manually$/,
|
||||
/^✅ Move straight to implementing T[1-9]\d* to T[1-9]\d* in the real repo$/,
|
||||
/^✅ No further review time on a three-task change$/,
|
||||
/^❌ Dashboard verdict stays NOT CLEARED until an eng review is logged$/,
|
||||
/^Net: gate discipline versus getting to the code faster on a change that is already tightly specified\.$/,
|
||||
];
|
||||
if (lines.length !== grammar.length + 1 || !grammar.every((re, i) => re.test(lines[i + 1]!))) return null;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z]:\s*/, '').replace(/\s*\(recommended\)$/, ''));
|
||||
const run = labels.indexOf('Run /plan-eng-review next'), manual = labels.indexOf('Skip, manual reviews');
|
||||
if (run < 0 || manual < 0 || run === manual ||
|
||||
q.options[run]!.description !== 'Required gate; runs after this plan is approved.' ||
|
||||
q.options[manual]!.description !== 'Proceed to implementation; eng gate remains open.') return null;
|
||||
return manual + 1;
|
||||
}
|
||||
|
||||
/** Closed CEO next-review navigation; native terminal/report checks prove completion separately. */
|
||||
function manualHandoffIndex(fp: AskUserQuestionFingerprint): number | null {
|
||||
const call = fp.nativeCall;
|
||||
// The capture path assigns this native identity only after matching the
|
||||
// active question. UI-only and mismatched pending records cannot steer it.
|
||||
if (!call || call.failed || fp.signature !== `${call.sessionId}:${call.toolUseId}` || call.questions.length !== 1) return null;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2) return null;
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:\s*([a-z0-9-]+)\s*>/gi)];
|
||||
if (ids.length > 1 || (q.question.match(/<gstack-qid/gi)?.length ?? 0) !== ids.length) return null;
|
||||
const id = ids[0]?.[1]?.toLowerCase();
|
||||
if (!id) {
|
||||
const sequenced = sequencedReviewNavigation(fp);
|
||||
if (sequenced !== null) return sequenced;
|
||||
}
|
||||
if (id && !/^(?:plan-ceo-(?:review-)?next-(?:steps?|review)|ceo-review-next-(?:steps?|review)|ceo-next-step-eng-review|ceo-plan-next-steps)$/.test(id)) return null;
|
||||
const declaration = q.question.replace(/^D\s*\d+\s*[—–:-]\s*/i, '')
|
||||
.replace(/^next\s+(?:review|steps?)\s*:\s*/i, '');
|
||||
const gateContext = [q.question, ...q.options.map(option => option.description ?? '')].join('\n');
|
||||
const explicitCompletion = /(?:^|[.!?]\s+)(?:ELI10:\s*)?(?:The\s+)?CEO\s+review\s+(?:is\s+)?(?:complete|cleared|clean|done(?:\s+and\s+the\s+plan\s+is\s+cleared)?)(?:\s+with\s+0\s+unresolved\s+decisions)?(?=\s*(?:[.!?—–]|$))/i.test(declaration);
|
||||
const genericCompletion = /(?:^|[.!?]\s+)(?:The\s+)?review\s+(?:is\s+)?(?:complete|cleared|clean|done)(?=\s*(?:[.!?—–]|$))/i.test(declaration);
|
||||
const questionText = declaration.replace(/<gstack-qid:[^>]+>/gi, '').trim();
|
||||
const unrunNavigation = id ? unrunEngNavigation(fp, questionText) : null;
|
||||
if (unrunNavigation !== null) return unrunNavigation;
|
||||
const explainedNavigation = id ? explainedRequiredEngNavigation(fp, questionText) : null;
|
||||
if (explainedNavigation !== null) return explainedNavigation;
|
||||
const recappedNavigation = Boolean(id) &&
|
||||
/^What(?:['’]s|\s+is)\s+the\s+next\s+(?:steps?|review)\s+after\s+(?:this|the)\s+CEO\s+review\?$/i.test(questionText) &&
|
||||
q.options.some(option => resolvedCeoRecap(option.description ?? ''));
|
||||
const unfinished = gateContext.replace(/\b(?:no|0)\s+unresolved\s+(?:decisions|gaps|issues|findings)\b/gi, '');
|
||||
const describedCompletion = (recappedNavigation || (genericCompletion && q.options.some(option => closedCeoRecap(option.description ?? '')))) &&
|
||||
!/\b(?:unresolved|outstanding|remains?|remaining|pending)\b/i.test(unfinished) &&
|
||||
!/(?:^|[.!?;]\s+|\b(?:please|must|need\s+to)\s+)(?:(?:please|first|then|also)\s+)*(?:add|fix|implement|resolve|decide)\b/im.test(gateContext);
|
||||
const metadataCompletion = Boolean(id) && (metadataClosedReviewNavigation(declaration, gateContext) ||
|
||||
(call.failed === false && q.options.length === 2 &&
|
||||
explainedMetadataNavigation(declaration, q.options.map(option => option.description ?? ''))));
|
||||
if (isMetadataNavigationQuestion(questionText) && /\n[ \t]*ELI10:/i.test(questionText) && !metadataCompletion) return null;
|
||||
const describedEngCompletion = q.options.length === 2 &&
|
||||
describedEngNavigation(questionText, q.options.map(option => option.description ?? ''), gateContext);
|
||||
const describedPostReviewCompletion = !id && q.options.length === 2 &&
|
||||
describedPostReviewNavigation(questionText, q.options.map(option => option.description ?? ''), gateContext);
|
||||
const countedCompletion = !id && call.failed === false && q.options.length === 2 &&
|
||||
countedCeoNavigation(questionText, q.options.map(option => option.description ?? ''));
|
||||
const clearCompletion = !id && call.failed === false && q.options.length === 2 &&
|
||||
clearRequiredEngNavigation(questionText, q.options.map(option => option.description ?? ''), gateContext);
|
||||
const completion = explicitCompletion || describedCompletion || metadataCompletion || describedEngCompletion || describedPostReviewCompletion || countedCompletion || clearCompletion;
|
||||
const bareNavigation = Boolean(id) && call.failed === false && q.options.length === 2 &&
|
||||
(fp.nativeQuestionIndex === undefined || fp.nativeQuestionIndex === 0) &&
|
||||
fp.options.length === 2 && fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) &&
|
||||
bareEngNavigation(questionText, q.options.map(o => o.description ?? ''));
|
||||
const requiredEng = /(?:\bEng(?:ineering)?\s+review|\/plan-eng-review)\b[^.!?]{0,180}\brequired(?:\s+shipping)?\s+gate\b/i.test(gateContext) ||
|
||||
/\brequired(?:\s+shipping)?\s+gate\s+is\s+(?:an?\s+)?(?:Eng(?:ineering)?\s+review|\/plan-eng-review)\b/i.test(gateContext) ||
|
||||
pronounEngGate(questionText, q.options.map(option => option.description ?? ''), gateContext);
|
||||
// These native next-review identities share a closed navigation contract;
|
||||
// the question or a following recap cannot hide a new repair obligation.
|
||||
if (id && /^(?:ceo-plan-next-steps|ceo-review-next-(?:steps?|review))$/.test(id) &&
|
||||
!closedReviewNavigation(declaration, gateContext)) return null;
|
||||
// A qid alone cannot authorize another fix. The bare navigation arm grants
|
||||
// no completion credit; native Exit, report freshness and finding floor remain independent.
|
||||
if (!/^next\s+(?:review|steps?)$/i.test(q.header.trim()) || !(completion || bareNavigation) || !requiredEng) return null;
|
||||
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).]\s*/i, '').replace(/\s*\(recommended\)\s*$/i, '').trim());
|
||||
if (clearCompletion && !labels.some(label => /^Run\s+\/plan-eng-review(?:\s+(?:next|now))?$/i.test(label))) return null;
|
||||
const runs = labels.map(label => /^Run\s+\/plan-(?:eng|design)-review(?:\s+(?:next|now))?(?:\s*\(required gate\))?$/i.test(label));
|
||||
const manual = labels.map(label => /^(?:Skip|Done)\s*[—–-]\s*(?:I['’]ll\s+)?handle\s+(?:reviews\s+)?manually$/i.test(label));
|
||||
// Deferring the next review until after already-approved implementation is
|
||||
// navigation too. A new fix/TODO/task choice remains substantive. The picker
|
||||
// always selects manual, never this implementation route.
|
||||
const deferred = labels.map((label, i) => /^Implement\s+now,\s+eng\s+review\s+later$/i.test(label) &&
|
||||
/^Proceed to implementation with (?:the )?(?:\d+ )?(?:already )?approved (?:tasks|plan|changes)(?: \([A-Z0-9–-]+\))?\. Run \/plan-eng-review before (?:the PR is merged|shipping)\.(?: Acceptable if implementation is expected to be fast with CC\.)?$/i.test(q.options[i]!.description?.trim() ?? ''));
|
||||
if (!runs.some(Boolean) || manual.filter(Boolean).length !== 1 || !labels.every((_, i) => runs[i] || manual[i] || deferred[i])) return null;
|
||||
return manual.findIndex(Boolean) + 1;
|
||||
}
|
||||
|
||||
/** Every offered explanation must remain a clause about this review handoff. */
|
||||
function closedNextReviewExplanations(question: string, descriptions: string[]): boolean {
|
||||
// Validate each complete sentence/line, rather than discarding prose under
|
||||
// an accepted heading. A new imperative has no navigation subject and
|
||||
// cannot borrow the preceding sentence's administrative classification.
|
||||
const navigation = [
|
||||
/^(?:The )?CEO review (?:is (?:done|complete|cleared)(?: and clears scope and strategy)?|cleared scope and strengthened both test assertions)$/i,
|
||||
/^The engineering review is the one gate that must pass before shipping \(skip_eng_review is false\)$/i,
|
||||
/^It checks architecture, code quality, and test design in depth$/i,
|
||||
/^gstack['’]s shipping gate is the eng review, which checks architecture and test design$/i,
|
||||
/^it has not run for this plan yet$/i,
|
||||
/^(?:There is no UI|No UI scope was found), so a design review does not apply$/i,
|
||||
/^Skipping (?:the )?eng review leaves the (?:required gate unmet, so the readiness dashboard stays NOT CLEARED until someone runs it later|ship dashboard NOT CLEARED)$/i,
|
||||
/^running it costs a few minutes on a two-test plan$/i,
|
||||
/^[A-Z] because eng review is the required (?:shipping gate and this plan is now precise enough for it to run quickly|gate and the plan changed since it was written \(two assertions strengthened\), so the tests deserve a second read)$/i,
|
||||
/^options differ in kind(?: \(which workflow runs next\))?, not coverage [—–-] no completeness score$/i,
|
||||
/^clear the (?:required )?gate now versus (?:handling reviews on your own schedule|implement first and review later)$/i,
|
||||
/^Clears the required (?:engineering gate while the plan and its two approved remedies are fresh|shipping gate on the review readiness dashboard)$/i,
|
||||
/^A second structured pass over the test design catches anything the scope review did not$/i,
|
||||
/^One more interactive review session before implementation starts$/i,
|
||||
/^Ends the review chain here$/i,
|
||||
/^you decide when the eng review runs$/i,
|
||||
/^No further (?:questions this session|review prompts in this session)$/i,
|
||||
/^(?:The required eng gate stays unmet and the dashboard remains NOT CLEARED|Dashboard stays NOT CLEARED until an eng review runs)$/i,
|
||||
/^Second read of the exact assertions and the await-then-count ordering before code is written$/i,
|
||||
/^A few extra minutes on a plan that is already two tests against existing probes$/i,
|
||||
/^Move straight to implementing T[1-9]\d* and T[1-9]\d* now$/i,
|
||||
/^Start the eng review against the updated plan after this review exits$/i,
|
||||
/^End here$/i,
|
||||
/^run reviews yourself later$/i,
|
||||
];
|
||||
const duration = String.raw`~?\d+(?:\.\d+)?\s*(?:minutes?|mins?|hours?|hrs?|days?|weeks?)`;
|
||||
const timing = new RegExp(String.raw`\s*\(human:\s*${duration}\s*/\s*CC:\s*${duration}\)$`, 'i');
|
||||
let metadata = 0;
|
||||
const body = question.split('\n').slice(1).concat(descriptions.flatMap(text => text.split('\n')));
|
||||
for (const raw of body) {
|
||||
const line = raw.trim();
|
||||
if (!line) continue;
|
||||
if (/^Project\/branch\/task:/.test(line)) {
|
||||
// Only the project/mode recap is metadata, never a repair paragraph.
|
||||
if (++metadata !== 1 || !/^Project\/branch\/task: (?:[\w-]+ on [\w/-]+, \/plan-ceo-review \(HOLD SCOPE\) finished on the payment test-coverage plan|`[\w/-]+`, CEO review of PLAN\.md complete \(HOLD SCOPE, 0 critical gaps, [1-9]\d* assertion fixes approved\))\.$/.test(line)) return false;
|
||||
continue;
|
||||
}
|
||||
if (line === 'Pros / cons:') continue;
|
||||
if (/^[A-Z][):] /.test(line)) {
|
||||
const offered = line.replace(/^[A-Z][):] /, '').replace(timing, '').replace(/ \(recommended\)$/, '');
|
||||
if (!/^(?:Run \/plan-eng-review next|Skip, handle reviews manually)$/.test(offered)) return false;
|
||||
continue;
|
||||
}
|
||||
const prose = line.replace(/^(?:ELI10|Stakes if we pick wrong|Recommendation|Note|Net):\s*/, '')
|
||||
.replace(/^[✅❌]\s*/, '').replace(timing, '');
|
||||
const clauses = prose.split(/[.;]\s+|[.]$/).map(s => s.trim()).filter(Boolean);
|
||||
if (!clauses.length || !clauses.every(clause => navigation.some(pattern => pattern.test(clause)))) return false;
|
||||
}
|
||||
return metadata === 1;
|
||||
}
|
||||
|
||||
/** Evidence-only next-review accounting; this never selects a pending option. */
|
||||
function completedNextReviewBrief(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.sessionId || !call.toolUseId || call.answered !== true || call.failed !== false ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` || call.questions.length !== 1 ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0) ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
Object.keys(call.answers ?? {}).length !== 1 || !Number.isFinite(Date.parse(call.answeredAt ?? ''))) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || !/^Next (?:step|review)$/i.test(q.header.trim()) || q.options.length !== 2 ||
|
||||
fp.options.length !== 2 || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question]) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const question = q.question.trim().replace(/^D\d+\s*[—–-]\s*/i, '');
|
||||
const lines = question.split('\n').map(line => line.trim()).filter(Boolean);
|
||||
// Numbered headings and echoed pros/cons are presentation. Require the
|
||||
// actual navigation query, explicit current CEO closure, and a final brief
|
||||
// boundary; a new question or directive after that boundary stays work.
|
||||
if (!/^(?:CEO review (?:is )?(?:complete|done|cleared)\. )?Which review runs next\?$/i.test(lines[0]!) ||
|
||||
!/^Net:\s+[^\n]+[.!]$/.test(lines.at(-1) ?? '') ||
|
||||
!/(?:^|[.!?]\s+|^ELI10:\s*)(?:The\s+)?CEO\s+review\s+(?:is\s+)?(?:complete|done|cleared)\b/im.test(question)) return false;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).:]\s*/i, '').replace(/\s*\(recommended\)\s*$/i, ''));
|
||||
if (labels.filter(label => /^Run \/plan-eng-review next$/i.test(label)).length !== 1 ||
|
||||
labels.filter(label => /^Skip\s*[,—–-]\s*(?:(?:I['’]ll\s+)?handle reviews manually|manual reviews)$/i.test(label)).length !== 1) return false;
|
||||
const context = [question, ...q.options.map(o => o.description ?? '')].join('\n')
|
||||
.replace(/^Stakes if we pick wrong:/m, 'Stakes:');
|
||||
// Only the next Eng gate can keep the readiness dashboard uncleared.
|
||||
// Its temporal explanation is not a condition on current CEO closure;
|
||||
// every other unfinished-work and conditional-closure guard still applies.
|
||||
const navigationContext = context.replace(
|
||||
/\b((?:(?:readiness|ship)\s+)?dashboard\s+(?:stays|remains)\s+NOT\s+CLEARED)\s+until\s+(?:someone\s+runs\s+it|(?:an?|the)\s+eng(?:ineering)?\s+review\s+runs)(?:\s+later)?(?=[.!]|\n|$)/gi,
|
||||
'$1',
|
||||
);
|
||||
return q.options.every(o => o.description?.trim()) &&
|
||||
closedNextReviewExplanations(question, q.options.map(o => o.description ?? '')) &&
|
||||
!/`{3}|~{3}|(?:^|\n)\s*>|\b(?:example|quoted source)\s*:/im.test(context) &&
|
||||
!/\bCEO\s+review\b[^.!?\n]{0,80}\b(?:not|never|incomplete|unfinished)\b/i.test(context) &&
|
||||
/\b(?:Eng|engineering) review\b[^.!?]{0,180}\bgate\b/i.test(context) && closedNavigationContext(navigationContext);
|
||||
}
|
||||
|
||||
/** Classification happens only after one real, successful, fully answered native call. */
|
||||
export function isCeoCompletionHandoff(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length) return false;
|
||||
if (completedNextReviewBrief(fp)) return true;
|
||||
if (manualHandoffIndex(fp) === null) return false;
|
||||
const q = call.questions[0]!;
|
||||
// A free-form answer can introduce a new substantive request. Do not
|
||||
// silently discard it merely because the menu itself was administrative.
|
||||
return q.options.some(option => option.label === call.answers?.[q.question]);
|
||||
}
|
||||
|
||||
/** Finish this CEO fixture instead of starting another skill; reuse the existing caller-pick hook. */
|
||||
export function pickCeoCompletionHandoff(
|
||||
fp: AskUserQuestionFingerprint,
|
||||
activeCapture: AskUserQuestionFingerprint = fp,
|
||||
): number | null {
|
||||
return activeCapture.nativeCall?.answered ? null : manualHandoffIndex(activeCapture);
|
||||
}
|
||||
@@ -0,0 +1,480 @@
|
||||
/** Match CEO mode labels after the PTY capture strips cursor-spacing escapes. */
|
||||
import {
|
||||
capturePlanCountQuestion,
|
||||
createPlanCountPermissionGuard,
|
||||
parseQuestionPrompt,
|
||||
parseNumberedOptions,
|
||||
auqFingerprint,
|
||||
classifyPlanCountFrame,
|
||||
isNumberedOptionListVisible,
|
||||
matchesNativePlanQuestion,
|
||||
planCountSubmissionInput,
|
||||
planCountPrerequisitePick,
|
||||
type AskUserQuestionFingerprint,
|
||||
} from './claude-pty-runner';
|
||||
import type { NativePlanQuestionCall, NativePublicToolEvent, PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
type CeoMode = 'HOLD SCOPE' | 'SCOPE EXPANSION' | 'SELECTIVE EXPANSION' | 'SCOPE REDUCTION';
|
||||
|
||||
function modeTitle(label: string): string | undefined {
|
||||
const title = label.split(/[│┌\r\n]/, 1)[0]!.trim().replace(/^[A-Z][).:]\s*/i, '').replace(/\s+/g, '').toUpperCase();
|
||||
return /^(HOLDSCOPE|SCOPEEXPANSION|SELECTIVEEXPANSION|SCOPEREDUCTION)(?:$|[^A-Z])/.exec(title)?.[1];
|
||||
}
|
||||
|
||||
export function findCeoModeOption(
|
||||
options: ReadonlyArray<{ index: number; label: string }>,
|
||||
targetMode: CeoMode,
|
||||
): number | null {
|
||||
const modes = options.map(option => {
|
||||
// The CLI renders a description pane beside the options. Its text may
|
||||
// mention a different mode, so match only the leading option title.
|
||||
return { index: option.index, mode: modeTitle(option.label) };
|
||||
});
|
||||
if (!modes.some(option => option.mode)) return null;
|
||||
const recognized = modes.map(option => option.mode).filter(Boolean);
|
||||
if (new Set(recognized).size !== recognized.length) {
|
||||
throw new Error('Mode AskUserQuestion has duplicate mode choices');
|
||||
}
|
||||
|
||||
const target = modes.find(option => option.mode === targetMode.replace(/\s+/g, ''));
|
||||
if (!target) {
|
||||
throw new Error(
|
||||
`Mode AskUserQuestion rendered but target "${targetMode}" not in option labels:\n` +
|
||||
options.map(option => ` ${option.index}. ${option.label}`).join('\n'),
|
||||
);
|
||||
}
|
||||
return target.index;
|
||||
}
|
||||
|
||||
type ModeNavigationAction =
|
||||
| { kind: 'wait' }
|
||||
| { kind: 'permission' | 'submission'; input: string }
|
||||
| { kind: 'question'; index: number; question: AskUserQuestionFingerprint }
|
||||
| { kind: 'mode'; index: number; question: AskUserQuestionFingerprint };
|
||||
|
||||
// Permission state follows the navigation session without entering its AUQ
|
||||
// dedup set. A file-tool result can reopen an identical permission prompt.
|
||||
const permissionStates = new WeakMap<Set<string>, {
|
||||
file: ReturnType<typeof createPlanCountPermissionGuard>;
|
||||
other: Set<string>;
|
||||
}>();
|
||||
function ceoPermissionAction(visible: string, seenQuestions: Set<string>, completionHistory = visible): 'grant' | 'handled' | null {
|
||||
let state = permissionStates.get(seenQuestions);
|
||||
if (!state) {
|
||||
state = { file: createPlanCountPermissionGuard(), other: new Set() };
|
||||
permissionStates.set(seenQuestions, state);
|
||||
}
|
||||
const file = state.file(visible, completionHistory);
|
||||
if (file !== null) return file;
|
||||
if (classifyPlanCountFrame(visible) !== 'permission') return null;
|
||||
const signature = auqFingerprint(parseQuestionPrompt(visible), parseNumberedOptions(visible));
|
||||
if (state.other.has(signature)) return 'handled';
|
||||
state.other.add(signature);
|
||||
return 'grant';
|
||||
}
|
||||
|
||||
/** Handle native controls before deduping actual navigation questions. */
|
||||
export function nextCeoModeNavigation(
|
||||
visible: string,
|
||||
targetMode: CeoMode,
|
||||
seenQuestions: Set<string>,
|
||||
pending?: NativePlanQuestionCall,
|
||||
completionHistory = visible,
|
||||
): ModeNavigationAction {
|
||||
const permission = pending && matchesNativePlanQuestion(visible, pending) ? null : ceoPermissionAction(visible, seenQuestions, completionHistory);
|
||||
if (permission !== null) return permission === 'grant'
|
||||
? { kind: 'permission', input: '1\r' } : { kind: 'wait' };
|
||||
const frame = classifyPlanCountFrame(visible);
|
||||
const submission = frame === null ? planCountSubmissionInput(visible) : null;
|
||||
if (submission !== null) return { kind: 'submission', input: submission };
|
||||
if (!isNumberedOptionListVisible(visible)) return { kind: 'wait' };
|
||||
const question = capturePlanCountQuestion(visible, seenQuestions, 0, true, pending);
|
||||
if (!question) return { kind: 'wait' };
|
||||
const index = findCeoModeOption(question.options, targetMode);
|
||||
// Preserve the seeded review plan by declining the existing optional
|
||||
// office-hours offer. Other navigation questions keep their first choice.
|
||||
return index === null
|
||||
? { kind: 'question', index: planCountPrerequisitePick(question) ?? 1, question }
|
||||
: { kind: 'mode', index, question };
|
||||
}
|
||||
|
||||
/** Match new assistant prose, never the native mode menu or answer echo. */
|
||||
export function hasPostAnswerCeoPosture(visible: string, posture: RegExp): boolean {
|
||||
let assistant: string[] | null = null;
|
||||
const matches = () => {
|
||||
if (!assistant?.length) return false;
|
||||
const compact = assistant.join('').replace(/\s+/g, '');
|
||||
// Tool headings use the same bullet as assistant messages. Their output
|
||||
// can quote the selected mode or the skill's posture instructions. Keep
|
||||
// word boundaries when detecting a call: ordinary prose can contain parentheses.
|
||||
if (/^(?:UseransweredClaude['’]squestions|(?:high|medium|low)·\/effort)/i.test(compact) ||
|
||||
/^[A-Za-z][\w.:_-]*[ \t]*\(/.test(assistant[0]!.trim())) return false;
|
||||
// A bare selected title gains no evidentiary value when the next terminal
|
||||
// update appends a spinner or other chrome to the same captured block.
|
||||
const prose = assistant.filter(line =>
|
||||
!/^(?:(?:You)?selected(?:option|mode)?[::]?)?(?:HOLDSCOPE|SCOPEEXPANSION|SELECTIVEEXPANSION|SCOPEREDUCTION)(?:\(recommended\))?\.?$/i.test(line.replace(/\s+/g, '')) &&
|
||||
!/^[✶✻✽✢·]/.test(line),
|
||||
).join('\n');
|
||||
return prose.search(posture) !== -1;
|
||||
};
|
||||
|
||||
for (const line of visible.replace(/\r\n?/g, '\n').split('\n')) {
|
||||
const text = line.trim();
|
||||
const message = /^[●⏺]\s*(.*)$/.exec(text);
|
||||
if (message) {
|
||||
if (matches()) return true;
|
||||
assistant = message[1] ? [message[1]] : [];
|
||||
} else if (/^(?:[☐☒❯⎿│┌└─⏸]|←|\d+[.)]\s*|Enter\s*to\s*select)/i.test(text)) {
|
||||
if (matches()) return true;
|
||||
assistant = null;
|
||||
} else if (assistant && text) {
|
||||
assistant.push(text);
|
||||
}
|
||||
}
|
||||
return matches();
|
||||
}
|
||||
|
||||
/** Apply the same echo/quotation exclusions to native prose and decision rationale. */
|
||||
function hasNativePostureProse(text: string, posture: RegExp): boolean {
|
||||
const prose = text.replace(/```[\s\S]*?```/g, '').split('\n').filter(line => {
|
||||
if (/^\s*>/.test(line)) return false;
|
||||
const plain = line.replace(/[*_`]/g, '').trim().replace(/^#+\s*/, '');
|
||||
// A repeated menu or bare confirmation is still only an answer echo.
|
||||
if (/^(?:[-+]|\d+[.)]|[A-D][.)])\s*(?:HOLD SCOPE|SCOPE EXPANSION|SELECTIVE EXPANSION|SCOPE REDUCTION)\b/i.test(plain)) return false;
|
||||
return !/^(?:(?:You\s+)?selected(?:\s+(?:option|mode))?\s*[::]?\s*)?(?:HOLD SCOPE|SCOPE EXPANSION|SELECTIVE EXPANSION|SCOPE REDUCTION)(?:\s+mode)?(?:\s+confirmed)?(?:\s*\(recommended\))?[.!]?$/i.test(plain);
|
||||
}).join('\n');
|
||||
return hasPostAnswerCeoPosture(`● ${prose}`, posture);
|
||||
}
|
||||
|
||||
/** Current scope lock + exclusion + hardening can apply HOLD without naming it. */
|
||||
function hasCurrentHoldScopePosture(text: string, selected: NativePlanQuestionCall): boolean {
|
||||
const plain = text.replace(/\*\*/g, '').replace(/’/g, "'").trim();
|
||||
// This gate checks adopted review posture, not completed review work. An
|
||||
// immediate first-person commitment carries the same scope obligations as
|
||||
// a present-tense declaration; deferred or conditional plans still do not.
|
||||
const opening = /^[\s\S]*?[.!?](?=\s|$)/.exec(plain)?.[0] ?? '';
|
||||
const sentence = opening.replace(/^(?:I'll|I will|We'll|We will) (lock|keep|hold)\b/i,
|
||||
(_match, verb: string) => `I am ${{ lock: 'locking', keep: 'keeping', hold: 'holding' }[verb.toLowerCase()]}`);
|
||||
if (/\b(?:not|never|won't|can't|don't|isn't|aren't|if|unless|until|may|might|could|would|will|later|tomorrow|eventually|future|example|hypothetical|after|once|when|whenever|following|pending|provided|assuming)\b|\b(?:next (?:week|month|year)|subject to)\b/i.test(sentence)) return false;
|
||||
const declaration = /^(?:I'm|I am|We're|We are) (?:locking|keeping|holding) (?:the )?scope (?:to|at) ([^,\n]+),\s*(?:flagging|treating|marking) (?:anything|everything) (?:beyond|outside) that(?: \([^()\n]+\))? as out of scope,? and (?:hunting|checking|looking) for (?:silent )?(?:failure modes|errors|edge cases)\b([^.!?\n]*)\.$/i.exec(sentence);
|
||||
const pressureTest = /^(?:I'm|I am|We're|We are) (?:locking|keeping|holding) (?:the )?scope fixed to ([^,\n]+),\s*pressure[- ]testing every stated behavior for failure modes, ([^.!?\n]+?) while deferring (?:anything|everything) extra rather than adding it silently\.$/i.exec(sentence);
|
||||
const ambiguityReview = /^(?:I'm|I am|We're|We are) (?:keeping|holding) strictly to (?:the )?(plan|[\w./-]+\.md)'s approved scope(?: \((Approach [A-Z])(?:, [^()\n]+)?\))? and (?:flagging|surfacing|raising) (?:any|the) ambiguities (?:the (?:sketch|plan) leaves undecided|in the (?:plan|sketch)) as targeted questions rather than expanding scope\.$/i.exec(sentence);
|
||||
const commitment = declaration ?? pressureTest;
|
||||
if (!commitment && !ambiguityReview) return false;
|
||||
// A later current correction can withdraw the declaration. Quoted examples
|
||||
// cannot; the opening declaration was matched before removing quoted blocks.
|
||||
// Only this recognized opening's explicit exclusion is negative expansion;
|
||||
// keep later corrections and every other scope statement in the withdrawal check.
|
||||
const current = ambiguityReview
|
||||
? opening.replace(/ rather than expanding scope\.$/i, '.') + plain.slice(opening.length) : plain;
|
||||
const currentProse = current.replace(/```[\s\S]*?(?:```|$)|~~~[\s\S]*?(?:~~~|$)/g, '')
|
||||
.replace(/^\s*>.*$/gm, '').replace(/"[^"\n]*"|“[^”\n]*”/g, '');
|
||||
if (ambiguityReview && (/\b(?:no longer|not)\s+(?:flagging|surfacing|raising)\s+(?:(?:any|the)\s+)?ambiguities\b/i.test(currentProse) ||
|
||||
/\b(?:this|the) posture (?:is|was|has been) (?:withdrawn|rejected|cancelled|canceled|superseded|(?:not|no longer) current)\b/i.test(currentProse))) return false;
|
||||
if (/\b(?:expand\w*|widen\w*|reduc(?:e|ing)|shrink\w*)\s+(?:the\s+)?scope\b/i.test(currentProse) ||
|
||||
/\b(?:add|adding)\b[^.!?\n]*\b(?:to|into)\s+(?:the\s+)?scope\b/i.test(currentProse) ||
|
||||
/\b(?:no longer|not)\s+(?:locking|keeping|holding|lock|keep|hold)\s+(?:(?:the\s+)?scope\b|strictly to (?:the )?(?:plan|[\w./-]+\.md)'s approved scope\b)/i.test(currentProse) ||
|
||||
/\b(?:previously|formerly) excluded\b[^.!?\n]*\b(?:now )?in scope\b/i.test(currentProse)) return false;
|
||||
const context = selected.questions.map(q => /^Project\/branch\/task:([^\n]*)/im.exec(q.question)?.[1] ?? '').join(' ');
|
||||
const plans = new Set(context.match(/\b[\w./-]+\.md\b/gi) ?? []);
|
||||
if (ambiguityReview) {
|
||||
if (plans.size !== 1 || (sentence.match(/\b[\w./-]+\.md\b/gi) ?? []).some(plan => !plans.has(plan))) return false;
|
||||
// Parenthetical approach labels must belong to the approved current plan.
|
||||
const approach = ambiguityReview[2];
|
||||
const approval = approach ? new RegExp(`\\b${approach}\\b[^.!?\\n]*\\bapproved\\b`, 'i') : /\bapproved\b/i;
|
||||
return approval.test(context) && !/\b(?:not|no|unapproved|rejected|hypothetical|if|unless|until|after|once|when|whenever|following|pending|provided|assuming)\b|\bsubject to\b/i.test(context);
|
||||
}
|
||||
const baseline = commitment![1]!.trim();
|
||||
const namedPlan = /^(?:the )?(?:(?:\d+|one|two|three|four|five|six|seven|eight|nine|ten) )?([\w./-]+\.md) (?:bullets|requirements|scope)(?: from approach [A-Z])?$/i.exec(baseline);
|
||||
const possessivePlan = /^(?:the )?([\w./-]+\.md)'s (?:(?:\d+|one|two|three|four|five|six|seven|eight|nine|ten) )?(?:bullets|requirements|scope)( plus the approved schema)?$/i.exec(baseline);
|
||||
const plan = namedPlan ?? possessivePlan;
|
||||
if (plan ? plans.size !== 1 || !plans.has(plan[1]!)
|
||||
: !/^the (?:current|agreed|approved|existing) (?:plan|scope)$/i.test(baseline)) return false;
|
||||
if (possessivePlan?.[2] && (!/\bschema\b[^.!?\n]{0,80}\bapproved\b|\bapproved\b[^.!?\n]{0,80}\bschema\b/i.test(context)
|
||||
|| /\b(?:not|no|unapproved|rejected|hypothetical|if|unless|until|after|once|when|whenever|following|pending|provided|assuming)\b|\bsubject to\b/i.test(context))) return false;
|
||||
// Concrete failure surfaces distinguish review rigor from merely retaining scope.
|
||||
const hardening = commitment![2]!;
|
||||
return [/\bconstraints\b/i, /\b(?:error handling|errors)\b/i, /\bedge cases\b/i,
|
||||
/\baccess(?:-rule)? leaks\b/i, /\bsecurity\b/i, /\btest(?:ing|s)\b/i, /\bproduction visibility\b/i]
|
||||
.filter(surface => surface.test(hardening)).length >= 2;
|
||||
}
|
||||
|
||||
/** A native successful answer, not the key we intended to send to the menu. */
|
||||
export function nativeCeoModeAnswer(
|
||||
transcript: PlanCountTranscript,
|
||||
targetMode: CeoMode,
|
||||
selectionStartedAt: number,
|
||||
): NativePlanQuestionCall | null {
|
||||
if (transcript.status !== 'ready') return null;
|
||||
const choices = transcript.calls.flatMap(call => {
|
||||
const at = Date.parse(call.answeredAt ?? '');
|
||||
if (!call.answered || call.failed || !Number.isFinite(at) || at < selectionStartedAt) return [];
|
||||
return call.questions.flatMap(question => {
|
||||
const recognized = question.options.map(option => modeTitle(option.label)).filter(Boolean);
|
||||
const modes = new Set(recognized);
|
||||
if (modes.size !== recognized.length) return [{ call, at, mode: undefined }];
|
||||
const answer = call.answers?.[question.question];
|
||||
return modes.size >= 2 && typeof answer === 'string'
|
||||
? [{ call, at, mode: modeTitle(answer) }] : [];
|
||||
});
|
||||
}).sort((a, b) => b.at - a.at);
|
||||
const latest = choices[0];
|
||||
return latest?.mode === targetMode.replace(/\s+/g, '') ? latest.call : null;
|
||||
}
|
||||
|
||||
/** An answered AUQ must have one matching native request and successful reply. */
|
||||
function completedQuestionTimes(call: NativePlanQuestionCall, events: ReadonlyArray<NativePublicToolEvent>) {
|
||||
if (!call.answered || call.failed || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length) return null;
|
||||
const own = events.filter(event => event.sessionId === call.sessionId && event.toolUseId === call.toolUseId);
|
||||
const requests = own.filter(event => event.kind === 'use');
|
||||
const replies = own.filter(event => event.kind === 'result');
|
||||
if (requests.length !== 1 || replies.length !== 1) return null;
|
||||
const request = requests[0]!;
|
||||
const reply = replies[0]!;
|
||||
const requestedAt = Date.parse(request.timestamp);
|
||||
const answeredAt = Date.parse(reply.timestamp);
|
||||
if (request.name !== 'AskUserQuestion' || reply.isError ||
|
||||
JSON.stringify(request.input?.questions) !== JSON.stringify(call.questions) ||
|
||||
reply.timestamp !== call.answeredAt || !Number.isFinite(requestedAt) ||
|
||||
!Number.isFinite(answeredAt) || requestedAt >= answeredAt) return null;
|
||||
return { requestedAt, answeredAt };
|
||||
}
|
||||
|
||||
/** New shorthand forms must be one complete decision, not a mode mention or extra question. */
|
||||
function singleScopeBrief(text: string, descriptions: readonly string[], comparison = true): boolean {
|
||||
if ([text, ...descriptions].some(value => /(?:^|[.!?]\s+|\n)\s*(?:Also|Separately|Additionally)\b|\b(?:Please|We must|You must|The plan must)\b/i.test(value))) return false;
|
||||
// Query parameter names such as ?view= are not another decision prompt.
|
||||
const questions = text.replace(/\?[A-Za-z_][\w-]*=/g, '=').match(/\?/g);
|
||||
if (questions?.length !== 1 || /```|~~~|^\s*>/m.test(text)) return false;
|
||||
const markers = [/Project\/branch\/task:/gi, /ELI10:/gi, /Stakes if (?:we pick )?wrong:/gi,
|
||||
/Recommendation:/gi, /Completeness:/gi, /Net:/gi];
|
||||
let previous = -1;
|
||||
const complete = markers.every(marker => {
|
||||
const matches = [...text.matchAll(marker)];
|
||||
if (matches.length !== 1 || matches[0]!.index! <= previous) return false;
|
||||
previous = matches[0]!.index!;
|
||||
return true;
|
||||
});
|
||||
// Net closes this decision brief. A following instruction is not part of its
|
||||
// comparison; this is not a general classifier of instructions inside prose.
|
||||
const net = text.slice(previous + 'Net:'.length).trim();
|
||||
return complete && (comparison ? /^[^.!?;\n]+ (?:vs|versus) [^.!?;\n]+\.$/.test(net)
|
||||
: /^[^.!?;\n]+\.$/.test(net.replace(/\bvs\./gi, 'vs')));
|
||||
}
|
||||
|
||||
/** HOLD can apply its boundary in a completed defer decision, before standalone prose. */
|
||||
function hasAnsweredHoldPosture(transcript: PlanCountTranscript, selected: NativePlanQuestionCall,
|
||||
posture: RegExp, events: ReadonlyArray<NativePublicToolEvent>): boolean {
|
||||
const modeTimes = completedQuestionTimes(selected, events);
|
||||
if (!modeTimes) return false;
|
||||
return transcript.calls.some(call => {
|
||||
if (call === selected || call.sessionId !== selected.sessionId || call.questions.length !== 1) return false;
|
||||
const times = completedQuestionTimes(call, events);
|
||||
if (!times || times.requestedAt <= modeTimes.answeredAt) return false;
|
||||
const q = call.questions[0]!;
|
||||
// A substantive review decision can apply HOLD in its rationale before
|
||||
// standalone prose is published. Metadata and answer echoes do not count.
|
||||
// This recognizes posture language; it does not validate every scope choice.
|
||||
const context = /Project\/branch\/task:([\s\S]*?)(?=ELI10:)/i.exec(q.question)?.[1] ?? '';
|
||||
const rationale = /ELI10:([\s\S]*?)(?=Stakes if (?:we pick )?wrong:)/i.exec(q.question)?.[1]?.trim() ?? '';
|
||||
const offered = q.options.map(o => o.label.trim());
|
||||
if (!q.multiSelect && q.options.length >= 2 && q.options.length <= 4 && new Set(offered).size === offered.length &&
|
||||
offered.includes(call.answers?.[q.question] ?? '') && /\bHOLD SCOPE\b/i.test(context) &&
|
||||
!/\b(?:SCOPE EXPANSION|SELECTIVE EXPANSION|SCOPE REDUCTION)\b/i.test(context) &&
|
||||
singleScopeBrief(q.question, q.options.map(o => o.description ?? ''), false) &&
|
||||
hasNativePostureProse(rationale, posture)) return true;
|
||||
if (q.multiSelect || q.options.length !== 3 || !singleScopeBrief(q.question, q.options.map(o => o.description ?? ''))) return false;
|
||||
const title = /^D\d+\s*[—–-]\s*Under HOLD SCOPE, keep or defer the (\w+)\b([\s\S]+?)\?\s+Project\/branch\/task:/i.exec(q.question);
|
||||
if (!title || !/\bnot in the plan text\b/i.test(title[2]!) ||
|
||||
!/\bwritten scope is the baseline\b/i.test(q.question) ||
|
||||
!/\bpure additions,\s*not repairs to meet a stated invariant\b/i.test(q.question)) return false;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-C][):.]\s*/i, '')
|
||||
.replace(/\s*\(recommended\)\s*$/i, '').toLowerCase());
|
||||
const count = title[1]!.toLowerCase();
|
||||
const defer = labels.findIndex(label => label === `defer all ${count} to todos` || label === `defer all ${count} to todos.md`);
|
||||
const subset = labels.find(label => /^keep .+ only$/.test(label));
|
||||
if (new Set(labels).size !== 3 || defer < 0 || !labels.includes(`keep all ${count}`) ||
|
||||
!subset || !title[2]!.toLowerCase().includes(subset.slice(5, -5)) ||
|
||||
call.answers?.[q.question] !== q.options[defer]!.label) return false;
|
||||
return hasPostAnswerCeoPosture(`● ${q.question}`, posture);
|
||||
});
|
||||
}
|
||||
|
||||
/** A concrete, opted-in expansion brief is itself assistant posture evidence. */
|
||||
function hasAnsweredExpansionPosture(
|
||||
transcript: PlanCountTranscript, selected: NativePlanQuestionCall,
|
||||
posture: RegExp, events: ReadonlyArray<NativePublicToolEvent>,
|
||||
): boolean {
|
||||
const modeTimes = completedQuestionTimes(selected, events);
|
||||
if (!modeTimes) return false;
|
||||
return transcript.calls.some(call => {
|
||||
if (call === selected || call.sessionId !== selected.sessionId || call.questions.length !== 1) return false;
|
||||
const times = completedQuestionTimes(call, events);
|
||||
if (!times || times.requestedAt <= modeTimes.answeredAt) return false;
|
||||
const question = call.questions[0]!;
|
||||
// This is evidence that the selected mode produced a concrete scope
|
||||
// decision, not authority to answer it. Numbering and heading names vary.
|
||||
const title = question.question.split('\n')[0]!
|
||||
.replace(/\s*<gstack-qid:[a-z0-9-]+>\s*$/i, '').replace(/^D\d+\s*[—–-]\s*/i, '');
|
||||
const context = /\nProject\/branch\/task:([^\n]+)/i.exec(question.question)?.[1] ?? '';
|
||||
if (question.multiSelect || question.options.length !== 3 ||
|
||||
!/^[\p{L}\p{N}][^?\n]+\?$/u.test(title) ||
|
||||
/\b(?:review\s+(?:mode|posture)|(?:selected|confirmed)\s+(?:mode|option))\b/i.test(title) ||
|
||||
/\b(?:HOLD SCOPE|SELECTIVE EXPANSION|SCOPE REDUCTION)\b/i.test(context) ||
|
||||
!/\b(?:SCOPE\s+EXPANSION|EXPANSION\s+(?:mode|opt[ -]in))\b/i.test(context) ||
|
||||
!singleScopeBrief(question.question, question.options.map(o => o.description ?? ''), false)) return false;
|
||||
const labels = question.options.map(option => option.label.trim()
|
||||
.replace(/^[A-C][):.]\s*/i, '').replace(/\s*\(recommended\)\s*$/i, '').toLowerCase()
|
||||
.replace(/^add to (?:(?:this|the) plan['’]s )?scope$/, 'add to scope'));
|
||||
if (new Set(labels).size !== 3 || !['add to scope', 'defer to todos.md', 'skip'].every(label => labels.includes(label)) ||
|
||||
!question.options.some(option => option.label === call.answers?.[question.question])) return false;
|
||||
// Never search quoted instructions, tool output or a menu for posture.
|
||||
const prose = question.question.replace(/```[\s\S]*?(?:```|$)|~~~[\s\S]*?(?:~~~|$)/g, '')
|
||||
.replace(/^\s*>.*$/gm, '');
|
||||
return hasPostAnswerCeoPosture(`● ${prose}`, posture);
|
||||
});
|
||||
}
|
||||
|
||||
/** Match finalized prose or a completed concrete scope decision after the actual mode answer. */
|
||||
export function hasNativePostAnswerCeoPosture(
|
||||
transcript: PlanCountTranscript,
|
||||
targetMode: CeoMode,
|
||||
posture: RegExp,
|
||||
selectionStartedAt: number,
|
||||
publicTools: ReadonlyArray<NativePublicToolEvent> = [],
|
||||
): boolean {
|
||||
const selected = nativeCeoModeAnswer(transcript, targetMode, selectionStartedAt);
|
||||
if (!selected) return false;
|
||||
const answeredAt = Date.parse(selected.answeredAt!);
|
||||
return transcript.assistantMessages.some(message => {
|
||||
if (message.sessionId !== selected.sessionId || Date.parse(message.timestamp) <= answeredAt) return false;
|
||||
return hasNativePostureProse(message.text, posture) ||
|
||||
(targetMode === 'HOLD SCOPE' && Number.isFinite(Date.parse(message.timestamp)) &&
|
||||
Date.parse(message.timestamp) <= Date.now() &&
|
||||
hasCurrentHoldScopePosture(message.text, selected));
|
||||
}) || (targetMode === 'SCOPE EXPANSION' && hasAnsweredExpansionPosture(transcript, selected, posture, publicTools)) ||
|
||||
(targetMode === 'HOLD SCOPE' && hasAnsweredHoldPosture(transcript, selected, posture, publicTools));
|
||||
}
|
||||
|
||||
type PosturePacket = { headers: string[]; screens: string[]; next: number; nativeId?: string; submitted: boolean };
|
||||
const postureContinuations = new WeakMap<Set<string>, { modeId: string; packet?: PosturePacket }>();
|
||||
|
||||
/** The complete native bar binds delayed-JSONL tabs to one bounded AUQ. */
|
||||
function posturePacketBar(visible: string): { headers: string[]; answered: boolean[] } | null {
|
||||
const bars = [...visible.matchAll(/←([^\r\n]+)✔\s*Submit\s*→/g)];
|
||||
const bar = bars.at(-1);
|
||||
if (!bar) return null;
|
||||
const tabs = [...bar[1]!.matchAll(/([☐☒])\s*([^☐☒]+)/g)];
|
||||
if (tabs.length < 2 || tabs.length > 4 || bar[1]!.slice(0, tabs[0]!.index).trim()) return null;
|
||||
const headers = tabs.map(tab => tab[2]!.trim().replace(/\s+/g, ' '));
|
||||
if (headers.some(header => !header) || new Set(headers).size !== headers.length) return null;
|
||||
return { headers, answered: tabs.map(tab => tab[1] === '☒') };
|
||||
}
|
||||
|
||||
const BARLESS_SUBMIT_END = 'Readytosubmityouranswers?❯1.Submitanswers2.Cancel';
|
||||
|
||||
/** The barless review must reproduce every observed question and chosen option. */
|
||||
function barlessPostureSubmit(visible: string, packet: PosturePacket): boolean {
|
||||
if (packet.next !== packet.headers.length || packet.screens.length !== packet.next) return false;
|
||||
const compact = (text: string) => text.replace(/^[\t │┃]*[●⏺][\t ]*/gm, '')
|
||||
.replace(/[│┃\s]/g, '');
|
||||
let prefix = '';
|
||||
const answers: string[] = [];
|
||||
for (const screen of packet.screens) {
|
||||
const bar = [...screen.matchAll(/←[^\r\n]+✔\s*Submit\s*→/g)].at(-1);
|
||||
const cursor = [...screen.matchAll(/❯\s*1\./g)].at(-1);
|
||||
const choice = parseNumberedOptions(screen).find(option => option.index === 1)?.label;
|
||||
if (!bar || !cursor || !choice || cursor.index! <= bar.index! + bar[0].length) return false;
|
||||
const question = compact(screen.slice(bar.index! + bar[0].length, cursor.index));
|
||||
if (!question || /[←☐☒❯]/.test(question)) return false;
|
||||
// The caller answers option 1 for each of this one bounded call's tabs.
|
||||
answers.push(`${question}→${compact(choice)}`);
|
||||
prefix = compact(screen.slice(0, bar.index));
|
||||
}
|
||||
const panel = compact(visible);
|
||||
if (!panel.startsWith(prefix)) return false;
|
||||
const body = panel.slice(prefix.length).replace(/^Reviewyouranswers/, '');
|
||||
return body === answers.join('') + BARLESS_SUBMIT_END;
|
||||
}
|
||||
|
||||
/**
|
||||
* Claude can defer persisting assistant prose until the next AUQ resolves.
|
||||
* Permit one fresh downstream call, including its remaining tabs and Submit.
|
||||
* Native completion still supplies all posture evidence; UI only drives input.
|
||||
*/
|
||||
export function nextCeoPostureContinuation(
|
||||
visible: string,
|
||||
transcript: PlanCountTranscript,
|
||||
targetMode: CeoMode,
|
||||
selectionStartedAt: number,
|
||||
seenQuestions: Set<string>,
|
||||
alreadyContinued: boolean,
|
||||
completionHistory = visible,
|
||||
pendingQuestion?: NativePlanQuestionCall & {source:'pre_tool_use'},
|
||||
): 'permission' | 'question' | 'submission' | null {
|
||||
const supplied = pendingQuestion?.source === 'pre_tool_use' && !pendingQuestion.answered && !pendingQuestion.failed &&
|
||||
!transcript.calls.some(call => call.sessionId === pendingQuestion.sessionId && call.toolUseId === pendingQuestion.toolUseId)
|
||||
? pendingQuestion : undefined;
|
||||
const pending = transcript.calls.find(call => !call.answered && !call.failed) ?? supplied;
|
||||
const permission = pending && matchesNativePlanQuestion(visible, pending) ? null : ceoPermissionAction(visible, seenQuestions, completionHistory);
|
||||
if (permission !== null) return permission === 'grant' ? 'permission' : null;
|
||||
const selected = nativeCeoModeAnswer(transcript, targetMode, selectionStartedAt);
|
||||
if (!selected) return null;
|
||||
if (pending && pending.sessionId !== selected.sessionId) return null;
|
||||
const modeId = `${selected.sessionId}:${selected.toolUseId}`;
|
||||
const state = postureContinuations.get(seenQuestions);
|
||||
if (state && state.modeId !== modeId) return null;
|
||||
const bar = posturePacketBar(visible);
|
||||
const packet = state?.packet;
|
||||
if (state || alreadyContinued) {
|
||||
if (!packet || packet.submitted) return null;
|
||||
const barless = !bar && barlessPostureSubmit(visible, packet);
|
||||
if (!barless && (!bar || JSON.stringify(bar.headers) !== JSON.stringify(packet.headers) ||
|
||||
!bar.answered.every((answered, i) => answered === (i < packet.next)))) return null;
|
||||
const sameHeaders = (call: NativePlanQuestionCall) => JSON.stringify(call.questions.map(q =>
|
||||
q.header.trim().replace(/\s+/g, ' '))) === JSON.stringify(packet.headers);
|
||||
const recorded = transcript.calls.slice(transcript.calls.indexOf(selected) + 1).find(sameHeaders);
|
||||
const call = packet.nativeId
|
||||
? transcript.calls.find(call => `${call.sessionId}:${call.toolUseId}` === packet.nativeId) ??
|
||||
(pending && `${pending.sessionId}:${pending.toolUseId}` === packet.nativeId ? pending : undefined)
|
||||
: pending ?? recorded;
|
||||
if (packet.nativeId && !call) return null;
|
||||
if (call) {
|
||||
if (call.sessionId !== selected.sessionId || call.answered || call.failed ||
|
||||
!sameHeaders(call) || (pending && pending !== call) ||
|
||||
!packet.screens.every((screen, i) => capturePlanCountQuestion(
|
||||
screen, new Set(), 0, false, call)?.nativeQuestionIndex === i)) return null;
|
||||
packet.nativeId = `${call.sessionId}:${call.toolUseId}`;
|
||||
}
|
||||
if (packet.next === packet.headers.length) {
|
||||
if (!barless && planCountSubmissionInput(visible) !== '\r') return null;
|
||||
packet.submitted = true;
|
||||
return 'submission';
|
||||
}
|
||||
} else if (bar && bar.answered.some(Boolean)) return null;
|
||||
// An unbound Submit panel is never a fresh question to answer with option 1.
|
||||
if (!bar && visible.replace(/[│┃\s]/g, '').endsWith(BARLESS_SUBMIT_END)) return null;
|
||||
// With native metadata present, require the same call and exact displayed
|
||||
// tab. Without it, the complete bar, footer and ordered answered transitions
|
||||
// are required; another menu cannot spend this call's remaining tab budget.
|
||||
if (bar) {
|
||||
if (!/Enter\s*to\s*select\s*·\s*Tab\/Arrow\s*keys\s*to\s*navigate\s*·\s*Esc\s*to\s*cancel/i.test(visible)) return null;
|
||||
if (pending && (pending.sessionId !== selected.sessionId ||
|
||||
!matchesNativePlanQuestion(visible, pending) ||
|
||||
JSON.stringify(pending.questions.map(q => q.header.trim().replace(/\s+/g, ' '))) !== JSON.stringify(bar.headers))) return null;
|
||||
}
|
||||
const action = nextCeoModeNavigation(visible, targetMode, seenQuestions, pending, completionHistory);
|
||||
if (action.kind !== 'question') return null;
|
||||
if (bar) {
|
||||
const current = packet ?? { headers: bar.headers, screens: [], next: 0, submitted: false };
|
||||
if (action.question.nativeCall) {
|
||||
const nativeId = `${action.question.nativeCall.sessionId}:${action.question.nativeCall.toolUseId}`;
|
||||
if ((current.nativeId && current.nativeId !== nativeId) || action.question.nativeQuestionIndex !== current.next) return null;
|
||||
current.nativeId = nativeId;
|
||||
}
|
||||
current.screens.push(visible);
|
||||
current.next++;
|
||||
postureContinuations.set(seenQuestions, { modeId, packet: current });
|
||||
} else postureContinuations.set(seenQuestions, { modeId });
|
||||
return 'question';
|
||||
}
|
||||
@@ -0,0 +1,815 @@
|
||||
/**
|
||||
* Section loading needs a complete, bounded plan, not a finding-count fixture.
|
||||
* The former two-bullet plan made a successful review invent key isolation,
|
||||
* error, concurrency, observability, rollout, and test contracts in a 38 KB
|
||||
* report. Those surrounding contracts are explicit here; the read/write sketch
|
||||
* still permits an old in-flight read to refill a key after write invalidation.
|
||||
*/
|
||||
export const CACHE_READ_WRITE_SKETCH = `async function readProfile(key) {
|
||||
const cached = cache.get(key);
|
||||
if (cached !== undefined) return cached;
|
||||
const value = await repository.read(key);
|
||||
cache.set(key, value);
|
||||
return value;
|
||||
}
|
||||
|
||||
async function writeProfile(key, update) {
|
||||
const saved = await repository.write(key, update);
|
||||
cache.delete(key);
|
||||
return saved;
|
||||
}`;
|
||||
|
||||
export const CEO_SECTION_CACHE_PLAN = `# Plan: cache profile summaries in one process
|
||||
|
||||
## Measured problem and accepted scope
|
||||
The existing profile-summary service has one active process. A one-week trace
|
||||
shows repeated reads of about 900 hot keys: DB CPU is 70%, with read p95 120 ms.
|
||||
Add a process-local LRU wrapper to the existing repository. Acceptance targets
|
||||
are at least 60% cache hits, DB CPU below 50%, and read p95 below 60 ms, with the
|
||||
existing error-rate and correctness SLOs unchanged. This is an internal backend
|
||||
change with no UI, API, schema, pricing, or developer onboarding change.
|
||||
|
||||
## Existing contracts retained
|
||||
- All reads and writes use this repository in the same process; there are no
|
||||
external DB writers. Multi-process operation remains unsupported and startup
|
||||
rejects that configuration while caching is enabled.
|
||||
- Authentication and authorization run before repository access. Keys encode
|
||||
the authenticated tenant ID and validated profile ID without ambiguity.
|
||||
Values are immutable profile-summary DTOs; secrets and cache keys are never
|
||||
logged. Cached results cannot bypass authorization.
|
||||
- The existing LRU adapter supports 1000 entries, a 16 MiB byte cap, and a
|
||||
30-second TTL. Recorded hot data fits those limits. Absent records use a
|
||||
distinct sentinel with a 10-second TTL; undefined means a cache miss.
|
||||
- Cache operations are synchronous and atomic in the single JS event loop.
|
||||
On any cache failure the existing adapter bypasses the cache until an empty
|
||||
cache is reinitialized; repository errors keep the current typed API error
|
||||
mapping. The existing per-key
|
||||
single-flight wrapper coalesces simultaneous misses and releases on failure.
|
||||
- A read already in progress when a write commits may return its earlier DB
|
||||
snapshot to that caller. Every read begun after that write completes must
|
||||
observe the committed version. TTL expiry is not a substitute for this rule.
|
||||
|
||||
## Proposed wrapper integration
|
||||
Keep the current read-through repository interface and shared adapters. These
|
||||
are the complete new read/write ordering rules; no additional version checks or
|
||||
coordination between a cache fill and a write are proposed:
|
||||
|
||||
\`\`\`javascript
|
||||
${CACHE_READ_WRITE_SKETCH}
|
||||
\`\`\`
|
||||
|
||||
## Verification and rollout
|
||||
Existing repository contract tests cover tenant isolation, key validation,
|
||||
absence, DB failures, and authorization. New wrapper tests cover hit/miss,
|
||||
eviction and byte limits, TTL, adapter-failure fallback, successful-write
|
||||
invalidation, failed-write preservation, and concurrent-miss coalescing.
|
||||
The rollout uses the existing runtime feature flag: enable for 10% of keys,
|
||||
then 50%, then all keys after one healthy hour at each stage. Monitor hit/miss,
|
||||
eviction, cache bytes, fallback errors, DB CPU, and read p95 without raw IDs.
|
||||
On error-rate or latency regression, disable the flag immediately; both reads
|
||||
and writes bypass the cache while disabled, and enabling creates an empty cache.
|
||||
Cold starts remain within the existing DB capacity. The service owner monitors
|
||||
the rollout and records the results against the acceptance targets.
|
||||
|
||||
## Out of scope
|
||||
Distributed caching, cross-process coherence, prewarming, changing consistency
|
||||
semantics, or adding new product surfaces. The repository interface preserves a
|
||||
future replacement path without introducing a general cache framework now.
|
||||
`;
|
||||
|
||||
/** All six events must form one ordered, same-key, post-write reader trace. */
|
||||
function hasNumberedStaleFillTrace(text: string): boolean {
|
||||
const events = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ')).filter(Boolean);
|
||||
if (events.length !== 6) return false;
|
||||
const arrow = String.raw`\s*(?:→|->)\s*`;
|
||||
const backArrow = String.raw`\s*(?:←|<-)\s*`;
|
||||
const identifier = String.raw`([A-Za-z_$][\w$]*)`;
|
||||
const read = String.raw`readProfile\(\s*${identifier}\s*\)`;
|
||||
const write = String.raw`writeProfile\(\s*${identifier}\s*\)`;
|
||||
const first = new RegExp(String.raw`^t1:\s*${read}${arrow}cache miss${arrow}(?:single-flight${arrow})?await DB read(?: \(suspends\))?$`, 'i').exec(events[0]!);
|
||||
if (!first) return false;
|
||||
// Prose keywords are case-insensitive; identifiers remain case-sensitive.
|
||||
const sameKey = (event: string, pattern: string) => new RegExp(pattern, 'i').exec(event)?.[1] === first[1];
|
||||
return sameKey(events[1]!, String.raw`^t2:\s*${write}${arrow}await DB write(?: \(suspends\))?$`)
|
||||
&& sameKey(events[2]!, String.raw`^t3:\s*DB write completes${arrow}cache\.delete\(\s*${identifier}\s*\)${arrow}writeProfile returns$`)
|
||||
&& new RegExp(String.raw`^t4:\s*DB read \(from t1\) completes${arrow}returns (?:old|stale) (?:snapshot|value)$`, 'i').test(events[3]!)
|
||||
&& sameKey(events[4]!, String.raw`^t5:\s*cache\.set\(\s*${identifier}\s*,\s*(?:(?:OLD|STALE)_VALUE|(?:old|stale) (?:snapshot|value))\s*\)${backArrow}(?:old|stale) (?:value|snapshot) (?:re-inserted|refilled) after invalidation!?$`)
|
||||
&& sameKey(events[5]!, String.raw`^t6:\s*(?:next|new|subsequent) ${read}${arrow}cache HIT${arrow}returns (?:old|stale) (?:value|snapshot)${backArrow}(?:INVARIANT|CONTRACT) (?:VIOLATED|VIOLATION)!?$`);
|
||||
}
|
||||
|
||||
/** Bind each column of an explicit execution to its reader, writer, key and version. */
|
||||
function columnarStaleFillTrace(text: string): { tail: string; reader: string; old: string } | undefined {
|
||||
const lines = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ')).filter(Boolean);
|
||||
const cells = (line: string) => line.replace(/^\|\s*|\s*\|$/g, '').split('|').map(cell => cell.trim());
|
||||
const header = cells(lines[0] ?? '');
|
||||
if (header.length !== 6 || header[0]!.toLowerCase() !== 't') return;
|
||||
const reader = /^(R[1-9]\d*) read \(begins before (W[1-9]\d*|W)\)$/i.exec(header[1]!);
|
||||
const later = /^(R[1-9]\d*) read \(begins after (W[1-9]\d*|W)\)$/i.exec(header[3]!);
|
||||
const cache = /^cache\[([A-Za-z_$][\w$]*)\]$/.exec(header[4]!);
|
||||
const db = /^DB\[([A-Za-z_$][\w$]*)\]$/.exec(header[5]!);
|
||||
if (!reader || !later || !cache || !db || reader[1] === later[1] ||
|
||||
reader[2] !== later[2] || header[2] !== `${reader[2]} write` || cache[1] !== db[1]) return;
|
||||
const events = lines.slice(1, 8).map(cells);
|
||||
if (events.length !== 7 || events.some((row, i) => row.length !== 6 || row[0] !== String(i + 1))) return;
|
||||
const old = events[0]![5]!, fresh = events[2]![5]!;
|
||||
if (!/^[A-Za-z][\w.-]*$/.test(old) || !/^[A-Za-z][\w.-]*$/.test(fresh) || old === fresh) return;
|
||||
const escape = (value: string) => value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
||||
const key = escape(cache[1]!), before = escape(old), after = escape(fresh);
|
||||
const arrow = String.raw`\s*(?:->|→)\s*`;
|
||||
const matches = (value: string, pattern: string) => new RegExp(`^(?:${pattern})$`).test(value);
|
||||
const empty = (value: string) => value === '-' || value === 'empty';
|
||||
if (!matches(events[0]![1]!, String.raw`get\(${key}\)${arrow}undefined`) ||
|
||||
!matches(events[1]![1]!, String.raw`await repository\.read(?:\(${key}\))?${arrow}${before}`) ||
|
||||
!matches(events[2]![2]!, String.raw`await write commits ${after}`) ||
|
||||
!matches(events[3]![2]!, String.raw`delete\(${key}\) \(no entry\)`) || events[4]![2] !== 'returns' ||
|
||||
!matches(events[5]![1]!, String.raw`resume: set\(${key},\s*${before}\); return ${before}`) ||
|
||||
!matches(events[6]![3]!, String.raw`get\(${key}\)${arrow}${before}; return ${before}`)) return;
|
||||
for (let i = 0; i < 7; i++) {
|
||||
const row = events[i]!;
|
||||
if (row[5] !== (i < 2 ? old : fresh) ||
|
||||
(i < 5 ? !empty(row[4]!) : row[4] !== (i === 5 ? `${old} STALE` : old)) ||
|
||||
(i < 6 && row[3] !== '') ||
|
||||
([0, 1, 5, 6].includes(i) && row[2] !== '') ||
|
||||
([2, 3, 4].includes(i) && row[1] !== 'paused') || (i === 6 && row[1] !== '')) return;
|
||||
}
|
||||
const violation = String.raw`VIOLATION t7: ${escape(later[1]!)} began after ${escape(reader[2]!)} completed \(t5\), observes ${before}(?: for up to [1-9]\d* (?:s|seconds))?\.`;
|
||||
if (!matches(lines[8] ?? '', violation)) return;
|
||||
const tail = lines.slice(8).join(' ');
|
||||
if (/\b(?:(?:this|that|the)\s+(?:trace|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(tail)) return;
|
||||
return { tail, reader: reader[1]!, old };
|
||||
}
|
||||
|
||||
/** An explicitly named alternate order can supply the later cache-hit reader. */
|
||||
function alternateOrderStaleFillTrace(text: string): string | undefined {
|
||||
const lines = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ')).filter(Boolean);
|
||||
const cells = (line: string) => line.replace(/^\|\s*|\s*\|$/g, '').split('|').map(cell => cell.trim());
|
||||
const header = cells(lines[0] ?? '');
|
||||
if (header.length !== 6 || header[0] !== 't') return;
|
||||
const identifier = String.raw`([A-Za-z_$][\w$]*)`;
|
||||
const read = new RegExp(String.raw`^(R[1-9]\d*) readProfile\(${identifier}\)$`);
|
||||
const first = read.exec(header[1]!);
|
||||
const writer = new RegExp(String.raw`^(W[1-9]\d*|W) writeProfile\(${identifier},\s*${identifier}\)$`).exec(header[2]!);
|
||||
const later = read.exec(header[3]!);
|
||||
if (!first || !writer || !later || first[1] === later[1] || first[2] !== writer[2] || first[2] !== later[2]
|
||||
|| header[4] !== `cache[${first[2]}]` || header[5] !== `inflight[${first[2]}]`) return;
|
||||
const rows = lines.slice(1, 6).map(cells);
|
||||
if (rows.length !== 5 || rows.some((row, i) => row.length !== 6 || row[0] !== String(i + 1))) return;
|
||||
const flight = new RegExp(String.raw`^miss; flight ${identifier}; await read$`).exec(rows[0]![1]!);
|
||||
const fill = new RegExp(String.raw`^read resolves ${identifier}; set\(${identifier},\s*${identifier}\)$`).exec(rows[4]![1]!);
|
||||
if (!flight || !fill || fill[2] !== first[2] || fill[1] !== fill[3] || fill[1] === writer[3]) return;
|
||||
const old = fill[1]!, fresh = writer[3]!, key = first[2]!;
|
||||
const expected = [
|
||||
[rows[0]![1]!, '', '', '-', flight[1]!],
|
||||
['', `await write ... commit ${fresh}`, '', '-', flight[1]!],
|
||||
['', `delete(${key}) no-op; return`, '', '-', flight[1]!],
|
||||
['', '', `begins; miss; joins ${flight[1]}`, '-', flight[1]!],
|
||||
[rows[4]![1]!, '', `receives ${old} VIOLATION`, `${old} BAD`, '-'],
|
||||
];
|
||||
if (rows.some((row, i) => row.slice(1).some((cell, j) => cell !== expected[i]![j]))) return;
|
||||
// The ordinary row 4 joins an old flight. It is not a later cache hit.
|
||||
// Order B explicitly moves that same reader after the stale fill at row 5.
|
||||
const order = `Order B: ${later[1]} begins after t5 -> cache hit ${old} VIOLATION (until TTL or next write)`;
|
||||
const contract = `Contract: ${later[1]} began after ${writer[1]} completed, so ${later[1]} must observe ${fresh}. Sketch has no preventing mechanism.`;
|
||||
if (lines[6]?.replace(/→/g, '->') !== order || lines[7] !== contract) return;
|
||||
// The separately labelled amended execution cannot supply original evidence.
|
||||
// Keep original assessment text, including any later dismissal, authoritative.
|
||||
const amended = lines.findIndex((line, i) => i > 7 && /^AMENDED \(D[1-9]\d*\):$/.test(line));
|
||||
const assessment = lines.slice(6, amended < 0 ? undefined : amended).join(' ');
|
||||
if (lines.slice(amended < 0 ? lines.length : amended + 1).some(line =>
|
||||
/\b(?:original|this|the)\s+(?:trace|schedule|scenario|race)\s+(?:is|was)\s+(?:impossible|not\s+(?:a\s+)?(?:bug|defect|violation))\b/i.test(line))) return;
|
||||
return assessment;
|
||||
}
|
||||
|
||||
/** Read an explicit original-sketch override without crediting the amended fill. */
|
||||
function originalSketchStaleFillTrace(text: string): { assessment: string; summary: string } | undefined {
|
||||
const lines = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ').replace(/→/g, '->')).filter(Boolean);
|
||||
const cells = (line: string) => line.split('|').map(cell => cell.trim());
|
||||
const header = cells(lines[0] ?? '');
|
||||
if (header.length !== 5 || header[0] !== 'step' || header[4] !== 'cache / pending') return;
|
||||
const first = /^(R[1-9]\d*) read \(began before commit\)$/.exec(header[1]!);
|
||||
const writer = /^(W[1-9]\d*|W) write$/.exec(header[2]!);
|
||||
const later = /^(R[1-9]\d*) read \(began after (W[1-9]\d*|W) resolves\)$/.exec(header[3]!);
|
||||
if (!first || !writer || !later || first[1] === later[1] || writer[1] !== later[2]) return;
|
||||
const rows = lines.slice(1, 9).map(cells);
|
||||
const steps = ['1', '2', '3', '4', "4'", '5', '6', '7'];
|
||||
if (rows.length !== 8 || rows.some((row, i) => row.length !== 5 || row[0] !== steps[i])) return;
|
||||
const token = /^get->miss; token ([A-Za-z_$][\w$]*)$/.exec(rows[0]![1]!);
|
||||
const next = /^get->miss; token ([A-Za-z_$][\w$]*); fresh DB read$/.exec(rows[5]![3]!);
|
||||
const write = /^await repo\.write -> ([A-Za-z_$][\w$]*) committed$/.exec(rows[2]![2]!);
|
||||
const read = /^read resolves ([A-Za-z_$][\w$]*); ([A-Za-z_$][\w$]*)✗ -> no fill$/.exec(rows[6]![1]!);
|
||||
if (!token || !next || !write || !read || token[1] === next[1] || read[2] !== token[1] || read[1] === write[1]) return;
|
||||
const old = read[1]!, fresh = write[1]!, pending = token[1]!, newPending = next[1]!;
|
||||
// One cache/pending column owns this execution. The writer cancels the
|
||||
// original reader's exact token; no other key, token or actor may borrow it.
|
||||
const expected = [
|
||||
[`get->miss; token ${pending}`, '', '', `∅ / {${pending}}`],
|
||||
['await singleFlight(repo.read)', '', '', ''],
|
||||
['', `await repo.write -> ${fresh} committed`, '', ''],
|
||||
['', `invalidate: cancel ${pending}, detach, delete`, '', `∅ / {${pending}✗}`],
|
||||
['', 'writeProfile resolves (write "complete")', '', ''],
|
||||
['', '', `get->miss; token ${newPending}; fresh DB read`, `∅ / {${pending}✗,${newPending}}`],
|
||||
[`read resolves ${old}; ${pending}✗ -> no fill`, '', '', `∅ / {${newPending}}`],
|
||||
['', '', `resolves ${fresh}; fill; return ${fresh}`, `${fresh} / ∅`],
|
||||
];
|
||||
if (rows.some((row, i) => row.slice(1).some((cell, j) => cell !== expected[i]![j]))) return;
|
||||
const alternate = `Alternate order (6 before 4): ${first[1]} fills ${old}, then step 4 deletes it; ${later[1]} misses and reads ${fresh}. Safe.`;
|
||||
if (lines[9] !== alternate) return;
|
||||
const original = /^Original sketch: step 6 fills ([A-Za-z_$][\w$]*) after step 4 -> (R[1-9]\d*) hits ([A-Za-z_$][\w$]*) -> VIOLATION \((S[1-9]\d*)\)\.$/.exec(lines[10] ?? '');
|
||||
if (!original || original[1] !== old || original[3] !== old || original[2] !== later[1]) return;
|
||||
const assessment = lines.slice(10).join(' ');
|
||||
if (/\b(?:(?:this|that|the|original)\s+(?:trace|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(assessment)) return;
|
||||
return {
|
||||
assessment,
|
||||
summary: `Schedule ${original[4]}: ${first[1]} misses, ${writer[1]} commits and deletes, ${first[1]} fills stale ${old}, ${later[1]} hits ${old}.`,
|
||||
};
|
||||
}
|
||||
|
||||
/** An original-order annotation can override the safe fill in its owned amended trace. */
|
||||
function originalOrderAStaleFillTrace(text: string): { assessment: string; ttl: string } | undefined {
|
||||
const lines = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ').replace(/→/g, '->')).filter(Boolean);
|
||||
const cells = (line: string) => line.split('|').map(cell => cell.trim());
|
||||
const header = cells(lines[0] ?? '');
|
||||
const first = /^(R[1-9]\d*) \(began before (W[1-9]\d*|W)\)$/.exec(header[1] ?? '');
|
||||
const cache = /^cache\[([A-Za-z_$][\w$]*)\]$/.exec(header[3] ?? '');
|
||||
if (header.length !== 5 || header[0] !== 't' || !first || !cache || header[2] !== first[2]
|
||||
|| header[4] !== `inflight[${cache[1]}]`) return;
|
||||
const rows = lines.slice(1, 9).map(cells);
|
||||
const token = /^([A-Za-z_$][\w$]*) stale=F$/.exec(rows[0]?.[4] ?? '');
|
||||
const write = /^DB write commits ([A-Za-z_$][\w$]*)$/.exec(rows[1]?.[2] ?? '');
|
||||
const read = /^DB returns ([A-Za-z_$][\w$]*) \(resume queued\)$/.exec(rows[2]?.[1] ?? '');
|
||||
const later = /^(R[1-9]\d*) begins: miss, new ([A-Za-z_$][\w$]*), DB->([A-Za-z_$][\w$]*), fill ([A-Za-z_$][\w$]*)$/.exec(rows[7]?.[1] ?? '');
|
||||
if (!token || !write || !read || !later || token[1] === later[2] || first[1] === later[1]
|
||||
|| write[1] === read[1] || later[3] !== write[1] || later[4] !== write[1]) return;
|
||||
const old = read[1]!, fresh = write[1]!, pending = token[1]!, writer = first[2]!;
|
||||
const expected = [
|
||||
['1', 'get->undef; run(); DB read sent', '', '-', `${pending} stale=F`],
|
||||
['2', '', `DB write commits ${fresh}`, '-', pending],
|
||||
['3', `DB returns ${old} (resume queued)`, '', '-', pending],
|
||||
['4', '', `resume: invalidate(${pending}), delete`, '-', `- (${pending} detached)`],
|
||||
['5', '', `settles -> ${writer} complete`, '-', '-'],
|
||||
['6', `resume: ${pending}.stale -> skip fill`, '', '-', '-'],
|
||||
['', `return ${old} (allowed: began < 5)`, '', '', ''],
|
||||
['7', later[0], `${fresh} OK`, later[2]!],
|
||||
];
|
||||
if (rows.length !== expected.length || rows.some((row, i) => row.length !== expected[i]!.length
|
||||
|| row.some((cell, j) => cell !== expected[i]![j]))) return;
|
||||
if (lines[9] !== `Order B (6 before 4): ${first[1]} fills ${old}, then ${writer} deletes at 4 -> ${later[1]} misses -> ${fresh} OK`) return;
|
||||
const joiner = /^Late joiner (R[1-9]\d*) arriving after 5: ([A-Za-z_$][\w$]*) detached -> new entry -> ([A-Za-z_$][\w$]*) OK$/.exec(lines[10] ?? '');
|
||||
if (!joiner || [first[1], later[1]].includes(joiner[1]) || joiner[2] !== pending || joiner[3] !== fresh) return;
|
||||
const original = /^Original sketch, order A: fill ([A-Za-z_$][\w$]*) at 6 after delete at 4 -> (R[1-9]\d*) reads ([A-Za-z_$][\w$]*) for <=([1-9]\d*) s VIOLATION$/.exec(lines[11] ?? '');
|
||||
if (!original || original[1] !== old || original[2] !== later[1] || original[3] !== old
|
||||
|| lines[12] !== `Original sketch, ${joiner[1]} after 5: joins ${pending} -> ${old} VIOLATION`) return;
|
||||
// The original annotation supplies the failing fill. The amended skip and
|
||||
// the already-started reader's allowed return cannot establish that defect.
|
||||
return { assessment: lines.slice(11).join(' '), ttl: original[4]! };
|
||||
}
|
||||
|
||||
/** Supplement a version-bound prose sequence with its named actors and write completion. */
|
||||
function versionedOriginalSchedule(text: string, key: string, old: string, fresh: string, finding: string): { schedule: string; assessment: string } | undefined {
|
||||
const raw = text.split('\n').filter(line => line.trim());
|
||||
const lines = raw.map(line => line.trim().replace(/\s+/g, ' '));
|
||||
const header = /^ *(S[1-9]\d*) late fill +(R[1-9]\d*) \(read, began before (W[1-9]\d*|W)\) +(W[1-9]\d*|W) \(write\) +(R[1-9]\d*) \(read, began after (W[1-9]\d*|W)\) +cache +gen$/.exec(raw[0] ?? '');
|
||||
if (!header || header[2] === header[5] || header[3] !== header[4] || header[4] !== header[6]) return;
|
||||
const operations = [
|
||||
'get->miss, seen=0', `DB SELECT -> ${old}`, `DB UPDATE commits ${fresh}`,
|
||||
'delete (no-op), return', `promise resolves, set(${old})`, `get -> ${old}`,
|
||||
];
|
||||
const expected = [
|
||||
`1 ${operations[0]} - 0`, `2 ${operations[1]}`, `3 ${operations[2]}`,
|
||||
`4 sketch ${operations[3]} - 0`, `5 sketch ${operations[4]} ${old}`, `6 sketch ${operations[5]} VIOLATION`,
|
||||
];
|
||||
if (expected.some((line, i) => lines[i + 1] !== line)) return;
|
||||
// Whitespace columns identify who performs each operation. A writer's
|
||||
// return in the original row 4 precedes the later reader's row 6 cache hit.
|
||||
const firstColumn = raw[0]!.indexOf(`${header[2]} (read`);
|
||||
const writerColumn = raw[0]!.indexOf(`${header[4]} (write`);
|
||||
const laterColumn = raw[0]!.indexOf(`${header[5]} (read`);
|
||||
const owners = [firstColumn, firstColumn, writerColumn, writerColumn, firstColumn, laterColumn];
|
||||
if (operations.some((operation, i) => Math.abs(raw[i + 1]!.indexOf(operation) - owners[i]!) > 1)) return;
|
||||
const guarded = new Set([
|
||||
"4'guarded gen=1, delete, return - 1",
|
||||
`5'guarded stamp 1 != seen 0: skip fill, metric++, return ${old} (allowed) - 1`,
|
||||
`6'guarded miss, seen=1, flight ${key}#1 -> ${fresh}, set ${fresh}`,
|
||||
]);
|
||||
const nextSchedule = lines.findIndex((line, i) => {
|
||||
const named = /^(S[1-9]\d*)\s/.exec(line);
|
||||
return i > 6 && named !== null && named[1] !== header[1];
|
||||
});
|
||||
// Only these exact separately-labelled guarded operations are an amendment,
|
||||
// not a dismissal of S1. Unknown same-schedule assessment text is retained.
|
||||
const sameOwner = new RegExp(`^(?:${header[1]}|${finding})\\b`);
|
||||
const assessment = lines.slice(7, nextSchedule < 0 ? undefined : nextSchedule)
|
||||
.filter(line => !guarded.has(line));
|
||||
// A later different schedule does not erase a subsequent explicit
|
||||
// assessment of this original schedule or its same finding.
|
||||
if (nextSchedule >= 0) assessment.push(...lines.slice(nextSchedule).filter(line => sameOwner.test(line)));
|
||||
return { schedule: header[1]!, assessment: assessment.join(' ') };
|
||||
}
|
||||
|
||||
/** Two adjacent original-schedule rows keep each operation in its named column. */
|
||||
function continuationStaleFillTrace(text: string, schedule: string, old: string, ttl: string): { key: string; writer: string } | undefined {
|
||||
const lines = text.split('\n').map(line => line.trim().replace(/\s+/g, ' ').replace(/→/g, '->')).filter(Boolean);
|
||||
const cells = (line: string) => line.split('|').map(cell => cell.trim());
|
||||
const header = cells(lines[0] ?? '');
|
||||
if (header.length !== 6 || header[0] !== 'Sched' || header[5] !== 'Result') return;
|
||||
const first = /^(R[1-9]\d*) \(miss, reads ([A-Za-z_$][\w$]*)\)$/.exec(header[1]!);
|
||||
const writer = /^(W[1-9]\d*|W) \(commits ([A-Za-z_$][\w$]*)\)$/.exec(header[2]!);
|
||||
const later = /^(R[1-9]\d*) \(begins after (W[1-9]\d*|W)\)$/.exec(header[3]!);
|
||||
const cache = /^cache\[([A-Za-z_$][\w$]*)\]$/.exec(header[4]!);
|
||||
if (!first || !writer || !later || !cache || first[1] === later[1] || writer[1] !== later[2]
|
||||
|| first[2] !== old || writer[2] === old) return;
|
||||
const rows = lines.slice(2).map(cells);
|
||||
if (cells(lines[1] ?? '').length !== 6 || !/^[-| ]+$/.test(lines[1] ?? '') || rows.some(row => row.length !== 6)) return;
|
||||
const matches = rows.flatMap((row, i) => row[0] === `${schedule}*` ? [i] : []);
|
||||
if (matches.length !== 1 || rows.some(row => row[0] === schedule)) return;
|
||||
const i = matches[0]!;
|
||||
const expected = [
|
||||
[`${schedule}*`, 'await read ...', 'write commits, delete(noop)', '', '-', ''],
|
||||
['', `resolves ${old} -> set ${old}`, '', `hit -> ${old}`, `${old} (${ttl} s)`, 'VIOLATION'],
|
||||
];
|
||||
if (expected.some((row, n) => row.some((cell, c) => rows[i + n]?.[c] !== cell))) return;
|
||||
// A third unlabeled row would still belong to this schedule and could
|
||||
// contradict the claimed late fill or later cache hit.
|
||||
if (rows[i + 2]?.[0] === '') return;
|
||||
return { key: cache[1]!, writer: writer[1]! };
|
||||
}
|
||||
|
||||
/** A named finding may put its ordering evidence in a trace, not one paragraph. */
|
||||
function assertedStructuredOwner(prose: string[], index: number): boolean {
|
||||
const owners: Array<{ level: number; title: string }> = [];
|
||||
for (const line of prose.slice(0, index + 1)) {
|
||||
const heading = /^(#{1,6})\s+(.+)$/.exec(line);
|
||||
if (!heading) continue;
|
||||
while (owners.length && owners.at(-1)!.level >= heading[1]!.length) owners.pop();
|
||||
owners.push({ level: heading[1]!.length, title: heading[2]! });
|
||||
}
|
||||
return !owners.some(owner => /\b(?:hypothetical|example|quoted|historical|template|source)\b/i.test(owner.title));
|
||||
}
|
||||
|
||||
/** A same-ID assessment survives intervening diagrams and named assessment headings. */
|
||||
function structuredFindingAssessment(prose: string[], finding: number, traceEnd: number, ids: string[], assertedOwner = assertedStructuredOwner): string[] {
|
||||
const assessment: string[] = [];
|
||||
const sameId = new RegExp(`^(?:(?:${ids.join('|')})\\b|\\|\\s*(?:${ids.join('|')})\\s*\\|)`);
|
||||
const ownsHeading = new RegExp(`\\b(?:${ids.join('|')})\\b`);
|
||||
let followingTrace = false, namedAssessment = false;
|
||||
for (let i = finding + 1; i < prose.length; i++) {
|
||||
const line = prose[i]!;
|
||||
if (i === traceEnd + 1) followingTrace = true;
|
||||
if (/^#{1,6}\s/.test(line)) {
|
||||
namedAssessment = ownsHeading.test(line);
|
||||
followingTrace = false;
|
||||
} else if (/^\||^[FSDA][1-9]\d*\b/.test(line) && !sameId.test(line)) {
|
||||
followingTrace = false;
|
||||
namedAssessment = false;
|
||||
}
|
||||
if (assertedOwner(prose, i) && (sameId.test(line) || namedAssessment || followingTrace)) {
|
||||
// A current table can put a scalar verdict in the cited finding's row.
|
||||
// Normalize that owned status only, never borrow a neighboring row.
|
||||
const status = /^\|\s*(F[1-9]\d*)\s*\|\s*["“']?(withdrawn|rejected|dismissed)\b/i.exec(line);
|
||||
assessment.push(status && ids.includes(status[1]!) ? `${status[1]} is ${status[2]}. ${line}` : line);
|
||||
}
|
||||
}
|
||||
return assessment;
|
||||
}
|
||||
|
||||
function hasStructuredStaleFillFinding(report: string): boolean {
|
||||
const lines = report.split('\n');
|
||||
const prose = lines.map(() => '');
|
||||
const traces: Array<{ start: number; end: number; text: string }> = [];
|
||||
let fence: { char: string; length: number; start: number; info: string } | null = null;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const line = lines[i]!;
|
||||
const delimiter = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (delimiter) {
|
||||
const run = delimiter[1]!;
|
||||
if (!fence) fence = { char: run[0]!, length: run.length, start: i, info: delimiter[2]!.trim() };
|
||||
else if (run[0] === fence.char && run.length >= fence.length && !delimiter[2]!.trim()) {
|
||||
if (!fence.info || fence.info === 'text') traces.push({ start: fence.start, end: i, text: lines.slice(fence.start + 1, i).join('\n') });
|
||||
fence = null;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
// Unclosed, tilde and longer fences remain source until their own real
|
||||
// closing delimiter. They cannot supply an asserted prose violation.
|
||||
if (!fence && !/^\s*>/.test(line) && !/^(?: {4}|\t)/.test(line)) prose[i] = line;
|
||||
}
|
||||
for (const trace of traces) {
|
||||
let heading = trace.start - 1;
|
||||
while (heading >= 0 && !prose[heading]!.trim()) heading--;
|
||||
const identity = /^Async ordering schedule \((F[1-9]\d*), invariant boundary = `writeProfile` settles\):$/.exec(prose[heading] ?? '');
|
||||
if (!identity || !assertedStructuredOwner(prose, heading)) continue;
|
||||
const framed = (line: string) => !line || /^(?:\||#{1,6}\s)/.test(line);
|
||||
let previous = heading - 1;
|
||||
while (previous >= 0 && !lines[previous]!.trim()) previous--;
|
||||
if (!framed(prose[previous] ?? '') || (!prose[previous] && previous >= 0 && !traces.some(other => other.end === previous))) continue;
|
||||
const findings = prose.slice(0, heading).map((line, index) => ({ index, cells: line.split('|').map(cell => cell.trim()) }))
|
||||
.filter(({ cells }) => cells.length === 8 && cells[0] === '' && cells[1] === identity[1] && cells[2] === 'CRITICAL');
|
||||
if (findings.length !== 1) continue;
|
||||
const finding = findings[0]!;
|
||||
if (!assertedStructuredOwner(prose, finding.index)) continue;
|
||||
let registryHeading = finding.index - 1;
|
||||
while (registryHeading >= 0 && !/^#{1,6}\s/.test(prose[registryHeading]!)) registryHeading--;
|
||||
if (!/^#{1,6} Findings registry$/.test(prose[registryHeading] ?? '')
|
||||
|| prose.slice(registryHeading + 1, finding.index).some(line => line.trim() && !/^\|/.test(line))) continue;
|
||||
const prefix = prose.slice(0, registryHeading).filter(line => line.trim()).at(-1) ?? '';
|
||||
const priorSection = prose.slice(0, registryHeading).filter(line => /^#{1,6}\s/.test(line)).at(-1) ?? '';
|
||||
const closesDecisions = /^Lake Score: [0-9]+\/[0-9]+ coverage-scored decisions \([A-Z0-9, -]+\) chose the complete option\.$/.test(prefix)
|
||||
&& /^#{1,6} Decision registry\b/.test(priorSection);
|
||||
if (!framed(prefix) && !closesDecisions) continue;
|
||||
const citation = /^Lines "no additional version checks or coordination between a cache fill and a write" vs invariant "every read begun after that write completes must observe the committed version"\. Schedule in Section [1-9]\d* shows a pre-write DB snapshot filled after `cache\.delete`, served up to ([1-9]\d*) s;/.exec(finding.cells[3] ?? '');
|
||||
const ordered = originalOrderAStaleFillTrace(trace.text);
|
||||
if (!citation || !ordered || citation[1] !== ordered.ttl) continue;
|
||||
const assessment = structuredFindingAssessment(prose, finding.index, trace.end, [identity[1]!]);
|
||||
const registry = [...finding.cells];
|
||||
// This exact residual allowance refers only to readers begun before W,
|
||||
// matching R1 in the validated header, never to the later R2 cache hit.
|
||||
if (registry[5] === 'Readers that began before the write may still see the old snapshot (permitted by contract)') {
|
||||
registry[5] = 'Allowed by contract: the original reader returns its earlier value.';
|
||||
}
|
||||
const context = [registry.join(' | '), ordered.assessment, ...assessment].join(' ');
|
||||
if (new RegExp(`\\b${identity[1]}\\s+(?:is|was|remains)\\s+(?:impossible|rejected|dismissed|withdrawn)\\b`, 'i').test(context)
|
||||
|| /\b(?:(?:this|that|the|original)\s+(?:trace|schedule|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(context)) continue;
|
||||
const claim = `Concurrent cache read fills an old value after the write committed and invalidated the same key. A new reader receives that stale value, which violates the read-after-write contract. ${context}`;
|
||||
if (hasProseStaleFillFinding(claim)) return true;
|
||||
}
|
||||
for (const trace of traces) {
|
||||
let heading = trace.start - 1;
|
||||
while (heading >= 0 && !prose[heading]!.trim()) heading--;
|
||||
const shared = /^\*\*[1-9]\d*\. Async ordering schedules \(shared state: cache\[([A-Za-z_$][\w$]*)\], writeGen\[([A-Za-z_$][\w$]*)\]\)\*\*$/.exec(prose[heading] ?? '');
|
||||
if (!shared || shared[1] !== shared[2] || !assertedStructuredOwner(prose, heading)) continue;
|
||||
const framed = (line: string) => !line || /^(?:\||#{1,6}\s|\*\*[0-9]+[A-Z]?\b)/.test(line);
|
||||
let previous = heading - 1;
|
||||
while (previous >= 0 && !lines[previous]!.trim()) previous--;
|
||||
if (!framed(prose[previous] ?? '') || (!prose[previous] && previous >= 0 && !traces.some(other => other.end === previous))) continue;
|
||||
const findings = prose.slice(0, heading).map((line, index) => ({ index, cells: line.split('|').map(cell => cell.trim()) }))
|
||||
.filter(({ cells }) => cells.length === 8 && cells[0] === '' && /^F[1-9]\d*$/.test(cells[1] ?? '') && /^P[0-3] CRITICAL$/.test(cells[2] ?? ''));
|
||||
for (const finding of findings) {
|
||||
if (!assertedStructuredOwner(prose, finding.index) || findings.filter(other => other.cells[1] === finding.cells[1]).length !== 1) continue;
|
||||
let registryHeading = finding.index - 1;
|
||||
while (registryHeading >= 0 && !/^#{1,6}\s/.test(prose[registryHeading]!)) registryHeading--;
|
||||
if (!/^#{1,6} Findings Registry$/.test(prose[registryHeading] ?? '')
|
||||
|| !framed(prose.slice(0, registryHeading).filter(line => line.trim()).at(-1) ?? '')) continue;
|
||||
const sequence = /^Late fill after write\. Read misses, DB returns ([A-Za-z][\w.-]*), write commits ([A-Za-z][\w.-]*) and deletes \(no-op\), read then fills ([A-Za-z][\w.-]*); every later read gets ([A-Za-z][\w.-]*) until TTL\.(?=\s|$)/.exec(finding.cells[3] ?? '');
|
||||
if (!sequence || sequence[1] === sequence[2] || sequence[1] !== sequence[3] || sequence[1] !== sequence[4]) continue;
|
||||
const ordered = versionedOriginalSchedule(trace.text, shared[1]!, sequence[1]!, sequence[2]!, finding.cells[1]!);
|
||||
if (!ordered || !new RegExp(`\\bschedules?\\s+${ordered.schedule}(?=[, .]|$)`).test(finding.cells[6] ?? '')) continue;
|
||||
const assessment = structuredFindingAssessment(prose, finding.index, trace.end, [finding.cells[1]!, ordered.schedule]);
|
||||
const context = [finding.cells.join(' | '), ordered.assessment, ...assessment].join(' ');
|
||||
if (new RegExp(`\\b(?:${finding.cells[1]}|${ordered.schedule})\\s+(?:is|was|remains)\\s+(?:impossible|rejected|dismissed|withdrawn)\\b`, 'i').test(context)
|
||||
|| /\b(?:(?:this|that|the|original)\s+(?:trace|schedule|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(context)) continue;
|
||||
const claim = `Concurrent cache read fills an old value after the write committed and invalidated the same key. A new reader receives that stale value, which violates the read-after-write contract. ${context}`;
|
||||
if (hasProseStaleFillFinding(claim)) return true;
|
||||
}
|
||||
}
|
||||
for (const trace of traces) {
|
||||
let heading = trace.start - 1;
|
||||
while (heading >= 0 && !/^#{1,6}\s/.test(prose[heading]!)) heading--;
|
||||
const section = /^#{1,6} Async Ordering Record \(Section ([1-9]\d*)\)$/.exec(prose[heading] ?? '');
|
||||
if (!section) continue;
|
||||
const framed = (prefix: string) => !prefix || /^(?:\||#{1,6}\s)/.test(prefix);
|
||||
if (!framed(prose.slice(0, heading).filter(line => line.trim()).at(-1) ?? '')) continue;
|
||||
const ownership = prose.slice(heading + 1, trace.start).join(' ').replace(/\s+/g, ' ').trim();
|
||||
const shared = /^Shared state: `cache\[([A-Za-z_$][\w$]*)\]`, `inflight\[([A-Za-z_$][\w$]*)\]`\. Invariant boundary: a read that \*begins\* after `writeProfile` resolves must return the committed version\.$/.exec(ownership);
|
||||
if (!shared || shared[1] !== shared[2]) continue;
|
||||
const findings = prose.slice(0, heading).map((line, index) => ({ index, cells: line.split('|').map(cell => cell.trim()) }))
|
||||
.filter(({ cells }) => cells[0] === '' && /^F[1-9]\d*$/.test(cells[1] ?? '') && cells[2] === 'CRITICAL GAP');
|
||||
for (const finding of findings) {
|
||||
let registryHeading = finding.index - 1;
|
||||
while (registryHeading >= 0 && !/^#{1,6}\s/.test(prose[registryHeading]!)) registryHeading--;
|
||||
const prefix = prose.slice(0, registryHeading).filter(line => line.trim()).at(-1) ?? '';
|
||||
const previousSection = prose.slice(0, registryHeading).filter(line => /^#{1,6}\s/.test(line)).at(-1) ?? '';
|
||||
// A completed decision registry's score closes that earlier section;
|
||||
// an unheaded example/hypothesis cannot introduce the current finding.
|
||||
const closesDecisions = /^Lake Score: [0-9]+\/[0-9]+ recommendations chose the complete option\.$/.test(prefix)
|
||||
&& /^#{1,6} Decision Registry(?: \(all auto-resolved to recommended option\))?$/.test(previousSection);
|
||||
if (!/^#{1,6} Findings Registry$/.test(prose[registryHeading] ?? '')
|
||||
|| (!framed(prefix) && !closesDecisions)
|
||||
|| prose.slice(registryHeading + 1, finding.index).some(line => line.trim() && !/^\|/.test(line))) continue;
|
||||
if (!finding.cells[3]?.split(/,\s*/).includes(section[1]!)) continue;
|
||||
if (!framed(prose.slice(0, finding.index).filter(line => line.trim()).at(-1) ?? '')) continue;
|
||||
const evidence = (finding.cells[4] ?? '').replace(/`/g, '');
|
||||
const citation = /^Original sketch: fill after await repository\.read has no guard; plan text says no fill\/write coordination\. Schedule (S[1-9]\d*) makes a post-write reader see ([A-Za-z_$][\w$]*) for ([1-9]\d*) s[;.]/.exec(evidence);
|
||||
if (!citation || !/(?:^|[.;]\s+)Violates retained invariant\.$/.test(evidence)) continue;
|
||||
if (findings.filter(other => other.cells[1] === finding.cells[1]).length !== 1) continue;
|
||||
const ordered = continuationStaleFillTrace(trace.text, citation[1]!, citation[2]!, citation[3]!);
|
||||
if (!ordered || ordered.key !== shared[1]) continue;
|
||||
const assessment: string[] = [];
|
||||
for (let i = trace.end + 1; i < prose.length; i++) {
|
||||
const named = /^([FSDA][1-9]\d*)\b/.exec(prose[i]!);
|
||||
if (/^(?:#{1,6}\s|\|)/.test(prose[i]!) || (named && named[1] !== citation[1] && named[1] !== finding.cells[1])) break;
|
||||
assessment.push(prose[i]!);
|
||||
}
|
||||
const tail = assessment.join(' ').replace(/\s+/g, ' ').trim();
|
||||
if (!/^`\*` = original sketch\.(?:\s|$)/.test(tail)) continue;
|
||||
const withdrawn = new RegExp(`\\b(?:${citation[1]}|${finding.cells[1]})\\s+(?:is|was|remains)\\s+(?:impossible|rejected|dismissed|withdrawn)\\b`, 'i');
|
||||
if (withdrawn.test(tail)) continue;
|
||||
if (/\b(?:(?:this|that|the|original)\s+(?:trace|schedule|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(tail)) continue;
|
||||
// Only the validated original rows establish the race. The registry and
|
||||
// following assessment still govern whether the report dismisses it.
|
||||
const claim = `Concurrent cache read fills an old value after the write committed and invalidated the same key. A new reader receives that stale value, which violates the read-after-write contract. ${finding.cells.join(' | ')} ${tail}`;
|
||||
if (hasProseStaleFillFinding(claim)) return true;
|
||||
}
|
||||
}
|
||||
for (const trace of traces) {
|
||||
let heading = trace.start - 1;
|
||||
while (heading >= 0 && !prose[heading]!.trim()) heading--;
|
||||
const identity = /^(?:#{1,6}\s+)?(?:\*\*)?[1-9]\d*\. Async schedule \((F[1-9]\d*)\)(?: with one column per operation and shared state)?(?:\*\*)?$/.exec(prose[heading] ?? '');
|
||||
if (!identity) continue;
|
||||
// Unheaded prose immediately introducing this finding or schedule owns
|
||||
// its assertion status, regardless of whether it ends in ':' or '.'.
|
||||
// Require a fresh structural boundary instead of dropping that context.
|
||||
const framed = (prefix: string) => !prefix || /^(?:\||#{1,6}\s|\*\*\d+\.\s)/.test(prefix);
|
||||
const prefix = prose.slice(0, heading).filter(line => line.trim()).at(-1) ?? '';
|
||||
if (!framed(prefix)) continue;
|
||||
const findings = prose.slice(0, heading).map((line, index) => ({ index, cells: line.split('|').map(cell => cell.trim()) }))
|
||||
.filter(({ cells }) => cells[0] === '' && cells[1] === identity[1] && /^CRITICAL(?: GAP(?: \(fixed by D[1-9]\d*\))?)?$/.test(cells[3] ?? ''));
|
||||
if (findings.length !== 1) continue;
|
||||
const findingPrefix = prose.slice(0, findings[0]!.index).filter(line => line.trim()).at(-1) ?? '';
|
||||
if (!framed(findingPrefix)) continue;
|
||||
const alternate = alternateOrderStaleFillTrace(trace.text);
|
||||
const original = originalSketchStaleFillTrace(trace.text);
|
||||
const evidence = findings[0]!.cells[4] ?? '';
|
||||
const ordered = alternate && /(?:^|[.;]\s+)Schedule below shows\b/.test(evidence) ? alternate
|
||||
: original && evidence.split(/(?<=\.)\s+/).includes(original.summary) ? original.assessment : undefined;
|
||||
if (!ordered) continue;
|
||||
const assessment: string[] = [];
|
||||
for (let i = trace.end + 1; i < prose.length; i++) {
|
||||
if (/^(?:#{1,6}\s|\*\*\d+\.|[DSF][1-9]\d*\b|\|)/.test(prose[i]!)) break;
|
||||
assessment.push(prose[i]!);
|
||||
}
|
||||
if (/\b(?:(?:this|that|the|original)\s+(?:trace|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test([ordered, ...assessment].join(' '))) continue;
|
||||
// A selected repair's invalidated-fill rule describes the amendment. It
|
||||
// does not assert that the original late fill was already impossible.
|
||||
const registry = [...findings[0]!.cells];
|
||||
const repair = /\(fixed by (D[1-9]\d*)\)$/.exec(registry[3] ?? '')?.[1];
|
||||
if (repair && registry[5]?.startsWith(`${repair}: `)) {
|
||||
registry[5] = registry[5].replace(/\binvalidated fills never `?set`?\b/g, 'discard invalidated fills');
|
||||
}
|
||||
const claim = `Concurrent cache read fills an old value after the write committed and invalidated the same key. A new reader receives that stale value, which violates the read-after-write contract. ${registry.join(' | ')} ${ordered} ${assessment.join(' ')}`;
|
||||
if (hasProseStaleFillFinding(claim)) return true;
|
||||
}
|
||||
for (const trace of traces) {
|
||||
let heading = trace.start - 1;
|
||||
while (heading >= 0 && !prose[heading]!.trim()) heading--;
|
||||
const identity = /^(S[1-9]\d*) Async ordering schedule \((F[1-9]\d*) evidence\):$/.exec(prose[heading] ?? '');
|
||||
if (!identity) continue;
|
||||
const previous = prose.slice(0, heading).filter(line => line.trim()).at(-1) ?? '';
|
||||
if (/^(?!\||#{1,6}\s).*:\s*$/.test(previous)) continue;
|
||||
const finding = prose.slice(0, heading).map((line, index) => ({ index, cells: line.split('|').map(cell => cell.trim()) }))
|
||||
.filter(({ cells }) => cells[0] === '' && cells[1] === identity[2] && cells[2] === 'CRITICAL GAP');
|
||||
if (finding.length !== 1 || !finding[0]!.cells[4]?.includes(`Schedule ${identity[1]} below`)) continue;
|
||||
const findingPrefix = prose.slice(0, finding[0]!.index).filter(line => line.trim()).at(-1) ?? '';
|
||||
if (/^(?!\||#{1,6}\s).*:\s*$/.test(findingPrefix)) continue;
|
||||
const ordered = columnarStaleFillTrace(trace.text);
|
||||
if (!ordered) continue;
|
||||
// The verified columns establish this claim. Keep the real finding and
|
||||
// trace assessment in the existing dismissal checks; an allowance for
|
||||
// R1's own pre-write return cannot authorize a stale cache or later reader.
|
||||
const allowance = `Allowed by contract: ${ordered.reader} itself returns ${ordered.old} (read in progress when write committed).`;
|
||||
const tail = ordered.tail.replace(allowance, 'Allowed by contract: the original reader returns its earlier value.');
|
||||
const assessment: string[] = [];
|
||||
for (let i = trace.end + 1; i < prose.length; i++) {
|
||||
if (/^(?:#{1,6}\s|[DSF][1-9]\d*\b|\|)/.test(prose[i]!)) break;
|
||||
assessment.push(prose[i]!);
|
||||
}
|
||||
// A scored, unselected alternative in an accepted decision is not the
|
||||
// verdict. Other quotes remain in the assessment, including a directly
|
||||
// quoted rejection of the finding itself.
|
||||
const registry = finding[0]!.cells.map(cell => /^Accepted\b/.test(cell)
|
||||
? cell.replace(/\bvs\s+\d+(?:\.\d+)?\/10\s+for\s+(?:"(?:[^"\\]|\\.)*"|“[^”]*”)/g, 'unselected alternative')
|
||||
: cell).join(' | ');
|
||||
const claim = `Concurrent cache read fills an old value after the write committed and invalidated the same key. A new reader receives that stale value, which violates the read-after-write contract. ${registry} ${tail} ${assessment.join(' ')}`;
|
||||
if (hasProseStaleFillFinding(claim)) return true;
|
||||
}
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const legacy = /^\*\*CRITICAL FINDING\s*[—–:-].*\*\*\s*$/.test(prose[i]!);
|
||||
let previousIndex = i - 1;
|
||||
while (previousIndex >= 0 && !prose[previousIndex]!.trim()) previousIndex--;
|
||||
const numbered = /^\*\*CRITICAL GAP\*\*\s*[—–:-]/.test(prose[i]!)
|
||||
&& /^#{1,6}\s+Critical Finding:\s+\S.*$/i.test(prose[previousIndex] ?? '');
|
||||
if (!legacy && !numbered) continue;
|
||||
const previous = prose.slice(0, numbered ? previousIndex : i).filter(value => value.trim()).at(-1) ?? '';
|
||||
if (/\b(?:example|template|source|quoted|format)\b[^.]*:\s*$/i.test(previous)) continue;
|
||||
let end = i + 1;
|
||||
while (end < lines.length && !/^(?:#{1,6}\s|\*\*(?:(?:CRITICAL|HIGH|MEDIUM|LOW)\s+)?(?:FINDING|GAP)\b)/i.test(prose[end]!)) end++;
|
||||
const claim = prose.slice(numbered ? i : i + 1, end).join('\n');
|
||||
if (numbered) {
|
||||
// A quoted requirement alone is insufficient: the same finding must
|
||||
// independently assert that the current wrapper violates it.
|
||||
if (!/^\*\*CRITICAL GAP\*\*\s*[—–:-]\s*The plan states: "Every read begun after that write completes must observe the committed version\.(?: TTL expiry is not a substitute for this rule\.)?" The proposed wrapper violates this invariant\.\s*$/m.test(claim)) continue;
|
||||
for (const trace of traces.filter(trace => trace.start > i && trace.end < end)) {
|
||||
if (!hasNumberedStaleFillTrace(trace.text)) continue;
|
||||
// Only this validated same-finding trace becomes prose evidence.
|
||||
// Reuse all existing dismissal/accepted-staleness checks unchanged;
|
||||
// this recognizes a finding, not the correctness of its proposed fix.
|
||||
if (hasProseStaleFillFinding((claim + '\n' + trace.text).replace(/\s+/g, ' '))) return true;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
if (!/^This\s+violates\s+the\s+stated\s+(?:invariant|contract):/m.test(claim) ||
|
||||
!/Every read begun after that write\s+completes must observe the committed version/.test(claim)) continue;
|
||||
if (/\b(?:not\s+(?:a\s+)?(?:gap|bug|defect|issue)|no\s+(?:fix|change|guard)\s+(?:is\s+)?(?:needed|required))\b/i.test(claim)) continue;
|
||||
for (const trace of traces.filter(trace => trace.start > i && trace.end < end)) {
|
||||
// All four ordered events and the post-write new reader must be shown.
|
||||
// A copied wrapper has neither this execution trace nor an independent
|
||||
// asserted violation in the same finding.
|
||||
if (/await\s+repository\.read[\s\S]*repository\.write[\s\S]*cache\.delete[\s\S]*cache\.set\([^\n]*(?:old|stale)[^\n]*\)[\s\S]*readProfile\([^\n]*started after[^\n]*[\s\S]*cache\.get[^\n]*(?:old|stale)/i.test(trace.text)) return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** Require an unresolved late-fill defect, not a keyword-bearing dismissal. */
|
||||
export function hasStaleFillRaceFinding(report: string): boolean {
|
||||
return hasStructuredStaleFillFinding(report) || hasProseStaleFillFinding(report);
|
||||
}
|
||||
|
||||
/** An ordered execution can establish overlap without naming it "in-flight". */
|
||||
function hasOrderedStaleFillOperations(text: string, sourceText = text): boolean {
|
||||
const separator = String.raw`\s*[,;.]\s*(?:then\s+)?`;
|
||||
const subject = String.raw`(?:(?:a|the)\s+)?`;
|
||||
const sameObject = String.raw`(?:\s+(?:(?:the\s+)?same\s+)?(?:cache\s+)?(?:key|entry))?`;
|
||||
const read = String.raw`${subject}read\s+(?:misses|gets\s+a\s+cache\s+miss)`;
|
||||
const write = String.raw`${subject}write\s+commits\s+(?:and|then)\s+(?:deletes|invalidates|evicts)${sameObject}(?:\s*\(no-op\))?`;
|
||||
const fill = String.raw`${subject}(?:(?:original|same)\s+)?reader\s+(?:(?:then|later)\s+)?(?:fills|refills|repopulates)\s+(?:the\s+)?(?:old|stale|pre[- ](?:write|commit))\s+(?:snapshot|value|data)`;
|
||||
const later = String.raw`(?:(?:every|all|the)\s+)?(?:later|next|new|subsequent)\s+readers?\s+(?:sees?|gets?|observes?|receives?)\s+(?:the\s+)?(?:stale|old|outdated)\s+(?:data|value|snapshot)`;
|
||||
const findingPrefix = String.raw`(?:(?:F[1-9]\d*|(?:Finding|Issue)\s+[1-9]\d*)\s*[—–:-]\s*)?(?:P[0-3]\s*[—–:-]\s*)?`;
|
||||
const sequence = new RegExp(String.raw`^${findingPrefix}${read}${separator}${write}${separator}${fill}${separator}${later}(?=[\s.!?;]|$)`, 'i');
|
||||
// An asserted schedule can name the read's resolution and store separately.
|
||||
// All four operations must remain in one cell and in their causal order;
|
||||
// quoted requirements may follow, but inline code cannot supply operations.
|
||||
const resolvedRead = String.raw`${subject}(?:(?:original|same)\s+)?read(?:er)?\s+resolves\s+and\s+stores\s+(?:the\s+)?(?:old|stale|pre[- ](?:write|commit))\s+(?:snapshot|value|data)`;
|
||||
const staleHit = String.raw`(?:(?:a|the)\s+)?(?:later|next|new|subsequent)\s+read\s+hits\s+(?:the\s+)?(?:stale|old)\s+(?:value|data|snapshot)`;
|
||||
const assertedSchedule = new RegExp(String.raw`^Original (?:plan|sketch|wrapper)\b[^.?]*\.\s+Schedule:\s*${read}${separator}${write}${separator}${resolvedRead}\.\s+${staleHit}(?=[\s.!?;]|$)`, 'i');
|
||||
const scheduleCells = sourceText.replace(/`[^`]*`/g, '[literal]').replace(/[*_]/g, '')
|
||||
.replace(/\s+/g, ' ').trim().split(/\s*\|\s*/);
|
||||
if (scheduleCells.some(cell => {
|
||||
const unquoted = cell.replace(/"(?:[^"\\]|\\.)*"|“[^”]*”/g, '[quotation]');
|
||||
return assertedSchedule.test(unquoted)
|
||||
&& !/\b(?:if|unless|whether|might|may|could|not|never|no\s+longer|example|template|hypothetical|historical|quoted|copied|source|earlier\s+review)\b/i.test(unquoted)
|
||||
&& !/\b(?:another|different|separate|unrelated|other)\s+(?:cache|key|entry|reader|read|request)\b/i.test(unquoted)
|
||||
&& !/\b(?:(?:this|that|the)\s+(?:trace|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(unquoted);
|
||||
})) return true;
|
||||
// Table cells cannot lend operation order to each other. Bare "reader"
|
||||
// refers back to the missed read; explicit foreign cache/key references,
|
||||
// quoted examples and conditional/negated executions cannot establish it.
|
||||
return text.split(/\s*\|\s*/).some(cell => {
|
||||
// A finding may describe a fill's lifetime instead of naming the first
|
||||
// cache miss. Its original-plan sequence must still place the same old
|
||||
// version in the cache after commit/delete and deliver it to a later read.
|
||||
// Quoted requirements can accompany that assertion, but quoted operations
|
||||
// cannot supply it; never join fragments across a quoted span or table cell.
|
||||
const unquoted = cell.replace(/"(?:[^"\\]|\\.)*"|“[^”]*”/g, '[quoted]');
|
||||
const compact = /^Original (?:plan|sketch|wrapper)\b[^.!?]*\.\s+(?:Schedule(?: Diagram)? [A-Za-z0-9][\w.-]*:\s*)?fill starts,\s*write commits,\s*write deletes(?:\s*\(no-op\))?,\s*fill sets pre-commit ([A-Za-z0-9][\w.-]*),\s*later read hits ([A-Za-z0-9][\w.-]*)\.(?=\s|$)/i.exec(unquoted);
|
||||
if (compact && compact[1] === compact[2]
|
||||
&& !/\b(?:if|unless|whether|might|may|could|never|no\s+longer|example|template|hypothetical|historical)\b/i.test(unquoted)
|
||||
&& !/\b(?:another|different|separate|unrelated|other)\s+(?:cache|key|entry|reader|read|request)\b/i.test(unquoted)
|
||||
&& !/\b(?:trace|scenario|execution|sequence)\s+(?:is|was|remains)\s+impossible\b/i.test(unquoted)) return true;
|
||||
return !/["“”?]|\b(?:if|unless|whether|might|may|could|not|never|no\s+longer|example|template|quoted)\b/i.test(cell)
|
||||
&& !/\b(?:another|different|separate|unrelated|other)\s+(?:cache|key|entry|reader|read|request)\b/i.test(cell)
|
||||
&& !/\b(?:(?:this|that|the)\s+(?:trace|scenario|execution|sequence)|this|that|it)\s+(?:is|was|remains)\s+impossible\b/i.test(cell)
|
||||
&& sequence.test(cell);
|
||||
});
|
||||
}
|
||||
|
||||
/** Explicit copied/example framing owns its section and descendant headings. */
|
||||
function assertedProseOwner(prose: string[], index: number): boolean {
|
||||
const owners = [{ level: 0, source: false }];
|
||||
for (const line of prose.slice(0, index + 1)) {
|
||||
const heading = /^(#{1,6})\s+(.+)$/.exec(line);
|
||||
if (heading) {
|
||||
while (owners.length > 1 && owners.at(-1)!.level >= heading[1]!.length) owners.pop();
|
||||
// An explicit fresh review ends an unheaded introductory source block.
|
||||
if (/^(?:Current|Actual)\s+(?:review|findings|assessment)\b/i.test(heading[2]!)) owners[0]!.source = false;
|
||||
owners.push({ level: heading[1]!.length, source: /\b(?:hypothetical|examples?|quoted|copied|historical|template|source)\b/i.test(heading[2]!) });
|
||||
} else if (/^(?:Source|Earlier review):\s*$/i.test(line.trim())
|
||||
|| /^Hypothetical scenario[.:](?:\s|$)/i.test(line.trim())
|
||||
|| /\b(?:unproven\s+hypothesis|(?:hypothetical|historical)\s+example|quoted\s+source)\b/i.test(line)
|
||||
|| /^(?:The\s+following\b|Below\b|This\s+(?:section|material|example)\b)[^.!?]*\b(?:copied|quoted|source|examples?|hypothetical|historical|template)\b/i.test(line.trim())) {
|
||||
owners.at(-1)!.source = true;
|
||||
}
|
||||
}
|
||||
return !owners.some(owner => owner.source);
|
||||
}
|
||||
|
||||
function hasProseStaleFillFinding(report: string): boolean {
|
||||
// Copied source, diagrams and quoted examples cannot supply a finding.
|
||||
let fence: { char: string; length: number } | null = null;
|
||||
const prose = report.split('\n').map(line => {
|
||||
const delimiter = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (delimiter) {
|
||||
const run = delimiter[1]!;
|
||||
if (!fence) fence = { char: run[0]!, length: run.length };
|
||||
else if (run[0] === fence.char && run.length >= fence.length && !delimiter[2]!.trim()) fence = null;
|
||||
return '';
|
||||
}
|
||||
return fence || /^\s*>/.test(line) || /^(?: {4}|\t)/.test(line) ? '' : line;
|
||||
}).join('\n');
|
||||
// Independent list items and table rows cannot borrow each other's words.
|
||||
const blocks = prose.split(/\n\s*\n|\n(?=\s*(?:#{1,6}\s|\||\d+\.\s|[-*]\s))/).map(block => block.trim());
|
||||
const lines = blocks.flatMap(block => block.split('\n'));
|
||||
const normalize = (text: string) => text.replace(/[*_`]/g, '').replace(/\s+/g, ' ').trim();
|
||||
return blocks.some((block, index) => {
|
||||
const text = normalize(block);
|
||||
const owners = blocks.slice(0, index + 1).flatMap(part => part.split('\n'));
|
||||
if (!assertedProseOwner(owners, owners.length - 1)) return false;
|
||||
const stale = /\b(?:stale|outdated)\b|\b(?:old(?:er)?|pre[- ]write)\s+(?:value|data|result|version|snapshot)\b/i.test(text);
|
||||
const inFlight = /\b(?:race|racing|concurrent|concurrency|in[- ]flight|pending)\b/i.test(text)
|
||||
|| hasOrderedStaleFillOperations(text, block);
|
||||
const read = /\b(?:read|fetch)\w*\b/i.test(text);
|
||||
const fillPattern = /\b(?:fill|refill|repopulat|populat|insert|stor|restor)\w*\b|\bcache\.set\b|\bcache(?:s|d)?\s+(?:the|an?|old|stale|same)\s+(?:\w+\s+){0,2}(?:value|data|result|snapshot)\b/i;
|
||||
const fill = fillPattern.test(text);
|
||||
const invalidation = /\b(?:invalidat|evict|write|commit|delet)\w*\b/i.test(text);
|
||||
const ordering = /\b(?:after|later|resum\w*)\b|out[- ]of[- ]order/i.test(text);
|
||||
// A review may identify the ordering defect directly as missing coordination
|
||||
// between cache fills and writes that violates read-after-write freshness.
|
||||
// That is independent evidence even when the old-value trace is a diagram.
|
||||
// Inline source cannot supply the assertion; the amendment label is metadata.
|
||||
const coordinationText = normalize(block.replace(/`([^`]*)`/g, (_span, body: string) =>
|
||||
/^\[Amended:[^\]]+\]$/.test(body) ? body : '[literal]'));
|
||||
const premise = /(?:^|[.;]\s+)(?:\[Amended:[^\]]{1,80}\]\s*)?(?:the\s+)?(?:original|current|proposed)\s+(sketch|wrapper|implementation)\s+(?:(?:had|has|proposed)\s+no\s+coordination\s+between\s+(?:an?\s+)?cache\s+fill\s+and\s+(?:an?\s+)?write\b|stated\s+that\s+no\s+coordination\s+between\s+(?:an?\s+)?cache\s+fill\s+and\s+(?:an?\s+)?write\s+was\s+proposed\b)/i.exec(coordinationText);
|
||||
const citedConclusion = /(?:^|[.;]\s+)Review\s+(?:showed|shows)\s+that\s+(sketch|wrapper|implementation)\s+(?:violates|breaks)\s+the\s+(?:retained\s+)?read[- ]after[- ]write\s+(?:rule|contract|guarantee|invariant)\s+\(see\s+(F[1-9]\d*)\)(?:[.!](?=\s|$)|$)/i.exec(coordinationText);
|
||||
// The reviewer can assert the original coordination violation directly,
|
||||
// immediately after its premise, without naming the old version 'stale'.
|
||||
const reportedViolation = /(?:^|[.;]\s+)(?:the\s+)?review\s+found\s+that\s+this\s+(?:violates|breaks)\s+the\s+read[- ]after[- ]write\s+(?:rule|contract|guarantee|invariant)(?:\s+above)?\s+\((F[1-9]\d*)\)(?=\s+and\s+omits\b|[.!](?:\s|$)|$)/i.exec(coordinationText);
|
||||
const conclusion = /(?:^|[.;]\s+)(?:finding\s+[\w.-]+\s+(?:showed|shows)\s+)?(?:this|that|it)\s+(?:violates|breaks)\s+the\s+read[- ]after[- ]write\s+(?:rule|contract|guarantee|invariant)(?:[.!](?=\s|$)|$)/i.exec(coordinationText) ?? citedConclusion ?? reportedViolation;
|
||||
const coordinationGap = premise !== null && conclusion !== null && premise.index < conclusion.index
|
||||
&& (conclusion !== citedConclusion || premise[1]!.toLowerCase() === citedConclusion![1]!.toLowerCase())
|
||||
&& (conclusion !== reportedViolation || /^\s*$/.test(coordinationText.slice(premise.index + premise[0].length, conclusion.index)))
|
||||
&& !coordinationText.slice(premise.index, conclusion.index).includes('|')
|
||||
&& !/["“”]|\b(?:if|example|template|quoted)\b/i.test(text)
|
||||
&& !/\b(?:example|template|source|quoted|format)\b[^.]*:\s*$/i.test(blocks[index - 1] ?? '');
|
||||
if ((!stale || !inFlight || !read || !fill || !invalidation || !ordering) && !coordinationGap) return false;
|
||||
|
||||
// A neighboring explanation/remedy belongs to this paragraph only until
|
||||
// another named finding/section/table row begins. In particular, a
|
||||
// following dismissal cannot turn a traced race into positive coverage.
|
||||
const next = blocks[index + 1] ?? '';
|
||||
const independent = /^(?:#{1,6}(?:\s|\d)|\d+\.\s|[-*]\s|\||(?:[*_]+)?(?:Finding\b|Section\s|P[0-3]\b))/i.test(next);
|
||||
const explicitId = /^\|\s*(F[1-9]\d*)\s*\|/.exec(text)?.[1]
|
||||
?? (coordinationGap && conclusion === citedConclusion ? citedConclusion?.[2] : undefined)
|
||||
?? (coordinationGap && conclusion === reportedViolation ? reportedViolation?.[1] : undefined);
|
||||
const assessment = explicitId
|
||||
? structuredFindingAssessment(lines, owners.length - 1, owners.length - 1, [explicitId], assertedProseOwner).join(' ')
|
||||
: '';
|
||||
const context = text + (independent ? '' : ' ' + normalize(next)) + ' ' + normalize(assessment);
|
||||
const findingId = explicitId ?? 'F[1-9]\\d*';
|
||||
if (new RegExp(`\\b(?:(?:this|that|the)\\s+(?:finding|issue|gap|race)|${findingId})\\s+(?:is|was|remains)\\s+["“'‘]?(?:withdrawn|rejected|dismissed)\\b`, 'i').test(context)) return false;
|
||||
|
||||
const finding = /\b(?:P[0-3]|missing|gap|bug|defect|violat\w*|unsafe|incorrect)\b|\bno\s+mention\s+of\s+(?:this|the)\s+race\b/i.test(context);
|
||||
const subsequentRead = /\b(?:next|later|subsequent|new|fresh|future)\s+(?:read\w*|request\w*|caller\w*)\b/i.test(context);
|
||||
const remedy = context.split(/[.!?]\s+/).some(sentence =>
|
||||
(/\b(?:guard|serialize|serialise|coordinate|prevent|reject|skip)\w*\b/i.test(sentence) &&
|
||||
/\b(?:cache|fill|refill|write|mutation|invalidation)\w*\b/i.test(sentence)) ||
|
||||
/\bper[- ]key\s+(?:epoch|generation|version)\b/i.test(sentence));
|
||||
// Table cells and semicolon-separated statements have separate owners;
|
||||
// retain an explicit "that return" continuation with the return it names.
|
||||
const claims = context.split(/\s*\|\s*/).flatMap(cell =>
|
||||
cell.split(/(?:[.!?]\s+|;\s+(?!that\s+return\b)|\b(?:but|however|nevertheless|yet)\s*[:,]?\s+)/i));
|
||||
const violation = claims.some(claim => /\b(?:violat\w*|break\w*)\b[^.!?]*\b(?:contract|guarantee|consistency|rule)\b/i.test(claim)
|
||||
&& !/\b(?:not|no|never)\b/i.test(claim));
|
||||
for (const [claimIndex, claim] of claims.entries()) {
|
||||
// "Not permitted" is a violation assertion, not permission. Scope a
|
||||
// permitted old result to its original caller; it cannot justify a
|
||||
// cache fill or a later reader observing that same old version.
|
||||
const allowanceText = claim.replace(/\b(?:not|never)\s+(?:an?\s+)?(?:permitted|allowed|acceptable|accepted)\b/gi, 'forbidden');
|
||||
// An imperative's purpose clause describes the proposed guard's goal,
|
||||
// not a claim that the current implementation already prevents the race.
|
||||
const proposedPrevention = /^(?:guard|serialize|serialise|coordinate|prevent|reject|skip)\b/i.test(claim.trim())
|
||||
&& /\b(?:cache|fill|refill|write|mutation|invalidation)\w*\b/i.test(claim)
|
||||
&& /\b(?:so(?:\s+that)?|to\s+ensure)\b/i.test(claim);
|
||||
// The model declaration must accept the stale consequence itself.
|
||||
// A normative freshness requirement called an accepted model is not a
|
||||
// dismissal. Bare "This" can refer only to the preceding stale claim.
|
||||
const modelDeclaration = /^(.+?)\s+(?:is|remains)\s+(?:(?:the|an?)\s+)?(?:accepted|expected|intentional|documented)\s+consistency\s+(?:model|contract|policy|semantics)\b/i.exec(allowanceText.trim())
|
||||
?? /^(.+?)\s+(?:is|remains)\s+(?:accepted|expected|intentional|documented)[.!?]?$/i.exec(allowanceText.trim());
|
||||
const subject = modelDeclaration?.[1] ?? '';
|
||||
const previousClaim = claims[claimIndex - 1] ?? '';
|
||||
const explicitStaleSubject = /^(?:this|the|an?)\s+(?:bounded\s+)?(?:inconsistency|staleness|stale[- ](?:read|fill)|stale\s+(?:read|fill|refill))(?:\s+(?:window|behavior|behaviour|race|consequence))?$/i.test(subject);
|
||||
const impliedStaleSubject = /^this$/i.test(subject)
|
||||
&& /\b(?:stale|outdated|old(?:er)?\s+(?:value|snapshot|data)|pre[- ]write\s+(?:value|data|result|version|snapshot))\b/i.test(previousClaim)
|
||||
&& /\b(?:read|fetch|fill|refill|repopulat)\w*\b/i.test(previousClaim)
|
||||
&& !/\b(?:must|shall|requires?|violat\w*|not|cannot|can't)\b/i.test(previousClaim);
|
||||
const acceptedStaleModel = Boolean(modelDeclaration) && (explicitStaleSubject || impliedStaleSubject);
|
||||
const dismissal = /\b(?:not|isn't)\s+(?:a\s+|an\s+)?(?:(?:stale|late)[- ]fill\s+)?(?:gap|bug|defect|issue|violation|problem|race)\b|\bno\s+(?:(?:stale|late)[- ]fill\s+)?(?:gap|bug|defect|issue|violation|race)\b/i.test(claim)
|
||||
|| /\b(?:accepted|expected|intentional|documented)\s+(?:invariant|behavior|trade[- ]off|stale[- ]read\s+window)\b|\b(?:allowed|permitted|acceptable)\b/i.test(allowanceText)
|
||||
|| acceptedStaleModel
|
||||
|| (!proposedPrevention && /\b(?:cannot|can't|never|does not|will not)\s+(?:\w+\s+){0,3}(?:refill|repopulate|populate|insert|store|cache|set|violate)\b/i.test(claim))
|
||||
|| (!proposedPrevention && /\b(?:cannot|can't|never|does not|doesn't|will not|won't|did not|didn't|is not|isn't|was not|wasn't|has not|hasn't|had not|hadn't)\s+(?:\w+\s+){0,3}restor\w*\b/i.test(claim))
|
||||
|| /\bno\s+(?:fix|change|coordination|guard)\s+(?:is\s+)?(?:needed|required)\b/i.test(claim);
|
||||
if (!dismissal) continue;
|
||||
const originalCaller = /\b(?:original|already[- ]pending)\s+(?:pending\s+)?(?:caller|reader|request)\b|\bpending\s+caller\b/i.test(claim);
|
||||
// A finding can name versions instead of calling them "old". Explicit
|
||||
// start-before-commit and return-to-own-caller evidence scopes this
|
||||
// allowance to that already-started call, never to cache/later readers.
|
||||
const explicitlyEarlierCall = originalCaller && /\bto\s+its\s+own\s+caller\b/i.test(claim)
|
||||
&& !/\b(?:if|unless|whether|might|may|could)\b/i.test(claim)
|
||||
&& /\b(?:it|(?:the\s+)?(?:original\s+)?(?:read|request|call))\s+(?:began|started)\s+before\s+(?:(?:the|that)\s+)?(?:write\s+)?commit\b/i.test(claim);
|
||||
const onlyEarlierReturn = originalCaller && /\b(?:return|receiv|observ)\w*\b/i.test(claim)
|
||||
&& (/\b(?:old|earlier|previous|pre[- ]write)\s+(?:snapshot|value|result|version)\b/i.test(claim) || explicitlyEarlierCall)
|
||||
&& !fillPattern.test(claim) && !/\b(?:next|later|subsequent|new|fresh|future)\s+(?:read\w*|request\w*|caller\w*)\b/i.test(claim);
|
||||
if (!(onlyEarlierReturn && subsequentRead && (violation || remedy || explicitlyEarlierCall))) return false;
|
||||
}
|
||||
return finding || subsequentRead || remedy;
|
||||
});
|
||||
}
|
||||
+2946
-212
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,411 @@
|
||||
/** Completed parent file delivery and seeded coverage-diagram evidence. */
|
||||
import { posix, win32 } from 'node:path';
|
||||
import type { SkillTestResult } from './session-runner';
|
||||
|
||||
export interface CoverageAuditFiles {
|
||||
cwd: string;
|
||||
source: { path: string; content: string };
|
||||
tests: { path: string; content: string };
|
||||
}
|
||||
const object = (v: unknown): v is Record<string, any> => v !== null && typeof v === 'object' && !Array.isArray(v);
|
||||
const normalized = (text: string) => text.replace(/\r\n?/g, '\n').trim();
|
||||
const outputText = (output: unknown): string => typeof output === 'string' ? output
|
||||
: Array.isArray(output) && output.every(b => object(b) && b.type === 'text' && typeof b.text === 'string')
|
||||
? output.map(b => b.text).join('\n') : '';
|
||||
const literal = (token: string): string | undefined => {
|
||||
if (/^'[^']*'$/.test(token) || /^"[^"$`\\]*"$/.test(token)) return token.slice(1, -1);
|
||||
return /^[^\s'"$`\\;|&<>]+$/.test(token) ? token : undefined;
|
||||
};
|
||||
// Recorded transcripts may come from another OS. Resolve their paths in the
|
||||
// recorded cwd's namespace, retaining the canonical and containment checks.
|
||||
const evidencePaths = (cwd: string) => /^(?:[A-Za-z]:[\\/]|\\\\)/.test(cwd) ? win32 : posix;
|
||||
|
||||
/** Closed literal cat/sed forms only; no shell execution or general shell parser. */
|
||||
function readsFile(command: unknown, file: string, cwd: string, output: unknown, owned: CoverageAuditFiles): boolean {
|
||||
const path = evidencePaths(cwd);
|
||||
if (typeof command !== 'string' || command.length > 16384 || /[\r\n]/.test(command)) return false;
|
||||
const parts: string[] = [];
|
||||
const separators: string[] = [];
|
||||
let part = '', quote = '', andList = false, semicolons = false;
|
||||
for (let index = 0; index < command.length; index++) {
|
||||
const char = command[index]!;
|
||||
if (quote) {
|
||||
if (quote !== "'" && /[`$]/.test(char)) return false;
|
||||
// These escapes remain literal regex characters in double quotes. A
|
||||
// neighboring grep may use them; only its closed display form below
|
||||
// accepts the backslashes. Shell expansion and escaped quotes stay out.
|
||||
if (char === '\\' && quote !== "'" && !/[.|]/.test(command[index + 1] ?? '')) return false;
|
||||
part += char; if (char === quote) quote = '';
|
||||
}
|
||||
else if (char === '\'' || char === '"') { quote = char; part += char; }
|
||||
else if (/[`$\\#<{}()]/.test(char)) return false; // Comments, heredocs, functions and grouped execution are unsupported.
|
||||
else if (char === ';') { semicolons = true; separators.push(';'); parts.push(part.trim()); part = ''; }
|
||||
else if (char === '&') {
|
||||
if (command[index + 1] !== '&') return false;
|
||||
index++; andList = true; separators.push('&&'); parts.push(part.trim()); part = '';
|
||||
}
|
||||
else part += char;
|
||||
}
|
||||
if (quote) return false;
|
||||
parts.push(part.trim());
|
||||
// A final Git display can hide a failed && prefix. Only the two owned reads
|
||||
// with their exact ordered output can establish delivery through this form.
|
||||
if (andList && semicolons && separators.at(-1) === ';' && separators.slice(0, -1).every(s => s === '&&') &&
|
||||
/^git log --oneline [A-Za-z0-9_][A-Za-z0-9_./~^-]*$/.test(parts.at(-2) ?? '') &&
|
||||
/^git diff [A-Za-z0-9_][A-Za-z0-9_./~^-]* --stat$/.test(parts.at(-1) ?? '')) {
|
||||
const files = [owned.source, owned.tests], readPaths: string[] = [], prefix: string[] = [];
|
||||
for (const segment of parts.slice(0, -2)) {
|
||||
const read = /^cat -n (.+)$/.exec(segment), target = read && literal(read[1]!);
|
||||
if (target) {
|
||||
const known = files.find(f => path.resolve(cwd, target) === f.path);
|
||||
if (!known || readPaths.includes(known.path)) return false;
|
||||
readPaths.push(known.path); prefix.push(known.content.replace(/\r\n?/g, '\n').replace(/\n$/, ''));
|
||||
} else if (/^echo [-=]+$/.test(segment)) prefix.push(segment.slice(5));
|
||||
else return false;
|
||||
}
|
||||
const actual = outputText(output), expected = normalized(prefix.join('\n'));
|
||||
const deliveredPrefix = normalized(actual.replace(/^ *\d+(?:\t|→)/gm, ''));
|
||||
return readPaths.length === 2 && readPaths.includes(file) && actual.length <= 4 * 1024 * 1024 &&
|
||||
(deliveredPrefix === expected || deliveredPrefix.startsWith(expected + '\n'));
|
||||
}
|
||||
const cd = /^cd\s+(.+)$/.exec(parts[0] ?? '');
|
||||
if (cd) {
|
||||
const target = literal(cd[1]!);
|
||||
if (target !== cwd) return false;
|
||||
parts.shift();
|
||||
}
|
||||
// A cwd change or shell control cannot turn a relative target into another
|
||||
// file, or leave a printed old command mistaken for an executed read.
|
||||
if (parts.some(p => /^(?:cd|pushd|popd|source|\.|eval|exec|exit|return|function|alias|if|then|else|for|while|until|case)\s/.test(p) ||
|
||||
/^(?:exit|return|fi|done)$/.test(p) || /^[A-Za-z_][A-Za-z0-9_]*=/.test(p))) return false;
|
||||
const readTarget = (p: string): string | undefined => {
|
||||
const cat = /^cat(?:\s+-n)?(?:\s+--)?\s+(.+)$/.exec(p);
|
||||
const sed = /^sed\s+-n\s+(?:'\d+(?:,\d+)?p'|"\d+(?:,\d+)?p"|\d+(?:,\d+)?p)\s+(.+)$/.exec(p);
|
||||
return literal((cat ?? sed)?.[1] ?? '');
|
||||
};
|
||||
// Unrelated reads may precede/follow a delivered file. They cannot mutate it
|
||||
// or print replacement content through another interpreter. Only discarded
|
||||
// stderr is allowed; a credited cat/sed itself still has no redirection.
|
||||
const readOnly = (part: string) => {
|
||||
// A neighboring optional file read may report absence. It never receives
|
||||
// source/test delivery credit; only earlier independent cat/sed segments do.
|
||||
const fallback = /^(cat(?:\s+-n)?(?:\s+--)?\s+.+)\s+2>\/dev\/null\s+\|\|\s+echo\s+(.+)$/.exec(part);
|
||||
if (fallback) return readTarget(fallback[1]!) !== undefined && literal(fallback[2]!) !== undefined && !part.includes('\\');
|
||||
const stages: string[] = [];
|
||||
let value = '', quoted = '';
|
||||
for (const char of part) {
|
||||
if (quoted) { value += char; if (char === quoted) quoted = ''; }
|
||||
else if (char === "'" || char === '"') { quoted = char; value += char; }
|
||||
else if (char === '|') { stages.push(value); value = ''; }
|
||||
else value += char;
|
||||
}
|
||||
stages.push(value);
|
||||
return stages.every(value => {
|
||||
const stage = value.trim().replace(/(?:^|\s)2>\/dev\/null(?=\s|$)/g, ' ').trim();
|
||||
if (/[<>]/.test(stage.replace(/'[^']*'|"[^"]*"/g, ''))) return false;
|
||||
// Backslashes are data only in these closed grep display patterns.
|
||||
// In particular, echo -e cannot print replacement fixture bodies.
|
||||
const grepRange = /^grep\s+-n(?:\s+-i)?(?:\s+-B\d{1,4})?(?:\s+-A\d{1,4})?\s+"(?:[^"\\$`]|\\[|.])*"\s+(.+)$/.exec(stage);
|
||||
const grepInput = grepRange && literal(grepRange[1]!);
|
||||
const displayGrep = Boolean(grepInput && !grepInput.startsWith('-'));
|
||||
// An awk range without actions only prints matching input lines.
|
||||
// Programs, BEGIN/END, output redirection and interpreter calls cannot
|
||||
// match this grammar, and its input path must be one literal operand.
|
||||
const awkRange = /^awk\s+'\/(?:[^/\\]|\\[./|])*\/,\/(?:[^/\\]|\\[./|])*\/'\s+(.+)$/.exec(stage);
|
||||
const awkInput = awkRange && literal(awkRange[1]!);
|
||||
// A single flag starts display at a heading and exits at the next one.
|
||||
// No other awk action, output destination or interpreter call is allowed.
|
||||
const awkHeadings = /^awk\s+'\/(?:[^/\\]|\\[./|])*\/\{f=1\} f&&\/(?:[^/\\]|\\[./|])*\/\{exit\} f'\s+(.+)$/.exec(stage);
|
||||
const headingInput = awkHeadings && literal(awkHeadings[1]!);
|
||||
const displayAwk = Boolean((awkInput && !awkInput.startsWith('-')) || (headingInput && !headingInput.startsWith('-')));
|
||||
if (stage.includes('\\') && !/^grep\s+-(?:E|cE)\s+'[^']*'(?:\s+[^\\]*)?$/.test(stage) && !displayGrep && !displayAwk) return false;
|
||||
const git = /^git\s+(log|diff)(?:\s+(.*))?$/.exec(stage);
|
||||
// These neighboring Git calls are display-only: literal revisions and the
|
||||
// observed display flag. Quoted/concatenated or unknown options may write
|
||||
// files or invoke helpers, so they cannot borrow a read-only classification.
|
||||
const gitDisplay = git !== null && (!git[2] || git[2].split(/\s+/).every(token =>
|
||||
token === (git[1] === 'log' ? '--oneline' : '--stat') ||
|
||||
(git[1] === 'log' && /^-[1-9]\d{0,4}$/.test(token)) || /^[A-Za-z0-9_][A-Za-z0-9_./~^-]*$/.test(token)));
|
||||
return /^(?:cat|grep|head|ls|echo)(?:\s|$)/.test(stage) || stage === 'pwd' || stage === 'wc -l' || stage === 'git ls-files' || stage === "sed 's/^/TESTFILES:/'" || /^\[ -f [A-Za-z0-9_.\/-]+ \]$/.test(stage) ||
|
||||
readTarget(stage) !== undefined || gitDisplay || displayAwk;
|
||||
});
|
||||
};
|
||||
if (parts.some(p => p && !readOnly(p))) return false;
|
||||
// A successful, unmixed && list may include literal display separators
|
||||
// and a closed diff-stat command. These segments never receive file credit.
|
||||
const andDisplay = (p: string) => {
|
||||
if (p === 'echo' || /^echo\s+[-=]+$/.test(p) || /^echo [-=]{2,} [A-Za-z0-9_.\/-]+ [-=]{2,}$/.test(p)) return true;
|
||||
const caption = /^echo\s+(.+)$/.exec(p), value = caption && literal(caption[1]!);
|
||||
if (value && /^[-=]{2,}\s+[A-Za-z0-9_][A-Za-z0-9_./-]*(?:\s+(?:vs|and)\s+[A-Za-z0-9_][A-Za-z0-9_./-]*)?\s+[-=]{2,}$/.test(value)) return true;
|
||||
return /^git\s+diff(?:\s+[A-Za-z0-9_][A-Za-z0-9_./~^-]*)?\s+--stat$/.test(p);
|
||||
};
|
||||
if (andList && semicolons) {
|
||||
const caption = /^echo (.+)$/.exec(parts[0] ?? ''), value = caption && literal(caption[1]!);
|
||||
if (!value || ![owned.source, owned.tests].some(f => value === `=== ${path.relative(cwd, f.path)} ===`)) return false;
|
||||
// Later display commands can mask an earlier exit code. Require both owned
|
||||
// reads in the initial && chain and their exact ordered stdout prefix.
|
||||
const prefix: string[] = [], readPaths: string[] = [];
|
||||
for (let i = 0; i < parts.length && (i === 0 || separators[i - 1] === '&&'); i++) {
|
||||
const segment = parts[i]!, target = readTarget(segment);
|
||||
const known = target && [owned.source, owned.tests].find(f => path.resolve(cwd, target) === f.path);
|
||||
if (known && /^cat -n /.test(segment) && !readPaths.includes(known.path)) {
|
||||
readPaths.push(known.path); prefix.push(known.content.replace(/\r\n?/g, '\n').replace(/\n$/, ''));
|
||||
} else if (andDisplay(segment) && /^echo(?: |$)/.test(segment)) {
|
||||
const value = segment.slice(5); prefix.push(literal(value) ?? value);
|
||||
} else return false;
|
||||
if (readPaths.length === 2) break;
|
||||
}
|
||||
const actual = outputText(output), expected = normalized(prefix.join('\n'));
|
||||
const deliveredPrefix = normalized(actual.replace(/^ *\d+(?:\t|→)/gm, ''));
|
||||
return readPaths.length === 2 && readPaths.includes(file) && actual.length <= 4 * 1024 * 1024 &&
|
||||
(deliveredPrefix === expected || deliveredPrefix.startsWith(expected + '\n'));
|
||||
}
|
||||
if (andList && parts.some(p => readTarget(p) === undefined && !andDisplay(p))) return false;
|
||||
return parts.some(p => {
|
||||
const target = readTarget(p);
|
||||
return target !== undefined && path.resolve(cwd, target) === file;
|
||||
});
|
||||
}
|
||||
function delivered(output: unknown, expected: string): boolean {
|
||||
const text = outputText(output);
|
||||
const body = normalized(expected);
|
||||
if (!body || text.length > 4 * 1024 * 1024) return false;
|
||||
if (normalized(text).includes(body)) return true;
|
||||
// Native Read gutters and cat -n use different separators; retain actual
|
||||
// code indentation and require the entire contiguous file, not filenames.
|
||||
return normalized(text.replace(/^ *\d+(?:\t|→)/gm, '')).includes(body);
|
||||
}
|
||||
|
||||
export function coverageAuditReadEvidence(transcript: unknown[], files: CoverageAuditFiles): { sourceRead: boolean; testsRead: boolean } {
|
||||
const path = evidencePaths(files.cwd);
|
||||
const found = { sourceRead: false, testsRead: false };
|
||||
if (!path.isAbsolute(files.cwd) || path.resolve(files.cwd) !== files.cwd || files.source.path === files.tests.path ||
|
||||
[files.source, files.tests].some(f => {
|
||||
const relative = path.relative(files.cwd, f.path);
|
||||
return !path.isAbsolute(f.path) || path.resolve(f.path) !== f.path || !relative || relative === '..' ||
|
||||
relative.startsWith('..' + path.sep) || path.isAbsolute(relative) || !normalized(f.content);
|
||||
})) return found;
|
||||
const init = transcript.filter((e): e is Record<string, any> =>
|
||||
object(e) && e.type === 'system' && e.subtype === 'init' && e.cwd === files.cwd);
|
||||
if (init.length !== 1 || typeof init[0].session_id !== 'string' || !init[0].session_id) return found;
|
||||
const session = init[0].session_id,
|
||||
uses = new Map<string, { name: string; input: Record<string, any> }>(),
|
||||
results = new Set<string>();
|
||||
for (const e of transcript.slice(transcript.indexOf(init[0]) + 1)) {
|
||||
if (!object(e) || e.session_id !== session || (e.parent_tool_use_id !== null && e.parent_tool_use_id !== undefined) ||
|
||||
!object(e.message) || !Array.isArray(e.message.content)) continue;
|
||||
for (const b of e.message.content) {
|
||||
if (!object(b)) continue;
|
||||
if (e.type === 'assistant' && e.message.role === 'assistant' && b.type === 'tool_use' && typeof b.id === 'string' && object(b.input)) {
|
||||
if (uses.has(b.id)) return { sourceRead: false, testsRead: false };
|
||||
uses.set(b.id, { name: b.name, input: b.input });
|
||||
} else if (e.type === 'user' && e.message.role === 'user' && b.type === 'tool_result' && typeof b.tool_use_id === 'string') {
|
||||
if (results.has(b.tool_use_id)) return { sourceRead: false, testsRead: false };
|
||||
results.add(b.tool_use_id);
|
||||
const u = uses.get(b.tool_use_id);
|
||||
if (!u || (b.is_error !== undefined && b.is_error !== false)) continue;
|
||||
for (const [key, file] of [['sourceRead', files.source], ['testsRead', files.tests]] as const) {
|
||||
const named = u.name === 'Read'
|
||||
? typeof u.input.file_path === 'string' && path.resolve(files.cwd, u.input.file_path) === file.path
|
||||
: u.name === 'Bash' && readsFile(u.input.command, file.path, files.cwd, b.content, files);
|
||||
if (named && delivered(b.content, file.content)) found[key] = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return found;
|
||||
}
|
||||
const exampleDiagram = (line: string) => /^(?:example|sample|illustration)\b/i.test(line.trim().replace(/^[#*]+\s*/, ''));
|
||||
/** Top-level ASCII, optionally fenced; an outer source/example fence owns its body. */
|
||||
function diagramBlocks(output: string): string[][] {
|
||||
const blocks: string[][] = [];
|
||||
let outside: string[] = [];
|
||||
let fence: { char: string; length: number; allowed: boolean; lines: string[] } | undefined;
|
||||
for (const line of output.split('\n')) {
|
||||
const marker = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (fence) {
|
||||
if (marker && marker[1]![0] === fence.char && marker[1]!.length >= fence.length && !marker[2]!.trim()) {
|
||||
if (fence.allowed) blocks.push(fence.lines);
|
||||
fence = undefined;
|
||||
} else fence.lines.push(line);
|
||||
} else if (marker) {
|
||||
if (outside.length) blocks.push(outside);
|
||||
fence = { char: marker[1]![0]!, length: marker[1]!.length,
|
||||
allowed: /^(?:text|ascii|plaintext)?$/.test(marker[2]!.trim()) && !outside.some(exampleDiagram), lines: [] };
|
||||
outside = [];
|
||||
} else outside.push(line);
|
||||
}
|
||||
if (outside.length) blocks.push(outside);
|
||||
return blocks.filter(lines => {
|
||||
const firstRow = lines.findIndex(line => treeRow(line) !== undefined);
|
||||
return firstRow >= 0 && !lines.slice(0, firstRow).some(exampleDiagram);
|
||||
});
|
||||
}
|
||||
|
||||
function treeRow(line: string): { depth: number; text: string } | undefined {
|
||||
const match = /^([ |│]*)(?:[├└]─+►?|[+|]-+)\s+(.+)$/.exec(line);
|
||||
if (!match) {
|
||||
// Unindented function roots own following branch rows. Other function
|
||||
// roots also end a subtree, so a sibling cannot lend a coverage marker.
|
||||
return /^[A-Za-z_$][A-Za-z0-9_$]*\([^()\n]*\)(?:[\t ]+.*)?$/.test(line)
|
||||
? { depth: -1, text: line } : undefined;
|
||||
}
|
||||
// A parallel USER FLOWS column cannot supply CODE PATHS coverage markers.
|
||||
return { depth: match[1]!.length, text: match[2]!.split(/ {3,}(?=[├└+|])/, 1)[0]! };
|
||||
}
|
||||
|
||||
function diagramLegend(lines: string[], firstRow: number): Map<string, boolean> {
|
||||
const meanings = new Map<string, boolean>();
|
||||
const pair = String.raw`\[([✓✔✗✘])\][\t ]+(TESTED|COVERED|GAP|UNTESTED)`;
|
||||
const legend = new RegExp(String.raw`^${pair}[\t |,;]+${pair}$`, 'i');
|
||||
const bareKey = new RegExp(String.raw`${pair}.*\[[✓✔✗✘]\]`, 'i');
|
||||
for (const [index, original] of lines.entries()) {
|
||||
// Footer keys govern this same block too; tree and annotation markers
|
||||
// describe paths. A bare pair remains a declaration even when indented.
|
||||
if (index >= firstRow && (treeRow(original) || (!/\bLegend\b/i.test(original) && !bareKey.test(original)))) continue;
|
||||
// Decorative branch keys and an explicit GAP explanation do not change
|
||||
// the two coverage meanings. All other qualifiers keep the closed grammar.
|
||||
const line = original.replace(/[\t ]+[─-]+►[\t ]+branch$/i, '')
|
||||
.replace(/(\[[✓✔✗✘]\][\t ]+(?:GAP|UNTESTED))[\t ]+\((?:no test|GAP)\)$/i, '$1')
|
||||
.replace(/(\[[✓✔✗✘]\][\t ]+COVERED)[\t ]+by a test\b/gi, '$1')
|
||||
.replace(/(\[[✓✔✗✘]\][\t ]+(?:GAP|UNTESTED))[\t ]+[—–-][\t ]+no test exercises this path$/i, '$1');
|
||||
const start = line.search(/\[[✓✔✗✘]\]/);
|
||||
if (start < 0) continue;
|
||||
if (/^\s*>|["“”]|\b(?:not|no|never|example|sample|false|incorrect|hypothetical)\b/i.test(line)) return new Map();
|
||||
// Only a legend label or a literal file's coverage-map caption may precede
|
||||
// the pair. Arbitrary prose must not be discarded into an affirmative key.
|
||||
const prefix = line.slice(0, start).trim();
|
||||
if (prefix && !/^(?:Legend:|[A-Za-z0-9_.-]+(?:\/[A-Za-z0-9_.-]+)*\.[A-Za-z0-9]+[\t ]+[—–-][\t ]+(?:test[\t ]+)?coverage[\t ]+map)$/i.test(prefix)) return new Map();
|
||||
const match = legend.exec(line.slice(start).trim());
|
||||
if (!match || match[1] === match[3]) return new Map();
|
||||
const entries = [[match[1]!, /^(?:TESTED|COVERED)$/i.test(match[2]!)],
|
||||
[match[3]!, /^(?:TESTED|COVERED)$/i.test(match[4]!)]] as const;
|
||||
if (entries[0][1] === entries[1][1]) return new Map();
|
||||
for (const [symbol, covered] of entries) {
|
||||
if (meanings.has(symbol) && meanings.get(symbol) !== covered) return new Map();
|
||||
meanings.set(symbol, covered);
|
||||
}
|
||||
}
|
||||
return currentDiagramLegend(lines) ? meanings : new Map();
|
||||
}
|
||||
|
||||
/** Text markers need an explicit, current legend in this same diagram block. */
|
||||
function diagramWordLegend(lines: string[]): Map<string, boolean> | undefined {
|
||||
const declarations = lines.filter(line => /^\s*Legend\b/i.test(line) && /\[\s*(?:OK|GAP)\s*\]/i.test(line));
|
||||
if (!declarations.length) return undefined;
|
||||
const meanings = new Map<string, boolean>();
|
||||
const pair = String.raw`\[\s*(OK|GAP)\s*\]\s+(covered|tested|no test|untested)`;
|
||||
const form = new RegExp(String.raw`^\s*Legend:?\s+${pair}(?:\s+[|,;]?\s*|[|,;]\s*)${pair}\s*$`, 'i');
|
||||
for (const line of declarations) {
|
||||
const match = form.exec(line);
|
||||
if (!match || match[1]!.toUpperCase() === match[3]!.toUpperCase()) return new Map();
|
||||
for (const [name, description] of [[match[1]!, match[2]!], [match[3]!, match[4]!]]) {
|
||||
const key = name.toUpperCase(), covered = /^(?:covered|tested)$/i.test(description);
|
||||
if ((key === 'OK') !== covered || (meanings.has(key) && meanings.get(key) !== covered)) return new Map();
|
||||
meanings.set(key, covered);
|
||||
}
|
||||
}
|
||||
if (lines.some(line => /^\s*(?:Correction:\s*)?(?:This|The|That)\s+legend\s+(?:is|has been)\s+['"‘’“”`]*(?:withdrawn|superseded|cancelled|canceled|rejected|retracted|not current|no longer current)\b/i.test(line))) return new Map();
|
||||
return meanings;
|
||||
}
|
||||
|
||||
function currentDiagramLegend(lines: string[]): boolean {
|
||||
const status = '(?:withdrawn|superseded|cancelled|canceled|rejected|retracted|incorrect|hypothetical|proposed|optional|not current|no longer current)';
|
||||
for (const line of lines) {
|
||||
if (/^\s*>/.test(line)) continue;
|
||||
const owner = '(?:this|the|that) legend (?:is|has been) ';
|
||||
const scalar = new RegExp(`((?:^|[.!?;]\\s+)[\\t ]*(?:Correction:\\s*)?${owner})["“'‘\x60](${status})["”'’\x60]`, 'gi');
|
||||
const current = line.replace(/\*\*/g, '').replace(scalar, '$1$2').replace(/"[^"\n]*"|“[^”\n]*”|'[^'\n]*'|‘[^’\n]*’|`[^`\n]*`/g, '');
|
||||
if (exampleDiagram(current) || /^\s*(?:Source|Quoted(?: source)?|Historical(?: note| assessment)?|Hypothetical|Example|If approved)\s*:/i.test(current) ||
|
||||
new RegExp(`(?:^|[.!?;]\\s+)[\\t ]*(?:Correction:\\s*)?${owner}${status}\\b`, 'i').test(current) ||
|
||||
/(?:^|[.!?;]\s+)[\t ]*(?:this|the|that) legend (?:applies|will apply) (?:only )?(?:if|once|when) approved\b/i.test(current)) return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Checkbox states are meaningful only under a current key in this block. */
|
||||
function diagramCheckboxLegend(lines: string[]): Map<string, boolean> | undefined {
|
||||
const declarations = lines.filter(line => /^\s*Legend\b/i.test(line) && /\[[x ]\]/i.test(line));
|
||||
if (!declarations.length) return undefined;
|
||||
const pair = String.raw`\[([x ])\]\s+(covered(?: by an existing test)?|tested|no test(?: reaches this path)?|untested|GAP)`;
|
||||
const form = new RegExp(String.raw`^\s*Legend:?\s+${pair}(?:\s+[|,;]?\s*|[|,;]\s*)${pair}\s*$`, 'i');
|
||||
const meanings = new Map<string, boolean>();
|
||||
for (const original of declarations) {
|
||||
const line = original.replace(/[\t ]+[─-] happy path[\t ]+✗ negative path$/i, '');
|
||||
const match = form.exec(line);
|
||||
if (!match || match[1]!.toLowerCase() === match[3]!.toLowerCase()) return new Map();
|
||||
for (const [symbol, description] of [[match[1]!, match[2]!], [match[3]!, match[4]!]]) {
|
||||
const key = symbol.toLowerCase(), covered = /^(?:covered|tested)\b/i.test(description);
|
||||
if ((key === 'x') !== covered || (meanings.has(key) && meanings.get(key) !== covered)) return new Map();
|
||||
meanings.set(key, covered);
|
||||
}
|
||||
}
|
||||
return currentDiagramLegend(lines) ? meanings : new Map();
|
||||
}
|
||||
|
||||
function seededDiagram(output: string): boolean {
|
||||
for (const lines of diagramBlocks(output)) {
|
||||
const rows = lines.map(treeRow);
|
||||
// A coverage marker aligned beneath a branch's label continues that row.
|
||||
// Stop at prose or another branch; a distant parallel column cannot lend it.
|
||||
let owner = -1;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
if (rows[i]) { owner = i; continue; }
|
||||
const continuation = /^([ |│]+)(\[[✓✔✗✘xX ]\].*)$/.exec(lines[i]!);
|
||||
if (owner >= 0 && continuation && [4, 6, 8].includes(continuation[1]!.length - rows[owner]!.depth))
|
||||
rows[owner]!.text += ' ' + continuation[2]!;
|
||||
else if (!/^[ |│]*$/.test(lines[i]!)) owner = -1;
|
||||
}
|
||||
const legend = diagramLegend(lines, rows.findIndex(row => row !== undefined));
|
||||
const wordLegend = diagramWordLegend(lines);
|
||||
const checkboxLegend = diagramCheckboxLegend(lines);
|
||||
if (rows.some(row => row && /\[[x ]\]/i.test(row.text)) && checkboxLegend?.size !== 2) continue;
|
||||
const marker = String.raw`(?:\[[✓✔✗✘xX ]\]|\[\s*(?:OK|GAP)\s*\])`;
|
||||
const marked = (line: string) => /\[[✓✔✗✘xX ]\]|\[\s*OK\s*\]/i.test(line) ||
|
||||
(wordLegend !== undefined && /\[\s*GAP\s*\]/i.test(line));
|
||||
const symbolMeans = (line: string, covered: boolean) => {
|
||||
if (new RegExp(String.raw`\b(?:not|never)\s+${marker}|(?:${marker}|\b(?:marker|symbol))\s+(?:is|are)\s+(?:false|incorrect|wrong)\b`, 'i').test(line)) return false;
|
||||
// A status correction [covered]→[gap] carries only its final marker.
|
||||
// Unrelated contradictory markers cannot supply both coverage states.
|
||||
const corrected = line.replace(new RegExp(marker + String.raw`[\t ]*(?:→|->)[\t ]*(?=` + marker + ')', 'gi'), '');
|
||||
const states = [...corrected.matchAll(/\[([✓✔✗✘])\]|\[\s*(OK|GAP)\s*\]|\[([x ])\]/gi)]
|
||||
.map(match => match[1] ? legend.get(match[1]) : match[2] ? wordLegend?.get(match[2].toUpperCase()) : checkboxLegend?.get(match[3]!.toLowerCase()));
|
||||
return states.includes(covered) && !states.includes(!covered);
|
||||
};
|
||||
const payment = rows.findIndex(row => row && /^processPayment\b/.test(row.text));
|
||||
const refund = rows.findIndex(row => row && /^refundPayment\b/.test(row.text));
|
||||
if (payment < 0 || refund < 0) continue;
|
||||
const subtree = (index: number): string[] => {
|
||||
const texts = [rows[index]!.text], depth = rows[index]!.depth;
|
||||
for (let i = index + 1; i < rows.length; i++) {
|
||||
const row = rows[i];
|
||||
if (row && row.depth <= depth) break;
|
||||
if (row) texts.push(row.text);
|
||||
}
|
||||
return texts;
|
||||
};
|
||||
const covered = subtree(payment).some(line =>
|
||||
(marked(line) ? symbolMeans(line, true) :
|
||||
/(?:\bTESTED\b|\bCOVERED\b)/i.test(line) || (/✓/.test(line) && (legend.get('✓') ?? true))) && /happy|success|valid|USD/i.test(line) &&
|
||||
!/untested|(?:not|never)\s+(?:yet\s+)?(?:tested|covered)|no\s+test/i.test(line));
|
||||
const missing = subtree(refund).some(line =>
|
||||
(marked(line) ? symbolMeans(line, false) : /(?:\[GAP\]|✗\s*GAP|\bUNTESTED\b)/i.test(line)) &&
|
||||
!/\b(?:not|never)\s+(?:\[)?(?:untested|gap)\b|\b(?:untested|gap)\]?\s+(?:is|are)\s+(?:false|incorrect|wrong)\b|\bno\s+(?:coverage\s+)?gaps?\b|\b(?:fully|completely)\s+(?:tested|covered)\b/i.test(line));
|
||||
if (covered && missing) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
export function coverageAuditVerdict(
|
||||
result: Pick<SkillTestResult, 'exitReason' | 'browseErrors' | 'output' | 'transcript'>,
|
||||
files: CoverageAuditFiles,
|
||||
) {
|
||||
const reads = coverageAuditReadEvidence(Array.isArray(result.transcript) ? result.transcript : [], files);
|
||||
const diagram = typeof result.output === 'string' && seededDiagram(result.output),
|
||||
failures: string[] = [];
|
||||
if (result.exitReason !== 'success') failures.push('capture did not complete successfully');
|
||||
if (result.browseErrors.length) failures.push('capture reported tool errors');
|
||||
if (!reads.sourceRead) failures.push('missing successful source-file read');
|
||||
if (!reads.testsRead) failures.push('missing successful test-file read');
|
||||
if (!diagram) failures.push('missing seeded covered-payment/refund-gap diagram');
|
||||
return { ...reads, diagram, passed: failures.length === 0, failures };
|
||||
}
|
||||
@@ -0,0 +1,91 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
|
||||
/** Recording or skipping an explicitly deferred typography TODO preserves the
|
||||
* current design. Selecting its build-now alternative remains a review choice.
|
||||
*/
|
||||
function deferredTypographyTodo(q: NonNullable<AskUserQuestionFingerprint['nativeCall']>['questions'][number], selected: string): boolean {
|
||||
const clean = (text: string) => text.trim().replace(/\s+/g, ' ');
|
||||
const lines = q.question.trim().split('\n').map(clean);
|
||||
const title = /^D[1-9]\d* [—–-] TODO proposal: (?:record|add) a deferred TODOS\.md (?:item|note) to (?:evaluate|explore|consider) [^?\n]+ \(replacing ([A-Za-z][A-Za-z0-9-]{0,39})\) (?:in a later|during a future) design pass\?$/.exec(lines[0] ?? '');
|
||||
if (!title || !/^TODO [1-9]\d*$/.test(q.header) || lines.length !== 7 ||
|
||||
!/^Project\/branch\/task: [A-Za-z0-9_./-]+, \/plan-design-review of PLAN\.md, post-pass TODOS\.md updates\.$/.test(lines[1]!)) return false;
|
||||
const font = title[1]!;
|
||||
const assessment = lines[2]!;
|
||||
// Bind the current scope before reading the optional explanation of future
|
||||
// value. Font examples, effort estimates and brand prose are not evidence.
|
||||
const scope = new RegExp(`^ELI10: DESIGN\\.md and (?:this|the current) plan (?:keep|retain|preserve) ${font} as the app font(?:, and you excluded visual exploration from this update|\\. Visual exploration (?:remains|is) out of scope for this update), so (?:nothing changes now|the current design remains unchanged)\\.`);
|
||||
const recordOnly = /\bThis question is only about whether to (?:write|record) [^.]+ in TODOS\.md [^.]*\bfuture \/design-consultation\b[^.]*, not about changing anything here\./.test(assessment)
|
||||
|| /\bThis question only (?:records|skips) a deferred TODOS\.md (?:item|note) for a future \/design-consultation; it does not change the current design\./.test(assessment);
|
||||
if (!scope.test(assessment) || !recordOnly ||
|
||||
!/^Stakes if we pick wrong: .+\.$/.test(lines[3]!) ||
|
||||
!/^Recommendation: A\b/.test(lines[4]!) || !/\bout[- ]of[- ]scope\b/.test(lines[4]!) ||
|
||||
!/^Note: options differ in kind, not coverage\b/.test(lines[5]!) ||
|
||||
!/^Net: .+\.$/.test(lines[6]!)) return false;
|
||||
const choices = ['A Add to TODOS.md', 'B Skip, not valuable enough', 'C Build it now in this PR'];
|
||||
const labels = q.options.map(o => clean(o.label).replace(/ \(recommended\)$/i, ''));
|
||||
if (labels.length !== 3 || choices.some(label => labels.filter(l => l === label).length !== 1)) return false;
|
||||
const description = (label: string) => clean(q.options[labels.indexOf(label)]!.description!);
|
||||
const [add, skip, build] = choices.map(description) as [string, string, string];
|
||||
const selectedLabel = labels[q.options.findIndex(o => o.label === selected)];
|
||||
if (selectedLabel !== choices[0] && selectedLabel !== choices[1]) return false;
|
||||
// A and B are record/skip-only choices; C is explicitly an implementation
|
||||
// alternative. Additional present-work instructions invalidate the boundary.
|
||||
const all = [...lines, add, skip, build].join('\n');
|
||||
if (/\b(?:Correction|Hypothetical|Source only|Example only)\s*:|\b(?:this (?:scope|proposal|deferment)|the (?:scope|proposal|deferment)) (?:is|was) (?:withdrawn|cancelled|rejected)\b/i.test(all) ||
|
||||
/\b(?:Also|Additionally|Instead|Now|Then)\s+(?:we\s+)?(?:must\s+|will\s+)?(?:fix|replace|change|implement|build|add|load|remove)\b/i.test(all) ||
|
||||
/\b(?:font replacement|visual exploration|typography work)\s+(?:is|becomes|remains)\s+(?:now\s+)?in scope\b/i.test(all)) return false;
|
||||
const unquoted = (text: string) => text.replace(/"[^"\n]*"|“[^”\n]*”/g, '[quoted]');
|
||||
const retained = unquoted([...lines, add, skip].join('\n'));
|
||||
if (/\bVisual exploration (?:is|remains) (?:no longer|not) out of scope\b/i.test(retained) ||
|
||||
new RegExp(`\\b(?:This|The current) plan (?:no longer|does not) (?:keeps?|retains?|preserves?) ${font}\\b`, 'i').test(retained)) return false;
|
||||
const recording = unquoted(add + '\n' + skip).split(/[.!?]\s+|[✅❌]|\n/).map(clean);
|
||||
if (recording.some(sentence => /^(?:Add|Create|Fix|Replace|Change|Implement|Build|Load|Remove|Set|Make)\b/i.test(sentence) &&
|
||||
(!/^(?:Add|Create) (?:a |the |one )?TODOS\.md (?:file|item|note)(?: for (?:the )?(?:future|deferred|later) [A-Za-z0-9 /_-]+)?\.?$/i.test(sentence) || /\b(?:and|then|plus)\s+(?:add|create|fix|replace|change|implement|build|load|remove|set|make)\b/i.test(sentence)))) return false;
|
||||
if (/\b(?:it is false that|not true that|if approved|no longer preserves?)\b/i.test(retained) ||
|
||||
/\b(?:replace|change|implement|build|load|remove)\b[^.!?\n]*\b(?:now|in this PR|in this update)\b/i.test(retained.replace(/not about changing anything here\./g, '').replace(/do not add it now/g, 'deferred'))) return false;
|
||||
const preservesFont = new RegExp(`(?:Nothing changes|No design changes) in this update; DESIGN\\.md and ${font} (?:stay as approved|remain unchanged)\\.`);
|
||||
return /\b(?:next|future) \/design-consultation\b/.test(add) && preservesFont.test(add) &&
|
||||
/\b(?:Adds a TODOS\.md file|Records only a TODOS\.md note)\b/.test(add) &&
|
||||
/\b(?:No TODOS\.md noise|No TODO is recorded)\b/.test(skip) && /\b(?:Zero follow-up work|No follow-up work)\./.test(skip) &&
|
||||
/\b(?:immediately|now|this PR)\b/i.test(build) && /\b(?:Fonts? load|Replace the font|Change the font)\b/i.test(build);
|
||||
}
|
||||
|
||||
/** Accepted rendering of existing decisions adds an artifact, not a finding.
|
||||
* It still changes the deliverable and therefore remains a freshness boundary.
|
||||
*/
|
||||
export function isDesignArtifactGeneration(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.sessionId || !call.toolUseId || call.answered !== true || call.failed !== false ||
|
||||
call.questions.length !== 1 || !Array.isArray(call.unansweredQuestionIndices) ||
|
||||
call.unansweredQuestionIndices.length || !Number.isFinite(Date.parse(call.answeredAt ?? '')) ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2 || q.options.length > 3 || Object.keys(call.answers ?? {}).length !== 1 ||
|
||||
q.options.some(o => typeof o.description !== 'string' || ('preview' in o && Boolean(o.preview))) ||
|
||||
fp.options.length !== q.options.length || fp.options.some((o, i) => o.index !== i + 1 || o.label !== q.options[i]!.label)) return false;
|
||||
const clean = (text: string) => text.trim().replace(/\s+/g, ' ');
|
||||
const label = (text: string) => clean(text).replace(/^[AB]\) /, '').replace(/ \(Recommended\)$/, '');
|
||||
const positive = q.options.find(o => call.answers?.[q.question] === o.label);
|
||||
if (!positive) return false;
|
||||
if (q.options.length === 3) return (fp.nativeQuestionIndex === undefined || fp.nativeQuestionIndex === 0) &&
|
||||
deferredTypographyTodo(q, positive.label);
|
||||
const other = q.options.find(o => o !== positive)!;
|
||||
const question = clean(q.question);
|
||||
const description = clean(positive.description!);
|
||||
const alternative = clean(other.description!);
|
||||
// Consume every sentence. A heading or "no new decisions" claim alone
|
||||
// cannot conceal an added requirement, omitted state, or actual design choice.
|
||||
if (/^D\d+ StateTable$/.test(q.header) &&
|
||||
/^D\d+ — Add a state coverage table to the plan body for implementer reference\? <gstack-qid:plan-design-review-states-\d+>$/.test(question)) {
|
||||
return label(positive.label) === 'Add state table' && label(other.label) === 'Leave states in prose only' &&
|
||||
/^Insert a feature × state table \(Form load \/ Save \/ Export \/ Dirty state × Loading \/ Empty \/ Error \/ Success \/ Pending\)\. No new design decisions — all cells derive from existing specs\. Completeness: \d+\/10 — implementers can verify each state against a single reference\.$/.test(description) &&
|
||||
/^Keep the existing prose descriptions without a structured table\. Completeness: \d+\/10 — specs are all there but scattered across paragraphs; edge cases like Export error during dirty-edit are harder to spot\.$/.test(alternative);
|
||||
}
|
||||
if (/^D\d+ Storyboard$/.test(q.header) &&
|
||||
/^D\d+ — Add a user journey storyboard to the plan\? <gstack-qid:plan-design-review-journey-\d+>$/.test(question)) {
|
||||
return label(positive.label) === 'Add storyboard' && label(other.label) === 'Keep one-sentence journey description' &&
|
||||
/^Render the accepted journey as a step\/user-does\/user-feels\/plan-specifies table \(\d+ rows covering happy path, save failure, cancel with dirty state, first-time new account\)\. No new design decisions — pure rendering of existing specs\. Completeness: \d+\/10 — implementers understand the emotional arc and can verify the spec covers each moment\.$/.test(description) &&
|
||||
/^Leave the current one-sentence happy-path description\. Completeness: \d+\/10 — the journey exists but reads like a state machine; error recovery arcs and first-time experience aren't visible without cross-referencing multiple paragraphs\.$/.test(alternative);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
|
||||
/** Seeded-count fixtures cover native review cadence; outside voices have separate evals. */
|
||||
export function pickDesignCountOutsideVoices(
|
||||
_routing: AskUserQuestionFingerprint,
|
||||
active: AskUserQuestionFingerprint,
|
||||
): number | null {
|
||||
const call = active.nativeCall;
|
||||
let question: string;
|
||||
let labels: string[];
|
||||
if (call) {
|
||||
if (call.answered || call.failed) return null;
|
||||
const index = active.nativeQuestionIndex ?? (call.questions.length === 1 ? 0 : undefined);
|
||||
if (index === undefined || !Number.isInteger(index) || index < 0 || index >= call.questions.length) return null;
|
||||
const identity = `${call.sessionId}:${call.toolUseId}` +
|
||||
(call.questions.length > 1 ? `:question:${index}` : '');
|
||||
if (active.signature !== identity) return null;
|
||||
const q = call.questions[index]!;
|
||||
if (q.multiSelect || !/^outside(?: design)? voices$/i.test(q.header.trim()) ||
|
||||
!/<gstack-qid:outside-voices-design>/.test(q.question)) return null;
|
||||
question = q.question;
|
||||
labels = q.options.map(option => option.label);
|
||||
} else {
|
||||
// Native JSONL can arrive after the answer. The caller supplies the
|
||||
// active viewport fingerprint; a known but unmatched packet is blocked
|
||||
// before this hook. Require the specific opt-in premise and both actions.
|
||||
question = active.promptSnippet;
|
||||
const packetBar = /^←[^→]*[☐☒]\s+Outside voices\b[^→]*✔\s*Submit\s*→\s*[│┃]?\s*/i.exec(question);
|
||||
if (packetBar) question = question.slice(packetBar[0].length);
|
||||
else if (!/^(?:[☐□]\s*)?outside(?: design)? voices\b/i.test(question)) return null;
|
||||
labels = active.options.map(option => option.label);
|
||||
while (labels.length > 2 && /^(?:Type something\.?|Chat about this)$/i.test(labels.at(-1)!.trim())) labels.pop();
|
||||
}
|
||||
if (!/\b(?:want|run|include|enable)\b[^?]{0,90}\boutside design voices\b/i.test(question) ||
|
||||
!/\b(?:before|for)\s+(?:the\s+)?(?:detailed\s+)?(?:design\s+)?review\b/i.test(question)) return null;
|
||||
labels = labels.map(label => label.trim().replace(/\s*\(recommended\)\s*$/i, ''));
|
||||
if (labels.length !== 2) return null;
|
||||
const yes = labels.map(label => /^Yes,?\s+run outside design voices$/i.test(label));
|
||||
const no = labels.map(label => /^No,?\s+proceed without$/i.test(label));
|
||||
if (yes.filter(Boolean).length !== 1 || no.filter(Boolean).length !== 1) return null;
|
||||
return no.findIndex(Boolean) + 1;
|
||||
}
|
||||
@@ -0,0 +1,599 @@
|
||||
import { designFirstReviewAUQ } from './claude-pty-runner';
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
import { pickDesignCountOutsideVoices } from './design-count-outside';
|
||||
|
||||
/** Choosing reviewer participation is setup, even when numbered or asked late. */
|
||||
export function isDesignCountSetup(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed || call.questions.length !== 1 ||
|
||||
call.unansweredQuestionIndices?.length || fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || !/^outside(?: design)? voices$/i.test(q.header.trim()) ||
|
||||
(q.question.match(/<gstack-qid:/g)?.length ?? 0) !== 1 ||
|
||||
!/<gstack-qid:(?:plan-design-review-outside-voices|outside-voices-design)>\s*$/.test(q.question) ||
|
||||
(q.question.match(/\?/g)?.length ?? 0) !== 1 ||
|
||||
!/^(?:D\s*\d+(?:\s*\(Step\s*0[A-Z]?\))?\s*[—–:-]\s*)?(?:Run|Want|Include|Enable)\s+outside(?: design)? voices\s+(?:before|for)\s+the\s+(?:detailed\s+)?(?:design\s+)?review(?:\s+passes)?\?/i.test(q.question.trim())) return false;
|
||||
const labels = q.options.map(option => option.label.trim().replace(/\s*\(recommended\)\s*$/i, ''));
|
||||
// Consume the entire menu, not just its opening question or action labels.
|
||||
// Unknown explanatory prose can contain a second product decision.
|
||||
const remainder = q.question.slice(q.question.indexOf('?') + 1).replace(/<gstack-qid:[^>]+>\s*$/, '').trim();
|
||||
if (remainder && !/^(?:Codex evaluates the design; a Claude subagent reviews completeness\.|Codex evaluates against OpenAI's design hard rules \+ litmus checks; a Claude subagent does an independent completeness review\. \(Requires Codex CLI to be installed\.\))$/.test(remainder)) return false;
|
||||
const descriptions = q.options.map(option => (option.description ?? '').trim().replace(/\s+/g, ' '));
|
||||
const noDescription = /^(?:Skip Codex \+ Claude subagent outside pass\. Best for this case: it's a scoped settings form update with a complete DESIGN\.md; hard-rejection checks apply to marketing surfaces, not OPERATE\/settings UI\.|Skip outside voices and go straight to the 7 review passes\. Faster; sufficient for most plans\.)$/;
|
||||
const yesDescription = /^(?:Run Codex against OpenAI design hard rules \+ litmus checks, and a separate Claude subagent for an independent completeness review\. Adds time but catches anything a single-model pass misses\.|Launches Codex design critique \+ Claude subagent completeness review in parallel before the 7 passes\. Adds 1[–-]2 minutes\.)$/;
|
||||
if (labels.some((label, index) => descriptions[index] &&
|
||||
!(/^No\b/.test(label) ? noDescription : yesDescription).test(descriptions[index]!))) return false;
|
||||
const no = labels.filter(label => /^No(?:\s*[,—–-]\s*|\s+)proceed without$/i.test(label));
|
||||
const yes = labels.filter(label => /^Yes(?:\s*[,—–-]\s*|\s+)run (?:outside(?: design)? voices|Codex \+ Claude subagent)$/i.test(label));
|
||||
return labels.length === 2 && no.length === 1 && yes.length === 1 &&
|
||||
q.options.some(option => option.label === call.answers?.[q.question]);
|
||||
}
|
||||
|
||||
/** A numbered design-system amendment can be the first review decision. */
|
||||
function numberedVisualHierarchyFinding(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.answered !== true || call.failed !== false || !call.sessionId || !call.toolUseId ||
|
||||
call.questions.length !== 1 || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return false;
|
||||
const q = call.questions[0]!;
|
||||
const finding = /^Gap ([1-9]\d*) of ([1-9]\d*)\s*[—–-]\s*([A-Za-z][A-Za-z0-9_-]{0,39}) button visual hierarchy: apply DESIGN\.md primary button style\?$/i.exec(q.question.trim());
|
||||
if (!finding || Number(finding[1]) > Number(finding[2]) ||
|
||||
!new RegExp(`^Gap ${finding[1]}: Button$`, 'i').test(q.header.trim()) || q.multiSelect || q.options.length !== 2 ||
|
||||
fp.options.length !== 2 || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question])) return false;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/\s*\(recommended\)\s*$/i, ''));
|
||||
const apply = labels.findIndex(s => /^Apply DESIGN\.md fix$/i.test(s));
|
||||
const defer = labels.findIndex(s => /^Defer to implementation$/i.test(s));
|
||||
if (apply < 0 || defer < 0 || apply === defer) return false;
|
||||
const control = '[A-Za-z][A-Za-z0-9_-]{0,39}';
|
||||
const amendment = new RegExp(`^Add to plan: ${finding[3]} gets #[0-9a-f]{6} filled \\+ (?:white|black) text \\(primary\\); ${control}(?:, ${control})*(?:,? and ${control})? get neutral ghost style\\. Closes the visual hierarchy gap exactly as DESIGN\\.md specifies\\. Implementation task T[1-9]\\d* becomes committed\\.$`, 'i');
|
||||
// Both offered bodies describe the actual style amendment or its deferral;
|
||||
// readiness, a source-selection question, or an example is not this finding.
|
||||
return amendment.test(q.options[apply]!.description?.trim() ?? '') &&
|
||||
/^Leave the gap named but unresolved\. Engineer decides the button styles at implementation time without a spec\. Risk: inconsistency with the design system or re-work after review\.$/i.test(q.options[defer]!.description?.trim() ?? '');
|
||||
}
|
||||
|
||||
/** A qidless Issue with its own design gap is a finding, independent of D numbering. */
|
||||
function ordinaryDesignIssue(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.answered !== true || call.failed !== false || !call.sessionId || !call.toolUseId ||
|
||||
call.questions.length !== 1 || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return false;
|
||||
const q = call.questions[0]!;
|
||||
const title = q.question.split('\n')[0]!.trim();
|
||||
// This primary-action decision can name the control in its Issue header.
|
||||
// F labels annotate findings; they do not establish review identity alone.
|
||||
const headerActionIssue = /^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*)(?: \(F[1-9]\d*\))?: (How should the header action group establish the primary action)\?$/i.exec(title);
|
||||
const signaledPrimaryIssue = /^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*) \(G[1-9]\d*\): (How should the header action group signal that [A-Za-z][A-Za-z0-9 _-]{0,39} is the primary action)\?$/i.exec(title);
|
||||
const distinguishedPrimaryIssue = /^D[1-9]\d*\s*[—–:-]\s*Issue ([1-9]\d*): How should ([A-Za-z][A-Za-z0-9 _-]{0,39}) (?:be distinguished|stand out) from ([A-Za-z][A-Za-z0-9 ,_-]{0,119}?)(?: in the header)?\?$/i.exec(title);
|
||||
const questionIssue = /^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*)(?: \((?:(?:G[1-9]\d*|Pass [1-7]), )?(?:Visual Hierarchy|Spacing|Color|Typography|Motion)\))?: ([^?]+)\?$/i.exec(title) ?? headerActionIssue ?? signaledPrimaryIssue;
|
||||
// A declaration can own the same primary-action decision. Its body and
|
||||
// native choices below must prove the gap, complete styling and deferral.
|
||||
const declaredPrimaryIssue = (!questionIssue || /\nELI10: (?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) header buttons currently /i.test(q.question)) &&
|
||||
/^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*)(?: \(G[1-9]\d*\))?: ([^?\n]+)\??$/i.exec(title);
|
||||
const issue = questionIssue || declaredPrimaryIssue;
|
||||
const declaredGap = declaredPrimaryIssue && /\(G([1-9]\d*)\)/.exec(title)?.[1];
|
||||
const descriptivePrimaryHeader = (signaledPrimaryIssue && /^(?!(?:focus|scope|setup|routing|learnings|outside voices|next steps?)$)[A-Za-z][A-Za-z _-]{0,39}$/i.test(q.header.trim())) ||
|
||||
(distinguishedPrimaryIssue && /^(?:Visual )?Hierarchy$/i.test(q.header.trim()));
|
||||
if (!issue || !(new RegExp(`^Issue ${issue[1]}(?:: [A-Za-z][A-Za-z0-9 _-]{0,39})?$`, 'i').test(q.header.trim()) || descriptivePrimaryHeader) ||
|
||||
/<gstack-qid:/i.test(q.question) || q.multiSelect ||
|
||||
q.options.length < 2 || new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
fp.options.length !== q.options.length || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question])) return false;
|
||||
// The numbered headline must ask about a concrete design requirement.
|
||||
// Reviewer participation or workflow navigation can also use Issue labels.
|
||||
if (!distinguishedPrimaryIssue && !/\b(?:buttons?|primary(?: header)? actions?|primary emphasis|hierarchy|spacing|contrast|colou?rs?|labels?|typography|fonts?|loading|spinner|skeleton|motion)\b/i.test(issue[2]!)) return false;
|
||||
const opposed = q.options.filter(o => /^(?:[1-9]\d*[A-Z](?:[).:]\s*|\s+))?(?:Defer|Decline|Leave|Keep|Accept the gap)\b/i.test(o.label) ||
|
||||
(distinguishedPrimaryIssue && /^(?:Keep|Leave)\b/i.test(o.description?.trim() ?? '')));
|
||||
const repair = !headerActionIssue && !distinguishedPrimaryIssue && !declaredPrimaryIssue && /\b(?:fix|resolve|address)\b/i.test(title) &&
|
||||
q.options.some(o => /\b(?:closing|closes|fixes|resolves?|applies?)\b/i.test(o.description ?? ''));
|
||||
// A source citation alone can describe a report or the next reviewer.
|
||||
// Bind the alternate wording to a named control's concrete style amendment
|
||||
// and the opposed choice that leaves the documented violation unresolved.
|
||||
const primary = /^Make ([A-Za-z][A-Za-z0-9 _-]{0,39}) the (?:visible|visually|(?:only|single)(?: filled| visually)?) primary (?:header )?action(?: in the header)?$/i.exec(issue[2]!) ??
|
||||
/^Give ([A-Za-z][A-Za-z0-9 _-]{0,39}) primary emphasis in the header action group$/i.exec(issue[2]!) ??
|
||||
/^How should the header action group signal that ([A-Za-z][A-Za-z0-9 _-]{0,39}) is the primary action$/i.exec(issue[2]!) ??
|
||||
/^How should (?:the )?(?:header )?actions establish that ([A-Za-z][A-Za-z0-9 _-]{0,39}) is the primary action$/i.exec(issue[2]!) ??
|
||||
(distinguishedPrimaryIssue && /^How should ([A-Za-z][A-Za-z0-9 _-]{0,39}) (?:be distinguished|stand out) from /i.exec(issue[2]!)) ??
|
||||
(headerActionIssue && new RegExp(`^Issue ${issue[1]}: ([A-Za-z][A-Za-z0-9 _-]{0,39})$`, 'i').exec(q.header.trim()));
|
||||
if (declaredPrimaryIssue && (!primary || q.options.length > 4 || Object.keys(call.answers ?? {}).length !== 1)) return false;
|
||||
const primaryEmphasisIssue = !!signaledPrimaryIssue || !!distinguishedPrimaryIssue || !!declaredPrimaryIssue || /^Give [A-Za-z][A-Za-z0-9 _-]{0,39} primary emphasis in the header action group$/i.test(issue[2]!);
|
||||
const scopedPrimaryStatus = !!headerActionIssue || primaryEmphasisIssue;
|
||||
const explicitStyle = primary && `${primary[1]} filled (?:primary )?#[0-9a-f]{6}(?:/| with )(?:white|black)(?: text)?; ` +
|
||||
'[A-Za-z][A-Za-z0-9 ,/_-]{0,99} neutral ghost(?: buttons)?\\.';
|
||||
const amendments = primary && [
|
||||
new RegExp(`^(?:✅\\s*)?Matches DESIGN\\.md exactly: ${primary[1]} filled #[0-9a-f]{6} with (?:white|black) text; ` +
|
||||
'[A-Za-z][A-Za-z0-9 ,_-]{0,99} as neutral ghost buttons\\.', 'i'),
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]} becomes the (?:single|one|only) filled(?: primary)?(?: button)? \\(#[0-9a-f]{6}, (?:white|black) text\\); ` +
|
||||
'[A-Za-z][A-Za-z0-9 /,_-]{0,99} (?:become|are) neutral ghost buttons (?:exactly as DESIGN\\.md specifies|per DESIGN\\.md)\\b', 'i'),
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]} is the (?:single|only) filled #[0-9a-f]{6} button; ` +
|
||||
'[A-Za-z][A-Za-z0-9 /,_-]{0,99} become neutral ghosts, exactly (?:per DESIGN\\.md|as DESIGN\\.md prescribes)\\b', 'i'),
|
||||
new RegExp(`^(?:✅\\s*)?Apply DESIGN\\.md(?: tokens)?: ${primary[1]} #[0-9a-f]{6} filled(?: with)? (?:white|black) text; ` +
|
||||
'[A-Za-z][A-Za-z0-9 ,/_-]{0,99} neutral ghost(?: buttons)?\\.', 'i'),
|
||||
// The same concrete style can cite DESIGN.md before or after its tokens.
|
||||
new RegExp(`^(?:✅\\s*)?(?:Apply DESIGN\\.md(?: tokens)?: ${explicitStyle}|${explicitStyle} Exact DESIGN\\.md\\.)`, 'i'),
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]}\\s*=\\s*filled #[0-9a-f]{6} with (?:white|black) text; ` +
|
||||
'[A-Za-z][A-Za-z0-9 ,/_-]{0,99}\\s*=\\s*neutral ghost(?: buttons?)?,? per DESIGN\\.md\\b', 'i'),
|
||||
];
|
||||
const primaryHeader = !q.header.includes(':') || q.header.split(':')[1]!.trim().toLowerCase() === primary?.[1]?.toLowerCase();
|
||||
const ownedStatus = (value: string, index: number, source: string) =>
|
||||
/^(?:withdrawn|superseded|resolved|closed|historical|hypothetical|rejected|cancelled|canceled|not current|no longer current)$/i.test(value) &&
|
||||
/(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:(?:This (?:issue|finding|question|amendment|deferral|style|fix|remedy|choice|option|(?:DESIGN\.md |token )?(?:requirement|contract))|(?:Issue |G)[1-9]\d*) (?:is|was|has been)|(?:these|the|this) (?:tokens?|styles?|primary treatment) (?:are|is|were|was|have been|has been)) $/i.test(source.slice(0, index));
|
||||
// The style wordings share one owned decision: a current equal-weight gap,
|
||||
// a named control's DESIGN.md amendment, and a different choice retaining it.
|
||||
// A following status assertion remains current after a parenthesized effort
|
||||
// estimate. Preserve the estimate and expose its boundary to the same guards.
|
||||
const currentText = (text: string) => (scopedPrimaryStatus
|
||||
? text.replace(/(\(human: ~?[0-9]+(?:\.[0-9]+)?(?:h|min) \/ CC: ~?[0-9]+(?:\.[0-9]+)?(?:h|min)\))(?=\s+\S)/g, '$1.')
|
||||
: text)
|
||||
.replace(/```[\s\S]*?(?:```|$)|~~~[\s\S]*?(?:~~~|$)/g, '')
|
||||
.replace(/^(?:\s*>| {4}|\t).*$/gm, '')
|
||||
.replace(/`([^`]+)`/g, (_, body: string, index: number, source: string) =>
|
||||
scopedPrimaryStatus && ownedStatus(body, index, source) ? body : /\s/.test(body) ? '' : body)
|
||||
// A quoted status scalar remains a current assertion when its unquoted
|
||||
// subject names this decision; whole quoted historical prose stays absent.
|
||||
.replace(scopedPrimaryStatus
|
||||
? /"[^"\n]*"|“[^”\n]*”|(?<!\w)'[^'\n]*'(?!\w)|‘[^’\n]*’/g
|
||||
: /"[^"\n]*"|“[^”\n]*”/g, (quoted: string, index: number, source: string) =>
|
||||
ownedStatus(quoted.slice(1, -1), index, source)
|
||||
? quoted.slice(1, -1) : '').replace(/\*\*/g, '');
|
||||
const questionText = currentText(q.question);
|
||||
const assessments = [...questionText.matchAll(/^ELI10: (.+)$/gm)];
|
||||
const prefix = questionText.slice(0, assessments[0]?.index ?? 0)
|
||||
.split('\n').filter(line => line.trim()).slice(1);
|
||||
const sourceAssessment = /\b(?:historical|hypothetical|quoted|source|earlier review)\s+(?:example|excerpt|assessment|material|text)\b|\bnot\s+(?:the\s+)?current\s+(?:UI|assessment|finding|amendment|deferral|remedy|choice|option)\b|\bthis (?:finding|amendment|deferral|remedy|choice|option) (?:applies only to|belongs to) (?:an? )?(?:another|different) (?:project|plan|review)\b/i;
|
||||
const assessment = assessments.length === 1 &&
|
||||
prefix.every(line => /^(?:Project\/branch\/task:|\[P[0-3]\])/.test(line)) &&
|
||||
!/^(?:Project\/branch\/task:|\[P[0-3]\])\s*(?:If|When|Unless|Provided|Assuming)\b/im.test(prefix.join('\n')) &&
|
||||
!sourceAssessment.test(prefix.join(' ')) && !sourceAssessment.test(assessments[0]![1]!)
|
||||
? assessments[0]![1]! : '';
|
||||
const headerPeers = primary && headerActionIssue && new RegExp(`^${primary[1]}, ([A-Za-z][A-Za-z0-9 _-]{0,39}(?:, [A-Za-z][A-Za-z0-9 _-]{0,39})*(?:,? and [A-Za-z][A-Za-z0-9 _-]{0,39})?) currently look (?:the same|identical)\\.`, 'i').exec(assessment);
|
||||
// Equal visual properties can establish the same current lack of hierarchy.
|
||||
// Shared geometry alone is not a claim that the actions look equally primary.
|
||||
const properties = '(?:size|weight|colou?r|fill|emphasis)(?:(?:, ?|,? and )(?:size|weight|colou?r|fill|emphasis))*';
|
||||
const equalProperties = distinguishedPrimaryIssue && new RegExp('^(?:Right now|Today) (?:all|the) (two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) header buttons (?:are|have|share) the same (' + properties + ')\\.', 'i').exec(assessment);
|
||||
const countedHeader = equalProperties && /\b(?:weight|colou?r|fill|emphasis)\b/i.test(equalProperties[2]!) && equalProperties ||
|
||||
distinguishedPrimaryIssue && /^(?:Right now|Today) (?:all|the) (two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) header buttons look (?:the same|identical)\./i.exec(assessment) ||
|
||||
declaredPrimaryIssue && /^(two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) header buttons currently (?:share one style|look identical)\./i.exec(assessment);
|
||||
const primaryAssessment = distinguishedPrimaryIssue || declaredPrimaryIssue ? countedHeader?.[0] : headerActionIssue ? headerPeers?.[0] :
|
||||
primary && new RegExp(`^(?:Right now|Today) ${primary[1]}(?:, [A-Za-z][A-Za-z0-9 _-]{0,39})+(?:,? and [A-Za-z][A-Za-z0-9 _-]{0,39})? (?:(?:all )?look (?:the same|identical)|are (?:all )?(?:(?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) )?identical buttons)\\b`, 'i').exec(assessment)?.[0];
|
||||
const premiseSentence = assessment.split(/[.!?](?:\s|$)/)[0] ?? '';
|
||||
const currentPrimary = !!primaryAssessment && !/\b(?:not|never|no longer)\b/i.test(primaryAssessment) &&
|
||||
!/\b(?:archived|historical|hypothetical|quoted|example|previous|earlier)\b/i.test(premiseSentence);
|
||||
// The current assessment can state the full token contract while an offered
|
||||
// amendment names the existing component variants that implement it.
|
||||
const numberValue = (value: string) => /^\d+$/.test(value) ? Number(value) :
|
||||
['zero', 'one', 'two', 'three', 'four', 'five', 'six', 'seven', 'eight', 'nine', 'ten'].indexOf(value.toLowerCase());
|
||||
const controlNames = (text: string) => text.toLowerCase().split(/,\s*(?:and\s+)?|\s+and\s+/).map(s => s.trim()).sort();
|
||||
const headerControls = distinguishedPrimaryIssue ? controlNames(distinguishedPrimaryIssue[3]!) : headerPeers ? controlNames(headerPeers[1]!) : [];
|
||||
const otherControls = headerActionIssue || distinguishedPrimaryIssue ? headerControls.length : primary && primaryAssessment
|
||||
? primaryAssessment.replace(new RegExp(`^(?:Right now|Today) ${primary[1]},\\s*`, 'i'), '')
|
||||
.replace(/\s+(?:(?:all )?look (?:the same|identical)|are (?:all )?(?:(?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) )?identical buttons)$/, '')
|
||||
.split(/,\s*(?:and\s+)?|\s+and\s+/).length : 0;
|
||||
const variantContract = primary && new RegExp(`(?:^|[.!?]\\s+)DESIGN\\.md already says ${primary[1]} is the only filled button ` +
|
||||
'\\(#[0-9a-f]{6} with (?:white|black) text(?:, about [0-9]+(?:\\.[0-9]+)?:1 contrast)?\\) and the other ' +
|
||||
'(two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) are neutral ghost buttons\\.', 'i').exec(assessment);
|
||||
const statusBoundary = scopedPrimaryStatus ? '[.!?;]' : '[.!?]';
|
||||
const invalidContract = new RegExp(`(?:^|${statusBoundary}\\s+|\\n)(?:Correction:\\s*)?(?:this|that|the) (?:(?:DESIGN\\.md|token) )?(?:requirement|contract) (?:is|was|has been) (?:withdrawn|superseded|rejected|cancelled|canceled|not current|no longer current)\\b`, 'i');
|
||||
const namedContract = primary && new RegExp(`(?:^|[.!?]\\s+)DESIGN\\.md already says ${primary[1]} is the only filled primary button and the other (two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) are neutral ghost buttons\\.`, 'i').exec(assessment);
|
||||
const headerContract = primary && headerActionIssue && new RegExp(`(?:^|[.!?]\\s+)DESIGN\\.md already answers it: ${primary[1]} is the only filled primary button, the other (two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) are neutral ghost buttons\\.`, 'i').exec(assessment);
|
||||
const conditionalHeader = (text: string) => /(?:^|[.!?;]\s+|\n)(?:[✅❌]\s*)?(?:Correction:\s*)?(?:If|When|Unless|Assuming|Provided)\b/i.test(text) || /\b(?:only if|unless|pending approval|subject to approval)\b/i.test(text);
|
||||
// Approval conditions suspend this offered decision; explanatory conditions
|
||||
// about user behavior do not make an otherwise current amendment optional.
|
||||
const pendingPrimaryApproval = (text: string) => primaryEmphasisIssue && (
|
||||
/(?:^|[.!?;]\s+|\n)(?:[✅❌]\s*)?(?:Correction:\s*)?(?:If|When|Once|Provided|Assuming|Pending)\s+(?:approval|approved|acceptance|accepted|(?:we|you)\s+(?:approve|accept))\b/i.test(text) ||
|
||||
(declaredPrimaryIssue && new RegExp(`(?:^|[.!?;]\\s+|\\n)(?:Correction:\\s*)?(?:This (?:issue|finding|amendment|deferral|option)|Issue ${issue[1]}${declaredGap ? `|G${declaredGap}` : ''}) (?:requires approval|applies only if approved)\\b`, 'i').test(text)));
|
||||
const currentHeaderContract = declaredPrimaryIssue ? countedHeader && !conditionalHeader(questionText) : distinguishedPrimaryIssue ? (countedHeader && headerControls.length > 0 &&
|
||||
new Set(headerControls).size === headerControls.length && !headerControls.includes(primary![1]!.toLowerCase()) &&
|
||||
numberValue(countedHeader[1]!) === headerControls.length + 1 && !conditionalHeader(questionText)) : !headerActionIssue || (headerContract && headerControls.length > 0 &&
|
||||
new Set(headerControls).size === headerControls.length && !headerControls.includes(primary![1]!.toLowerCase()) &&
|
||||
numberValue(headerContract[1]!) === headerControls.length && !conditionalHeader(questionText));
|
||||
const statedVariant = (variantContract || namedContract) &&
|
||||
numberValue((variantContract || namedContract)![1]!) === otherControls &&
|
||||
!/\b(?:proposed|hypothetical|quoted|historical|source)\s+(?:example|contract|requirement)\b/i.test(assessment.slice(0, (variantContract || namedContract)!.index)) &&
|
||||
!invalidContract.test(questionText);
|
||||
const withdrawn = new RegExp(`(?:^|${statusBoundary}\\s+|\\n)(?:Correction:\\s*)?(?:(?:This (?:issue|finding|question|amendment|deferral|style|fix|remedy|choice|option)|Issue ${issue[1]}${declaredGap ? `|G${declaredGap}` : ''}) (?:is|was|has been) (?:withdrawn|superseded|resolved|closed|historical|hypothetical|rejected|cancelled|canceled|not current|no longer current)|We have (?:resolved|closed|withdrawn) this (?:issue|finding)|No current (?:issue|finding|gap|violation) (?:remains|exists))\\b`, 'i');
|
||||
const closedGap = /(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:this|the|that) (?:gap|violation) (?:is|was|has been) (?:already\s+|now\s+)?(?:resolved|fixed|closed)\b/i;
|
||||
const cancelledStyle = /(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:do not|don't|never|skip|cancel|withdraw)\s+(?:apply|use|add|keep)\s+(?:(?:these|the|this)\s+)?(?:tokens?|styles?|primary treatment)\b/i;
|
||||
const withdrawnStyles = /(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:these|the|this) (?:tokens?|styles?|primary treatment) (?:are|is|were|was|have been|has been) (?:withdrawn|rejected|cancelled|canceled|not current|no longer current)\b/i;
|
||||
const choiceIds = q.options.map(o => /^([1-9]\d*)[A-Z](?:[).:]?\s+)/.exec(o.label));
|
||||
const primaryRepair = primaryHeader && amendments && currentPrimary && currentHeaderContract &&
|
||||
prefix.filter(line => /^Project\/branch\/task:/.test(line)).length === 1 &&
|
||||
!!call.answeredAt && Number.isFinite(Date.parse(call.answeredAt)) &&
|
||||
choiceIds.every(id => id?.[1] === issue[1]) &&
|
||||
!pendingPrimaryApproval(questionText) &&
|
||||
!sourceAssessment.test(questionText) &&
|
||||
!withdrawn.test(questionText) && !closedGap.test(questionText) && !withdrawnStyles.test(questionText) && !invalidContract.test(questionText) &&
|
||||
q.options.some(amendment => {
|
||||
const body = currentText(amendment.description ?? '');
|
||||
if (pendingPrimaryApproval(body)) return false;
|
||||
// Roles and their concrete tokens belong to one native option; a familiar
|
||||
// label alone cannot supply the style or borrow DESIGN.md from a peer.
|
||||
const roleLabel = currentText(amendment.label);
|
||||
const propertyStyle = primary && distinguishedPrimaryIssue && new RegExp(`^(?:✅\\s*)?${primary[1]} (?:is|becomes) the (?:only|single) filled(?: primary)? #[0-9a-f]{6}(?: button)? with (?:white|black) text; ([A-Za-z][A-Za-z0-9 ,/_-]{0,119}) (?:are|become) neutral ghost(?: buttons)?\\.`, 'i').exec(body);
|
||||
const roleAuthority = new RegExp(`^${issue[1]}[A-Z][).:]?\\s+(?:(?:Apply|Use|Reuse) )?DESIGN\\.md\\b`, 'i').test(roleLabel) ||
|
||||
/(?:^|[.;]\s+)(?:Matches DESIGN\.md exactly|Per DESIGN\.md)\b/i.test(body);
|
||||
const roleStyle = propertyStyle && roleAuthority && !conditionalHeader(roleLabel) &&
|
||||
!/\b(?:not|never|no|if|historical|hypothetical|source|quoted|withdrawn|superseded|cancelled|canceled)\b/i.test(roleLabel) &&
|
||||
!headerControls.some(peer => new RegExp(`\\b(?:primary(?: button| action)? ${peer}|${peer} (?:as )?(?:the )?(?:filled )?primary)\\b`, 'i').test(roleLabel)) ? propertyStyle : null;
|
||||
const declaredStyle = primary && declaredPrimaryIssue &&
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]} filled #[0-9a-f]{6}(?: with)? (?:white|black)(?: text)?; ([A-Za-z][A-Za-z0-9 ,/_-]{0,119}) neutral ghost(?: buttons)?\\.`, 'i').exec(body);
|
||||
if (declaredPrimaryIssue) {
|
||||
const peers = declaredStyle && controlNames(declaredStyle[1]!.replaceAll('/', ','));
|
||||
if (!peers || peers.length !== numberValue(countedHeader![1]!) - 1 ||
|
||||
new Set(peers).size !== peers.length || peers.includes(primary![1]!.toLowerCase()) ||
|
||||
!new RegExp(`^${issue[1]}[A-Z][).:]?\\s+Apply DESIGN\\.md tokens?(?: \\(recommended\\))?$`, 'i').test(amendment.label) ||
|
||||
conditionalHeader(body) || invalidContract.test(body)) return false;
|
||||
}
|
||||
const headerStyle = primary && headerActionIssue && new RegExp(`^(?:✅\\s*)?${primary[1]} becomes the only filled #[0-9a-f]{6} button with (?:white|black) text; ([A-Za-z][A-Za-z0-9 ,_-]{0,119}) become neutral ghost buttons, exactly as DESIGN\\.md states\\.`, 'i').exec(body);
|
||||
// A descriptive header still owns a concrete primary and every peer.
|
||||
// Its native option supplies the primary/secondary roles and tokens.
|
||||
const distinguishedStyle = primary && distinguishedPrimaryIssue && (
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]} is #[0-9a-f]{6} with (?:white|black) text; ([A-Za-z][A-Za-z0-9 ,_-]{0,119}) are neutral ghost buttons per DESIGN\\.md\\.`, 'i').exec(body) ??
|
||||
new RegExp(`^(?:✅\\s*)?${primary[1]} becomes the only filled button \\(#[0-9a-f]{6}, (?:white|black) text\\); ([A-Za-z][A-Za-z0-9 ,/_-]{0,119}) use the existing neutral ghost variant` +
|
||||
'(?: \\(human: ~?[0-9]+(?:\\.[0-9]+)?(?:h|min) / CC: ~?[0-9]+(?:\\.[0-9]+)?(?:h|min)\\))?\\. (?:✅\\s*)?Matches DESIGN\\.md exactly\\b', 'i').exec(body) ?? roleStyle);
|
||||
const style = declaredPrimaryIssue ? declaredStyle?.[0] : distinguishedPrimaryIssue ? distinguishedStyle?.[0] : headerActionIssue ? headerStyle?.[0] : amendments.map(pattern => pattern.exec(body)).find(Boolean)?.[0];
|
||||
if ((declaredPrimaryIssue || distinguishedPrimaryIssue) &&
|
||||
/(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:the|this) (?:current )?(?:amendment|fix) keeps (?:all )?(?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) (?:header )?buttons identical\b/i.test(body)) return false;
|
||||
if (distinguishedPrimaryIssue && (!distinguishedStyle || conditionalHeader(body) || invalidContract.test(body) ||
|
||||
!(roleStyle || new RegExp(`^[1-9]\\d*[A-Z][).:]?\\s+Filled primary (?:(?:\\+|and|with) ghosts|${primary![1]})(?: \\(recommended\\))?$`, 'i').test(amendment.label)) ||
|
||||
JSON.stringify(controlNames(distinguishedStyle[1]!.replaceAll('/', ','))) !== JSON.stringify(headerControls))) return false;
|
||||
if (headerActionIssue && (!headerStyle || conditionalHeader(body) || invalidContract.test(body) ||
|
||||
JSON.stringify(controlNames(headerStyle[1]!)) !== JSON.stringify(headerControls))) return false;
|
||||
const variantLine = /^✅\s*Uses the existing Button primary and ghost variants from DESIGN\.md; no new styles\./m.exec(body);
|
||||
const benefits = variantLine ? body.slice(0, variantLine.index).trim().split('\n').filter(Boolean) : [];
|
||||
const labelledRoles = /^✅ Matches DESIGN\.md exactly: one filled primary, (two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) neutral ghosts, [1-9]\d*px targets kept\./.exec(body);
|
||||
const labelledStyle = statedVariant && namedContract && labelledRoles &&
|
||||
numberValue(labelledRoles[1]!) === otherControls &&
|
||||
new RegExp(`^[1-9]\\d*[A-Z]\\) ${primary![1]} filled #[0-9a-f]{6}/(?:white|black), others ghost(?: \\(recommended\\))?$`, 'i').test(amendment.label);
|
||||
const variantRepair = !headerActionIssue && (labelledStyle || (statedVariant && variantContract && variantLine &&
|
||||
new RegExp(`^[1-9]\\d*[A-Z] Filled ${primary![1]}, ghost others(?: \\(recommended\\))?$`, 'i').test(amendment.label) &&
|
||||
benefits.every(line => /^✅\s*(?!(?:If|When|Unless|Historical|Hypothetical|Quoted|Source|Example)\b)\S/i.test(line)) &&
|
||||
!/\b(?:archived|historical|hypothetical|quoted|previous|earlier)\b/i.test(benefits.join(' ')))) &&
|
||||
!/(?:^|[.!?]\s+|\n)(?:Correction:\s*)?(?:do not|don't|never|skip|cancel|withdraw) (?:apply|use|add|keep) (?:the |these )?(?:Button )?primary and ghost variants\b/i.test(body) &&
|
||||
!/(?:^|[.!?]\s+|\n)(?:Correction:\s*)?(?:these|the|this) (?:tokens?|styles?|variants?) (?:do|does) not match DESIGN\.md\b/i.test(body) &&
|
||||
!/(?:^|[.!?]\s+|\n)(?:Correction:\s*)?(?:the|this) (?:current )?amendment keeps all (?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) buttons identical\b/i.test(body);
|
||||
// A named primary cannot simultaneously occur in the ghost-control list.
|
||||
if ((!style && !variantRepair) || (style && new RegExp(`\\b${primary![1]}\\b`, 'i').test(style.slice(style.indexOf(';') + 1))) ||
|
||||
sourceAssessment.test(body) || withdrawn.test(body) || closedGap.test(body) || cancelledStyle.test(body) || withdrawnStyles.test(body)) return false;
|
||||
return opposed.some(defer => {
|
||||
const declined = currentText(defer.description ?? '');
|
||||
if (pendingPrimaryApproval(declined)) return false;
|
||||
if (declaredPrimaryIssue) return defer !== amendment &&
|
||||
new RegExp(`^${issue[1]}[A-Z][).:]?\\s+Defer(?: \\(recommended\\))?$`, 'i').test(defer.label) &&
|
||||
new RegExp(`^Leave ${declaredGap ? `G${declaredGap}` : `Issue ${issue[1]}`} open and record it as unresolved\\.`, 'i').test(declined) &&
|
||||
!new RegExp(`(?:^|[.!?;]\\s+|\\n)(?:Correction:\\s*)?(?:do not|don't|never|skip|cancel|withdraw) (?:leave|keep|defer) (?:${declaredGap ? `G${declaredGap}|` : ''}Issue ${issue[1]})\\b`, 'i').test(declined) &&
|
||||
!conditionalHeader(declined) && !sourceAssessment.test(declined) && !withdrawn.test(declined) &&
|
||||
!closedGap.test(declined) && !invalidContract.test(declined) && !cancelledStyle.test(declined) && !withdrawnStyles.test(declined);
|
||||
// Native menus can list current benefits before the gap retained by
|
||||
// declining. Only consume a complete affirmative pro/con prefix; prose
|
||||
// framing a source example or a future condition cannot expose an icon.
|
||||
const headerDeferral = /^(?:Leave|Keep) the header unchanged and record the gap as (?:debt|an open issue)\.\s*/i.exec(declined);
|
||||
const deferralBody = headerDeferral ? declined.slice(headerDeferral[0].length) : declined;
|
||||
const cancelledHeaderDeferral = headerDeferral && /(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:do not|don't|never|skip|cancel|withdraw) (?:keep|leave) (?:the )?header unchanged\b/i.test(declined);
|
||||
const pros = /^(?:✅(?!\s*(?:If|When|Unless|Historical|Hypothetical|Quoted|Source|Example)\b)\s*[^✅❌]+)+❌\s*/i.exec(deferralBody);
|
||||
const remaining = pros && !sourceAssessment.test(pros[0]) ? deferralBody.slice(pros[0].length) : deferralBody;
|
||||
const retainedEmphasis = distinguishedPrimaryIssue && new RegExp(`^(?:Keep|Leave) identical (?:header )?buttons, bold (?:the )?${primary![1]} (?:text|label)\\. (?:Weak(?: visual)? signal, )?off-token\\.`, 'i').test(remaining);
|
||||
const retainedHierarchyGap = distinguishedPrimaryIssue && /^(?:❌\s*)?Ships a (?:known|documented) DESIGN\.md violation and the plan['’]s own Visual Hierarchy gap (?:stays|remains) open\./i.test(remaining);
|
||||
const retainedRoleGap = roleStyle && /^(?:No change[.;]\s*)?(?:the |this )?(?:finding|issue|gap) (?:stays|remains) (?:open|unresolved)\b/i.test(remaining);
|
||||
if (distinguishedPrimaryIssue) return defer !== amendment && !!(retainedEmphasis || retainedHierarchyGap || retainedRoleGap) &&
|
||||
!conditionalHeader(declined) && !sourceAssessment.test(declined) && !withdrawn.test(declined) &&
|
||||
!/(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:do not|don't|never|skip|cancel|withdraw) (?:keep|leave) identical (?:header )?buttons\b/i.test(declined) &&
|
||||
!closedGap.test(declined) && !invalidContract.test(declined) && !cancelledStyle.test(declined) && !withdrawnStyles.test(declined);
|
||||
const retainedButtons = /^(?:❌\s*)?Keep all (two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) (?:header )?buttons identical; gap stays documented\./i.exec(remaining);
|
||||
const cancelledRetainedButtons = /(?:^|[.!?;]\s+|\n)(?:Correction:\s*)?(?:do not|don't|never|skip|cancel|withdraw) (?:keep|leave) (?:all )?(two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) (?:header )?buttons identical\b/i.exec(declined);
|
||||
if (headerActionIssue) {
|
||||
const keep = /^[1-9]\d*[A-Z]: Keep all (two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) identical(?: \(recommended\))?$/i.exec(defer.label);
|
||||
return defer !== amendment && keep && numberValue(keep[1]!) === headerControls.length + 1 &&
|
||||
!conditionalHeader(declined) && !sourceAssessment.test(declined) && !withdrawn.test(declined) && !invalidContract.test(declined) &&
|
||||
!closedGap.test(declined) && !cancelledRetainedButtons &&
|
||||
/^(?:❌\s*)?Ships a (?:known|documented) DESIGN\.md violation; the review score stays capped and users keep scanning a flat row\./i.test(remaining);
|
||||
}
|
||||
return defer !== amendment && !sourceAssessment.test(declined) && !withdrawn.test(declined) && !closedGap.test(declined) && !cancelledHeaderDeferral &&
|
||||
((retainedButtons && numberValue(retainedButtons[1]!) === otherControls + 1 &&
|
||||
(!cancelledRetainedButtons || numberValue(cancelledRetainedButtons[1]!) !== otherControls + 1)) ||
|
||||
/^(?:❌\s*)?(?:Leaves a documented DESIGN\.md violation in place|Keeps the documented DESIGN\.md violation and the scan problem|Violates DESIGN\.md and leaves the mis-click on [A-Za-z][A-Za-z /_-]{0,79} unaddressed|Ships a (?:known|documented) DESIGN\.md violation and the primary action (?:stays|remains) undiscoverable|Ships a header with no primary action; PLAN\.md['’]s own gap stays open|Ships the documented violation;[^.\n]*\bthe gap remains open|Primary-action ambiguity ships; documented DESIGN\.md violation remains|Decline the fix; gap stays documented and lowers the score|Decline the fix; document the violation as accepted|Keep all (?:two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) identical; record as an open DESIGN\.md violation)\b/i.test(remaining) ||
|
||||
(variantRepair && /^(?:Violates DESIGN\.md and leaves users guessing which action is primary; Pass [1-7] stays at [0-9](?:\.[0-9]+)?\/10|Documented DESIGN\.md violation ships and Pass [1-7] stays at [0-9](?:\.[0-9]+)?\/10)\.$/i.test(remaining)));
|
||||
});
|
||||
});
|
||||
return opposed.length > 0 && !!(repair || primaryRepair);
|
||||
}
|
||||
|
||||
/** A design-system choice can name the gap without using an imperative repair verb. */
|
||||
function designSystemChoiceIssue(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.answered !== true || call.failed !== false || !call.sessionId || !call.toolUseId ||
|
||||
call.questions.length !== 1 || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0) ||
|
||||
!call.answeredAt || !Number.isFinite(Date.parse(call.answeredAt))) return false;
|
||||
const q = call.questions[0]!;
|
||||
const lines = q.question.trim().split('\n');
|
||||
const issue = /^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*): (.+)\?$/.exec(lines[0]!);
|
||||
if (!issue || q.header.trim() !== `Issue ${issue[1]}` || lines.length !== 7 ||
|
||||
!/^Project\/branch\/task: [^\n,]+ on [^\n,]+, PLAN\.md design review, Pass [1-7] [A-Za-z][A-Za-z &()-]+\.$/.test(lines[1]!) ||
|
||||
!/^ELI10: \S/.test(lines[2]!) || !/\bDESIGN\.md\b/.test(lines[2]!) ||
|
||||
!/^Stakes if we pick wrong: \S/.test(lines[3]!) || !/^Recommendation: \S/.test(lines[4]!) ||
|
||||
!/^Completeness: \S/.test(lines[5]!) || !/^Net: \S/.test(lines[6]!) ||
|
||||
/<gstack-qid:|```|^ELI10: (?:Example|Hypothetical|Quoted)\b/im.test(q.question) || q.multiSelect ||
|
||||
q.options.length < 2 || q.options.length > 4 || new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
!q.options.every(o => new RegExp(`^${issue[1]}[A-Z]: \\S`).test(o.label)) ||
|
||||
fp.options.length !== q.options.length || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question])) return false;
|
||||
// These are current visual/interaction choices, not reviewer participation or next-step routing.
|
||||
const subjects = [
|
||||
/^How should [A-Z][A-Za-z0-9 _/-]{0,79} be distinguished from [A-Z][A-Za-z0-9 ,/_-]{0,119}$/,
|
||||
/^What does the user see while [A-Z][A-Za-z0-9 _/-]{0,79} is pending for [1-9]\d*(?:[-–][1-9]\d*)? seconds$/,
|
||||
/^What type scale should (?:form )?labels(?: and section headings)? use$/,
|
||||
/^What vertical spacing rhythm should the form use$/,
|
||||
/^How should (?:the )?error message meet WCAG AA contrast$/,
|
||||
];
|
||||
const subject = subjects.findIndex(pattern => pattern.test(issue[2]!));
|
||||
if (subject < 0) return false;
|
||||
const assessments = [/^ELI10: The header shows\b/, /^ELI10: After clicking\b/,
|
||||
/^ELI10: Labels on the form are set\b/, /^ELI10: Gaps between sections are\b/, /^ELI10: The error message is\b/];
|
||||
if (!assessments[subject]!.test(lines[2]!) ||
|
||||
/(?:^|[.!?]\s+)(?:This (?:issue|finding) (?:is|has been) (?:withdrawn|resolved|closed)|We have (?:resolved|closed|withdrawn) this (?:issue|finding)|No current (?:issue|finding|gap|defect|violation) (?:remains|exists))\b/i.test(lines[2]!.slice(7))) return false;
|
||||
const control = /^How should (.+) be distinguished from /.exec(issue[2]!)?.[1];
|
||||
const concrete = [new RegExp(`^${control}\\b[^\\n]*\\b(?:filled|ghost|outlined|primary)\\b`, 'i'),
|
||||
/^(?:Spinner|InlineStatus|Static indicator)\b/i, /^[1-9]\d*px\b/i, /^[1-9]\d*px\b/i, /^#[0-9a-f]{6}\b/i][subject]!;
|
||||
const conforming = q.options.filter(o => concrete.test(o.label.replace(/^[1-9]\d*[A-Z]: /, '')) &&
|
||||
(/^✅ Exact(?:ly)? (?:the (?:two )?)?DESIGN\.md\b/.test(o.description ?? '') ||
|
||||
(subject === 0 && new RegExp(`^✅ ${control} is [^\\n]+\\bexactly per DESIGN\\.md\\b`).test(o.description ?? ''))));
|
||||
return conforming.some(choice => q.options.some(o => o !== choice &&
|
||||
/^❌ (?:Ships (?:the documented violation|a known WCAG AA failure)\b|Deviates from the DESIGN\.md\b)/m.test(o.description ?? '')));
|
||||
}
|
||||
|
||||
/** Named decision fields may be compact prose; native choices still own the finding. */
|
||||
function compactPrimaryDecision(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.answered !== true || call.failed !== false || !call.sessionId || !call.toolUseId ||
|
||||
!call.answeredAt || !Number.isFinite(Date.parse(call.answeredAt)) || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length !== 2 || new Set(q.options.map(o => o.label)).size !== 2 ||
|
||||
fp.options.length !== 2 || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question]) || /<gstack-qid:/i.test(q.question)) return false;
|
||||
const headline = /^(?:D[1-9]\d*\s*[—–:-]\s*)?Issue ([1-9]\d*): ([A-Za-z][A-Za-z0-9 _-]{0,39}) (?:has no|lacks) primary[- ]action hierarchy\./i.exec(q.question.trim());
|
||||
if (!headline || q.header.trim() !== `Issue ${headline[1]}`) return false;
|
||||
const issue = headline[1]!, control = headline[2]!;
|
||||
const inactive = 'withdrawn|superseded|resolved|closed|hypothetical|unproven|rejected|cancelled|canceled|deferred|not current|no longer current';
|
||||
const owner = `(?:This (?:issue|finding|question|amendment|deferral|style|fix|remedy|choice|option)|Issue ${issue}|(?:These|The|This) (?:tokens?|styles?|primary treatment))`;
|
||||
const boundary = '(?:^|[.!?;]\\s+|\\n)(?:Correction:\\s*)?';
|
||||
const scalarPrefix = new RegExp(`${boundary}${owner} (?:is|was|are|were|has been|have been) $`, 'i');
|
||||
const current = (value: string) => value
|
||||
.replace(/```[\s\S]*?(?:```|$)|~~~[\s\S]*?(?:~~~|$)/g, '')
|
||||
.replace(/^(?:\s*>| {4}|\t).*$/gm, '')
|
||||
.replace(/`[^`\n]*`|"[^"\n]*"|“[^”\n]*”|(?<!\w)'[^'\n]*'(?!\w)|‘[^’\n]*’/g, (quoted, index, source) =>
|
||||
new RegExp(`^(?:${inactive})$`, 'i').test(quoted.slice(1, -1)) && scalarPrefix.test(source.slice(0, index)) ? quoted.slice(1, -1) : '')
|
||||
.replace(/\*\*/g, '');
|
||||
const invalid = (value: string) =>
|
||||
new RegExp(`${boundary}${owner} (?:is|was|are|were|has been|have been) (?:${inactive})\\b`, 'i').test(value) ||
|
||||
new RegExp(`${boundary}(?:(?:This|The) (?:gap|violation) (?:is|was|has been) (?:already |now )?(?:resolved|fixed|closed)|No current (?:gap|issue|finding|violation) (?:remains|exists))\\b`, 'i').test(value) ||
|
||||
new RegExp(`${boundary}(?:If|When|Once|Provided|Assuming|Pending) (?:approval|approved|acceptance|accepted|(?:we|you) (?:approve|accept))\\b`, 'i').test(value) ||
|
||||
new RegExp(`${boundary}(?:Do not|Don't|Never|Skip|Cancel|Withdraw) (?:apply|use|add|keep) (?:this (?:fix|amendment)|(?:the |these )?(?:tokens?|styles?|primary treatment))\\b`, 'i').test(value) ||
|
||||
new RegExp(`${boundary}(?:${control} (?:already (?:is|has)|is already) (?:the (?:only |visible )?primary action|primary[- ]action hierarchy)|This (?:issue|finding) has no current (?:gap|defect)|(?:This|The) (?:amendment|fix) keeps (?:all )?(?:[a-z]+|[1-9]\\d*) buttons identical)\\b`, 'i').test(value) ||
|
||||
/(?:^|[.!?;]\s+|\n)(?:Historical|Hypothetical|Quoted|Source|Archived|Example)(?:\s+(?:review|example|excerpt|assessment|material|text))?\s*:/i.test(value);
|
||||
const text = current(q.question);
|
||||
if (invalid(text)) return false;
|
||||
// These are the skill's existing decision fields, not a particular sentence
|
||||
// or line layout. Duplicate/missing fields cannot borrow a neighboring issue.
|
||||
const fields = ['Project/branch/task:', 'ELI10:', 'Stakes if we pick wrong:', 'Recommendation:', 'Completeness:', 'Net:'];
|
||||
const positions = fields.map(field => text.indexOf(field));
|
||||
if (positions.some((position, i) => position < 0 || text.lastIndexOf(fields[i]!) !== position ||
|
||||
(i > 0 && position <= positions[i - 1]!)) ||
|
||||
text.slice(0, positions[0]).trim() !== headline[0] ||
|
||||
(text.match(/\?/g)?.length ?? 0) !== 1 || !/\?\s*$/.test(text)) return false;
|
||||
const values = fields.map((field, i) => text.slice(positions[i]! + field.length, positions[i + 1] ?? text.length).trim());
|
||||
if (values.some(value => !value) || values.some(value => /^(?:If|When|Once|Unless|Assuming|Provided|Historical|Hypothetical|Quoted|Source|Example)\b/i.test(value)) ||
|
||||
!/\bDESIGN\.md\b/.test(values[3]!)) return false;
|
||||
const count = (value: string) => /^\d+$/.test(value) ? Number(value) :
|
||||
['zero', 'one', 'two', 'three', 'four', 'five', 'six', 'seven', 'eight', 'nine', 'ten'].indexOf(value.toLowerCase());
|
||||
const names = (value: string) => value.toLowerCase().split(/\s*[,/]\s*(?:and\s+)?|\s+and\s+/).map(s => s.trim()).sort();
|
||||
const same = (a: string[], b: string[]) => JSON.stringify(a) === JSON.stringify(b);
|
||||
const assessment = /^The header shows ([A-Za-z][A-Za-z0-9 ,/_-]{0,159}) as (two|three|four|five|six|seven|eight|nine|ten|[1-9]\d*) identical buttons\./i.exec(values[1]!);
|
||||
if (!assessment) return false;
|
||||
const actors = names(assessment[1]!);
|
||||
if (new Set(actors).size !== actors.length || actors.length !== count(assessment[2]!) || !actors.includes(control.toLowerCase())) return false;
|
||||
const peers = actors.filter(actor => actor !== control.toLowerCase());
|
||||
const ids = q.options.map(o => new RegExp(`^(${issue}[A-Z])[).:]\\s+`).exec(o.label)?.[1]);
|
||||
if (ids.some(id => !id) || new Set(ids).size !== 2 || !ids.some(id => values[3]!.startsWith(`${id} `))) return false;
|
||||
const offered = ids.map(id => [...values[4]!.matchAll(new RegExp(`(?:^|\\s)${id}[).:]\\s+`, 'g'))]);
|
||||
if (offered.some(matches => matches.length !== 1)) return false;
|
||||
return q.options.some((option, index) => {
|
||||
const body = current(option.description ?? ''), other = q.options[1 - index]!, declined = current(other.description ?? '');
|
||||
const style = /^([A-Za-z][A-Za-z0-9 _-]{0,39}): filled (#[0-9a-f]{6}) with (white|black) text\. ([A-Za-z][A-Za-z0-9 ,/_-]{0,159}): neutral ghost(?: buttons)?\./i.exec(body);
|
||||
const keep = new RegExp(`^${ids[1 - index]}[).:] Keep (two|three|four|five|six|seven|eight|nine|ten|[1-9]\\d*) equal buttons(?: \\(recommended\\))?$`, 'i').exec(other.label);
|
||||
if (!style || style[1]!.toLowerCase() !== control.toLowerCase() || !same(names(style[4]!), peers) ||
|
||||
!new RegExp(`^${ids[index]}[).:] Filled primary ${control}(?: \\(recommended\\))?$`, 'i').test(option.label) ||
|
||||
!keep || count(keep[1]!) !== actors.length || invalid(body) || invalid(declined) ||
|
||||
!/^No change\. Documented as a declined fix; Pass [1-7] stays below 10\./i.test(declined)) return false;
|
||||
// The detailed offered action must agree with its native menu's tokens and
|
||||
// actors; prose about another control cannot lend this choice a remedy.
|
||||
const start = offered[index]![0]!.index!, next = offered[1 - index]![0]!.index!;
|
||||
const action = values[4]!.slice(start, next > start ? next : undefined);
|
||||
const detail = new RegExp(`(?:^|[✅]\\s*)${control} becomes the only filled button \\((#[0-9a-f]{6}), (white|black) text\\); ([A-Za-z][A-Za-z0-9 ,/_-]{0,159}) become neutral ghost buttons`, 'i').exec(action);
|
||||
return !!detail && detail[1]!.toLowerCase() === style[2]!.toLowerCase() && detail[2]!.toLowerCase() === style[3]!.toLowerCase() && same(names(detail[3]!), peers);
|
||||
});
|
||||
}
|
||||
|
||||
/** A completed finding can start the passes when the caller already supplied the focus. */
|
||||
export function isDesignCountFirstReview(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed) return false;
|
||||
if (isDesignCountSetup(fp)) return false;
|
||||
if (numberedVisualHierarchyFinding(fp) || ordinaryDesignIssue(fp) || designSystemChoiceIssue(fp) || compactPrimaryDecision(fp)) return true;
|
||||
if (designFirstReviewAUQ(fp)) return true;
|
||||
return call.questions.some(q => {
|
||||
if (!call.answers?.[q.question] || q.options.length < 2) return false;
|
||||
if (/^(?:focus|scope|learnings|routing|next steps?|outside(?: design)? voices)$/i.test(q.header.trim())) return false;
|
||||
const id = /<gstack-qid:\s*([a-z0-9-]+)\s*>/i.exec(q.question)?.[1] ?? '';
|
||||
if (/(?:^|-)(?:focus|scope|setup|routing|learnings|onboarding|next-steps?|posture|mockups?|target)(?:-|$)/i.test(id)) return false;
|
||||
// Native fingerprints prepend the menu header. Inspect the actual question
|
||||
// for an explicit finding that offers a plan amendment and deferral.
|
||||
if (call.answered === true && call.failed === false && /^Pass\s*[1-7]\s*\([^)]*\)\s*[—–:]\s*Finding\s*[1-9]\d*:\s+\S/i.test(q.question.trim()) &&
|
||||
/^plan-design-review-[a-z0-9-]+$/i.test(id) &&
|
||||
(q.question.match(/<gstack-qid/gi)?.length ?? 0) === 1 &&
|
||||
/\b(?:Apply|Add|Fix|Specify|Define|Restore)\b[^?\n]*\b(?:to|in) the plan\?\s*<gstack-qid:[^>]+>\s*$/i.test(q.question) &&
|
||||
q.options.some(option => /^(?:Apply|Add|Fix|Specify|Define|Restore)\b/i.test(option.label)) &&
|
||||
q.options.some(option => /^(?:Defer|Leave|Keep as-is|Accept the gap)\b/i.test(option.label)) &&
|
||||
q.options.some(option => option.label === call.answers?.[q.question]) &&
|
||||
Array.isArray(call.unansweredQuestionIndices) && !call.unansweredQuestionIndices.length &&
|
||||
fp.signature === `${call.sessionId}:${call.toolUseId}`) return true;
|
||||
// A named or scored pass can ask for a missing design requirement before a
|
||||
// numbered finding heading appears. Its actual decision and opposed
|
||||
// choices establish review; a score or familiar qid alone cannot.
|
||||
const scoredPass = /^(?:D\s*\d+\s*[—–:-]\s*)?Pass\s*[1-7]\s*\([^)]*\)\s*[—–:-]\s*(?:10|[0-9])(?:\.[0-9]+)?\/10[.!:]/i.test(q.question.trim());
|
||||
const namedPass = /^(?:D\s*\d+\s*[—–:-]\s*)?Pass\s*[1-7]\s*[—–:-]\s*[A-Za-z][A-Za-z ]{3,60}:\s+/i.test(q.question.trim());
|
||||
const chosen = q.options.some(option => option.label === call.answers?.[q.question]);
|
||||
const fixChoice = q.options.some(option => /^(?:Add|Fix|Specify|Define|Restore)\b/i.test(option.label));
|
||||
const leaveChoice = q.options.some(option => /^(?:Leave as-is|Keep as-is|Defer|Accept the gap)\b/i.test(option.label) ||
|
||||
/^Skip\s*[—–-]\s*implied by\s+[^.!?]+\bgap$/i.test(option.label));
|
||||
if ((scoredPass || namedPass) && /^plan-design-review-[a-z0-9-]+$/i.test(id) &&
|
||||
(q.question.match(/<gstack-qid/gi)?.length ?? 0) === 1 &&
|
||||
/\b(?:gap|problem|defect|missing|inconsisten\w*)\b|\b(?:doesn['’]t|does not)\s+(?:record|specify|define|describe)\b/i.test(q.question) &&
|
||||
/\bShould I (?:add|fix|specify|define|restore)\b[^?]+\?\s*<gstack-qid:[^>]+>\s*$/i.test(q.question) &&
|
||||
fixChoice && leaveChoice && chosen &&
|
||||
!(call.unansweredQuestionIndices?.length) &&
|
||||
fp.signature === `${call.sessionId}:${call.toolUseId}`) return true;
|
||||
// Native pass decisions can carry a D-number before the pass title and
|
||||
// use plan-design-passN rather than plan-design-review-... identities.
|
||||
// Bind both forms to the same explicit pass and an offered choice that
|
||||
// leaves a named gap unresolved. Pass readiness is only setup.
|
||||
const numberedPass = /^D\s*\d+\s*[—–:-]\s*Pass\s*([1-7])\s*\([^)]*\)\s*:/i.exec(q.question.trim());
|
||||
const passId = /^plan-design-pass([1-7])-/i.exec(id);
|
||||
const unresolvedChoice = q.options.some(option =>
|
||||
/\b(?:leave|keep|defer|accept)\b/i.test(option.label) &&
|
||||
/\b(?:gap|problem|defect|inconsisten\w*)\b/i.test(`${option.label} ${option.description ?? ''}`));
|
||||
if (numberedPass && passId && numberedPass[1] === passId[1] && unresolvedChoice && /\?/.test(q.question)) return true;
|
||||
// These are issue-bearing pass statements in actual answered calls,
|
||||
// not a setup request that merely mentions the seven review passes.
|
||||
return /^Pass\s*[1-7]\s+(?:surfaces|(?:also\s+)?(?:found|flagged))\b/i.test(q.question.trim()) &&
|
||||
/\?/.test(q.question);
|
||||
});
|
||||
}
|
||||
|
||||
/** A closed recap may explain why Eng is next; it cannot request another fix. */
|
||||
function closedDesignGateRecap(tail: string, descriptions: string[]): boolean {
|
||||
const navigation = /\bWhat(?:['’]s)?\s+next\?\s*<gstack-qid:[a-z0-9-]+>\s*$/i.exec(tail);
|
||||
if (!navigation) return false;
|
||||
const body = tail.slice(0, navigation.index).trim();
|
||||
const gate = /^(?:Eng(?:ineering)? Review is (?:the )?required (?:shipping gate|gate before shipping))[.!]?$/i;
|
||||
const sentences = (text: string) => text.split(/[.!]\s+|[.!]$/).map(s => s.trim()).filter(Boolean);
|
||||
const recap = (text: string): boolean => {
|
||||
// Each count describes completed or explicitly absent work. A positive
|
||||
// deferred/open count is not a closed review, regardless of its title.
|
||||
const count = /^(?:(?:\d+|all)\s+(?:design\s+)?(?:decisions|findings|issues)\s+(?:(?:are|were)\s+)?(?:resolved|approved|addressed|closed)|\d+\s+(?:implementation\s+)?tasks\s+(?:(?:are|were)\s+)?(?:added|recorded|ready)|(?:no|zero|0)\s+(?:deferred(?:\s+(?:decisions|findings|issues|tasks|items))?|(?:unresolved|open|pending|outstanding)\s+(?:decisions|findings|issues|tasks|items)))$/i;
|
||||
if (text.split(/,\s*(?:and\s+)?|\s+and\s+/i).every(part => count.test(part))) return true;
|
||||
// Only a declarative completed-review subject can introduce explanatory
|
||||
// content. Separate clauses, questions and conditional/future work fail.
|
||||
if (!/^(?:The|This)\s+(?:design\s+)?review\s+(?:has\s+)?(?:added|recorded|approved|addressed|specified|covered|resolved)\s+\S/i.test(text)) return false;
|
||||
if (/[;?<>]|\b(?:if|unless|until|once|when|should|must|need|needs|will|would|could|please|then|also|still|missing|unresolved)\b|\b(?:and|but)\s+(?:first\s+)?(?:do|add|fix|repair|implement|resolve|decide|configure|remove|delete|pick|choose)\b/i.test(text)) return false;
|
||||
const clauses = text.split(/\s+[—–]\s+/);
|
||||
return clauses.length <= 2 && (clauses.length === 1 || /^(?:architectural|engineering|implementation)\s+(?:implications|considerations|details)\b/i.test(clauses[1]!));
|
||||
};
|
||||
const parts = sentences(body);
|
||||
if (parts.filter(part => gate.test(part)).length !== 1 ||
|
||||
!parts.every(part => gate.test(part) || recap(part))) return false;
|
||||
return descriptions.every(description => sentences(description).every(part =>
|
||||
gate.test(part) || recap(part) ||
|
||||
/^Exit plan mode and proceed on your own$/i.test(part) ||
|
||||
/^You have \d+ (?:concrete )?(?:implementation )?tasks ready to build from$/i.test(part)));
|
||||
}
|
||||
|
||||
/** A qidless closed handoff must consume every question/description clause. */
|
||||
function resolvedDesignHandoff(q: NonNullable<AskUserQuestionFingerprint['nativeCall']>['questions'][number]): number | null {
|
||||
if (!/^next review$/i.test(q.header.trim()) || q.options.length !== 2) return null;
|
||||
const completed = /^Design review complete [—–-] (?:10|[0-9](?:\.\d+)?)\/10 (?:→|->) (?:10|[0-9](?:\.\d+)?)\/10\. All ([1-9]\d*) decisions resolved\. The plan is design-complete; next is the required shipping gate\. What['’]s next\?$/.exec(q.question.trim());
|
||||
if (!completed) return null;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/\s*\(recommended\)\s*$/i, ''));
|
||||
const review = labels.findIndex(label => /^Run \/plan-eng-review$/i.test(label));
|
||||
const manual = labels.findIndex(label => /^Skip\s*[—–-]\s*I['’]ll handle next steps manually$/i.test(label));
|
||||
if (review < 0 || manual < 0 || review === manual) return null;
|
||||
const description = (index: number) => (q.options[index]!.description ?? '').trim().replace(/\s+/g, ' ');
|
||||
const topics = '(?:spinner|skeleton|(?:button|switch|field) (?:keyboard|focus|loading|error|disabled|pending|success)|(?:keyboard|focus|loading|error|disabled|pending|success) (?:states?|behavior|navigation))';
|
||||
const recap = new RegExp('^Eng review is the required shipping gate\\. It validates architecture, component wiring, tests, and accessibility implementation against the ' + completed[1] + ' approved design decisions\\. This design review added interaction specs \\(' + topics + '(?:, ' + topics + ')*\\), so eng review needs to validate their architectural fit\\.$');
|
||||
if (!recap.test(description(review)) ||
|
||||
!/^End the review workflow here\. The improved plan is at the (?:e2e output|approved plan) path; implementation can begin\. Run \/plan-eng-review later before shipping\.$/.test(description(manual))) return null;
|
||||
return manual + 1;
|
||||
}
|
||||
|
||||
function designHandoff(fp: AskUserQuestionFingerprint): { manualIndex: number | null } | null {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.failed || call.questions.length !== 1 ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}`) return null;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2) return null;
|
||||
const pending = call.answered === false && call.answers === undefined && call.answeredAt === undefined &&
|
||||
(call.unansweredQuestionIndices === undefined || (Array.isArray(call.unansweredQuestionIndices) &&
|
||||
call.unansweredQuestionIndices.length === 1 && call.unansweredQuestionIndices[0] === 0));
|
||||
const resolvedManual = call.failed === false && (call.answered === true || pending) ? resolvedDesignHandoff(q) : null;
|
||||
if (resolvedManual !== null) return { manualIndex: resolvedManual };
|
||||
if (!/^next\s+steps?$/i.test(q.header.trim())) return null;
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:\s*([a-z0-9-]+)\s*>/gi)].map(match => match[1]);
|
||||
if ((q.question.match(/<gstack-qid\b/gi) ?? []).length !== 1 || ids.length !== 1 || !/^plan-design-(?:review-)?next-steps?$/i.test(ids[0]!)) return null;
|
||||
const declaration = q.question.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '')
|
||||
.replace(/^next\s+steps?\s*:\s*/i, '');
|
||||
// Scores and a completed decision count describe a closed review. A
|
||||
// condition or unresolved gap cannot masquerade as its next-step menu.
|
||||
const completed = /^Design\s+review\s+(?:is\s+)?complete(?:[.!]|\s+\((?:\d+(?:\.\d+)?(?:\/10)?\s*(?:→|->|to)\s*)?\d+(?:\.\d+)?\/10(?:,\s*\d+\s+decisions?(?:\s+(?:made|added))?)?\)[.!])(?:\s|$)/i.exec(declaration);
|
||||
if (!completed) return null;
|
||||
const requiredGateOffer = /^The required next gate is Eng(?:ineering)? Review\s*[—–-]\s*want me to run it now\?\s*<gstack-qid:[a-z0-9-]+>\s*$/i.test(declaration.slice(completed[0].length).trim());
|
||||
const closedRecap = q.options.length === 2 && closedDesignGateRecap(
|
||||
declaration.slice(completed[0].length).trim(), q.options.map(option => option.description ?? ''));
|
||||
const requiredGateQuestion = requiredGateOffer || closedRecap || /^(?:\d+ implementation tasks ready\.\s*)?Eng(?:ineering)? Review is the required shipping gate\.\s*What next\?\s*<gstack-qid:[a-z0-9-]+>\s*$/i.test(declaration.slice(completed[0].length).trim());
|
||||
// The offered Eng action can carry the required-gate declaration while the
|
||||
// closed question asks only what is next. Its descriptions remain part of
|
||||
// the decision, so they cannot conceal a new repair or conditional closure.
|
||||
const describedRequiredGate = /^What['’]s\s+next\?\s*<gstack-qid:[a-z0-9-]+>\s*$/i.test(declaration.slice(completed[0].length).trim()) &&
|
||||
q.options.some(option => /^Run \/plan-eng-review(?:\s*\(recommended\))?$/i.test(option.label.trim()) &&
|
||||
/^Required gate before shipping[.!]/i.test(option.description ?? ''));
|
||||
const guardedNavigation = requiredGateQuestion || describedRequiredGate;
|
||||
if (!requiredGateQuestion && !/\bWhat['’]s\s+next\?\s*<gstack-qid:[a-z0-9-]+>\s*$/i.test(declaration)) return null;
|
||||
// A routing label cannot conceal a new repair in its description.
|
||||
if (guardedNavigation && q.options.some(option =>
|
||||
/(?:^|[.!?;]\s+|\b(?:proceed to|continue to|must|need to)\s+)(?:(?:please|first|then|also)\s+)*(?:add|fix|repair|implement|resolve|decide)\b|\b(?:(?:should|could|can|would)\s+(?:we|I)|(?:we|I)\s+(?:should|could|can|would))\s+(?:add|fix|repair|implement|resolve|decide)\b/i.test(option.description ?? '') ||
|
||||
/\b(?:Design|the|this)\s+review\s+(?:(?:is|remains)\s+)?(?:not\s+(?:complete|done|resolved)|incomplete|unfinished)\b|\bnot\s+all\s+(?:decisions|findings|issues|gaps)\s+(?:are\s+)?(?:resolved|complete|done)\b|\b(?:decisions|findings|issues|gaps)\s+(?:are\s+)?not\s+(?:resolved|complete|done)\b/i.test(option.description ?? '') ||
|
||||
/\b(?:once|after|when|if|unless|until)\b[^.!?]*\b(?:review|decisions?|findings?|issues?|gaps?)\b[^.!?]*\b(?:complete|done|resolved)\b|\b(?:review|decisions?|findings?|issues?|gaps?)\b[^.!?]*\b(?:complete|done|resolved)\b[^.!?]*\b(?:once|after|when|if|unless|until)\b/i.test(option.description ?? ''))) return null;
|
||||
// A closed heading does not override an affirmative outstanding-work claim
|
||||
// in its recap. Zero/no outstanding work is a compatible completion claim.
|
||||
const outstanding = (guardedNavigation ? [declaration, ...q.options.map(o => o.description ?? '')].join('\n') : declaration)
|
||||
.replace(/\b(?:no|zero|0)\s+(?:unresolved|open|pending|unaddressed|remaining|outstanding)\s+(?:[a-z-]+\s+){0,3}(?:gaps?|issues?|decisions?|requirements?|work)\b/gi, '')
|
||||
.replace(/\bno\s+(?:gaps?|issues?|decisions?|requirements?|work)\s+remains?\b/gi, '');
|
||||
if (/\b(?:unresolved|open|pending|unaddressed|remaining|outstanding)\s+(?:[a-z-]+\s+){0,3}(?:gaps?|issues?|decisions?|requirements?|work)\b|\b(?:gaps?|issues?|decisions?|requirements?|work)\s+(?:still\s+)?remains?\b|\b(?:gaps?|issues?|decisions?|requirements?|work)\s+(?:is|are)\s+still\s+(?:unresolved|open|pending|unaddressed)\b/i.test(outstanding)) return null;
|
||||
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).]\s*/i, '')
|
||||
.replace(/\s*\(recommended\)\s*$/i, '').trim());
|
||||
const manual = labels.map(label => /^(?:Handle next steps manually|Skip\s*[—–-]\s*I['’]ll handle next steps manually)$/i.test(label) ||
|
||||
(guardedNavigation && /^Skip\s*[—–-]\s*handle (?:next steps )?manually$/i.test(label)));
|
||||
const review = labels.map(label => /^Run \/plan-eng-review(?: next)?(?: \(required gate\))?$/i.test(label));
|
||||
const navigation = labels.map(label => /^(?:Skip to implementation|Run \/plan-ceo-review(?: first)?|Run \/design-(?:shotgun|html))$/i.test(label));
|
||||
if (manual.filter(Boolean).length > 1 || !review.some(Boolean) ||
|
||||
!labels.every((_, i) => manual[i] || review[i] || navigation[i])) return null;
|
||||
// Classification does not invent a missing stop option. Only an offered
|
||||
// manual action can steer a pending question away from another workflow.
|
||||
const index = manual.findIndex(Boolean);
|
||||
return { manualIndex: index < 0 ? null : index + 1 };
|
||||
}
|
||||
|
||||
/** Completed handoffs retain raw evidence and their own administrative count. */
|
||||
export function isDesignCompletionHandoff(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed || !Array.isArray(call.unansweredQuestionIndices) ||
|
||||
call.unansweredQuestionIndices.length || designHandoff(fp) === null) return false;
|
||||
const q = call.questions[0]!;
|
||||
return q.options.some(option => call.answers?.[q.question] === option.label);
|
||||
}
|
||||
|
||||
/** Preserve the native-only outside opt-out, then finish this review at its actual handoff. */
|
||||
export function pickDesignCountQuestion(
|
||||
routing: AskUserQuestionFingerprint,
|
||||
active: AskUserQuestionFingerprint,
|
||||
): number | null {
|
||||
const outside = pickDesignCountOutsideVoices(routing, active);
|
||||
if (outside !== null) return outside;
|
||||
return active.nativeCall?.answered ? null : designHandoff(active)?.manualIndex ?? null;
|
||||
}
|
||||
@@ -0,0 +1,682 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './plan-count-transcript';
|
||||
|
||||
/** Concrete product decisions, separate from the skill's mandatory Step-0 confirmations. */
|
||||
export const DEVEX_COUNT_FILES: Record<string, string> = {
|
||||
'README.md': `# EvalKit SDK
|
||||
|
||||
EvalKit is a Python SDK for ML engineers evaluating LLM responses. The primary
|
||||
developer writes Python daily, uses a terminal, and wants a local result before
|
||||
connecting the SDK to production CI. The agreed review posture is DX POLISH:
|
||||
improve the existing SDK's touchpoints within the beta release scope.
|
||||
|
||||
## Getting started
|
||||
|
||||
Install with \`python -m pip install evalkit==2.0.0b1\`,
|
||||
then follow the quickstart's command: \`python examples/first_eval.py\`.
|
||||
The published package inventory is in docs/package-contents.txt.
|
||||
|
||||
The chosen first-success experience is an included, copy-paste demo command:
|
||||
\`python -m evalkit.demo\`. It evaluates bundled sample responses and prints
|
||||
real per-example scores plus an overall score. It needs no hosted playground
|
||||
or new interactive UI. Like every first evaluation, it currently waits for the
|
||||
mandatory CI check described in docs/current-contracts.md.
|
||||
|
||||
The bundled demo already works without a developer API key. Its sample evaluation
|
||||
uses the shipped mock transport; its mandatory remote CI check uses the included
|
||||
sample-project binding. No credentials step precedes this first demo result.
|
||||
The keyless demo still waits for that CI check and has no skip or offline bypass.
|
||||
|
||||
After the demo, developers obtain a key for their first live evaluation at
|
||||
https://console.evalkit.example/settings/api-keys: select the project, choose
|
||||
Create key, copy the value once, and export EVALKIT_API_KEY in their terminal.
|
||||
The page also lists existing keys and provides revoke/rotate controls. The
|
||||
bundled demo does not use this key; live evaluations do.
|
||||
|
||||
Expected completed demo output for the bundled sample responses is documented
|
||||
here; the shipped demo prints this per-example and aggregate score format:
|
||||
|
||||
example 1: score=0.80
|
||||
example 2: score=1.00
|
||||
overall: score=0.90
|
||||
|
||||
See docs/api.md for public API and upgrade behavior, and docs/benchmarks.md for
|
||||
the completed onboarding study. These documents describe the existing SDK's
|
||||
behavior; its runtime is maintained separately from this release-planning repo.
|
||||
`,
|
||||
'docs/benchmarks.md': `# Completed onboarding study
|
||||
|
||||
The internal comparison measured Python SDK onboarding with the same developer
|
||||
and machine. Peer SDK A took 2 minutes, B took 4 minutes, and C took 3 minutes.
|
||||
EvalKit took 6 minutes, including the mandatory 5-minute CI wait. The measurement
|
||||
starts before installation and ends at the first real evaluation result.
|
||||
|
||||
The agreed target is under 2 minutes. The study, target persona, and terminal
|
||||
demo delivery vehicle are already approved. Timing instrumentation and the
|
||||
post-beta feedback survey exist and will continue unchanged.
|
||||
`,
|
||||
'docs/current-contracts.md': `# Existing SDK contracts
|
||||
|
||||
On a developer's first local evaluation, the SDK requires a successful remote
|
||||
CI check and blocks for five minutes before returning an evaluation result.
|
||||
There is no skip flag or offline first-run path. The beta plan retains this gate.
|
||||
|
||||
During the required wait, the existing SDK writes a progress line to stderr
|
||||
every 30 seconds, such as "Waiting for CI check: 90s elapsed of 300s", and reports
|
||||
when the check finishes. Progress does not bypass the check or return evaluation
|
||||
results before its required successful completion.
|
||||
|
||||
Before the countdown, the SDK already prints what the check verifies and where
|
||||
to inspect it: "Verifying the sample-project binding with EvalKit CI; inspect
|
||||
https://ci.evalkit.example/checks/<check-id>; normally completes within 300s."
|
||||
The URL identifies the check without exposing credentials. If it has not
|
||||
succeeded at 300s, the SDK reports EVALKIT_CI_TIMEOUT, the check URL, and the
|
||||
instruction to inspect that check and retry after CI recovers. Its help link
|
||||
explains the check states and recovery steps. Success is still required before
|
||||
the first local result; these messages do not change the mandatory wait.
|
||||
|
||||
Authentication errors behave exactly as documented in docs/api.md. All other
|
||||
errors already identify the cause, relevant argument or file, and an actionable
|
||||
fix. Errors redact secrets. API timeouts, cancellation, rate limits, and retries
|
||||
are bounded and documented; evaluation IDs prevent duplicate submitted jobs.
|
||||
|
||||
The SDK supports Python 3.10+, macOS, Linux, and Windows without Docker. Its
|
||||
type annotations, offline sample data, mock transport, noninteractive CI mode,
|
||||
API reference, support contact, changelog, and contributor guide already work.
|
||||
Telemetry is opt-in. No new hosted service, language binding, or community
|
||||
program is proposed in this release.
|
||||
`,
|
||||
'docs/api.md': `# Public API retained by the beta plan
|
||||
|
||||
The two evaluation functions accept positional arguments:
|
||||
|
||||
- \`run_eval(dataset, evaluator)\`
|
||||
- \`run_batch(evaluator, dataset)\`
|
||||
|
||||
Both argument names describe the same concepts. The reversed positional order
|
||||
is intentional in the current draft; neither function requires keyword arguments.
|
||||
|
||||
For an invalid API key, the SDK raises \`AuthError("request failed")\`.
|
||||
There is no authentication error code, explanation of the cause, or instruction
|
||||
for replacing the key. The plan retains this message.
|
||||
|
||||
Version 1 exposes \`Client.evaluate()\`. Version 2 replaces it with
|
||||
\`Client.run()\` and removes the old name immediately. No compatibility alias,
|
||||
deprecation warning, migration guide, or codemod is provided. Other public APIs
|
||||
retain their existing behavior, and the release changelog is otherwise complete.
|
||||
`,
|
||||
'docs/package-contents.txt': `Published evalkit 2.0.0b1 package inventory:
|
||||
evalkit/__init__.py
|
||||
evalkit/client.py
|
||||
evalkit/demo.py
|
||||
evalkit/sample_responses.json
|
||||
README.md
|
||||
|
||||
The quickstart references examples/first_eval.py, but that file is absent from
|
||||
both the published package and the release examples archive. The terminal demo
|
||||
module and its sample data are included and work as documented.
|
||||
`,
|
||||
};
|
||||
|
||||
export function planDevexCountFixture(planPath: string): string {
|
||||
return [
|
||||
`Please review this plan thoroughly. As you go, write your plan-mode plan to ${planPath} (use Edit/Write to that exact path).`,
|
||||
'This is an interactive review; a responder is available for AskUserQuestion.',
|
||||
'Present each unresolved decision and wait for its answer.',
|
||||
'',
|
||||
'# Plan: EvalKit SDK beta release polish',
|
||||
'',
|
||||
'The primary developer, onboarding benchmark, and terminal demo experience',
|
||||
'are settled in README.md and docs/benchmarks.md. Use DX POLISH for the',
|
||||
'existing release scope. Review the actual documented contracts and proposed',
|
||||
'behavior, including the first-run CI requirement, public function signatures,',
|
||||
'authentication error, packaged quickstart, and v1-to-v2 client upgrade.',
|
||||
'',
|
||||
'The current draft ships the behavior in docs/current-contracts.md and',
|
||||
'docs/api.md unchanged, using the package inventory in docs/package-contents.txt.',
|
||||
'Recommendations that repair those developer-facing contracts belong in this',
|
||||
'plan. Existing working contracts remain the baseline for the review.',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
type QuestionRecord = { header: string; question: string; options?: Array<{ label: string; description?: string }> };
|
||||
|
||||
function questionRecords(fp: AskUserQuestionFingerprint, answeredOnly = false): QuestionRecord[] {
|
||||
if (!fp.nativeCall) return [{ header: '', question: fp.promptSnippet }];
|
||||
return fp.nativeCall.questions.filter(q => !answeredOnly
|
||||
|| (fp.nativeCall!.answered && Boolean(fp.nativeCall!.answers?.[q.question])));
|
||||
}
|
||||
|
||||
const ADMINISTRATIVE_HEADERS = new Set([
|
||||
'design doc', 'prerequisite', 'routing rules', 'routing setup', 'cross-project',
|
||||
'target persona', 'developer persona', 'persona selection', 'empathy check',
|
||||
'narrative check', 'tthw target', 'competitive benchmark', 'benchmark confirmation',
|
||||
'magic delivery', 'review mode', 'fix scope', 'confusion scope',
|
||||
]);
|
||||
|
||||
/** The structured accuracy frame approves an observation, never a proposed repair. */
|
||||
function structuredEmpathyAccuracy(header: string, question: string, options: QuestionRecord['options']): boolean {
|
||||
if (!/^Empathy$/i.test(header.trim()) || !options || options.length !== 3 || /<gstack-qid/i.test(question)) return false;
|
||||
const compact = (text: string) => text.trim().replace(/\s+/g, ' ');
|
||||
const clean = (text: string) => compact(text).replace(/\s*\(recommended\)$/i, '');
|
||||
// Consume complete descriptions too: an accurate recap cannot conceal an
|
||||
// additional approval in the explanation of an option.
|
||||
const descriptions = new Map([
|
||||
['accurate, proceed', /^✅ Every beat is grounded in a documented contract, not a guess about the runtime\. ✅ Lets the review move to friction-point decisions immediately\. ❌ If the runtime differs from the docs, the scores inherit that gap\.$/i],
|
||||
['some of this is wrong', /^✅ You correct specific beats \(for example, the demo may not need an API key\) before scoring\. ✅ Keeps the narrative honest for the implementer who reads it\. ❌ Costs one round-trip before friction-point questions begin\.$/i],
|
||||
['way off, actual experience is...', /^✅ Replaces the narrative entirely with your account of the real first run\. ✅ Prevents a review built on a wrong premise\. ❌ Discards the traced path and requires you to describe the flow from scratch\.$/i],
|
||||
]);
|
||||
const labels = options.map(option => clean(option.label).toLowerCase());
|
||||
if (new Set(labels).size !== 3 || options.some((option, i) => !option.description ||
|
||||
!descriptions.get(labels[i]!)?.test(compact(option.description)))) return false;
|
||||
const parts = question.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '').split(/\n\s*\n/);
|
||||
if (parts.length !== 3) return false;
|
||||
const role = String.raw`(?:(?:ML|backend|frontend|full-stack) )?(?:developer|engineer)`;
|
||||
const preamble = new RegExp(String.raw`^Does this first-person narrative match what your ${role} experiences today\? Project/branch/task: [\w-]+ on [\w/-]+, [\w.-]+ SDK beta polish\. ELI10: Before scoring anything, I walk the actual README path as the target developer and describe what they see and feel\. If I have the experience wrong, every score downstream is wrong too, so please correct me here\. Stakes: this narrative becomes the Developer Perspective section the implementer reads\.$`, 'i');
|
||||
if (!preamble.test(compact(parts[0]!)) ||
|
||||
!/^Stakes if we pick wrong: the review polishes the wrong pain\. Recommendation: A because every step above traces to a specific line in README\.md, docs\/api\.md, docs\/current-contracts\.md, or docs\/package-contents\.txt\. Note: options differ in kind, not coverage [—–-] no completeness score\. Net: proceed on the traced path vs\. correct it before scoring\.$/i.test(compact(parts[2]!))) return false;
|
||||
const journey = parts[1]!.split('\n');
|
||||
if (!new RegExp(String.raw`^NARRATIVE \(${role}, terminal, wants a local result before CI\):$`, 'i').test(journey.shift() ?? '')) return false;
|
||||
// Quoted commands/messages are source evidence. Every unquoted sentence
|
||||
// must consume one known observation form; a heading alone cannot turn
|
||||
// arbitrary instructions, deontic clauses or imperatives into evidence.
|
||||
const sentences = compact(journey.join(' ')).replace(/`[^`]*`|"(?:[^"\\]|\\.)*"|“[^”]*”/g, '[source]').split(/(?<=[.!?])\s+/);
|
||||
const observations = [
|
||||
/^I open the README\.$/i,
|
||||
/^Heading one is \[source\], and the first paragraph describes me exactly, so I keep reading\.$/i,
|
||||
/^Under \[source\] I copy \[source\], export [A-Z][A-Z_]+, and run \[source\] as instructed\.$/,
|
||||
/^Python says \[source\]\.$/,
|
||||
/^I check site-packages: \w+ has \w+\.py, \w+\.py, \w+\.json, no examples folder\.$/i,
|
||||
/^(?:\d+|Thirty) seconds lost, some trust lost\.$/i,
|
||||
/^The next paragraph mentions \[source\], so I try that\.$/i,
|
||||
/^It starts, then stderr prints \[source\]\.$/i,
|
||||
/^I wanted a local score on bundled sample data; instead I['’]m waiting (?:\d+|five) minutes on a remote check I never configured, at \d+-second updates, with no flag to skip it\.$/i,
|
||||
/^Peer SDK [A-Z] gave me a number in (?:\d+|two) minutes total\.$/i,
|
||||
/^I alt-tab\.$/i,
|
||||
/^Later the scores appear: \d+(?:\.\d+)?, \d+(?:\.\d+)?, \d+(?:\.\d+)?\.$/i,
|
||||
/^Fine\.$/i,
|
||||
/^I write my own call: \[source\]\.$/i,
|
||||
/^Then I try \[source\] and it fails, because run_batch takes \(evaluator, dataset\)\.$/i,
|
||||
/^I paste a typo['’]d key and get \[source\]: no code, no hint that the key is the problem\.$/i,
|
||||
/^On my existing v\d+ code, \[source\] is now simply gone with no warning or migration note\.$/i,
|
||||
];
|
||||
return sentences.length > 0 && sentences.every(sentence => observations.some(pattern => pattern.test(sentence)));
|
||||
}
|
||||
|
||||
/** Confirming a quoted developer journey authorizes understanding, not its repairs. */
|
||||
function empathyAccuracyConfirmation(header: string, question: string, options: QuestionRecord['options']): boolean {
|
||||
if (!/^(?:Empathy(?: narrative| trace)?|Narrative)$/i.test(header.trim()) ||
|
||||
!options || options.length < 2 || options.length > 4) return false;
|
||||
const clean = (value: string) => value.trim().replace(/\s*\(recommended\)\s*$/i, '').trim();
|
||||
const confirm = (label: string) => /^(?:Accurate|Yes\s*[—–-]\s*accurate)\s*[—–-]\s*proceed(?: with this understanding)?$/i.test(clean(label));
|
||||
const correct = (label: string) => /^(?:Part(?:ly|ially) wrong\s*[—–-]\s*let me correct it|Mostly right\s*[—–-]\s*minor corrections|Wrong path\s*[—–-]\s*the actual flow is different|Wrong\s*[—–-]\s*actual experience differs|The experience is different\s*[—–-]\s*let me describe it)$/i.test(clean(label));
|
||||
const labels = options.map(option => clean(option.label));
|
||||
if (new Set(labels).size !== labels.length || labels.filter(confirm).length !== 1 ||
|
||||
!labels.some(correct) || !labels.every(label => confirm(label) || correct(label))) return false;
|
||||
// Consume each description completely: an accuracy label must not also
|
||||
// approve a remedy hidden in a subsequent sentence or clause.
|
||||
const description = /^(?:(?:The (?:narrative|trace) is (?:correct|accurate)\.[ ]*)?Proceed with this understanding(?: for the full DX review)?\.|Some details are off; I['’]ll clarify (?:before we continue|the actual experience)\.|This matches the actual developer experience; use it as the basis for the review\.|The (?:real|actual) (?:getting-started path|flow|experience) differs(?: significantly)? from what was traced\.)$/i;
|
||||
if (options.some(option => option.description && !description.test(clean(option.description)))) return false;
|
||||
const ids = question.match(/<gstack-qid:[^>]+>/gi) ?? [];
|
||||
if (ids.length > 1 || (question.match(/<gstack-qid/gi)?.length ?? 0) !== ids.length) return false;
|
||||
const text = question.replace(/\s*<gstack-qid:[^>]+>\s*$/i, '').trim()
|
||||
.replace(/^D\s*\d+\s*[—–:-]\s*/i, '');
|
||||
const paragraphs = text.split(/\n\s*\n/);
|
||||
const opening = paragraphs.shift() ?? '';
|
||||
const closing = paragraphs.pop() ?? '';
|
||||
if (!/^(?:Empathy (?:narrative|trace): does this match (?:(?:the [\w.-]+ (?:getting-started|onboarding|first-run) )?reality|your actual developer experience)\?|Does (?:this|the) (?:empathy narrative|first-person developer trace) match reality\?)$/i.test(opening) ||
|
||||
!/^Does this match (?:reality|the actual experience)\?(?: Where am I wrong\?)?$/i.test(closing)) return false;
|
||||
// Only quoted journey evidence and an observational preface may intervene.
|
||||
// Additional questions or instructions outside the quote remain decisions.
|
||||
const source = String.raw`(?:the docs|[\w-]+(?:[/.][\w-]+)+)`;
|
||||
const role = String.raw`(?:(?:Python|JavaScript|TypeScript|Go|Rust|Java|Ruby) )?(?:(?:ML|backend|frontend|full-stack) )?(?:developer|engineer)`;
|
||||
// A first-person journey may be delimited with horizontal rules instead
|
||||
// of blockquotes. Keep its observation preface and both boundaries exact;
|
||||
// an obligation outside that evidence is still a substantive decision.
|
||||
const narrated = new RegExp(String.raw`^Here['’]s what I think a ${role} experiences today with [\w.-]+:$`, 'i');
|
||||
if (narrated.test(paragraphs[0] ?? '')) {
|
||||
const journey = paragraphs.slice(2, -1);
|
||||
const observed = /^(?:I (?:find|found|open|read|run|try|install|look|wait|see|notice|receive|got|get|check|search|browse|start|follow)\b|After (?:scanning|reading|checking|searching|browsing)\b[^.!?\n]*\bI (?:find|spot|see|notice)\b)/i;
|
||||
const decision = /\b(?:approv\w*|recommend\w*|suggest\w*|propos\w*|authoriz\w*|consent\w*|decid\w*|request\w*)\b|\b(?:should|could|can|may|must|shall|would) (?:we|you|I)\b|\b(?:we|you|I) (?:should|could|must|shall|will|would|need to|want to)\b|\blet['’]s\b|(?:^|[.!?;:]\s+|\b(?:please|also|then|and)\s+)(?:add|fix|package|remove|change|implement|enable|disable|repair|rewrite|apply|replace)\b/i;
|
||||
// Every unquoted sentence must still describe an observation. Delimiters
|
||||
// cannot turn a new imperative (including an unknown action verb) into
|
||||
// quoted evidence. Explicit requests and obligations fail independently
|
||||
// of which action they name.
|
||||
const obligation = /\b(?:please|must|should|shall|ought|need(?:s)? to|ha(?:ve|s) to|required to)\b/i;
|
||||
const sentences = journey.flatMap(part => part
|
||||
.replace(/`[^`]*`|"(?:[^"\\]|\\.)*"|“[^”]*”/g, quote =>
|
||||
'[source]' + (/[.!?]["”]$/.test(quote) ? quote.at(-2) : ''))
|
||||
.split(/(?<=[.!?;])\s+/));
|
||||
const observation = /^(?:(?:(?:Fine,|But)\s+)?I (?:find|found|open|read|run|try|install|look|wait|see|notice|receive|got|get|check|search|browse|start|follow|go|sit|lost|burned|don['’]t know)\b|After (?:scanning|reading|checking|searching|browsing)\b[^.!?\n]*\bI (?:find|spot|see|notice)\b|(?:The )?README (?:then says:|pointed me at)\s|First thing I see: install with \[source\]\.?$|Then: (?:set )?\[source\]\.?$|It starts [—–-] nothing happens\.?$|[\w]+ (?:seconds?|minutes?) (?:later: \[source\]|pass)\.?$|Wait, what\?$|A local demo needs a CI check\?$|Is something broken\?$|\[source\]\.?$)/i;
|
||||
return paragraphs.length >= 4 && paragraphs[1] === '---' && paragraphs.at(-1) === '---' &&
|
||||
journey.every(part => observed.test(part) && !decision.test(part) && !obligation.test(part)) &&
|
||||
sentences.every(sentence => observation.test(sentence));
|
||||
}
|
||||
const goal = String.raw`(?: who just heard about [\w.-]+ and wants to verify it works locally before integrating it into their team['’]s CI pipeline)?`;
|
||||
const preface = new RegExp(String.raw`^(?:Here['’]s what I (?:traced|observed) from ${source}(?:, ${source})*(?: and ${source})?\.\s*)?(?:The persona: ${role}${goal}\.)?$`, 'i');
|
||||
let quoted = false;
|
||||
for (const paragraph of paragraphs) {
|
||||
if (paragraph.split('\n').every(line => /^\s*>/.test(line))) { quoted = true; continue; }
|
||||
if (quoted || !preface.test(paragraph)) return false;
|
||||
}
|
||||
return quoted;
|
||||
}
|
||||
|
||||
function administrativeQuestion(header: string, question: string, options: QuestionRecord['options']): boolean {
|
||||
// These decisions establish the review's evidence and scope. Mentioning a
|
||||
// defect in their recap does not turn a confirmation into a finding.
|
||||
if (ADMINISTRATIVE_HEADERS.has(header.toLowerCase().replace(/\s+/g, ' ').trim())) return true;
|
||||
if (empathyAccuracyConfirmation(header, question, options)) return true;
|
||||
if (structuredEmpathyAccuracy(header, question, options)) return true;
|
||||
if (/^empathy(?:\s*\(0B\))?$/i.test(header.trim()) &&
|
||||
/^Does (?:this|the) empathy narrative match\b/i.test(question.replace(/^D\s*\d+\s*[—–:-]\s*/i, ''))) {
|
||||
const labels = options?.map(option => option.label.trim().replace(/\s*\(recommended\)\s*$/i, '')) ?? [];
|
||||
const confirm = (label: string) => /^Yes\s*[—–-]\s*accurate, proceed with this understanding$/i.test(label);
|
||||
const correct = (label: string) => /^The experience is different\s*[—–-]\s*let me describe it$/i.test(label) ||
|
||||
(/^Partially\s*[—–-]\s*(?:the [^;.!?]+? (?:does|is|has)|it (?:does|is|has)|there (?:is|are))\s+[^;.!?]+$/i.test(label) &&
|
||||
!/\b(?:should|must|needs?|shall|will|would|could)\b|(?:[,::]|\b(?:and|then)\b)\s*(?:add|fix|package|remove|change|implement|enable|disable)\b/i.test(label));
|
||||
if (labels.filter(confirm).length === 1 && labels.some(correct) && labels.every(label => confirm(label) || correct(label))) return true;
|
||||
}
|
||||
const narrativeHeader = header.trim().replace(/^D\s*\d+\s*(?:[—–:-]\s*)?/i, '');
|
||||
const narrativeQuestion = question.replace(/^D\s*\d+\s*[—–:-]\s*/i, '');
|
||||
if (/^Narrative$/i.test(narrativeHeader) &&
|
||||
/^Does (?:this|the) first-person developer trace match reality\?/i.test(narrativeQuestion) &&
|
||||
!/<gstack-qid/i.test(question)) {
|
||||
const labels = options?.map(option => option.label.trim().replace(/\s*\(recommended\)\s*$/i, '')) ?? [];
|
||||
const confirm = (label: string) => /^Accurate\s*[—–-]\s*proceed$/i.test(label);
|
||||
const correct = (label: string) => /^(?:Mostly right\s*[—–-]\s*minor corrections|Wrong\s*[—–-]\s*actual experience differs)$/i.test(label);
|
||||
const repair = /(?:^|[.!?]\s+|\b(?:and|then|also|please|must|should|will|need to|proceed to|continue to)\s+)(?:add|fix|package|remove|change|implement|enable|disable|repair|rewrite)\b/i;
|
||||
// The captured trace has only its opening and closing accuracy questions.
|
||||
// An additional question asks for another decision, even with accuracy labels.
|
||||
const confirmationOnly = /^Does (?:this|the) first-person developer trace match reality\?[^?]*Does this match the actual experience\?\s*$/i.test(narrativeQuestion);
|
||||
if (labels.filter(confirm).length === 1 && labels.some(correct) &&
|
||||
new Set(labels).size === labels.length && labels.every(label => confirm(label) || correct(label)) &&
|
||||
confirmationOnly && !repair.test(narrativeQuestion) &&
|
||||
options!.every(option => !repair.test(option.description ?? ''))) return true;
|
||||
}
|
||||
const id = [...question.matchAll(/<gstack-qid:([^>]+)>/gi)].at(-1)?.[1];
|
||||
if (id && /^(?:routing-injection|cross-project-learnings|plan-devex-review-(?:office-hours-preflight|prereq|persona|empathy(?:-check|-narrative)?|tthw-tier|competitive-tier|benchmark-tier|magical-moment|mode|confusion-report))$/i.test(id)) return true;
|
||||
return /how deep should this dx review|which (?:dx )?review mode|\b(?:can|shall|should) we (?:continue|proceed|begin)(?: (?:the )?(?:setup|review)| now)?\?\s*$/i.test(question);
|
||||
}
|
||||
|
||||
/** The answered native call proves a decision; its content must identify a concrete problem. */
|
||||
function substantiveIssue({ header, question, options }: QuestionRecord): boolean {
|
||||
if (administrativeQuestion(header, question, options)) return false;
|
||||
const normalized = `${header} ${question}`.replace(/\s+/g, ' ');
|
||||
const ciGate = /\b(?:CI|continuous integration)\b/i.test(normalized)
|
||||
&& /\b(?:first[- ](?:local[- ])?runs?|first eval(?:uation)?|local eval(?:uation)?|hello world)\b/i.test(normalized)
|
||||
&& /\b(?:mandatory|required|blocks?|five[- ]minute|5[- ]min(?:ute)?|wait|gate)\b/i.test(normalized);
|
||||
const argumentsReversed = /\brun_eval\b/i.test(normalized) && /\brun_batch\b/i.test(normalized)
|
||||
&& /\b(?:revers\w*|inconsisten\w*|swapp\w*|different|order|positional)\b/i.test(normalized);
|
||||
const opaqueAuth = /\b(?:AuthError|API[- ]?key|authentication|invalid key)\b/i.test(normalized)
|
||||
&& /request failed|\b(?:opaque|generic|unactionable|cryptic)\b|no (?:cause|guidance|fix|explanation|instruction)|doesn.t (?:explain|guide)/i.test(normalized);
|
||||
const missingExample = /examples\/first_eval\.py|\b(?:packaged|quickstart|quick-start) example\b/i.test(normalized)
|
||||
&& /\b(?:missing|absent|omitted|FileNotFoundError)\b|not (?:included|packaged|shipped)|doesn.t (?:exist|ship)/i.test(normalized);
|
||||
const breakingRename = /Client\.evaluate|Client\.run|\bmethod rename\b/i.test(normalized)
|
||||
&& /\b(?:breaking|remov\w*|renam\w*)\b/i.test(normalized)
|
||||
&& /\b(?:migration|deprecation|compatibility|alias|codemod)\b/i.test(normalized);
|
||||
// Expected-output documentation is separate from whether its command
|
||||
// exists. Count the actual gap plus offered documentation remedy, not a
|
||||
// generic navigation question that merely names output in its options.
|
||||
const outputSubject = String.raw`(?:(?:expected|sample|example)(?: demo)?|demo) output`;
|
||||
// Consume the complete noun phrase, including a negating determiner,
|
||||
// before judging its absence. A nested "demo output" suffix cannot
|
||||
// escape "no sample demo output is missing" and become a finding.
|
||||
const missingState = [...normalized.matchAll(new RegExp(String.raw`\b(?:(no|not any)\s+)?${outputSubject}\s+(?:(?:is|are|was|were)\s+)?(?:missing|absent|omitted|unspecified)\b`, 'gi'))];
|
||||
const missingSubject = [...normalized.matchAll(new RegExp(String.raw`\b(?:(no|not any)\s+)?missing\s+${outputSubject}\b`, 'gi'))];
|
||||
const noOutput = new RegExp(String.raw`\bno\s+${outputSubject}\s*(?:[,.;!?]|\b(?:in|from|for|yet)\b)`, 'i');
|
||||
const outputGap = missingState.some(match => !match[1]) || missingSubject.some(match => !match[1]) || noOutput.test(normalized);
|
||||
const missingOutput = /\b(?:README|quick[- ]?start|documentation)\b/i.test(normalized)
|
||||
&& (outputGap || /\b(?:README|quick[- ]?start|documentation)\b[^.!?;]{0,50}\b(?:doesn['’]t|does not)\s+(?:show|include)\b[^.!?;]{0,25}\boutput\b/i.test(normalized))
|
||||
&& Boolean(options?.some(option => /^(?:[A-Z][.:)]\s*)?Add\s+(?:to\s+(?:the\s+)?plan:\s*include\s+)?(?:an?\s+)?(?:expected|sample|example)(?:\s+demo)?\s+output\b[^.!?]*\b(?:README|quick[- ]?start|documentation)\b/i.test(option.label)));
|
||||
return ciGate || argumentsReversed || opaqueAuth || missingExample || breakingRename || missingOutput;
|
||||
}
|
||||
|
||||
/** A setup heading cannot hide a positively selected repair to the existing behavior. */
|
||||
function answeredSetupRepair(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call || call.failed || !call.answered || call.questions.length !== 1 ||
|
||||
call.unansweredQuestionIndices?.length || fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const question = call.questions[0]!;
|
||||
if (question.multiSelect) return false;
|
||||
const selected = question.options.filter(option => option.label === call.answers?.[question.question]);
|
||||
if (selected.length !== 1) return false;
|
||||
const ids = [...question.question.matchAll(/<gstack-qid:([a-z0-9-]+)>/gi)];
|
||||
if (ids.length !== 1 || (question.question.match(/<gstack-qid/gi)?.length ?? 0) !== 1) return false;
|
||||
const id = ids[0]![1]!.toLowerCase();
|
||||
const header = question.header.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '');
|
||||
const text = question.question.replace(/\s+/g, ' ');
|
||||
const label = selected[0]!.label.replace(/^[A-Z][.):]\s*/i, '');
|
||||
if (id === 'plan-devex-review-tthw-tier' && /^TTHW target$/i.test(header)) {
|
||||
return /TTHW|Time-to-Hello-World/i.test(text) && /\bCI\b/i.test(text) &&
|
||||
/\b(?:mandatory|blocks?|retains? the CI block)\b/i.test(text) &&
|
||||
/(?:^|[—–:]\s*)add\s+(?:an?\s+)?(?:skip flag|--skip-ci|offline(?:[- ]first[- ]run)? path)\b/i.test(label);
|
||||
}
|
||||
if (id === 'plan-devex-review-tthw-ci-block' && /^TTHW target$/i.test(header)) {
|
||||
// A confirmed benchmark does not approve a new CI bypass. This captured
|
||||
// menu asserts the broken target and selects an explicit repair.
|
||||
const headline = question.question.split('\n')[0]!.replace(/<gstack-qid:[^>]+>/i, '').trim();
|
||||
return Array.isArray(call.unansweredQuestionIndices) && call.unansweredQuestionIndices.length === 0 &&
|
||||
/^D\s*\d+\s*[—–:-]\s*Journey Stage HELLO WORLD:\s*The \d+[- ]minute mandatory CI block makes the under-\d+[- ]minute TTHW target unreachable\.\s*$/i.test(headline) &&
|
||||
/^Add (?:a )?demo-mode CI skip flag(?:\s*\(Recommended\))?$/i.test(label);
|
||||
}
|
||||
if (id === 'plan-devex-review-magical-moment' && /^Magical moment$/i.test(header)) {
|
||||
// The selected option adds progress feedback beyond the already chosen
|
||||
// demo vehicle and prior CI-bypass decision. An unselected remedy or
|
||||
// a confirmation of that vehicle alone remains setup.
|
||||
return /\bdemo\b/i.test(text) && /\bsilently blocks?\b|\bsilent (?:CI )?wait\b/i.test(text) &&
|
||||
/(?:^|[—–:]\s*)add\s+[^.!?;]{0,80}\bprogress (?:output|indicator)\b/i.test(label);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** A current first-pass repair can name the broken contract without its file path. */
|
||||
function answeredContractRepair(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length || fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2 || new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
q.options.filter(o => o.label === call.answers?.[q.question]).length !== 1 ||
|
||||
administrativeQuestion(q.header, q.question, q.options)) return false;
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:([^>]+)>/gi)];
|
||||
if (ids.length !== 1 || (q.question.match(/<gstack-qid/gi)?.length ?? 0) !== 1) return false;
|
||||
if (/^(?:plan-)?devex-(?:review-)?[a-z0-9-]+$/i.test(ids[0]![1]!) &&
|
||||
!/(?:^|-)(?:mode|setup|scope|routing|prerequisite|next-steps?)(?:-|$)/i.test(ids[0]![1]!) &&
|
||||
/^TTHW block$/i.test(q.header.trim())) {
|
||||
// The retained CI wait contradicts an agreed target; this is an accepted
|
||||
// repair decision, not selection or confirmation of the target itself.
|
||||
if (call.answered !== true || call.failed !== false ||
|
||||
fp.options.length !== q.options.length || !fp.options.every((o, i) =>
|
||||
o.index === i + 1 && o.label === q.options[i]!.label)) return false;
|
||||
const body = q.question.replace(/\s*<gstack-qid:[^>]+>\s*$/i, '').trim();
|
||||
const timing = /^D\s*\d+\s*[—–:-]\s*Pass 1 \(Getting Started\): The agreed <(\d+(?:\.\d+)?) min TTHW target is mathematically impossible with the retained (\d+(?:\.\d+)?)[- ](?:min|minute) CI block\. Which resolution belongs in the plan\?$/i.exec(body);
|
||||
if (!timing) return false;
|
||||
const [target, wait] = timing.slice(1).map(Number);
|
||||
const selected = call.answers![q.question]!.replace(/\s*\(Recommended\)\s*$/i, '').trim();
|
||||
return [target, wait].every(n => Number.isFinite(n) && n! > 0) && wait! >= target! &&
|
||||
/^(?:Demo-only CI bypass|Add --offline flag to [a-z_$][\w$.-]*|Update TTHW target to reflect reality)$/i.test(selected);
|
||||
}
|
||||
if (ids[0]![1] === 'devex-demo-ci-bypass') {
|
||||
// A demo is a first result too. Require an affirmative measured timing
|
||||
// contradiction and a direct bypass decision, not benchmark confirmation.
|
||||
if (call.failed !== false || !/^Demo CI gate$/i.test(q.header.trim()) ||
|
||||
/(?:^|\n)[ \t]*(?:>|`{3}|~{3}|example:)/im.test(q.question)) return false;
|
||||
const headline = /^D\s*\d+\s*[—–:-]\s*[a-z][a-z0-9 -]{0,60} demo command: should it bypass the mandatory CI check to reach the <(\d+(?:\.\d+)?) min TTHW target\?$/i.exec(q.question.split('\n')[0]!.trim());
|
||||
const timing = /^ELI10:\s*The agreed onboarding target is under (\d+(?:\.\d+)?) minutes(?: \([^\n)]+\))?\.\s+Today `[^`\n]+` blocks for (\d+(?:\.\d+)?) minutes waiting for a CI check, giving a measured TTHW of (\d+(?:\.\d+)?) minutes(?: [—–-] Red Flag tier vs\. Competitor [A-Z]['’]s \d+(?:\.\d+)? minutes)?\.(?:\s|$)/im.exec(q.question);
|
||||
if (!headline || !timing) return false;
|
||||
const [target, wait, measured] = timing.slice(1).map(Number);
|
||||
return [target, wait, measured].every(n => Number.isFinite(n) && n! > 0) &&
|
||||
Number(headline[1]) === target && wait! >= target! && measured! >= wait!;
|
||||
}
|
||||
if (
|
||||
!/^plan-devex-(?:review-)?[a-z0-9-]+$/i.test(ids[0]![1]!) ||
|
||||
/(?:^|-)(?:mode|setup|scope|routing|prerequisite|next-steps?)(?:-|$)/i.test(ids[0]![1]!)) return false;
|
||||
const body = q.question.replace(/<gstack-qid:[^>]+>/i, '').trim().replace(/\s+/g, ' ');
|
||||
if (!/^D\s*\d+\s*[—–:-]\s*Pass\s+1\s*\(Getting Started\):/i.test(body)) return false;
|
||||
const statement = body.replace(/^D\s*\d+\s*[—–:-]\s*Pass\s+1\s*\(Getting Started\):\s*/i, '');
|
||||
const absentPackageFile = /^(?:The )?(?:README )?quickstart points to a file that doesn['’]t exist in the (?:published )?package\b/i.test(statement) &&
|
||||
/\bhow should (?:the plan|we) fix (?:it|this)\?$/i.test(body);
|
||||
const conflictingGate = /^(?:The )?plan targets TTHW\b[^.!?]*\bbut retains a mandatory\b[^.!?]*\bCI gate with no skip path\b/i.test(statement) &&
|
||||
/\b(?:these are mutually exclusive|these contradict each other)\b/i.test(body) &&
|
||||
/\bhow should (?:the plan|we) resolve (?:this|it)\?$/i.test(body);
|
||||
// A first-run decision may describe shipment, or compare the measured gate
|
||||
// directly with the benchmark. Require the complete affirmative claim and
|
||||
// its repair question; setup/quoted/negated recaps still fail above/below.
|
||||
const completedNative = call.answered === true && call.failed === false;
|
||||
const unshippedQuickstart = completedNative &&
|
||||
/^(?:The )?(?:README )?quickstart points to a file that doesn['’]t ship in the (?:published )?package\. Should we fix the quickstart path in the plan\?$/i.test(statement);
|
||||
const unreachableBenchmark = completedNative &&
|
||||
/^(?:The )?benchmarks set an? <\d+(?:\.\d+)? min TTHW target, but the mandatory \d+(?:\.\d+)?[- ]minute CI gate makes that unreachable\. The plan retains the gate\. How should this plan handle the contradiction\?$/i.test(statement);
|
||||
return absentPackageFile || conflictingGate || unshippedQuickstart || unreachableBenchmark;
|
||||
}
|
||||
|
||||
/** An explicitly quoted developer account plus accuracy-only choices adds no repair. */
|
||||
function answeredQuotedAccuracy(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.sessionId || !call.toolUseId || call.answered !== true || call.failed !== false ||
|
||||
call.questions.length !== 1 || fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
Object.keys(call.answers ?? {}).length !== 1 || !Number.isFinite(Date.parse(call.answeredAt ?? ''))) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.header !== 'Narrative' || q.multiSelect || q.options.length !== 3 || fp.options.length !== 3 ||
|
||||
!fp.options.every((o,i) => o.index === i+1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => call.answers?.[q.question] === o.label) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const expectedOptions = [
|
||||
['This is accurate, proceed', 'Use this narrative as the Developer Perspective section and continue to friction-point decisions.'],
|
||||
['Some of this is wrong, let me correct it', 'Tell me which steps differ; I will fold corrections in before scoring.'],
|
||||
['This is way off, the actual experience is...', 'Describe the real flow and I will rebuild the narrative from it.'],
|
||||
];
|
||||
if (!q.options.every((o,i) => o.label.replace(/ \(recommended\)$/i, '') === expectedOptions[i]![0] && o.description === expectedOptions[i]![1])) return false;
|
||||
const parts = q.question.replace(/^D\d+\s*[—–-]\s*/, '').split(/\n\s*\n/);
|
||||
if (parts.length < 5 || parts[0] !== 'Empathy narrative: does this match what your ML engineer experiences today?' ||
|
||||
!/^Project\/branch\/task: [\w/-]+ branch, [\w. -]+ beta polish, tracing the README getting-started path as written\.$/.test(parts[1]!) ||
|
||||
parts[2] !== 'Here is what I think your ML engineer experiences today:') return false;
|
||||
const quoted = parts.slice(3,-1).join('\n\n');
|
||||
// These are source words in an explicitly bounded quotation, not approval
|
||||
// of any action they mention. No unquoted paragraph may intervene.
|
||||
if (!/^"I [\s\S]+"$/.test(quoted) || (quoted.match(/"/g)?.length ?? 0) !== 2) return false;
|
||||
const explanatory = [
|
||||
"ELI10: This narrative becomes the 'Developer Perspective' section the implementer reads. If it is wrong, the whole review is calibrated against a fake developer.",
|
||||
'Stakes if we pick wrong: we fix friction your developer never hits, or miss the one that actually loses them.',
|
||||
'Recommendation: A because every step above quotes a documented contract in README.md, docs/api.md, docs/current-contracts.md, or docs/package-contents.txt rather than a guess.',
|
||||
'Note: options differ in kind, not coverage — no completeness score.',
|
||||
'A) This is accurate, proceed with this understanding (recommended)',
|
||||
'✅ Every friction point is grounded in a specific documented line, not hypothesized',
|
||||
'✅ Lets the review move straight to per-friction-point decisions with shared context',
|
||||
'❌ If the docs lag the real runtime, a fixed contract could be reviewed as if still broken',
|
||||
'B) Some of this is wrong, let me correct it',
|
||||
'✅ Corrections get folded into the narrative before any scoring happens',
|
||||
'✅ Catches doc-versus-runtime drift the repo cannot show me',
|
||||
'❌ Requires you to spell out which steps differ and how',
|
||||
'C) This is way off, the actual experience is...',
|
||||
'✅ Resets the review against your real onboarding flow',
|
||||
'✅ Prevents scoring against contracts that no longer exist',
|
||||
'❌ Discards a trace that matches the docs line for line, so the docs would also need fixing',
|
||||
'Net: trading trust in the checked-in docs against knowledge only you have about the live SDK.',
|
||||
];
|
||||
const tail = parts.at(-1)!.split('\n').map(line => line.trim());
|
||||
return tail.length === explanatory.length && tail.every((line,i) => line === explanatory[i]);
|
||||
}
|
||||
|
||||
/** A missing release measurement is new work even though its benchmark already exists. */
|
||||
function answeredMeasurementGate(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.sessionId || !call.toolUseId || call.answered !== true || call.failed !== false ||
|
||||
call.questions.length !== 1 || fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
Object.keys(call.answers ?? {}).length !== 1 || !Number.isFinite(Date.parse(call.answeredAt ?? ''))) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.header !== 'Measurement' || q.multiSelect || q.options.length < 2 || q.options.length > 4 ||
|
||||
new Set(q.options.map(o => o.label)).size !== q.options.length || fp.options.length !== q.options.length ||
|
||||
!fp.options.every((o,i) => o.index === i+1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => call.answers?.[q.question] === o.label) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const title = q.question.split('\n')[0]!.replace(/^D\d+\s*[—–-]\s*/, '');
|
||||
return /^Pass \d+ \(DX Measurement\): the < \d+(?:\.\d+)? min target is asserted but never re-measured after the fixes\.$/.test(title) &&
|
||||
/^Evidence: [^\n]+\. Nothing in the plan re-runs that same study after D\d+[–-]D\d+ land, so the beta could ship with the target still unmet and nobody would know until the survey\.$/m.test(q.question) &&
|
||||
q.options.some(o => /^Fix in plan: re-run study as ship gate, record demo and live TTHW(?: \(recommended\))?$/.test(o.label) &&
|
||||
/^Same protocol as docs\/benchmarks\.md on the release candidate; demo TTHW < \d+(?:\.\d+)? min required before tagging\.$/.test(o.description ?? ''));
|
||||
}
|
||||
|
||||
/** A recap can confirm existing approvals, but its text cannot manufacture them. */
|
||||
function answeredRoleplayRecap(fp: AskUserQuestionFingerprint, priorCalls: readonly NativePlanQuestionCall[]): boolean {
|
||||
const call = fp.nativeCall;
|
||||
const completed = (c: NativePlanQuestionCall) => c.answered === true && c.failed === false &&
|
||||
Boolean(c.sessionId && c.toolUseId) && c.questions.length === 1 && !c.questions[0]!.multiSelect &&
|
||||
Array.isArray(c.unansweredQuestionIndices) && c.unansweredQuestionIndices.length === 0 &&
|
||||
Object.keys(c.answers ?? {}).length === 1 && Number.isFinite(Date.parse(c.answeredAt ?? '')) &&
|
||||
c.questions[0]!.options.filter(o => o.label === c.answers?.[c.questions[0]!.question]).length === 1;
|
||||
if (!call || !completed(call) || fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.header !== 'Roleplay' || q.options.length !== 4 || fp.options.length !== 4 ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
new Set(q.options.map(o => o.label)).size !== 4 || /<gstack-qid/i.test(q.question)) return false;
|
||||
const labels = q.options.map(o => o.label.replace(/ \(recommended\)$/i, ''));
|
||||
if (labels.join('|') !== 'All of them, fix every confusion point|Let me pick which ones matter|Critical ones only (#1, #2, #5)|This is unrealistic, our developers already know the context' ||
|
||||
call.answers?.[q.question] !== q.options[0]!.label) return false;
|
||||
const mapping = /^Address #1 through #(\d+), matching the D(\d+)[–-]D(\d+) decisions\.$/.exec(q.options[0]!.description ?? '');
|
||||
if (!mapping) return false;
|
||||
const [size, first, last] = mapping.slice(1).map(Number);
|
||||
if (size !== 5 || last! - first! + 1 !== size || first! < 1 || last! > 1000) return false;
|
||||
if (q.options[1]!.description !== 'Tell me which numbers to keep and which to drop.' ||
|
||||
q.options[2]!.description !== 'Fix quickstart, CI gate, and upgrade; leave signature order and auth error.' ||
|
||||
q.options[3]!.description !== 'Skip the confusion points; keep contracts as drafted.') return false;
|
||||
const prior: NativePlanQuestionCall[] = [];
|
||||
for (let decision = first!; decision <= last!; decision++) {
|
||||
const matches = priorCalls.filter(c => c.sessionId === call.sessionId && completed(c) &&
|
||||
c.toolUseId !== call.toolUseId && Date.parse(c.answeredAt!) < Date.parse(call.answeredAt!) &&
|
||||
new RegExp(`^D${decision}\\s*[—–-]\\s*`).test(c.questions[0]!.question));
|
||||
if (matches.length !== 1 || !/^Fix in plan:/.test(matches[0]!.answers![matches[0]!.questions[0]!.question]!)) return false;
|
||||
prior.push(matches[0]!);
|
||||
}
|
||||
// Each observed confusion point refers to the same already-approved contract.
|
||||
// The fixture's five independent defects remain explicit; new measurement,
|
||||
// documentation or TODO decisions do not enter this confirmation path.
|
||||
const subjects = [/examples\/first_eval\.py/, /\bCI\b/, /\brun_eval\b[\s\S]*\brun_batch\b|\brun_batch\b[\s\S]*\brun_eval\b/, /\bAuthError\b/, /Client\.evaluate\(\)/i];
|
||||
if (prior.some((c, i) => !subjects[i]!.test(c.questions[0]!.question))) return false;
|
||||
// Sharing a subject or a "fix" prefix is not approval of this remedy. Bind
|
||||
// each chosen option and its entire consequence to the contract recapped.
|
||||
const approvedRepairs = [
|
||||
['Fix in plan: demo-first quickstart + resolve first_eval.py', 'README leads with python -m evalkit.demo; ship or remove first_eval.py; add a packaging check for documented paths.'],
|
||||
['Fix in plan: no CI check on mock-transport runs; gate the first live eval instead', 'Demo returns immediately; CI check with existing progress/timeout messaging moves to the first keyed evaluation.'],
|
||||
['Fix in plan: align order + keyword-only + clear TypeError', 'run_batch(dataset, evaluator) matching run_eval; keyword-only enforcement; positional misuse raises a TypeError naming the expected call.'],
|
||||
['Fix in plan: coded, causal AuthError with fix and redaction', 'Error code, key source, cause, console fix URL, redacted key prefix, help link. Matches the existing error pattern.'],
|
||||
['Fix in plan: alias + DeprecationWarning + migration guide + codemod', 'evaluate() delegates to run() with a warning through 2.x betas; changelog and docs/api.md gain a migration section; sed/codemod recipe shipped.'],
|
||||
];
|
||||
if (prior.some((c, i) => {
|
||||
const question = c.questions[0]!;
|
||||
const selected = question.options.find(o => o.label === c.answers![question.question])!;
|
||||
return selected.label.replace(/ \(recommended\)$/i, '') !== approvedRepairs[i]![0] ||
|
||||
selected.description !== approvedRepairs[i]![1];
|
||||
})) return false;
|
||||
const parts = q.question.replace(/^D\d+\s*[—–-]\s*/, '').split(/\n\s*\n/);
|
||||
if (parts.length !== 5 || parts[0] !== 'First-time developer roleplay: which confusion points should the plan address?' ||
|
||||
!/^Project\/branch\/task: [\w/-]+ branch, [\w. -]+ beta polish; roleplayed your ML engineer through the README as written\.$/.test(parts[1]!) ||
|
||||
parts[2] !== 'I roleplayed as your ML engineer attempting the getting started flow. Here is what confused me, with timestamps:') return false;
|
||||
const observed = parts[3]!.split('\n');
|
||||
const source = String.raw`[\w./-]+:\d+(?:-\d+)?`;
|
||||
const observation = [
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +#1 \x60python examples/first_eval\.py\x60 fails: file not in package or archive \(${source}, ${source}\)\. "[^"\n]+"$`),
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +#2 Keyless demo starts a remote CI check on a sample-project binding I never created \(${source}, ${source}\)\. "[^"\n]+"$`),
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +Scores print\. Works, but \d+ min vs the \d+ min target \(${source}\)\. Impression: slow\.$`),
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +#3 run_batch fails inside the evaluator because its argument order is the reverse of run_eval \(${source}\)\. "[^"\n]+"$`),
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +#4 \x60AuthError: request failed\x60 on a wrong-project key; I check network and server status first because nothing says "key" \(${source}\)\.$`),
|
||||
new RegExp(String.raw`^T\+\d+:\d+ +#5 v1 project upgraded: every client\.evaluate\(\) raises AttributeError; changelog has no migration entry \(${source}\)\. Final state: file an issue or pin v1\.$`),
|
||||
];
|
||||
if (observed.length !== observation.length || observed.some((line, i) => !observation[i]!.test(line))) return false;
|
||||
// Consume the complete decision explanation too. Additional work under a
|
||||
// valid heading or in a choice description must remain substantive.
|
||||
const range = `D${first}–D${last}`;
|
||||
const tail = parts[4]!.replace(new RegExp(`D${first}[–-]D${last}`, 'g'), range).split('\n');
|
||||
const expected = [
|
||||
'ELI10: Each numbered point is a place a real first-time user stops and asks a question nobody is there to answer. The plan should remove every one it reasonably can.',
|
||||
'Stakes if we pick wrong: leave one in and that is the step where the developer\'s session ends; each maps to a contract PLAN.md explicitly asked to be reviewed.',
|
||||
`Recommendation: A because all five map one-to-one to the ${range} decisions you already resolved as "fix in plan", so addressing all of them is consistent with those calls.`,
|
||||
'Completeness: A=10/10, B=depends on selection, C=6/10, D=1/10',
|
||||
'A) All of them, fix every confusion point (recommended)',
|
||||
`✅ Consistent with ${range}; every confusion point already has an agreed fix`,
|
||||
'✅ Leaves no known dead end in the first 30 minutes of use',
|
||||
'❌ Full set of fixes touches README, client.py, demo gate, error class, and changelog (human: ~3 days / CC: ~1.5 hours)',
|
||||
'B) Let me pick which ones matter',
|
||||
'✅ Lets you drop a point if you know something the docs do not show',
|
||||
'✅ Keeps the plan focused on what you consider blocking',
|
||||
`❌ Reopens decisions ${range} that were just settled`,
|
||||
'C) The critical ones only (#1, #2, #5), skip #3 and #4',
|
||||
'✅ Covers the broken quickstart, the TTHW blocker, and the upgrade break',
|
||||
'✅ Smaller diff to review',
|
||||
'❌ Ships an inconsistent API and an undiagnosable auth error in a DX polish release',
|
||||
'D) This is unrealistic, our developers already know the context',
|
||||
'✅ Zero work now',
|
||||
'✅ Valid if every beta user is internal and already trained',
|
||||
'❌ README.md:3-5 describes an external ML engineer meeting the SDK fresh, which contradicts this',
|
||||
'Net: trading a known, already-scoped set of fixes against leaving a documented dead end in the first session.',
|
||||
];
|
||||
return tail.length === expected.length && tail.every((line, i) => line.trim() === expected[i]);
|
||||
}
|
||||
|
||||
/** A batched native call remains one decision; the caller owns call-ID deduplication. */
|
||||
export function isDevexReviewIssue(fp: AskUserQuestionFingerprint, priorCalls: readonly NativePlanQuestionCall[] = []): boolean {
|
||||
if (answeredQuotedAccuracy(fp) || answeredRoleplayRecap(fp, priorCalls)) return false;
|
||||
return answeredMeasurementGate(fp) || answeredSetupRepair(fp) || answeredContractRepair(fp) || answeredKeylessDemoRepair(fp) || answeredDocumentationFollowup(fp) || questionRecords(fp, true).some(substantiveIssue);
|
||||
}
|
||||
|
||||
/** Key acquisition docs and eliminating the demo's key requirement are distinct work. */
|
||||
function answeredKeylessDemoRepair(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (call?.answered !== true || call.failed !== false || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2 || new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
fp.options.length !== q.options.length || fp.options.some((o, i) => o.index !== i + 1 || o.label !== q.options[i]!.label) ||
|
||||
!/^Golden path$/i.test(q.header.trim()) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const selected = q.options.filter(o => o.label === call.answers?.[q.question]);
|
||||
if (selected.length !== 1 || !/^Install, demo, then key(?: \(recommended\))?$/i.test(selected[0]!.label) ||
|
||||
!/\bDemo path is guaranteed keyless and offline; if the runtime currently insists on a key for the demo, remove that check\b/.test(selected[0]!.description ?? '')) return false;
|
||||
const lines = q.question.split('\n');
|
||||
return /^D\s*\d+\s*[—–:-]\s*Pass 1 Getting Started \((?:10|[0-9])\/10 today\): should the golden path put the demo BEFORE the API key step\?$/i.test(lines[0]!) &&
|
||||
/^ELI10: Today README "Getting started" \(lines \d+-\d+\) reads install, set [A-Z][A-Z_]+, run a missing file\./m.test(q.question);
|
||||
}
|
||||
|
||||
/** New documentation and example obligations are separate from the original repairs. */
|
||||
function answeredDocumentationFollowup(fp: AskUserQuestionFingerprint): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.answered || call.failed !== false || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}`) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.options.length < 2 || new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
administrativeQuestion(q.header, q.question, q.options)) return false;
|
||||
const selected = q.options.filter(o => o.label === call.answers?.[q.question]);
|
||||
const ids = [...q.question.matchAll(/<gstack-qid:([^>]+)>/gi)];
|
||||
if (selected.length !== 1 || ids.length !== 1 || (q.question.match(/<gstack-qid/gi)?.length ?? 0) !== 1) return false;
|
||||
const label = selected[0]!.label.trim().replace(/^[A-Z][.):]\s*/i, '').replace(/\s*\(recommended\)$/i, '');
|
||||
const headline = q.question.split('\n')[0]!.replace(/<gstack-qid:[^>]+>/i, '').trim();
|
||||
if (/^plan-devex-review-todo\d+-migration-guide$/i.test(ids[0]![1]!) && /^TODO[- ]\d+ Migration$/i.test(q.header.trim())) {
|
||||
// A written upgrade guide is additional work beyond the accepted runtime
|
||||
// compatibility shim. Require that distinct gap and the selected doc task;
|
||||
// a recap, hypothetical example or unselected guide cannot supply it.
|
||||
const parts = q.question.replace(/\s*<gstack-qid:[^>]+>\s*$/i, '').trim().split(/\n\s*\n/);
|
||||
const compact = (text: string | undefined) => (text ?? '').replace(/\s+/g, ' ').trim();
|
||||
return call.answered === true && parts.length === 5 &&
|
||||
/^D\s*\d+\s*[—–:-]\s*TODO: should the plan include a v\d+→v\d+ written migration guide\?$/i.test(compact(parts[0])) &&
|
||||
/^The deprecation shim \(T\d+\) handles the runtime experience: v\d+ callers get a DeprecationWarning naming `[a-z_]\w*\(\)` as the replacement\. But there is currently no written migration guide in docs\/\.$/i.test(compact(parts[1])) &&
|
||||
/^A one-page migration guide covers: - What changed \(`[a-z_]\w*\(\)` → `[a-z_]\w*\(\)`\) - What stayed the same \(all other APIs\) - How to find and update callsites \(grep for `[^`\n]+`\) - When the shim is removed \(e\.g\., v\d+(?:\.\d+)?\)$/i.test(compact(parts[2])) &&
|
||||
/^Without it, developers upgrading a large codebase need to discover the change at each call site rather than planning the migration upfront\. The changelog has the what; the guide provides the how and the timeline\.$/i.test(compact(parts[3])) &&
|
||||
/^Completeness: A=(?:10|[0-9])\/10 \(complete\), B=(?:10|[0-9])\/10 \(runtime-only, no planning\), C=(?:10|[0-9])\/10$/i.test(compact(parts[4])) &&
|
||||
/^Add to TODOS\.md [—–-] include migration guide in plan$/i.test(label) &&
|
||||
/^Add docs\/migration-v\d+-v\d+\.md as a P[0-3] task\. One page covering the rename, unchanged APIs, grep command to find callsites, and shim removal timeline\. Completeness: (?:10|[0-9])\/10\.$/i.test(compact(selected[0]!.description));
|
||||
}
|
||||
if (ids[0]![1] === 'devex-api-key-docs' && /^API key docs$/i.test(q.header.trim())) {
|
||||
// The earlier auth-error decision changes runtime diagnostics. This one
|
||||
// adds the missing acquisition instructions to the README itself.
|
||||
return /^D\s*\d+\s*[—–:-]\s*Pass\s+\d+:\s*Documentation\s*[—–:-]\s*README says ['"][^'"]+['"] but never says where to get one\.$/i.test(headline) &&
|
||||
/^Add key acquisition link to README$/i.test(label);
|
||||
}
|
||||
if (ids[0]![1] === 'devex-todo-real-world-examples' && /^TODO examples$/i.test(q.header.trim())) {
|
||||
// The quickstart repair supplies one missing file. These additional
|
||||
// custom-data examples are an independently accepted follow-up obligation.
|
||||
return /^D\s*\d+\s*[—–:-]\s*TODO check:\s*Real-world examples beyond the bundled sample data\?$/i.test(headline) &&
|
||||
/^\*\*What:\*\* Add \d+(?:-\d+)? additional examples\/ files showing real use cases\b/m.test(q.question) &&
|
||||
/^(?:Add to TODOS\.md for post-beta|Build it now as part of this plan)$/i.test(label);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** Select POLISH only on the recognized mode menu; leave all other answers unchanged. */
|
||||
export function devexReviewModePick(fp: AskUserQuestionFingerprint): number | null {
|
||||
if (fp.nativeCall && fp.nativeCall.questions.length !== 1) return null;
|
||||
const record = questionRecords(fp)[0];
|
||||
const text = record ? `${record.header} ${record.question}` : '';
|
||||
if (!/<gstack-qid:plan-devex-review-mode>/i.test(text)
|
||||
&& !/how\s*deep\s*should\s*this\s*dx\s*review|which\s*(?:dx\s*)?review\s*mode/i.test(text)) return null;
|
||||
const modes = fp.options.map(option => ({
|
||||
index: option.index,
|
||||
mode: /^(?:[A-C][.)])?DX(POLISH|EXPANSION|TRIAGE)(?:$|[^A-Z])/.exec(
|
||||
option.label.split(/[│┌\r\n]/, 1)[0]!.replace(/\s+/g, '').toUpperCase(),
|
||||
)?.[1],
|
||||
}));
|
||||
if (!['POLISH', 'EXPANSION', 'TRIAGE'].every(mode => modes.filter(option => option.mode === mode).length === 1)) return null;
|
||||
return modes.find(option => option.mode === 'POLISH')!.index;
|
||||
}
|
||||
@@ -0,0 +1,293 @@
|
||||
import type { NativePlanQuestion, PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
export const DEVEX_SEEDED_GAPS = [
|
||||
'local-ci-gate', 'missing-quickstart', 'reversed-arguments', 'opaque-auth-error', 'breaking-upgrade',
|
||||
] as const;
|
||||
export type DevexSeededGap = typeof DEVEX_SEEDED_GAPS[number];
|
||||
|
||||
/** Bind an unnamed signature question to its own first asserted explanation. */
|
||||
function explainedReversedSignatures(q: NativePlanQuestion, title: string): boolean {
|
||||
const question = /^(?:Journey stage [A-Z ]+: )?the two public functions take the same two arguments in (?:opposite|reversed) positional order\. How should (?:the plan|we) (?:fix|align|unify) the signatures\?$/i.test(title);
|
||||
const traced = /^Journey stage: REAL USAGE\. Two sibling functions take the same two arguments in (?:opposite|reversed) order\.$/i.test(title);
|
||||
const declared = traced || /^Journey stage(?: REAL USAGE:|: REAL USAGE\.) The two public evaluation functions take the same two arguments in (?:opposite|reversed) order\.$/i.test(title);
|
||||
// A dedicated assertion can put its named signatures in its own Evidence
|
||||
// field. Bind subject, source identities and repair instead of menu wording.
|
||||
const subject = title.replace(/^Journey stage(?: REAL USAGE:|: REAL USAGE\.)\s*/i, '');
|
||||
const evidenced = !question && !declared &&
|
||||
/^(?:the )?(?:two|both) public (?:evaluation )?functions take\b/i.test(subject) &&
|
||||
/\bthe same two arguments\b/i.test(subject) && /\b(?:opposite|reversed) (?:positional )?order\.?$/i.test(subject);
|
||||
const declaration = declared || evidenced;
|
||||
if (!question && !declaration) return false;
|
||||
const lines = q.question.split('\n');
|
||||
if (lines[0]!.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '') !== title) return false;
|
||||
const explanation = lines.findIndex(line => line.startsWith('ELI10: '));
|
||||
const context = lines.slice(1, explanation).filter(line => line.trim());
|
||||
const project = traced ? /^Project\/branch\/task: [^;\n]+; ([\w./-]+):\d+(?:[-–]\d+)?\.$/.exec(context[0] ?? '') : declaration && context.length === 1
|
||||
? /^Project\/branch\/task: [^;\n]+; ([\w./-]+) lines? \d+(?: to |[-–])\d+\.$/.exec(context[0]!) : null;
|
||||
if (explanation < 1 || (declared && !project) || (!traced && !evidenced && context.some(line =>
|
||||
(!declaration && !/^Project\/branch\/task: [^;\n]+; reviewing the public function signatures in [\w./-]+\.$/.test(line)) ||
|
||||
/\b(?:quoted|source excerpt|source example|hypothetical|historical|not (?:a )?current|if approved)\b/i.test(line)))) return false;
|
||||
// Inline code may name each signature; a quoted/fenced explanation, earlier
|
||||
// unrelated sentence, past definition or hypothetical definition cannot.
|
||||
const declaredSignatures = traced
|
||||
? /^I traced the first real integration after the demo\. ([\w./-]+) lists the two evaluation functions: (`?)run_eval\(\s*dataset\s*,\s*evaluator\s*\)\2 and (`?)run_batch\(\s*evaluator\s*,\s*dataset\s*\)\3\./.exec(context[1] ?? '')
|
||||
: declaration && /^ELI10: ([\w./-]+) documents (`?)run_eval\(\s*dataset\s*,\s*evaluator\s*\)\2 and (`?)run_batch\(\s*evaluator\s*,\s*dataset\s*\)\3\. Same two concepts, reversed positional order, and neither function requires keywords\./.exec(lines[explanation]!);
|
||||
if (!evidenced && (declaration ? !declaredSignatures || declaredSignatures[1] !== project?.[1]
|
||||
: !/^ELI10: [\w./-]+(?: lines? \d+(?:\s*[-–]\s*\d+)?)? define (`?)run_eval\(\s*dataset\s*,\s*evaluator\s*\)\1 and (`?)run_batch\(\s*evaluator\s*,\s*dataset\s*\)\2\./.test(lines[explanation]!))) return false;
|
||||
const currentProse = (text: string) => {
|
||||
let fence = false;
|
||||
return text.split('\n').filter(line => {
|
||||
if (/^\s*(?:```|~~~)/.test(line)) { fence = !fence; return false; }
|
||||
return !fence && !/^\s*>/.test(line);
|
||||
}).join('\n')
|
||||
.replace(/(^|[.!?\n]\s*)((?:Correction:\s*)?(?:this|that|the) (?:evidence|trace) (?:is|was|has been) )["“'‘`](withdrawn|rejected|cancelled|canceled|superseded|historical|hypothetical|(?:not|no longer) current)["”'’`]/gi, '$1$2$3')
|
||||
.replace(/`[^`\n]*`|"[^"\n]*"|“[^”\n]*”/g, '');
|
||||
};
|
||||
// The traced declaration owns its named signatures before ELI10, so its
|
||||
// currentness must include that same source paragraph.
|
||||
const current = currentProse(lines.slice(traced || evidenced ? 1 : explanation).join('\n'));
|
||||
if ((current.match(/^ELI10:/gm)?.length ?? 0) !== 1) return false;
|
||||
if (declaration && /\b(?:if|once|when|unless) (?:approved|accepted)|\b(?:after|pending) approval\b/i.test(current)) return false;
|
||||
if (declaration && /(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:these|the) (?:functions|signatures) (?:are (?:now|already)|have been) (?:aligned|consistent)\b/i.test(current)) return false;
|
||||
if ((traced || evidenced) && /(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:this|that|the) (?:trace|evidence) (?:is|was|has been) (?:withdrawn|rejected|cancelled|canceled|superseded|historical|hypothetical|(?:not|no longer) current)\b/i.test(current)) return false;
|
||||
if (/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:(?:this|that|the) (?:finding|explanation)|(?:(?:this|that|the) )?argument[- ]order (?:issue|defect)|these signatures)\b[^.\n]*\b(?:withdrawn|rejected|(?:already )?(?:fixed|resolved)|historical|(?:not|no longer) current)\b/i.test(current) ||
|
||||
/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:there is|there's) no argument[- ]order (?:issue|defect)\b/i.test(current) ||
|
||||
/(?:^|[.!?\n]\s*)(?:Correction:\s*)?run_eval and run_batch now (?:use|take) the same positional order\b/i.test(current)) return false;
|
||||
if (evidenced) {
|
||||
// Only an asserted citation at the start of this decision's field owns
|
||||
// the pair; quoted examples, later borrowed prose and split fields do not.
|
||||
const fields = lines.slice(1, explanation + 1).filter(line => /^(?:Evidence|ELI10):/.test(line));
|
||||
const pair = /^(?:Evidence|ELI10):\s*[\w./-]+(?: lines? \d+(?:[-–]\d+)?|:\d+(?:[-–]\d+)?)?:\s*(`?)run_eval\(\s*dataset\s*,\s*evaluator\s*\)\1 and (`?)run_batch\(\s*evaluator\s*,\s*dataset\s*\)\2(?:[.;]|$)/;
|
||||
if (!fields.some(line => pair.test(line)) || fields.some(line =>
|
||||
/^(?:Evidence|ELI10):\s*(?:>|`|"|“|Source\b|Quoted\b|Historical\b|Earlier\b|Example\b|Hypothetical\b|If\b|Assuming\b|Provided\b)/i.test(line))) return false;
|
||||
return q.options.some(option => {
|
||||
const label = currentProse(option.label.replace(/`(\(\s*dataset\s*,\s*evaluator\s*\))`/g, '$1'));
|
||||
const remedy = currentProse(option.description ?? '');
|
||||
const first = remedy.split(/[.!?\n]/)[0] ?? '';
|
||||
return /^(?:[A-D]\)\s*)?(?:Align|Unify|Standardize)\b/i.test(label) && /\(\s*dataset\s*,\s*evaluator\s*\)/.test(label) &&
|
||||
/\bsame (?:positional )?order\b/i.test(first) && /\bboth functions\b/i.test(first) &&
|
||||
/\bkeywords? (?:accepted|supported)\b|\baccept keywords\b/i.test(remedy) &&
|
||||
/\bswaps? (?:is |are )?(?:detected|caught|rejected)\b/i.test(remedy) && /\b(?:clear|actionable) (?:error|message)\b/i.test(remedy) &&
|
||||
!/\b(?:if|unless|when|once|after|pending)\b|\b(?:no|not|never|without|do not|don't)\b|\b(?:other|another|foreign|different) (?:functions?|API|pair|project|issue)\b/i.test(`${label}\n${remedy}`) &&
|
||||
!/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:this|that|the) (?:option|action|correction) (?:is|was|has been) (?:historical|withdrawn|rejected|cancelled|canceled|superseded|(?:not|no longer) current)\b/i.test(remedy) &&
|
||||
!/\brun_(?!eval\b|batch\b)\w+\b/.test(remedy);
|
||||
});
|
||||
}
|
||||
// A declared reversal may offer a keyword-only repair instead of a swap
|
||||
// guard. It must bind both arguments to both functions in the same option.
|
||||
if (declaration) return q.options.some(option =>
|
||||
(traced ? /^Fix in plan: same order \+ keyword-only for both(?: \(recommended\))?$/i.test(option.label) &&
|
||||
/^✅\s*run_eval\(\*\s*,\s*dataset\s*,\s*evaluator\s*\) and run_batch\(\*\s*,\s*dataset\s*,\s*evaluator\s*\); wrong order becomes a TypeError naming the parameter at the call site\b/i.test(option.description ?? '')
|
||||
: /^(?:Align|Unify|Standardize) order \+ keyword-only(?: \(recommended\))?$/i.test(option.label) &&
|
||||
/^Both functions (?:take|accept|use) dataset and evaluator as keyword-only in the same order\./i.test(option.description ?? '')) &&
|
||||
!/\b(?:if|once|when|unless) (?:approved|accepted)|\b(?:after|pending) approval\b/i.test(currentProse(option.description ?? '')) &&
|
||||
!/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:(?:do not|don't|never) (?:change|align|unify) (?:either|both|the|these) (?:functions?|signatures?)\b|(?:do not|don't|never) (?:make|require) (?:either|both|the) (?:functions?|signatures?|arguments?) keyword-only\b|(?:this|the) (?:option|correction|action) is (?:withdrawn|rejected|cancelled)\b)/i.test(currentProse(option.description ?? '')));
|
||||
// The same offered action must align both functions and retain the call-site
|
||||
// swap guard. Selecting an offered alternate or deferral is still a decision.
|
||||
return q.options.some(option => /^Same order\s*\+\s*swap guard(?: \(recommended\))?$/i.test(option.label) &&
|
||||
/^(?:✅\s*)?Both (?:become|use|take) `?\(\s*dataset\s*,\s*evaluator\s*\)`?, accept keywords, and raise a call-site `?TypeError`? naming the swapped argument and the fix if types are reversed\./i.test(option.description ?? '') &&
|
||||
!/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:(?:do not|don't|never) (?:change|align|unify) (?:either|both|the) signatures?\b|(?:do not|don't|never|skip) (?:add|require|implement) (?:a |the )?swap guard\b|(?:this|the) (?:option|correction|action) is (?:withdrawn|rejected|cancelled)\b)/i.test(currentProse(option.description ?? '')));
|
||||
}
|
||||
|
||||
/** Identify a dedicated seed decision by its subject and meaningful alternatives. */
|
||||
function decisionGaps(q: NativePlanQuestion): DevexSeededGap[] {
|
||||
const rawTitle = q.question.split('\n')[0]!.trim().replace(/^D\s*\d+\s*[—–:-]\s*/i, '');
|
||||
const title = rawTitle.replace(/`([^`\n]+)`/g, '$1');
|
||||
const questionMarks = title.match(/\?/g)?.length ?? 0;
|
||||
const upgradeVocabulary = /\b(?:alias|warning|compatibility|deprecat\w*|migration|remov\w*|rename|keep)\b/i.test(title);
|
||||
// A named method becoming its replacement is a transition even when the
|
||||
// title asks about a soft landing. Its own explanation must establish the gap.
|
||||
const upgradeTransition = !upgradeVocabulary && /\bClient\.evaluate\(\) becomes Client\.run\(\)/i.test(title);
|
||||
// Journey labels, possessives and a positive inclusive aside format the
|
||||
// asserted subject. Keep the original title for all meaning/currentness checks.
|
||||
// These six stages come from the skill's journey trace. A decision may span
|
||||
// adjacent touchpoints without changing the subject or who asserts it.
|
||||
const journeyStage = '(?:DISCOVER|INSTALL|HELLO WORLD|REAL USAGE|DEBUG|UPGRADE)';
|
||||
const stage = title.match(new RegExp(`^Journey stage ${journeyStage}(?:\\s*\\/\\s*${journeyStage})?: (.+)$`, 'i'))
|
||||
?? title.match(new RegExp(`^Journey stage: ${journeyStage}(?:\\s*\\/\\s*${journeyStage})?\\. (.+)$`, 'i'));
|
||||
// New field declarations require a canonical stage. Existing direct
|
||||
// questions can still name another touchpoint without normalizing it.
|
||||
if (!stage && (/^Journey stage:/i.test(title) ||
|
||||
(/^Journey stage\b/i.test(title) && !title.includes('?')))) return [];
|
||||
if (stage && /^(?:Assuming|Provided)\b/i.test(stage[1]!.trim())) return [];
|
||||
const assertionTitle = stage ? stage[1]!
|
||||
.replace(/\b([A-Za-z0-9_.]+)['’]s\b/g, '$1')
|
||||
.replace(/, including ([A-Za-z0-9_-]+(?: [A-Za-z0-9_-]+){0,6}),/gi, (aside, subject: string) =>
|
||||
/\b(?:if|unless|assuming|provided|except|excluding|only|no|not|never|without|was|were|is|are|has|had|may|might|could|would|historical|earlier|quoted|source|example|hypothetical|fixed|resolved|cancelled|canceled|withdrawn|rejected|superseded)\b/i.test(subject) ? aside : '')
|
||||
: title;
|
||||
const opaqueAuthentication = /^(?:The )?authentication error says nothing[.?]?$/i.test(assertionTitle);
|
||||
const vanishingUpgrade = /^v\d+ Client\.evaluate\(\) vanishes in v\d+ with no warning, alias, or guide[.?]?$/i.test(assertionTitle);
|
||||
// A defect heading can assert a prerequisite or compare named signatures
|
||||
// without a finite verb. Keep these semantic families narrow: a topic label,
|
||||
// healthy signature pair or optional check is not an asserted defect.
|
||||
const nominalDefect = /^(?:The )?(?:Mandatory|Required) (?:\d+(?:\.\d+)?[- ](?:minute|second) )?(?:remote )?CI (?:check|gate) before (?:the )?first local (?:result|evaluation|run)[.?]?$/i.test(assertionTitle) ||
|
||||
/^run_eval\(\s*dataset\s*,\s*evaluator\s*\) (?:vs\.?|versus|and) run_batch\(\s*evaluator\s*,\s*dataset\s*\): (?:reversed|opposite|swapped) (?:positional|argument) order[.?]?$/i.test(assertionTitle);
|
||||
const nominalSubject = /^(?:Mandatory|Required|Optional)\b[^?!\n]*\bCI (?:check|gate)\b/i.test(assertionTitle) ||
|
||||
/^run_eval\([^)]+\) (?:vs\.?|versus|and) run_batch\([^)]+\):/i.test(assertionTitle);
|
||||
if (nominalSubject && !nominalDefect) return [];
|
||||
const signatureDeclaration = /^run_eval\(\s*dataset\s*,\s*evaluator\s*\) and run_batch\(\s*evaluator\s*,\s*dataset\s*\) (?:take|takes)\b/i.test(assertionTitle);
|
||||
if (/^run_eval\([^)]+\) and run_batch\([^)]+\) (?:take|takes)\b/i.test(assertionTitle) && !signatureDeclaration) return [];
|
||||
// The named tuples can establish the reversal without an adjective. Keep
|
||||
// their identities and order together; malformed or negated comparisons
|
||||
// cannot fall through to the broader direct-question path.
|
||||
const tupleSubject = /^run_eval (?:takes?|does not take)\b[^\n]*\brun_batch\b/i.test(assertionTitle);
|
||||
const tuples = /^run_eval takes\s*\(\s*(\w+)\s*,\s*(\w+)\s*\) (?:but|while) run_batch takes\s*\(\s*(\w+)\s*,\s*(\w+)\s*\)(?:[.?]|\. Fix in plan\?)?$/i.exec(assertionTitle);
|
||||
const reversedTuples = Boolean(tuples && tuples[1] !== tuples[2] &&
|
||||
[tuples[1], tuples[2]].sort().join(',') === 'dataset,evaluator' &&
|
||||
tuples[1] === tuples[4] && tuples[2] === tuples[3]);
|
||||
if (tupleSubject && !reversedTuples) return [];
|
||||
const finiteTitle = signatureDeclaration ? assertionTitle.replace(/\([^)]*\)/g, '') : assertionTitle;
|
||||
// Negative availability asserts a missing referenced file. Bind it to that
|
||||
// object; do not erase a negation of the quickstart's own reference or gate.
|
||||
const absentReference = /\b(?:points?|references?) (?:at|to) (?:examples\/first_eval\.py|(?:a|the) (?:file|example)),? (?:which|that) (?:is not (?:shipped|in (?:the )?(?:package|wheel)(?: or (?:the )?(?:release )?examples archive)?)|does not (?:ship|exist))[.?]?$/i.test(assertionTitle);
|
||||
const newAssertion = nominalDefect || signatureDeclaration || reversedTuples || absentReference || opaqueAuthentication || vanishingUpgrade;
|
||||
const guardedDeclaration = Boolean(stage || newAssertion || upgradeTransition);
|
||||
const polarityTitle = absentReference ? title.replace(/\bdoes not (ship|exist)([.?]?)$/i, 'is absent$2') : title;
|
||||
// Punctuation cannot route a newly admitted asserted family around its
|
||||
// ownership checks; an offered alternate still resolves the same decision.
|
||||
const declaration = (questionMarks === 0 || (questionMarks === 1 && title.endsWith('?'))) &&
|
||||
(/^(?:[A-Za-z0-9_.]+\s+){1,12}(?:points?|references?|blocks?|requires?|takes?|raises?|removes?|drops?)\b/i.test(finiteTitle) || nominalDefect || opaqueAuthentication || vanishingUpgrade) &&
|
||||
!/^`[^`]*`$/.test(rawTitle) &&
|
||||
!/\b(?:if|unless|suppose|might|may|could|would|previously|earlier|historical|hypothetical|example|quoted|source|never|no longer|does not|do not|did not)\b/i.test(polarityTitle);
|
||||
if (newAssertion && !declaration) return [];
|
||||
if ((!declaration && (!title.endsWith('?') || questionMarks !== 1)) ||
|
||||
/^`[^`]*`[.?]?$/.test(rawTitle) ||
|
||||
/^(?:>|"|“|Example\b|Quoted\b|Source(?: excerpt| example)?[,:.]|Historical\b|Earlier review\b|If (?:approved|accepted)\b|Assuming\b|Provided\b|Suppose\b)|\bhypothetical\b/i.test(title) ||
|
||||
/\b(?:if|once|when|unless) (?:approved|accepted)|\b(?:after|pending) approval\b/i.test(title) ||
|
||||
/\b(?:continue|proceed|next section|move on|format|already (?:fixed|resolved))\b/i.test(title) ||
|
||||
/\b(?:have|did)\b[^?]*\bread\b|\b(?:narrative|trace|recap|summary)\b[^?]*\b(?:accurate|match|confirm)\b/i.test(title) ||
|
||||
/\b(?:report|summary|recap)\b[^?]*\b(?:mention|include|reference|list)\b|\b(?:mention|include|reference|list)\b[^?]*\b(?:report|summary|recap)\b/i.test(title)) return [];
|
||||
let offered = q.options;
|
||||
let signatureOptions = q.options;
|
||||
if (declaration || upgradeTransition) {
|
||||
const currentProse = (text: string, offeredAction = false) => {
|
||||
// In a tuple decision, a semicolon also separates current assertions.
|
||||
// Quotations and fenced examples are still removed as whole statements.
|
||||
if (reversedTuples || upgradeTransition) text = text.replace(/;/g, '.');
|
||||
// An option's trailing effort estimate separates its prose from an owned
|
||||
// status even without punctuation. Keep it on the same line so a quoted
|
||||
// historical sentence is still removed as one quotation below.
|
||||
const bounded = guardedDeclaration && offeredAction ? text.replace(/(\(human:[^()\n]{1,80}\/ CC:[^()\n]{1,80}\))[ \t]+(?=(?:Correction:\s*)?(?:(?:this|that|the) (?:option|action|correction)|D\s*[1-9]\d*) (?:is|was|has been) ["“'‘`]?(?:cancelled|canceled|superseded|withdrawn|rejected|(?:not|no longer) current)\b)/gi, '$1. ') : text;
|
||||
// Preserve a scalar status asserted by a current, unquoted owner before
|
||||
// removing source quotations. The owner must still match this decision.
|
||||
const owned = guardedDeclaration ? bounded.replace(/(^|[.!?\n]\s*)((?:Correction:\s*)?(?:(?:this|that|the) (?:finding|issue|gap|defect|explanation|option|action|correction)|D\s*[1-9]\d*) (?:is|was|has been) )["“'‘`](cancelled|canceled|superseded|withdrawn|rejected|(?:not|no longer) current)["”'’`]/gim, '$1$2$3') : bounded;
|
||||
let fence = false;
|
||||
return owned.split('\n').filter(line => {
|
||||
if (/^\s*(?:```|~~~)/.test(line)) { fence = !fence; return false; }
|
||||
return !fence && !/^\s*>/.test(line);
|
||||
}).join('\n')
|
||||
.replace(/(^|[.!?\n]\s*)((?:Correction:\s*)?(?:this|that|the) (?:finding|issue|gap|defect|explanation|option|action|correction) (?:is|was|has been) )["“](withdrawn|rejected|(?:already )?(?:fixed|resolved)|historical|(?:not|no longer) current|cancelled)["”]/gim, '$1$2$3')
|
||||
.replace(/`[^`\n]*`|"[^"\n]*"|“[^”\n]*”/g, '');
|
||||
};
|
||||
const sourceFrame = /(?:^|[.!?\n;]\s*)(?:(?:ELI10|Project\/branch\/task):\s*)?(?:(?:Source(?: excerpt| example)?|Quoted(?: source| example)?|Historical(?: example| assessment)?(?: only)?|Earlier(?: review)? assessment|Example|Hypothetical(?: example| assessment| scenario)?|If approved|If accepted)[,:.]|(?:The following|This assessment|This explanation)\b[^.\n]*\b(?:quoted|source|historical|hypothetical|example)\b|Historically,)/i;
|
||||
const lines = q.question.split('\n'), explanation = lines.findIndex(line => /^ELI10:/.test(line));
|
||||
const preface = lines.slice(0, explanation < 0 ? undefined : explanation + 1).join('\n');
|
||||
if (guardedDeclaration && /^(?:Project\/branch\/task|ELI10):\s*(?:Assuming|Provided)\b/im.test(currentProse(preface))) return [];
|
||||
if (/^\s*(?:```|~~~)/m.test(preface) || sourceFrame.test(currentProse(preface)) ||
|
||||
/\bnot (?:a )?current (?:finding|issue|defect)\b/i.test(currentProse(preface))) return [];
|
||||
const current = currentProse(q.question);
|
||||
const approval = /\b(?:if|once|when|unless) (?:approved|accepted)|\b(?:after|pending) approval\b/i;
|
||||
if (upgradeTransition) {
|
||||
if (/^ELI10:\s*>/.test(lines[explanation] ?? '')) return [];
|
||||
const namedCurrent = currentProse(q.question.replace(/`([A-Za-z_$][\w.$]*(?:\(\))?)`/g, '$1'));
|
||||
if (/(?:^|[.!?\n]\s*)(?:Correction:\s*)?Client\.evaluate\(\) (?:is (?:now|already|still)|now remains) (?:a |an )?(?:deprecated |compatibility )?alias\b/i.test(namedCurrent)) return [];
|
||||
const first = currentProse((lines[explanation] ?? '').replace(/`([A-Za-z_$][\w.$]*(?:\(\))?)`/g, '$1'))
|
||||
.replace(/^ELI10:\s*/, '').split(/(?<=[.!?])\s/)[0] ?? '';
|
||||
if (explanation < 1 || lines.filter(line => /^ELI10:/.test(line)).length !== 1 || approval.test(current) ||
|
||||
/\b(?:if|unless|assuming|provided|suppose|might|may|could|would|previously|earlier|historical|hypothetical|never|no longer|does not|do not|did not)\b/i.test(`${title} ${first}`) ||
|
||||
!/\brenames Client\.evaluate\(\) to Client\.run\(\)/i.test(first) ||
|
||||
!/\b(?:deletes|removes|drops) (?:the )?old (?:name|method)\b/i.test(first) ||
|
||||
!/\b(?:no |without (?:a )?)(?:compatibility )?alias\b/i.test(first)) return [];
|
||||
}
|
||||
if (reversedTuples && (approval.test(current) ||
|
||||
/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:these (?:functions|signatures)|run_eval and run_batch) (?:are (?:now|already) aligned|(?:now )?(?:use|take) the same (?:positional )?order)\b/i.test(current))) return [];
|
||||
const decision = guardedDeclaration && /^D\s*([1-9]\d*)\s*[—–:-]/i.exec(q.question);
|
||||
if (decision && new RegExp(`(?:^|[.!?\\n]\\s*)(?:Correction:\\s*)?D\\s*${decision[1]} (?:is|was|has been) (?:withdrawn|rejected|cancelled|canceled|superseded|(?:not|no longer) current)\\b`, 'i').test(current)) return [];
|
||||
if (guardedDeclaration && /(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:this|that|the) (?:finding|issue|gap|defect|explanation) (?:is|was|has been) (?:cancelled|canceled|superseded)\b/i.test(current)) return [];
|
||||
if (/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:(?:this|that|the) (?:finding|issue|gap|defect|explanation) (?:is|was|has been) (?:withdrawn|rejected|(?:already )?(?:fixed|resolved)|historical|(?:not|no longer) current|(?:a |only a )?source example)|there is no (?:current )?(?:finding|issue|gap|defect))\b/i.test(current)) return [];
|
||||
// A declaration's action evidence must belong to a current offered option,
|
||||
// rather than an example or an explicitly withdrawn correction.
|
||||
const action = (text: string) => currentProse(text.replace(/`([A-Za-z_$][\w.$/-]*(?:\([^`\n]*\))?)`/g, '$1'), true);
|
||||
signatureOptions = offered.filter(option => {
|
||||
const text = `${option.label}\n${option.description ?? ''}`, prose = currentProse(text, true);
|
||||
if (upgradeTransition && (approval.test(prose) || /(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:do not|don't|never) (?:keep|add|preserve|provide|retain) (?:the |a |an )?(?:compatibility )?alias\b/i.test(prose))) return false;
|
||||
if (reversedTuples && (approval.test(prose) ||
|
||||
/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:do not|don't|never) (?:align|unify|standardize|change|make|require) (?:either|both|the|these) (?:functions?|signatures?|arguments?)\b/i.test(prose))) return false;
|
||||
if (guardedDeclaration && /(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:this|that|the) (?:option|action|correction) (?:is|was|has been) (?:cancelled|canceled|superseded|withdrawn|rejected|(?:not|no longer) current)\b/i.test(prose)) return false;
|
||||
if (guardedDeclaration && /^(?:Assuming|Provided)\b/im.test(prose)) return false;
|
||||
if (decision && new RegExp(`(?:^|[.!?\\n]\\s*)(?:Correction:\\s*)?D\\s*${decision[1]} (?:is|was|has been) (?:withdrawn|rejected|cancelled|canceled|superseded|(?:not|no longer) current)\\b`, 'i').test(prose)) return false;
|
||||
return !/^(?:>|"|“)|^`[^`]*`$/.test(option.label.trim()) && !sourceFrame.test(prose) &&
|
||||
!/(?:^|[.!?\n]\s*)(?:Correction:\s*)?(?:this|the) (?:option|action|correction) (?:is|was|has been) (?:withdrawn|rejected|cancelled)\b/i.test(prose);
|
||||
});
|
||||
// Preserve inline tuple evidence for the stricter signature parser.
|
||||
offered = signatureOptions.map(option => ({ ...option, label: action(option.label), description: action(option.description ?? '') }));
|
||||
}
|
||||
const options = offered.map(o => `${o.label} ${o.description ?? ''}`);
|
||||
const ownUpgradeAlias = (option: string) => !(upgradeTransition || vanishingUpgrade) || (
|
||||
/(?<![\w.])(?:Client\.)?evaluate\(\) (?:stays|remains) (?:as )?(?:a |an )?(?:thin |deprecated |compatibility )?alias\b|\b(?:keep|retain|preserve) (?<![\w.])(?:Client\.)?evaluate\(\) as (?:a |an )?(?:deprecated |compatibility )?alias\b/i.test(option) &&
|
||||
!/\b(?:no |without (?:a )?)(?:compatibility )?alias\b|\b(?:do not|don't|never) (?:keep|retain|preserve) (?:Client\.)?evaluate\(\)/i.test(option));
|
||||
const labels = q.options.map(o => o.label.trim().replace(/\s*\(recommended\)$/i, '').toLowerCase());
|
||||
const yesNo = labels.length === 2 && labels.includes('yes') && labels.includes('no');
|
||||
// A terse Yes/No panel still resolves an action explicitly asked in the
|
||||
// main question; action words in background prose never supply this arm.
|
||||
const directAction = (verbs: string) => yesNo && new RegExp(
|
||||
`^(?:should|shall|can|do|would) (?:we|I) (?:${verbs})\\b[^?]*\\?$`, 'i').test(title);
|
||||
|
||||
const found: DevexSeededGap[] = [];
|
||||
if (/\bCI\b/i.test(title) && /\b(?:local|demo|first)\b/i.test(title) &&
|
||||
/\b(?:gate|check|blocks?|waits?|bypass|mandatory|required)\b/i.test(title) &&
|
||||
(options.some(o => /\b(?:no CI gate|remove|move|skip|bypass|gate)\b/i.test(o) && /\b(?:CI|check|gate|local|demo)\b/i.test(o)) || directAction('remove|move|skip|bypass|gate'))) found.push('local-ci-gate');
|
||||
if (/\bquickstart\b|examples\/first_eval\.py/i.test(title) &&
|
||||
/\b(?:README|file|example|demo|missing|absent|package|wheel|ship|point)\b|first_eval\.py/i.test(title) &&
|
||||
(!declaration || absentReference || /\b(?:not (?:shipped|included|available|present)|missing|absent|nonexistent|does not exist)\b/i.test(title)) &&
|
||||
(options.some(o => /\b(?:point|ship|add|demo is)\b/i.test(o) && /\bquickstart\b|first_eval\.py/i.test(o)) || directAction('point|ship|add|replace|fix'))) found.push('missing-quickstart');
|
||||
if (explainedReversedSignatures({ ...q, options: signatureOptions }, title) || (/\brun_eval\b/i.test(title) && /\brun_batch\b/i.test(title) &&
|
||||
/\b(?:arguments?|order|positional|reversed|opposite|consistent|align|unify|dataset|evaluator)\b/i.test(title) &&
|
||||
(!declaration || reversedTuples || /\b(?:reversed|opposite|swapped|inconsistent)\b/i.test(title)) &&
|
||||
(options.some(o => (!reversedTuples || /\bboth functions\b|\brun_eval\b[^\n]*\brun_batch\b/i.test(o)) &&
|
||||
/\b(?:align|unify|standardize|keyword|swap)\b/i.test(o) && /\b(?:order|dataset|arguments?|positional)\b/i.test(o)) || directAction('align|unify|standardize|enforce|make')))) found.push('reversed-arguments');
|
||||
if ((opaqueAuthentication || /\bAuthError\b|\binvalid API key\b/i.test(title)) &&
|
||||
/\b(?:error|message|code|cause|fix|guidance|opaque|explain)\b|request failed/i.test(title) &&
|
||||
(!declaration || opaqueAuthentication || /\b(?:no (?:cause|fix|explanation|code)|opaque)\b|request failed/i.test(title)) &&
|
||||
(options.some(o => (!opaqueAuthentication || /\bAuthError\b/i.test(o)) && (/\bcodes?\b/i.test(o) || /^(?:[A-D]\)\s*)?Coded\b/i.test(o)) && /\b(?:cause|fix|link)\b/i.test(o)) || directAction('add|include|explain|replace|report|give'))) found.push('opaque-auth-error');
|
||||
if (/Client\.evaluate\b/i.test(title) &&
|
||||
/Client\.run\b|\b(?:v\d+|version \d+|alias|deprecation|migration)\b/i.test(title) &&
|
||||
(upgradeVocabulary || upgradeTransition) &&
|
||||
(!declaration || /\b(?:no |without (?:a )?)(?:compatibility )?(?:alias|warning|migration (?:guide|path))\b/i.test(title)) &&
|
||||
(options.some(o => ownUpgradeAlias(o) && /\balias\b/i.test(o) && /\b(?:warning|DeprecationWarning|migration)\b/i.test(o)) || directAction('keep|add|preserve|provide|retain'))) found.push('breaking-upgrade');
|
||||
return found;
|
||||
}
|
||||
|
||||
/** Extra real decisions are permitted; each seeded gap needs its own completed native call. */
|
||||
export function devexSeedCoverage(transcript: PlanCountTranscript) {
|
||||
const decisions = Object.fromEntries(DEVEX_SEEDED_GAPS.map(gap => [gap, []])) as Record<DevexSeededGap, string[]>;
|
||||
const batched: string[] = [];
|
||||
const invalid: string[] = [];
|
||||
const sessions = new Set(transcript.calls.map(c => c.sessionId));
|
||||
if (transcript.status !== 'ready' || sessions.size !== 1 || sessions.has('')) invalid.push('missing or mixed native session');
|
||||
const ids = new Set<string>();
|
||||
for (const call of transcript.calls) {
|
||||
const id = `${call.sessionId}:${call.toolUseId}`;
|
||||
if (!call.toolUseId || ids.has(id)) { invalid.push(`missing or repeated native call: ${id}`); continue; }
|
||||
ids.add(id);
|
||||
const gaps = call.questions.flatMap(decisionGaps);
|
||||
if (!gaps.length) continue;
|
||||
if (call.questions.length !== 1 || gaps.length !== 1 || call.questions[0]!.multiSelect) {
|
||||
batched.push(id); continue;
|
||||
}
|
||||
const q = call.questions[0]!;
|
||||
const labels = q.options.map(o => o.label);
|
||||
const complete = call.answered === true && call.failed === false &&
|
||||
Array.isArray(call.unansweredQuestionIndices) && call.unansweredQuestionIndices.length === 0 &&
|
||||
Number.isFinite(Date.parse(call.answeredAt ?? '')) && q.options.length >= 2 && q.options.length <= 4 &&
|
||||
new Set(labels).size === labels.length && labels.every(Boolean) &&
|
||||
Object.keys(call.answers ?? {}).length === 1 && labels.includes(call.answers?.[q.question] ?? '');
|
||||
if (complete) decisions[gaps[0]!]!.push(id);
|
||||
}
|
||||
const missing = DEVEX_SEEDED_GAPS.filter(gap => decisions[gap].length === 0);
|
||||
const matchedIds = new Set(Object.values(decisions).flat());
|
||||
return {
|
||||
complete: invalid.length === 0 && batched.length === 0 && missing.length === 0 && matchedIds.size >= DEVEX_SEEDED_GAPS.length,
|
||||
missing, decisions, batched, invalid,
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
/** Isolated real-parent fixture and execution oracle for the plan-review off switch. */
|
||||
import { mkdirSync, readFileSync, symlinkSync, unlinkSync, writeFileSync } from 'node:fs';
|
||||
import { delimiter, join } from 'node:path';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import { extractSkillSections } from './skill-fixture';
|
||||
import { claudeOutsideExecutions } from './outside-voice-evidence';
|
||||
|
||||
export const OUTSIDE_PLAN_SECTION = 'Outside Voice — Independent Plan Challenge (default-on)';
|
||||
|
||||
/** Extract generated instructions; runtime paths are the only content substitution. */
|
||||
export function installDisabledPlanReviewFixture(rendered: string, repo: string, runtimeRoot: string) {
|
||||
const source = join(rendered, 'plan-eng-review');
|
||||
const main = readFileSync(join(source, 'SKILL.md'), 'utf8');
|
||||
const frontmatter = main.match(/^---\r?\n[\s\S]*?\r?\n---\r?\n/)?.[0];
|
||||
if (!frontmatter) throw new Error('Generated plan-eng-review has no frontmatter');
|
||||
// The shared plan challenge is lazily loaded. Give the established,
|
||||
// fence-aware extractor its frontmatter without copying any workflow prose.
|
||||
const input = join(repo, 'outside-plan-source.md');
|
||||
writeFileSync(input, frontmatter + readFileSync(join(source, 'sections/review-sections.md'), 'utf8'));
|
||||
let generated: string;
|
||||
try { generated = extractSkillSections(input, [OUTSIDE_PLAN_SECTION]); }
|
||||
finally { unlinkSync(input); }
|
||||
symlinkSync(runtimeRoot, join(repo, 'runtime'), 'dir');
|
||||
const instructions = generated
|
||||
.replaceAll('$HOME/.claude/skills/gstack', '$PWD/runtime')
|
||||
.replaceAll('~/.claude/skills/gstack', './runtime');
|
||||
const workflowPath = join(repo, 'OUTSIDE-PLAN.md');
|
||||
writeFileSync(workflowPath, instructions);
|
||||
|
||||
const stateDir = join(repo, 'gstack-state');
|
||||
const configDir = join(repo, 'claude-config');
|
||||
const spyDir = join(repo, 'cli-bin');
|
||||
const cliDispatchLog = join(repo, 'outside-cli-dispatch.log');
|
||||
for (const dir of [stateDir, configDir, spyDir]) mkdirSync(dir, { recursive: true });
|
||||
// A forbidden invocation is observable without buying another model call.
|
||||
// No prompt/credentials are recorded; even --version/auth probes count.
|
||||
writeFileSync(join(spyDir, 'codex'), '#!/bin/sh\nprintf "codex invoked\\n" >> "$GSTACK_DISABLED_CLI_LOG"\nexit 73\n', { mode: 0o755 });
|
||||
const env = {
|
||||
PATH: `${spyDir}${delimiter}${process.env.PATH ?? ''}`,
|
||||
CLAUDE_CONFIG_DIR: configDir,
|
||||
GSTACK_HOME: stateDir,
|
||||
GSTACK_STATE_ROOT: stateDir,
|
||||
GSTACK_DISABLED_CLI_LOG: cliDispatchLog,
|
||||
GSTACK_ACTIVE_HOST: 'claude',
|
||||
GSTACK_PROJECT_SLUG: 'disabled-plan-fixture',
|
||||
};
|
||||
const config = spawnSync(join(runtimeRoot, 'bin/gstack-config'), ['set', 'codex_reviews', 'disabled'], {
|
||||
cwd: repo, env: { ...process.env, ...env }, encoding: 'utf8', timeout: 5_000,
|
||||
});
|
||||
if (config.status !== 0) throw new Error(`Cannot seed isolated review control: ${config.stderr}`);
|
||||
// Seed real historical coverage so the new disabled record must replace it
|
||||
// in the dashboard's latest-record view, not merely appear in final prose.
|
||||
const prior = {
|
||||
skill: 'codex-plan-review', timestamp: new Date(Date.now() - 60_000).toISOString(),
|
||||
status: 'clean', source: 'codex', host: 'claude', outside_provider: 'codex',
|
||||
outside_status: 'completed', phase: 'plan-review',
|
||||
};
|
||||
const logged = spawnSync(join(runtimeRoot, 'bin/gstack-review-log'), [JSON.stringify(prior)], {
|
||||
cwd: repo, env: { ...process.env, ...env }, encoding: 'utf8', timeout: 5_000,
|
||||
});
|
||||
if (logged.status !== 0) throw new Error(`Cannot seed historical review: ${logged.stderr}`);
|
||||
const slug = spawnSync(join(runtimeRoot, 'bin/gstack-slug'), [], {
|
||||
cwd: repo, env: { ...process.env, ...env }, encoding: 'utf8', timeout: 5_000,
|
||||
});
|
||||
const branch = /^BRANCH=([a-zA-Z0-9._-]+)$/m.exec(slug.stdout)?.[1];
|
||||
if (slug.status !== 0 || !branch) throw new Error('Cannot resolve isolated review-log branch');
|
||||
const reviewLogPath = join(stateDir, 'projects', env.GSTACK_PROJECT_SLUG, `${branch}-reviews.jsonl`);
|
||||
const priorRecord = JSON.parse(readFileSync(reviewLogPath, 'utf8').trim());
|
||||
return { workflowPath, instructions, generated, stateDir, cliDispatchLog, reviewLogPath, priorRecord, env };
|
||||
}
|
||||
|
||||
/** A dated log value belongs to its record, not to the workflow quoting it. */
|
||||
function preRunLogRecordValue(before: string, nextClause: string): boolean {
|
||||
const owner = /^(?:the\s+)?review\s+log\s+(?:already\s+)?(?:held|contained)\s+an?\s+(?:record|entry|line)\b/i.exec(before.trim());
|
||||
if (!owner) return false;
|
||||
const value = before.trim().slice(owner[0].length);
|
||||
// This route requires both an explicit pre-run date and a reported review
|
||||
// value. Arbitrary intervening prose cannot switch the reporting subject.
|
||||
const datedValue = /^\s*,?\s*(?:timestamped|recorded|written)\s+(?:about\s+)?(?:a|an|one|two|\d+)\s+(?:minute|hour|day|week)s?\s+before\s+(?:this|my)\s+(?:run|session|workflow)\s*,?\s+(?:marking|recording)\s+an?\s+(?:clean|successful|completed)\s+(?:[a-z][a-z-]*\s+)?(?:plan\s+)?review\s+with\s*$/i;
|
||||
if (!datedValue.test(value)) return false;
|
||||
// The immediately following assertion must keep the same record as its
|
||||
// subject and explicitly exclude this workflow as its origin. A later
|
||||
// current completion occurrence is still checked independently below.
|
||||
return /^(?:that|the|this)\s+(?:record|entry|line)\s+(?:was|is)\s+not\s+(?:produced|created|written|recorded)\s+(?:by|during|in)\s+(?:this|my)\s+(?:run|session|workflow)\b/i.test(nextClause.trim()) ||
|
||||
/^(?:that|the|this)\s+(?:record|entry|line)\s+predates\s+(?:this|my)\s+(?:run|session|workflow)\s+and\s+was\s+not\s+(?:produced|created|written|recorded)\s+by\s+it\b/i.test(nextClause.trim()) ||
|
||||
/^(?:that|the|this)\s+(?:record|entry|line)\s+(?:does not|doesn't|cannot)\s+(?:reflect|establish|provide|supply)\s+(?:current\s+)?outside\s+(?:review\s+)?coverage\s+(?:from|for)\s+(?:this|my)\s+(?:run|session|workflow)\b/i.test(nextClause.trim());
|
||||
}
|
||||
|
||||
/** Check each completion mention; a historical record cannot excuse a later current claim. */
|
||||
function hasUnattributedOutsideCompletion(output: string): boolean {
|
||||
const marker = /\bboth reviewers agree\b|["']?\boutside_status["']*\s*[:=]\s*["']*completed\b/gi;
|
||||
const clauses = output.replace(/[*`]/g, '').split(/\r?\n|(?<=[.!?;])\s+|\b(?:but|however|nevertheless|yet)\b[:,]?\s*/i);
|
||||
return clauses.some((clause, clauseIndex) => [...clause.matchAll(marker)].some(match => {
|
||||
const before = clause.slice(0, match.index).trimEnd();
|
||||
// A quoted phrase is not automatically a denial. Require the local no-claim
|
||||
// statement, so a second positive assertion in the same paragraph still fails.
|
||||
const denied = /\bno\s*["'“”‘’]*\s*$/i.test(before)
|
||||
|| /\b(?:(?:do|did|will|would|can|could)\s+not|cannot|can't|won't)\s+(?:claim|say|state|report)\s*["'“”‘’]*\s*$/i.test(before)
|
||||
|| /\b(?:am|is|are)\s+not\s+(?:claiming|saying|stating|reporting)\s*["'“”‘’]*\s*$/i.test(before);
|
||||
if (denied) return false;
|
||||
// Agreement is a current prose claim unless explicitly denied; an old
|
||||
// log entry only establishes the provenance of its recorded status value.
|
||||
if (/^both reviewers agree$/i.test(match[0])) return true;
|
||||
// A record explicitly dated before this run is historical even when its
|
||||
// subject is "that record" rather than "the earlier record".
|
||||
const datedBeforeRun = String.raw`\s+is\s+timestamped\s+(?:about\s+)?(?:a|an|one|two|\d+)\s+(?:minute|hour|day|week)s?\s+before\s+(?:this|my)\s+(?:run|session|workflow)`;
|
||||
const recordPattern = new RegExp(String.raw`\b(?:(?:earlier|prior|historical|old(?:er)?)\s+(?:entry|record|line)|(?:that|the)\s+(?:entry|record|line)(?=${datedBeforeRun}))\b`, 'gi');
|
||||
const record = [...before.matchAll(recordPattern)].at(-1);
|
||||
if (!record) return !preRunLogRecordValue(before, clauses[clauseIndex + 1] ?? '');
|
||||
// Bind this occurrence to an old record's reported value. A mere mention
|
||||
// of a record, a second status, or a new reporting subject cannot inherit
|
||||
// its historical attribution, even without a sentence boundary.
|
||||
const prefix = before.slice(record.index + record[0].length);
|
||||
// Date metadata still describes this record's own value. Admit explicit
|
||||
// clock or pre-run timestamps, not arbitrary prose that can change subjects.
|
||||
const clock = String.raw`\s+from\s+(?:[01]\d|2[0-3]):[0-5]\d,?`;
|
||||
const beforeRun = String.raw`\s*,?\s*(?:written|recorded)\s+(?:about\s+)?(?:a|an|one|\d+)\s+(?:minute|hour|day|week)s?\s+before\s+this\s+(?:run|session|workflow)(?:\s+(?:started|began))?,?`;
|
||||
const timestamp = String.raw`(?:${datedBeforeRun}\s+and|\s*,?\s*timestamped\b[^,;.!?]{1,160},?|${clock}|${beforeRun})`;
|
||||
const report = new RegExp(String.raw`^(?:${timestamp})?\s*(?:(?:that\s+)?(?:claims?|claiming|shows?|showed|says?|said|records?|recorded|reported)\b|:)\s*`, 'i').exec(prefix);
|
||||
// Only intervening review-log metadata belongs to this reported value.
|
||||
// Arbitrary prose could switch to a new subject without an earlier status.
|
||||
const field = String.raw`["']?(?:status|source|host|outside_provider|phase|timestamp)["']?\s*[:=]\s*["']?[a-z0-9_.:+-]+["']?`;
|
||||
const metadata = new RegExp(String.raw`^(?:${field}\s*(?:,\s*|(?:with|and)\s*))*$`, 'i');
|
||||
const reportsOldValue = report !== null && metadata.test(prefix.slice(report[0].length).trim());
|
||||
const attribution = clause.slice(record.index).replace(/\bbefore\s+(?:this|my)\s+(?:run|session|workflow)\b/gi, 'beforehand');
|
||||
const current = /\b(?:now|currently|current|today|new|updat\w*|append\w*|chang\w*|mark\w*|set|write|wrote)\b|\bthis\s+(?:run|session|workflow)\b/i.test(attribution);
|
||||
return !reportsOldValue || current;
|
||||
}));
|
||||
}
|
||||
|
||||
export function disabledPlanReviewEvidence(result: {
|
||||
exitReason: string; output: string; transcript: any[];
|
||||
}, cliDispatchLog: string, reviewLog = '', priorRecord?: Record<string, unknown>) {
|
||||
const init = result.transcript.find(event => event?.type === 'system' && event.subtype === 'init');
|
||||
const terminal = result.transcript.filter(event => event?.type === 'result').at(-1);
|
||||
const toolCalls = result.transcript.flatMap(event => event?.type === 'assistant' && Array.isArray(event.message?.content)
|
||||
? event.message.content.filter((block: any) => block.type === 'tool_use') : []);
|
||||
const fallbackCalls = toolCalls.filter((call: any) => call.name === 'Agent' || call.name === 'Task');
|
||||
// Command text may contain an unexecuted guarded branch. Keep it for
|
||||
// diagnostics; only the executable sentinel establishes CLI dispatch.
|
||||
const outsideCommandMentions = toolCalls.filter((call: any) => call.name === 'Bash'
|
||||
&& /\bcodex\s+(?:exec|review)\b|\bgstack-claude-code(?:['"])?\s+--/.test(String(call.input?.command ?? '')));
|
||||
const disabledPreflight = claudeOutsideExecutions(result.transcript).filter(call => call.succeeded
|
||||
&& /gstack-config\b/.test(call.command) && /codex_reviews\b/.test(call.command)
|
||||
&& /^CODEX_MODE: disabled\s*$/m.test(call.output));
|
||||
const completed = result.exitReason === 'success' && terminal?.subtype === 'success'
|
||||
&& terminal.is_error !== true && typeof terminal.result === 'string' && terminal.result.trim().length > 0;
|
||||
const disabledAttribution = /outside_status["'`*]*\s*[:=]\s*["'`*]*disabled\b/i.test(result.output)
|
||||
|| result.output.split('\n').some(line => /\b(?:outside(?:\s+(?:voice|review))?|codex(?:\s+review)?)\b/i.test(line)
|
||||
&& /\bdisabled\b/i.test(line) && !/\bnot\s+disabled\b/i.test(line));
|
||||
const falseCompletion = hasUnattributedOutsideCompletion(result.output);
|
||||
// Native CLI releases expose the requested subagent as Agent or Task.
|
||||
// Availability never permits dispatch: fallbackCalls rejects both names.
|
||||
const agentAvailable = Array.isArray(init?.tools) && init.tools.some((tool: unknown) => tool === 'Agent' || tool === 'Task');
|
||||
let records: any[] = [];
|
||||
let malformedLog = false;
|
||||
try { records = reviewLog.trim() ? reviewLog.trim().split('\n').map(line => JSON.parse(line)) : []; }
|
||||
catch { malformedLog = true; }
|
||||
const planRecords = records.filter(record => record && typeof record === 'object' && record.skill === 'codex-plan-review');
|
||||
const priorIndex = priorRecord ? planRecords.findIndex(record => JSON.stringify(record) === JSON.stringify(priorRecord)) : -1;
|
||||
const latestRecord = planRecords.at(-1);
|
||||
const persistedDisabled = !malformedLog && priorIndex >= 0 && priorIndex < planRecords.length - 1
|
||||
&& priorRecord?.outside_status === 'completed' && latestRecord.status === 'skipped'
|
||||
&& latestRecord.source === 'none' && latestRecord.host === 'claude'
|
||||
&& latestRecord.outside_provider === 'codex' && latestRecord.outside_status === 'disabled'
|
||||
&& latestRecord.phase === 'plan-review' && typeof latestRecord.timestamp === 'string'
|
||||
&& typeof priorRecord.timestamp === 'string' && Number.isFinite(Date.parse(latestRecord.timestamp))
|
||||
&& Date.parse(latestRecord.timestamp) > Date.parse(priorRecord.timestamp);
|
||||
return {
|
||||
passed: completed && agentAvailable && disabledPreflight.length > 0 && fallbackCalls.length === 0
|
||||
&& cliDispatchLog.trim() === '' && disabledAttribution && !falseCompletion && persistedDisabled,
|
||||
completed, agentAvailable, disabledAttribution, falseCompletion, persistedDisabled, latestRecord, malformedLog,
|
||||
disabledPreflight, fallbackCalls, outsideCommandMentions, cliDispatchLog,
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,57 @@
|
||||
import type { NativePlanQuestionCall } from './plan-count-transcript';
|
||||
|
||||
/** A selected manual exit cannot modify the decisions already in the report.
|
||||
* Inspect the current recap and chosen action; unchosen follow-up reviews may
|
||||
* still have gates to clear. Report freshness and Exit ownership stay with the
|
||||
* caller, so this recognition alone never establishes review completion.
|
||||
*/
|
||||
export function isRecordedDxManualNavigation(call: NativePlanQuestionCall): boolean {
|
||||
if (!call.answered || call.failed !== false || !call.sessionId || !call.toolUseId ||
|
||||
!Number.isFinite(Date.parse(call.answeredAt ?? '')) || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
Object.keys(call.answers ?? {}).length !== 1) return false;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.header.trim() !== 'Next steps' || q.options.length !== 3 ||
|
||||
new Set(q.options.map(o => o.label)).size !== 3 || /```|~~~|<gstack-qid/i.test(q.question) ||
|
||||
(q.question.match(/\?/g)?.length ?? 0) !== 1) return false;
|
||||
const label = (s: string) => s.trim().replace(/\s*\(recommended\)$/i, '');
|
||||
const manual = (s: string) => /^Skip, handle (?:next steps )?manually$/i.test(label(s));
|
||||
const eng = (s: string) => /^Run \/plan-eng-review(?: next)?$/i.test(label(s));
|
||||
const implement = (s: string) => /^Ready to implement$/i.test(label(s));
|
||||
if (q.options.filter(o => manual(o.label)).length !== 1 ||
|
||||
q.options.filter(o => eng(o.label)).length !== 1 ||
|
||||
q.options.filter(o => implement(o.label)).length !== 1) return false;
|
||||
const selected = q.options.find(o => o.label === call.answers?.[q.question]);
|
||||
if (!selected || !manual(selected.label) ||
|
||||
!/^(?:End|Finish|Stop) the DX review here[.;]\s*you (?:will )?(?:run|handle) (?:subsequent|later) reviews (?:yourself|manually)\.$/i
|
||||
.test(selected.description?.trim() ?? '')) return false;
|
||||
|
||||
// Native options, not the prose letters, identify the chosen action. Separate
|
||||
// the offered pros/cons from assertions about the current review, retaining
|
||||
// the final Net paragraph so it cannot hide an additional plan instruction.
|
||||
const parts = q.question.split(/\nPros \/ cons:\s*\n/);
|
||||
if (parts.length !== 2) return false;
|
||||
const net = /(?:^|\n)Net:\s*([^]*)$/.exec(parts[1]!);
|
||||
if (!net) return false;
|
||||
const current = `${parts[0]}\nNet: ${net[1]}`;
|
||||
const lines = current.split(/\r?\n/).filter(line =>
|
||||
!/^\s*>/.test(line) && !/^\s*(["'`]).*\1\s*$/.test(line));
|
||||
const prose = lines.join('\n');
|
||||
if (!/^(?:D[1-9]\d*\s*[—–-]\s*)?DX review (?:is )?complete\. What(?:['’]s)? next\?\n/i.test(prose)) return false;
|
||||
const contexts = [...prose.matchAll(/^Project\/branch\/task:\s*([^\n]+)$/gim)];
|
||||
const recaps = [...prose.matchAll(/^ELI10:\s*([^\n]+)$/gim)];
|
||||
if (contexts.length !== 1 || recaps.length !== 1) return false;
|
||||
const context = contexts[0]![1]!, recap = recaps[0]![1]!;
|
||||
// Only the single current context may introduce the recap. An intervening
|
||||
// source heading or hypothetical introduction changes its authority.
|
||||
const introduction = prose.slice(0, recaps[0]!.index).split('\n').slice(1).filter(line => line.trim());
|
||||
if (introduction.length !== 1 || !/^Project\/branch\/task:/.test(introduction[0]!) ||
|
||||
/^(?:source|example|historical|previous|earlier|if|unless|when|once|after|provided|proposed|optional)\b|^(?:for historical context|for example|from (?:a )?source excerpt)\b|^["'“‘`]/i.test(context) ||
|
||||
!/\/plan-devex-review\b[^.!?\n]*\bis finished; plan written to [^\s;]+\.md\./i.test(context) ||
|
||||
!/^The DX review is done:\s/i.test(recap) ||
|
||||
!/\bdecisions recorded\b/i.test(recap)) return false;
|
||||
return !/(?:^|\n|[.!;]\s+)(?:Source|Example|Historical(?: review)?|Previously|Earlier review(?: assessment)?):/i.test(prose) &&
|
||||
!/\b(?:DX review|DX findings?|DX decisions?|DX tasks?|plan|report|handoff)\b[^.!?\n]{0,70}\b(?:unresolved|outstanding|pending|remaining|withdrawn|retracted|superseded|cancelled|canceled|not current)\b/i.test(prose) &&
|
||||
!/\b(?:not|never)\s+(?:all\s+)?(?:done|complete|completed|recorded|written|resolved)\b|\b(?:done|complete|completed|recorded|written|resolved)\s+(?:only\s+)?(?:after|if|when|once|unless|until)\b/i.test(prose) &&
|
||||
!/(?:^|[.!?;]\s+|\n|\b(?:should|must|need to|will)\s+)(?:(?:we|you|please|first|then|also)\s+)*(?:add|fix|edit|update|rewrite|remove|implement|resolve|decide|change|approve|start|run)\b/im.test(prose);
|
||||
}
|
||||
@@ -0,0 +1,109 @@
|
||||
import type { NativePlanQuestion } from './plan-count-transcript';
|
||||
|
||||
/** A current cache-writer decision may name its actors in the plan context.
|
||||
* Called only after engNumberedFindingAUQ validates completed native metadata. */
|
||||
export function engCacheWriterDecision(q: NativePlanQuestion): boolean {
|
||||
const lines = q.question.split('\n');
|
||||
// Severity, confidence and source citations annotate an owned issue; they
|
||||
// never replace its current defect, assessment or opposed choices.
|
||||
const annotated = /^D([1-9]\d*) [—–:-] Issue ([1-9]\d*) \[P[0-3]\] \(confidence (?:10|[1-9])\/10\) [A-Za-z][\w./-]*:[1-9]\d*(?:-[1-9]\d*)?(?: \+ :[1-9]\d*(?:-[1-9]\d*)?)? [—–:-] ([A-Za-z_$][\w$]*) and ([A-Za-z_$][\w$]*) both mutate one module-level ([A-Za-z_$][\w$]*) that does not serialize mutations\. How should the shared cache be wired\?$/.exec(lines[0] ?? '');
|
||||
const architecture = /^D([1-9]\d*) [—–:-] Architecture issue ([1-9]\d*): global mutable ([A-Za-z_$][\w$]*) shared by two services with unserialized mutations\.?$/.exec(lines[0] ?? '');
|
||||
const declared = annotated || architecture;
|
||||
const ordinal = declared?.[1] ?? /^D([1-9]\d*) [—–:-] Who is allowed to write to the auth cache\?$/.exec(lines[0] ?? '')?.[1];
|
||||
if (!ordinal || (declared ? q.header !== `Arch ${declared[2]}` : q.header !== 'Cache writes')) return false;
|
||||
const context = declared ? /^Project\/branch\/task: (\S[^\n]*)\.$/.exec(lines[1] ?? '')
|
||||
: /^Project\/branch\/task: (\S[^\n]*), ([A-Za-z_$][\w$]*) and ([A-Za-z_$][\w$]*) both mutating one backing cache \(([\w./-]+\.md):\d+(?:, \d+(?:[-–]\d+)?)?\)\.$/.exec(lines[1] ?? '');
|
||||
const declaredWriters = architecture && /^ELI10: [A-Za-z][\w./-]*\.md:[1-9]\d*(?:-[1-9]\d*)? has ([A-Za-z_$][\w$]*) and ([A-Za-z_$][\w$]*) both import one module-level ([A-Za-z_$][\w$]*) and both write to it, and [A-Za-z][\w./-]*\.md:[1-9]\d* says nothing serializes those writes\./.exec(lines[2] ?? '');
|
||||
const actors = architecture ? [declaredWriters?.[1] ?? '', declaredWriters?.[2] ?? ''] : annotated ? [annotated[3]!, annotated[4]!] : [context?.[2] ?? '', context?.[3] ?? ''];
|
||||
const assessment = architecture ? declaredWriters && declaredWriters[3] === architecture[3] : annotated
|
||||
? /^ELI10: Two services share one global cache object exported from a module, and both write to it\. /.test(lines[2] ?? '')
|
||||
: /^ELI10: Two services write to the same cache and nothing orders their writes\. /.test(lines[2] ?? '');
|
||||
if (!context || actors[0] === actors[1] || !assessment ||
|
||||
lines.filter(line => /^Project\/branch\/task:/.test(line)).length !== 1 ||
|
||||
lines.filter(line => /^ELI10:/.test(line)).length !== 1) return false;
|
||||
const boundary = '(?:^|[.!?;]\\s+|\\n|[✅❌]\\s*)(?:Correction:\\s*)?';
|
||||
const owner = `(?:(?:this|the|that) (?:finding|issue|gap|assessment|option|action|remedy|race|single-writer requirement)|D\\s*${ordinal}${declared ? `|${architecture ? '(?:Architecture )?' : ''}Issue ${declared[2]}` : ''})`;
|
||||
const status = `(?:withdrawn|superseded|rejected|cancelled|canceled|resolved|closed|hypothetical|not current|no longer current${declared ? '|optional|unproven|proposed|conditional on approval' : ''})`;
|
||||
const current = (text: string) => (declared ? text.replace(/\*\*/g, '').replace(/\((?:human|CC):[^)\n]*\)[ \t]+(?=(?:Correction:\s*)?(?:this|the|that) (?:option|action|remedy)\b)/gi, '$&. ') : text)
|
||||
.replace(/```[\s\S]*?(?:```|$)|~~~[\s\S]*?(?:~~~|$)/g, '')
|
||||
.replace(/^(?:\s*>| {4}|\t).*$/gm, '')
|
||||
.replace(new RegExp(`(${boundary}${owner} (?:is|was|has been) )["“'‘\x60](${status})["”'’\x60]`, 'gim'), '$1$2')
|
||||
.replace(/"[^"\n]*"|“[^”\n]*”|(?<!\w)'[^'\n]*'(?!\w)|‘[^’\n]*’/g, '')
|
||||
.replace(/`([^`\n]*)`/g, (_, value: string) => /^[A-Za-z_$][\w$]*$/.test(value) ? value : '');
|
||||
const framed = new RegExp(`${boundary}(?:Source(?: excerpt| example)?|Quoted(?: source| example)?|Historical(?: assessment| example)?|Hypothetical(?: scenario| example)?|Earlier review)(?:[.:,]|\\s)|${boundary}(?:if|unless|when|assuming|provided|suppose|imagine)\\b`, 'i');
|
||||
const closed = new RegExp(`${boundary}${owner} (?:is|was|has been) ${status}\\b|${boundary}(?:there is )?no current (?:gap|risk|finding|race) (?:remains|exists)\\b`, 'i');
|
||||
// The first assertion is current; its following revoked-token scenario is
|
||||
// a causal explanation, not a condition on whether this review occurs.
|
||||
if (/\b(?:if|when|once|unless) (?:approved|accepted)|\b(?:after|pending) approval\b/i.test(current(context[1]!)) || framed.test(current(context[1]!)) || framed.test(current(lines.slice(0, 3).join('\n').split('ELI10:')[0]!)) ||
|
||||
closed.test(current(q.question))) return false;
|
||||
const escape = (value: string) => value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
||||
const contradiction = new RegExp(`${boundary}(?:(?:the|these|both) services (?:no longer (?:writes?|mutates?)|now serialize)|(?:${actors.map(escape).join('|')}) (?:no longer|does not) (?:writes?|mutates?)|(?:the |this )?(?:auth )?cache (?:is no longer shared|now serializes)|(?:the )?(?:race is (?:resolved|closed)|(?:writes|writers) are (?:now )?(?:ordered|serialized)))\\b`, 'i');
|
||||
if (contradiction.test(current(q.question))) return false;
|
||||
if (architecture) {
|
||||
const cache = escape(architecture[3]!);
|
||||
const text = current(q.question);
|
||||
const namedResolved = new RegExp(`${boundary}${cache} (?:is (?:now |already )?(?:serialized|ordered)|now serializes|no longer (?:shares|has) (?:mutable )?state)\\b`, 'i');
|
||||
const onlyWriter = new RegExp(`${boundary}only (?:${actors.map(escape).join('|')}) writes\\b`, 'i');
|
||||
const conditional = new RegExp(`\\b(?:if|when|once|unless) (?:approved|accepted)|\\b(?:after|pending) approval\\b|${boundary}${owner} (?:requires (?:approval|acceptance)|is (?:conditional|contingent|dependent) on (?:approval|acceptance))\\b`, 'i');
|
||||
const cancelled = new RegExp(`${boundary}(?:do not|don't|never|skip|cancel|withdraw) (?:inject|use|keep|accept|adopt|choose|proceed|reject|write|serialize|document)\\b`, 'i');
|
||||
if (framed.test(text) || conditional.test(text) || cancelled.test(text) || namedResolved.test(text) || onlyWriter.test(text)) return false;
|
||||
const race = /Picture ([A-Za-z_$][\w$]*) invalidating tenant ([A-Za-z][\w-]*) on suspension at the same instant ([A-Za-z_$][\w$]*) writes a freshly refreshed token for tenant ([A-Za-z][\w-]*)\. Last writer wins, the suspended tenant keeps a valid session, and nothing logs it\./.exec(current(lines[2] ?? ''));
|
||||
if (!race || race[1] === race[3] || !actors.includes(race[1]!) || !actors.includes(race[3]!) || race[2] !== race[4]) return false;
|
||||
const rows = q.options.map(option => ({ id: new RegExp(`^${architecture[2]}([A-D]): (.+?)(?: \\(recommended\\))?$`).exec(option.label), text: current(option.description ?? '').trim() }));
|
||||
if (rows.some(row => !row.id || framed.test(row.text) || closed.test(row.text) || conditional.test(row.text) || cancelled.test(row.text)) || new Set(rows.map(row => row.id![1])).size !== rows.length) return false;
|
||||
const remedy = rows.find(row => new RegExp(`^Inject ${cache} by constructor; ${cache} owns all writes and serializes per tenant key; every method requires tenant context$`).test(row.id![2]!));
|
||||
const unchanged = rows.find(row => row.id![2] === 'Do nothing; document that mutations are unserialized');
|
||||
if (!remedy || !unchanged ||
|
||||
!/^✅\s*Invalidation can never be overwritten by a concurrent refresh: writes for one tenant key run in order through one owner\./.test(remedy.text) ||
|
||||
!new RegExp(`✅\\s*Tests construct a fresh ${cache} per case, so no cross-test state leaks\\.`).test(remedy.text) ||
|
||||
!/✅\s*No method accepts a call without a tenant ID, so tenant isolation is enforced at the type boundary\./.test(remedy.text) ||
|
||||
!/❌\s*Leaves a silent security failure mode \(suspended tenant stays valid\) in a multi-tenant auth path\./.test(unchanged.text) ||
|
||||
contradiction.test(unchanged.text) || namedResolved.test(unchanged.text) || onlyWriter.test(unchanged.text)) return false;
|
||||
const reversed = new RegExp(`${boundary}(?:${cache} (?:does not|no longer) (?:owns?|serializes?)|(?:the |this )?(?:queue|mutex|serialization|tenant context) (?:is|was|has been) (?:removed|disabled|optional)|(?:either|a) service (?:still )?writes directly|(?:${actors.map(escape).join('|')}) (?:also |still )?(?:writes|mutates) (?:the cache )?directly)\\b`, 'i');
|
||||
if (reversed.test(remedy.text)) return false;
|
||||
return rows.every(row => row === remedy || row === unchanged ||
|
||||
new RegExp(`^Keep module-level export but add a per-key write lock inside ${cache}$`).test(row.id![2]!) &&
|
||||
/❌\s*Tests still share one global instance and need manual reset hooks\./.test(row.text) && !namedResolved.test(row.text) && !onlyWriter.test(row.text));
|
||||
}
|
||||
if (annotated) {
|
||||
const text = current(q.question), explanation = current(lines[2] ?? '');
|
||||
const conditional = new RegExp(`\\b(?:if|when|once|unless) (?:approved|accepted)|\\b(?:after|pending) approval\\b|${boundary}${owner} (?:requires (?:approval|acceptance)|is (?:conditional|contingent|dependent) on (?:approval|acceptance))\\b`, 'i');
|
||||
const namedResolved = new RegExp(`${boundary}${escape(annotated[5]!)} (?:is (?:now |already )?(?:serialized|ordered)|no longer (?:shares|has) (?:mutable )?state)\\b`, 'i');
|
||||
const scenario = /while ([A-Za-z_$][\w$]*) is halfway through minting, the mint can land after the invalidation and a suspended tenant keeps a live session\./.exec(explanation);
|
||||
if (!scenario || !actors.includes(scenario[1]!) || framed.test(text) || conditional.test(text) || namedResolved.test(text)) return false;
|
||||
const rows = q.options.map(option => ({ label: option.label.replace(/ \(recommended\)$/, ''), text: current(option.description ?? '').trim() }));
|
||||
const cancelled = new RegExp(`${boundary}(?:do not|don't|never|skip|cancel|withdraw) (?:inject|use|keep|accept|adopt|choose|proceed|reject|write|serialize|document)\\b`, 'i');
|
||||
if (rows.some(row => framed.test(row.text) || closed.test(row.text) || conditional.test(row.text) || cancelled.test(row.text))) return false;
|
||||
const remedy = rows.find(row => row.label === 'Inject + single-writer + version-checked writes');
|
||||
const unchanged = rows.find(row => row.label === 'Keep module-level export as planned');
|
||||
if (!remedy || !unchanged) return false;
|
||||
const writer = new RegExp(`^✅\\s*${escape(annotated[5]!)} passed into both services by constructor from one composition root; ([A-Za-z_$][\\w$]*) is the only writer, ([A-Za-z_$][\\w$]*) reads and invalidates\\.`).exec(remedy.text);
|
||||
if (!writer || writer[1] === writer[2] || !actors.includes(writer[1]) || !actors.includes(writer[2]) ||
|
||||
!/✅\s*Writes carry the policy version and are rejected if the entry was invalidated since read \(compare-and-set\), with a unit test for the interleaving\./.test(remedy.text) ||
|
||||
!/❌\s*Global mutable state shared by two writers, no serialization, order-dependent tests, and a silent tenant-isolation hole\./.test(unchanged.text) ||
|
||||
contradiction.test(unchanged.text) || namedResolved.test(unchanged.text)) return false;
|
||||
const override = new RegExp(`${boundary}(?:${escape(writer[2]!)} (?:also |still )?writes|${escape(writer[1]!)} (?:does not|no longer) writes|(?:the )?adapter (?:accepts stale writes|does not reject stale writes)|(?:the )?(?:version check|single-writer requirement) is (?:removed|disabled|optional))\\b`, 'i');
|
||||
if (override.test(remedy.text) || new RegExp(`${boundary}only (?:${actors.map(escape).join('|')}) writes\\b`, 'i').test(unchanged.text)) return false;
|
||||
return rows.every(row => row === remedy || row === unchanged || row.label === 'Constructor injection only' &&
|
||||
/❌\s*Both services still write freely; the write-after-invalidate race stays open\b/.test(row.text) && !contradiction.test(row.text) &&
|
||||
!namedResolved.test(row.text) && !new RegExp(`${boundary}only (?:${actors.map(escape).join('|')}) writes\\b`, 'i').test(row.text));
|
||||
}
|
||||
const rows = q.options.map(option => ({
|
||||
id: new RegExp(`^${ordinal}([A-D]) (.+?)(?: \\(recommended\\))?$`).exec(option.label),
|
||||
text: current(option.description ?? '').trim(),
|
||||
}));
|
||||
if (rows.some(row => !row.id) || new Set(rows.map(row => row.id![1])).size !== rows.length) return false;
|
||||
const cancelled = new RegExp(`${boundary}(?:do not|don't|never|skip|cancel|withdraw) (?:use|keep|accept|adopt|choose|proceed|reject|write|serialize|document)\\b`, 'i');
|
||||
if (rows.some(row => framed.test(row.text) || closed.test(row.text) || cancelled.test(row.text))) return false;
|
||||
const remedy = rows.find(row => row.id![2] === 'Single writer + version');
|
||||
const unchanged = rows.find(row => row.id![2] === 'Accept the race');
|
||||
if (!remedy || !unchanged) return false;
|
||||
const writers = /^([A-Za-z_$][\w$]*) writes with policy-version tag; adapter rejects stale writes; ([A-Za-z_$][\w$]*) reads\/invalidates\./.exec(remedy.text);
|
||||
if (!writers || writers[1] === writers[2] || !actors.includes(writers[1]!) || !actors.includes(writers[2]!)) return false;
|
||||
const override = new RegExp(`${boundary}(?:${escape(writers[2]!)} (?:also |still )?writes|${escape(writers[1]!)} (?:does not|no longer) writes|(?:the )?adapter (?:accepts stale writes|does not reject stale writes)|(?:the )?(?:version check|single-writer requirement) is (?:removed|disabled|optional))\\b`, 'i');
|
||||
const unchangedOverride = new RegExp(`${boundary}(?:only (?:${actors.map(escape).join('|')}) writes|(?:the |both )?writers no longer (?:write|mutate)|(?:the )?race is (?:no longer current|resolved|closed))\\b`, 'i');
|
||||
if (override.test(remedy.text) || !/^Keep both writers as planned and document the known race\./.test(unchanged.text) || contradiction.test(unchanged.text) || unchangedOverride.test(unchanged.text)) return false;
|
||||
return rows.every(row => row === remedy || row === unchanged ||
|
||||
(row.id![2] === 'Single writer only' &&
|
||||
new RegExp(`^${escape(writers[1]!)} writes, ${escape(writers[2]!)} reads and invalidates\\. No version check\\.`).test(row.text) && !override.test(row.text)));
|
||||
}
|
||||
@@ -0,0 +1,173 @@
|
||||
import type { AskUserQuestionFingerprint } from './claude-pty-runner';
|
||||
import type { NativePlanQuestionCall } from './plan-count-transcript';
|
||||
|
||||
/** Count a completed review-navigation choice separately; never choose pending input. */
|
||||
export function isEngCompletionHandoff(fp: AskUserQuestionFingerprint, reviewedPlan: string,
|
||||
priorCalls: readonly NativePlanQuestionCall[] = []): boolean {
|
||||
const call = fp.nativeCall;
|
||||
if (!call?.sessionId || !call.toolUseId || call.answered !== true || call.failed !== false ||
|
||||
fp.signature !== `${call.sessionId}:${call.toolUseId}` || call.questions.length !== 1 ||
|
||||
!Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
|
||||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0) ||
|
||||
Object.keys(call.answers ?? {}).length !== 1 || !Number.isFinite(Date.parse(call.answeredAt ?? ''))) return false;
|
||||
if (isApprovedMaintenanceRecap(fp, reviewedPlan, priorCalls) || isPublishedPrerequisiteHandoff(fp, reviewedPlan)) return true;
|
||||
const q = call.questions[0]!;
|
||||
if (q.multiSelect || q.header.trim() !== 'Next steps' || q.options.length !== 2 ||
|
||||
fp.options.length !== 2 || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question]) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const body = q.question.trim().replace(/^D[1-9]\d*\s*[—–:-]\s*/i, '');
|
||||
const lines = body.split('\n').map(s => s.trim()).filter(Boolean);
|
||||
if (lines.length !== 4 ||
|
||||
!/^Next steps: Eng Review is CLEAR\. This is a backend auth refactor with no UI scope, so \/plan-design-review does not apply\. No CEO review exists, but the plan changes no product direction\. What next\?$/.test(lines[0]!) ||
|
||||
!/^Recommendation: [1-9]\d*[A-Z] because the plan is implementation-ready and a CEO review would add little to a pure infrastructure refactor\.$/.test(lines[1]!) ||
|
||||
!/^Note: options differ in kind, not coverage [—–-] no completeness score\.$/.test(lines[2]!) ||
|
||||
!/^Net: start building now versus one more optional review pass on scope\.$/.test(lines[3]!)) return false;
|
||||
const label = (s: string) => s.trim().replace(/^[1-9]\d*[A-Z]\)\s*/, '').replace(/\s*\(recommended\)$/, '');
|
||||
const ready = q.options.find(o => /^Ready to implement [—–-] run \/ship when done$/.test(label(o.label)));
|
||||
const ceo = q.options.find(o => label(o.label) === 'Run /plan-ceo-review first');
|
||||
if (!ready || !ceo) return false;
|
||||
const task = /^Exit plan mode with the reviewed plan; implement T([1-9]\d*)[–-]T([1-9]\d*) \(record T([1-9]\d*) regression fixtures first\)\. ✅ All required reviews complete and logged\. ✅ Tasks JSONL and QA test plan are already written for \/autoplan and \/qa\. ❌ No second-opinion pass since codex reviews are disabled\.$/.exec(ready.description ?? '');
|
||||
// Bind implementation references to the published reviewed task catalog.
|
||||
// A new task or a newly proposed regression step is still substantive work.
|
||||
if (!task || Number(task[1]) !== 1 || Number(task[2]) < Number(task[3]) || Number(task[3]) < Number(task[1])) return false;
|
||||
const tasks = [...reviewedPlan.matchAll(/^- \[ \] \*\*T([1-9]\d*)\b[^\n]+$/gm)];
|
||||
const ids = tasks.map(t => Number(t[1])).sort((a, b) => a - b);
|
||||
if (ids.length !== Number(task[2]) || ids.some((id, i) => id !== i + 1)) return false;
|
||||
const regression = tasks.find(t => t[1] === task[3])?.[0] ?? '';
|
||||
if (!/ [—–] Record regression(?: characterization)? fixtures before\b/i.test(regression)) return false;
|
||||
return ceo.description === 'Optional scope and strategy pass before implementing. ✅ Catches product-level questions the eng review does not ask. ✅ Adds a CEO row to the dashboard. ❌ Little product surface here; likely confirms the current scope.';
|
||||
}
|
||||
|
||||
/** Navigation may repeat already-approved post-review bookkeeping, but cannot
|
||||
* authorize it afresh or hide new implementation work behind a completion label. */
|
||||
function isApprovedMaintenanceRecap(fp: AskUserQuestionFingerprint, plan: string,
|
||||
prior: readonly NativePlanQuestionCall[]): boolean {
|
||||
const call = fp.nativeCall!, q = call.questions[0]!;
|
||||
const label = (s: string) => s.replace(/^[1-9]\d*[A-Z]\)\s*/, '').replace(/\s*\(recommended\)$/, '').trim();
|
||||
const current = (s: string) => !/^(?:\s*>|\s*`{3,}|\s*~{3,})|(?:^|\n)\s*(?:source|example|historical|quoted)\b|["“”]|\b(?:withdrawn|cancelled|canceled|rejected|superseded|no longer current|not approved|no longer approved|pending approval|if approved|once approved|assuming approval|provided approval)\b/i.test(s);
|
||||
if (q.multiSelect || !/^Next steps?$/i.test(q.header) || q.options.length !== 2 || fp.options.length !== 2 ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!/^D[1-9]\d*\s*[—–:-]\s*Next steps?[.:]/i.test(q.question) ||
|
||||
!/\bEng(?:ineering)? review is (?:clear(?:ed)?|complete[d]?)\b/i.test(q.question) ||
|
||||
!/\bno UI scope\b/i.test(q.question) || !/\bCEO review is optional\b/i.test(q.question) ||
|
||||
/\b(?:add|implement|rewrite|build|require|if|unless|assuming|provided|pending)\b/i.test(q.question) ||
|
||||
!current([q.question, ...q.options.map(o => `${o.label}\n${o.description ?? ''}`)].join('\n'))) return false;
|
||||
const ready = q.options.find(o => /^Ready to implement(?: [—–-] run \/ship when done)?$/i.test(label(o.label)));
|
||||
const ceo = q.options.find(o => /^Run \/plan-ceo-review(?: first)?$/i.test(label(o.label)));
|
||||
if (!ready || !ceo || call.answers?.[q.question] !== ready.label) return false;
|
||||
const recap = /^Exit plan mode with the reviewed plan\. (?:Post-exit|After exiting): (.+)\.$/i.exec(ready.description ?? '');
|
||||
const actions = recap?.[1]?.split(/\s+and\s+|;\s*/);
|
||||
if (!actions || actions.length !== 2) return false;
|
||||
const routing = actions.filter(a => /^(?:append|add) (?:the )?(?:gstack )?routing rules to CLAUDE\.md$/i.test(a));
|
||||
const todos = actions.map(a => /^(?:create|write) TODOS\.md with (?:the )?(one|two|three|four|five|six|seven|eight|nine|[1-9]) accepted items?$/i.exec(a)).filter(Boolean);
|
||||
if (routing.length !== 1 || todos.length !== 1) return false;
|
||||
const count = Number(todos[0]![1]) || ['one','two','three','four','five','six','seven','eight','nine'].indexOf(todos[0]![1]!.toLowerCase()) + 1;
|
||||
const identities = prior.map(c => `${c.sessionId}:${c.toolUseId}`);
|
||||
if (!prior.length || new Set(identities).size !== prior.length || prior.some(c => c.sessionId !== call.sessionId ||
|
||||
c.toolUseId === call.toolUseId || !c.toolUseId || c.answered !== true || c.failed !== false ||
|
||||
!Number.isFinite(Date.parse(c.answeredAt ?? '')) || Date.parse(c.answeredAt!) >= Date.parse(call.answeredAt!) ||
|
||||
!Array.isArray(c.unansweredQuestionIndices) || c.unansweredQuestionIndices.length ||
|
||||
!c.questions.length || c.questions.length > 4 || Object.keys(c.answers ?? {}).length !== c.questions.length ||
|
||||
new Set(c.questions.map(q => q.question)).size !== c.questions.length ||
|
||||
c.questions.some(q => q.multiSelect || q.options.length < 2 || q.options.length > 4 ||
|
||||
new Set(q.options.map(o => o.label)).size !== q.options.length ||
|
||||
!q.options.some(o => o.label === c.answers?.[q.question])))) return false;
|
||||
const approved = prior.flatMap(c => c.questions.map(q => ({ q, selected: label(c.answers![q.question]!) })));
|
||||
const routingCalls = approved.filter(({q}) => q.header === 'Routing');
|
||||
if (routingCalls.length !== 1 || routingCalls.filter(({q, selected}) => current(q.question) &&
|
||||
/\bskill routing rules\b/.test(q.question) && /CLAUDE\.md/.test(q.question) &&
|
||||
/^(?:Add|Append) (?:gstack )?routing rules to CLAUDE\.md$/.test(selected)).length !== 1) return false;
|
||||
const accepted = approved.filter(({q, selected}) => /^TODO [1-9]\d*$/.test(q.header) &&
|
||||
selected === 'Add to TODOS.md' && current(q.question) &&
|
||||
/\bCaptured in the plan's TODOS section now\b/.test(q.options.find(o => label(o.label) === selected)?.description ?? ''));
|
||||
if (accepted.length !== count || approved.filter(({q}) => /^TODO [1-9]\d*$/.test(q.header)).length !== count ||
|
||||
new Set(accepted.map(a => a.q.header)).size !== count) return false;
|
||||
// Only current, unfenced TODO headings in the published report bind the recap.
|
||||
let fence: string | undefined;
|
||||
const lines = plan.split(/\r?\n/).filter(line => {
|
||||
const mark = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (mark) { if (!fence) fence = mark[1]; else if (mark[1]![0] === fence[0] && mark[1]!.length >= fence.length && !mark[2]!.trim()) fence = undefined; return false; }
|
||||
return !fence && !/^(?: {0,3}>| {4}|\t)/.test(line);
|
||||
});
|
||||
if (fence) return false;
|
||||
const starts = lines.flatMap((line, i) => /^## TODOS?(?:\.md)?\b/i.test(line) ? [i] : []);
|
||||
if (starts.length !== 1 || !current(lines[starts[0]!]!) || /:\s*$/.test(lines.slice(0, starts[0]).filter(s => s.trim()).at(-1) ?? '')) return false;
|
||||
const tail = lines.slice(starts[0]! + 1), end = tail.findIndex(line => /^## /.test(line));
|
||||
const blocks = (end < 0 ? tail : tail.slice(0, end)).join('\n').trim().split(/\n(?=### )/);
|
||||
return blocks.length === count && accepted.every(({q}) => {
|
||||
const subjects = q.question.split('\n')[0]!.match(/\b(?:[a-z]+|[A-Z][a-z]+)(?:[A-Z][A-Za-z0-9]*)+\b/g) ?? [];
|
||||
return subjects.length === 1 && blocks.filter(b => /^### /.test(b) && current(b) &&
|
||||
new RegExp(`\\b${subjects[0]}\\b`).test(b.split('\n')[0]!)).length === 1;
|
||||
});
|
||||
}
|
||||
|
||||
/** A finished backend review may recap an already-published author prerequisite.
|
||||
* This classifies only its completed navigation; the runner still independently
|
||||
* requires the owned report, fresh modifying decisions and a later native exit.
|
||||
*/
|
||||
function isPublishedPrerequisiteHandoff(fp: AskUserQuestionFingerprint, reviewedPlan: string): boolean {
|
||||
const call = fp.nativeCall!, q = call.questions[0]!;
|
||||
if (q.multiSelect || q.header.trim() !== 'Next' || q.options.length !== 2 || fp.options.length !== 2 ||
|
||||
!fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label) ||
|
||||
!q.options.some(o => o.label === call.answers?.[q.question]) || /<gstack-qid/i.test(q.question)) return false;
|
||||
const lines = q.question.trim().split('\n').map(s => s.trim()).filter(Boolean);
|
||||
if (lines.length !== 7 || !/^D[1-9]\d* [—–-] Next step after this eng review\?$/.test(lines[0]!) ||
|
||||
!/^Project\/branch\/task: [\w.-]+ on [\w./-]+, reviewed plan written to gstack-test-plan-eng\.md\.$/.test(lines[1]!) ||
|
||||
lines[2] !== 'ELI10: The eng review is the only gate that blocks shipping and it is clear. No UI is touched, so a design review does not apply. A CEO review is optional and normally for product-direction changes; this is a backend auth refactor, so it is a soft mention only.' ||
|
||||
lines[3] !== 'Stakes if we pick wrong: Low either way; the CEO review would cost time on a change with no product-facing scope decision left open.' ||
|
||||
lines[5] !== 'Note: options differ in kind, not coverage — no completeness score.' ||
|
||||
lines[6] !== 'Net: start implementing versus an optional strategy pass on a backend refactor.') return false;
|
||||
const recommendation = /^Recommendation: ([A-Z]) because all required reviews are complete and the plan's remaining blocker is the author confirming the Context section, not another review\.$/.exec(lines[4]!);
|
||||
const ready = q.options.find(o => /^[A-Z]\) Ready to implement \(recommended\)$/.test(o.label));
|
||||
const ceo = q.options.find(o => /^[A-Z]\) Run \/plan-ceo-review$/.test(o.label));
|
||||
if (!recommendation || !ready || !ceo || ready.label[0] !== recommendation[1] || ready.label[0] === ceo.label[0]) return false;
|
||||
const task = /^✅ All relevant reviews complete; run \/ship when the work is done\. ✅ The first task is the author confirming Context, then T([1-9]\d*) characterization tests\. ❌ No second strategic opinion on whether the refactor is the right thing to build now\.$/.exec(ready.description ?? '');
|
||||
if (!task || ceo.description !== '✅ Adds a scope-and-strategy pass before any code is written. ✅ Useful if the refactor\'s business motivation is contested. ❌ Backend-only refactor with no product-direction choice; likely low yield for the time.') return false;
|
||||
|
||||
// Only the published Context and Implementation Tasks sections own the
|
||||
// references. Quoted or fenced examples cannot supply a prerequisite/task.
|
||||
const published: string[] = [];
|
||||
let fence: { marker: string; length: number } | undefined;
|
||||
for (const line of reviewedPlan.split(/\r?\n/)) {
|
||||
if (fence) {
|
||||
const close = /^ {0,3}(`{3,}|~{3,})[ \t]*$/.exec(line);
|
||||
if (close && close[1]![0] === fence.marker && close[1]!.length >= fence.length) fence = undefined;
|
||||
continue;
|
||||
}
|
||||
const open = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (open && (open[1]![0] !== '`' || !open[2]!.includes('`'))) {
|
||||
fence = { marker: open[1]![0]!, length: open[1]!.length }; continue;
|
||||
}
|
||||
if (!/^(?: {4}|\t| {0,3}>)/.test(line)) published.push(line);
|
||||
}
|
||||
if (fence) return false;
|
||||
const section = (heading: string): string | undefined => {
|
||||
const hits = published.flatMap((line, i) => line === `## ${heading}` ? [i] : []);
|
||||
if (hits.length !== 1) return undefined;
|
||||
const start = hits[0]!;
|
||||
const preceding = published.slice(0, start).filter(s => s.trim()).at(-1) ?? '';
|
||||
if (/[::]$|\b(?:example|sample|hypothetical|template|quoted)\b/i.test(preceding)) return undefined;
|
||||
const end = published.findIndex((line, i) => i > start && /^#{1,2} /.test(line));
|
||||
return published.slice(start + 1, end < 0 ? undefined : end).join('\n');
|
||||
};
|
||||
const context = section('Context'), tasks = section('Implementation Tasks');
|
||||
if (!context || !tasks) return false;
|
||||
const prerequisite = /^### Prerequisite P(\d+) \(decision D[1-9]\d* → [1-9]\d*[A-Z]\)\nImplementation does not start until the author confirms or edits the Problem, Goal,\nInvariants and Latency target above\./m.exec(context);
|
||||
if (!prerequisite || !context.includes(`author must confirm — see Prerequisite P${prerequisite[1]} below`)) return false;
|
||||
const entries = [...tasks.matchAll(/^- \[ \] \*\*T([1-9]\d*)\b[^\n]+$/gm)];
|
||||
const referenced = entries.filter(t => t[1] === task[1]);
|
||||
if (referenced.length !== 1 || !/ — auth\/legacy — Write characterization tests for `legacyAuthFlow\(\)` before any rewrite$/.test(referenced[0]![0])) return false;
|
||||
const prerequisiteTail = context.slice(prerequisite.index!);
|
||||
const nextPrerequisiteHeading = prerequisiteTail.indexOf('\n### ', 1);
|
||||
const prerequisiteOwner = nextPrerequisiteHeading < 0 ? prerequisiteTail : prerequisiteTail.slice(0, nextPrerequisiteHeading);
|
||||
const taskStart = referenced[0]!.index!;
|
||||
const nextTask = entries.find(entry => entry.index! > taskStart)?.index ?? tasks.length;
|
||||
// A later correction inside the same owner can withdraw its earlier rule.
|
||||
// Unrelated prerequisite/task bodies cannot supply or revoke this reference.
|
||||
const withdrawn = (owner: string, id: string) => new RegExp(
|
||||
`\\b(?:Prerequisite )?${id}\\b(?: (?:requirement|task))? (?:is |was |has been )?(?:cancelled|canceled|withdrawn|rejected|not (?:required|needed|necessary))\\b`, 'i').test(owner);
|
||||
return !withdrawn(prerequisiteOwner, `P${prerequisite[1]}`) &&
|
||||
!/\b(?:the )?author no longer needs to confirm Context\b|\bContext confirmation is (?:not required|cancelled|withdrawn)\b/i.test(prerequisiteOwner) &&
|
||||
!withdrawn(tasks.slice(taskStart, nextTask), `T${task[1]}`) &&
|
||||
!/\bno characterization tests (?:are )?required\b|\bcharacterization tests are (?:not required|cancelled|withdrawn)\b/i.test(tasks.slice(taskStart, nextTask));
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
/** A recorded legacy oracle can survive removal of the implementation it sampled. */
|
||||
export function hasRetainedLegacyCorpus(
|
||||
current: ReadonlyArray<{ title: string; body: string[] }>, snapshot: string,
|
||||
): boolean {
|
||||
const flat = (s: string) => s.replace(/\s+/g, ' ').trim();
|
||||
const quoted = (s: string) => s.replace(/"[^"\n]*"|“[^”\n]*”|(?<![\w])'[^'\n]*'(?![\w])|‘[^’\n]*’/g, '');
|
||||
const source = (s: string) => /(?:^|\n)\s*(?:source(?: excerpt| material)?|quoted(?: source)?|historical(?: example| assessment)?|if approved|once approved|when approved|pending approval|assuming approval|provided approval)\s*[,.:—-]/i.test(quoted(s));
|
||||
const status = '(?:withdrawn|rejected|declined|cancelled|canceled|superseded|deferred|optional|proposed|not current|no longer current|not required|no longer required|conditional on approval)';
|
||||
const owner = '(?:(?:this|the) (?:(?:legacy|recorded|baseline) )?(?:(?:regression|parity) )?(?:suite|test|requirement|verification|oracle|corpus))';
|
||||
const scalar = (s: string, id: string) => quoted(s.replace(new RegExp(
|
||||
`((?:^|[.!?;]\\s+|\\n)\\s*(?:Correction:\\s*)?(?:${id}(?: verification)?|${owner}) (?:is|are|was|were|has been|have been) )["“'‘](${status})["”'’]`, 'gim'), '$1$2'));
|
||||
const inactive = (s: string, id: string) => source(s) || new RegExp(
|
||||
`\\b(?:${id}(?: verification)?|${owner}) (?:is|are|was|were|has been|have been) ${status}\\b|\\b(?:only if|unless|pending) (?:user )?approv`, 'i').test(scalar(s, id));
|
||||
const negated = (s: string) => /\b(?:do not|don't|never|skip|omit|defer|cancel) (?:add|write|implement|record|capture|pin|run|assert|retain|keep)\b/i.test(quoted(s));
|
||||
const owned = (s: string) => !source(s) && !negated(s) && snapshot.includes(flat(s));
|
||||
|
||||
for (const declaration of current) {
|
||||
if (!/^CRITICAL regression test \([^)]*\bmandatory\b[^)]*\)$/i.test(declaration.title)
|
||||
|| /\b(?:not|never|no longer) mandatory\b/i.test(declaration.title)) continue;
|
||||
const body = declaration.body.join('\n').trim();
|
||||
const add = /^Add ([A-Za-z][\w/.-]*\.test\.[jt]s):$/m.exec(body);
|
||||
if (!add || !owned(body) || inactive(body, 'T[1-9]\\d*')) continue;
|
||||
const bullets = body.split(/\n(?=- )/).slice(1).map(flat);
|
||||
const capture = bullets.findIndex(s => /^- (?:Record|Capture|Pin) a corpus of .+ with the legacy decision for each\.$/i.test(s));
|
||||
const parity = bullets.findIndex(s => /^- Run the (?:same )?corpus through the new [A-Za-z][\w]* path and assert identical allow\/deny and reason code for every entry\.$/i.test(s));
|
||||
if (capture < 0 || parity <= capture || !bullets.some(s => /^- This test is also the shadow-mode oracle; it stays after legacy deletion, re-pointed at the recorded decisions\.$/i.test(s))) continue;
|
||||
// Retention makes the old behavior the oracle. Building independent new
|
||||
// modules before recording it does not itself rewrite that old behavior.
|
||||
const retention = current.filter(s => s.title === 'What already exists').flatMap(s => s.body.join('\n').split(/\n(?=- )/))
|
||||
.find(s => /^- legacyAuthFlow\(\): retained behind the per-tenant flag as the shadow oracle and regression baseline until deletion\.$/i.test(flat(s)) && owned(s) && !inactive(s, 'T[1-9]\\d*'));
|
||||
if (!retention) continue;
|
||||
for (const section of current.filter(s => s.title === 'Implementation Tasks')) {
|
||||
const taskBody = section.body.join('\n').trim();
|
||||
const tasks = taskBody.split(/\n(?=- )/);
|
||||
const rows = tasks.map(body => ({ body, match: /^- (?:\[[ xX]\] )?(T[1-9]\d*)(?: \([^\n)]*\))? [—–-] (.+)(?:\n|$)/.exec(body) })).filter(row => row.match);
|
||||
if (rows.length !== new Set(rows.map(row => row.match![1])).size) continue;
|
||||
for (const row of rows) {
|
||||
const id = row.match![1]!, title = row.match![2]!;
|
||||
const action = row.body.replace(/^ - Surfaced by:.*$/gm, '');
|
||||
const preceding = taskBody.slice(0, taskBody.indexOf(row.body)).trim().split('\n').at(-1) ?? '';
|
||||
const files = [...row.body.matchAll(/^ - Files: ([^\n]+)$/gm)];
|
||||
const verifies = [...row.body.matchAll(/^ - Verify: ([^\n]+)$/gm)];
|
||||
if (!/^[A-Za-z][\w -]* [—–-] CRITICAL: recorded-corpus parity test legacy vs [A-Za-z][\w]*$/i.test(title)
|
||||
|| files.length !== 1 || verifies.length !== 1 || !snapshot.includes(flat(row.body)) || source(action) || negated(action) || source(preceding)
|
||||
|| inactive(action, id) || !files[0]![1]!.split(',').map(s => s.trim()).includes(add[1]!)) continue;
|
||||
if (!/^100% decision \+ reason-code parity across the corpus$/i.test(verifies[0]![1]!)) continue;
|
||||
const cancelled = current.some(s => {
|
||||
if (/\b(?:history|historical|source|quoted|example)\b/i.test(s.title)) return false;
|
||||
const raw = s.body.join('\n'), value = scalar(raw, id);
|
||||
const named = /^(.*?)\b(?:regression|characterization|parity)\s+(?:suite|test)/i.exec(s.title)?.[1]?.trim();
|
||||
const foreign = Boolean(named && !/^(?:(?:critical|recorded|required|current|final|updated)\s*)*(?:legacy(?:AuthFlow\(\))?)?[\s:—-]*$/i.test(named));
|
||||
if (new RegExp(`^\\s*\\|\\s*${id}\\s*\\|\\s*["“'‘]?${status}["”'’]?\\s*\\|`, 'im').test(raw)) return true;
|
||||
return value.split(/\n|[.!?;]\s+/).some(line => {
|
||||
if (source(line) || /^\s*(?:if|unless|assuming|provided)\b/i.test(line)) return false;
|
||||
if (foreign && !new RegExp(`\\b${id}\\b|legacyAuthFlow|\\blegacy (?:regression|parity) (?:suite|test|corpus|baseline)\\b`, 'i').test(line)) return false;
|
||||
return inactive(line, id) || new RegExp(`^\\s*(?:Correction:\\s*)?legacyAuthFlow\\(\\) (?:is|was|has been|will be) (?:modified|changed|rewritten|refactored|removed|deleted) before ${id}\\b`, 'i').test(line)
|
||||
|| /^\s*(?:Correction:\s*)?legacyAuthFlow\(\) (?:is|was|has been|will be) (?:modified|changed|rewritten|refactored|removed|deleted) before (?:the )?(?:(?:legacy|recorded|regression) )?(?:corpus|baseline)(?: is)? (?:recorded|captured|pinned)\b/i.test(line)
|
||||
|| /^\s*(?:the|this) (?:(?:legacy|regression) )?(?:corpus|baseline) is (?:recorded|captured|pinned) (?:only )?after legacyAuthFlow\(\) is (?:modified|changed|rewritten|refactored|removed|deleted)\b/i.test(line)
|
||||
|| /\b(?:the|this) legacy (?:regression baseline|shadow oracle) (?:is|was|has been) (?:changed|rewritten|removed|deleted|replaced)\b/i.test(line);
|
||||
});
|
||||
});
|
||||
if (!cancelled) return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -30,7 +30,8 @@ export const PTY_MS = 900_000;
|
||||
/**
|
||||
* Chained/judged PTY observation — the ceiling tier. 1200s leaves the
|
||||
* 1800s shard wall real overhead; anything that genuinely needs more
|
||||
* should be SPLIT, not budgeted past the wall.
|
||||
* should be split or use an explicitly registered workflow exception with
|
||||
* corresponding runner and CI walls; never inflate an ordinary tier.
|
||||
*/
|
||||
export const PTY_LONG_MS = 1_200_000;
|
||||
|
||||
@@ -41,3 +42,31 @@ export const ALL_TIERS = {
|
||||
PTY_MS,
|
||||
PTY_LONG_MS,
|
||||
} as const;
|
||||
|
||||
/**
|
||||
* Explicit exception for one uninterrupted four-phase workflow. These are
|
||||
* specified allowances, not measured latency or a conservative confidence bound.
|
||||
* The historical 900-second failure remains a failure. Ordinary tiers do not grow.
|
||||
*/
|
||||
export const AUTOPLAN_CHAIN_BUDGET = {
|
||||
id: 'autoplan-four-native-phases-v1',
|
||||
file: 'test/skill-e2e-autoplan-chain.test.ts',
|
||||
workMs: 4 * PTY_LONG_MS,
|
||||
sessionMs: 84 * 60_000,
|
||||
testMs: 85 * 60_000,
|
||||
shardMs: 172 * 60_000,
|
||||
retries: 1,
|
||||
shardReserveMs: 2 * 60_000,
|
||||
ciJobMs: 200 * 60_000,
|
||||
ciReserveMs: 28 * 60_000,
|
||||
reason: 'One command must complete CEO, Design, DX and Eng, including native reviews and amendment handoffs.',
|
||||
} as const;
|
||||
|
||||
/** The only registered over-tier test budget; arbitrary per-file escapes fail. */
|
||||
export function assertPaidTestBudget(file: string, ms: number): void {
|
||||
if (!Number.isSafeInteger(ms) || ms <= 0 ||
|
||||
(ms > PTY_LONG_MS * 1.25 &&
|
||||
(file !== AUTOPLAN_CHAIN_BUDGET.file || ms !== AUTOPLAN_CHAIN_BUDGET.testMs))) {
|
||||
throw new Error(`Unregistered paid test budget: ${file}: ${ms}`);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
import * as fs from 'node:fs';
|
||||
|
||||
/** Keep fixture CLIs native on Windows, where an executable shebang is unsupported. */
|
||||
export function createFakeBunCli(file: string, source: string, native = process.platform === 'win32'): string {
|
||||
const script = native ? `${file}.ts` : file;
|
||||
fs.writeFileSync(script, `#!${process.execPath}\n` + source.replace(/^#![^\n]*\n/, ''), { mode: 0o755 });
|
||||
if (!native) return script;
|
||||
const executable = `${file}.exe`;
|
||||
const built = Bun.spawnSync([process.execPath, 'build', '--compile', script, '--outfile', executable], {
|
||||
stdout: 'pipe', stderr: 'pipe', timeout: 30_000,
|
||||
});
|
||||
if (built.exitCode !== 0) throw new Error(`Could not compile fake CLI: ${built.stderr.toString()}`);
|
||||
return executable;
|
||||
}
|
||||
@@ -249,12 +249,73 @@ export function getHermeticDirs(): HermeticDirs {
|
||||
|
||||
let cachedSkillsConfigDir: string | null = null;
|
||||
|
||||
/**
|
||||
* Canonical paths without exposing the checkout to Claude's recursive markdown
|
||||
* file index. Directory links would also expose .context, dependencies, and
|
||||
* generated host registries. Expand owned runtime directories and link files
|
||||
* individually, preserving their source realpaths and executable bits.
|
||||
*/
|
||||
export function seedHermeticRuntimeView(root: string, destination: string): void {
|
||||
const sourceRoot = fs.realpathSync(root);
|
||||
const excluded = new Set(['node_modules', 'test', 'tests']);
|
||||
const roots = new Set([
|
||||
'SKILL.md', 'ETHOS.md', 'VERSION', 'package.json', 'bin', 'lib', 'scripts',
|
||||
'browse', 'browser-skills', 'design', 'extension', 'model-overlays', 'agents',
|
||||
// On-demand protocol/reference files explicitly read by shipped skills.
|
||||
'docs/askuserquestion-split.md', 'docs/askuserquestion-cjk.md',
|
||||
'docs/designs/PLAN_TUNING_V0.md', 'docs/designs/PLAN_TUNING_V1.md',
|
||||
...skillCensus(root).physicalSkillFiles.filter(file => file !== 'SKILL.md').map(file => path.dirname(file)),
|
||||
]);
|
||||
const allowed = [...roots].filter(name => fs.existsSync(path.join(root, name))).map(name => ({
|
||||
path: fs.realpathSync(path.join(root, name)), directory: fs.statSync(path.join(root, name)).isDirectory(),
|
||||
}));
|
||||
function link(source: string, target: string, ancestors: Set<string>): void {
|
||||
const real = fs.realpathSync(source);
|
||||
if (real !== sourceRoot && !real.startsWith(sourceRoot + path.sep))
|
||||
throw new Error(`Runtime asset leaves the owned checkout: ${source}`);
|
||||
if (path.relative(sourceRoot, real).split(path.sep).some(name =>
|
||||
name.startsWith('.') || excluded.has(name) || name.endsWith('.tmpl')))
|
||||
throw new Error(`Runtime asset resolves into an excluded tree: ${source}`);
|
||||
if (real !== sourceRoot && !allowed.some(asset => real === asset.path ||
|
||||
asset.directory && real.startsWith(asset.path + path.sep)))
|
||||
throw new Error(`Runtime asset resolves into an excluded tree: ${source}`);
|
||||
const stat = fs.statSync(source);
|
||||
if (stat.isDirectory()) {
|
||||
if (ancestors.has(real)) throw new Error(`Circular runtime asset: ${source}`);
|
||||
fs.mkdirSync(target);
|
||||
const next = new Set([...ancestors, real]);
|
||||
for (const name of fs.readdirSync(source)) {
|
||||
if (name.startsWith('.') || excluded.has(name) || name.endsWith('.tmpl')) continue;
|
||||
link(path.join(source, name), path.join(target, name), next);
|
||||
}
|
||||
} else if (stat.isFile()) {
|
||||
fs.symlinkSync(source, target, 'file');
|
||||
}
|
||||
}
|
||||
fs.mkdirSync(destination);
|
||||
try {
|
||||
for (const name of roots) {
|
||||
const source = path.join(root, name);
|
||||
if (fs.existsSync(source)) {
|
||||
const target = path.join(destination, name);
|
||||
fs.mkdirSync(path.dirname(target), { recursive: true });
|
||||
link(source, target, new Set([sourceRoot]));
|
||||
}
|
||||
}
|
||||
} catch (error) {
|
||||
fs.rmSync(destination, { recursive: true, force: true });
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A hermetic CLAUDE_CONFIG_DIR with the repo's shipped skills REGISTERED in
|
||||
* user scope, mirroring ./setup's registration exactly: each discovered skill
|
||||
* gets a REAL directory `<configDir>/skills/<registryName>/` containing a
|
||||
* SYMLINK to that skill's SKILL.md (absolute path), plus a `sections/`
|
||||
* symlink when the skill has one. registryName is the frontmatter `name:`
|
||||
* SYMLINK to that skill's SKILL.md (absolute path), plus symlinks to its
|
||||
* runtime assets using setup's exclusions. An owned runtime view also
|
||||
* lives at `<configDir>/skills/gstack`, so canonical
|
||||
* lazy-section paths work alongside flattened discovery. registryName is the frontmatter `name:`
|
||||
* (dir-name fallback), NO gstack- prefix; the root SKILL.md router registers
|
||||
* as `_gstack-command`. skillCensus().registryEntries is the authoritative
|
||||
* set of what must appear here.
|
||||
@@ -268,12 +329,11 @@ let cachedSkillsConfigDir: string | null = null;
|
||||
* cover it. Ends in `/.claude` for the same plan-path anchoring reason as
|
||||
* HermeticDirs.configDir.
|
||||
*
|
||||
* Two intentional non-hermetic edges:
|
||||
* - Seeding reads the LIVE repo tree BY DESIGN — the skills ARE the subject
|
||||
* under test; a snapshot would measure stale copies.
|
||||
* - HOME is not hermeticized, so the ~64 absolute
|
||||
* `~/.claude/skills/gstack/...` preamble references inside each SKILL.md
|
||||
* still resolve to the operator install (same limitation as CI).
|
||||
* Seeding reads the live repo tree: the skills are the subject under test.
|
||||
* For default seeded PTY sessions, launchClaudePty also supplies an owned HOME
|
||||
* via hermetic-skill-runtime so literal runtime and lazy-section paths reach
|
||||
* this same checkout. Explicit HOME/config overrides remain caller-owned;
|
||||
* this registration helper itself does not change their environment.
|
||||
*/
|
||||
export function hermeticSkillsConfigDir(): string {
|
||||
if (cachedSkillsConfigDir) return cachedSkillsConfigDir;
|
||||
@@ -303,13 +363,23 @@ export function hermeticSkillsConfigDir(): string {
|
||||
safeUnlink(path.join(target, 'SKILL.md'));
|
||||
fs.symlinkSync(skillMd, path.join(target, 'SKILL.md'));
|
||||
if (rel !== 'SKILL.md') {
|
||||
const sections = path.join(root, skillDir, 'sections');
|
||||
if (fs.existsSync(sections)) {
|
||||
safeUnlink(path.join(target, 'sections'));
|
||||
fs.symlinkSync(sections, path.join(target, 'sections'));
|
||||
// Mirror setup's _link_skill_runtime_assets, including references and
|
||||
// helpers beside sections. Missing assets can send a live agent looking
|
||||
// outside its installed fixture and into the operator's stale checkout.
|
||||
const source = path.join(root, skillDir);
|
||||
for (const name of fs.readdirSync(source)) {
|
||||
if (name.startsWith('.') || ['SKILL.md', 'node_modules', 'dist', 'test'].includes(name) || name.endsWith('.tmpl')) continue;
|
||||
const asset = path.join(source, name);
|
||||
if (!fs.existsSync(asset)) continue;
|
||||
const destination = path.join(target, name);
|
||||
safeUnlink(destination);
|
||||
fs.symlinkSync(asset, destination, fs.statSync(asset).isDirectory() ? 'dir' : 'file');
|
||||
}
|
||||
}
|
||||
}
|
||||
// Canonical lazy paths remain available without letting native file-index
|
||||
// discovery recursively read the source checkout's historical artifacts.
|
||||
seedHermeticRuntimeView(root, path.join(skillsDir, 'gstack'));
|
||||
cachedSkillsConfigDir = configDir;
|
||||
return configDir;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
/**
|
||||
* Runtime paths for PTY sessions that seed the checkout's Claude skills.
|
||||
*
|
||||
* Claude's registry uses CLAUDE_CONFIG_DIR, but generated skills also read
|
||||
* ~/.claude/skills/gstack and execute $HOME/.claude/skills/gstack/bin tools.
|
||||
* Point those unchanged paths at the same checkout, never an operator install.
|
||||
* This HOME is owned by the existing hermetic runRoot and its exit/GC cleanup.
|
||||
*
|
||||
* Two intentional shared resources survive the HOME change: nested Codex's
|
||||
* configured auth/model directory, and Playwright's existing browser cache.
|
||||
* Claude auth stays in the existing seeded config/API env; gstack state stays
|
||||
* at the caller's GSTACK_HOME. No operator skill, hooks, or settings are copied.
|
||||
*/
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import * as os from 'node:os';
|
||||
import { getHermeticDirs, hermeticSkillsConfigDir } from './hermetic-env';
|
||||
|
||||
let cachedRuntime: { home: string; root: string; stateRoot: string } | undefined;
|
||||
|
||||
export function hermeticSkillRuntime(): { home: string; root: string; stateRoot: string } {
|
||||
if (cachedRuntime) return cachedRuntime;
|
||||
const home = fs.mkdtempSync(path.join(getHermeticDirs().runRoot, 'skill-home-'));
|
||||
const root = path.join(home, '.claude', 'skills', 'gstack');
|
||||
const stateRoot = path.join(home, '.gstack');
|
||||
try {
|
||||
// Some generated workflows keep snapshots under ~/.gstack independently
|
||||
// of GSTACK_HOME. This is owned by the same disposable HOME and cleanup.
|
||||
fs.mkdirSync(stateRoot);
|
||||
fs.mkdirSync(path.dirname(root), { recursive: true });
|
||||
fs.symlinkSync(path.resolve(import.meta.dir, '..', '..'), root, 'dir');
|
||||
// Setup exposes both the gstack runtime checkout and flattened skill
|
||||
// entries. Keep HOME discovery consistent with CLAUDE_CONFIG_DIR: native
|
||||
// tools may resolve ~/paths even though slash commands use the latter.
|
||||
const registry = path.join(hermeticSkillsConfigDir(), 'skills');
|
||||
for (const name of fs.readdirSync(registry)) {
|
||||
if (name === 'gstack') continue; // canonical checkout already linked above
|
||||
fs.symlinkSync(path.join(registry, name), path.join(path.dirname(root), name), 'dir');
|
||||
}
|
||||
} catch (error) {
|
||||
fs.rmSync(home, { recursive: true, force: true });
|
||||
throw error;
|
||||
}
|
||||
cachedRuntime = { home, root, stateRoot };
|
||||
return cachedRuntime;
|
||||
}
|
||||
|
||||
/** Keep the cache the child used before HOME isolation (Playwright defaults). */
|
||||
function browserCache(env: Record<string, string>, originalHome: string): string {
|
||||
if (env.PLAYWRIGHT_BROWSERS_PATH) return env.PLAYWRIGHT_BROWSERS_PATH;
|
||||
const base = process.platform === 'darwin'
|
||||
? path.join(originalHome, 'Library', 'Caches')
|
||||
: process.platform === 'win32'
|
||||
? env.LOCALAPPDATA || path.join(originalHome, 'AppData', 'Local')
|
||||
: env.XDG_CACHE_HOME || path.join(originalHome, '.cache');
|
||||
return path.join(base, 'ms-playwright');
|
||||
}
|
||||
|
||||
export function withHermeticSkillRuntime(
|
||||
env: Record<string, string>,
|
||||
operatorEnv: NodeJS.ProcessEnv = process.env,
|
||||
): { env: Record<string, string>; root: string; stateRoot: string } {
|
||||
const originalHome = env.HOME || os.homedir();
|
||||
const runtime = hermeticSkillRuntime();
|
||||
// An explicit empty per-test value resets Codex to its original HOME default.
|
||||
const codexHome = env.CODEX_HOME !== undefined
|
||||
? env.CODEX_HOME || path.join(originalHome, '.codex')
|
||||
: operatorEnv.CODEX_HOME || path.join(originalHome, '.codex');
|
||||
return {
|
||||
root: runtime.root,
|
||||
stateRoot: runtime.stateRoot,
|
||||
env: {
|
||||
...env,
|
||||
HOME: runtime.home,
|
||||
CODEX_HOME: codexHome,
|
||||
PLAYWRIGHT_BROWSERS_PATH: browserCache(env, originalHome),
|
||||
},
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
import type { NativePublicToolEvent, PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
export interface NativeAutoDecision {
|
||||
sessionId: string;
|
||||
skillToolUseId: string;
|
||||
timestamp: string;
|
||||
summary: string;
|
||||
option: string;
|
||||
annotation: string;
|
||||
}
|
||||
|
||||
const plain = (text: string) => text.replace(/\*\*([^*]+)\*\*/g, '$1').trim();
|
||||
const annotationLine = /^Auto-decided ([^\r\n→]{1,240}) → ([^\r\n→]{1,200}) \(your (?:preference|saved preference on `([a-z][a-z0-9-]*)`)\)\. Change with \/plan-tune\.$/;
|
||||
|
||||
/** Asserted prose only; later quoted examples cannot retract a current decision. */
|
||||
function publicProse(text: string): string {
|
||||
const lines: string[] = [];
|
||||
let fence: { char: string; length: number } | undefined;
|
||||
for (const line of text.split(/\r?\n/)) {
|
||||
const marker = /^ {0,3}(`{3,}|~{3,})/.exec(line);
|
||||
if (marker) {
|
||||
if (!fence) fence = { char: marker[1]![0]!, length: marker[1]!.length };
|
||||
else if (marker[1]![0] === fence.char && marker[1]!.length >= fence.length && /^\s*$/.test(line.slice(marker[0].length))) fence = undefined;
|
||||
continue;
|
||||
}
|
||||
if (fence || /^(?: {4}|\t|\s*>)/.test(line)) continue;
|
||||
const current = plain(line);
|
||||
// A direct correction owns its quoted verdict; an attributed example does not.
|
||||
if (/^(?:Correction|Actually|Update):\s*/i.test(current)) lines.push(current.replace(/["“”`]/g, ''));
|
||||
else lines.push(current.replace(/"(?:[^"\\]|\\.)*"|“[^”]*”|`[^`]*`/g, '""'));
|
||||
}
|
||||
return lines.join('\n');
|
||||
}
|
||||
|
||||
function withdrawn(text: string, option: string): boolean {
|
||||
const prose = publicProse(text);
|
||||
if (/\b(?:I|we)\s+(?:retract|withdraw|revoke|cancel)\b[^.!?\n]{0,100}\b(?:auto[- ]decision|annotation|decision|selection|choice)\b/i.test(prose) ||
|
||||
/\b(?:I|we)\s+(?:did not|didn't|have not|haven't|will not|won't|no longer)\s+auto-decide\b/i.test(prose) ||
|
||||
/\b(?:I|we)\s+(?:did not|didn't|have not|haven't)\s+make\s+(?:this|that|the)\s+(?:decision|selection|choice)\b/i.test(prose) ||
|
||||
/\b(?:this|that|the)\s+(?:auto[- ]decision|annotation|statement|decision|selection|choice)\b[^.!?\n]{0,100}\b(?:withdrawn|retracted|revoked|cancelled|canceled|hypothetical|conditional|example)\b/i.test(prose)) return true;
|
||||
return [...prose.matchAll(/^Review mode:\s*([^\n.]+)\.?$/gmi)].some(m => plain(m[1]!).toLowerCase() !== option.toLowerCase());
|
||||
}
|
||||
|
||||
function assertedAnnotation(text: string, skillName: string): RegExpExecArray | null {
|
||||
const paragraphs = text.replace(/^(?:[ \t]*\r?\n)+|(?:\r?\n[ \t]*)+$/g, '').split(/\r?\n\s*\r?\n/);
|
||||
let index = 0;
|
||||
// A preamble notice is independent of the immediately following current
|
||||
// mode declaration. No arbitrary source/example prefix is skipped.
|
||||
const preambleNotice = "Heads-up from the preamble: unshipped work on this branch, so `/review` then `/ship` when you're ready. Also, gstack follows the **Boil the Ocean** principle: do the complete thing when AI makes the marginal cost near zero. Read more at https://garryslist.org/posts/boil-the-ocean if you'd like.";
|
||||
const decisionNotice = "Heads-up from gstack: there is unshipped work on this branch, so `/review` then `/ship` when you get to it.";
|
||||
if (/^Heads-up from gstack: this branch has unshipped work\. Run `\/review` then `\/ship` when you're ready\.$/.test(paragraphs[0] ?? '') || paragraphs[0] === preambleNotice || paragraphs[0] === decisionNotice) index++;
|
||||
const mode = /^\*\*Review mode:\s*([^*\n.]+)\.\*\*$/.exec(paragraphs[index] ?? '');
|
||||
const decisionHeading = /^\*\*D[1-9]\d* [—–-] Review mode for the ([^*\n]+) draft\*\*$/.exec(paragraphs[index] ?? '');
|
||||
if (decisionHeading && /\b(?:example|hypothetical|historical|previous|quoted)\b/i.test(decisionHeading[1]!)) return null;
|
||||
if (mode || decisionHeading) index++;
|
||||
if (index && !mode && !decisionHeading) return null;
|
||||
const paragraph = paragraphs[index];
|
||||
if (!paragraph || /^(?: {4}|\t)/.test(paragraph) || paragraph.includes('\n')) return null;
|
||||
const match = annotationLine.exec(paragraph);
|
||||
if (!match || !plain(match[1]!) || !plain(match[2]!)) return null;
|
||||
if (decisionHeading && (!/^"Review mode:[^"]+\?"$/.test(match[1]!) ||
|
||||
!/^(?:HOLD SCOPE|SCOPE EXPANSION|SELECTIVE EXPANSION|SCOPE REDUCTION)$/.test(plain(match[2]!)))) return null;
|
||||
// The printed skill template is not a concrete observed choice.
|
||||
if (/<[^>\r\n]+>/.test(match[1]!) || /<[^>\r\n]+>/.test(match[2]!)) return null;
|
||||
// A named saved preference belongs to the invoked skill's mode, and its
|
||||
// concrete choice must agree with the adjacent current mode declaration.
|
||||
if (match[3] && (match[3] !== `${skillName}-mode` || !mode ||
|
||||
!/^(?:HOLD SCOPE|SCOPE EXPANSION|SELECTIVE EXPANSION|SCOPE REDUCTION)$/.test(plain(match[2]!)))) return null;
|
||||
if (mode && (plain(mode[1]!).toLowerCase() !== plain(match[2]!).toLowerCase() ||
|
||||
(!/^(?:review mode|"Review mode:[^"]+\?")$/i.test(match[1]!) &&
|
||||
!(match[3] && /^"Select review mode"$/.test(match[1]!))))) return null;
|
||||
return match;
|
||||
}
|
||||
|
||||
/** Exact public annotation after a successful invocation in the owned native session. */
|
||||
export function findNativeAutoDecision(
|
||||
transcript: PlanCountTranscript,
|
||||
tools: NativePublicToolEvent[],
|
||||
opts: { skillName: string; sessionId: string; commandStartedAt: number; now: number },
|
||||
): NativeAutoDecision | null {
|
||||
if (transcript.status !== 'ready' || !opts.sessionId || !opts.skillName ||
|
||||
!Number.isFinite(opts.commandStartedAt) || !Number.isFinite(opts.now) || opts.now < opts.commandStartedAt) return null;
|
||||
const at = (timestamp: string) => Date.parse(timestamp);
|
||||
const timely = (timestamp: string) => Number.isFinite(at(timestamp)) && at(timestamp) >= opts.commandStartedAt && at(timestamp) <= opts.now;
|
||||
const uses = tools.filter(e => e.kind === 'use' && e.sessionId === opts.sessionId && e.name === 'Skill' &&
|
||||
[opts.skillName, `gstack:${opts.skillName}`].includes(String(e.input?.skill ?? '')) && timely(e.timestamp));
|
||||
if (uses.length !== 1 || !uses[0]!.toolUseId) return null;
|
||||
const use = uses[0]!;
|
||||
const results = tools.filter(e => e.kind === 'result' && e.sessionId === opts.sessionId && e.toolUseId === use.toolUseId);
|
||||
if (results.length !== 1 || results[0]!.isError !== false || !timely(results[0]!.timestamp) || at(results[0]!.timestamp) < at(use.timestamp)) return null;
|
||||
// A native question actually surfaced; the annotation cannot erase it.
|
||||
if (transcript.calls.some(call => call.sessionId === opts.sessionId)) return null;
|
||||
const messages = transcript.assistantMessages.filter(m => m.sessionId === opts.sessionId);
|
||||
if (messages.some(m => !Number.isFinite(at(m.timestamp)) || at(m.timestamp) > opts.now)) return null;
|
||||
const loadedAt = at(results[0]!.timestamp);
|
||||
for (const message of messages) {
|
||||
if (at(message.timestamp) < loadedAt) continue;
|
||||
const match = assertedAnnotation(message.text, opts.skillName);
|
||||
if (!match) continue;
|
||||
const option = plain(match[2]!);
|
||||
const current = messages.filter(m => at(m.timestamp) >= at(message.timestamp)).map(m => m.text).join('\n\n');
|
||||
if (withdrawn(current, option)) continue;
|
||||
return { sessionId: opts.sessionId, skillToolUseId: use.toolUseId, timestamp: message.timestamp,
|
||||
summary: plain(match[1]!), option, annotation: match[0] };
|
||||
}
|
||||
return null;
|
||||
}
|
||||
@@ -0,0 +1,207 @@
|
||||
/** Inspect execution events, never prose, for completed outside-review evidence. */
|
||||
export interface OutsideExecution {
|
||||
command: string;
|
||||
output: string;
|
||||
succeeded: boolean;
|
||||
background?: {
|
||||
toolUseId: string;
|
||||
taskId: string;
|
||||
outputFile: string;
|
||||
outputToolUseId?: string;
|
||||
completion: 'pending' | 'task_notification' | 'TaskOutput';
|
||||
};
|
||||
}
|
||||
|
||||
function toolText(content: unknown): string {
|
||||
return typeof content === 'string' ? content : Array.isArray(content)
|
||||
? content.filter(block => block?.type === 'text' && typeof block.text === 'string').map(block => block.text).join('\n') : '';
|
||||
}
|
||||
|
||||
function sessionIdentity(event: any): string | null {
|
||||
if (event?.isSidechain === true || event?.parent_tool_use_id != null) return null;
|
||||
const id = event?.session_id ?? event?.sessionId;
|
||||
if (typeof id !== 'string' || !id || (event.sessionId && event.sessionId !== id)) return null;
|
||||
return id;
|
||||
}
|
||||
|
||||
/** Only native task events carry completion; quoted/user-authored notifications do not. */
|
||||
function taskNotification(event: any) {
|
||||
if (event?.type === 'system' && event.subtype === 'task_notification') return {
|
||||
taskId: event.task_id, toolUseId: event.tool_use_id, outputFile: event.output_file,
|
||||
status: event.status, summary: event.summary,
|
||||
};
|
||||
if (event?.type !== 'attachment' || event.attachment?.type !== 'queued_command' ||
|
||||
event.attachment.commandMode !== 'task-notification') return null;
|
||||
const text = event.attachment.prompt;
|
||||
if (typeof text !== 'string') return null;
|
||||
const match = /^<task-notification>\s*<task-id>([^<>\r\n]+)<\/task-id>\s*<tool-use-id>([^<>\r\n]+)<\/tool-use-id>\s*<output-file>([^<>\r\n]+)<\/output-file>\s*<status>(completed|failed|stopped)<\/status>\s*<summary>([^<>]+)<\/summary>\s*<\/task-notification>$/.exec(text);
|
||||
return match ? { taskId: match[1], toolUseId: match[2], outputFile: match[3], status: match[4], summary: match[5] } : null;
|
||||
}
|
||||
|
||||
/** A complete native Read starts at line one and includes the task's exit footer. */
|
||||
function completedTaskRead(text: string): string | null {
|
||||
const lines = text.replace(/\r?\n$/, '').split(/\r?\n/);
|
||||
const numbered = lines.map(line => /^\s*(\d+)\t(.*)$/.exec(line));
|
||||
if (!numbered.length || numbered.some((line, index) => !line || Number(line[1]) !== index + 1)) return null;
|
||||
const output = numbered.map(line => line![2]).join('\n');
|
||||
return /(?:^|\n)\[exited with code 0\]\s*$/.test(output) ? output : null;
|
||||
}
|
||||
|
||||
/** A literal cat cannot replace, truncate or redirect the acknowledged task file. */
|
||||
function literalTaskCat(command: string, outputFile: string): boolean {
|
||||
if (!/^\/[A-Za-z0-9_./-]+$/.test(outputFile)) return false;
|
||||
return ['', '-- '].some(option => [outputFile, `"${outputFile}"`, `'${outputFile}'`]
|
||||
.some(file => command.trim() === `cat ${option}${file}`));
|
||||
}
|
||||
|
||||
export function claudeOutsideExecutions(transcript: unknown[]): OutsideExecution[] {
|
||||
const calls = new Map<string, { command?: string; name: string; input: any; session: string | null }>();
|
||||
const tasks = new Map<string, {
|
||||
command: string; session: string; taskId: string; outputFile: string; failed: boolean;
|
||||
completion?: 'task_notification' | 'TaskOutput';
|
||||
outputs: Array<{ toolUseId: string; output: string; complete: boolean }>;
|
||||
}>();
|
||||
const results: OutsideExecution[] = [];
|
||||
for (const event of transcript as any[]) {
|
||||
const notice = taskNotification(event);
|
||||
const notified = notice && tasks.get(notice.toolUseId);
|
||||
if (notified && sessionIdentity(event) === notified.session && notice.taskId === notified.taskId && notice.outputFile === notified.outputFile) {
|
||||
if (notice.status === 'failed' || notice.status === 'stopped') notified.failed = true;
|
||||
else if (notice.status === 'completed') {
|
||||
const terminal = /^Background command [\s\S]+ completed \(exit code (-?\d+)\)$/.exec(notice.summary ?? '');
|
||||
if (terminal?.[1] === '0') notified.completion = 'task_notification';
|
||||
else if (terminal) notified.failed = true;
|
||||
}
|
||||
}
|
||||
const blocks = event?.message?.content;
|
||||
if (!Array.isArray(blocks)) continue;
|
||||
for (const block of blocks) {
|
||||
if (event.type === 'assistant' && block.type === 'tool_use' && typeof block.id === 'string' &&
|
||||
['Bash', 'Read', 'TaskOutput'].includes(block.name)) {
|
||||
calls.set(block.id, { command: block.input?.command, name: block.name, input: block.input, session: sessionIdentity(event) });
|
||||
}
|
||||
if (event.type === 'user' && block.type === 'tool_result' && calls.has(block.tool_use_id)) {
|
||||
const call = calls.get(block.tool_use_id)!;
|
||||
const output = toolText(block.content);
|
||||
if (call.name === 'Bash' && typeof call.command === 'string') {
|
||||
const readTask = [...tasks.values()].find(task => call.session === task.session &&
|
||||
sessionIdentity(event) === task.session && literalTaskCat(call.command!, task.outputFile));
|
||||
if (readTask) {
|
||||
const complete = block.is_error !== true && /(?:^|\n)\[exited with code 0\]\s*$/.test(output);
|
||||
readTask.outputs.push({ toolUseId: block.tool_use_id, output, complete });
|
||||
continue;
|
||||
}
|
||||
const acknowledgment = /^Command running in background with ID: ([a-zA-Z0-9_-]+)\. Output is being written to: ([^\r\n]+)\. You will be notified when it completes\. To check interim output, use Read on that file path\.(?:\nSession cwd remains (\/[^;\r\n]+); directory changes made by the backgrounded command do not apply to subsequent commands\.)?$/.exec(output);
|
||||
if (!acknowledgment) {
|
||||
const existing = tasks.get(block.tool_use_id);
|
||||
if (existing) {
|
||||
if (block.is_error === true) existing.failed = true;
|
||||
results.push({ command: existing.command, output, succeeded: false,
|
||||
background: { toolUseId: block.tool_use_id, taskId: existing.taskId,
|
||||
outputFile: existing.outputFile, completion: 'pending' } });
|
||||
continue;
|
||||
}
|
||||
const pending = /^Command running in background\b/.test(output) || event.toolUseResult?.backgroundTaskId;
|
||||
results.push({ command: call.command, output, succeeded: block.is_error !== true && !pending });
|
||||
continue;
|
||||
}
|
||||
const [, taskId, outputFile] = acknowledgment;
|
||||
const background = { toolUseId: block.tool_use_id, taskId: taskId!, outputFile: outputFile!, completion: 'pending' as const };
|
||||
results.push({ command: call.command, output, succeeded: false, background });
|
||||
const existing = tasks.get(block.tool_use_id);
|
||||
if (block.is_error === true || !call.session || sessionIdentity(event) !== call.session ||
|
||||
(acknowledgment[3] && typeof event.cwd === 'string' && acknowledgment[3] !== event.cwd) ||
|
||||
!outputFile!.endsWith(`/tasks/${taskId}.output`) || !outputFile!.startsWith('/')) {
|
||||
if (existing) existing.failed = true;
|
||||
continue;
|
||||
}
|
||||
if (existing) {
|
||||
// Replayed acknowledgments cannot reset completed output or a
|
||||
// terminal failure. One original call cannot launch a second task.
|
||||
if (existing.command !== call.command || existing.session !== call.session ||
|
||||
existing.taskId !== taskId || existing.outputFile !== outputFile) existing.failed = true;
|
||||
continue;
|
||||
}
|
||||
tasks.set(block.tool_use_id, { command: call.command, session: call.session, taskId: taskId!, outputFile: outputFile!, failed: false, outputs: [] });
|
||||
continue;
|
||||
}
|
||||
if (!call.session || sessionIdentity(event) !== call.session) continue;
|
||||
for (const task of tasks.values()) {
|
||||
if (task.session !== call.session) continue;
|
||||
if (call.name === 'Read' && call.input?.file_path === task.outputFile) {
|
||||
const complete = block.is_error !== true && (call.input.offset === undefined || call.input.offset === 1)
|
||||
? completedTaskRead(output) : null;
|
||||
task.outputs.push({ toolUseId: block.tool_use_id, output: complete ?? output, complete: complete !== null });
|
||||
}
|
||||
if (block.is_error === true) {
|
||||
if (call.name === 'TaskOutput' && call.input?.task_id === task.taskId) task.failed = true;
|
||||
continue;
|
||||
}
|
||||
if (call.name === 'TaskOutput' && call.input?.task_id === task.taskId) {
|
||||
// CLI's native TaskOutput envelope puts terminal status before the
|
||||
// output. Findings containing an exit-code string cannot supply it.
|
||||
const terminal = /^<retrieval_status>success<\/retrieval_status>\s*<task_id>([^<>]+)<\/task_id>\s*<task_type>local_bash<\/task_type>\s*<status>(completed|failed|stopped)<\/status>\s*<exit_code>(-?\d+)<\/exit_code>\s*<output>\n([\s\S]*)\n<\/output>\s*$/.exec(output);
|
||||
if (!terminal || terminal[1] !== task.taskId) continue;
|
||||
const complete = terminal[2] === 'completed' && terminal[3] === '0';
|
||||
if (!complete) task.failed = true;
|
||||
else {
|
||||
task.completion = 'TaskOutput';
|
||||
}
|
||||
task.outputs.push({ toolUseId: block.tool_use_id, output: terminal[4]!, complete });
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for (const [toolUseId, task] of tasks) {
|
||||
for (const output of task.outputs) {
|
||||
results.push({ command: task.command, output: output.output,
|
||||
succeeded: output.complete && Boolean(task.completion) && !task.failed,
|
||||
background: { toolUseId, taskId: task.taskId, outputFile: task.outputFile,
|
||||
outputToolUseId: output.toolUseId, completion: task.completion ?? 'pending' } });
|
||||
}
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
export function codexOutsideExecutions(lines: string[]): OutsideExecution[] {
|
||||
return lines.flatMap(line => {
|
||||
try {
|
||||
const event = JSON.parse(line);
|
||||
const item = event.item;
|
||||
if (event.type !== 'item.completed' || item?.type !== 'command_execution') return [];
|
||||
if (typeof item.command !== 'string' || typeof item.aggregated_output !== 'string') return [];
|
||||
return [{ command: item.command, output: item.aggregated_output, succeeded: item.exit_code === 0 }];
|
||||
} catch { return []; }
|
||||
});
|
||||
}
|
||||
|
||||
/** Host command diagnostics retain wrapper executions without granting outside credit. */
|
||||
export function codexExecutionTranscript(executions: OutsideExecution[]) {
|
||||
return executions.map(({ command, output, succeeded }) => ({
|
||||
type: 'host_execution' as const, host: 'codex' as const, command, output, succeeded,
|
||||
}));
|
||||
}
|
||||
|
||||
function outsideInvocation(provider: 'codex' | 'claude-code'): RegExp {
|
||||
return provider === 'codex' ? /\bcodex\s+(?:exec|review)\b/
|
||||
: /\bgstack-claude-code(?:['"])?\s+--/;
|
||||
}
|
||||
|
||||
/** Preserve provider results, including failures, without unrelated shell/config events. */
|
||||
export function outsideExecutionTranscript(executions: OutsideExecution[], provider: 'codex' | 'claude-code') {
|
||||
return executions.filter(({ command }) => outsideInvocation(provider).test(command))
|
||||
.map(({ command, output, succeeded, background }) => ({
|
||||
type: 'outside_execution' as const, provider, command, output, succeeded,
|
||||
...(background ? { background } : {}),
|
||||
}));
|
||||
}
|
||||
|
||||
export function foundInvoiceAuthorizationDefect(executions: OutsideExecution[], provider: 'codex' | 'claude-code'): boolean {
|
||||
return executions.some(({ command, output, succeeded }) => succeeded
|
||||
&& outsideInvocation(provider).test(command)
|
||||
&& /invoice/i.test(output)
|
||||
&& /owner|ownership|unauthori[sz]ed|another user|cross[- ](?:tenant|user)|access control|authorization/i.test(output)
|
||||
&& !/(?:^|\n)\s*(?:I|We)\s+(?:cannot|can['’]t|won['’]t|am unable to|are unable to)\s+(?:review|perform|complete|provide|conduct|help with)\b/i.test(output)
|
||||
&& !/OUTSIDE_STATUS:\s*(?:unavailable|disabled|skipped)/.test(output));
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
/** Extract the installed review workflow, handling carved and inline host renders. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { extractSkillSections } from './skill-fixture';
|
||||
|
||||
export function installOutsideReviewFixture(rendered: string, host: 'claude' | 'codex', repo: string, runtimeRoot: string): string {
|
||||
const source = host === 'claude' ? join(rendered, 'review') : join(rendered, '.agents', 'skills', 'gstack-review');
|
||||
const name = host === 'claude' ? 'review' : 'gstack-review';
|
||||
const destination = join(repo, host === 'claude' ? '.claude' : '.agents', 'skills', name);
|
||||
mkdirSync(destination, { recursive: true });
|
||||
const head = extractSkillSections(source, ['Step 0: Detect platform and base branch', 'Step 3: Get the diff']);
|
||||
const sectionPath = join(source, 'sections', 'adversarial.md');
|
||||
const section = existsSync(sectionPath) ? readFileSync(sectionPath, 'utf8')
|
||||
: extractSkillSections(source, ['Step 5.7: Adversarial review (always-on)']).replace(/^---\r?\n[\s\S]*?\r?\n---\r?\n/, '');
|
||||
if (!section.includes('Adversarial review (always-on)')) throw new Error(`Missing adversarial workflow: ${source}`);
|
||||
// Runtime paths are the only fixture substitution. Provider selection,
|
||||
// caller controls, prompt, probes, and execution code stay generated verbatim.
|
||||
const content = (head + '\n' + section)
|
||||
.replaceAll('~/.claude/skills/gstack', runtimeRoot)
|
||||
.replaceAll('$HOME/.claude/skills/gstack', runtimeRoot)
|
||||
.replaceAll('${GSTACK_BIN}', join(runtimeRoot, 'bin'))
|
||||
.replaceAll('$GSTACK_BIN', join(runtimeRoot, 'bin'))
|
||||
.replaceAll('$GSTACK_ROOT', runtimeRoot);
|
||||
writeFileSync(join(destination, 'SKILL.md'), content);
|
||||
return destination;
|
||||
}
|
||||
|
||||
// Each paid invocation must begin at the same committed authorization defect.
|
||||
function git(cwd: string, ...args: string[]) {
|
||||
const result = Bun.spawnSync(['git', ...args], { cwd, stdout: 'pipe', stderr: 'pipe', timeout: 10_000 });
|
||||
if (result.exitCode !== 0) throw new Error(`git ${args[0]}: ${result.stderr.toString()}`);
|
||||
}
|
||||
|
||||
export function createOutsideReviewRepo(fixtureRoot: string, host: 'claude' | 'codex'): string {
|
||||
// Bun retries reuse beforeAll state; each attempt needs a new git repository.
|
||||
const dir = fs.mkdtempSync(path.join(fixtureRoot, `${host}-`));
|
||||
git(dir, 'init', '-b', 'main');
|
||||
git(dir, 'config', 'user.email', 'eval@example.com');
|
||||
git(dir, 'config', 'user.name', 'Outside Voice Eval');
|
||||
const safe = `export async function readPrivateInvoice(db, actor, invoiceId) {
|
||||
const invoice = await db.invoice.findUnique({ where: { id: invoiceId } });
|
||||
if (!invoice) return null;
|
||||
if (invoice.ownerId !== actor.id) throw new Error('Forbidden');
|
||||
return { amount: invoice.amount, bankAccount: invoice.bankAccount };
|
||||
}
|
||||
`;
|
||||
fs.writeFileSync(path.join(dir, 'invoice.ts'), safe);
|
||||
git(dir, 'add', 'invoice.ts');
|
||||
git(dir, 'commit', '-m', 'Protect private invoices by owner');
|
||||
git(dir, 'update-ref', 'refs/remotes/origin/main', 'HEAD');
|
||||
git(dir, 'checkout', '-b', 'feature/invoice-lookup');
|
||||
fs.writeFileSync(path.join(dir, 'invoice.ts'), safe.replace(" if (invoice.ownerId !== actor.id) throw new Error('Forbidden');\n", ''));
|
||||
git(dir, 'add', 'invoice.ts');
|
||||
git(dir, 'commit', '-m', 'Simplify invoice lookup');
|
||||
return dir;
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
/** Fixture-only witness of the real launcher; not a same-UID tamper boundary. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { createHash, randomUUID } from 'node:crypto';
|
||||
import { spawn } from 'node:child_process';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
import { validateOutsideReview } from '../../lib/outside-review-result';
|
||||
import type { OutsideExecution } from './outside-voice-evidence';
|
||||
|
||||
const LIMIT = 32 * 1024 * 1024;
|
||||
const hash = (data: string | Buffer) => createHash('sha256').update(data).digest('hex');
|
||||
type Binding = { nonce: string; cwd: string; launcher: string; launcherHash: string; records: string; createdAt: number };
|
||||
type Receipt = Binding & { id: string; pid: number; startTicks: string | null; argv: string[];
|
||||
actualCwd: string; startedAt: number; endedAt: number; exitCode: number | null; signal: string | null;
|
||||
inputHash: string; inputBytes: number; stdout: string; stderr: string; stdoutHash: string; stderrHash: string };
|
||||
|
||||
/** Called only by the fixture's executable shim, never by the review model. */
|
||||
export async function recordOutsideInvocation(binding: Binding, argv: string[]): Promise<void> {
|
||||
// Do not substitute argv, cwd, environment, access, model, tools or deadlines.
|
||||
const startedAt = Date.now();
|
||||
const child = spawn(binding.launcher, argv, { cwd: process.cwd(), env: process.env, stdio: ['pipe', 'pipe', 'pipe'] });
|
||||
let startTicks: string | null = null;
|
||||
try { startTicks = fs.readFileSync(`/proc/${child.pid}/stat`, 'utf8').split(') ').at(-1)!.split(' ')[19]!; } catch { /* Unknown off Linux or after an early exit. */ }
|
||||
const input = createHash('sha256'); let inputBytes = 0; let inputEnded = false; let bytes = 0; let overflow = false;
|
||||
const stdout: Buffer[] = [], stderr: Buffer[] = [];
|
||||
process.stdin.on('data', (chunk: Buffer) => { input.update(chunk); inputBytes += chunk.length; });
|
||||
process.stdin.once('end', () => { inputEnded = true; });
|
||||
process.stdin.pipe(child.stdin!);
|
||||
child.stdin!.on('error', () => { /* The real launcher may reject arguments before reading stdin. */ });
|
||||
for (const [stream, output, saved] of [[child.stdout!, process.stdout, stdout], [child.stderr!, process.stderr, stderr]] as const) {
|
||||
stream.on('data', (chunk: Buffer) => {
|
||||
output.write(chunk); bytes += chunk.length;
|
||||
if (bytes <= LIMIT) saved.push(Buffer.from(chunk)); else overflow = true;
|
||||
});
|
||||
}
|
||||
const forward = (signal: NodeJS.Signals) => { try { child.kill(signal); } catch { /* Already exited. */ } };
|
||||
const interrupt = () => forward('SIGINT'), terminate = () => forward('SIGTERM');
|
||||
process.on('SIGINT', interrupt); process.on('SIGTERM', terminate);
|
||||
const ended = await new Promise<{ code: number | null; signal: NodeJS.Signals | null; error?: string }>(resolve => {
|
||||
child.once('error', error => resolve({ code: null, signal: null, error: error.message }));
|
||||
child.once('close', (code, signal) => resolve({ code, signal }));
|
||||
});
|
||||
process.removeListener('SIGINT', interrupt); process.removeListener('SIGTERM', terminate);
|
||||
process.stdin.unpipe(child.stdin!); process.stdin.pause();
|
||||
if (!ended.error && child.pid && !overflow && inputEnded) try {
|
||||
const outputBytes = Buffer.concat(stdout), diagnosticBytes = Buffer.concat(stderr);
|
||||
const output = outputBytes.toString(), diagnostic = diagnosticBytes.toString();
|
||||
if (!Buffer.from(output).equals(outputBytes) || !Buffer.from(diagnostic).equals(diagnosticBytes)) throw new Error('Non-UTF8 result');
|
||||
const receipt: Receipt = { ...binding, id: randomUUID(), pid: child.pid, startTicks, argv,
|
||||
actualCwd: process.cwd(), startedAt, endedAt: Date.now(), exitCode: ended.code, signal: ended.signal,
|
||||
inputHash: input.digest('hex'), inputBytes, stdout: output, stderr: diagnostic,
|
||||
stdoutHash: hash(output), stderrHash: hash(diagnostic) };
|
||||
const file = path.join(binding.records, receipt.id + '.json');
|
||||
const temporary = file + '.tmp';
|
||||
fs.writeFileSync(temporary, JSON.stringify(receipt), { flag: 'wx', mode: 0o600 }); fs.renameSync(temporary, file);
|
||||
} catch { /* Observation failure never changes the real launcher's result. */ }
|
||||
if (ended.signal) process.kill(process.pid, ended.signal);
|
||||
else process.exitCode = ended.code ?? 1;
|
||||
}
|
||||
|
||||
export function createOutsideReceiptRuntime(runtimeRoot: string, fixtureRoot: string, cwd: string) {
|
||||
const folder = fs.mkdtempSync(path.join(fixtureRoot, 'outside-receipt-'));
|
||||
const facade = path.join(folder, 'runtime'), records = path.join(folder, 'records');
|
||||
fs.mkdirSync(facade); fs.mkdirSync(records, { mode: 0o700 }); fs.mkdirSync(path.join(facade, 'bin'));
|
||||
for (const entry of fs.readdirSync(runtimeRoot)) if (entry !== 'bin')
|
||||
fs.symlinkSync(path.join(runtimeRoot, entry), path.join(facade, entry));
|
||||
for (const entry of fs.readdirSync(path.join(runtimeRoot, 'bin'))) if (entry !== 'gstack-claude-code')
|
||||
fs.symlinkSync(path.join(runtimeRoot, 'bin', entry), path.join(facade, 'bin', entry));
|
||||
const launcher = path.join(runtimeRoot, 'bin/gstack-claude-code');
|
||||
const binding: Binding = { nonce: randomUUID(), cwd, launcher, launcherHash: hash(fs.readFileSync(launcher)), records, createdAt: Date.now() };
|
||||
const shim = path.join(facade, 'bin/gstack-claude-code');
|
||||
fs.writeFileSync(shim, `#!${process.execPath}\nimport {recordOutsideInvocation} from ${JSON.stringify(pathToFileURL(import.meta.path).href)};\nawait recordOutsideInvocation(${JSON.stringify(binding)},process.argv.slice(2));\n`, { mode: 0o755 });
|
||||
const shimHash = hash(fs.readFileSync(shim));
|
||||
return {
|
||||
runtimeRoot: facade,
|
||||
read() {
|
||||
const receipts: Receipt[] = [], executions: OutsideExecution[] = [];
|
||||
let names: string[];
|
||||
try {
|
||||
if (hash(fs.readFileSync(shim)) !== shimHash || hash(fs.readFileSync(launcher)) !== binding.launcherHash) return { receipts, executions };
|
||||
names = fs.readdirSync(records);
|
||||
} catch { return { receipts, executions }; }
|
||||
for (const name of names) {
|
||||
if (!/^[a-f0-9-]{36}\.json$/.test(name)) continue;
|
||||
const file = path.join(records, name);
|
||||
try {
|
||||
const info = fs.lstatSync(file);
|
||||
if (!info.isFile() || info.size > LIMIT * 3 || fs.realpathSync(file) !== file) continue;
|
||||
const r: Receipt = JSON.parse(fs.readFileSync(file, 'utf8'));
|
||||
if (Object.entries(binding).some(([key, value]) => r[key as keyof Receipt] !== value) ||
|
||||
r.id + '.json' !== name || r.actualCwd !== cwd || !Number.isSafeInteger(r.pid) || r.pid <= 0 ||
|
||||
(process.platform === 'linux' && !/^\d+$/.test(r.startTicks ?? '')) ||
|
||||
!Number.isSafeInteger(r.startedAt) || !Number.isSafeInteger(r.endedAt) ||
|
||||
r.startedAt < binding.createdAt || r.endedAt < r.startedAt || r.endedAt > Date.now() ||
|
||||
!Array.isArray(r.argv) || !r.argv.every(x => typeof x === 'string') ||
|
||||
r.argv.filter(x => x === '--cwd').length !== 1 || r.argv[r.argv.indexOf('--cwd') + 1] !== cwd ||
|
||||
typeof r.stdout !== 'string' || typeof r.stderr !== 'string' ||
|
||||
hash(r.stdout) !== r.stdoutHash || hash(r.stderr) !== r.stderrHash ||
|
||||
!/^[a-f0-9]{64}$/.test(r.inputHash) || !Number.isSafeInteger(r.inputBytes) || r.inputBytes <= 0) continue;
|
||||
receipts.push(r);
|
||||
const envelope = JSON.parse(r.stdout);
|
||||
if (r.exitCode !== 0 || r.signal !== null || envelope?.status !== 'completed' ||
|
||||
envelope.provider !== 'claude-code' || envelope.exit_code !== 0 ||
|
||||
typeof envelope.session_id !== 'string' || !/^[A-Za-z0-9_-]+$/.test(envelope.session_id) ||
|
||||
typeof envelope.result !== 'string' || !validateOutsideReview(envelope.result, 'review').completed) continue;
|
||||
// This command describes the recorded real argv, not the outer shell text.
|
||||
const quote = (word: string) => /^[A-Za-z0-9_\/.:-]+$/.test(word) ? word : "'" + word.replaceAll("'", "'\\''") + "'";
|
||||
executions.push({ command: [launcher, ...r.argv].map(quote).join(' '), output: envelope.result, succeeded: true });
|
||||
} catch { /* Malformed, foreign or incomplete receipt supplies no evidence. */ }
|
||||
}
|
||||
return { receipts, executions };
|
||||
},
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { getProjectEvalDir } from './eval-store';
|
||||
|
||||
interface PlanCountSnapshot {
|
||||
skillName: string;
|
||||
observation: object;
|
||||
raw: string;
|
||||
visible: string;
|
||||
viewport?: string;
|
||||
cwd: string;
|
||||
claudeConfigDir: string | null;
|
||||
}
|
||||
|
||||
/** One owned directory per count attempt; periodic captures replace files atomically. */
|
||||
export function createPlanCountSnapshotWriter(env: NodeJS.ProcessEnv = process.env):
|
||||
(input: PlanCountSnapshot) => { artifactDir?: string; artifactError?: string } {
|
||||
let artifactDir: string | undefined;
|
||||
return (input) => {
|
||||
if (!env.EVALS_RUN_ID) return {};
|
||||
try {
|
||||
if (!artifactDir) {
|
||||
const segment = (text: string) => text.replace(/[^a-zA-Z0-9_-]/g, '_').slice(0, 120) || 'run';
|
||||
const root = path.resolve(env.GSTACK_EVAL_DIR || getProjectEvalDir(), 'pty-count', segment(env.EVALS_RUN_ID));
|
||||
fs.mkdirSync(root, { recursive: true, mode: 0o700 });
|
||||
artifactDir = fs.mkdtempSync(path.join(root, `${segment(input.skillName)}-${Date.now()}-`));
|
||||
}
|
||||
const write = (name: string, content: string) => {
|
||||
const target = path.join(artifactDir!, name);
|
||||
fs.writeFileSync(`${target}.tmp`, content, { mode: 0o600 });
|
||||
fs.renameSync(`${target}.tmp`, target);
|
||||
};
|
||||
write('terminal.raw.log', input.raw);
|
||||
write('terminal.visible.log', input.visible);
|
||||
if (input.viewport !== undefined) write('terminal.screen.log', input.viewport);
|
||||
write('observation.json', JSON.stringify({
|
||||
...input.observation, artifactDir,
|
||||
capture: { skill: input.skillName, runId: env.EVALS_RUN_ID, cwd: input.cwd,
|
||||
claudeConfigDir: input.claudeConfigDir, at: new Date().toISOString() },
|
||||
}, null, 2) + '\n');
|
||||
return { artifactDir };
|
||||
} catch (error) {
|
||||
// Preserve any partial evidence and the original test outcome; make the
|
||||
// write failure visible instead of claiming diagnostics were retained.
|
||||
return { artifactDir, artifactError: String(error) };
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/** Keep a single snapshot outside the temporary fixture that setup later removes. */
|
||||
export function persistPlanCountSnapshot(input: PlanCountSnapshot, env: NodeJS.ProcessEnv = process.env) {
|
||||
return createPlanCountSnapshotWriter(env)(input);
|
||||
}
|
||||
@@ -0,0 +1,180 @@
|
||||
/** Content-free native file-permission identity for disposable count fixtures. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import type { PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
const MAX_RECORD_BYTES = 64 * 1024;
|
||||
export interface FilePermissionEpoch { pendingId: string; completedId: string | null; completedIds?: string[] }
|
||||
const identifier = (v: unknown): v is string => typeof v === 'string' && /^[A-Za-z0-9_-]{1,160}$/.test(v);
|
||||
const quote = (v: string) => `'${(process.platform === 'win32' ? v.replaceAll('\\', '/') : v).replaceAll("'", "'\\''")}'`;
|
||||
const scoped = (file: unknown, config: string, session: string) => {
|
||||
if (typeof file !== 'string' || !path.isAbsolute(file)) return false;
|
||||
const rel = path.relative(path.join(config, 'projects'), file).split(path.sep);
|
||||
return rel.length === 2 && rel[0] !== '..' && rel[0] !== '.' && rel[1] === `${session}.jsonl`;
|
||||
};
|
||||
export function createFilePermissionRecorder(cwd: string, config: string, expected: string) {
|
||||
const relative = path.relative(os.tmpdir(), expected);
|
||||
if (!path.isAbsolute(expected) || !relative || relative === '..' || relative.startsWith('..' + path.sep) || path.isAbsolute(relative)) return undefined;
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-file-permission-'));
|
||||
const file = path.join(dir, 'state.json');
|
||||
const command = [process.execPath, import.meta.path, '--record', file, cwd, config, expected].map(quote).join(' ');
|
||||
const hook = { matcher: '^(Write|Edit)$', hooks: [{ type: 'command', command, timeout: 5 }] };
|
||||
return { file, hooks: { PreToolUse: [hook], PostToolUse: [hook], PostToolUseFailure: [hook] },
|
||||
dispose: () => fs.rmSync(dir, {recursive:true,force:true}) };
|
||||
}
|
||||
|
||||
/** No stdout, permission decision, input rewrite, model context, or file content. */
|
||||
export function recordFilePermission(input: string, file: string, cwd: string, config: string, expected: string) {
|
||||
try {
|
||||
if (Buffer.byteLength(input) > 4 * 1024 * 1024) throw Error('oversized hook');
|
||||
const e = JSON.parse(input);
|
||||
if (e?.agent_id !== undefined || e?.cwd !== cwd) return;
|
||||
if (!['PreToolUse','PostToolUse','PostToolUseFailure'].includes(e.hook_event_name) ||
|
||||
!['Write','Edit'].includes(e.tool_name) || !identifier(e.session_id) || !identifier(e.tool_use_id) ||
|
||||
!scoped(e.transcript_path,config,e.session_id) || e.tool_input?.file_path !== expected) return;
|
||||
let old: any = {};
|
||||
if (fs.existsSync(file)) {
|
||||
const stat = fs.lstatSync(file);
|
||||
if (!stat.isFile() || stat.size > MAX_RECORD_BYTES) throw Error('invalid record');
|
||||
old = JSON.parse(fs.readFileSync(file,'utf8'));
|
||||
}
|
||||
const id = `${e.session_id}:${e.tool_use_id}`;
|
||||
const timestamp = new Date().toISOString();
|
||||
const same = old.cwd === cwd && old.expected === expected && old.sessionId === e.session_id;
|
||||
const state = { cwd, expected, sessionId: e.session_id, transcriptPath: e.transcript_path,
|
||||
seenIds: same && Array.isArray(old.seenIds) ? old.seenIds : [],
|
||||
pendingId: same ? old.pendingId ?? null : null,
|
||||
completedId: same ? old.completedId ?? null : null,
|
||||
completedIds: same && Array.isArray(old.completedIds) ? old.completedIds : [],
|
||||
timestamp };
|
||||
if (e.hook_event_name === 'PreToolUse') {
|
||||
// Replayed requests, including failed and older completed IDs, never reopen.
|
||||
if (state.seenIds.includes(id)) return;
|
||||
if (state.seenIds.length >= 128) throw Error('too many file requests');
|
||||
state.seenIds.push(id);
|
||||
state.pendingId = id;
|
||||
} else {
|
||||
// An unrelated/late result cannot overwrite the current request epoch.
|
||||
if (state.pendingId !== id) return;
|
||||
state.pendingId = null;
|
||||
if (e.hook_event_name === 'PostToolUse') {
|
||||
state.completedId = id;
|
||||
// Polling may miss automatically permitted edits between two menus.
|
||||
// Keep each exact success, bounded by the same 128-request limit.
|
||||
if (!state.completedIds.includes(id)) state.completedIds.push(id);
|
||||
}
|
||||
}
|
||||
fs.writeFileSync(file+'.tmp',JSON.stringify(state)+'\n',{mode:0o600});
|
||||
fs.renameSync(file+'.tmp',file);
|
||||
} catch { try { fs.rmSync(file,{force:true}); } catch {} }
|
||||
}
|
||||
|
||||
/** A long diff can crop its path header; the native access choice repeats the directory. */
|
||||
function croppedEditTarget(screen: string, cwd: string, expected: string): string | undefined {
|
||||
const text = screen.replace(/\r+\n?/g, '\n');
|
||||
// Cropping may begin inside a wrapped added/deleted diff row (four/five-space gutter).
|
||||
// Still require numbered rows below and the full native footer; never a quoted AUQ.
|
||||
// The heading can be cropped one row earlier, leaving the complete path.
|
||||
// Keep it only when the existing menu independently identifies that target.
|
||||
const header = /^ {0,3}([^\n]+)\n[╌─━]{3,}[ \t]*\n/.exec(text);
|
||||
const headerPath = header?.[1]?.trim();
|
||||
const pathOnly = headerPath && (path.isAbsolute(headerPath) || /^\.\.?[/\\]/.test(headerPath));
|
||||
// A crop can start on the single native rule immediately above the diff.
|
||||
let diff = pathOnly ? text.slice(header![0].length) : text.replace(/^[╌─━]{3,}[ \t]*\n/, '');
|
||||
// A wrapped unchanged row has no +/- marker. Its visible tail must belong
|
||||
// to the preceding line of the exact current owned file, not arbitrary prose.
|
||||
let continuation = /^( +)([^+\-\s][^\n]*)\n(?=( {0,3}[1-9]\d* ))/.exec(diff);
|
||||
// A normal numbered diff row is not a newly recognized wrapped tail.
|
||||
if (continuation && continuation[1]!.length !== 6 &&
|
||||
/^[1-9]\d* [ +\-]/.test(continuation[2]!)) continuation = null;
|
||||
// Preserve the existing six-space crop. Other native gutters must align
|
||||
// with the next unchanged row's actual padding and line-number width.
|
||||
if (continuation && continuation[1]!.length !== 6 &&
|
||||
continuation[1]!.length !== continuation[3]!.length) return undefined;
|
||||
if (continuation) diff = diff.slice(continuation[0].length);
|
||||
if (!/^(?:\s*\d+\s+[ +\-]?| {4,5}[+\-])/.test(diff) || /[☐□]|^\s*(?:>|`{3}|~{3})/m.test(text)) return undefined;
|
||||
const prompt = [...text.matchAll(/^ {0,3}Do you want to make this edit to ([^\n?\/\\]+)\?[ \t]*\n([\s\S]*)$/gm)].at(-1);
|
||||
if (!prompt || (text.slice(0, prompt.index).match(/^\s*\d+\s+/gm)?.length ?? 0) < 2) return undefined;
|
||||
// The unselected option supplies path identity only. Input remains one-time Yes.
|
||||
// A redraw can leave this exact keyboard-hint tail on the unselected No row.
|
||||
// It does not change the selected one-time Yes or authorize another action.
|
||||
const choices = /^ {0,3}❯[ \t]*1\.[ \t]*Yes[ \t]*\n\s*2\.[ \t]*Yes,\s+and\s+switch\s+to\s+accept\s+edits\s+\(auto-approve\s+file\s+edits\s+and\s+common\s+file\s+commands\)\s+for\s+this\s+session;\s+Yes,\s+and\s+always\s+allow\s+access\s+to\s+([^\r\n]+?)\s+for\s+this\s+session(?:\s*\(shift\+tab\))?\s*\n\s*3\.[ \t]*No(?:hift\+tab\))?[ \t]*\n\s*Esc to cancel [·•] Tab to amend\s*$/.exec(prompt[2]!);
|
||||
const directory = choices?.[1]?.trim();
|
||||
if (!directory || !path.isAbsolute(directory)) return undefined;
|
||||
const target = path.join(directory, prompt[1]!.trim());
|
||||
if (continuation) {
|
||||
if (target !== expected) return undefined;
|
||||
const nextLine = Number(continuation[3]!.trim());
|
||||
if (!Number.isSafeInteger(nextLine) || nextLine < 2) return undefined;
|
||||
try {
|
||||
const stat = fs.lstatSync(target);
|
||||
if (!stat.isFile() || stat.size > MAX_RECORD_BYTES) return undefined;
|
||||
const fd = fs.openSync(target, fs.constants.O_RDONLY | (fs.constants.O_NOFOLLOW ?? 0));
|
||||
try {
|
||||
const opened = fs.fstatSync(fd);
|
||||
if (!opened.isFile() || opened.dev !== stat.dev || opened.ino !== stat.ino || opened.size > MAX_RECORD_BYTES) return undefined;
|
||||
const bytes = Buffer.alloc(MAX_RECORD_BYTES + 1);
|
||||
const length = fs.readSync(fd, bytes, 0, bytes.length, 0);
|
||||
if (length !== opened.size || length > MAX_RECORD_BYTES) return undefined;
|
||||
const prior = bytes.subarray(0, length).toString('utf8').split(/\r?\n/)[nextLine - 2];
|
||||
if (!prior?.trimEnd().endsWith(continuation[2]!.trimEnd())) return undefined;
|
||||
} finally { fs.closeSync(fd); }
|
||||
} catch { return undefined; }
|
||||
}
|
||||
return !pathOnly || path.resolve(cwd, headerPath!) === target ? target : undefined;
|
||||
}
|
||||
|
||||
/** Undefined leaves other permissions alone; null keeps this report pane waiting. */
|
||||
export function currentFilePermissionEpoch(file: string | undefined, expected: string | undefined,
|
||||
cwd: string, config: string | null, startedAt: number, transcript: PlanCountTranscript,
|
||||
screen: string): FilePermissionEpoch | null | undefined {
|
||||
if (!file || !expected || !config) return undefined;
|
||||
const panel = [...screen.matchAll(/(?:^|\n) {0,3}(?:Edit|Write) file[ \t]*\n {0,3}([^\n]+)\n/g)].at(-1);
|
||||
const target = panel ? path.resolve(cwd,panel[1]!.trim()) : croppedEditTarget(screen, cwd, expected);
|
||||
if (target !== expected) {
|
||||
// A foreign path with this report's basename cannot fall back to a stale
|
||||
// owned grant. An incomplete owned menu also waits for full path identity.
|
||||
const prompt = [...screen.matchAll(/^ {0,3}Do you want to make this edit to ([^\n?\/\\]+)\?[ \t]*$/gm)].at(-1);
|
||||
return (target && path.basename(target) === path.basename(expected)) ||
|
||||
prompt?.[1]?.trim() === path.basename(expected) ? null : undefined;
|
||||
}
|
||||
try {
|
||||
const stat = fs.lstatSync(file);
|
||||
if (!stat.isFile() || stat.size > MAX_RECORD_BYTES || transcript.status !== 'ready') return null;
|
||||
const r = JSON.parse(fs.readFileSync(file,'utf8'));
|
||||
const sessions = new Set([...transcript.calls.map(c=>c.sessionId),...transcript.assistantMessages.map(m=>m.sessionId)]);
|
||||
const time = Date.parse(r.timestamp);
|
||||
const validId = (id: unknown) => typeof id === 'string' && id.startsWith(r.sessionId+':') && identifier(id.slice(r.sessionId.length+1));
|
||||
if (r.cwd !== cwd || r.expected !== expected || !identifier(r.sessionId) || sessions.size !== 1 || !sessions.has(r.sessionId) ||
|
||||
!scoped(r.transcriptPath,config,r.sessionId) || !Number.isFinite(time) || time < startedAt || time > Date.now() ||
|
||||
!validId(r.pendingId) || (r.completedId !== null && !validId(r.completedId)) || r.pendingId === r.completedId ||
|
||||
!Array.isArray(r.seenIds) || r.seenIds.length > 128 || !r.seenIds.every(validId) ||
|
||||
new Set(r.seenIds).size !== r.seenIds.length || !r.seenIds.includes(r.pendingId) ||
|
||||
!Array.isArray(r.completedIds) || r.completedIds.length > 128 || !r.completedIds.every(validId) ||
|
||||
new Set(r.completedIds).size !== r.completedIds.length ||
|
||||
r.completedIds.some((id: string) => id === r.pendingId || !r.seenIds.includes(id)) ||
|
||||
(r.completedId === null ? r.completedIds.length !== 0 : r.completedIds.at(-1) !== r.completedId)) return null;
|
||||
return {pendingId:r.pendingId,completedId:r.completedId,completedIds:r.completedIds};
|
||||
} catch { return null; }
|
||||
}
|
||||
|
||||
/** Check every owned path before a same-basename block closes the whole menu. */
|
||||
export function currentFilePermissionBinding<T extends {file: string; expected: string}>(
|
||||
bindings: readonly T[], cwd: string, config: string | null, startedAt: number,
|
||||
transcript: PlanCountTranscript, screen: string,
|
||||
): {binding: T; epoch: FilePermissionEpoch} | null | undefined {
|
||||
let blocked = false;
|
||||
for (const binding of bindings) {
|
||||
const epoch = currentFilePermissionEpoch(binding.file, binding.expected, cwd, config, startedAt, transcript, screen);
|
||||
if (epoch) return {binding, epoch};
|
||||
if (epoch === null) blocked = true;
|
||||
}
|
||||
return blocked ? null : undefined;
|
||||
}
|
||||
|
||||
if (import.meta.main && process.argv[2] === '--record') {
|
||||
try { const [file,cwd,config,expected] = process.argv.slice(3);
|
||||
if (file && cwd && config && expected) recordFilePermission(await Bun.stdin.text(),file,cwd,config,expected);
|
||||
} catch { /* silent observation never changes native permission decisions */ }
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import { getHermeticDirs } from './hermetic-env';
|
||||
|
||||
/** Disposable config for evals that explicitly cover native review only. */
|
||||
export function createNativeReviewState(): {
|
||||
env: Record<string, string>;
|
||||
cleanup(): void;
|
||||
} {
|
||||
const sharedState = getHermeticDirs().gstackHome;
|
||||
const stateRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-native-review-state-'));
|
||||
const cleanup = () => fs.rmSync(stateRoot, { recursive: true, force: true });
|
||||
try {
|
||||
for (const entry of fs.readdirSync(sharedState, { withFileTypes: true })) {
|
||||
// Keep onboarding seeds; never copy sibling review logs/artifacts.
|
||||
if (entry.isFile() && (entry.name === '.activated' ||
|
||||
/^\..*(?:-seen|-prompted|-shown)$/.test(entry.name) ||
|
||||
entry.name.startsWith('.feature-prompted-'))) {
|
||||
fs.copyFileSync(path.join(sharedState, entry.name), path.join(stateRoot, entry.name));
|
||||
}
|
||||
}
|
||||
const config = fs.readFileSync(path.join(sharedState, 'config.yaml'), 'utf8')
|
||||
.replace(/^codex_reviews:.*(?:\r?\n|$)/gm, '');
|
||||
fs.writeFileSync(path.join(stateRoot, 'config.yaml'), config + '\ncodex_reviews: disabled\n');
|
||||
// Readers and onboarding writers must agree on the owned state.
|
||||
return { env: { GSTACK_HOME: stateRoot, GSTACK_STATE_ROOT: stateRoot }, cleanup };
|
||||
} catch (error) {
|
||||
cleanup();
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Count evals review a seeded plan, never the checkout that supplies skills.
|
||||
* Put the complete request in Claude's initial project context before the
|
||||
* bare slash command starts: a later message can remain queued behind the
|
||||
* skill's first AskUserQuestion and leave it reviewing the live branch.
|
||||
*/
|
||||
export function createPlanCountFixture(prompt: string, opts: { nativeReviewOnly?: boolean; files?: Record<string, string> } = {}): {
|
||||
cwd: string;
|
||||
env: Record<string, string>;
|
||||
cleanup(): void;
|
||||
} {
|
||||
const cwd = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-plan-count-'));
|
||||
let nativeState: ReturnType<typeof createNativeReviewState> | undefined;
|
||||
const env: Record<string, string> = {};
|
||||
const cleanup = () => {
|
||||
try {
|
||||
fs.rmSync(cwd, { recursive: true, force: true });
|
||||
} finally {
|
||||
nativeState?.cleanup();
|
||||
}
|
||||
};
|
||||
try {
|
||||
const files = Object.entries(opts.files ?? {});
|
||||
for (const [name] of files) {
|
||||
const parts = name.split(/[\\/]/);
|
||||
// The fixture owns its seed plan, instructions and Git metadata. Extra
|
||||
// context must be an ordinary relative file, never an overwrite/escape.
|
||||
if (path.isAbsolute(name) || path.win32.isAbsolute(name) || parts.some(part =>
|
||||
part === '' || part === '.' || part === '..' || part.toLowerCase() === '.git') ||
|
||||
/^(?:plan|claude)\.md$/i.test(name)) {
|
||||
throw new Error(`Invalid plan-count fixture file: ${name}`);
|
||||
}
|
||||
}
|
||||
if (opts.nativeReviewOnly) {
|
||||
// Seeded-N bands cover native finding cadence; mode fixtures keep defaults.
|
||||
nativeState = createNativeReviewState();
|
||||
Object.assign(env, nativeState.env);
|
||||
}
|
||||
fs.writeFileSync(path.join(cwd, 'PLAN.md'), prompt);
|
||||
fs.writeFileSync(path.join(cwd, 'CLAUDE.md'), [
|
||||
'# Plan review fixture',
|
||||
'',
|
||||
'This repository contains the plan under review. Use PLAN.md as the',
|
||||
'current plan for the requested plan-review skill. The skill installation',
|
||||
'supplies the workflow; its source checkout is not the review target.',
|
||||
'',
|
||||
'The complete user request is available from the start of this session:',
|
||||
'',
|
||||
prompt,
|
||||
'',
|
||||
].join('\n'));
|
||||
for (const [name, content] of files) {
|
||||
const target = path.join(cwd, name);
|
||||
fs.mkdirSync(path.dirname(target), { recursive: true });
|
||||
fs.writeFileSync(target, content);
|
||||
}
|
||||
|
||||
const git = (args: string[]) => {
|
||||
const result = spawnSync('git', args, {
|
||||
cwd,
|
||||
encoding: 'utf8',
|
||||
timeout: 10_000,
|
||||
});
|
||||
if (result.error || result.status !== 0) {
|
||||
throw new Error(`Could not initialize plan-count fixture: ${result.error?.message ?? result.stderr}`);
|
||||
}
|
||||
};
|
||||
git(['init', '-b', 'main']);
|
||||
git(['add', '--', 'PLAN.md', 'CLAUDE.md', ...files.map(([name]) => name)]);
|
||||
git(['-c', 'user.name=Plan Count Fixture', '-c', 'user.email=plan-count@example.test',
|
||||
'-c', 'commit.gpgsign=false', 'commit', '--no-verify', '-m', 'Seed review plan']);
|
||||
git(['update-ref', 'refs/remotes/origin/main', 'HEAD']);
|
||||
return { cwd, env, cleanup };
|
||||
} catch (error) {
|
||||
cleanup();
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
/** Read-only pending ExitPlanMode identity for counting evals. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import type { PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
const MAX_RECORD_BYTES = 4096;
|
||||
const identifier = (value: unknown): value is string =>
|
||||
typeof value === 'string' && /^[A-Za-z0-9_-]{1,160}$/.test(value);
|
||||
const shellQuote = (value: string) => `'${(process.platform === 'win32' ? value.replaceAll('\\', '/') : value).replaceAll("'", "'\\''")}'`;
|
||||
|
||||
export function createPendingExitRecorder(cwd: string, configDir: string) {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-pending-exit-'));
|
||||
const file = path.join(dir, 'pending.json');
|
||||
const command = [process.execPath, import.meta.path, '--record', file, cwd, configDir].map(shellQuote).join(' ');
|
||||
const settings = JSON.stringify({ hooks: { PreToolUse: [{ matcher: '^ExitPlanMode$',
|
||||
hooks: [{ type: 'command', command, timeout: 5 }],
|
||||
}] } });
|
||||
return { file, settings, dispose: () => fs.rmSync(dir, { recursive: true, force: true }) };
|
||||
}
|
||||
|
||||
function scopedTranscript(file: unknown, configDir: string, session: string): boolean {
|
||||
if (typeof file !== 'string' || !path.isAbsolute(file)) return false;
|
||||
const relative = path.relative(path.join(configDir, 'projects'), file).split(path.sep);
|
||||
return relative.length === 2 && relative[0] !== '..' && relative[0] !== '.' &&
|
||||
relative[1] === `${session}.jsonl`;
|
||||
}
|
||||
|
||||
/** No stdout, permission decision, updated input, or model-visible context. */
|
||||
export function recordPendingExit(input: string, file: string, cwd: string, configDir: string): void {
|
||||
try {
|
||||
if (Buffer.byteLength(input) > 64 * 1024) throw new Error('hook input too large');
|
||||
const event = JSON.parse(input);
|
||||
// Subagents and other fixtures must never supply this main session's gate.
|
||||
if (event?.agent_id !== undefined || event?.cwd !== cwd) return;
|
||||
fs.rmSync(file, { force: true });
|
||||
if (event.hook_event_name !== 'PreToolUse' || event.tool_name !== 'ExitPlanMode' ||
|
||||
!identifier(event.session_id) || !identifier(event.tool_use_id) ||
|
||||
!scopedTranscript(event.transcript_path, configDir, event.session_id)) return;
|
||||
const record = { sessionId: event.session_id, toolUseId: event.tool_use_id,
|
||||
cwd, transcriptPath: event.transcript_path, timestamp: new Date().toISOString() };
|
||||
const temporary = `${file}.tmp`;
|
||||
fs.writeFileSync(temporary, JSON.stringify(record) + '\n', { mode: 0o600 });
|
||||
fs.renameSync(temporary, file);
|
||||
} catch {
|
||||
// A failed recorder cannot turn an earlier gate into current evidence.
|
||||
try { fs.rmSync(file, { force: true }); } catch { /* no evidence */ }
|
||||
}
|
||||
}
|
||||
|
||||
/** Exact current native gate; ordinary prose or a still-streaming menu is insufficient. */
|
||||
export function isCurrentPlanApprovalScreen(screen: string): boolean {
|
||||
const gate = /(?:^|\n) {0,3}─{5,}[ \t]*\n {0,3}Claude has written up a plan and is ready to execute\. Would you like to proceed\?[ \t]*\n\s*❯[ \t]*1\.[ \t]*Yes, and use auto mode[ \t]*\n[ \t]*2\.[ \t]*Yes, manually approve edits[ \t]*\n[ \t]*3\.[ \t]*Tell Claude what to change[ \t]*(?:\n[ \t]*shift\+tab to approve with this feedback)?\s*$/i.exec(screen) ??
|
||||
/(?:^|\n) {0,3}Exit plan mode\?[ \t]*\n(?:[ \t]*\n)* {0,4}Claude wants to exit plan mode[ \t]*\n\s*❯[ \t]*1\.[ \t]*Yes, and switch to default \(ask each time\) for this session[ \t]*\n[ \t]*2\.[ \t]*No\s*$/i.exec(screen);
|
||||
if (!gate) return false;
|
||||
const before = screen.slice(0, gate.index);
|
||||
if (/\b(?:example|sample|template|quote)\b.*[::]\s*$/i.test(before.trimEnd().split('\n').at(-1) ?? '')) return false;
|
||||
let fence: { marker: string; length: number } | undefined;
|
||||
for (const line of before.split('\n')) {
|
||||
const delimiter = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (!delimiter) continue;
|
||||
if (!fence) fence = { marker: delimiter[1]![0]!, length: delimiter[1]!.length };
|
||||
else if (delimiter[1]![0] === fence.marker && delimiter[1]!.length >= fence.length && !delimiter[2]!.trim()) fence = undefined;
|
||||
}
|
||||
return fence === undefined;
|
||||
}
|
||||
|
||||
/** Adds pending identity only; successful AUQ coverage still comes from JSONL results. */
|
||||
export function withPendingExit(
|
||||
transcript: PlanCountTranscript, file: string | undefined, cwd: string,
|
||||
configDir: string | null, startedAt: number, screen: string,
|
||||
): PlanCountTranscript {
|
||||
if (!file || !configDir || transcript.status !== 'ready' ||
|
||||
!isCurrentPlanApprovalScreen(screen)) return transcript;
|
||||
try {
|
||||
const stat = fs.lstatSync(file);
|
||||
if (!stat.isFile() || stat.size > MAX_RECORD_BYTES) return transcript;
|
||||
const record = JSON.parse(fs.readFileSync(file, 'utf8'));
|
||||
const time = Date.parse(record.timestamp);
|
||||
const sessions = new Set([...transcript.calls.map(call => call.sessionId),
|
||||
...transcript.assistantMessages.map(message => message.sessionId)]);
|
||||
if (record.cwd !== cwd || !identifier(record.sessionId) || !identifier(record.toolUseId) ||
|
||||
sessions.size !== 1 || !sessions.has(record.sessionId) ||
|
||||
!scopedTranscript(record.transcriptPath, configDir, record.sessionId) ||
|
||||
!Number.isFinite(time) || time < startedAt || time > Date.now()) return transcript;
|
||||
// A zero-question review still needs its owned gate for failure diagnostics.
|
||||
// Bind it to actual current-session assistant output; this adds no coverage.
|
||||
if (!transcript.calls.length && !transcript.assistantMessages.some(message =>
|
||||
message.sessionId === record.sessionId && message.text.trim() &&
|
||||
Number.isFinite(Date.parse(message.timestamp)) && Date.parse(message.timestamp) >= startedAt &&
|
||||
Date.parse(message.timestamp) <= time)) return transcript;
|
||||
const requests = transcript.planReadyRequests ?? [];
|
||||
// A flushed native record, including a failed result, always wins.
|
||||
if (requests.some(request => request.sessionId === record.sessionId && request.toolUseId === record.toolUseId)) return transcript;
|
||||
return { ...transcript, planReadyRequests: [...requests, {
|
||||
sessionId: record.sessionId, toolUseId: record.toolUseId,
|
||||
timestamp: record.timestamp, failed: false, source: 'pre_tool_use',
|
||||
}] };
|
||||
} catch { return transcript; }
|
||||
}
|
||||
|
||||
if (import.meta.main && process.argv[2] === '--record') {
|
||||
try {
|
||||
const [file, cwd, configDir] = process.argv.slice(3);
|
||||
if (file && cwd && configDir) recordPendingExit(await Bun.stdin.text(), file, cwd, configDir);
|
||||
} catch { /* observation never changes the native permission outcome */ }
|
||||
}
|
||||
@@ -0,0 +1,216 @@
|
||||
/** Pending native AUQ identity only; answers and coverage always come from JSONL. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as os from 'node:os';
|
||||
import * as path from 'node:path';
|
||||
import type { NativePlanQuestion, NativePlanQuestionCall, PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
const MAX_BYTES = 128 * 1024;
|
||||
const MAX_IDS = 128;
|
||||
const REASONS = ['invalid_event', 'input_overflow', 'stdin_timeout', 'record_error', 'concurrent_pending',
|
||||
'conflicting_replay', 'record_overflow', 'lock_conflict', 'hook_error', 'unknown'] as const;
|
||||
type RecorderReason = typeof REASONS[number];
|
||||
const identifier = (v: unknown): v is string => typeof v === 'string' && /^[A-Za-z0-9_-]{1,160}$/.test(v);
|
||||
const object = (v: unknown): v is Record<string, any> => v !== null && typeof v === 'object' && !Array.isArray(v);
|
||||
const quote = (v: string) => `'${v.replaceAll("'", "'\\''")}'`;
|
||||
const keysOnly = (v: Record<string, unknown>, keys: string[]) => Object.keys(v).every(k => keys.includes(k));
|
||||
|
||||
function questions(value: unknown): value is NativePlanQuestion[] {
|
||||
return Array.isArray(value) && value.length >= 1 && value.length <= 4 && value.every(q =>
|
||||
object(q) && keysOnly(q, ['header', 'question', 'options', 'multiSelect']) &&
|
||||
typeof q.header === 'string' && q.header.trim() && typeof q.question === 'string' && q.question.trim() &&
|
||||
(q.multiSelect === undefined || typeof q.multiSelect === 'boolean') &&
|
||||
Array.isArray(q.options) && q.options.length >= 2 && q.options.length <= 4 &&
|
||||
q.options.every((o: unknown) => object(o) && keysOnly(o, ['label', 'description']) &&
|
||||
typeof o.label === 'string' && o.label.trim() && (o.description === undefined || typeof o.description === 'string')) &&
|
||||
new Set(q.options.map((o: {label:string}) => o.label)).size === q.options.length) &&
|
||||
new Set(value.map(q => q.header)).size === value.length && new Set(value.map(q => q.question)).size === value.length;
|
||||
}
|
||||
|
||||
/** Completion-only fields from Claude's permission UI; never retained as answers. */
|
||||
function completionInput(input: Record<string, any>): boolean {
|
||||
if (!keysOnly(input, ['questions', 'answers', 'annotations'])) return false;
|
||||
const texts = new Set(input.questions.map((q: NativePlanQuestion) => q.question));
|
||||
if (input.answers !== undefined && (!object(input.answers) || Object.entries(input.answers).some(([key, value]) =>
|
||||
!texts.has(key) || typeof value !== 'string'))) return false;
|
||||
if (input.annotations !== undefined && (!object(input.annotations) || Object.entries(input.annotations).some(([key, value]) =>
|
||||
!texts.has(key) || !object(value) || !keysOnly(value, ['preview', 'notes']) ||
|
||||
Object.values(value).some(field => typeof field !== 'string')))) return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
function questionIdentity(value: NativePlanQuestion[]): string {
|
||||
// Claude normalizes object-key order after permission collection; question
|
||||
// order, option order and every supported field must still match.
|
||||
return JSON.stringify(value.map(q => ({header:q.header, question:q.question, multiSelect:q.multiSelect,
|
||||
options:q.options.map(o => ({label:o.label, description:o.description}))})));
|
||||
}
|
||||
|
||||
/** One real parent JSONL, inside the owned config; never a subagent or symlink. */
|
||||
function scopedTranscript(file: unknown, configDir: string, session: string): file is string {
|
||||
if (typeof file !== 'string' || !path.isAbsolute(file)) return false;
|
||||
const projects = path.join(configDir, 'projects');
|
||||
const rel = path.relative(projects, file).split(path.sep);
|
||||
if (rel.length !== 2 || !rel[0] || rel[0] === '..' || rel[0] === '.' || rel[1] !== `${session}.jsonl`) return false;
|
||||
try {
|
||||
return fs.lstatSync(file).isFile() && !fs.lstatSync(projects).isSymbolicLink() &&
|
||||
!fs.lstatSync(path.dirname(file)).isSymbolicLink() &&
|
||||
fs.realpathSync(file) === path.join(fs.realpathSync(configDir), 'projects', ...rel);
|
||||
} catch { return false; }
|
||||
}
|
||||
|
||||
interface Pending {
|
||||
sessionId: string; toolUseId: string; transcriptPath: string; timestamp: string; questions: NativePlanQuestion[];
|
||||
}
|
||||
interface State {
|
||||
version: 1; cwd: string; configDir: string; sessionId?: string; seenIds: string[]; pending: Pending | null;
|
||||
}
|
||||
|
||||
function readState(file: string, cwd: string, configDir: string): State {
|
||||
const stat = fs.lstatSync(file);
|
||||
if (!stat.isFile() || stat.size > MAX_BYTES) throw Error('invalid recorder file');
|
||||
const s = JSON.parse(fs.readFileSync(file, 'utf8'));
|
||||
if (!object(s) || s.version !== 1 || s.cwd !== cwd || s.configDir !== configDir ||
|
||||
(s.sessionId !== undefined && !identifier(s.sessionId)) || !Array.isArray(s.seenIds) ||
|
||||
s.seenIds.length > MAX_IDS || !s.seenIds.every(identifier) || new Set(s.seenIds).size !== s.seenIds.length ||
|
||||
(s.pending !== null && (!object(s.pending) || !identifier(s.pending.sessionId) ||
|
||||
s.pending.sessionId !== s.sessionId || !identifier(s.pending.toolUseId) ||
|
||||
!s.seenIds.includes(s.pending.toolUseId) || !questions(s.pending.questions) ||
|
||||
typeof s.pending.timestamp !== 'string' || !scopedTranscript(s.pending.transcriptPath, configDir, s.pending.sessionId)))) {
|
||||
throw Error('invalid recorder state');
|
||||
}
|
||||
return s as State;
|
||||
}
|
||||
|
||||
export function createPendingQuestionRecorder(cwd: string, configDir: string) {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-pending-question-'));
|
||||
const file = path.join(dir, 'state.json');
|
||||
fs.writeFileSync(file, JSON.stringify({version:1, cwd, configDir, seenIds:[], pending:null}) + '\n', {mode:0o600});
|
||||
// Git Bash needs slash-separated command paths. Recorder arguments retain
|
||||
// their native spelling because the owned event/state identity is exact.
|
||||
const shellPath = (value: string) => process.platform === 'win32' ? value.replaceAll('\\', '/') : value;
|
||||
const command = [shellPath(process.execPath), shellPath(import.meta.path), '--record', file, cwd, configDir].map(quote).join(' ');
|
||||
const hook = {matcher:'^AskUserQuestion$', hooks:[{type:'command', command, timeout:5}]};
|
||||
return {file, hooks:{PreToolUse:[hook], PostToolUse:[hook], PostToolUseFailure:[hook]},
|
||||
dispose: () => fs.rmSync(dir, {recursive:true, force:true})};
|
||||
}
|
||||
|
||||
/** Malformed/ambiguous owned input stays closed; later replay cannot revive it. */
|
||||
function poison(file: string, reason: RecorderReason) {
|
||||
try { fs.writeFileSync(file + '.invalid', JSON.stringify({reason})+'\n', {mode:0o600, flag:'wx'}); } catch { /* already invalid or disposed */ }
|
||||
}
|
||||
|
||||
/** Content-free diagnostic survives in caller snapshots even after close removes the recorder. */
|
||||
export function pendingQuestionRecorderStatus(file: string | undefined, cwd: string, configDir: string | null):
|
||||
{status:'disabled'|'missing'|'busy'|'idle'|'pending'|'invalid'; reason?:RecorderReason} {
|
||||
if (!file || !configDir) return {status:'disabled'};
|
||||
try {
|
||||
if (fs.existsSync(file+'.invalid')) {
|
||||
const stat = fs.lstatSync(file+'.invalid');
|
||||
const reason = stat.isFile() && stat.size <= 1024 ? JSON.parse(fs.readFileSync(file+'.invalid','utf8')).reason : 'unknown';
|
||||
return {status:'invalid', reason:REASONS.includes(reason) ? reason : 'unknown'};
|
||||
}
|
||||
if (fs.existsSync(file+'.lock')) return {status:'busy'};
|
||||
if (!fs.existsSync(file)) return {status:'missing'};
|
||||
return {status:readState(file,cwd,configDir).pending ? 'pending' : 'idle'};
|
||||
} catch { return {status:'invalid',reason:'record_error'}; }
|
||||
}
|
||||
|
||||
/** Silent observation: no approval, answers, input rewrite, or model context. */
|
||||
export function recordPendingQuestion(input: string, file: string, cwd: string, configDir: string): void {
|
||||
let lock: number | undefined;
|
||||
let reason: RecorderReason = 'input_overflow';
|
||||
const temporary = `${file}.${process.pid}.tmp`;
|
||||
try {
|
||||
if (Buffer.byteLength(input) > MAX_BYTES) throw Error('oversized hook input');
|
||||
reason = 'invalid_event';
|
||||
const e = JSON.parse(input);
|
||||
// Foreign and sidechain events cannot cancel or supply the parent request.
|
||||
if (object(e) && (e.agent_id !== undefined || (typeof e.cwd === 'string' && e.cwd !== cwd))) return;
|
||||
if (fs.existsSync(file + '.invalid')) return;
|
||||
reason = 'lock_conflict';
|
||||
lock = fs.openSync(file + '.lock', 'wx', 0o600);
|
||||
reason = 'record_error';
|
||||
const old = readState(file, cwd, configDir);
|
||||
reason = 'invalid_event';
|
||||
if (!object(e) || e.cwd !== cwd || !['PreToolUse', 'PostToolUse', 'PostToolUseFailure'].includes(e.hook_event_name) ||
|
||||
e.tool_name !== 'AskUserQuestion' || !identifier(e.session_id) || !identifier(e.tool_use_id) ||
|
||||
!scopedTranscript(e.transcript_path, configDir, e.session_id) ||
|
||||
!object(e.tool_input) || !questions(e.tool_input.questions) ||
|
||||
!(e.hook_event_name === 'PreToolUse' ? keysOnly(e.tool_input, ['questions']) : completionInput(e.tool_input))) {
|
||||
throw Error('invalid owned hook event');
|
||||
}
|
||||
if (old.sessionId !== undefined && old.sessionId !== e.session_id) return;
|
||||
const state: State = {...old, sessionId:e.session_id};
|
||||
const pending = old.pending;
|
||||
if (e.hook_event_name !== 'PreToolUse' && pending?.toolUseId === e.tool_use_id &&
|
||||
(pending.transcriptPath !== e.transcript_path || questionIdentity(pending.questions) !== questionIdentity(e.tool_input.questions))) {
|
||||
throw Error('completion does not match pending request');
|
||||
}
|
||||
if (e.hook_event_name === 'PreToolUse') {
|
||||
if (old.seenIds.includes(e.tool_use_id)) {
|
||||
if (pending?.toolUseId === e.tool_use_id && JSON.stringify(pending.questions) !== JSON.stringify(e.tool_input.questions)) {
|
||||
reason = 'conflicting_replay';
|
||||
throw Error('conflicting replay');
|
||||
}
|
||||
return;
|
||||
}
|
||||
if (pending) { reason = 'concurrent_pending'; throw Error('concurrent pending questions'); }
|
||||
state.pending = {sessionId:e.session_id, toolUseId:e.tool_use_id, transcriptPath:e.transcript_path,
|
||||
timestamp:new Date().toISOString(), questions:e.tool_input.questions};
|
||||
} else {
|
||||
// A late result cannot clear another request. Remember completion even
|
||||
// if its Pre was never observed, so replay cannot reopen that old ID.
|
||||
if (pending?.toolUseId === e.tool_use_id) state.pending = null;
|
||||
}
|
||||
if (!state.seenIds.includes(e.tool_use_id)) state.seenIds = [...state.seenIds, e.tool_use_id];
|
||||
const serialized = JSON.stringify(state) + '\n';
|
||||
reason = 'record_overflow';
|
||||
if (state.seenIds.length > MAX_IDS || Buffer.byteLength(serialized) > MAX_BYTES) throw Error('recorder overflow');
|
||||
reason = 'record_error';
|
||||
fs.writeFileSync(temporary, serialized, {mode:0o600, flag:'wx'});
|
||||
fs.renameSync(temporary, file);
|
||||
} catch { poison(file, reason); }
|
||||
finally {
|
||||
try { fs.rmSync(temporary, {force:true}); } catch { /* no evidence */ }
|
||||
if (lock !== undefined) {
|
||||
try { fs.closeSync(lock); fs.unlinkSync(file + '.lock'); } catch { poison(file, 'record_error'); }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Pending-only source. A published JSONL call, including failure, always wins. */
|
||||
export function readPendingQuestion(file: string | undefined, cwd: string, configDir: string | null,
|
||||
startedAt: number, transcript: PlanCountTranscript,
|
||||
): (NativePlanQuestionCall & {source:'pre_tool_use'}) | undefined {
|
||||
if (!file || !configDir || transcript.status !== 'ready' || !Number.isFinite(startedAt)) return undefined;
|
||||
try {
|
||||
if (fs.existsSync(file + '.invalid') || fs.existsSync(file + '.lock')) return undefined;
|
||||
const p = readState(file, cwd, configDir).pending;
|
||||
if (!p) return undefined;
|
||||
const time = Date.parse(p.timestamp);
|
||||
const sessions = new Set([...transcript.calls.map(c => c.sessionId), ...transcript.assistantMessages.map(m => m.sessionId)]);
|
||||
if (sessions.size !== 1 || !sessions.has(p.sessionId) || !Number.isFinite(time) || time < startedAt || time > Date.now() ||
|
||||
transcript.calls.some(c => c.sessionId === p.sessionId && c.toolUseId === p.toolUseId)) return undefined;
|
||||
if (!transcript.calls.length && !transcript.assistantMessages.some(m => m.sessionId === p.sessionId && m.text.trim() &&
|
||||
Number.isFinite(Date.parse(m.timestamp)) && Date.parse(m.timestamp) >= startedAt && Date.parse(m.timestamp) <= time)) return undefined;
|
||||
return {sessionId:p.sessionId, toolUseId:p.toolUseId, questions:p.questions, answered:false, failed:false, source:'pre_tool_use'};
|
||||
} catch { return undefined; }
|
||||
}
|
||||
|
||||
if (import.meta.main && process.argv[2] === '--record') {
|
||||
const [file, cwd, configDir] = process.argv.slice(3);
|
||||
if (file && cwd && configDir) {
|
||||
const timer = setTimeout(() => { poison(file, 'stdin_timeout'); process.exit(0); }, 4000);
|
||||
try {
|
||||
const chunks: Uint8Array[] = [];
|
||||
let size = 0;
|
||||
for await (const chunk of Bun.stdin.stream()) {
|
||||
size += chunk.byteLength;
|
||||
if (size > MAX_BYTES) { poison(file, 'input_overflow'); process.exit(0); }
|
||||
chunks.push(chunk);
|
||||
}
|
||||
recordPendingQuestion(Buffer.concat(chunks).toString('utf8'), file, cwd, configDir);
|
||||
} catch { poison(file, 'hook_error'); }
|
||||
finally { clearTimeout(timer); }
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,266 @@
|
||||
/** Lossless, read-only question metadata from one isolated Claude fixture. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
|
||||
export interface NativePlanQuestion {
|
||||
header: string;
|
||||
question: string;
|
||||
options: Array<{ label: string; description?: string }>;
|
||||
multiSelect?: boolean;
|
||||
}
|
||||
|
||||
export interface NativePlanQuestionCall {
|
||||
sessionId: string;
|
||||
toolUseId: string;
|
||||
questions: NativePlanQuestion[];
|
||||
answered: boolean;
|
||||
failed?: boolean;
|
||||
failure?: string;
|
||||
answers?: Record<string, string>;
|
||||
unansweredQuestionIndices?: number[];
|
||||
answeredAt?: string;
|
||||
}
|
||||
|
||||
/** Optional public tool projection for the Autoplan delivery audit; never thinking. */
|
||||
export interface NativePublicToolEvent {
|
||||
sessionId: string;
|
||||
timestamp: string;
|
||||
toolUseId: string;
|
||||
kind: 'use' | 'result';
|
||||
name?: string;
|
||||
/** Exact native message/request identity, used only for owned queued tools. */
|
||||
messageId?: string;
|
||||
requestId?: string;
|
||||
input?: Record<string, unknown>;
|
||||
content?: unknown;
|
||||
file?: unknown;
|
||||
isError?: boolean;
|
||||
}
|
||||
|
||||
export interface PlanCountTranscript {
|
||||
status: 'missing' | 'ready' | 'error';
|
||||
calls: NativePlanQuestionCall[];
|
||||
assistantMessages: Array<{ sessionId: string; text: string; timestamp: string }>;
|
||||
/** Actual native plan-mode approval requests; pending is the UI gate, never an AUQ. */
|
||||
planReadyRequests?: Array<{ sessionId: string; toolUseId: string; timestamp: string; failed: boolean; source?: 'pre_tool_use' }>;
|
||||
error?: string;
|
||||
}
|
||||
|
||||
/** A rejected/refused call needs an actual later answer, not unrelated progress. */
|
||||
export function unresolvedPlanQuestionCalls(calls: NativePlanQuestionCall[]): NativePlanQuestionCall[] {
|
||||
return calls.filter((call, index) => call.failed && !call.questions.every(q =>
|
||||
calls.slice(index + 1).some(later => later.answered && later.answers?.[q.question])));
|
||||
}
|
||||
|
||||
const MAX_BYTES = 32 * 1024 * 1024;
|
||||
const MAX_FILES = 64;
|
||||
const object = (value: unknown): value is Record<string, any> =>
|
||||
value !== null && typeof value === 'object' && !Array.isArray(value);
|
||||
const validTimestamp = (value: unknown): value is string =>
|
||||
typeof value === 'string' && Number.isFinite(Date.parse(value));
|
||||
|
||||
/** Read one length-delimited protobuf field, rejecting malformed/ambiguous input. */
|
||||
function signatureField(bytes: Uint8Array | undefined, wanted: number): Uint8Array | undefined {
|
||||
if (!bytes) return;
|
||||
let cursor = 0;
|
||||
let result: Uint8Array | undefined;
|
||||
let seen = false;
|
||||
const integer = () => {
|
||||
let value = 0;
|
||||
for (let shift = 0; shift < 70; shift += 7) {
|
||||
if (cursor >= bytes.length) throw new Error('truncated signature');
|
||||
const byte = bytes[cursor++]!;
|
||||
value += (byte & 127) * 2 ** shift;
|
||||
if (!Number.isSafeInteger(value)) throw new Error('signature integer overflow');
|
||||
if (!(byte & 128)) return value;
|
||||
}
|
||||
throw new Error('overlong signature integer');
|
||||
};
|
||||
while (cursor < bytes.length) {
|
||||
const key = integer();
|
||||
const field = Math.floor(key / 8);
|
||||
if (field < 1 || field > 0x1fffffff) throw new Error('invalid signature field');
|
||||
if (field === wanted) {
|
||||
if (seen) throw new Error('duplicate signature field');
|
||||
seen = true;
|
||||
}
|
||||
switch (key % 8) {
|
||||
case 0: integer(); break;
|
||||
case 1: cursor += 8; break;
|
||||
case 2: {
|
||||
const length = integer();
|
||||
if (length > bytes.length - cursor) throw new Error('truncated signature field');
|
||||
if (field === wanted) result = bytes.subarray(cursor, cursor + length);
|
||||
cursor += length;
|
||||
break;
|
||||
}
|
||||
case 5: cursor += 4; break;
|
||||
default: throw new Error('unsupported signature wire type');
|
||||
}
|
||||
if (cursor > bytes.length) throw new Error('truncated signature field');
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* Claude's public narration renderer classifies signature fields 2→1→8 as
|
||||
* block_kind="narration": summaries of inter-tool prose, not private reasoning.
|
||||
* Match that metadata only in this already-owned native transcript. This is
|
||||
* classification, not cryptographic signature verification. Never read the
|
||||
* thinking text of an untagged, unknown, malformed or legacy block.
|
||||
*/
|
||||
function publicNarrationText(block: Record<string, any>): string | undefined {
|
||||
if (block.type !== 'thinking' || typeof block.signature !== 'string' ||
|
||||
block.signature.length > 64 * 1024 || !/^[A-Za-z0-9+/]+={0,2}$/.test(block.signature)) return;
|
||||
try {
|
||||
const bytes = Buffer.from(block.signature, 'base64');
|
||||
const canonical = bytes.toString('base64');
|
||||
if (block.signature !== canonical && block.signature !== canonical.replace(/=+$/, '')) return;
|
||||
const tag = signatureField(signatureField(signatureField(bytes, 2), 1), 8);
|
||||
if (!tag || Buffer.from(tag).toString('utf8') !== 'narration') return;
|
||||
return typeof block.thinking === 'string' && block.thinking.trim() ? block.thinking : undefined;
|
||||
} catch { return; }
|
||||
}
|
||||
|
||||
function validQuestions(value: unknown): value is NativePlanQuestion[] {
|
||||
return Array.isArray(value) && value.length > 0 && value.every(q =>
|
||||
object(q) && typeof q.header === 'string' && typeof q.question === 'string' && q.question.trim() &&
|
||||
Array.isArray(q.options) && q.options.length >= 2 && q.options.every((o: unknown) =>
|
||||
object(o) && typeof o.label === 'string' && o.label.trim()));
|
||||
}
|
||||
|
||||
/**
|
||||
* Count callers consume each answered (sessionId, toolUseId) once, regardless
|
||||
* of questions[].length. A batched tool call must never become N findings.
|
||||
* Partial final lines remain pending; missing/foreign/sidechain records add
|
||||
* no coverage. Traversal stays inside the owned config's projects directory.
|
||||
*/
|
||||
export function readPlanCountTranscript(configDir: string, cwd: string,
|
||||
onPublicToolEvent?: (event: NativePublicToolEvent) => void,
|
||||
/** Optional exact parent journal, already validated by the owning native hook. */
|
||||
ownedParentTranscript?: string,
|
||||
): PlanCountTranscript {
|
||||
const calls = new Map<string, NativePlanQuestionCall>();
|
||||
const assistantMessages: PlanCountTranscript['assistantMessages'] = [];
|
||||
const planReadyRequests = new Map<string, NonNullable<PlanCountTranscript['planReadyRequests']>[number]>();
|
||||
let matched = false;
|
||||
let bytes = 0;
|
||||
let files = 0;
|
||||
const projects = path.join(configDir, 'projects');
|
||||
try {
|
||||
if (!fs.existsSync(projects)) return { status: 'missing', calls: [], assistantMessages: [] };
|
||||
const dirs = fs.readdirSync(projects, { withFileTypes: true }).filter(d => d.isDirectory());
|
||||
if (dirs.length > MAX_FILES) throw new Error('too many project directories');
|
||||
for (const dir of dirs) {
|
||||
const project = path.join(projects, dir.name);
|
||||
for (const entry of fs.readdirSync(project, { withFileTypes: true })) {
|
||||
if (!entry.isFile() || !entry.name.endsWith('.jsonl')) continue;
|
||||
if (++files > MAX_FILES) throw new Error('too many transcript files');
|
||||
const file = path.join(project, entry.name);
|
||||
if (ownedParentTranscript !== undefined && file !== ownedParentTranscript) continue;
|
||||
bytes += fs.statSync(file).size;
|
||||
if (bytes > MAX_BYTES) throw new Error('transcript exceeds 32 MiB read limit');
|
||||
const text = fs.readFileSync(file, 'utf8');
|
||||
// Native sessions retain their original journal after Bash changes cwd.
|
||||
// Admit that continuation only through UUID ancestry rooted in this
|
||||
// fixture's first parent user message; legacy records keep exact-cwd scoping.
|
||||
let originSeen = false;
|
||||
const ancestry = new Set<string>();
|
||||
const nativeUuid = (value: unknown): value is string =>
|
||||
typeof value === 'string' && /^[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}$/i.test(value);
|
||||
// Claude appends JSONL during rendering; an unfinished record is not
|
||||
// evidence of a call or an answer until its newline has been written.
|
||||
for (const line of text.slice(0, text.lastIndexOf('\n') + 1).split('\n')) {
|
||||
if (!line.trim()) continue;
|
||||
const record = JSON.parse(line);
|
||||
if (!object(record) || typeof record.sessionId !== 'string' ||
|
||||
entry.name !== `${record.sessionId}.jsonl` ||
|
||||
(ownedParentTranscript !== undefined && record.agentId != null)) continue;
|
||||
const parentMetadata = record.isSidechain === false && record.agentId == null &&
|
||||
typeof record.cwd === 'string' && path.isAbsolute(record.cwd) &&
|
||||
nativeUuid(record.uuid) && validTimestamp(record.timestamp);
|
||||
const continuation = parentMetadata && nativeUuid(record.parentUuid) &&
|
||||
ancestry.has(record.parentUuid) && !ancestry.has(record.uuid);
|
||||
if (!originSeen && object(record.message) && ['user', 'assistant'].includes(record.message.role)) {
|
||||
originSeen = true;
|
||||
if (parentMetadata && record.cwd === cwd && record.message.role === 'user' &&
|
||||
record.parentUuid === null) ancestry.add(record.uuid);
|
||||
}
|
||||
if (continuation) ancestry.add(record.uuid);
|
||||
if ((record.cwd !== cwd && !continuation) || record.isSidechain !== false ||
|
||||
!object(record.message) || !Array.isArray(record.message.content)) continue;
|
||||
matched = true;
|
||||
for (const block of record.message.content) {
|
||||
if (!object(block)) continue;
|
||||
if (onPublicToolEvent && validTimestamp(record.timestamp)) {
|
||||
if (record.message.role === 'assistant' && block.type === 'tool_use' &&
|
||||
typeof block.id === 'string' && typeof block.name === 'string' && object(block.input)) {
|
||||
const batch = typeof record.message.id === 'string' && /^msg_[A-Za-z0-9_-]{1,160}$/.test(record.message.id) &&
|
||||
typeof record.requestId === 'string' && /^req_[A-Za-z0-9_-]{1,160}$/.test(record.requestId)
|
||||
? { messageId: record.message.id, requestId: record.requestId } : {};
|
||||
onPublicToolEvent({ sessionId: record.sessionId, timestamp: record.timestamp,
|
||||
toolUseId: block.id, kind: 'use', name: block.name, input: block.input, ...batch });
|
||||
} else if (record.message.role === 'user' && block.type === 'tool_result' &&
|
||||
typeof block.tool_use_id === 'string') {
|
||||
onPublicToolEvent({ sessionId: record.sessionId, timestamp: record.timestamp,
|
||||
toolUseId: block.tool_use_id, kind: 'result', content: block.content,
|
||||
file: record.toolUseResult?.file, isError: block.is_error === true });
|
||||
}
|
||||
}
|
||||
|
||||
if (record.message.role === 'assistant' && validTimestamp(record.timestamp)) {
|
||||
const text = block.type === 'text' && typeof block.text === 'string' && block.text.trim()
|
||||
? block.text : publicNarrationText(block);
|
||||
if (text) assistantMessages.push({ sessionId: record.sessionId, text, timestamp: record.timestamp });
|
||||
}
|
||||
if (record.message.role === 'assistant' && block.type === 'tool_use' && block.name === 'ExitPlanMode' &&
|
||||
typeof block.id === 'string' && validTimestamp(record.timestamp)) {
|
||||
const key = `${record.sessionId}:${block.id}`;
|
||||
if (!planReadyRequests.has(key)) planReadyRequests.set(key, { sessionId: record.sessionId,
|
||||
toolUseId: block.id, timestamp: record.timestamp, failed: false });
|
||||
}
|
||||
if (record.message.role === 'assistant' && block.type === 'tool_use' && block.name === 'AskUserQuestion' &&
|
||||
typeof block.id === 'string' && object(block.input) && validQuestions(block.input.questions)) {
|
||||
const key = `${record.sessionId}:${block.id}`;
|
||||
const prior = calls.get(key);
|
||||
if (prior && JSON.stringify(prior.questions) !== JSON.stringify(block.input.questions)) {
|
||||
throw new Error('conflicting question metadata for one tool call');
|
||||
}
|
||||
if (!prior) calls.set(key, { sessionId: record.sessionId, toolUseId: block.id,
|
||||
questions: block.input.questions, answered: false, failed: false });
|
||||
} else if (record.message.role === 'user' && block.type === 'tool_result' &&
|
||||
typeof block.tool_use_id === 'string') {
|
||||
const ready = planReadyRequests.get(`${record.sessionId}:${block.tool_use_id}`);
|
||||
if (ready && block.is_error === true) ready.failed = true;
|
||||
const call = calls.get(`${record.sessionId}:${block.tool_use_id}`);
|
||||
const answers = record.toolUseResult?.answers;
|
||||
const validAnswers = call && object(answers) ? Object.fromEntries(call.questions
|
||||
.filter(q => typeof answers[q.question] === 'string' && answers[q.question].trim())
|
||||
.map(q => [q.question, answers[q.question]])) : {};
|
||||
if (call && block.is_error !== true && Object.keys(validAnswers).length > 0) {
|
||||
// The CLI allows submitting a multi-question packet with
|
||||
// unanswered tabs. This completes ONE call, not N questions.
|
||||
call.answered = true;
|
||||
call.failed = false;
|
||||
delete call.failure;
|
||||
call.answers = validAnswers;
|
||||
call.unansweredQuestionIndices = call.questions.flatMap((q, i) => q.question in validAnswers ? [] : [i]);
|
||||
call.answeredAt = validTimestamp(record.timestamp) ? record.timestamp : undefined;
|
||||
} else if (call) {
|
||||
if (call.answered) throw new Error('conflicting successful and failed results for one question call');
|
||||
call.failed = true;
|
||||
call.failure = block.is_error === true ? 'Native question tool returned is_error' : 'Native question returned no matching nonempty answers';
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return { status: matched ? 'ready' : 'missing', calls: [...calls.values()], assistantMessages,
|
||||
...(planReadyRequests.size ? { planReadyRequests: [...planReadyRequests.values()] } : {}) };
|
||||
} catch (error) {
|
||||
// A failed read cannot silently turn an incomplete transcript into a
|
||||
// complete review. Keep the diagnostic explicit and return no coverage.
|
||||
return { status: 'error', calls: [], assistantMessages: [], error: `Claude question transcript: ${String(error)}` };
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,65 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { createHash } from 'node:crypto';
|
||||
|
||||
export interface PlanFloorTargetDelivery {
|
||||
status: 'missing' | 'ready' | 'error';
|
||||
reason?: string;
|
||||
sessionId: string;
|
||||
command: string;
|
||||
targetPath: string;
|
||||
targetSha256: string;
|
||||
acknowledgedAt?: string;
|
||||
}
|
||||
|
||||
/** The first owned native user command must name the already-seeded plan. */
|
||||
export function readPlanFloorTarget(configDir: string | null, cwd: string, opts: {
|
||||
seed: string; sessionId: string; slashCommand: string; startedAt: number; now: number;
|
||||
}): PlanFloorTargetDelivery {
|
||||
const delivery: PlanFloorTargetDelivery = {
|
||||
status: 'missing', sessionId: opts.sessionId, command: `${opts.slashCommand} PLAN.md`,
|
||||
targetPath: path.join(cwd, 'PLAN.md'), targetSha256: createHash('sha256').update(opts.seed).digest('hex'),
|
||||
};
|
||||
try {
|
||||
if (!configDir || !/^[\da-f-]{36}$/i.test(opts.sessionId) || !/^\/[\w-]+$/.test(opts.slashCommand) ||
|
||||
!Number.isFinite(opts.startedAt) || !Number.isFinite(opts.now) || opts.now < opts.startedAt) return delivery;
|
||||
if (!fs.lstatSync(delivery.targetPath).isFile() || fs.readFileSync(delivery.targetPath, 'utf8') !== opts.seed) {
|
||||
return { ...delivery, status: 'error', reason: 'Owned seed plan is missing or changed' };
|
||||
}
|
||||
const projects = path.join(configDir, 'projects');
|
||||
if (!fs.existsSync(projects)) return delivery;
|
||||
if (!fs.lstatSync(projects).isDirectory()) throw Error('Native project root is not a directory');
|
||||
const dirs = fs.readdirSync(projects, { withFileTypes: true }).filter(dir => dir.isDirectory());
|
||||
if (dirs.length > 128) throw Error('Too many native project directories');
|
||||
const files = dirs.map(dir => path.join(projects, dir.name, `${opts.sessionId}.jsonl`)).filter(file => fs.existsSync(file));
|
||||
if (files.length === 0) return delivery;
|
||||
if (files.length !== 1 || !fs.lstatSync(files[0]!).isFile()) throw Error('Ambiguous or nonregular native session');
|
||||
if (fs.statSync(files[0]!).size > 32 * 1024 * 1024) throw Error('Native session exceeds 32 MiB');
|
||||
const raw = fs.readFileSync(files[0]!, 'utf8');
|
||||
const expected = `<command-message>${opts.slashCommand.slice(1)}</command-message>\n` +
|
||||
`<command-name>${opts.slashCommand}</command-name>\n<command-args>PLAN.md</command-args>`;
|
||||
for (const line of raw.slice(0, raw.lastIndexOf('\n') + 1).split('\n')) {
|
||||
if (!line.trim()) continue;
|
||||
const record = JSON.parse(line), at = Date.parse(record.timestamp ?? '');
|
||||
if (record.type !== 'user' || record.isSidechain !== false || record.cwd !== cwd ||
|
||||
record.sessionId !== opts.sessionId || record.parent_tool_use_id || record.message?.role !== 'user' ||
|
||||
!Number.isFinite(at) || at < opts.startedAt || at > opts.now) continue;
|
||||
const content = record.message.content;
|
||||
const text = typeof content === 'string' ? content : Array.isArray(content) && content.length === 1 &&
|
||||
content[0]?.type === 'text' && typeof content[0].text === 'string' ? content[0].text : undefined;
|
||||
if (text === undefined) {
|
||||
if (Array.isArray(content) && content.some(block => block?.type === 'text')) {
|
||||
return { ...delivery, status: 'error', reason: 'First owned native user text was not one target command' };
|
||||
}
|
||||
continue; // Tool-result-only user envelopes are not submitted commands.
|
||||
}
|
||||
// A later queued target cannot repair a skill that already started bare.
|
||||
return text === expected
|
||||
? { ...delivery, status: 'ready', acknowledgedAt: record.timestamp }
|
||||
: { ...delivery, status: 'error', reason: 'First owned native user input did not name the seeded plan' };
|
||||
}
|
||||
return delivery;
|
||||
} catch (error) {
|
||||
return { ...delivery, status: 'error', reason: error instanceof Error ? error.message : String(error) };
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,134 @@
|
||||
import type { NativePublicToolEvent, PlanCountTranscript } from './plan-count-transcript';
|
||||
|
||||
function deicticPlanSelection(text: string): RegExpExecArray | null {
|
||||
return /^(?:I'll|I will) (?:review|(?:run|invoke) (?:the )?\/?([\w:-]+(?:[ \t]+[\w:-]+)*) skill (?:to review|on|against)) (?:this|your|the) (?:draft[ \t]+)?([\p{L}\p{N}]+(?:[ \t\u2010-\u2015-]+[\p{L}\p{N}]+)*[ \t]+)?plan\.$/iu.exec(text);
|
||||
}
|
||||
|
||||
function describesTitle(descriptor: string | undefined, title: string): boolean {
|
||||
const normalize = (value: string) => value.toLowerCase().replace(/[\u2010-\u2015-]+/g, ' ').replace(/\s+/g, ' ').trim();
|
||||
return !descriptor || ` ${normalize(title)} `.includes(` ${normalize(descriptor)} `);
|
||||
}
|
||||
|
||||
/** An asserted correction can retract a declaration; quoted source cannot. */
|
||||
function withdrawsPlanSelection(message: string, title: string): boolean {
|
||||
let fence: { char: string; length: number } | undefined;
|
||||
let source = false;
|
||||
const assertions: string[] = [];
|
||||
for (const line of message.split(/\r?\n/)) {
|
||||
const mark = /^ {0,3}(`{3,}|~{3,})(.*)$/.exec(line);
|
||||
if (mark) {
|
||||
if (!fence) fence = { char: mark[1]![0]!, length: mark[1]!.length };
|
||||
else if (mark[1]![0] === fence.char && mark[1]!.length >= fence.length && !mark[2]!.trim()) fence = undefined;
|
||||
continue;
|
||||
}
|
||||
if (fence || /^(?:\s*>| {4}|\t)/.test(line)) continue;
|
||||
const text = line.trim();
|
||||
if (/^(?:#{1,6}\s+)?(?:Current|Actual)\s+(?:assessment|scope|selection|status)\b/i.test(text)) source = false;
|
||||
else if (/^(?:#{1,6}\s+)?(?:Source|Example|Historical|Quoted|Original message|Expected output)\b/i.test(text)
|
||||
|| /^(?:The following|This is)\b[^.!?]*\b(?:source|example|hypothetical|quoted)\b/i.test(text)) source = true;
|
||||
if (!source) assertions.push(text);
|
||||
}
|
||||
for (const statement of assertions.join('\n').split(/(?<=[.!?;])\s+|\n+/).map(line => line.trim())) {
|
||||
if (statement.endsWith('?')) continue;
|
||||
const claim = statement.replace(/^Correction:\s*/i, '');
|
||||
const plain = claim.replace(/"[^"\n]*"|“[^”\n]*”|'[^'\n]*'|‘[^’\n]*’|`[^`\n]*`/g, (quoted, index) =>
|
||||
/^(?:withdrawn|retracted|cancelled|canceled|hypothetical|no longer current|superseded|rejected)$/i.test(quoted.slice(1, -1)) &&
|
||||
/^(?:The|This|That|My)\s+(?:(?:scope|target)\s+)?(?:selection|declaration)\s+(?:is|was|has been|remains)\s+(?:now\s+)?$/i.test(claim.slice(0, index))
|
||||
? quoted.slice(1, -1) : '[quoted]');
|
||||
if (/^(?:The|This|That|My)\s+(?:(?:scope|target)\s+)?(?:selection|declaration)\s+(?:is|was|has been|remains)\s+(?:now\s+)?(?:withdrawn|retracted|cancelled|canceled|hypothetical|no longer current|superseded|rejected)\b/i.test(plain)
|
||||
|| /^(?:(?:I|We)\s+(?:have\s+)?)?(?:withdrawn?|withdrew|retract(?:ed)?|cancel(?:led|ed)?|disregard(?:ed)?|ignore(?:d)?)\s+(?:this|that|the|my)\s+(?:selection|declaration)\b/i.test(plain)) return true;
|
||||
const deictic = deicticPlanSelection(claim);
|
||||
if (deictic && !describesTitle(deictic[2], title)) return true;
|
||||
const reviewing = /^(?:I'll|I will|I'm|I am|We will|We're|We are) (?:now )?(?:review|reviewing) (?:the )?(?:branch diff|(?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`) (?:draft(?: plan)?|plan))(?: instead)?\.$/i.exec(claim);
|
||||
if (reviewing && (reviewing[1] ?? reviewing[2] ?? reviewing[3] ?? 'branch diff').toLowerCase() !== title.toLowerCase()) return true;
|
||||
const changedTarget = /^(?:The|This|My)\s+(?:selected|review)\s+target\s+is\s+(?:now\s+)?(.+?)[.!?]?$/i.exec(claim);
|
||||
if (changedTarget) {
|
||||
const target = changedTarget[1]!.replace(/^(?:the\s+)/i, '').replace(/["“”`]/g, '').replace(/\s+(?:draft(?:\s+plan)?|plan)[.!?]?$/i, '').trim();
|
||||
if (target.toLowerCase() !== title.toLowerCase()) return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** A selected pasted plan may be announced by name instead of the menu letter. */
|
||||
export function nativeSeededPlanSelection(
|
||||
transcript: PlanCountTranscript,
|
||||
tools: NativePublicToolEvent[],
|
||||
opts: { seed: string; skillName: string; sessionId: string; commandStartedAt: number },
|
||||
): boolean {
|
||||
if (transcript.status !== 'ready' || !opts.sessionId || !Number.isFinite(opts.commandStartedAt)) return false;
|
||||
const headings = [...opts.seed.matchAll(/^#\s+(?:Plan:\s*)?([^\r\n]+)$/gmi)];
|
||||
if (headings.length !== 1) return false;
|
||||
const title = headings[0]![1]!.trim();
|
||||
if (!title || title.length > 200) return false;
|
||||
const at = (timestamp: string) => Date.parse(timestamp);
|
||||
// Loading must succeed in this invocation; the public target declaration
|
||||
// may come immediately before it, so the user can interrupt before work.
|
||||
const calls = tools.filter(event => event.kind === 'use' && event.sessionId === opts.sessionId &&
|
||||
event.name === 'Skill' && [opts.skillName, `gstack:${opts.skillName}`].includes(String(event.input?.skill ?? '')) &&
|
||||
Number.isFinite(at(event.timestamp)) && at(event.timestamp) >= opts.commandStartedAt);
|
||||
const loaded = calls.flatMap(call => tools.filter(event => event.kind === 'result' &&
|
||||
event.sessionId === opts.sessionId && event.toolUseId === call.toolUseId && event.isError === false &&
|
||||
Number.isFinite(at(event.timestamp)) && at(event.timestamp) >= at(call.timestamp)));
|
||||
if (calls.length !== 1 || loaded.length !== 1) return false;
|
||||
// Explicit Skill arguments must select this seed, not merely mention its
|
||||
// title while requesting another target. Unknown argument forms fail closed.
|
||||
const args = calls[0]!.input?.args;
|
||||
if (args !== undefined && args !== '') {
|
||||
if (typeof args !== 'string') return false;
|
||||
const target = /^Review (?:the )?(?:draft plan|plan) (?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`)(?: (?:provided|pasted) in the conversation above)?(?: \([^()\n]*\))?\.?$/i.exec(args)
|
||||
?? /^Review (?:the )?pasted (?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`) (?:draft plan|plan)\.?$/i.exec(args);
|
||||
const pasted = /^Review this draft plan:\s*([\s\S]+)$/i.exec(args);
|
||||
const sameDraft = pasted && pasted[1]!.replace(/\s+/g, ' ').trim() === opts.seed.replace(/\s+/g, ' ').trim();
|
||||
if (!sameDraft && (!target || (target[1] ?? target[2] ?? target[3])!.toLowerCase() !== title.toLowerCase()
|
||||
|| /\b(?:if|unless|instead|not|pending|assuming)\b/i.test(args))) return false;
|
||||
}
|
||||
const work = tools.filter(event => event.kind === 'use' && event.sessionId === opts.sessionId && event.toolUseId !== calls[0]!.toolUseId);
|
||||
const beforeWork = (time: number) => work.every(event => Number.isFinite(at(event.timestamp)) &&
|
||||
(at(event.timestamp) < opts.commandStartedAt || at(event.timestamp) > time));
|
||||
if (!beforeWork(at(loaded[0]!.timestamp))) return false;
|
||||
const remainsSelected = (timestamp: string) => !transcript.assistantMessages.some(later =>
|
||||
later.sessionId === opts.sessionId && Number.isFinite(at(later.timestamp)) &&
|
||||
at(later.timestamp) >= at(timestamp) && withdrawsPlanSelection(later.text, title));
|
||||
for (const message of transcript.assistantMessages) {
|
||||
if (message.sessionId !== opts.sessionId || !Number.isFinite(at(message.timestamp)) ||
|
||||
at(message.timestamp) <= opts.commandStartedAt || !beforeWork(at(message.timestamp))) continue;
|
||||
// Only a first asserted line can select the target. A source, quote or
|
||||
// hypothesis introduction owns its following text regardless of wording.
|
||||
const line = message.text.split(/\r?\n/).find(value => value.trim());
|
||||
if (!line || /^(?: {4}|\t)/.test(line)) continue;
|
||||
const text = line.trim();
|
||||
// A deictic target binds to the single pasted plan. Any descriptor must
|
||||
// occur as contiguous whole words in its title, never merely in its body.
|
||||
const draft = deicticPlanSelection(text);
|
||||
const names = [opts.skillName, `gstack:${opts.skillName}`, opts.skillName.replace(/^plan-/, '')];
|
||||
// Human role names remain tied to this successfully loaded skill.
|
||||
names.push(opts.skillName.replace(/-/g, ' '), opts.skillName.replace(/^plan-/, '').replace(/-/g, ' '));
|
||||
if (opts.skillName === 'plan-eng-review') names.push('eng-manager plan review');
|
||||
if (draft && describesTitle(draft[2], title) && (!draft[1] || names.includes(draft[1].toLowerCase().replace(/[ \t]+/g, ' '))) && remainsSelected(message.timestamp)) return true;
|
||||
const automatic = /^(?:I\'ll|I will) auto[- ]select option B and review\s+(?:the\s+)?(.+?)\s+(?:draft(?:\s+plan)?|plan)\s+(?:you shared|you pasted|pasted here)(.*)$/i.exec(text);
|
||||
const automaticTarget = automatic?.[1]?.replace(/^(?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`)$/, (_, straight, curly, code) => straight ?? curly ?? code);
|
||||
const selectedNow = /^(?:I've|I have) selected (?:option B, )?(?:reviewing|to review)\s+(?:the\s+)?pasted\s+(?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`)\s+(?:draft(?:\s+plan)?|plan)(.*)$/i.exec(text);
|
||||
const selected = selectedNow ?? /^(?:Scope gate confirms plan mode, so )?(?:I'll review|I will review|I'll go with reviewing|I will go with reviewing|I'll proceed with reviewing|I will proceed with reviewing|I'm proceeding with reviewing|I am proceeding with reviewing)\s+(?:the\s+)?(?:pasted\s+)?(?:"([^"\n]+)"|“([^”\n]+)”|`([^`\n]+)`)\s+(?:draft(?:\s+plan)?|plan)(?:\s+(?:you pasted|pasted here))?(.*)$/i.exec(text);
|
||||
const selectedTarget = automaticTarget ?? (selected && (selected[1] ?? selected[2] ?? selected[3]));
|
||||
if (!selectedTarget || selectedTarget.trim().toLowerCase() !== title.toLowerCase()) continue;
|
||||
const tail = automatic?.[2] ?? selected![4]!;
|
||||
if (/\b(?:if|unless|assuming|pending|only after|instead|not|won't|cannot)\b/i.test(tail)) continue;
|
||||
if (/\b(?:retract|withdraw|cancel|disregard|ignore)\s+(?:that|this|the|my)\s+(?:selection|declaration)\b/i.test(tail)) continue;
|
||||
if (/\b(?:treat|consider|regard)\s+(?:that|this|the|my)\s+(?:selection|declaration)\s+as\s+(?:a\s+)?(?:hypothetical|example|proposal)\b/i.test(tail)) continue;
|
||||
if (/\b(?:that|this|the|my)\s+(?:selection|declaration)\s+(?:is|was)\s+(?:withdrawn|cancelled|canceled|hypothetical|retracted)\b/i.test(tail)) continue;
|
||||
if (automatic) {
|
||||
if (/^(?:\.|,\s*running\s+(?:the\s+)?(?:pre-review\s+)?audit\b[^?]*\.)$/i.test(tail) && remainsSelected(message.timestamp)) return true;
|
||||
continue;
|
||||
}
|
||||
if (selectedNow) {
|
||||
// A completed selection may name the pasted target before its plan-mode
|
||||
// reason. Keep the first assertion bound; later work is not a new target.
|
||||
// Internal token dots (DESIGN.md) do not open another sentence.
|
||||
if (/^(?:\s+(?:since|because)\s+(?:we're|we are|I'm|I am)\s+in plan mode)?\.(?:\s+(?:Now|Next,?|Then)\s+(?:I'll|I will)\s+(?:run|start|begin)\s+(?:the\s+)?(?:pre-review\s+)?(?:audit|Design Doc Check)\b(?:[^.!?]|\.(?=\S))*\.)?$/i.test(tail) && remainsSelected(message.timestamp)) return true;
|
||||
continue;
|
||||
}
|
||||
if (/^(?:\.(?:\s+(?:Next,|Then\b).*)?|,\s*(?:starting|beginning)\s+(?:with|by)\b.*|,\s*and\s+now\s+(?:I'm|I am)\s+(?:running|starting|beginning)\s+(?:the\s+)?(?:pre-review\s+)?audit\b[^?]*\.|\. Running (?:the )?(?:pre-review )?audit\b(?:[^.!?]|\.(?=\S))*\.|)$/.test(tail) && remainsSelected(message.timestamp)) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
@@ -0,0 +1,93 @@
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { createRequire } from 'node:module';
|
||||
|
||||
/** The existing xterm package ships this headless source beside its browser bundle. */
|
||||
let terminalConstructor: Promise<any> | undefined;
|
||||
function loadTerminal(): Promise<any> {
|
||||
return terminalConstructor ??= (async () => {
|
||||
try {
|
||||
const source = path.join(path.dirname(createRequire(import.meta.url).resolve('xterm/package.json')), 'src');
|
||||
const entry = path.join(source, 'headless/public/Terminal.ts');
|
||||
if (!fs.existsSync(entry)) throw new Error(`Installed xterm headless source is missing: ${entry}`);
|
||||
const transpiler = new Bun.Transpiler({ loader: 'ts', target: 'bun', trimUnusedImports: true,
|
||||
tsconfig: JSON.stringify({ compilerOptions: { experimentalDecorators: true, useDefineForClassFields: false } }) });
|
||||
const build = await Bun.build({ entrypoints: [entry], target: 'bun', format: 'esm',
|
||||
// xterm5 detects Node by navigator absence. This changes only the
|
||||
// compiled module; Bun's global navigator and other tests stay intact.
|
||||
define: { navigator: 'undefined' },
|
||||
plugins: [{ name: 'installed-xterm-headless', setup(builder) {
|
||||
builder.onResolve({ filter: /^(common|headless)\// }, args => {
|
||||
const stem = path.join(source, args.path);
|
||||
return { path: fs.existsSync(`${stem}.ts`) ? `${stem}.ts` : `${stem}.d.ts` };
|
||||
});
|
||||
builder.onLoad({ filter: /[/\\]xterm[/\\]src[/\\].*\.ts$/ }, async args => ({
|
||||
contents: transpiler.transformSync(await Bun.file(args.path).text()), loader: 'js',
|
||||
}));
|
||||
} }],
|
||||
});
|
||||
if (!build.success) throw new AggregateError(build.logs, 'Installed xterm headless build failed');
|
||||
const encoded = Buffer.from(await build.outputs[0].text()).toString('base64');
|
||||
return (await import(`data:text/javascript;base64,${encoded}`)).Terminal;
|
||||
} catch (cause) {
|
||||
throw new Error('PTY screen unavailable; cannot safely observe terminal input.', { cause });
|
||||
}
|
||||
})();
|
||||
}
|
||||
|
||||
export interface PtyScreen {
|
||||
write(text: string): void;
|
||||
read(): Promise<string>;
|
||||
dispose(): Promise<void>;
|
||||
}
|
||||
|
||||
/** One terminal per session; read only the actual viewport, never scrollback. */
|
||||
export async function createPtyScreen(cols: number, rows: number): Promise<PtyScreen> {
|
||||
const Terminal = await loadTerminal();
|
||||
const terminal = new Terminal({ cols, rows, scrollback: 0, allowProposedApi: true });
|
||||
// Match current CLI scalar column widths instead of xterm5's Unicode 6
|
||||
// default. xterm still handles combining cells; this is not grapheme shaping.
|
||||
terminal.unicode.register({
|
||||
version: 'bun-scalar',
|
||||
wcwidth(codepoint: number) {
|
||||
const width = Bun.stringWidth(String.fromCodePoint(codepoint));
|
||||
if (width !== 0 && width !== 1 && width !== 2) throw new Error('Invalid terminal scalar width.');
|
||||
return width;
|
||||
},
|
||||
});
|
||||
terminal.unicode.activeVersion = 'bun-scalar';
|
||||
let pending = 0;
|
||||
let failure: unknown;
|
||||
let final: string | undefined;
|
||||
let closing: Promise<void> | undefined;
|
||||
const waiting = new Set<() => void>();
|
||||
const settled = () => { if (pending === 0) { for (const done of waiting) done(); waiting.clear(); } };
|
||||
const drain = async () => {
|
||||
while (pending > 0) await new Promise<void>(resolve => waiting.add(resolve));
|
||||
if (failure) throw new Error('PTY screen parse failed.', { cause: failure });
|
||||
};
|
||||
const viewport = () => {
|
||||
const buffer = terminal.buffer.active;
|
||||
return Array.from({ length: rows }, (_, i) => buffer.getLine(buffer.baseY + i)?.translateToString(true) ?? '').join('\n');
|
||||
};
|
||||
return {
|
||||
write(text) {
|
||||
if (closing) throw new Error('Cannot write to a disposed PTY screen.');
|
||||
if (!text) return;
|
||||
pending++;
|
||||
try { terminal.write(text, () => { pending--; settled(); }); }
|
||||
catch (error) { failure = error; pending--; settled(); }
|
||||
},
|
||||
async read() {
|
||||
if (closing) { await closing; return final!; }
|
||||
await drain();
|
||||
return viewport();
|
||||
},
|
||||
dispose() {
|
||||
return closing ??= (async () => {
|
||||
try { await drain(); final = viewport(); }
|
||||
finally { terminal.dispose(); }
|
||||
})();
|
||||
},
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,37 @@
|
||||
import { stripVTControlCharacters } from 'node:util';
|
||||
|
||||
/** Select the explicit affirmative entry in a fully rendered startup menu. */
|
||||
export function trustDialogInput(visible: string): string | null {
|
||||
const clean = stripVTControlCharacters(visible);
|
||||
const headers = [...clean.matchAll(/Accessing\s*workspace:|Do\s*you\s*trust\s*(?:the\s*files\s*in\s*)?this\s*folder\?|Trust\s*this\s*folder\?/gi)];
|
||||
const header = headers.at(-1);
|
||||
if (!header) return null;
|
||||
const dialog = clean.slice(header.index);
|
||||
|
||||
type Option = { selected: boolean; affirmative: boolean };
|
||||
const options: Option[] = [];
|
||||
for (const line of dialog.split(/[\r\n]+/)) {
|
||||
const match = line.replace(/\s+/g, '').match(
|
||||
/^(❯)?(?:\d+[.)])?(Yes(?:,Itrustthisfolder|,andalwaysallowaccessto.+)?|No(?:,exit)?)(?:\(Recommended\))?$/i,
|
||||
);
|
||||
if (match) options.push({ selected: !!match[1], affirmative: /^Yes/i.test(match[2]!) });
|
||||
}
|
||||
|
||||
// Cursor-positioning escapes can collapse the two current menu rows into
|
||||
// one logical line. Their complete labels still identify both choices.
|
||||
if (options.length < 2) {
|
||||
options.length = 0;
|
||||
for (const match of dialog.replace(/\s+/g, '').matchAll(/(❯)?(?:\d+[.)])?(Yes,Itrustthisfolder|No,exit)/gi)) {
|
||||
options.push({ selected: !!match[1], affirmative: /^Yes/i.test(match[2]!) });
|
||||
}
|
||||
}
|
||||
|
||||
const target = options.findIndex((option) => option.affirmative);
|
||||
const selected = options.findIndex((option) => option.selected);
|
||||
// Wait for both choices and a single cursor. An incomplete or unfamiliar
|
||||
// frame must never turn a guessed default into an implicit rejection.
|
||||
if (target < 0 || selected < 0 || !options.some((option) => !option.affirmative) ||
|
||||
options.filter((option) => option.selected).length !== 1) return null;
|
||||
const distance = target - selected;
|
||||
return (distance < 0 ? '\x1b[A' : '\x1b[B').repeat(Math.abs(distance)) + '\r';
|
||||
}
|
||||
@@ -11,6 +11,7 @@ import * as path from 'path';
|
||||
import * as os from 'os';
|
||||
import { spawn } from 'child_process';
|
||||
import { Readable } from 'node:stream';
|
||||
import { createHash, type Hash } from 'node:crypto';
|
||||
import { getProjectEvalDir } from './eval-store';
|
||||
import { hermeticChildEnv, isHermeticEnabled } from './hermetic-env';
|
||||
import { killProcessGroup } from '../../scripts/test-strict-output';
|
||||
@@ -127,13 +128,70 @@ function truncate(s: string, max: number): string {
|
||||
return s.length > max ? s.slice(0, max) + '…' : s;
|
||||
}
|
||||
|
||||
/** Diagnostic-only projection. Partial input never becomes a complete tool call. */
|
||||
function publicStreamProjection(startTime: number): (line: string) => string {
|
||||
let messageId: string | undefined;
|
||||
const blocks = new Map<number, { type: string; tool?: string; bytes: number; hash: Hash }>();
|
||||
return (line) => {
|
||||
let row: any;
|
||||
try { row = JSON.parse(line); } catch {
|
||||
// A truncated line may contain private reasoning or unfinished tool input.
|
||||
return JSON.stringify({ type: 'public_stream_diagnostic', kind: 'unparseable_line',
|
||||
elapsedMs: Date.now() - startTime, bytes: Buffer.byteLength(line) });
|
||||
}
|
||||
if (row === null || typeof row !== 'object' || Array.isArray(row)) {
|
||||
return JSON.stringify({ type: 'public_stream_diagnostic', kind: 'non_object_line',
|
||||
elapsedMs: Date.now() - startTime, bytes: Buffer.byteLength(line) });
|
||||
}
|
||||
if (row.type === 'stream_event') {
|
||||
const event = row.event ?? {};
|
||||
if (event.type === 'message_start') {
|
||||
messageId = event.message?.id;
|
||||
blocks.clear();
|
||||
}
|
||||
const index = event.index;
|
||||
if (event.type === 'content_block_start' && Number.isInteger(index)) {
|
||||
blocks.set(index, { type: event.content_block?.type, tool: event.content_block?.name,
|
||||
bytes: 0, hash: createHash('sha256') });
|
||||
}
|
||||
const block = blocks.get(index);
|
||||
if (event.type === 'content_block_delta' && block?.type === 'tool_use'
|
||||
&& event.delta?.type === 'input_json_delta' && typeof event.delta.partial_json === 'string') {
|
||||
const chunk = Buffer.from(event.delta.partial_json);
|
||||
block.bytes += chunk.length;
|
||||
block.hash.update(chunk);
|
||||
}
|
||||
const diagnostic = { type: 'public_stream_diagnostic', kind: event.type,
|
||||
session_id: row.session_id, messageId, index, elapsedMs: Date.now() - startTime,
|
||||
blockType: block?.type, toolName: block?.tool, deltaType: event.delta?.type,
|
||||
...(block?.type === 'tool_use' ? { inputBytes: block.bytes,
|
||||
inputSha256: block.hash.copy().digest('hex') } : {}),
|
||||
...(event.type === 'message_delta' ? { stopReason: event.delta?.stop_reason } : {}),
|
||||
};
|
||||
if (event.type === 'content_block_stop') blocks.delete(index);
|
||||
return JSON.stringify(diagnostic);
|
||||
}
|
||||
if (Array.isArray(row.message?.content)) {
|
||||
row.message.content = row.message.content.map((block: any) =>
|
||||
block.type === 'thinking' || block.type === 'redacted_thinking'
|
||||
? { type: block.type, omitted: true } : block);
|
||||
}
|
||||
return JSON.stringify(row);
|
||||
};
|
||||
}
|
||||
|
||||
// --- Main runner ---
|
||||
|
||||
export async function runSkillTest(options: {
|
||||
prompt: string;
|
||||
workingDirectory: string;
|
||||
maxTurns?: number;
|
||||
/** Approval allowlist; does not restrict which tools the model can see. */
|
||||
allowedTools?: string[];
|
||||
/** Optional built-in tool availability. Omit to preserve the CLI defaults. */
|
||||
tools?: string[];
|
||||
/** Opt-in public block timing/input-size diagnostics; never completion evidence. */
|
||||
publicStreamDiagnostics?: boolean;
|
||||
timeout?: number;
|
||||
testName?: string;
|
||||
runId?: string;
|
||||
@@ -198,6 +256,11 @@ export async function runSkillTest(options: {
|
||||
'--max-turns', String(maxTurns),
|
||||
'--allowed-tools', ...allowedTools,
|
||||
];
|
||||
// --allowed-tools controls approval, including when permissions are skipped;
|
||||
// only --tools removes unrelated built-ins such as Agent, Bash, and Skill.
|
||||
// Keep this opt-in: existing workflow evals intentionally use CLI defaults.
|
||||
if (options.tools !== undefined) args.push('--tools', options.tools.join(','));
|
||||
if (options.publicStreamDiagnostics) args.push('--include-partial-messages');
|
||||
// Hermetic children get zero MCP servers (no --mcp-config is passed).
|
||||
// Gated on the same call-time check as the env scrub so EVALS_HERMETIC=0
|
||||
// restores operator MCP along with the operator env.
|
||||
@@ -294,6 +357,7 @@ export async function runSkillTest(options: {
|
||||
const reader = stdoutWeb.getReader();
|
||||
const decoder = new TextDecoder();
|
||||
let buf = '';
|
||||
const projectLine = options.publicStreamDiagnostics ? publicStreamProjection(startTime) : (line: string) => line;
|
||||
|
||||
try {
|
||||
while (true) {
|
||||
@@ -302,8 +366,9 @@ export async function runSkillTest(options: {
|
||||
buf += decoder.decode(value, { stream: true });
|
||||
const lines = buf.split('\n');
|
||||
buf = lines.pop() || '';
|
||||
for (const line of lines) {
|
||||
if (!line.trim()) continue;
|
||||
for (const rawLine of lines) {
|
||||
if (!rawLine.trim()) continue;
|
||||
const line = projectLine(rawLine);
|
||||
collectedLines.push(line);
|
||||
|
||||
// Track time to first NDJSON line (measures latency from spawn to first Claude response)
|
||||
@@ -376,7 +441,11 @@ export async function runSkillTest(options: {
|
||||
|
||||
// Flush remaining buffer
|
||||
if (buf.trim()) {
|
||||
collectedLines.push(buf);
|
||||
const line = projectLine(buf);
|
||||
collectedLines.push(line);
|
||||
if (options.publicStreamDiagnostics && runDir && safeName) {
|
||||
try { fs.appendFileSync(path.join(runDir, `${safeName}.ndjson`), line + '\n'); } catch { /* non-fatal */ }
|
||||
}
|
||||
}
|
||||
|
||||
// Same orphan hazard as stdout: an orphaned grandchild holding stderr open
|
||||
@@ -427,9 +496,9 @@ export async function runSkillTest(options: {
|
||||
if (resultLine.subtype === 'success' && resultLine.is_error) {
|
||||
// claude -p can return subtype=success with is_error=true (e.g. API connection failure)
|
||||
exitReason = 'error_api';
|
||||
} else if (resultLine.subtype === 'success') {
|
||||
} else if (resultLine.subtype === 'success' && exitCode === 0 && !timedOut) {
|
||||
exitReason = 'success';
|
||||
} else if (resultLine.subtype) {
|
||||
} else if (resultLine.subtype && resultLine.subtype !== 'success') {
|
||||
// Preserve known subtypes like error_max_turns even if is_error is set
|
||||
exitReason = resultLine.subtype;
|
||||
}
|
||||
|
||||
@@ -13,7 +13,8 @@
|
||||
* order given. For tests that exercise specific workflow steps.
|
||||
* - extractSkillBody(skillDir)
|
||||
* frontmatter + intro + everything AFTER the shared generated preamble
|
||||
* ("## Preamble (run first)" .. end of "## Plan Status Footer").
|
||||
* ("## Preamble (run first)" or "## Preamble (after scope gate)"
|
||||
* .. end of "## Plan Status Footer").
|
||||
* For tests that exercise the skill's ENTIRE specific flow but never
|
||||
* touch the ~780-line shared preamble.
|
||||
* - extractSkillHead(skillDir, bodyLineCount)
|
||||
@@ -101,7 +102,7 @@ export const CODEX_REVIEW_E2E_SECTIONS = [
|
||||
|
||||
/** First/last H2 headings of the shared preamble block that gen-skill-docs
|
||||
* emits into every tier >= 2 skill. extractSkillBody drops this range. */
|
||||
const SHARED_PREAMBLE_FIRST = 'Preamble (run first)';
|
||||
const SHARED_PREAMBLE_FIRST = ['Preamble (run first)', 'Preamble (after scope gate)'];
|
||||
const SHARED_PREAMBLE_LAST = 'Plan Status Footer';
|
||||
|
||||
interface H2Section {
|
||||
@@ -220,15 +221,25 @@ export function extractSkillSections(skillDir: string, sections: string[]): stri
|
||||
}
|
||||
|
||||
/**
|
||||
* Frontmatter + intro (everything before "## Preamble (run first)") + the
|
||||
* Frontmatter + intro/scope gate (everything before either exact preamble heading) + the
|
||||
* full skill-specific body (everything after the "## Plan Status Footer"
|
||||
* section). Use when a test exercises the whole skill flow: this drops the
|
||||
* ~780-line shared generated preamble and nothing else.
|
||||
*/
|
||||
export function extractSkillBody(skillDir: string): string {
|
||||
const { file, frontmatter, bodyLines, sections: all } = loadSkill(skillDir);
|
||||
const first = findSection(all, SHARED_PREAMBLE_FIRST, file);
|
||||
const last = findSection(all, SHARED_PREAMBLE_LAST, file);
|
||||
const boundary = (names: string[]): H2Section => {
|
||||
const matches = all.filter(section => names.includes(section.heading));
|
||||
const label = names.map(name => `"## ${name}"`).join(' or ');
|
||||
if (!matches.length) throw new Error(`skill-fixture: section ${label} not found in ${file}.`);
|
||||
if (matches.length !== 1) throw new Error(`skill-fixture: ambiguous section ${label} in ${file}.`);
|
||||
return matches[0];
|
||||
};
|
||||
const first = boundary(SHARED_PREAMBLE_FIRST);
|
||||
const last = boundary([SHARED_PREAMBLE_LAST]);
|
||||
if (first.start >= last.start) {
|
||||
throw new Error(`skill-fixture: "## ${SHARED_PREAMBLE_LAST}" precedes the preamble in ${file}.`);
|
||||
}
|
||||
const intro = bodyLines.slice(0, first.start).join('\n').trimEnd();
|
||||
const tail = bodyLines.slice(last.end).join('\n').trimEnd();
|
||||
if (!tail) {
|
||||
|
||||
+562
-83
File diff suppressed because one or more lines are too long
@@ -0,0 +1,86 @@
|
||||
/** Preserve source-file boundaries when a workflow judge reads carved skills. */
|
||||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
|
||||
export interface WorkflowJudgeFile {
|
||||
path: string;
|
||||
kind: 'entrypoint' | 'section';
|
||||
content: string;
|
||||
startLine: number;
|
||||
endLine: number;
|
||||
}
|
||||
|
||||
export interface WorkflowJudgeInput {
|
||||
files: WorkflowJudgeFile[];
|
||||
text: string;
|
||||
}
|
||||
|
||||
export function readWorkflowJudgeInput(opts: {
|
||||
root: string;
|
||||
skillPath: string;
|
||||
startMarker: string;
|
||||
endMarker: string | null;
|
||||
}): WorkflowJudgeInput {
|
||||
const sources = [{
|
||||
path: opts.skillPath,
|
||||
kind: 'entrypoint' as const,
|
||||
content: fs.readFileSync(path.join(opts.root, opts.skillPath), 'utf8'),
|
||||
}];
|
||||
const sectionDir = path.join(path.dirname(opts.skillPath), 'sections');
|
||||
const sectionRoot = path.join(opts.root, sectionDir);
|
||||
const sections = fs.existsSync(sectionRoot)
|
||||
? fs.readdirSync(sectionRoot).sort().filter(name => name.endsWith('.md'))
|
||||
.map(name => ({
|
||||
path: path.join(sectionDir, name),
|
||||
kind: 'section' as const,
|
||||
content: fs.readFileSync(path.join(sectionRoot, name), 'utf8'),
|
||||
}))
|
||||
: [];
|
||||
const allSources = [...sources, ...sections];
|
||||
|
||||
// Preserve the existing marker window, including markers that moved into
|
||||
// section files. Offsets identify the source of each slice; prose prefixes
|
||||
// cannot reliably identify a file after its generated header was sliced off.
|
||||
const union = allSources.map(file => file.content).join('\n');
|
||||
const start = union.indexOf(opts.startMarker);
|
||||
if (start < 0) throw new Error(`Start marker not found in ${opts.skillPath}: "${opts.startMarker}"`);
|
||||
const end = opts.endMarker === null ? union.length : union.indexOf(opts.endMarker, start);
|
||||
if (end < 0) throw new Error(`End marker not found in ${opts.skillPath}: "${opts.endMarker}"`);
|
||||
|
||||
const files: WorkflowJudgeFile[] = [];
|
||||
let offset = 0;
|
||||
for (const file of allSources) {
|
||||
// Every section was already supplied in full by the old judge input. Keep
|
||||
// that coverage, but include each file once even when the marker window
|
||||
// also covers part of it. Only the entrypoint retains the requested slice.
|
||||
const from = file.kind === 'section' ? 0 : Math.max(0, start - offset);
|
||||
const to = file.kind === 'section' ? file.content.length : Math.min(file.content.length, end - offset);
|
||||
if (from < to) {
|
||||
files.push({
|
||||
...file,
|
||||
path: file.path.split(path.sep).join('/'),
|
||||
content: file.content.slice(from, to),
|
||||
startLine: file.content.slice(0, from).split('\n').length,
|
||||
endLine: file.content.slice(0, to - 1).split('\n').length,
|
||||
});
|
||||
}
|
||||
offset += file.content.length + 1;
|
||||
}
|
||||
|
||||
const context = [
|
||||
'The material below is a bundle of source-file excerpts, with each original file and line range labeled.',
|
||||
'SKILL.md is the entry point; the labeled ranges identify which excerpts are supplied.',
|
||||
...(sections.length > 0 ? [
|
||||
'The section files remain separate on disk and are read at the points and conditions specified by the skill\'s Read directives.',
|
||||
'They are supplied here as on-demand references; their order in this bundle is not execution order.',
|
||||
] : []),
|
||||
].join('\n');
|
||||
return {
|
||||
files,
|
||||
text: context + '\n\n' + files.map(file => [
|
||||
`--- BEGIN FILE ${JSON.stringify(file.path)} (lines ${file.startLine}-${file.endLine}; ${file.kind}) ---`,
|
||||
file.content,
|
||||
`--- END FILE ${JSON.stringify(file.path)} ---`,
|
||||
].join('\n')).join('\n\n'),
|
||||
};
|
||||
}
|
||||
Reference in New Issue
Block a user