v1.91.12.0 v1.91.12.0: audit fix wave, ~11-minute paid eval lanes, eval reliability policy (#2999)

* test: delete test-infrastructure dead code (G)

- exit-propagation drives the runner's real strict verdict
  (BunTestOutputClassifier + strictTestExitCode); delete the unused
  shardRunLooksTruncated predicate.
- delete skill-coverage-matrix registry + its gate (nothing reads it; the
  floor already iterates skillCensus()).
- delete touchfiles-facade export-parity tests (Bun fails missing imports
  at link time) and the duplicated E2E_TIERS tier-value test.
- delete brain-cache-spec TRANSPORT_DEFAULT_POLICY, SKILL_RUN_RETENTION_DAYS
  and the now-unused BrainTrustPolicy type with their literal tests.
  AUTOPLAN_PREFLIGHT_BUDGET_BYTES stays: skill-preflight-budget enforces it
  against real resolver output.
- delete audit-compliance's JSDoc-comment grep.

* test: replace product tests that fake the product with real-boundary tests (F)

- design: serve.test.ts drove an inline mirror server; now two tests run the
  real serve() on an ephemeral port (reload confinement, submit exit 0).
- setup-gbrain: rollback + voyage tests execute the template-extracted init
  blocks (3 sites) instead of drifted local bash copies.
- terminal-agent: internalHandler source greps replaced by a behavioral
  /internal/grant + /internal/revoke auth matrix (no/wrong/valid token).
- /health: server-security-surface and the server-auth / security-audit-r2 /
  sidebar-tabs source greps fold into one liveness-only check on the real
  body; the L4 sidecar wiring gets a behavioral /pty-inject-scan test.
- delete tautologies (browser-manager onDisconnect, memory-command #12),
  ios swiftui tap fixture self-check, memory-ingest put_page grep, detach
  source greps, sidebar-agent absence pins, dead-CSS pins + the dead CSS,
  security-audit-r2 Task 1 + the test-only meta-commands re-export,
  duplicate generated-SKILL.md checks.
- make-pdf coverage-gaps cases move into their owner test files.

* test: delete tests of dead eval code (A)

- A1: the retired Eng lexical oracle (evaluateEngSeedCoverage,
  isEngSeedDecisionAUQ), the completion-handoff detector and the retained
  corpus had no paid caller since v1.87.6; delete their 26 replay files,
  ~2.6k helper LOC and fixtures, and the dead blocks in 8 mixed files
  (live hasNativePlanTerminal / batching assertions stay).
- A2: dead viewport approvers in autoplan-artifact-permission and their 11
  replay files + fixtures; recorder/launcher cases stay.
- A3: never-wired oracles and seeders (autoplan-phase-order,
  eng-finding-fixture, ceo-paired-fixture, design-ui-scope,
  plan-skill-completion, pty-current-screen, required-reads,
  transcript-section-logger); plan-seed-submission now decodes through the
  production createPtyScreen; section manifests name their actual guard.
- A4: zero-reference helper exports, plus execGit and invokeAndObserve
  found by the reachability pass.
- 52 fixtures orphaned by the deletions; touchfile and selection-table
  entries for every deleted path.

* test: clean up the paid eval lane (B1-B4, B6, B7)

- B1: delete paid files that assert nothing or cannot pass meaningfully:
  skill-llm-eval-spec and skill-e2e-spec-execute (test.todo), gemini-e2e
  (+ gemini-session-runner; no gemini CLI in CI), ship-idempotency (red
  since v1.63), the two opus-4-7 *-sonnet overlay wrappers, conductor-prose
  (+ its source-evaluation replay), codex-e2e-plan-format; drop their keys,
  scripts and census rows.
- B2: skill-llm-eval grades browse/sections/command-list.md with one union
  judge that also carries the baseline score pin; regression-vs-baseline
  deleted (paid run: pass, c4/c4/a4).
- B3: memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral
  make no model calls; renamed out of the paid glob so they run on every
  PR. Swift builds need GSTACK_TEST_SWIFT=1; device stub deleted.
- B4: codex-e2e*, outside-voice, aside and ios-device cannot run in the CI
  image; excluded from the weekly lane with a tracked re-entry condition.
- B6: fold opus-47's negative routing controls into skill-routing-e2e
  journey-negatives (paid run: 3/3 unrouted) and delete the file.
- B7: delete the never-green brain-privacy-gate eval; a free
  gstack-skill-start test now proves consent precedes artifacts egress.

* test: retire the finding-count cluster and trim its helpers (C)

- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
  skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
  budget, never on skill behavior; delete them, their touchfile/tier ids,
  AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
  (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
  closure, and delete the free replay tests whose assertions exercised only
  that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
  only as input for a live subject keep their assertions: the multiSelect
  default moved to plan-review-decisions, runner PTY tests use inline caller
  policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
  (its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
  and per-finding cadence coverage with their re-entry tests.

* test: fold per-incident replay series into their detector owners (D)

Twelve detector families move into one owner test each: 73 incident files
become describe blocks in ceo-section-loading-fixture (stale-fill race),
model-overlays, coverage-audit-evidence, autoplan-phase-observer,
native-auto-decide, outside-voice-evidence, eng-first-review,
plan-count-completion, plan-count-file-permission, ceo-mode-option,
plan-scope-selection and plan-count-prerequisite. Each block keeps its original
code and fixture, so every case still runs; only tests asserting the incident
file's own touchfile registration are dropped (41). Touchfile lists that named
an incident now name its owner.

* test: start the plan-count history PTY on its readiness marker (H)

The fake CLI prints a startup marker and the runner waits for it instead of the
fixed 8 s startup sleep (8.6 s -> 0.9 s locally). eng-semantic-terminal's
sleeping registration cases went with C; plan-count-timeout keeps the fixed wait
because it asserts deadline behavior.

* test: derive paid touchfiles from each eval's static closure (E)

touchfiles.test.ts now checks, per key, that the paid file's static
test/helpers and test/fixtures closure (plus fixture paths it names in string
literals) is covered, and names the file, path, chain and key to fix when it is
not. Free *.test.ts files are no longer touchfiles, so editing a free replay
test stops selecting paid evals: 950 entries removed, 653 real closure paths
added. The hand-copied inventories go: periodic-fixture-selection,
fake-impeccable-touchfiles and 45 per-file selection examples. Selection for
the sample edits (plan-eng-review template, claude-pty-runner,
plan-count-fixture, gstack-config) loses no case under either profile.
CONTRIBUTING documents the rule and its lower bound.

* test: skip hollow tier shards and census judges in the paid planner (B5)

A paid file is now skipped for a tier lane only when every E2E id it registers
is known statically and none has that tier; ids come from the touchfile
registrations and literal testName/*IfSelected arguments, so a comment or
skill path that quotes another id cannot unschedule it, and computed names
keep today's scheduling. --list and the manifest show each skip as
"skipped: no E2E_TIERS id has tier <tier>". The weekly gate census drops the
LLM judges (--skip-judges); they still run in the periodic census and PR gate
lanes. Gate lane 52 -> 42 files, census 41; periodic 77 -> 69.

* test: run seven paid evals on the current default capture model (B8)

skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned
claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned
claude-sonnet-4-6; none tests a historical model, so they now capture with
resolveEvalModel('capture'), and the free harness tests that execute these
registrations receive the same resolver. The paid re-pin run passed all of
them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep
claude-opus-4-7: six of their cases failed on the default model (three
timeouts, a missing report file, a format miss and a posture score of 3), so
per the plan's fallback they keep their pins with a TODOS entry. The pre-spend
estimate and drop threshold are in docs/test-audit-2026-09.md.

* test: guard the reduced suite against new test-of-test files

- test/test-of-test-ratchet.test.ts records the 228 free tests that import only
  test/ code and fails on a new one, naming the owner test to extend instead;
  a stale baseline entry fails with the remove instruction.
- test/helpers/resolve-repo-path.ts is the one specifier/literal resolver for
  the ratchet and the touchfile closure invariant, with its own unit tests.
- CONTRIBUTING "Test tiers" describes the paid-failure workflow (fix, then one
  row in the detector's owner test) and the ratchet; TEST_PORTFOLIO gains the
  detector -> owner-test table and no longer claims an Autoplan chain eval.
- TODOS: automatic exclusion policy for chronically red periodic files (P3),
  the deferred native-completion table collapse, the unused CEO payment
  seeder; the PTY readiness item is narrowed to the paid runner.
- docs/test-audit-2026-09.md collects the triage, security mapping, inventories,
  selection proof, behavior-commit decisions and retained false positives.

* v1.91.8.0 test: smaller suite, derived paid selection, retired never-green evals

Release metadata for the test-reduction branch: VERSION 1.91.8.0 (1.91.7.0 is
claimed by #2983), CHANGELOG with the measured before/after table and a
contributor section, durations re-recorded on Ubicloud standard-16 (857 files,
0 failures), the agents digest, CONTRIBUTING's after-measurement row, the B8
fallback TODOS entry, and the after metrics, kept-vs-plan notes, B8 run and
census estimate in docs/test-audit-2026-09.md.

* fix(ubicloud): skip retrieval globs that match nothing instead of reporting a failed pull

* test: pin DISABLE_AUTOUPDATER in hermetic env and capture corrupt-seed warning

Both EVALS_HERMETIC branches of buildHermeticEnv now carry
DISABLE_AUTOUPDATER=1 (the allowlist scrubbed the workflow's copy, so every
PTY screen showed the updater's npm-prefix failure). Per-test overrides
still win. The corrupt durations-seed test now captures its expected
warning and restores the console spy.

* style(cso): format lib/cso TypeScript with pinned Prettier

Mechanical reformat only. Minified transpile output is byte-identical for
21 of 22 files; witness.ts differs only in three regex flag orders
(/mi -> /im), which JavaScript canonicalizes. Source-text assertions over
lib/cso now compare whitespace-insensitively with the same tokens.

* fix(cso): import join for compiled-launcher assertion witnesses

Compiled installs always take the non-Bun branch, which called an unimported
join and threw before any runtime-tested assertion could be witnessed. The
child command selection is now a pure, platform-aware function; a missing
sibling launcher fails with its expected path.

* fix(browse): make connect --supervise actually respawn a crashed server

The supervisor respawned with a block-scoped env that no longer existed, so
every attempt threw and the loop gave up after five tries. The headed env is
now one pure helper used by connect and respawn, the loop is an injectable
runHeadedSupervisor with behavioral tests, failures name the daemon log and
relaunch command, and connect's usage advertises --supervise.

* test: one finite PR world for the shared-libs fixture; name dual-voice probe evidence

The shared-libs shim served 2 PRs for pulls?state=all and endless full pages
for state=open. gh pr list, pulls?state=open|all|closed (per_page/page,
short last page, direction) and search/issues now page one deterministic
table: PR 7, 600 older open PRs, PR 42 and 3 closed PRs, so five 100-item
open-metadata pages still leave older open PRs unchecked. The Contents API
lists pinned directories (the captured attempt got 404 for contents/ and
contents/src while files resolved, then fell back to a raw host), unknown
endpoints return 404 instead of repo metadata, and the read-only detector
is unchanged. Free tests cover view agreement, the budget bound, gh/curl
agreement and the empty world.

Dual-voice outside-voice failures now report probeToolUseId, probeMode and
the canonical-match result with the reason the probe output was rejected.

* feat: require a zero-error product typecheck and a test type-debt ratchet

Adds tsconfig.json (strict) over product code, fixes its remaining 90
diagnostics (type-only, interface corrections, and explicit narrowing),
and adds a typecheck job to the required free-tests aggregate running
bun run typecheck, the test-code ratchet (identity -> count baseline, fails
on new, repeated, or unlocked fixed diagnostics), and the lib/cso format
check. Reuses fixes from #2447 where they still applied.

* test: follow the headed env helper and the typecheck gate in source-shape checks

* fix(test): pin the package.json change kind in shared-input selection tests

computePaidCaseSelection read the version-only exemption from git even when
changed files were injected, so the shared-input test failed on main and on
version-only branches. The exemption is now an optional input; the test pins
a real package.json change and covers the version-only case.

* test: judge plan-count completion on structured evidence, not wording

Replaying run 36385945043's two Design attempts showed the existing routes
rejected correct endings: attempt 1 at the typed-completion path field
('- Reviewed plan written to …' is not a 'Plan written to' line), attempt 2
at the leading-fence veto (its final message opens with the dashboard).

nativePlanTerminalPreconditions is the structural prefix of
hasNativePlanTerminal (behavior unchanged). structuredPlanCompletion adds,
inside the existing nativeSummary branch: a complete report (Design
binding for Design), a completed review-log row for the expected skill
appended during this attempt under the child's GSTACK_HOME/project slug
(resolved with bin/gstack-slug) and stamped with the fixture commit, timed
between the report/last answer (second resolution) and the final native
message, a final message with stop_reason end_turn (now carried on public
transcript messages), and no visible question or permission prompt.

Timeout summaries add idleFor and lastTerminalCandidate. Terminal and throw
captures copy the plan file and review-log rows into the artifact
directory; copies are best-effort and recorded in evidence-copy.json.
Free regressions: both captured Design endings (trimmed fixture with
provenance; report, row and end_turn reconstructed and labelled), the
negative controls, and real-PTY completion/timeout runs through the real
review logger.

* test: structural Design count boundary; TODO proposals are not findings

Replaying run 36385945043 through the Design count predicates: routing,
focus and learnings setup was not recognized as setup, Issue 1 was counted
pre-review in both attempts (the boundary fired on it), and attempt 2
counted the Font TODO proposal as a finding (review=4 and review=5 for five
issues). The paid caller now starts review at the first answered native
decision that is not setup (recognized packet, or setup header/question ID),
a completion handoff, artifact rendering or a TODO proposal (the review's
Add to TODOS.md / Skip / Build it now menu). TODO proposals are recorded as
administrative extra decisions. The replay asserts each counted call: both
attempts review=5 (Issues 1-5). isDesignCountFirstReview and its controls
are unchanged.

* test: CEO classifier throws name the question and matched predicates

Replaying run 36385945043's FAN-1 and ERR-1 throws (ledger rows
reconstructed from rendered diffs) through ceoPaymentFinding: the email
obligation's row, subject, option and proposal predicates pass and the
ELI10 explanation-defect predicate fails first ('lets that exception fly
out', 'the error bubbles up').

Binding the defect to the named ledger row instead (the planned fix) was
tried and reverted: scoped to the email seed it flips 30+ existing cf74
still-rejects replays, which require a vocabulary-free, ledger-bound email
question to earn credit only through a complete saved comparison. With
FAN-1's rendered currentDecision payload reconstructed, the recorded-
decision path counts it, so the real saved plan (not uploaded) must have
differed; failure artifacts now retain it.

The classifier stays fail-closed and unchanged. Its throw now prints the
header, the first 200 question characters and each obligation's predicate
results. Free regressions with provenance and negative controls: an
unrelated question, an email question whose row says it is already
rescued, and a ledger ID whose row belongs to another seed.

* chore: regenerate the test type-debt baseline on top of #2994

* fix(typecheck): strip the checkout root from ratchet diagnostic identities

* fix(test): recognize ledger row-ID split candidates so collection stops at the last ACK

Run 36385945043's split-overflow case asked all five candidate decisions by
8m55s, but the live candidate check required the question to open with
"E1:" and every option to be a known disposition. The skill cited ledger
row IDs ("D2.1 — R-E1: …") and offered "Hold, discuss first", so no
candidate was recognized and the attempt ran the whole review (1302s).

Identity now comes from the native header; the question must open with that
candidate's ledger reference, name only that candidate, and offer exactly one
include, defer and cut disposition. The selected answer must still be one of
those three. The semantic evaluator and every existing negative control are
unchanged; a trimmed capture from the run adds the positive case and four
row-ID negative controls.

* fix(test): stop the eng batching eval once its floor is proven

The case's only verdict is reviewCount >= FLOOR (3). Run 36385945043 had
three distinct acknowledged review decisions at 6m41s but kept answering
until the ceiling (7) at 12m13s. The registration now passes the runner's
existing isCollectionComplete stop once FLOOR non-setup, non-administrative
review decisions are acknowledged; the floor check, ceiling, budget and
counter are unchanged. A child-process registration test proves the stop
predicate and that below-floor and timeout outcomes still fail.

* test: add the non-blocking 'marathon' E2E tier

Full start-to-finish flows move out of the blocking lanes. E2E_TIERS and
E2ETier gain 'marathon'; describeE2ETier('marathon') is enabled only when
EVALS_TIER=marathon, so the gate/PR and periodic lanes (and the gate census)
never run those cases. The PR profile accepts marathon ids as scheduled
elsewhere and defers them with their own reason, even on full fallback.

* test: move the full office-hours workflow to marathon; add a periodic design-draft checkpoint

The full startup workflow runs 1–3 real spec-review rounds (~280s each) and
hit its 1200s capture in run 36385945043 at finalize. Review depth is the
product's loop, so the case cannot fit a blocking lane without cutting
rounds. It is now marathon tier with every assertion unchanged.

skill-e2e-office-hours-design-draft.test.ts (periodic) runs the same fixed
interview only through the Write that creates the design (269s in that run)
and applies the full validator's design-draft checks, the required section
reads and the launch/foreign-skill-read guards. validateOfficeHoursDesignDraft
is extracted from validateOfficeHoursCompletion, which still applies it.

Selection: office-hours-design-draft is registered periodic; the marathon-only
file is already excluded from the gate and periodic plans by the B5 planner
rule. Tier-alignment regexes and the valid-tier check accept 'marathon'.
A type-only cast in plan-scope-selection.test.ts removes a diagnostic whose
union print order made the ratchet identity unstable; baseline tightened.

* test: supply the split-overflow fixture's HOLD SCOPE mode as a prerequisite

The split actor always answered 0E's mode question with HOLD SCOPE. The
skill skips that question on an explicit choice, so the fixture now states
it and the attempt starts at the five candidate decisions (about 1.5 min
earlier in run 36385945043). Candidates, actor policy, floor and semantic
evaluation are unchanged; the fixture test pins the supplied choice.

* test: start the eng batching eval with its setup prerequisites supplied

Routing setup and cross-project learnings (D1/D2 in run 36385945043) are
never counted and are not what the case measures. The registration now uses
the runner's existing preconfiguredReviewActor so the attempt starts at the
review; engSetupAUQ still vetoes any late setup question. The registration
test pins the option.

* test: count the design-draft paid file and defer marathon ids in PR selection pins

The discovered paid-file census grows by one (skill-e2e-office-hours-design-draft).
Full-fallback PR selection defers every non-gate id; the shared-input pins now
expect periodic and marathon ids there.

* fix(review): resolve the judged revalidation, setup-authority, plan-gate and findings-record ambiguities

The census review workflow judge scored clarity/actionability 3 on both
attempts: smoke-clock limits appeared to forbid post-repair revalidation,
the caller deadline was undefined, 'ask for setup' conflicted with the
report-only browser rule, fallback-sourced HIGH discrepancies had no gate
decision, and the Step 5.8 record omitted adversarial findings.

* fix(office-hours): load the builder section for every builder-mode reply

Both census builder-wildness attempts answered a direct request for
adjacent unlocks without reading phase-2b-builder-brainstorm.md, whose
trigger read as applying only to the generative questions.

* fix(sync-gbrain): define Step 4 helper args and one atomic write path

Both census read-ready attempts spent turns reading the helper source to
resolve <user-args>, inspecting fixture internals kept inside the repo,
and reconciling 'Read + Edit' with the tmp+mv atomic write, then hit
max turns before the verdict.

* refactor(evals): share the import-closure walker and add the E2E shard reuse identity

sourceDependencyClosure moves from the workflow-judge adapter into
scripts/eval-input-cache.ts unchanged, so judge keys stay byte-identical.
scripts/e2e-shard-reuse.ts builds the consumed-input identity of one PR-lane
E2E shard (test import closure, every registered case's touchfiles, globals,
runner/workflow/setup actions, child env pins, CI image, Claude CLI) and fails
closed on anything unknown. Marathon joins the always-fresh purposes.

* feat(evals): ~12-minute blocking paid lanes and a non-blocking marathon lane

- Planner budget mode (--slice-budget S --jobs J): recorded per-tier wall
  times pack into as many ~9-minute executors as the work needs; the plan
  records per-slice estimates and the CI job timeout (supervised worst case
  + 20 min). evals.yml and evals-periodic.yml derive matrix size and
  timeout-minutes from it; max-parallel covers every slice at once.
- Case shards: plan/design/review-army/shared-libs(-paths) run one registered
  case per process (<file>#<case id>, exact name pattern, exactly one case).
- Retry rule: a timed-out attempt is a verdict. Only files whose every case
  budget is CAPTURE tier or shorter keep one retry; walls shrink to match.
- Marathon tier: positive selection, excluded from gate/periodic planners,
  run by the new evals-marathon.yml (weekly + dispatch, fresh, own report).
- PR-lane E2E reuse of verified first-attempt passes on identical inputs;
  the report rejects reuse outside the fast PR profile.
- Duration seed from census run 36385945043, per tier and per case shard.

* docs: blocking lane budget, marathon lane, retry policy and E2E reuse

* chore(typecheck): lock in two fixed test diagnostics

* fix(ci): drop a duplicated env/jobs block in evals-marathon.yml

* test(ship-docsync): shard the doc-sync lifecycle by case and drop the duplicate dispatch-only case

ship-docsync ran the same fixture and prompt as ship-docsync-completion and
asserted a subset of it. The file now runs one case per process, so its lane
wall is its longest case instead of half the sum of thirteen.

* fix(evals): plan CI-unrunnable cases as excluded entries, not empty case shards

design-review-fix drives the Aside browser and registers test.skip on Linux
runners, so its case shard executed zero cases and failed the exact-one-case
check in proof census 36597762183 (eval-slices 6). CASE_CI_EXCLUDE (reason +
tracking, beside PERIODIC_CI_EXCLUDE) now turns such cases into excluded
manifest entries that --list and the manifest surface; every planned case
shard still must execute exactly its case.

* docs(todos): list the case-level Aside exclusion with the CI-unrunnable evals

* fix(plan-ceo-review): restore experience-first expansion framing, require the mode handoff, skip pacing menus

Census 36597762183: both mode-routing runs logged provenance and moved on
without the mandated handoff chat; the EXPANSION run asked an unauthorized
batch/narrow pacing menu instead of the first per-addition question; the
expansion-energy proposals led with the spec because v1.87.6.0 dropped
'lead with the felt experience'. The HOLD review detector also rejected a
decision whose grounding line named no plan file although the owned source
Read binds it.

* test(outside-plan-disabled): bind quoted prior-record values by their sentence, not phrase order

The parent obeyed the off switch and twice named the seeded completed
record as pre-existing, once with the quotation after its owner and once
with slash separators; the order-specific stripper counted both as current
completion. Timestamp, location, current-claim and value-match controls
still reject.

* test(outside-plan-disabled): compare named record timestamps as instants; negated authorship is not a current claim

The repair rerun named the seeded record by its ISO second
(2026-09-29T16:58:52Z vs .727Z) and said 'I did not write'; both were
misread as a foreign timestamp and a current write.

* test(ceo-section-loading): recognize an arrow-ordered stale-fill execution by event roles

The census review traced the seeded race as 'R1 miss -> R1 store read (v1)
-> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2
(begun after W) hits v1', but the in-flight gate only accepted race
vocabulary or fixed sentence shapes. Order, actor, version and dismissal
mutations still fail.

* test(design-floor): answer the seed-declared all-seven 0D focus menu while it is pending

The actor declares 'Design: review all seven dimensions', but its picker
reused designReviewSetupAUQ, which only matches already-answered calls
(and a narrower header/label set), so the pending D1 focus menu was never
answered and the case waited out its 609 s deadline. The skill's Step 0D
requires asking; the fixture now answers it.

* test(ceo-mode-routing): accept the skill-mandated Note form and Recommendation reason as HOLD posture

HOLD Defer/Keep briefs must use 'Note: options differ in kind' (preamble),
but the answered-HOLD path demanded a Completeness score, rejected a
one-line Net with a semicolon, and read posture only from ELI10. The rerun's
brief applied HOLD SCOPE in its Recommendation reason. Revert the
ineffective 'always'/'handoff chat' wording: two runs still skipped the
mode handoff.

* test(qa-bugs): keep claude-opus-4-7 after qa-b6-static stalled on the default model

qa-b6-static timed out on claude-fable-5-1 in census 36597762183 and in one
of two targeted reruns. Both times the stream stopped mid-message with no
pending tool, right after the model found the disabled submit button, and
stayed silent until the 300 s deadline. Per the B8 fallback, re-pin with a
TODOS entry; budgets and retries are unchanged. A rerun on opus-4-7 passed
(125 s, 5/5 detected).

* test(evals): add E2E_KINDS, BEHAVIOR_WHY, EVAL_POLICY and CASE_QUARANTINE skeletons

Every E2E_TIERS and LLM_JUDGE_TOUCHFILES key starts as 'rule'; BEHAVIOR_WHY
and CASE_QUARANTINE start empty. EVAL_POLICY pre-registers the approved
panel (3, majority 2), quarantine entry 0.95/10 and exit 0.97/10, 10% cap,
8-weekly-run expiry, Fisher drift alarm and one INFRA re-dispatch.

* test(evals): add trial records, panelVerdict, expectContract and trial-outcomes JSONL

EvalTestEntry gains case_id, kind, trial, panel, failure_class and
policy_version, stamped from the runner's TRIAL_ENV on isolated trial
shards. panelVerdict() is the single verdict function (INCOMPLETE on
missing or duplicate trials, contract veto at any count, quarantine
hard-break rule, INFRA/INCOMPLETE machine classification). expectContract()
records failure_class 'contract' on the collector entry and a sidecar
before throwing. trial-outcomes JSONL has a fail-closed writer and a
data-only reader.

* test(evals): pin the fail-closed rule-shard gate through the real --report path

Synthetic slice artifacts for rule fail, timeout, missing slice, unreported
entry, hollow, never-started, collector failure and wrong-slice reports all
exit red before the panel-verdict gate change lands.

* test(evals): retire every paid automatic retry

Paid evals never retry (approved 2026-09-29): delete SHORT_CASE_RETRY_FILES
and retriesWithinCaseCap, drop the retry fields from the registered wall rows
(walls now cover one run plus reserve), make retriesForFiles return 0, pass
--retry 0 explicitly, and drop --retry 1 from the package.json paid scripts.
Add the eval:pass-rates alias. Tests that pinned the old retry allowance are
updated as a policy change; review-finalization-budget now proves late-result
recording under the production zero-retry arguments.

* test(llm-judge): sample every judge as a pre-registered 3-sample panel

Each of the 24 skill-llm-eval judges now draws EVAL_POLICY.judge.samples
independent samples of the same prompt concurrently inside the unchanged
JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean
against the unchanged threshold; booleans (would_browse, consistent) on a
strict majority. An erroring sample fails the whole panel and is never
resampled; a refusal is an unscored panel only when every sample refused.
callJudge's 429 backoff stays: it is transport before any model output.

The workflow-judge cache stores and validates only complete panels, and its
identity now records the panel and zero file retries. Harness tests that
pinned one provider call per case now pin the panel size.

* test(evals): classify every live case and re-select a case when its kind changes

E2E_KINDS: rule by default (191 E2E ids), 22 behavior cases whose verdict is
a live model choice with an acceptable sub-100% per-trial rate, each with a
BEHAVIOR_WHY tolerance, and 25 judge entries (the 24 workflow judges plus the
fixed-fixture llm-judge-recommendation rubric check). Contract-shaped cases
(ask-before-decide, plan-mode no-writes, mandated steps, secrets, the batching
floor) stay rule. Behavior requires a known literal registration and an exact
Bun test name so the case runs as its own trial shard.

Map-diff selection now diffs E2E_KINDS and BEHAVIOR_WHY per key, and a base
revision without them selects every key, so a kind flip runs the panel it
introduces. test/eval-kinds.test.ts enforces coverage, tolerances,
isolatability and the reviewed counts, printing the literal to add.

* feat(evals): per-case pass rates with Wilson intervals, identity series and quarantine policy

scripts/eval-flake-rank.ts becomes eval:pass-rates (eval:flake-rank stays an
alias, and the legacy aggregate stays exported). It reads eval-store's
trial-outcomes JSONL from the last N completed evals-periodic runs on this
branch and main (gh, downloading only the trial-outcomes artifact, cached and
size-capped, parsed as data), plus local eval dirs, and prints per-case
per-trial pass rates with 95% Wilson intervals.

A series is a case's own touchfiles minus GLOBAL_TOUCHFILES
(caseSeriesIdentities, for the report job to stamp), per model, CLI version
and policy version. Labels: INCONCLUSIVE, BROKEN, FLAKY, FAILING, PASSING.
--backfill imports legacy slice artifacts as pre-policy trials (first
attempt only, attributed by registry id, never guessed) for display only.

--gate fails with ACTION REQUIRED on post-policy evidence only: drift below
the quarantine entry rule, a rule case behaving like behavior, a one-sided
Fisher drop against the previous identity (Holm-controlled), and quarantine
entries that met their exit rule, expired after 8 weekly runs, broke the
10% tier cap or are invalid. CASE_QUARANTINE entries now carry a
failureClass (detector, harness or model-latency); a product defect has no
class and is never quarantined. The policy test pins EVAL_POLICY's approved
constants.

* feat(eval-pass-rates): attribute legacy records by the exact slug of their display name

* ci(image): pin Claude Code 2.1.284 so the eval model is recognized

2.1.251 logs [claude-code:unrecognized_model] for claude-fable-5-1, the
eval capture/judge default. 2.1.284 does not. The gate PTY smoke subset
(plan-ceo/plan-devex plan-mode, plan-mode-no-op) parses on the new TUI;
plan-design-review-plan-mode passed at 293 s on 2.1.284 and timed out at
300 s on 2.1.251 on the same tree.

* test(eng-batching): grade the floor once the review report is complete

A completed GSTACK REVIEW REPORT ends the review, so the review-question
count is final there. Run 36606688266 wrote its report at 1,248 s and
closed the session at 1,318 s; the case now stops collection and applies
the unchanged floor at the report instead of waiting out the session.
No budget changes.

* test(eng-batching): bind unsourced native briefs through the report's target

Run 36606688266 asked ten separate native review questions (D1-D9 bound
to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs
named the plan by title instead of citing PLAN.md, its report declared
'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and
it kept an unfenced copy of the plan's own H1. The named-source route now
accepts those spellings and non-inline ledger briefs. The same replay
rejects a foreign, mixed, duplicate or missing target, another plan's
title or copied H1, a brief naming another plan or file, a mismatched
saved brief, and re-asks. The run-36597762183 capture still counts 3.

* fix(plan-design-review): treat a designer with no API key as unavailable

Both proof runs (36597762183, 36606688266) printed DESIGN_READY, hit
'No OpenAI API key found' on the first $D variants call, then hand-built
HTML/CSS wireframes, screenshots and a comparison board for ~195-245 s
before the first review question; the second run timed out at 600 s.
A failed first generation now takes the existing text-only path, and the
skill forbids substituting hand-built mockups.

* fix(deslop-shared-libs): read related sources together within the turn limit

Run 36606688266's opportunity audit read sixteen sources one per turn and
stopped at error_max_turns; the passing run 36597762183 read the same
files in three batched commands. The skill now says turns are bounded and
asks for parallel reads or one read-only command per step.

* test(ceo-mode-routing): submit a mode review that scrolled past the viewport

Run 36606688266 bundled routing, learnings and the mode choice into one
native call. Its review panel was taller than the terminal, so the tab
bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD
SCOPE was never submitted ('no posture match'). With no bar on screen the
viewport must still end at the focused Submit prompt, and the accumulated
screen text supplies the one complete panel, authenticated exactly as
before. Replay controls reject another mode, an unoffered answer, an
altered question, a quoted panel, trailing output, a moved cursor and an
answered or changed call.

* docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic

AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.

* feat(evals): trial planner, slice exit split and panel-verdict report

Planner: behavior and quarantined cases become panels of isolated trial
shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact
test name; the file shard excludes them by name. Trials of one case never
share a slice, result slugs are unique, panels are validated whole, unknown
registrations throw, and the planner prints a capacity preflight.

Executor: each trial shard gets its TRIAL_ENV identity and a trial record
(outcome, failure class, cause, cost); every shard writes a JUnit report.
The slice exit now means execution completeness: a failed rule shard or a
trial without a record reds the runner, a failed trial does not.

Report: panelVerdict() decides every panel of the first run attempt (later
attempts are reported, never replacing it); rule shards keep the unchanged
fail-closed checks; collector records all count (no last-attempt wins);
census runs enforce the quarantine cap and expiry. It writes
collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge
cases), report-summary.md, and one headline + failure block with rerun
commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch.

The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green
with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red,
quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later
attempt never replacing the first.

* chore(evals): refresh paid duration seeds from proof runs 36597762183 and 36606688266

Both tiers, merged in run order (the later run wins). Notable: split-overflow
1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s;
multi-finding-batching 734s -> 1318s (its red path in run 36606688266).

* feat(evals): stamp trial series identities and fit panels to the live registry

- scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's
  caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step,
  keeping the history tool out of the paid runner's closure;
  TrialOutcomeRecord gains the optional series_identity field.
- Slice-count plans let a registered trial spill into an ordinary lane when
  its siblings hold every long lane, so panels never share a runner.
- Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY
  map-diff; no new module loading) and repinned its hash.
- Detach and release floors now count trial shards (66 periodic trials in
  22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s.
- Coordination fixtures supply the executor's trial records.

* ci(evals): attempt-scoped artifacts, verdict-v2 PR comment, weekly pass-rate gate and one INFRA re-dispatch

- Slice, census and marathon artifacts carry -a<run_attempt>; reports
  download them per artifact (no merge), so records never overwrite and a
  re-run never replaces the first attempt's verdict.
- Planners pass --max-parallel for the capacity preflight (24/16 unchanged:
  the refreshed periodic plan needs 24 slices, the gate census 12).
- PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized
  failure block); the group_by(.name)|last recomputation is gone.
- Reports stamp series identities, upload trial-outcomes-* for history, and
  shard logs upload always (a failed trial no longer reds its runner).
- Weekly report: headline + failure block of both lanes in the issue body,
  the eval:pass-rates --gate step (fails closed without history), close the
  issue on a green run, and UC-E1: when every red is machine-classified
  INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency
  group (redispatch_of), both runs reported.

* feat(evals): planner-side whole-panel reuse and negative receipts

The planner job restores this PR's receipt store once and ships a single
filtered set with the plan: a pass or panel receipt with a same-or-newer
FAIL for its input identity is dropped, and a panel receipt ships only as
a whole PASS panel (re-verified with panelVerdict) from one run. Executors
read only that set (no per-slice cache restore or save), so every trial of
a panel sees the same receipts; a trial reuses its own record from the
panel receipt, keeping a split PASS's failed trial.

Trial identities drop the trial index (run-scoped) and bind the panel
policy. Executed shards carry their input identity; the report turns a
whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule
shard into a negative receipt, and marks a panel that mixes reused and
fresh trials INCOMPLETE. The report job merges plan, slice and report
receipts (newest per file) and saves one store per run.

Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts
whose diagnostic text drifted with program order (baseline locked, fix only).

* feat(evals): --case/--trials local diagnosis and panels in local sharded runs

bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N
independent trials of one case through the CI panel runner (trial shards,
TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict();
N defaults to the case's policy panel and CI never reads it. The local
sharded path (test:gate:sharded, test:periodic:sharded) now plans the same
trial shards and exclusions as CI and exits on execution completeness plus
panel verdicts.

* test(pty): grant an owned Create pane whose title row is cropped

The targeted batching rerun on Claude Code 2.1.284 left its first report
Write unanswered for 1,372 s and timed out: the viewport began at the
pane's relative file row and rule, with the 'Create file' title cropped
above, so the preview parser rejected the file row as foreign. That row
must now resolve to the owned path and is skipped before the unchanged
line-by-line preview match. Replay controls reject another file, another
directory and an edited preview row.

* fix(evals): tsx-safe generics in eval-flake-rank, legacy artifact names, no-retry wall docs

* test(evals): record the read-only and detector-row invariants as contracts

shared-libs-opportunity-judgment and review-design-lite are behavior
cases: their recommendation and checklist judgments may vary, but the
read-only invariant (commands, provider requests, fixture bytes, hooks,
state) and the deterministic fake-engine detector rows are contracts.
Both now go through expectContract, so any failure vetoes the panel.

* test(judges): sample the recommendation rubric as a panel; never re-ask armJudge

llm-judge-recommendation is a judge case: each fixture now draws a
3-sample judgePanel, gates reason_substance on the panel mean and the
present/commits/has_because checks on a 2-of-3 majority, thresholds
unchanged. armJudge no longer re-asks on a malformed verdict; it is a
failed sample, as the judge policy requires.

* test(evals): record a pre-turn API or CLI failure as infra

recordE2E sets failure_class 'infra' on a failed session whose runner
reports error_api, timeout_startup, error_output_stream or a non-zero CLI
exit with zero turns and no assistant event. A model refusal, a timeout
after model work, max turns, or an explicit caller pass/class keeps its
ordinary classification.

* test: pin every-record outcome counts and the twelve doc-sync callbacks

* test(eng-batching): read the report target as a field, not a spelling

The next targeted rerun (Claude Code 2.1.284) again asked eleven separate
native questions and again counted zero: its briefs named no plan and its
report declared '- **Review target (fixed):** `/abs/PLAN.md`' under
'# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one
current target field that names a PLAN.md file, whatever its list or
emphasis markup; its ledger record still supplies the cited finding and
must reproduce the brief exactly. A brief that names its plan must still
match the report title. Replays of all three captures count 9, 9 and 3;
controls reject a foreign, duplicate or missing target and an archived
title.

* fix(evals): --case list mode and name precheck; case-shard qa-callers; refresh batching and design-with-ui seeds

* chore(release): v1.91.9.0

* test: settle the post-response composer before seeding; give the TPA recorder adapter its infra helper

submitPlanSeed accepted a stale empty composer when the transcript recorded
end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the
old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E
without isPreTurnInfraFailure, so every failed case threw before recording.

* test(autoplan-dual-voice): unwrap Claude Code 2.1.284 subagent hand-back frames; accept read-only probe diagnostics; record before asserting

Census run 36626737820: the native CEO report arrived framed and indented, so
its INPUT line never matched, and the model's exact probe plus two variable
echoes was not canonical. A column-zero line inside a frame, command
substitution, backticks, redirects, assignments, CODEX_MODE echoes and output
line-count mismatches stay rejected. The failure now records before asserting.

* ci(image): keep Claude Code 2.1.251; test(ceo-mode-routing): keep HOLD's own deferrals in scope before assessing its rigor decision

2.1.284 enables per-turn effort for claude-fable-5-1: in gate census
36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time,
+32% thinking tokens) and 11 cases timed out on unchanged budgets.

HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it
Defer and the assessment then judged that scope question as the rigor
decision. The actor now answers that menu Keep and assesses the next one.

* test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon

Census 36629958451 reds:
- outside-plan-disabled-no-fallback: the model quoted the pre-existing record
  as a parenthesized field list with its exact timestamp; attribution now
  requires that exact timestamp and the record's own field values.
- plan-devex-peer-comparison-classification: the judge correctly returned
  missing but wrote a 1069-character reason, voiding the judgment; structured
  outputs cannot enforce maxLength, so the bound is stated on the field.
- plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the
  periodic lane's wall clock; it now runs weekly in the marathon lane.

* test: supply holdDeferKeepIndex to the CEO routing mocks and follow split-overflow into the marathon lane

The registered-callback fixtures mock ceo-mode-option and lacked the new
export; the split fixtures asserted the periodic tier; the registered-budget
check looked for split-overflow only in the periodic manifest.

* fix(qa): checkpoint receipts print the report link for their exploration file

qa-functional-webhook-report failed in two of three censuses because the
report linked .qa-evidence/NNN capture folders as "checkpoints" and never
linked exploration-NNN.json. The checkpoint receipt now prints
link: [checkpoint NNN](exploration-NNN.json), and the functional report
template says capture folders are not checkpoints.

* docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups

* ci(evals): name the PR-comment loop's unused fields so shellcheck passes (SC2034)

* fix(plan-ceo-review): tighten expansion pacing wording to fit the skeleton cap after the main merge

The merged skeleton measured 80,166 bytes against its unchanged 80,150 cap.
Same instructions: ask separately for each addition, in turn, with no pacing
menu; lead each proposal with the felt experience, then shape, effort and impact.

* fix(eval-pass-rates): match trial-outcome files by basename so Windows backslash paths are read

* fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures

- design-consultation Phase 1 asks one brief that confirms context and decides
  research; the confirm-only first question scored substance 2.
- document-release defines ship-owned inputs, exact steps and the JSON result,
  and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4).
- plan-design-with-ui accepts the Step 0D focus menu the same way the shared
  picker does ("focus on specific ones?").
- plan-design-review plan-mode saves in three Edits instead of one final Write.
- QA functional annotations ask for the full 40-character revision.
- Outside-disabled attribution judges quoted prior-record data by its exact
  timestamp or a dated, pre-existing-record sentence; four captured phrasings
  replay clean and current claims still fail.
- --case can select autoplan-dual-voice by its literal test name.

* test(design): revert the three-Edit plan-mode flow

A focused paid run still timed out at 300 s: the first three passes alone took
150 s of thinking. The case stays a named timeout red rather than cutting review depth.

* test: accept 'review mode = X' auto-decide declarations and parenthetical scope exclusions in the shared-libs actor

auto-decide-preserved: the product auto-decided HOLD SCOPE and said
"Decision: review mode = HOLD SCOPE"; the grammar knew only "is" and ":".
shared-libs-plan-callers: the recommended option said "(no hardening)" and the
actor read "hardening" as an expansion. Both replay the captured text, keep
negative controls, and passed focused paid runs.

* fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture

- review-army-perf-n-plus-one: the parent copied full checklists into agent
  prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
  without probing; the probe is mandatory and its first line is reported, and
  the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
  invocation or status vocabulary; the model ran it through bun and wrote
  status "blocked". The prompt states both and the validator rejects
  out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.

* docs(changelog): proof-run product fixes

* fix(ship): always run the design-lite detector probe; test(shared-libs): credit a failed first file view and deferred-reuse Skip wording

- /ship design-lite: the probe is mandatory and any non-ready first line is
  stated, matching /review (5 of 6 captured /review trials had skipped it).
- shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error,
  so the one refetch is a legitimate recovery, charged to the same budget.
- shared-libs-review-prior-coverage: the Skip option said a future review can
  "reuse it once snapshot coverage holds"; a conditional tail on the recorded
  decision is not product work. Captured-text regressions and negative controls.

* fix(ship,qa,document-release): repair proof-run regressions and fixture gaps

- ship-docsync-completion: yesterday's audit-scope result dropped the section's
  status, so /ship spliced one in; the section now opens with **Status:**.
- ship-docsync-missing-asset: a missing section or old Ship-owned mode blocks
  before launch.
- ship-docsync-late-result: the invocation record says prepare already saves
  the candidate selection (no extra Read; budget unchanged).
- qa exploratory: await the method Reads before the first probe.
- qa-callers fixture: quote the real review-log record template; allow the
  git log command plan-completion prescribes.
- qa functional observer: a receipt caught mid-link(2) is checked at stop
  instead of failing with ENOENT (reproduced from CI).
Each repaired case passed a focused paid run.

* ci(image): pin Claude Code 2.1.284, the version users run

Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to
claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284
census was mostly API latency: its SDK-only judges were 25% slower too. Nine
previously slow cases pass on 2.1.284 within unchanged budgets.

* test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors

- plan-design-review-plan-mode was registered by two files; the PTY smoke is
  now plan-design-review-plan-mode-smoke, and a registry test requires one
  owner per case in case-sharded files.
- plan-devex-finding-floor: the template's 0B narrative-confirmation question
  is classified as setup structurally instead of timing out a Haiku assessor.
- setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration
  question the test asserts; it now accepts that question and declines others.
- design-review-plugin-handoff: the fake engine cited a file absent from the
  fixture repo and index.html linked a missing styles.css.
Captured-question regressions with negative controls; each case passed a
focused paid run.

* test: PTY harness handles clipped reviews and bundled setup tabs; AUQ judge uses structured output; design-consultation carve declines optional outside voices

- ceo mode routing: a Submit review taller than the viewport, a setup tab
  bundled after the mode tab, and a clip through the mode question each hung
  or misread the run; the native answer is still verified after Submit.
- judgeRecommendation requests a 1-5 enum schema; a malformed Haiku reply had
  scored substance 0 for a 4/5 brief. Judge failures now propagate.
- carve section-loading for design-consultation declines the optional outside
  voices (a supported path) and treats DESIGN.md as the report; timeout unchanged.
The Step 0E handoff defect is not fixed (0/15 samples across four wordings,
none shipped) and is filed in TODOS.

* test: fold the design-consultation completion replay into carve-section-sharding (test-of-test ratchet)

* docs(todos): record the pre-push hook shard-order hang

* test(qa-callers): disable git auto maintenance in the fixture repo (same guard as shared-libs; from #3002)

* test(office-hours-attempt): the fake judge SDK response carries stop_reason like the real API (structured judge requires end_turn)

* fix(qa): the caller STOP line says to await the method Reads before any probe

ship-exploratory-plan-checks: the model read exploratory.md and sent a capture
in the same response, before seeing the section's own await rule.

* fix(qa): number the qa value-bar questions from 1 and say reproduced bugs already answer the first two

* fix(qa): define evidence.json where it is built, point the preparation gate at the next section, name measured command durations in the report template

Recurring qa/qa-only workflow-judge complaints in CI (clarity/actionability 3.33).

* fix(plan-eng-review,review): a disallowed question tool is not headless; report kept tests only when some were skipped

* fix(plan-eng-review): keep the headless-rule contract phrases adjacent

* fix(evals): cut path variance at its measured sources

- gstack-qa-evidence capture prints startedAt/completedAt/durationMs and, for
  --deadline captures, remainingMs; the functional report takes durations from
  them. The section clock notice asks for one clock read up front instead of one
  after every checkpoint (QA runs spent 7-14% of tool calls on date -u).
- ship plan-completion: skip the audit dispatch when discovery already found no
  plan (the dispatch-vs-skip conflict produced an optional 60-100 s subagent).
- materialize/checkpoint validation errors state the expected schema, so a
  rejected annotations file is fixable in one call instead of blocking the phase.
- session-runner counts turns from the transcript when a run times out, so
  timeouts stop reporting 'turn 0'.

* fix(evals): count timeout turns only from object transcript events

* test(qa-callers): deterministic child transport, completion-time handoff reads, compact phase report

The exploratory caller cases exist to prove the caller starts and bounds
exploratory QA. Their native adversarial reviewer (review) and plan audit
(ship plan-checks) now come from recorded child outputs instead of a live
subagent, handoff freshness reads are required before completion records
rather than every bookkeeping log, and the phase report is compact. Measured:
194-257 s per case against 208-284 s before, no subagent calls.

* test(ship-docsync): seed fault cases at their gate instead of replaying attempt 1

The post-dispatch fault cases (missing-marker, launch-failure, timeout-unsettled,
late-result, stale-before, stale-after, recovery) now start from a fixture-owned
attempt 1: the real actor prepares and dispatches it, its verbatim output is saved
once, and the invocation journal carries its pre-dispatch entry with the child
asset hashes. The model resumes at Parent processing with a trimmed read list,
inspect named as the authoritative repository observation, and recovery's
intermediate checkpoint folded into the next attempt's pre-dispatch entry.
Assertions count only parent-issued transport events and require a read of the
saved attempt-1 output; missing-asset and the legacy failure case keep the full
model-driven first attempt, and their prompts are byte-identical.

* test(ship-docsync): name the seeded read list and cap journal/report length

The first seeded stale-before run spent calls locating documentation.md (two ls
sweeps), reading through cat and re-Reading the record before Edit, and ~40 s
composing 1.5-2.2 KB entries and report. Name every seeded read path, ask for
native Read, and bound entry/report length.

* test(ship-docsync): trim the seeded parent's measured model time

Measured on the seeded runs: one read the 78 KB ship/SKILL.md, the post-child
freshness comparison spent 18-32 s of thinking over full inspect contents, and
the final response restated the report (~1.1 KB). Say the phase excerpt stands
in for ship/SKILL.md, compare hashes first and read content only for changed
paths, and end with one status line.

* feat(qa-evidence): enforce the checkpoint sequence and fill report bookkeeping in code

- capture refuses to run another probe until a checkpoint anchored on the
  latest complete capture names this capture as its next command, and every
  complete capture prints that requirement.
- materialize fills revision, runtime, cwd and learning (checkpoints whose next
  native command differs) when omitted and prints the reportLinks the report
  must include; the QA section shrinks accordingly.

* test(qa-callers): hand the caller phase its invocation-start observations and review token; fix(next-version): fetch without auto maintenance

- Every caller case receives the diff, status, log, untracked list, HEAD and an
  already-captured review start token, so the phase spends its budget on the
  contract under test instead of re-running setup reads.
- gstack-next-version's fetches pass --no-auto-maintenance. On git 2.55 a
  completed fetch forks detached maintenance in the caller's repository; the
  free suite's live smoke test ran it inside the CI checkout, and every
  shard-12 pre-push hook hang so far followed a completed smoke fetch.

* feat(deslop-shared-libs): route every Git read through bin/gstack-safe-git

The skill made the model retype a long safe-Git prefix on each call and a
dropped flag failed shared-libs-read-only. bin/gstack-safe-git applies the
fixed env + flag prefix, adds --no-ext-diff --no-textconv to log/show/diff,
allows diff only between two explicit object IDs and ls-files only in the
NUL-delimited overlay form, and refuses every other shape with one line
naming the allowed forms. The template now points at the installed helper
(host global runtime via {{SAFE_GIT}}) and drops the prose it enforces.

Fixtures resolve the helper to this checkout, the git shim records the safety
environment, and isGuardedGitRequest requires the complete prefix (env
included) for every repository read.

* test(shared-libs): tee to a discard device is not a file write

Paid shared-libs-opportunity-judgment t1 on 1213b01 failed read-only on
'... | tee /dev/null | sha256sum'. The detector flagged any tee operand while
the same devices are allowed for redirection. tee now fails only when an
operand is a real file; tee to a file, -a file and -- -a stay violations.

* fix(qa-evidence,observer): reject placeholder metadata and replay-only learning; declare the docs atomic-write target

- materialize measures revision, runtime and cwd itself and rejects supplied
  values that differ (CI run wrote revision "HEAD" and runtime "bun"), and
  refuses learning checkpoints that replay the same probe, naming the fix.
- The docs write observer treats Claude Code's atomic temp for the authorized
  doc target as transient, so a temp renamed before its per-file watch no
  longer marks the observation incomplete (ship-docsync-completion flake).
  Per-file monitoring outside declared targets stays fail-closed.

* test(qa-functional): fix mode requires only the happy scenario from the model (carried byte-identical from #3002 183b01f4..3e6074b4)

verifyQANativeRegression already reruns all eight webhook scenarios on the
repaired source, so the model-side eight-scenario requirement in fix mode
duplicated harness coverage and pushed qa-functional-webhook-fix past its
budget. qa-only still requires every scenario.

* fix(deslop-shared-libs): probe the audited repository with -C <repo>

A CI run probed safe-git from the session directory above the target repo, so
the capability probe never touched the repository and the run fell back to the
API without a local attempt. The probe (and any call from elsewhere) now names
the audited repository.

* test(qa-deadline): never attach a reader to the full-pipe fixture's stdout

The full-pipe receipt test attached a 'data' listener (flowing mode) and then
paused; on CI the reader could drain the 2 MB write before the pause, so the
receipt write never blocked and the helper exited 0 in ~126 ms. The stdout pipe
now stays unread until the assertion, which is what the test means to model.

* feat(qa): helpers answer --help, and the QA eval interfaces declare it

Approved by Garry: asking gstack-qa-evidence or gstack-qa-deadline for usage
is read-only, so both helpers print usage and exit 0 on --help (the evidence
usage now names the annotation shape), and the functional and caller command
allowlists accept exactly 'bun <path>/bin/gstack-qa-{evidence,deadline} --help'.
Two CI runs failed only on that call.

* fix(qa): after an input change, a probe is affected unless shown otherwise

CI late-input run finished in time but revalidated only the happy probe after
the locale input changed and reported the stale adverse probe green. The
revalidation step now treats any probe not shown to be unaffected as affected.

* test(shared-libs): seed the lifecycle replay's first Step 3 pass instead of replaying it

shared-libs-review-lifecycle ran ~88% of its 300 s session budget (12-run
census median 265 s, 4/24 sessions timed out). The fixture now executes pass 1's
Step 3 once with the real logger and Git: a real unused REVIEW_START, then the
diff, inventories, attributes/config/index flags, gstack-review-read output and
every file's bytes and sha256, saved to one observation. The model resumes at
Step 4 with an exact four-file first read, the observation named as the
authoritative pass-1 repository read, one post-fix verification, an explicit
pass-2 read list and a twelve-line summary. Pass 2 still runs its own --start,
diff, reads, fingerprint and stage actor before --finish.

The actor scope now states that a current settled final-pass actor result
supplies the replaced QA/adversarial prerequisites and that the no-credit
disclosure is a reporting label: one r1 session persisted completed:false
from that ambiguity.

New assertions: the final binding never uses the seeded token's start or tree,
and the observation was read; free controls finish the seeded token (binding
changed) and omit the observation read, and both fail.

* test(shared-libs): trim the resumed review replays' setup and report

Every sibling review session (revalidation, path-eligibility, index-flags,
prior-coverage) loaded qa/sections/exploratory.md and often scope.md although
its QA and native adversarial results are supplied synthetic inputs, then spent
a second request on shared-code-reuse.md and base metadata. The resumed scope
now states that the supplied results replace Step 4's QA method loading; the
revalidation contract names one first response (workflow, checklist, finding,
prerequisites, shared-code-reuse.md, base metadata) and caps the summary at
twelve lines. Receipt order, direct source reads, the checker, the question and
final persistence are unchanged.

* fix(review): define what a Step 5c Skip option says

Step 5c named "B) Skip" without saying what its description may claim. Two
CI captures (path-eligibility on 131d43be, index-flags on 4643cb85) offered a
Skip whose description added effects beyond declining: "The extraction can be
applied in a later editing review pass" and "replacing the invalidated prior
Skip". Those read as change commitments, so the no-change actor refused both.
Step 5c now says to describe Skip only as no code/index change with the Skip
recorded; adjacent lines are compacted so the review parity caps hold
unchanged. Both exact packets are kept as a free regression: still refused,
and accepted once Skip follows the rule. The actor's classifier is unchanged.

* fix(qa-evidence): every complete capture needs an evidence row; test(tpa): accept the hyphenated app-specific-password spelling

- materialize refuses when a complete capture has no evidence row and is not
  named in limits (CI cli-report omitted capture 004), naming the missing IDs.
- tpa-apple-ban's detector required 'app-specific password' with a space; the
  CI answer said 'app-specific-password path' and was otherwise correct.

* test(qa-observer): fix mode treats atomic temps of authorized src/test writes as transient

CI webhook-fix failed with 'Could not watch test/worker.regression-1.test.ts.tmp...':
Claude Code's Write renamed its temp before the per-file watch was added. The
functional eval now tells the observer its mode, and a temp whose target that
mode may write is observed through its directory watch. Report-only mode and
undeclared paths keep failing closed.

* feat(qa-evidence): refuse evidence observed on an older input snapshot than the latest capture

When native probe output declares a top-level input snapshot, materialize
compares each evidence row with the latest capture's snapshot and refuses
stale rows unless they are classified superseded, naming the captures to
rerun. ship-exploratory-late-input kept reporting a pre-change adverse probe
green after the input changed.

* test(qa-functional): point the fixture at the helper's --help instead of its source

A CI webhook-fix run spent three turns reading lib/qa-evidence.ts to learn the
interface and timed out just before materialize (agreed with #3002's owner).

* feat(qa-evidence): captures list the caller's declared-but-unrun required probes

GSTACK_QA_REQUIRED_PROBES (a JSON array of native child commands) makes every
capture print requiredRemaining; it never judges pass or fail. The functional
eval passes the webhook list from QA_WEBHOOK_REQUIRED_SCENARIOS, which the
verdict now reads too, so the nudge and the verdict share one source (agreed
with #3002's owner). CI webhook-report kept stopping with scenarios unrun.

* test(review-army): record N+1's pre-dispatch stages and scope the session to Step 4.5

review-army-perf-n-plus-one timed out in 7 of 13 CI runs on this branch (passing
245-280 s of 300). Each session spent ~95 s on setup (the full extracted SKILL,
checklist, section greps, exploratory.md, diff-scope/stats/learnings, tooling
checks), ran Step 4's core pass, a search-before-recommending WebSearch, and
wrote a 10-16 KB report (~100 s after the Red Team returned).

The fixture now stages only review/sections/review-army.md plus the performance
and red-team checklists, and hands the session the recorded detect-scope,
specialist-stats and learnings outputs and the diff. The caller passes
--performance (every CI parent already treated the prompt as that force flag
against the <50-line skip), declares the core pass, QA, adversarial review, web
research, Fix-First and persistence out of scope, and caps the report at the
selection line, the SPECIALIST REVIEW block and the Red Team result (30 lines).
The Performance specialist and the conditional Red Team are still real
foreground subagents, and the report still has to surface the N+1.

New assertion: a foreground Performance specialist dispatch precedes the Red
Team dispatch. Free controls omit the Performance dispatch or background it, and
both fail; the budget lifecycle adapter supplies the current result shape.
Touchfiles now include the .rb fixture the case reads.

* test(review-army): share the recorded Step 4.5 staging with consensus and supply its Red Team

review-army-consensus (periodic) timed out in 2 of 13 census sessions; passing
runs took 213-297 s of 300. Like N+1 it spent ~30-50 s reading the whole
extracted SKILL, checklist and every specialist file, sometimes dispatched an
unrequested Maintainability specialist, then ran a Red Team (60-70 s) and a
second merge before writing a 9-15 KB report.

The N+1 staging and scope text move into stageReviewArmySession /
reviewArmyScope / reviewArmyChecklists (the N+1 prompt renders byte-identical).
Consensus now records its detect-scope, stats, learnings and diff, stages the
Review Army section with the security and testing checklists, forces
--security --testing, and caps the report like N+1. Its Red Team is outside
the multi-specialist contract, so the fixture supplies a labeled synthetic
NO FINDINGS result instead of a dispatch. The existing SQL-finding and
browser-error assertions are unchanged; the lifecycle adapter's spawnSync now
returns the git output the staging reads.

* docs(changelog): v1.91.10.0 records the flake census and its repairs

* test(strict-output): give the spool-prefix child time to finish before the pending stream times out

windows-free-tests failed on 9a7a7e54: the 150 ms shared deadline raced Bun
startup on Windows, so the child was killed mid-write and the spool held a
partial payload. Only the never-released extra stream should time out; the
child now has 3 s.

* fix(qa-evidence): accept a single limits string; test(qa-callers): read the handoff first when a probe snapshot changes

CI late-input spent a turn rewriting limits as an array after materialize
refused a string, and a ten-read sweep hunting for the changed input before it
read reports/HANDOFF.md, then timed out at 300 s.

* test(autoplan-dual-voice): unwrap the framed native report before Claude Code 2.1.284's agentId/usage trailer

* test(section-loading): credit a Bash print that contains every line of the carved section

* test(auto-decide): ask for the selected mode in the skill's mode handoff line, not a separate public decision

* test(plan-ceo floor): scope preservation approves no premise, approach or remedy

* test(autoplan-dual-voice): the fixture declares that delivered bash blocks run alone, diagnostics separately

* test(coverage-audit): a fenced plain-word caption in a successful && read chain is display only

Census 36776104571 plan-eng capture read both owned files with cat -n in one
successful && chain; the caption 'echo "=== git diff main --stat ==="' fell
outside the two-token caption grammar, so both reads lost credit. Accept a fenced
caption of plain words; unfenced command strings, expansions, redirection,
-e escapes and ; / || tails stay rejected.

* test(office-hours): a fork whose outer options are the seeded shapes is the Phase 4 question

Census trials 1-2 captured complete Phase 4 forks (A) Server-side B) Client-side
C) Hybrid, recommendation with because) whose prose used none of the vocabulary
words. Accept two seeded shapes as outer options as Phase 4 specificity; the
earlier-phase, nested, fenced and single-shape controls still fail.

* fix(review): design-lite rows keep the detector's [rule-id]; the e2e detector rows point at the diff

The output template had no rule-id slot, so rows merged with checklist items
dropped the detector id (census t2, local t1). Rows now carry [rule-id]. The
fake engine's sample rows named a foreign fixture path at line 0; the e2e remaps
them to landing.html/styles.css so trials stop spending turns reconciling it.

* test(shared-libs): the plan actor reads scheduler parity and unchanged-scope lists

Census 36776104571's question preserved the contract ('behaving exactly like the
scheduler', 'scheduler parity holds by construction') and excluded work with
'Existing copies and helper hardening stay unchanged'. Accept exactly/parity as
preservation (negated forms refuse) and a bare noun list that stays unchanged as
an exclusion for the expansion scan only; verb-led clauses still refuse.

* fix(qa-only,qa): name the exploratory read point and finalization order; judge qa with its browser assets

qa-only judges cited 'next section' pointing at the wrong heading, an exploratory
trigger that contradicted its read point, clock ownership in mixed runs and the
unstated order of exploratory section 4 vs reporting. The qa judge penalized the
absent qa-report-template and issue-taxonomy that qa-patterns loads; with them
in, it found issue-taxonomy's dangling 'rule 13' (the consent rule is browser rule 3).

* test(ship-docsync): seeded attempt 1 counts toward the limit; transport counts ignore calls that never reached the state file

- CI launch-failure retried after the seeded attempt 1 as if that attempt were
  the fixture's; the seeded prompt now says attempt 1 is this invocation's and
  a further attempt needs what Blocked recovery requires.
- A late-result run typo'd the state path once (ENOENT, the actor never ran),
  then repeated the call correctly; the per-action count compared both calls
  with one actor event. Only calls naming the real state file are counted.

* fix(plan-eng-review): show the accepted dedicated read form for coverage-diagram sources

CI plan-eng-coverage-audit mixed package/config and git diff into the source
read; the review variant, whose prompt shows the && display form, does not.
The plan trace step now shows it too, within the unchanged size cap.

* test(sync-gbrain-readiness): a negation earlier in the claim clause is not a search/write readiness claim

The census unknown actor wrote 'nothing about read, search, or write capability
is confirmed either way' after a YELLOW/WARN verdict. The claim window started
at 'write', so the leading 'nothing' was outside it. Check the clause subject for
nothing/neither/none/no; keep the original in-claim negations. Replay of the
captured output passes; positive controls still flag an unnegated claim.

* fix(office-hours): a forcing question's recommendation takes the position the founder's words support

auq-matrix office-hours asked D1 Demand as options about the founder's own
evidence and, with no rule for that shape, recommended 'answer whichever is
TRUE — A is marked recommended only because it is the strongest position'
(substance 2). Say what such a recommendation is: the option the founder's own
words support, why it matters for the next step, and what would change it.

* fix(plan-ceo-review): name the mode preference command and the exact handoff line

auto-decide-preserved at 6fcb0981: the model never ran the preference check,
read 'check ... through the preamble' as already done, auto-selected 'per your
preference setting', and wrote 'Selected mode: HOLD SCOPE, auto-decided from
your tuned preference' instead of the AUTO_DECIDE handoff line. At 9a7a7e54 it
ran the check but wrote 'Decision: HOLD SCOPE is the review mode for ...'.
Neither matched the handoff template the observer recognizes. Name
gstack-question-preference --check at the point of use and say the handoff
begins with the exact matching line. Collapse the audit block's comment
padding to stay within the unchanged 80150-byte skeleton cap.

* test(section-loading): record the CEO capture's report and transcript

The 6fcb0981 census failed hasStaleFillRaceFinding (line 98), but the case
records nothing beyond junit, so the report the detector judged is gone.
Return the SkillTestResult from captureSectionReads and record it, with the
full saved report, through the eval collector on pass and fail.

* test(design): plan-mode names its read list and caps its additions and summary

At 6fcb0981 plan-design-review-plan-mode timed out at 300 s (9 turns): 22 cat/sed
chunk reads (~50 s), then a 28 KB plan Write (~150 s), before the read-back
finished. The 9a7a7e54 pass took 240 s with a 24.6 KB Write. Read SKILL.md,
review-sections.md and plan.md natively in one response, keep additions under
14,000 characters and the summary within ten lines. Budgets unchanged.

* test(plan-mode-no-op): require prose evidence before a waiting verdict ends eng/design runs (carried byte-identical from #3002)

With the prose fallback forced, the gate renders as a lettered menu; a judge
'waiting' verdict on a spinner-only frame ended the run as 'asked' before the
menu rendered, so the scope-gate check failed on unchanged behavior.

* feat(qa-evidence): materialize computes the phase verdict; callers must report it

Approved by Garry: the helper, not the model, decides whether evidence can
pass. materialize writes verdict {status, open} into evidence.json and prints
it: fail or blocked from row classifications, inconclusive while any row is
superseded, a complete capture is withheld, a declared required probe is
unrun or there is no evidence, else pass. The caller fixture requires
receipt.status to equal that verdict. CI late-input kept reporting pass with a
superseded happy probe.

* test(qa-callers): compare the receipt with the helper verdict only when evidence.json was materialized

The producer free tests run captures without materialize; evidence.json is
optional for callers, so its absence is not a verdict mismatch.

* test(llm-judge): run the ship workflow judge at medium effort so its panel fits JUDGE_MS

claude-fable-5-1 accepts only adaptive thinking (thinking.type.enabled with
budget_tokens returns 400), so effort is the available thinking control.
Measured on the exact ship judge request (105,301 input tokens):

- default effort, 18 samples: thinking 5,086-10,881 tokens, 75.9-144.7 s;
  3 of 18 passed the 120 s deadline (about 42% of 3-sample panels).
- medium effort, 18 samples: thinking 2,749-5,762, output at most 6,144
  tokens, 43.1-77.9 s; scores 4/4/4 in 16 of 18 (clarity 3 in two), versus
  14 of 18 at default.

callJudge gains an effort option sent as output_config.effort; only the ship
judge sets it. Rubric, floors, panel size, deadline, model and max_tokens are
unchanged. The cache identity records effort.

* test(llm-judge): ask frontier workflow judges for 120-word reasoning under the unchanged 150-word check

Told "under 150 words", the ship judge's reasoning landed at 130-156 words
(3 of 18 probe samples at 152-156), so the structured-response check failed
about one panel in three independent of effort. The prompt's frontier block
and the response schema description now say under 120 words; the validator
still rejects 150 words or more. The changed prompt bytes reach only the two
frontier judges: ship/SKILL.md workflow (prompt and schema) and
review/SKILL.md workflow (prompt).

* test(llm-judge): type the stream transport mock call

* test(plan-ceo floor): the request answers only the questions it names

PR lane 36794871032 (head 20d6e98f): the CEO floor ran 608 s without a
question. Its Step 0 recorded the premise gap and approach choice as
unresolved ledger rows, then said "this session supplies all answers up
front, so no decision brief was dispatched" and wrote Sections 1-11.
2734e203 stopped scope preservation from approving the premise; this time
the actor block (declined setup, recall, outside voices, HOLD SCOPE) and the
fixture's "complete user request is available from the start" were read
as pre-answering every review question. The CEO actor now states that the
request answers only the routing, recall, outside-reviewer and review-mode
questions it names.

* test(plan-devex floor): a 'Partly wrong' 0B answer is the narrative confirmation

PR lane 36794871032: the DX floor asked its D1 narrative confirmation
(Accurate, proceed / Partly wrong, correct it / Way off, actual is...). The
deterministic setup rule accepted only 'Some ... wrong', so the question went
to the LLM assessor, which hit its 30 s spawnSync cap (ETIMEDOUT) and ended
the case as assessment_error at 141 s, the same failure as census
36641820398. The rule now accepts 'partly' beside 'some'; the captured
question is a free regression and the remedy-option controls still go to
the assessor.

* test(design-review plugin handoff): quoted report text is not an install command

PR lane 36794871032: every behavioral check passed except noInstallOrOverride,
which matched "no `npx impeccable`" inside the quoted heredoc that wrote
detector-output.md. Nothing was installed or downloaded. The check now drops
quoted-delimiter heredoc bodies (literal data) before matching; unquoted
bodies, which can expand $(...), and unterminated bodies stay checked. Free
controls cover the captured write, bare npx, an IMPECCABLE_BIN override, an
unquoted $(npx ...), npx after the delimiter and an unterminated body.

* test(review-army delivery audit): stage only the plan-completion section and record its git reads

PR lane 36794871032: the case timed out at its 120 s budget after 7 turns
(previous lane passed in 45 s). The session read the 46 KB extracted SKILL in
three passes (cat to persisted output, grep, sed), ran its own git reads,
wrote a 74-line report, then inspected and ran gstack-learnings-log and
rewrote the report's Learnings section. As in the Step 4.5 cases
(17ee2e54/2bd4651c), the fixture now stages only
review/sections/plan-completion.md, hands the session the recorded
git log and diff, declares the HIGH-impact question, its Scope Check,
learnings logging and later steps outside the capture, and caps the report
at the audit block and its DISCREPANCY entries (30 lines). The NOT DONE and
email assertions are unchanged.

* feat(qa-evidence): one capture call records the causal note for the previous capture

capture R NNN [--public] (--deadline D|--timeout-ms MS) --after PREV --hypothesis 'TEXT' -- CMD
publishes exploration-NNN.json {observationCapture, observationArgv, observed, hypothesis,
nextCapture, nextArgv} before running CMD, refusing unless PREV is the latest complete capture.
The receipt carries checkpoint/checkpointSha256; validators bind the note to the transcript's
capture calls by capture ID and receipt hash instead of exact command strings. The separate
checkpoint command and the capture guard keep working; materialize learning accepts both note
shapes and still rejects same-probe replays. Prose and eval fixture prompts teach the merged form.

* fix(qa-evidence): a superseded row stops holding the verdict open once its probe is rerun on current inputs

materialize requires an old-snapshot row to be classified superseded, and its verdict kept every
superseded row open, so rerunning the probe (what its own error tells the model to do) could never
reach pass; late-input reran 3 and 9 on the new snapshot and still got inconclusive. A superseded
row now closes only when a non-superseded row with the same captured argv observed the current
snapshot. Re-materializing an already-published evidence.json names the cause instead of failing
generically.

* test(plan-eng batching): count saved decisions whose label drops the (recommended) marker or whose report is titled 'Eng Review Report — <plan>'

* fix(qa): browser-only runs skip annotations/materialize; only Q captures can anchor evidence rows

* test(design): plan-mode length is a drafting target, not a check to measure and trim

* test(llm-judge): structured output for doc, outcome and posture judges so reasoning quotes cannot break JSON

* test(ship-docsync): steer skill file reads to Read; large cat output becomes an unpageable preview

* docs(changelog): browser-only QA evidence and structured judge output

* test(qa-only cleanup): refusal scenarios get a 1 s budget and an absolute worker deadline; 300 ms starved under parallel load

* fix(office-hours, design-consultation): ask the goal question and read the mode section first; ask the memorable-thing question on its own

* test(outside-disabled): a record named by the retained record's own clock and then disowned owns its completed status

* test(context-skills): install gstack-paths in the fixture bin; without it the model guessed the checkpoint root

* test(ceo mode routing): SCOPE EXPANSION posture credits plural 'expansions'

* test(ship-docsync): name the unmet atomic-replacement check on a forbidden temp-file write

* fix(qa): browser-only runs materialize an empty evidence list with checkpoints in limits, matching /qa-only

* test(qa callers): an accepted review-log record may cite checkpoints as finding evidence

* fix(plan-eng-review): state that a disallowed question tool never qualifies as headless before the headless action

* merge follow-up: re-record paid CLI parity for #2999's flags; trim merged review, qa-only and plan-eng wording toward the size caps

* test(golden): refresh codex/factory ship goldens for the trimmed caller QA wording

* test(coverage-audit fixture): disable git auto maintenance so cleanup is not racing a detached git writer

* test(parity): raise review, qa and plan-eng caps to the measured merged size of #2999 and #3002 (each fit alone), documented per cap

* fix(qa-evidence): materialize rejects an unrecognized classification before publishing, so the one-shot verdict cannot be locked inconclusive by a descriptive label
This commit is contained in:
Garry Tan authored and GitHub committed 2026-10-01 13:55:16 -07:00
1 parent df89475b17
commit 7fca42ad8b
340 files changed
+34342 -6789

No files matched your search

+33 -16
View File
@@ -53,7 +53,8 @@ export function scoreAuqFormat(text: string): { present: number; total: number;
* whether the ORIGINAL used the literal "because" — a soft style signal, since
* the format spec prefers it and the voice rule forbids the em-dash form.
*
* This does NOT touch judgeRecommendation or its pinned fixtures.
* This does NOT touch judgeRecommendation or its pinned fixtures. A judge
* failure propagates with its cause; it is never reported as substance 0.
*/
export async function gradeAuqRecommendation(
text: string,
@@ -75,12 +76,8 @@ export async function gradeAuqRecommendation(
}
}
try {
const r = await judgeRecommendation(graded);
return { substance: r.reason_substance, present: r.present, hadLiteralBecause, reason: r.reason_text };
} catch {
return { substance: 0, present: !!recLine, hadLiteralBecause, reason: '' };
}
const r = await judgeRecommendation(graded);
return { substance: r.reason_substance, present: r.present, hadLiteralBecause, reason: r.reason_text };
}
/**
@@ -212,6 +209,29 @@ export function hasDisabledOutsideReview(output: string): boolean {
return false;
}
/**
* Sections a capture loaded: a Read of the section file, or a Bash print of it
* (cat/sed ranges, as in run 36776104571) whose outputs together contain every
* line of the section as it stood before the run. A command without that
* printed content, such as head or grep, is not a read.
*/
export function detectSectionReads(toolCalls: SkillTestResult['toolCalls'], sections: Map<string, string>): Set<string> {
const readSections = new Set<string>();
for (const c of toolCalls) {
if (c.tool !== 'Read') continue;
const fp = String(c.input?.file_path ?? '');
const m = fp.match(/(?:^|[\\/])sections[\\/]([A-Za-z0-9._-]+\.md)(?=$|[?#])/);
if (m) readSections.add(m[1]);
}
for (const [name, content] of sections) {
const lines = content.split('\n').map(line => line.trimEnd()).filter(Boolean);
const printed = new Set(toolCalls.filter(c => c.tool === 'Bash' && String(c.input?.command ?? '').includes(`sections/${name}`))
.flatMap(c => c.output.split('\n').map(line => line.trimEnd())));
if (lines.length && lines.every(line => printed.has(line))) readSections.add(name);
}
return readSections;
}
export async function captureSectionReads(opts: {
planDir: string;
skillName: string;
@@ -233,7 +253,7 @@ export async function captureSectionReads(opts: {
nativeReviewOnly?: boolean;
}): Promise<{ readSections: Set<string>; reportProduced: boolean; reportWritten: boolean;
exitReason: SkillTestResult['exitReason']; toolCalls: SkillTestResult['toolCalls'];
transcript: SkillTestResult['transcript']; output: string }> {
transcript: SkillTestResult['transcript']; output: string; result: SkillTestResult }> {
const outFile = path.join(opts.planDir, opts.reportFile ?? 'REPORT.md');
const timeout = opts.timeout ?? 300_000;
const fullPlanReview = opts.skillName === 'plan-ceo-review' || opts.skillName === 'plan-eng-review';
@@ -269,6 +289,9 @@ export async function captureSectionReads(opts: {
};
const beforeReport = readReport();
const skillPath = path.join(opts.planDir, opts.skillName, 'SKILL.md');
const sectionsDir = path.join(opts.planDir, opts.skillName, 'sections');
const sections = new Map(fs.existsSync(sectionsDir) ? fs.readdirSync(sectionsDir)
.filter(name => name.endsWith('.md')).map(name => [name, fs.readFileSync(path.join(sectionsDir, name), 'utf-8')]) : []);
// Outside-review dispatch has separate behavioral coverage. Native-only
// captures use the real supported control in state owned by this call;
// never mutate the operator's or another capture's gstack configuration.
@@ -323,13 +346,7 @@ ${fullPlanReview ? `- Save the evolving plan and review outputs to ${outFile} wi
if (stateDir) fs.rmSync(stateDir, { recursive: true, force: true });
}
const readSections = new Set<string>();
for (const c of result.toolCalls) {
if (c.tool !== 'Read') continue;
const fp = String(c.input?.file_path ?? '');
const m = fp.match(/(?:^|[\\/])sections[\\/]([A-Za-z0-9._-]+\.md)(?=$|[?#])/);
if (m) readSections.add(m[1]);
}
const readSections = detectSectionReads(result.toolCalls, sections);
const afterReport = readReport();
const reportWritten = afterReport !== undefined
@@ -341,7 +358,7 @@ ${fullPlanReview ? `- Save the evolving plan and review outputs to ${outFile} wi
// Keep successful terminal-output captures, but a draft left by a failed run
// must never satisfy callers that use reportProduced as their completion gate.
return { readSections, reportProduced, reportWritten, exitReason: result.exitReason, toolCalls: result.toolCalls, transcript: result.transcript, output };
return { readSections, reportProduced, reportWritten, exitReason: result.exitReason, toolCalls: result.toolCalls, transcript: result.transcript, output, result };
}
/** A completed CEO review needs its artifact and every summary outcome. */
+44 -6
View File
@@ -8,6 +8,18 @@ import { claudeOutsideExecutions } from './outside-voice-evidence';
const sha = (value: string) => createHash('sha256').update(value).digest('hex');
const text = (content: unknown): string => typeof content === 'string' ? content : Array.isArray(content)
? content.flatMap(block => block?.type === 'text' && typeof block.text === 'string' ? [block.text] : []).join('\n') : '';
// Claude Code 2.1.284 frames a subagent report with one header line and indents
// every report line by two spaces. Only a fully indented report is unwrapped;
// a column-zero line inside the frame stays framed and earns no credit.
// The same release appends its own column-zero agentId/usage trailer after the
// indented report (run 36776104571); only that exact final trailer is removed.
const trailer = /\nagentId: ([0-9a-f]{8,}) \(use SendMessage with to: '\1', summary: '<5-10 word recap>' to continue this agent\)\n<usage>(?:[a-z_]+: \d+\n)*[a-z_]+: \d+<\/usage>$/;
const report = (content: string): string => {
const header = /^\[Subagent hand-back\] [^\n]*The report follows:\n/.exec(content);
if (!header) return content;
const lines = content.slice(header[0].length).replace(trailer, '').split('\n');
return lines.every(line => line === '' || line.startsWith(' ')) ? lines.map(line => line.slice(2)).join('\n') : content;
};
const object = (value: unknown): value is Record<string, any> => value !== null && typeof value === 'object' && !Array.isArray(value);
const parent = (event: any) => event?.parent_tool_use_id == null && event?.agentId == null && (event?.isSidechain == null || event?.isSidechain === false);
// These are delivered executable blocks, not a shell interpreter. Only blank
@@ -160,13 +172,35 @@ export function autoplanDualVoiceEvidence(transcript: unknown[], options: Autopl
}
return seen.size > 0;
};
// The exact probe may be followed by read-only diagnostics: double-quoted
// echoes of literal text and plain variables, one output line each, never
// naming CODEX_MODE. Their lines are the only output allowed after the mode.
const diagnosticEchoes = (command: string): number | null => {
const actual = code(command), contract = code(options.commands.probe);
if (!actual.startsWith(contract)) return null;
const suffix = actual.slice(contract.length);
if (suffix.includes('CODEX_MODE') ||
!/^(?:(?:;[ \t]*|\n)echo "(?:[^"$`\\\n]|\$[A-Za-z_][A-Za-z0-9_]*|\$\{[A-Za-z_][A-Za-z0-9_]*(?::-[A-Za-z0-9_ .,:=\/-]*)?\})*")+$/.test(suffix)) return null;
return suffix.match(/(?:;|\n)[ \t]*echo "/g)!.length;
};
let probeResult = 'no Bash call matched the canonical probe block';
let nonCanonicalProbes = 0;
for (const call of calls.values()) {
if (call.name !== 'Bash' || typeof call.input.command !== 'string' || !canonical(call.input.command, options.commands.probe)) continue;
if (call.name !== 'Bash' || typeof call.input.command !== 'string') continue;
const echoes = canonical(call.input.command, options.commands.probe) ? 0 : diagnosticEchoes(call.input.command);
if (echoes === null) {
if (call.input.command.includes('CODEX_MODE')) nonCanonicalProbes++;
continue;
}
result.probeToolUseId = call.id; delete result.probeMode;
if (!call.result || call.result.error) continue;
if (!call.result || call.result.error) { probeResult = call.result ? 'probe result is an error' : 'probe has no result'; continue; }
const modes = [...call.result.content.matchAll(/^CODEX_MODE: ([a-z_]+)\r?$/gm)];
if (modes.length !== 1 || !call.result.content.trimEnd().endsWith(modes[0]![0])) continue;
result.probeToolUseId = call.id; result.probeMode = modes[0]![1];
const trailing = modes.length === 1 ? call.result.content.slice(modes[0]!.index! + modes[0]![0].length).trimEnd() : '';
if (modes.length !== 1 || (trailing ? trailing.replace(/^\r?\n/, '').split(/\r?\n/).length : 0) !== echoes) {
probeResult = `probe output has ${modes.length} CODEX_MODE line(s) and ${modes.length === 1 ? 'does not end with it' : 'needs exactly one'}`;
continue;
}
result.probeToolUseId = call.id; result.probeMode = modes[0]![1]; probeResult = 'mode recorded';
}
const native: Array<{ call: Call; snapshot: any; content: string }> = [];
for (const call of calls.values()) {
@@ -182,7 +216,7 @@ export function autoplanDualVoiceEvidence(transcript: unknown[], options: Autopl
if (sha(content) !== snapshot.sha256 ||
!readFileSync(nativePath, 'utf8').includes(content)) continue;
if (!/^Async agent launched successfully\./.test(call.result.content) &&
!new RegExp('^INPUT: ceo ' + snapshot.sha256 + '(?:\\r?\\n|$)').test(call.result.content.trimStart())) continue;
!new RegExp('^INPUT: ceo ' + snapshot.sha256 + '(?:\\r?\\n|$)').test(report(call.result.content).trimStart())) continue;
native.push({ call, snapshot, content });
} catch { /* Unowned, spec-only, foreign-phase and forged snapshots earn no voice credit. */ }
}
@@ -226,6 +260,10 @@ export function autoplanDualVoiceEvidence(transcript: unknown[], options: Autopl
}
result.codexUnavailable ||= result.claudeVoiceFired && ['not_installed', 'not_authed', 'broken_install', 'model_unusable'].includes(result.probeMode ?? '');
if (!result.claudeVoiceFired) result.reasons.push('No acknowledged current CEO phase dispatch');
if (!result.codexVoiceFired && !result.codexUnavailable) result.reasons.push('No acknowledged outside execution or actual unavailable probe result');
if (!result.codexVoiceFired && !result.codexUnavailable) {
result.reasons.push('No acknowledged outside execution or actual unavailable probe result');
result.reasons.push(`probeToolUseId=${result.probeToolUseId ?? 'none'} probeMode=${result.probeMode ?? 'none'} ` +
`canonicalMatch=${result.probeToolUseId ? 'yes' : 'no'} (${probeResult}; ${nonCanonicalProbes} non-canonical Bash call(s) mention CODEX_MODE)`);
}
return result;
}
+5 -5
View File
@@ -221,7 +221,7 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
// hardening: host-anchored mode signal, precedence, passing-mention
// guards) and the plan-mode preamble reword land the union at 1.092.
maxSizeRatio: 1.174, // + clarity rules for saved decisions/setup gates + the Aside probe's failure reason; measured 1.1504. + test value bar and Tests to Retire in the lazy Test review section (~2.6KB); measured 1.168 + W1 guarded state-root resolution (`eval gstack-paths; : "${GSTACK_STATE_ROOT:?…}"`) in the Context Recovery preamble, the eureka log and each state-writing bash block; measured 1.173 (2026-09-30).
maxSizeRatio: 1.175, // + clarity rules for saved decisions/setup gates + the Aside probe's failure reason; measured 1.1504. + test value bar and Tests to Retire in the lazy Test review section (~2.6KB); measured 1.168 + W1 guarded state-root resolution (`eval gstack-paths; : "${GSTACK_STATE_ROOT:?…}"`) in the Context Recovery preamble, the eureka log and each state-writing bash block; measured 1.173 (2026-09-30). + v1.91.12.0 merge of #2999 (headless rule: a disallowed question tool never qualifies) with #3002; measured 1.1741 (2026-10-01).
},
'plan-design-review': {
skill: 'plan-design-review',
@@ -387,7 +387,7 @@ do not launch the downstream skill or open a browser.`,
expectedSections: ['proposal-and-preview.md'],
requiredReads: ['proposal-and-preview.md'],
scenario:
'The user gave product context (a B2B analytics dashboard for ops teams) and declined the research phase. Skip browser/design tool setup. Proceed to build the complete design-system proposal, then write DESIGN.md. Produce the proposal and the DESIGN.md content.',
'The user gave product context (a B2B analytics dashboard for ops teams), declined the research phase and declined the optional outside design voices. Skip browser/design tool setup. Proceed to build the complete design-system proposal, then write DESIGN.md and its CLAUDE.md guidance.',
staticInvariants: {
mustStayInSkeleton: ['## Phase 0: Pre-checks', '## Phase 1: Product Context', '## Phase 2: Research'],
mustMoveToSection: ['## Phase 3: The Complete Proposal', '## Phase 6: Write DESIGN.md'],
@@ -481,10 +481,10 @@ do not launch the downstream skill or open a browser.`,
gateAfterStop: undefined, // operational multi-STOP skill, like ship
},
behavioral: 'plan',
maxSkeletonBytes: 74_600, // Shared-code identity/skip/action rules + critical-severity validation; measured 74,493 (2026-09-17).
maxSkeletonBytes: 74_881, // Shared-code identity/skip/action rules + critical-severity validation; measured 74,493 (2026-09-17). + v1.91.12.0 merge of #2999 (review clarity repairs: await reads, research alongside dispatch, /review deadline and setup authority, findings sources) with #3002 (guarded state-root lines, plan-check checkpoints); each fit alone; measured 74,881 (2026-10-01).
minUnionBytes: 89_000, // Phase 4 wave 1; measured union 93,357
mustContain: ['confidence', 'P1', 'P2', 'Review Army', 'adversarial'],
maxSizeRatio: 1.18, // Shared-code feature + critical-severity validation: 128,042 union bytes / 108,523 baseline = 1.1799; preserves content floors.
maxSizeRatio: 1.185, // Shared-code feature + critical-severity validation: 128,042 union bytes / 108,523 baseline = 1.1799; preserves content floors. + v1.91.12.0 merge of #2999 (above, plus plan-completion fallback intent and specialist checklist-by-path) with #3002; measured 1.1843 (2026-10-01).
},
codex: {
skill: 'codex',
@@ -667,7 +667,7 @@ do not launch the downstream skill or open a browser.`,
},
behavioral: 'prompt',
maxSkeletonBytes: 63_500, // + v2.0 {{ASIDE_SETUP}}/{{BROWSE_FALLBACK}} (replaces the browse setup block); measured 61_253
maxSizeRatio: 1.102, // + v1.81 Aside contract + gstack-browser fallback block (1.080 on v1.91.7.0) + the shared test value bar at 8a.5 ({{TEST_VALUE_BAR:qa}}); measured 1.094 + W1 guarded state-root resolution (`eval gstack-paths; : "${GSTACK_STATE_ROOT:?…}"`) in the Context Recovery preamble, the eureka log and each state-writing bash block; measured 1.101 (2026-09-30)
maxSizeRatio: 1.103, // + v1.81 Aside contract + gstack-browser fallback block (1.080 on v1.91.7.0) + the shared test value bar at 8a.5 ({{TEST_VALUE_BAR:qa}}); measured 1.094 + W1 guarded state-root resolution (`eval gstack-paths; : "${GSTACK_STATE_ROOT:?…}"`) in the Context Recovery preamble, the eureka log and each state-writing bash block; measured 1.101 (2026-09-30) + v1.91.12.0 merge of #2999 (await scope/method Reads, capture --after checkpoints, browser-only empty evidence list) with #3002; measured 1.1028 (2026-10-01).
minUnionBytes: 69_500, // measured union 70,385
// 'aside repl' pins the Aside contract; '$B goto' pins the fallback block in the always-loaded skeleton.
mustContain: ['bug', 'aside repl', '$B goto', 'fix', 'Health Score Rubric', 'regression'],
+6 -2
View File
@@ -108,11 +108,15 @@ export function registerCarveSectionCase(skill: string): void {
? '- Proceed directly with the requested engineering review; skip the optional /office-hours prerequisite. You represent the plan author, whose scope and proposed steps are in PLAN.md. At each decision, choose the complete alternative that preserves those requirements and existing contracts; choose the recommended option only among alternatives within that scope. Do not authorize optional scope, extra public input guarantees, arbitrary size limits, or optional proof projects. Decline work explicitly listed out of scope, including creating TODOs for it. Record the decision and its actual authority as the skill requires, then continue without asking a human. A demonstrated incompatibility or missing required proof still requires resolution; do not hide it or claim approval when no offered alternative meets these constraints.'
: undefined,
// Both plan reviews persist their required report in the reviewed plan.
reportFile: ['plan-devex-review', 'plan-eng-review'].includes(guard.skill) ? 'PLAN.md' : undefined,
// design-consultation's final output is DESIGN.md itself; a second
// REPORT.md only duplicated the proposal (census 36641820398 timeout).
reportFile: ['plan-devex-review', 'plan-eng-review'].includes(guard.skill) ? 'PLAN.md'
: guard.skill === 'design-consultation' ? 'DESIGN.md' : undefined,
// This scenario produces an HTML implementation, whose complete
// document need not contain any of the prose report keywords.
reportMarker: guard.skill === 'design-html'
? /<!doctype\s+html\s*>\s*<html\b[^>]*>[\s\S]*?<head\b[^>]*>[\s\S]*?<\/head\s*>[\s\S]*?<body\b[^>]*>[\s\S]*?<\/body\s*>\s*<\/html\s*>/i
: guard.skill === 'design-consultation' ? /^# gstack: design-md-format=spec$/m
: /report|review|summary|design doc|handoff/i,
testName: `${guard.skill} section-loading`,
runId,
@@ -125,7 +129,7 @@ export function registerCarveSectionCase(skill: string): void {
});
// Require the HTML artifact itself; a terminal-only claim is insufficient.
// captureSectionReads already requires a successful native completion.
const reportProduced = completionMarked && (guard.skill !== 'design-html' || reportWritten);
const reportProduced = completionMarked && (!['design-html', 'design-consultation'].includes(guard.skill) || reportWritten);
const missing = guard.requiredReads.filter((s) => !readSections.has(s));
// Named failure output (codex #2): skill + expected + observed.
+2 -1
View File
@@ -64,7 +64,8 @@ export function buildCeoHoldPostureReview(input: CeoHoldPostureReviewInput): Pla
for (const call of [mode, decision]) {
const context = /Project\/branch\/task:([^\n]*)/i.exec(call.questions[0]!.question)?.[1] ?? '';
const plans = [...new Set(context.match(/(?<![\w.:/\\-])[\w.:/\\-]+\.md(?![\w.:/\\-])/gi) ?? [])];
if (plans.length !== 1 || (plans[0] !== name && plans[0] !== source.path)) fail('native source context differs from original plan');
// Naming no plan leaves the owned source Read below as the binding; naming another or several plans does not.
if (plans.length > 1 || (plans.length === 1 && plans[0] !== name && plans[0] !== source.path)) fail('native source context differs from original plan');
}
const decisionContext = /Project\/branch\/task:([^\n]*)/i.exec(decision.questions[0]!.question)?.[1] ?? '';
if (!/\bHOLD SCOPE\b/.test(decisionContext) ||
+116 -13
View File
@@ -146,10 +146,81 @@ function hasNativePostureProse(text: string, posture: RegExp): boolean {
return hasPostAnswerCeoPosture(`● ${prose}`, posture);
}
const CLIPPED_PREFIX_MIN = 120;
/**
* A review taller than the viewport can clip its heading and earlier questions
* before they ever render, and it truncates a long question with "…". Authenticate
* the visible tail from the Submit prompt backwards: every visible answer is an
* offered option, each question below the clip matches its native text (or a long
* native prefix before "…"), the mode question's target answer is visible, and only
* the topmost segment may be cut off above the viewport; a cut mode question must
* still show a long native tail. The native answer is verified again after Submit.
*/
function clippedReviewMatches(visible: string, selected: NativePlanQuestionCall,
modeQuestion: NativePlanQuestionCall['questions'][number], targetMode: CeoMode): boolean {
const compact = (text: string) => text.replace(/\s+/g, '');
let body = compact(visible.replace(/^[ \t]*[│┃] ?/gm, '').replace(/^[ \t]*[●⏺] ?/gm, ''));
if (!body.endsWith(BARLESS_SUBMIT_END) || /[←☐☒]/.test(body)) return false;
body = body.slice(0, -BARLESS_SUBMIT_END.length);
const modeIndex = selected.questions.indexOf(modeQuestion);
for (let i = selected.questions.length - 1; i >= 0; i--) {
const question = selected.questions[i]!;
const answers = (i === modeIndex
? question.options.filter(o => modeTitle(o.label) === targetMode.replace(/\s+/g, ''))
: question.options).map(o => `→${compact(o.label)}`).filter(answer => body.endsWith(answer));
if (answers.length !== 1) return false;
body = body.slice(0, -answers[0]!.length);
const text = compact(question.question);
let shown = 0;
if (body.endsWith(text)) shown = text.length;
else if (body.endsWith('…')) {
for (let length = text.length - 1; length >= CLIPPED_PREFIX_MIN && !shown; length--) {
if (body.slice(0, -1).endsWith(text.slice(0, length))) shown = length + 1;
}
}
if (!shown) {
if (i > modeIndex || (i === modeIndex && body.replace(/…$/, '').length < CLIPPED_PREFIX_MIN)) return false;
return body.endsWith('…') ? text.includes(body.slice(0, -1)) : text.endsWith(body);
}
body = body.slice(0, -shown);
if (!body) return i <= modeIndex;
}
return 'Reviewyouranswers'.endsWith(body);
}
/**
* A packet can bundle setup tabs after the mode tab. Once the mode tab is
* answered, answer each later non-mode tab of the same unsubmitted call once,
* with the navigation rule (prerequisite pick, else option 1), so Submit is reachable.
*/
export function ceoModePacketTabAnswer(
visible: string, selected: NativePlanQuestionCall | undefined, transcript: PlanCountTranscript, answered: Set<string>,
): { question: AskUserQuestionFingerprint; index: number } | null {
if (!selected || !selected.sessionId || !selected.toolUseId || transcript.status !== 'ready' ||
selected.questions.length < 2 || selected.questions.length > 4 || selected.questions.some(q => q.multiSelect)) return null;
const id = `${selected.sessionId}:${selected.toolUseId}`;
const current = transcript.calls.filter(call => `${call.sessionId}:${call.toolUseId}` === id);
if (current.length !== 1 || current[0]!.answered || current[0]!.failed ||
JSON.stringify(current[0]!.questions) !== JSON.stringify(selected.questions)) return null;
const bar = posturePacketBar(visible);
if (!bar || JSON.stringify(bar.headers) !== JSON.stringify(selected.questions.map(q => q.header.trim().replace(/\s+/g, ' ')))) return null;
const modeIndex = selected.questions.findIndex(q => q.options.filter(o => modeTitle(o.label)).length >= 2);
if (modeIndex < 0 || !bar.answered[modeIndex]) return null;
const question = capturePlanCountQuestion(visible, new Set(), 0, true, selected);
const tab = question?.nativeQuestionIndex;
if (!question || question.nativeCall !== selected || tab === undefined || tab <= modeIndex || bar.answered[tab] ||
JSON.stringify(question.options.map(o => o.label)) !== JSON.stringify(selected.questions[tab]!.options.map(o => o.label))) return null;
const key = `${id}:${tab}`;
if (answered.has(key)) return null;
answered.add(key);
return { question, index: planCountPrerequisitePick(question) ?? 1 };
}
/** Finish the selected native mode packet before waiting for its answer. */
export function ceoModeSubmissionInput(
visible: string, selected: NativePlanQuestionCall | undefined, targetMode: CeoMode,
transcript: PlanCountTranscript, submitted: Set<string>,
transcript: PlanCountTranscript, submitted: Set<string>, screenText = '',
): string | null {
if (!selected || selected.answered || selected.failed || !selected.sessionId || !selected.toolUseId ||
transcript.status !== 'ready' || selected.questions.length < 2 ||
@@ -161,16 +232,33 @@ export function ceoModeSubmissionInput(
const modeQuestions = selected.questions.filter(q => q.options.filter(o => modeTitle(o.label)).length >= 2);
if (modeQuestions.length !== 1 || findCeoModeOption(modeQuestions[0]!.options.map((o, i) =>
({index:i + 1, label:o.label})), targetMode) === null) return null;
const bar = posturePacketBar(visible);
if (!bar || !bar.answered.every(Boolean) || JSON.stringify(bar.headers) !== JSON.stringify(
selected.questions.map(q => q.header.trim().replace(/\s+/g, ' '))) ||
planCountSubmissionInput(visible) !== '\r') return null;
const rawBar = [...visible.matchAll(/←[^\r\n]+✔\s*Submit\s*→/g)].at(-1)!;
const preceding = visible.slice(0, rawBar.index);
if (/```|~~~|^\s*>|\b(?:example|quoted|source)[^:\n]*:\s*$/im.test(preceding)) return null;
const compact = (text: string) => text.replace(/\s+/g, '');
const panel = compact(visible.slice(rawBar.index! + rawBar[0].length)
.replace(/^[ \t]*[│┃] ?/gm, '').replace(/^[ \t]*[●⏺] ?/gm, ''));
const quotedContext = /```|~~~|^\s*>|\b(?:example|quoted|source)[^:\n]*:\s*$/im;
const bar = posturePacketBar(visible);
let review: string;
if (bar) {
if (!bar.answered.every(Boolean) || JSON.stringify(bar.headers) !== JSON.stringify(
selected.questions.map(q => q.header.trim().replace(/\s+/g, ' '))) ||
planCountSubmissionInput(visible) !== '\r') return null;
const rawBar = [...visible.matchAll(/←[^\r\n]+✔\s*Submit\s*→/g)].at(-1)!;
if (quotedContext.test(visible.slice(0, rawBar.index))) return null;
review = visible.slice(rawBar.index! + rawBar[0].length);
} else {
// A review taller than the terminal scrolls its tab bar and heading off
// the viewport (run 36606688266). The viewport must still end at the
// focused Submit prompt; the accumulated screen text then supplies the
// one complete review panel, authenticated below exactly as with a bar.
const heading = screenText.lastIndexOf('Review your answers');
if (heading < 0 && compact(screenText).endsWith(BARLESS_SUBMIT_END) &&
clippedReviewMatches(visible, selected, modeQuestions[0]!, targetMode)) {
submitted.add(id);
return '\r';
}
if (heading < 0 || !compact(visible).endsWith(BARLESS_SUBMIT_END) ||
quotedContext.test(screenText.slice(0, heading).split('\n').slice(-3).join('\n'))) return null;
review = screenText.slice(heading);
}
const panel = compact(review.replace(/^[ \t]*[│┃] ?/gm, '').replace(/^[ \t]*[●⏺] ?/gm, ''));
// Authenticate the complete review panel against native questions and
// offered answers. An intended keypress or a selected-mode echo is not an ACK.
let prefixes = ['Reviewyouranswers'];
@@ -291,7 +379,8 @@ function singleScopeBrief(text: string, descriptions: readonly string[], compari
quote => quote.replace(/\?/g, '')) : text;
const questions = questionText.replace(/\?[A-Za-z_][\w-]*=/g, '=').match(/\?/g);
if ((questions?.length ?? 0) !== (proposalHeading ? 0 : 1) || /```|~~~|^\s*>/m.test(text)) return false;
const comparisonMarker = expansion
// The preamble requires the Note form for different-kind menus (Add/Defer/Skip, Defer/Keep).
const comparisonMarker = expansion || !comparison
? /Completeness:|Note:\s*options differ in kind, not coverage\s*[—–-]\s*no completeness score\./gi
: /Completeness:/gi;
const markers = [/Project\/branch\/task:/gi, /ELI10:/gi, /Stakes if (?:we pick )?wrong:/gi,
@@ -312,7 +401,7 @@ function singleScopeBrief(text: string, descriptions: readonly string[], compari
if (!ratings.length || ratings.some(score => Number(score[1]) > 10)) return false;
}
return complete && (comparison ? /^[^.!?;\n]+ (?:vs|versus) [^.!?;\n]+\.$/.test(net)
: /^[^.!?;\n]+\.$/.test(net.replace(/\bvs\./gi, 'vs')));
: /^[^.!?\n]+\.$/.test(net.replace(/\bvs\./gi, 'vs')));
}
/** Fixture-owned baseline for a completed scope-preservation decision. */
@@ -388,7 +477,9 @@ function hasAnsweredHoldPosture(transcript: PlanCountTranscript, selected: Nativ
// standalone prose is published. Metadata and answer echoes do not count.
// This recognizes posture language; it does not validate every scope choice.
const context = /Project\/branch\/task:([\s\S]*?)(?=ELI10:)/i.exec(q.question)?.[1] ?? '';
const rationale = /ELI10:([\s\S]*?)(?=Stakes if (?:we pick )?wrong:)/i.exec(q.question)?.[1]?.trim() ?? '';
// The ELI10 and the Recommendation's reason are both the brief's own rationale.
const rationale = [/ELI10:([\s\S]*?)(?=Stakes if (?:we pick )?wrong:)/i, /Recommendation:[^\n]*?\bbecause\b([^\n]*)/i]
.map(part => part.exec(q.question)?.[1]?.trim() ?? '').join('\n');
const offered = q.options.map(o => o.label.trim());
if (!q.multiSelect && q.options.length >= 2 && q.options.length <= 4 && new Set(offered).size === offered.length &&
offered.includes(call.answers?.[q.question] ?? '') && /\bHOLD SCOPE\b/i.test(context) &&
@@ -961,3 +1052,15 @@ export function nextCeoPostureContinuation(
} else postureContinuations.set(seenQuestions, { modeId });
return 'question';
}
/** HOLD SCOPE's own "Deferring current scope" menu: one question, exactly a
* Defer-to-TODOS option and a Keep-in-scope option. Returns the Keep index. */
export function holdDeferKeepIndex(call: NativePlanQuestionCall | undefined): number | null {
if (call?.questions.length !== 1) return null;
const q = call.questions[0]!;
if (q.multiSelect || q.options.length !== 2) return null;
const labels = q.options.map(option => option.label.trim().replace(/^[A-Z][).:]\s+/, '').replace(/\s*\(recommended\)\s*$/i, ''));
const defer = labels.findIndex(label => /^Defer\b[^\n]*\bTODOS(?:\.md)?$/i.test(label));
const keep = labels.findIndex(label => /^Keep\b[^\n]*\bin scope$/i.test(label));
return defer >= 0 && keep >= 0 && defer !== keep ? keep + 1 : null;
}
+35 -1
View File
@@ -758,6 +758,40 @@ function hasOrderedStaleFillOperations(text: string, sourceText = text): boolean
}
/** Explicit copied/example framing owns its section and descendant headings. */
/**
* An arrow-separated execution order establishes the overlap by event roles, not
* wording: a reader misses before a writer commits and invalidates, that same
* reader then fills its pre-write value, and a reader begun after the write gets it.
*/
function hasArrowOrderedStaleFill(block: string): boolean {
return block.replace(/[*_`]/g, '').split(/\s*\|\s*|\n/).some(cell => {
if (/\b(?:impossible|cannot\s+happen|not\s+(?:a|an)\s+(?:bug|defect|violation|race|gap))\b/i.test(cell)) return false;
const events = cell.split(/\s*(?:->|→)\s*/).map(event => event.slice(event.lastIndexOf(':') + 1).trim());
if (events.length < 5) return false;
const actor = (event: string) => /^(R[1-9]\d*|W[1-9]\d*|W)\b/.exec(event)?.[1];
const version = (event: string) => /\b(v[0-9]+)\b/i.exec(event)?.[1]?.toLowerCase();
const find = (from: number, test: (event: string, who: string | undefined) => boolean) =>
events.findIndex((event, index) => index > from && test(event, actor(event)));
const miss = find(-1, (event, who) => /^R/.test(who ?? '') && /\bmiss(?:es)?\b/i.test(event));
if (miss < 0) return false;
const reader = actor(events[miss]!)!;
const commit = find(miss, (event, who) => /^W/.test(who ?? '') && /\bcommit(?:s|ted)?\b/i.test(event));
if (commit < 0) return false;
const writer = actor(events[commit]!)!;
const invalidate = find(commit, (event, who) => who === writer && /\b(?:delete|invalidat\w*|evict\w*)\b/i.test(event));
const fill = find(invalidate, (event, who) => who === reader && /\b(?:cache\.set|set|fills?|refills?|stores?|caches)\b/i.test(event));
const later = find(fill, (event, who) => /^R/.test(who ?? '') && who !== reader
&& new RegExp(String.raw`\b(?:begun|began|begins|started|starts)\s+after\s+${writer}\b`).test(event)
&& /\b(?:hits?|gets?|reads?|sees?|observes?|returns?)\b/i.test(event));
if (invalidate < 0 || fill < 0 || later < 0) return false;
const old = version(events[fill]!) ?? events.slice(miss, commit).map(version).find(Boolean);
const fresh = version(events[commit]!);
if (old && fresh && old === fresh) return false;
const seen = version(events[later]!);
return !(old && seen && seen !== old) && (Boolean(old) || /\b(?:old|stale|pre[- ]write)\b/i.test(events[fill]! + events[later]!));
});
}
function assertedProseOwner(prose: string[], index: number): boolean {
const owners = [{ level: 0, source: false }];
for (const line of prose.slice(0, index + 1)) {
@@ -855,7 +889,7 @@ function hasProseStaleFillFinding(report: string): boolean {
if (!assertedProseOwner(owners, owners.length - 1)) return false;
const stale = /\b(?:stale|outdated)\b|\b(?:old(?:er)?|pre[- ]write)\s+(?:value|data|result|version|snapshot)\b/i.test(text);
const inFlight = /\b(?:race|racing|concurrent|concurrency|in[- ]flight|pending)\b/i.test(text)
|| hasOrderedStaleFillOperations(text, block);
|| hasOrderedStaleFillOperations(text, block) || hasArrowOrderedStaleFill(block);
const read = /\b(?:read|fetch)\w*\b/i.test(text);
const fillPattern = /\b(?:fill|refill|repopulat|populat|insert|stor|restor)\w*\b|\bcache\.set\b|\bcache(?:s|d)?\s+(?:the|an?|old|stale|same)\s+(?:\w+\s+){0,2}(?:value|data|result|snapshot)\b/i;
const fill = fillPattern.test(text);
+14 -10
View File
@@ -18,20 +18,24 @@ export function ceoSplitOptionAction(label: string): 'include' | 'defer' | 'cut'
}
/** Candidate-shaped menus for live progress only. Final coverage, subject and
* independence are established by evaluatePlanReviewDecisions over every call. */
* independence are established by evaluatePlanReviewDecisions over every call.
* Identity comes from the native header. The question opens with that
* candidate's ledger reference (E1 or a row ID ending in it), names only that
* candidate, and offers exactly one include, defer and cut disposition. */
export function ceoSplitCandidate(question: NativeQuestion): string | null {
const header = /^E([1-5])\s+(.+)$/.exec(question.header.trim());
if (!header || question.multiSelect || question.options.length < 3 || question.options.length > 4) return null;
const index = Number(header[1]) - 1;
const lead = question.question.split(/\r?\n/, 1)[0]!
.replace(/^D[1-9]\d*(?:\.[1-9]\d*)?\s*[—–:-]\s*/, '');
const target = /^E([1-5])[):]\s+(.+\?)$/.exec(lead);
if (!target || question.multiSelect || question.options.length < 3 || question.options.length > 4) return null;
const id = `E${target[1]}`;
const platform = platforms[Number(target[1]) - 1]!;
if (!new RegExp(`^${id}\\s+${platform}$`, 'i').test(question.header.trim()) ||
!new RegExp(`\\b${platform}\\b`, 'i').test(target[2]!) ||
/\bE[1-5][):]/.test(target[2]!)) return null;
const names = (platform: string) => new RegExp(`\\b${platform}\\b`, 'i').test(lead);
if (!new RegExp(`^${platforms[index]}$`, 'i').test(header[2]!) ||
!new RegExp(`^\\S*\\bE${header[1]}[):]\\s+.+\\?$`).test(lead) || !names(platforms[index]!) ||
platforms.some((platform, i) => i !== index && names(platform)) ||
[...lead.matchAll(/\bE([1-9]\d*)\b/g)].some(match => match[1] !== header[1])) return null;
const actions = question.options.map(option => ceoSplitOptionAction(option.label));
return actions.every(Boolean) && new Set(actions).size === actions.length &&
['include', 'defer', 'cut'].every(action => actions.includes(action)) ? id : null;
return ['include', 'defer', 'cut'].every(action => actions.filter(found => found === action).length === 1)
? `E${header[1]}` : null;
}
export function isCeoSplitCandidateCall(fp: AskUserQuestionFingerprint): boolean {
@@ -17,6 +17,7 @@ import {
nativePlanCallFingerprint,
devexStep0Boundary,
type AskUserQuestionFingerprint,
pickDesignFocusAll,
} from './claude-pty-runner';
describe('Step0BoundaryPredicate per-skill', () => {
@@ -1772,3 +1773,21 @@ describe('explicit Step 0 complexity gate with size in native choices', () => {
}
});
});
describe('pickDesignFocusAll: the seed-declared all-seven answer for the pending 0D focus menu', () => {
const captured = require('../fixtures/design-floor-focus-36597762183.json');
const q = () => structuredClone(captured.question);
test('census 36597762183 pending menu selects the all-seven option', () => {
expect(pickDesignFocusAll(q())).toBe(1);
});
test.each([
['a different title', (x: any) => { x.question = x.question.replace('Review all 7 design dimensions, or focus?', 'Which fixes should I apply?'); }],
['a non-narrowing alternative', (x: any) => { x.options[1].label = 'Approve every fix now'; }],
['two all-seven options', (x: any) => { x.options[1].label = 'All seven dimensions'; }],
['a bundled product approval', (x: any) => { x.question = x.question.replace('Net:', 'Also approve the CTA redesign.\nNet:'); }],
['a foreign plan context', (x: any) => { x.question = x.question.replace('plan-design-review of PLAN.md', 'plan-design-review of OTHER.md'); }],
['multi select', (x: any) => { x.multiSelect = true; }],
])('%s is not answered', (_name, mutate) => {
const x = q(); mutate(x); expect(pickDesignFocusAll(x)).toBeNull();
});
});
+2 -2
View File
@@ -19,9 +19,9 @@ export { isProseAUQVisible, isScopeGateQuestionVisible, isScopeGateAutoSelectVis
export type { ClassifyResult } from './pty/classify';
export { nativePlanCallFingerprint, planCountQuestionPhase, parseQuestionPrompt, auqFingerprint, planCountQuestionInput, matchesNativePlanQuestion, capturePlanCountQuestion, createPlanCountPermissionGuard, planCountPrerequisitePick } from './pty/auq';
export type { AskUserQuestionFingerprint, Step0BoundaryPredicate } from './pty/auq';
export { assertReviewReportAtBottom, hasNativePlanCompletion, isQuestionlessNativePlanExit, evaluateOwnedNativePlanTerminal, hasNativePlanTerminal, assertReportAtBottomIfPlanWritten } from './pty/plan-native';
export { assertReviewReportAtBottom, hasCompletePlanReport, hasNativePlanCompletion, isQuestionlessNativePlanExit, evaluateOwnedNativePlanTerminal, hasNativePlanTerminal, assertReportAtBottomIfPlanWritten } from './pty/plan-native';
export type { ReviewReportAtBottomResult, NativePlanTerminalReview, NativePlanTerminalAssessment, NativePlanTerminalEvaluator } from './pty/plan-native';
export { ceoStep0Boundary, engSetupAUQ, engFirstReviewAUQ, engStep0Boundary, designReviewSetupAUQ } from './pty/boundaries';
export { ceoStep0Boundary, engSetupAUQ, engFirstReviewAUQ, engStep0Boundary, pickDesignFocusAll } from './pty/boundaries';
export { runPlanSkillObservation } from './pty/runners/observation';
export type { PlanSkillObservation, PlanSkillObservationOptions } from './pty/runners/observation';
export { runPlanSkillCounting, countingCapture, isNativeCompletionSummary } from './pty/runners/counting';
+9 -1
View File
@@ -133,6 +133,7 @@ export function installSkillToTempHome(
skillName: string,
tempHome?: string,
sections?: string[],
runtimeRoot?: string,
): string {
const home = tempHome || fs.mkdtempSync(path.join(os.tmpdir(), 'codex-e2e-'));
const destDir = path.join(home, '.codex', 'skills', skillName);
@@ -149,6 +150,11 @@ export function installSkillToTempHome(
// nonexistent skill in its response and otherwise pass discovery checks.
fs.copyFileSync(srcSkill, path.join(destDir, 'SKILL.md'));
}
if (runtimeRoot) {
// The temp HOME has no installed gstack runtime; point runtime helpers at the one under test.
const installed = path.join(destDir, 'SKILL.md');
fs.writeFileSync(installed, fs.readFileSync(installed, 'utf8').replaceAll('~/.codex/skills/gstack', runtimeRoot));
}
const srcOpenAIYaml = path.join(skillDir, 'agents', 'openai.yaml');
if (fs.existsSync(srcOpenAIYaml)) {
@@ -180,6 +186,7 @@ export async function runCodexSkill(opts: {
configOverrides?: string[]; // TOML key=value overrides (passed with -c)
ignoreUserConfig?: boolean; // Add --ignore-user-config; auth still comes from CODEX_HOME
signal?: AbortSignal; // Abort the process group when an enclosing eval expires
runtimeRoot?: string; // gstack runtime that ~/.codex/skills/gstack helper paths resolve to
}): Promise<CodexResult> {
const {
skillDir,
@@ -193,6 +200,7 @@ export async function runCodexSkill(opts: {
configOverrides = [],
ignoreUserConfig = false,
signal,
runtimeRoot,
} = opts;
const startTime = Date.now();
@@ -223,7 +231,7 @@ export async function runCodexSkill(opts: {
const realHome = os.homedir();
try {
installSkillToTempHome(skillDir, name, tempHome, sections);
installSkillToTempHome(skillDir, name, tempHome, sections, runtimeRoot);
// Copy authentication only. Copying the whole operator ~/.codex tree leaks
// plugins, MCP servers, rules, memories, and skills into a supposedly
+3 -1
View File
@@ -174,7 +174,9 @@ function readsFile(command: unknown, file: string, cwd: string, output: unknown,
const andDisplay = (p: string) => {
if (p === 'echo' || /^echo\s+[-=]+$/.test(p) || /^echo [-=]{2,} [A-Za-z0-9_.\/-]+ [-=]{2,}$/.test(p)) return true;
const caption = /^echo\s+(.+)$/.exec(p), value = caption && literal(caption[1]!);
if (value && /^[-=]{2,}(?:\s*[A-Za-z0-9_][A-Za-z0-9_./-]*(?:\s+(?:vs|and)\s+[A-Za-z0-9_][A-Za-z0-9_./-]*)?\s*)?[-=]{2,}$/.test(value)) return true;
// A fenced caption may name the next display in plain words, such as
// "=== git diff main --stat ==="; an unfenced command string stays data.
if (value && /^[-=]{2,}(?:\s*[A-Za-z0-9_][A-Za-z0-9_./-]*(?:\s+[A-Za-z0-9_./-]+)*\s*)?[-=]{2,}$/.test(value)) return true;
return /^git\s+diff(?:\s+[A-Za-z0-9_][A-Za-z0-9_./~^-]*)?\s+--stat$/.test(p) ||
/^git\s+log\s+--oneline\s+[A-Za-z0-9_][A-Za-z0-9_./~^-]*$/.test(p);
};
+87 -2
View File
@@ -81,11 +81,26 @@ function preRunLogRecordValue(before: string, nextClause: string): boolean {
// The immediately following assertion must keep the same record as its
// subject and explicitly exclude this workflow as its origin. A later
// current completion occurrence is still checked independently below.
return disownsRun(nextClause);
}
/** The next assertion keeps the record as its subject and excludes this run as its origin. */
function disownsRun(nextClause: string): boolean {
return /^(?:that|the|this)\s+(?:record|entry|line)\s+(?:was|is)\s+not\s+(?:produced|created|written|recorded)\s+(?:by|during|in)\s+(?:this|my)\s+(?:run|session|workflow)\b/i.test(nextClause.trim()) ||
/^(?:that|the|this)\s+(?:record|entry|line)\s+predates\s+(?:this|my)\s+(?:run|session|workflow)\s+and\s+was\s+not\s+(?:produced|created|written|recorded)\s+by\s+it\b/i.test(nextClause.trim()) ||
/^(?:that|the|this)\s+(?:record|entry|line)\s+(?:does not|doesn't|cannot)\s+(?:reflect|establish|provide|supply)\s+(?:current\s+)?outside\s+(?:review\s+)?coverage\s+(?:from|for)\s+(?:this|my)\s+(?:run|session|workflow)\b/i.test(nextClause.trim());
}
/** A record named by the retained prior record's own clock, then disowned, owns its reported value. */
function priorClockRecordValue(before: string, nextClause: string, priorRecord?: Record<string, unknown>): boolean {
const at = /T(\d{2}):(\d{2}):(\d{2})/.exec(String(priorRecord?.timestamp ?? ''));
if (!at || priorRecord?.outside_status !== 'completed') return false;
const owner = new RegExp(String.raw`\b(?:earlier|prior|previous|old(?:er)?|historical|pre[- ]existing)\s+(?:review[- ]log\s+)?(?:record|entry|line)\s+(?:from|at|dated|timestamped)\s+${at[1]}:${at[2]}(?::${at[3]}(?:\.\d+)?)?(?![\d:])`, 'i').exec(before);
if (!owner) return false;
const value = before.slice(owner.index + owner[0].length);
return !/\b(?:this|my)\s+(?:run|session|workflow)\b|\b(?:now|currently|current|new|updat\w*|append\w*)\b/i.test(value) && disownsRun(nextClause);
}
/** Structured quotations must belong to the exact retained prior record. */
function withoutAttributedPriorRecordData(output: string, priorRecord?: Record<string, unknown>): string {
if (!priorRecord || priorRecord.outside_status !== 'completed') return output;
@@ -100,7 +115,7 @@ function withoutAttributedPriorRecordData(output: string, priorRecord?: Record<s
}
else {
record = {};
for (const part of text.split(',')) {
for (const part of text.split(/[,/;]/)) {
const field = /^\s*["']?([a-z_]+)["']?\s*[:=]\s*["']?([a-z0-9_.:+-]+)["']?\s*$/i.exec(part);
if (!field || Object.hasOwn(record, field[1]!)) return false;
record[field[1]!] = field[2]!;
@@ -142,12 +157,81 @@ function withoutAttributedPriorRecordData(output: string, priorRecord?: Record<s
spans.push({ start: match.index, end: match.index + match[0].length });
}
}
// An inline quotation of the retained record's exact status/source/outside_status
// values is that record when its own sentence names it as pre-existing and
// makes no current claim; wording order around the quotation does not matter.
// A named record timestamp must denote the retained record's instant at the precision written.
const priorMs = typeof priorRecord.timestamp === 'string' ? Date.parse(priorRecord.timestamp) : NaN;
const sameInstant = (stamp: string): boolean => {
if (!Number.isFinite(priorMs)) return false;
const iso = priorMs ? new Date(priorMs).toISOString() : '';
const clock = /^(\d{2}:\d{2}(?::\d{2}(?:\.\d{1,3})?)?)Z?$/.exec(stamp);
if (clock) return iso.slice(11, 11 + clock[1]!.length) === clock[1];
const at = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}(?::\d{2}(?:\.\d+)?)?Z$/.test(stamp) ? Date.parse(stamp) : NaN;
return Number.isFinite(at) && iso.slice(0, stamp.includes('.') ? 23 : stamp.length - 1) === new Date(at).toISOString().slice(0, stamp.includes('.') ? 23 : stamp.length - 1);
};
const sentenceOwnsPriorValue = (index: number, length: number): boolean => {
const start = Math.max(output.lastIndexOf('\n', index - 1), ...['. ', '! ', '? ', '; '].map(end => output.lastIndexOf(end, index - 1) + 1)) + 1;
const ends = ['\n', '. ', '! ', '? ', '; '].map(end => output.indexOf(end, index + length)).filter(at => at >= 0);
const sentence = (output.slice(start, index) + ' ' + output.slice(index + length, ends.length ? Math.min(...ends) : output.length))
.replace(/[*`]/g, '').replace(/\b(?:predates|before)\s+(?:this|my)\s+(?:run|session|workflow)(?:\s+(?:started|began))?\b/gi, 'beforehand')
.replace(/\b(?:I|we)\s+(?:did\s+not|didn't|never)\s+(?:write|create|produce|record)\b/gi, 'unauthored');
const stamps = [...sentence.matchAll(/\btimestamp(?:ed)?\s+([0-9T:.Z-]+)/gi)].map(stamp => stamp[1]!.replace(/[.,;:]+$/, ''));
if (stamps.some(stamp => !sameInstant(stamp))) return false;
return !/\b(?:after|another|other|if|unless)\b/i.test(sentence)
&& /\b(?:earlier|prior|previous|historical|old(?:er)?|pre[- ]existing|stale)\s+(?:(?:review[- ]log|review|log)\s+)?(?:entry|record|line|row)\b/i.test(sentence)
&& !/\b(?:now|currently|current|today|new|updat\w*|append\w*|chang\w*|mark\w*|set|write|wrote|reports?|conclud\w*)\b|\bthis\s+(?:run|session|workflow)\b|\boutside_status\b|\bboth reviewers agree\b/i.test(sentence);
};
for (const match of output.matchAll(/`([^`\r\n]+)`/g)) {
if (spans.some(span => span.start <= match.index && match.index < span.end)) continue;
if (ownsPriorValue(output.slice(0, match.index), false) && matchesPrior(match[1]!, false)) {
if ((ownsPriorValue(output.slice(0, match.index), false) || sentenceOwnsPriorValue(match.index, match[0].length)) && matchesPrior(match[1]!, false)) {
spans.push({ start: match.index, end: match.index + match[0].length });
}
}
// A parenthesized field list right after a pre-existing-record owner is that
// record when it quotes the record's exact ISO timestamp and every other item
// is one of its own field values; a current mutation before the owner fails.
const owned = /\b(?:earlier|prior|previous|historical|old(?:er)?|pre[- ]existing)\s+(?:(?:review[- ]log|review|log)\s+)?(?:entry|record|line|row)\s*\(([^()\r\n]+)\)/gi;
for (const match of output.matchAll(owned)) {
const lineStart = output.lastIndexOf('\n', match.index) + 1;
const local = output.slice(lineStart, match.index).split(/(?<=[.!?;])\s+/).at(-1) ?? '';
if (/\b(?:now|currently|current|today|new|updat\w*|append\w*|chang\w*|mark\w*|set|write|wrote)\b/i.test(local.replace(/[*`]/g, ''))) continue;
const items = match[1]!.split(',').map(item => item.replace(/[*`]/g, '').trim());
const fieldsOk = items.every(item => {
if (item === priorRecord.timestamp) return true;
const field = /^["']?([a-z_]+)["']?\s*[:=]\s*["']?([a-z0-9_.:+-]+)["']?$/i.exec(item);
return !!field && fields.has(field[1]!) && field[1] !== 'timestamp' && priorRecord[field[1]!] === field[2];
});
if (!fieldsOk || !items.includes(String(priorRecord.timestamp)) || !items.some(item => /^["']?outside_status\b/i.test(item))) continue;
const start = match.index + match[0].length - match[1]!.length - 1;
spans.push({ start, end: start + match[1]!.length + 2 });
}
// A quoted fragment carrying the retained record's exact timestamp is that
// record's data when every field it quotes has that record's value.
for (const match of output.matchAll(/`([^`\r\n]+)`/g)) {
if (typeof priorRecord.timestamp !== 'string' || !match[1]!.includes(priorRecord.timestamp)) continue;
const pairs = [...match[1]!.matchAll(/["']?([a-z_]+)["']?\s*[:=]\s*["']?([^"',}\s]+)["']?/gi)].filter(pair => fields.has(pair[1]!));
if (!pairs.some(pair => pair[1] === 'outside_status') || pairs.some(pair => priorRecord[pair[1]!] !== pair[2])) continue;
spans.push({ start: match.index!, end: match.index! + match[0].length });
}
// A whole sentence that names the pre-existing record, dates it before this
// run (its exact instant or an explicit "before this run"), quotes only that
// record's own field values and makes no current claim is that record's
// report, however its fields are quoted or split.
for (const sentence of output.matchAll(/[^\n.!?;]*(?:[.!?;](?=\S)[^\n.!?;]*)*(?:[.!?;](?=\s|$)|\n|$)/g)) {
const plain = sentence[0].replace(/[*`]/g, '');
if (!/\boutside_status["']*\s*[:=]\s*["']*completed\b/i.test(plain)) continue;
if (!/\b(?:earlier|prior|previous|historical|old(?:er)?|pre[- ]existing|stale|seeded)\s+(?:(?:review[- ]log|review|log)\s+)?(?:entry|record|line|row)\b/i.test(plain)) continue;
const beforeRun = /\b(?:predates|before)\s+(?:this|my)\s+(?:run|session|workflow)(?:\s+(?:started|began))?\b/i;
const stamps = [...plain.matchAll(/\b(?:\d{4}-\d{2}-\d{2}T)?\d{2}:\d{2}(?::\d{2}(?:\.\d{1,3})?)?Z?\b/g)].map(m => m[0]);
if (stamps.some(stamp => !sameInstant(stamp)) || (!stamps.length && !beforeRun.test(plain))) continue;
const quoted = [...plain.matchAll(/\b([a-z_]+)["']?\s*[:=]\s*["']?([a-z0-9_.+-]+)["']?/gi)].filter(m => fields.has(m[1]!) && m[1] !== 'timestamp');
if (!['status', 'source', 'outside_status'].every(key => quoted.some(m => m[1] === key))
|| quoted.some(m => priorRecord[m[1]!] !== m[2])) continue;
if (/\b(?:now|currently|current|today|new|updat\w*|append\w*|chang\w*|wrote|recorded by me)\b|\bboth reviewers agree\b/i
.test(plain.replace(beforeRun, ''))) continue;
spans.push({ start: sentence.index!, end: sentence.index! + sentence[0].length });
}
for (const span of spans.sort((a, b) => b.start - a.start)) {
output = output.slice(0, span.start) + output.slice(span.start, span.end).replace(/[^\r\n]/g, ' ') + output.slice(span.end);
}
@@ -174,6 +258,7 @@ function hasUnattributedOutsideCompletion(output: string, priorRecord?: Record<s
// subject is "that record" rather than "the earlier record".
const datedBeforeRun = String.raw`\s+is\s+timestamped\s+(?:about\s+)?(?:a|an|one|two|\d+)\s+(?:minute|hour|day|week)s?\s+before\s+(?:this|my)\s+(?:run|session|workflow)`;
const recordPattern = new RegExp(String.raw`\b(?:(?:earlier|prior|historical|old(?:er)?)\s+(?:entry|record|line)|(?:that|the)\s+(?:entry|record|line)(?=${datedBeforeRun}))\b`, 'gi');
if (priorClockRecordValue(before, clauses[clauseIndex + 1] ?? '', priorRecord)) return false;
const record = [...before.matchAll(recordPattern)].at(-1);
if (!record) return !preRunLogRecordValue(before, clauses[clauseIndex + 1] ?? '');
// Bind this occurrence to an old record's reported value. A mere mention
+34 -31
View File
@@ -18,36 +18,38 @@ type DocsWriteContext = {
readOnly?: boolean;
};
function docsAtomicSources(observation: QAWriteObservation, allowed: string[], context?: DocsWriteContext): Set<string> {
const denied = new Set<string>();
/** Atomic temp files proven to be native replacements of DOC_PATH; `deniedAt` names the first unmet check. */
function docsAtomicSources(observation: QAWriteObservation, allowed: string[], context?: DocsWriteContext): Set<string> & { deniedAt?: string } {
const denied: Set<string> & { deniedAt?: string } = new Set<string>();
const deny = (line: number) => { denied.deniedAt = `docsync-observer.ts:${line}`; return denied; };
if (!context || context.readOnly || !allowed.includes(DOC_PATH) || !observation.complete || observation.failures.length) return denied;
const { result, fixture, scripts = [] } = context;
if (result.exitReason !== 'success' || !Array.isArray(result.transcript) || docsToolFailures(result, fixture, scripts).length) return denied;
if (result.exitReason !== 'success' || !Array.isArray(result.transcript) || docsToolFailures(result, fixture, scripts).length) return deny(27);
const failures: string[] = [];
const target = path.join(fixture.repo, DOC_PATH);
const native = nativeCalls(result.transcript, failures);
const calls = native.filter(call => ['Write', 'Edit'].includes(call.name)
&& typeof call.input.file_path === 'string' && path.resolve(fixture.repo, call.input.file_path) === target);
if (failures.length || !calls.length || calls.some((call, index) => call.failed || call.end <= call.start || (index > 0 && call.start <= calls[index - 1].end))) return denied;
if (failures.length || !calls.length || calls.some((call, index) => call.failed || call.end <= call.start || (index > 0 && call.start <= calls[index - 1].end))) return deny(33);
const before = observation.before[DOC_PATH];
const after = observation.after[DOC_PATH];
if (!/^\d+:[a-f0-9]{64}$/.test(before ?? '') || !/^\d+:[a-f0-9]{64}$/.test(after ?? '') || before.split(':')[0] !== after.split(':')[0]) return denied;
if (!/^\d+:[a-f0-9]{64}$/.test(before ?? '') || !/^\d+:[a-f0-9]{64}$/.test(after ?? '') || before.split(':')[0] !== after.split(':')[0]) return deny(36);
const hash = (text: string) => createHash('sha256').update(text).digest('hex');
const encoded = fixture.before?.contents[DOC_PATH];
if (typeof encoded !== 'string') return denied;
if (typeof encoded !== 'string') return deny(39);
const baseline = Buffer.from(encoded, 'base64');
let content = baseline.toString('utf8');
let contentHash = before.split(':')[1];
if (baseline.toString('base64') !== encoded || !Buffer.from(content).equals(baseline) || hash(content) !== contentHash) return denied;
if (baseline.toString('base64') !== encoded || !Buffer.from(content).equals(baseline) || hash(content) !== contentHash) return deny(43);
const seen = new Set([contentHash]);
for (const call of calls) {
const event = result.transcript[call.end];
const payload = event.tool_use_result;
const results = event.message.content.filter((block: any) => block?.type === 'tool_result');
if (results.length !== 1 || (results[0].is_error !== undefined && results[0].is_error !== false)) return denied;
if (results.length !== 1 || (results[0].is_error !== undefined && results[0].is_error !== false)) return deny(49);
const omitted = !Object.hasOwn(event, 'tool_use_result');
if (omitted) {
if (call.parent === null) return denied;
if (call.parent === null) return deny(52);
let child = call;
const ancestors = new Set<typeof call>();
while (child.parent !== null) {
@@ -55,71 +57,71 @@ function docsAtomicSources(observation: QAWriteObservation, allowed: string[], c
const blocks = result.transcript[candidate.end]?.message?.content?.filter((block: any) => block?.type === 'tool_result');
return blocks?.length === 1 && blocks[0].tool_use_id === child.parent;
});
if (parents.length !== 1) return denied;
if (parents.length !== 1) return deny(60);
const parent = parents[0];
const completion = result.transcript[parent.end];
const block = completion.message.content.find((block: any) => block?.type === 'tool_result');
if (!['Agent', 'Task'].includes(parent.name) || parent.failed || parent.input.run_in_background === true
|| parent.start >= child.start || parent.end <= child.end || ancestors.has(parent)
|| (block.is_error !== undefined && block.is_error !== false)) return denied;
|| (block.is_error !== undefined && block.is_error !== false)) return deny(66);
if (Object.hasOwn(completion, 'tool_use_result')) {
if (completion.tool_use_result?.status !== 'completed') return denied;
} else if (parent.parent === null) return denied;
if (completion.tool_use_result?.status !== 'completed') return deny(68);
} else if (parent.parent === null) return deny(69);
ancestors.add(parent);
child = parent;
}
} else if (!payload || payload.filePath !== target || payload.userModified !== false || payload.originalFile !== content) return denied;
} else if (!payload || payload.filePath !== target || payload.userModified !== false || payload.originalFile !== content) return deny(73);
if (call.name === 'Write') {
if (typeof call.input.content !== 'string' || (!omitted && (payload.type !== 'update' || payload.content !== call.input.content))) return denied;
if (typeof call.input.content !== 'string' || (!omitted && (payload.type !== 'update' || payload.content !== call.input.content))) return deny(75);
content = call.input.content;
} else {
const { old_string: old, new_string: replacement, replace_all: all = false } = call.input;
if (typeof old !== 'string' || !old || typeof replacement !== 'string' || typeof all !== 'boolean'
|| (!omitted && (payload.oldString !== old || payload.newString !== replacement || payload.replaceAll !== all))) return denied;
|| (!omitted && (payload.oldString !== old || payload.newString !== replacement || payload.replaceAll !== all))) return deny(80);
const parts = content.split(old);
if (parts.length < 2 || (!all && parts.length !== 2)) return denied;
if (parts.length < 2 || (!all && parts.length !== 2)) return deny(82);
content = parts.join(replacement);
}
contentHash = hash(content);
if (seen.has(contentHash)) return denied;
if (seen.has(contentHash)) return deny(86);
seen.add(contentHash);
}
if (contentHash !== after.split(':')[1]) return denied;
if (contentHash !== after.split(':')[1]) return deny(89);
const events = observation.events;
const destinations = events.flatMap((event, index) => event.path === DOC_PATH && event.mask === 0x80 ? [index] : []);
if (destinations.length !== calls.length) return denied;
if (destinations.length !== calls.length) return deny(92);
const sources = new Set<string>();
let previous = -1;
for (const destination of destinations) {
const move = events[destination];
if (!Number.isInteger(move.cookie) || move.cookie <= 0 || move.cookie > 0xffffffff) return denied;
if (!Number.isInteger(move.cookie) || move.cookie <= 0 || move.cookie > 0xffffffff) return deny(97);
const pair = events.flatMap((event, index) => event.cookie === move.cookie ? [index] : []);
if (pair.length !== 2 || pair[1] !== destination) return denied;
if (pair.length !== 2 || pair[1] !== destination) return deny(99);
const source = events[pair[0]];
if (source.mask !== 0x40 || source.path === DOC_PATH || path.dirname(source.path) !== path.dirname(DOC_PATH)
|| Object.hasOwn(observation.before, source.path) || Object.hasOwn(observation.after, source.path) || sources.has(source.path)) return denied;
|| Object.hasOwn(observation.before, source.path) || Object.hasOwn(observation.after, source.path) || sources.has(source.path)) return deny(102);
const lifecycle = events.flatMap((event, index) => event.path === source.path ? [{ event, index }] : []);
if (lifecycle[0]?.event.mask !== 0x100 || lifecycle[0].index <= previous || lifecycle.at(-1)?.index !== pair[0]) return denied;
if (lifecycle[0]?.event.mask !== 0x100 || lifecycle[0].index <= previous || lifecycle.at(-1)?.index !== pair[0]) return deny(104);
let modified = false;
let closed = false;
for (const { event, index } of lifecycle) {
if (index === pair[0]) { if (!modified || !closed) return denied; continue; }
if (event.cookie !== 0) return denied;
if (index === pair[0]) { if (!modified || !closed) return deny(108); continue; }
if (event.cookie !== 0) return deny(109);
if (index === lifecycle[0].index) continue;
if (event.mask === 0x2 && !closed) modified = true;
else if (event.mask === 0x4 && !closed) continue;
else if (event.mask === 0x8 && modified) closed = true;
else return denied;
else return deny(114);
}
sources.add(source.path);
previous = destination;
}
if (events.some((event, index) => event.path === DOC_PATH && (index < destinations[0]
|| ![0x80, 0x4, 0x400, 0x800].includes(event.mask) || (event.mask !== 0x80 && event.cookie !== 0)))) return denied;
|| ![0x80, 0x4, 0x400, 0x800].includes(event.mask) || (event.mask !== 0x80 && event.cookie !== 0)))) return deny(120);
for (const [index, destination] of destinations.entries()) {
const replaced = events.slice(destination + 1, destinations[index + 1]).filter(event => event.path === DOC_PATH);
if (replaced.filter(event => event.mask === 0x4).length !== 1 || replaced.filter(event => event.mask === 0x400).length !== 1
|| replaced.filter(event => event.mask === 0x800).length > 1) return denied;
|| replaced.filter(event => event.mask === 0x800).length > 1) return deny(124);
}
return sources;
}
@@ -129,7 +131,8 @@ export function docsWriteFailures(observation: QAWriteObservation, allowed: stri
if (!observation.complete) failures.push('incomplete docs write observation');
const atomicSources = docsAtomicSources(observation, allowed, context);
for (const file of new Set([...observation.events.map(e => e.path), ...observation.changed])) {
if (file !== '.qa-state/.observer-check' && !allowed.includes(file) && !atomicSources.has(file)) failures.push(`forbidden docs write: ${file}`);
if (file !== '.qa-state/.observer-check' && !allowed.includes(file) && !atomicSources.has(file))
failures.push(`forbidden docs write: ${file}${atomicSources.deniedAt && path.dirname(file) === path.dirname(DOC_PATH) ? ` (atomic replacement unproven at ${atomicSources.deniedAt})` : ''}`);
if (allowed.includes(file) && observation.before[file] && observation.after[file] &&
observation.before[file].split(':')[0] !== observation.after[file].split(':')[0]) failures.push(`document mode changed: ${file}`);
}
@@ -223,7 +226,7 @@ export function docsCommandAllowed(command: string, fixture: ReturnType<typeof f
export function docsNativeInterface(fixture: Pick<ReturnType<typeof fixtureDocs>, 'home' | 'repo' | 'skills'>, scripts: string[] = [], transport = false): string {
const skills = fixture.skills.split(path.sep).join('/');
return `Fixture observation interface (applies to parent and every child; include this interface in child prompts): Bash may execute only separate literal pwd, ls, cat, stat, sha256sum, Git read commands (status, diff, show, log, ls-files, rev-parse, merge-base, hash-object without -w, branch --show-current), the exact generated Preamble block with its spawned prefix, or literal installed gstack-skill-start/gstack-skill-end commands for document-release (start requires GSTACK_SESSION_KIND=spawned). No shell composition, custom interpreters, arbitrary scripts, inline eval or memory-mapped writes. The only additional scripts are ${scripts.length ? scripts.join(', ') : 'none'}. Read/Glob/Grep remain available. Use Write/Edit for permitted docs and private JSON/Markdown artifacts under ${fixture.home}; do not rewrite installed skills, config, actor state or scripts. No effects outside the owned fixture. The owner preserves evidence and cleans up. Missing observer coverage blocks acceptance; the Linux kernel monitor covers syscall writes in the product tree, not hostile processes or arbitrary external destinations.
return `Fixture observation interface (applies to parent and every child; include this interface in child prompts): Bash may execute only separate literal pwd, ls, cat, stat, sha256sum, Git read commands (status, diff, show, log, ls-files, rev-parse, merge-base, hash-object without -w, branch --show-current), the exact generated Preamble block with its spawned prefix, or literal installed gstack-skill-start/gstack-skill-end commands for document-release (start requires GSTACK_SESSION_KIND=spawned). No shell composition, custom interpreters, arbitrary scripts, inline eval or memory-mapped writes. The only additional scripts are ${scripts.length ? scripts.join(', ') : 'none'}. Read/Glob/Grep remain available; Read skill and section files with Read (offset/limit for ranges), because Bash output over 30KB becomes a preview that no permitted Bash command can page. Use Write/Edit for permitted docs and private JSON/Markdown artifacts under ${fixture.home}; do not rewrite installed skills, config, actor state or scripts. No effects outside the owned fixture. The owner preserves evidence and cleans up. Missing observer coverage blocks acceptance; the Linux kernel monitor covers syscall writes in the product tree, not hostile processes or arbitrary external destinations.
The working directory for parent and child Bash calls is already ${fixture.repo}. Run Git reads directly, for example: git status, git diff --cached, git merge-base main HEAD, git rev-parse HEAD. Do not use Git global options such as -C, -c, --git-dir or --work-tree, and do not prepend cd or another shell wrapper. The literal git subcommand must immediately follow git; an absolute owned repository path does not make git -C an allowed command.
+3 -2
View File
@@ -24,11 +24,12 @@
import { describe } from 'bun:test';
export type E2ETier = 'gate' | 'periodic';
export type E2ETier = 'gate' | 'periodic' | 'marathon';
/**
* True when this process should run whole-file-gated paid tests of `tier`:
* EVALS=1 AND EVALS_TIER exactly equals the tier.
* EVALS=1 AND EVALS_TIER exactly equals the tier. 'marathon' cases (full
* end-to-end flows) therefore never run in the gate/PR or periodic lanes.
*
* Deliberate consequence: EVALS=1 with EVALS_TIER unset is false for BOTH
* tiers. Tierless runs (`test:evals` / `eval:bg` / `eval:bg:all`) skip every
+13
View File
@@ -64,6 +64,19 @@ describe('e2e-gate: env matrix (read at call time)', () => {
expect(describeE2ETier('periodic')).toBe(describe.skip);
});
test('marathon runs only in its own lane; gate and periodic lanes skip it', () => {
process.env.EVALS = '1';
for (const lane of ['gate', 'periodic']) {
process.env.EVALS_TIER = lane;
expect(e2eTierEnabled('marathon')).toBe(false);
expect(describeE2ETier('marathon')).toBe(describe.skip);
}
process.env.EVALS_TIER = 'marathon';
expect(describeE2ETier('marathon')).toBe(describe);
expect(describeE2ETier('gate')).toBe(describe.skip);
expect(describeE2ETier('periodic')).toBe(describe.skip);
});
test('EVALS=1 + EVALS_TIER unset → skip both tiers (the tierless test:evals / eval:bg:all trap)', () => {
process.env.EVALS = '1';
expect(e2eTierEnabled('gate')).toBe(false);
+17 -2
View File
@@ -116,9 +116,10 @@ export let selectedTests: string[] | null = resolveModuleSelection(
// EVALS_TIER: filter tests by tier after diff-based selection.
// 'gate' = gate tests only (CI default — blocks merge)
// 'periodic' = periodic tests only (weekly cron / manual)
// 'marathon' = full end-to-end flows only (non-blocking marathon lane)
// not set = run all selected tests (local dev default, backward compat)
if (evalsEnabled && process.env.EVALS_TIER) {
const tier = process.env.EVALS_TIER as 'gate' | 'periodic';
const tier = process.env.EVALS_TIER as 'gate' | 'periodic' | 'marathon';
const tierTests = Object.entries(E2E_TIERS)
.filter(([, t]) => t === tier)
.map(([name]) => name);
@@ -228,6 +229,18 @@ export function createEvalCollector(suite: string): EvalCollector | null {
}
/** DRY helper to record an E2E test result into the eval collector. */
/** Exit reasons for an API or transport failure (session-runner.ts). */
const INFRA_EXIT_REASONS = new Set(['error_api', 'timeout_startup', 'error_output_stream']);
/** API/transport error or CLI crash before the first model turn: INFRA, never a
* verdict on the product. Any assistant event or counted turn means the model
* ran, so its refusal, timeout or wrong answer stays an ordinary failure. */
export function isPreTurnInfraFailure(result: Pick<SkillTestResult, 'exitReason' | 'transcript' | 'costEstimate'>): boolean {
return result.costEstimate.turnsUsed === 0
&& (INFRA_EXIT_REASONS.has(result.exitReason) || /^exit_code_\d+$/.test(result.exitReason))
&& !result.transcript.some(event => event?.type === 'assistant');
}
export function recordE2E(
evalCollector: EvalCollector | null,
name: string,
@@ -240,9 +253,11 @@ export function recordE2E(
? `${result.toolCalls[result.toolCalls.length - 1].tool}(${JSON.stringify(result.toolCalls[result.toolCalls.length - 1].input).slice(0, 60)})`
: undefined;
const passed = extra?.passed ?? (result.exitReason === 'success' && result.browseErrors.length === 0);
evalCollector?.addTest({
name, suite, tier: 'e2e',
passed: result.exitReason === 'success' && result.browseErrors.length === 0,
passed,
...(!passed && isPreTurnInfraFailure(result) ? { failure_class: 'infra' as const } : {}),
duration_ms: result.duration,
cost_usd: result.costEstimate.estimatedCost,
transcript: result.transcript,
+23 -13
View File
@@ -117,6 +117,9 @@ export function isEngBatchingIssueAUQ(fp: AskUserQuestionFingerprint, priorCalls
return !priorCalls.some(prior => batchingIssueNumber(prior) === issue);
}
// The report's target declaration field (Target / Review target / Reviewed target, optionally qualified).
const TARGET_FIELD = /^(?:Reviewed |Review )?target(?: \([^)\n]*\))?:/i;
/** A native brief can use its D number and topic while its stable R identity
* lives in the required saved ledger. Count that owned choice, not a title
* spelling. This does not approve the row or validate the implementation. */
@@ -160,26 +163,31 @@ function recordedBatchingIssue(call: NativePlanQuestionCall, savedPlan: string):
const rawSourceNames = [...(lines[1] ?? '').matchAll(/\b[\w./-]+\.md\b/g)];
const directSource = sourceNames.length > 0 && sourceNames.every(name => name === 'PLAN.md') &&
new Set([...metadata.matchAll(/\bPLAN\.md:([1-9]\d*(?:[-–][1-9]\d*)?)\b/g)].map(match => match[1])).size <= 1;
const targetName = (s: string) => clean(s).replace(/^Eng(?:ineering)? review:\s*/i, '')
const targetName = (s: string) => clean(s).replace(/^Eng(?:ineering)? review(?: report)?\s*[:—–-]\s*/i, '')
.replace(/^Plan\s*[:—–-]\s*/i, '').toLowerCase();
const named = [...(lines[1] ?? '').matchAll(/"(Plan:\s*[^"\n]+)"|“(Plan:\s*[^”\n]+)”/g)]
.map(match => targetName(match[1] ?? match[2]!));
const named = [...(lines[1] ?? '').matchAll(/"(Plan:\s*[^"\n]+)"|“(Plan:\s*[^”\n]+)”|\b[Pp]lan\s+"([^"\n]+)"|\b[Pp]lan\s+“([^”\n]+)”/g)]
.map(match => targetName(match[1] ?? match[2] ?? match[3] ?? match[4]!));
const titles = tokens.slice(0, start).filter(token => token.type === 'heading' && token.depth === 1);
// Target declarations are fields, whatever their list or emphasis markup.
const targetFields = tokens.slice(0, start).flatMap((token, at) => {
if (token.type !== 'paragraph' || !currentHeading(at)) return [];
if ((token.type !== 'paragraph' && token.type !== 'list') || !currentHeading(at)) return [];
const previous = tokens.slice(0, at).filter(t => t.type !== 'space').at(-1);
const quotedContext = /\b(?:quoted|copied|historical|example|hypothetical|archived)\b[^\n]*:\s*$/i;
if (previous?.type === 'paragraph' && quotedContext.test(previous.raw)) return [];
const parts = token.raw.split('\n');
return parts.filter((line, i) => /^Reviewed target:/.test(line) &&
const parts = token.raw.split('\n').map(line => line.replace(/^\s*(?:[-*+]|\d+[.)])\s+/, '').replace(/[*_]/g, '').trim());
return parts.filter((line, i) => TARGET_FIELD.test(line) &&
!parts.slice(0, i).some(part => quotedContext.test(part)));
});
const namedSource = !rawSourceNames.length && named.length === 1 && titles.length === 1 &&
const targetFiles = targetFields.length === 1 ? [...targetFields[0]!.matchAll(/[\w./-]*[\w-]+\.md\b/g)].map(match => match[0]) : [];
// An unsourced brief inherits the report's one current PLAN.md target; its
// ledger record still supplies the cited finding. A brief that names its plan
// must name the report title's plan, and an unfenced copy of that plan may
// add its own H1 only when it names that same plan.
const namedSource = !rawSourceNames.length && named.length <= 1 && titles.length >= 1 &&
titles[0]!.type === 'heading' && currentHeading(tokens.indexOf(titles[0]!)) &&
/^Eng(?:ineering)? review:\s*Plan\s*[:—–-]/i.test(clean(titles[0]!.text)) &&
targetName(titles[0]!.text) === named[0] && targetFields.length === 1 &&
/^Reviewed target:\s*`?PLAN\.md`?(?:\s|$)/.test(targetFields[0]!) &&
[...targetFields[0]!.matchAll(/\b[\w./-]+\.md\b/g)].length === 1;
targetFiles.length === 1 && targetFiles[0]!.split('/').at(-1) === 'PLAN.md' &&
(named.length === 0 || /^Eng(?:ineering)? review(?: report)?\s*[:—–-]\s*\S/i.test(clean(titles[0]!.text)) &&
titles.every(title => title.type === 'heading' && targetName(title.text) === named[0]));
if (!directSource && !namedSource) return;
const withdrawn = (value: string, owners: string) => new RegExp(
`(?:^|[.!?;]\\s+|\\n)(?:Correction:\\s*)?(?:${owners}) (?:is|was|has been) ["“'‘]?(?:withdrawn|cancelled|canceled|rejected|superseded|resolved|closed|hypothetical|not current|no longer current)\\b`, 'i').test(prose(value, true));
@@ -265,7 +273,7 @@ function recordedBatchingIssue(call: NativePlanQuestionCall, savedPlan: string):
if (questions.length !== 1) continue;
const inlineBrief = fields[questions[0]!]!.slice(marker.length).trim();
const inline = Boolean(inlineBrief);
if (!inline && (namedSource || clean(fields[questions[0]! + 1] ?? '') !== clean(title))) continue;
if (!inline && clean(fields[questions[0]! + 1] ?? '') !== clean(title)) continue;
const sources = [...finding[0]!.matchAll(/\b([\w./-]+\.md)(?::([1-9]\d*(?:[-–][1-9]\d*)?))?\b/g)];
if (sources.length !== 1 || sources[0]![1] !== 'PLAN.md' ||
!inline && !sources[0]![2] || source && sources[0]![2] !== source) continue;
@@ -316,6 +324,8 @@ function recordedBatchingIssue(call: NativePlanQuestionCall, savedPlan: string):
const explicitSelectors = nativeOptions.flatMap(option => option.selector ? [option.selector] : []);
if (new Set(explicitSelectors).size !== explicitSelectors.length ||
nativeOptions.some(option => option.selector && !selectors.includes(option.selector) || !option.label || /^[A-D][).:]\s+/.test(option.label))) continue;
// The preamble's `(recommended)` suffix marks the recommendation; it is not part of the choice.
const unmarked = (label: string) => clean(label).replace(/\s*\(recommended\)$/i, '');
const readOptions = (lines: string[]) => {
const records: Array<{ selector: string; label: string; description: string[] }> = [];
for (const line of lines) {
@@ -327,7 +337,7 @@ function recordedBatchingIssue(call: NativePlanQuestionCall, savedPlan: string):
if (records.length !== nativeOptions.length || new Set(records.map(record => record.selector)).size !== records.length ||
records.some(record => !selectors.includes(record.selector))) return undefined;
const matches = records.map(record => nativeOptions.flatMap((native, at) =>
(!native.selector || native.selector === record.selector) && clean(record.label) === native.label &&
(!native.selector || native.selector === record.selector) && unmarked(record.label) === unmarked(native.label) &&
clean(record.description.join('\n')) === clean(native.description) ? [at] : []));
return matches.every(match => match.length === 1) && new Set(matches.flat()).size === records.length ? records : undefined;
};
+42 -31
View File
@@ -46,9 +46,18 @@ export const ALL_TIERS = {
/** Supervision reserve added to every registered whole-file wall. */
export const SHARD_RESERVE_MS = 2 * 60_000;
/** Whole-file supervision must cover each existing attempt and its retry.
* These fixtures already allow 25 minutes per case; the old 30-minute
* wall could kill a second attempt after five minutes. No case budget grows.
/**
* Retry policy (approved 2026-09-29, eval reliability wave): paid evals never
* retry. Each case's kind (E2E_KINDS) fixes its trials before the run: `rule`
* one trial, `behavior` a panel of EVAL_POLICY.panel independent trials, and
* `judge` one case that samples its judge panel internally. A failed verdict
* is final for that run; a manual re-run adds trials under a new run attempt
* and never replaces the original verdict. Rows below keep only wall
* supervision; per-case budgets never change with this rule.
*/
/** Whole-file supervision for one run of every case.
* These fixtures allow 25 minutes per case.
* Reserve the sequential upper bound even when Bun runs sibling cases together.
*/
export const FINDING_RETRY_BUDGETS = [
@@ -58,57 +67,59 @@ export const FINDING_RETRY_BUDGETS = [
file, cases,
id: `${file.slice('test/skill-e2e-'.length, -'.test.ts'.length)}-existing-retry-v1`,
testMs: 1_500_000,
retries: 1,
caseMs: 1_500_000,
shardReserveMs: SHARD_RESERVE_MS,
shardMs: cases * 1_500_000 * 2 + SHARD_RESERVE_MS,
shardMs: cases * 1_500_000 + SHARD_RESERVE_MS,
}));
/** Three existing captures and one configured retry; only supervision grows. */
/** Three existing captures in one 16-minute case. */
export const AUQ_CONSISTENCY_RETRY_BUDGET = {
file: 'test/skill-e2e-auq-consistency.test.ts',
id: 'auq-consistency-existing-retry-v1',
cases: 1,
testMs: 3 * CAPTURE_MS + 60_000,
retries: 1,
caseMs: 3 * CAPTURE_MS + 60_000,
shardReserveMs: SHARD_RESERVE_MS,
shardMs: (3 * CAPTURE_MS + 60_000) * 2 + SHARD_RESERVE_MS,
shardMs: 3 * CAPTURE_MS + 60_000 + SHARD_RESERVE_MS,
} as const;
/** These fixtures have a fixed case count in every supported tier. */
export const STRICT_RETRY_CASE_BUDGETS = [...FINDING_RETRY_BUDGETS, AUQ_CONSISTENCY_RETRY_BUDGET];
/** Whole-file walls cover all existing cases and retries, even if Bun runs them
/** Whole-file walls cover all existing cases, even if Bun runs them
* sequentially. Mixed-tier files reserve their larger complete tier, never a
* currently selected subset. These rows add no case-count or model-work policy.
* The 10-second terms preserve the existing Codex/recording finalization grace.
* currently selected subset. caseMs is the longest single case budget, the
* wall of one isolated case shard. These rows add no case-count or model-work
* policy. The 10-second terms preserve the existing Codex/recording
* finalization grace.
*/
export const FILE_RETRY_BUDGETS = [
...STRICT_RETRY_CASE_BUDGETS,
...[
{ file: 'test/skill-e2e-qa-callers.test.ts', attemptMs: 5 * (CAPTURE_MS + 15_000), retries: 1 },
{ file: 'test/skill-e2e-shared-libs-paths.test.ts', attemptMs: 3 * CAPTURE_LONG_MS, retries: 1 },
{ file: 'test/skill-e2e-ship-docsync.test.ts', attemptMs: 5 * CAPTURE_LONG_MS + 8 * CAPTURE_MS, retries: 1 },
{ file: 'test/skill-e2e-qa-callers.test.ts', attemptMs: 5 * (CAPTURE_MS + 15_000), caseMs: CAPTURE_MS + 15_000 },
{ file: 'test/skill-e2e-shared-libs-paths.test.ts', attemptMs: 3 * CAPTURE_LONG_MS, caseMs: CAPTURE_LONG_MS },
{ file: 'test/skill-e2e-ship-docsync.test.ts', attemptMs: 4 * CAPTURE_LONG_MS + 8 * CAPTURE_MS, caseMs: CAPTURE_LONG_MS },
// Seventeen workflow judges include their 10s recording grace; the other
// seven judges retain 120s. Supervise all 24 and the existing one retry.
{ file: 'test/skill-llm-eval.test.ts', attemptMs: 17 * (JUDGE_MS + 10_000) + 7 * JUDGE_MS, retries: 1 },
{ file: 'test/skill-e2e-auq-matrix.test.ts', attemptMs: 6 * CAPTURE_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-format.test.ts', attemptMs: 4 * (CAPTURE_MS + 10_000), retries: 1 },
{ file: 'test/skill-e2e-auto-decide-preserved.test.ts', attemptMs: PTY_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-ceo-finding-floor.test.ts', attemptMs: PTY_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-eng-finding-floor.test.ts', attemptMs: PTY_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-design-finding-floor.test.ts', attemptMs: PTY_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-devex-finding-floor.test.ts', attemptMs: PTY_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-mode-no-op.test.ts', attemptMs: 5 * CAPTURE_LONG_MS, retries: 2 },
{ file: 'test/skill-e2e-plan-ceo-mode-routing.test.ts', attemptMs: 2 * CAPTURE_LONG_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-eng-plan-mode.test.ts', attemptMs: 2 * CAPTURE_LONG_MS, retries: 1 },
{ file: 'test/skill-e2e-plan-prosons.test.ts', attemptMs: 4 * (CAPTURE_MS + 10_000), retries: 1 },
// seven judges retain 120s. Supervise all 24.
{ file: 'test/skill-llm-eval.test.ts', attemptMs: 17 * (JUDGE_MS + 10_000) + 7 * JUDGE_MS, caseMs: JUDGE_MS + 10_000 },
{ file: 'test/skill-e2e-auq-matrix.test.ts', attemptMs: 6 * CAPTURE_MS, caseMs: CAPTURE_MS },
{ file: 'test/skill-e2e-plan-format.test.ts', attemptMs: 4 * (CAPTURE_MS + 10_000), caseMs: CAPTURE_MS + 10_000 },
{ file: 'test/skill-e2e-auto-decide-preserved.test.ts', attemptMs: PTY_MS, caseMs: PTY_MS },
{ file: 'test/skill-e2e-plan-ceo-finding-floor.test.ts', attemptMs: PTY_MS, caseMs: PTY_MS },
{ file: 'test/skill-e2e-plan-eng-finding-floor.test.ts', attemptMs: PTY_MS, caseMs: PTY_MS },
{ file: 'test/skill-e2e-plan-design-finding-floor.test.ts', attemptMs: PTY_MS, caseMs: PTY_MS },
{ file: 'test/skill-e2e-plan-devex-finding-floor.test.ts', attemptMs: PTY_MS, caseMs: PTY_MS },
{ file: 'test/skill-e2e-plan-mode-no-op.test.ts', attemptMs: 5 * CAPTURE_LONG_MS, caseMs: CAPTURE_LONG_MS },
{ file: 'test/skill-e2e-plan-ceo-mode-routing.test.ts', attemptMs: 2 * CAPTURE_LONG_MS, caseMs: CAPTURE_LONG_MS },
{ file: 'test/skill-e2e-plan-eng-plan-mode.test.ts', attemptMs: 2 * CAPTURE_LONG_MS, caseMs: CAPTURE_LONG_MS },
{ file: 'test/skill-e2e-plan-prosons.test.ts', attemptMs: 4 * (CAPTURE_MS + 10_000), caseMs: CAPTURE_MS + 10_000 },
// Gate: six 300s cases + one 610s case; periodic: two 900s + three 600s.
{ file: 'test/skill-e2e-plan.test.ts', attemptMs: Math.max(6 * CAPTURE_MS + CAPTURE_LONG_MS + 10_000, 2 * PTY_MS + 3 * CAPTURE_LONG_MS), retries: 1 },
].map(({ file, attemptMs, retries }) => ({
file, attemptMs, retries,
{ file: 'test/skill-e2e-plan.test.ts', attemptMs: Math.max(6 * CAPTURE_MS + CAPTURE_LONG_MS + 10_000, 2 * PTY_MS + 3 * CAPTURE_LONG_MS), caseMs: PTY_MS },
].map(({ file, attemptMs, caseMs }) => ({
file, attemptMs, caseMs,
id: `${file.slice('test/'.length, -'.test.ts'.length)}-existing-retry-v1`,
shardReserveMs: SHARD_RESERVE_MS,
shardMs: attemptMs * (retries + 1) + SHARD_RESERVE_MS,
shardMs: attemptMs + SHARD_RESERVE_MS,
})),
];
+201 -1
View File
@@ -14,8 +14,20 @@ import {
formatComparison,
generateCommentary,
judgePassed,
CONTRACT_VIOLATIONS_FILE,
ContractViolation,
TRIAL_ENV,
TRIAL_OUTCOME_SCHEMA,
expectContract,
failureClassOf,
formatTrialOutcomes,
panelVerdict,
parseTrialOutcomes,
sanitizeTrialError,
trialContextFromEnv,
} from './eval-store';
import type { EvalResult, EvalTestEntry, ComparisonResult } from './eval-store';
import type { EvalResult, EvalTestEntry, ComparisonResult, PanelTrial, TrialOutcomeRecord } from './eval-store';
import { EVAL_POLICY } from './periodic-exclude-data';
import { manualReviewFixture } from './manual-judge-review-fixture';
let tmpDir: string;
@@ -957,3 +969,191 @@ describe('generateCommentary', () => {
expect(notes.some(n => n.includes('Stable run'))).toBe(true);
});
});
// --- Trials, panel verdicts and contract vetoes (eval reliability policy) ---
const PANEL = EVAL_POLICY.panel;
const pass = (trial: number, extra: Partial<PanelTrial> = {}): PanelTrial => ({ trial, outcome: 'passed', ...extra });
const fail = (trial: number, extra: Partial<PanelTrial> = {}): PanelTrial => ({ trial, outcome: 'failed', ...extra });
const behavior = (trials: PanelTrial[], quarantined = false) =>
panelVerdict({ case: 'case-x', kind: 'behavior', panel: PANEL, trials, quarantined });
describe('panelVerdict', () => {
test('policy constants are the approved pre-registration', () => {
expect(EVAL_POLICY.panel).toEqual({ n: 3, k: 2 });
expect(EVAL_POLICY.quarantine).toEqual({
entry: { rate: 0.95, minTrials: 10 },
exit: { rate: 0.97, minTrials: 10 },
capFraction: 0.1,
expiryWeeklyRuns: 8,
});
expect(EVAL_POLICY.infraRedispatch).toBe(1);
});
test('rule: one trial, any failure fails the lane', () => {
const ok = panelVerdict({ case: 'r', kind: 'rule', panel: { n: 1, k: 1 }, trials: [pass(1)] });
expect(ok).toMatchObject({ status: 'PASS', split: false, failsLane: false, coverage: true, marks: '✓' });
const bad = panelVerdict({ case: 'r', kind: 'rule', panel: { n: 1, k: 1 }, trials: [fail(1)] });
expect(bad).toMatchObject({ status: 'FAIL', failsLane: true, coverage: false, redClass: 'VERDICT', marks: '✗' });
});
test('behavior 3/3 is a clean PASS', () => {
expect(behavior([pass(1), pass(2), pass(3)])).toMatchObject({ status: 'PASS', split: false, passed: 3, failsLane: false });
});
test('behavior 2/3 is a split PASS that shows its failed trial', () => {
const v = behavior([pass(1), fail(2, { exit_reason: 'timeout' }), pass(3)]);
expect(v).toMatchObject({ status: 'PASS', split: true, passed: 2, failed: 1, failsLane: false, coverage: true, marks: '✓✗✓', reason: 'PASS 2/3' });
expect(v.trials[1].exit_reason).toBe('timeout');
});
test('behavior 1/3 and 0/3 fail the lane', () => {
expect(behavior([pass(1), fail(2), fail(3)])).toMatchObject({ status: 'FAIL', failsLane: true, redClass: 'VERDICT' });
expect(behavior([fail(1), fail(2), fail(3)])).toMatchObject({ status: 'FAIL', failsLane: true });
});
test('a contract trial fails the panel even at 2/3', () => {
const v = behavior([pass(1), pass(2), fail(3, { failure_class: 'contract' })]);
expect(v).toMatchObject({ status: 'FAIL', contract: true, failsLane: true, redClass: 'VERDICT', reason: 'contract violation' });
});
test('a missing trial is INCOMPLETE and fails the lane', () => {
const v = behavior([pass(1), pass(3)]);
expect(v).toMatchObject({ status: 'INCOMPLETE', failsLane: true, coverage: false, redClass: 'INCOMPLETE', marks: '✓·✓' });
expect(v.reason).toContain('missing trial t2');
});
test('duplicate or out-of-range trial records are INCOMPLETE, never deduplicated', () => {
expect(behavior([pass(1), pass(2), pass(2), fail(3)]).status).toBe('INCOMPLETE');
expect(behavior([pass(1), pass(2), pass(3), pass(4)]).reason).toContain('unexpected trial t4');
const v = panelVerdict({ case: 'c', kind: 'behavior', panel: { n: 11, k: 6 }, trials: Array.from({ length: 10 }, (_, i) => pass(i + 2)) });
expect(v.reason).toContain('missing trial t1');
});
test('timeout and infra trials count as failed, never passing', () => {
const v = behavior([pass(1), fail(2, { exit_reason: 'timeout' }), fail(3, { failure_class: 'infra' })]);
expect(v).toMatchObject({ status: 'FAIL', passed: 1, failed: 2, failsLane: true, redClass: 'VERDICT' });
const infra = behavior([pass(1), fail(2, { failure_class: 'infra' }), fail(3, { failure_class: 'infra' })]);
expect(infra).toMatchObject({ status: 'FAIL', redClass: 'INFRA' });
expect(failureClassOf({ exit_reason: 'timeout' })).toBe('timeout');
expect(failureClassOf({})).toBe('assertion');
});
test('quarantined: 1/3 does not fail the lane, 0/3 and contract do, no coverage credit', () => {
expect(behavior([pass(1), fail(2), fail(3)], true)).toMatchObject({ status: 'FAIL', failsLane: false, coverage: false, redClass: null });
expect(behavior([fail(1), fail(2), fail(3)], true)).toMatchObject({ status: 'FAIL', failsLane: true });
expect(behavior([pass(1), pass(2), fail(3, { failure_class: 'contract' })], true)).toMatchObject({ status: 'FAIL', failsLane: true });
expect(behavior([pass(1), pass(2), pass(3)], true)).toMatchObject({ status: 'PASS', coverage: false, failsLane: false });
expect(behavior([pass(1), pass(2)], true)).toMatchObject({ status: 'INCOMPLETE', failsLane: true });
});
test('quarantined rule keeps rule meaning (k = n)', () => {
const v = panelVerdict({ case: 'r', kind: 'rule', panel: { n: 3, k: 3 }, trials: [pass(1), pass(2), fail(3)], quarantined: true });
expect(v).toMatchObject({ status: 'FAIL', failsLane: false });
});
test('all-skipped panel is SKIPPED with no credit; partly skipped is INCOMPLETE', () => {
const skip = (trial: number): PanelTrial => ({ trial, outcome: 'skipped' });
expect(behavior([skip(1), skip(2), skip(3)])).toMatchObject({ status: 'SKIPPED', coverage: false, failsLane: false });
expect(behavior([pass(1), pass(2), skip(3)])).toMatchObject({ status: 'INCOMPLETE', failsLane: true });
});
test('trials of different run attempts are never merged into one verdict', () => {
expect(() => behavior([pass(1), pass(2), fail(3, { attempt: 2 })])).toThrow(/run attempts/);
expect(behavior([pass(1, { attempt: 2 }), pass(2, { attempt: 2 }), pass(3, { attempt: 2 })]).attempt).toBe(2);
});
test('invalid panels throw', () => {
expect(() => panelVerdict({ case: 'c', kind: 'behavior', panel: { n: 3, k: 4 }, trials: [] })).toThrow(/invalid panel/);
expect(() => panelVerdict({ case: 'c', kind: 'nope' as any, panel: { n: 1, k: 1 }, trials: [] })).toThrow(/unknown kind/);
});
});
describe('trial context and expectContract', () => {
let dir: string;
const saved: Record<string, string | undefined> = {};
const keys = [...Object.values(TRIAL_ENV), 'GSTACK_EVAL_DIR'];
beforeEach(() => {
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'panel-verdict-'));
for (const key of keys) saved[key] = process.env[key];
});
afterEach(() => {
for (const key of keys) {
if (saved[key] === undefined) delete process.env[key];
else process.env[key] = saved[key];
}
fs.rmSync(dir, { recursive: true, force: true });
});
const setTrial = () => Object.assign(process.env, {
[TRIAL_ENV.caseId]: 'case-x', [TRIAL_ENV.kind]: 'behavior', [TRIAL_ENV.trial]: '2',
[TRIAL_ENV.panelN]: '3', [TRIAL_ENV.panelK]: '2', [TRIAL_ENV.policyVersion]: String(EVAL_POLICY.version),
GSTACK_EVAL_DIR: dir,
});
test('trialContextFromEnv: absent, complete, and malformed', () => {
for (const key of Object.values(TRIAL_ENV)) delete process.env[key];
expect(trialContextFromEnv()).toBeNull();
setTrial();
expect(trialContextFromEnv()).toEqual({ case_id: 'case-x', kind: 'behavior', trial: 2, panel: { n: 3, k: 2 }, policy_version: EVAL_POLICY.version });
process.env[TRIAL_ENV.trial] = '4';
expect(() => trialContextFromEnv()).toThrow(/Malformed trial context/);
});
test('passing contract is a no-op', () => {
setTrial();
expect(() => expectContract(true, 'fine')).not.toThrow();
expect(fs.existsSync(path.join(dir, CONTRACT_VIOLATIONS_FILE))).toBe(false);
});
test('failed contract stamps the recorded entry and the sidecar before throwing', () => {
setTrial();
const collector = new EvalCollector('e2e', dir);
collector.addTest({ name: 'case-x', suite: 's', tier: 'e2e', passed: true, duration_ms: 1, cost_usd: 0 });
expect(() => expectContract(false, 'handoff missing', { collector, name: 'case-x' })).toThrow(ContractViolation);
const partial = JSON.parse(fs.readFileSync(path.join(dir, '_partial-e2e.json'), 'utf-8'));
expect(partial.tests[0]).toMatchObject({ passed: false, failure_class: 'contract', case_id: 'case-x', trial: 2, kind: 'behavior', panel: { n: 3, k: 2 } });
const sidecar = fs.readFileSync(path.join(dir, CONTRACT_VIOLATIONS_FILE), 'utf-8').trim().split('\n').map((l) => JSON.parse(l));
expect(sidecar).toEqual([expect.objectContaining({ case_id: 'case-x', trial: 2, message: 'handoff missing' })]);
});
test('a contract marked before recording stamps the later record, or becomes its own at finalize', async () => {
setTrial();
const collector = new EvalCollector('e2e', dir);
expect(() => expectContract(0, 'no question asked', { collector, name: 'later' })).toThrow('CONTRACT: no question asked');
collector.addTest({ name: 'later', suite: 's', tier: 'e2e', passed: true, duration_ms: 1, cost_usd: 0 });
expect(() => expectContract(null, 'never recorded', { collector, name: 'orphan' })).toThrow();
const file = await collector.finalize();
const tests = JSON.parse(fs.readFileSync(file, 'utf-8')).tests;
expect(tests.find((t: any) => t.name === 'later')).toMatchObject({ passed: false, failure_class: 'contract' });
expect(tests.find((t: any) => t.name === 'orphan')).toMatchObject({ passed: false, failure_class: 'contract', error: 'never recorded' });
});
});
describe('trial-outcomes JSONL', () => {
const record = (extra: Partial<TrialOutcomeRecord> = {}): TrialOutcomeRecord => ({
schema: TRIAL_OUTCOME_SCHEMA, case: 'case-x', file: 'test/x.test.ts', tier: 'gate', kind: 'behavior',
trial: 1, panel: { n: 3, k: 2 }, attempt: 1, outcome: 'passed', duration_ms: 10, cost_usd: 0.1,
policy_version: EVAL_POLICY.version, quarantined: false, execution: 'executed', source: 'shard', ...extra,
});
test('round-trips valid records', () => {
const records = [record(), record({ trial: 2, outcome: 'failed', failure_class: 'timeout', exit_reason: 'timeout' })];
expect(parseTrialOutcomes(formatTrialOutcomes(records))).toEqual({ records, errors: [] });
});
test('writer fails closed; reader reports bad lines as data errors', () => {
expect(() => formatTrialOutcomes([record({ outcome: 'failed' })])).toThrow(/failed without failure_class/);
expect(() => formatTrialOutcomes([record({ trial: 4 })])).toThrow(/trial invalid/);
const text = `${JSON.stringify(record())}\nnot json\n${JSON.stringify({ ...record(), schema: 'other' })}\n`;
const parsed = parseTrialOutcomes(text);
expect(parsed.records).toHaveLength(1);
expect(parsed.errors).toEqual(['line 2: not JSON', 'line 3: schema other']);
expect(parseTrialOutcomes(text, { maxBytes: 10 }).errors[0]).toContain('exceed');
});
test('sanitizeTrialError keeps one capped line without mentions', () => {
expect(sanitizeTrialError('\n expected @garrytan to `see`\nsecond')).toBe("expected @\u200bgarrytan to 'see'");
expect(sanitizeTrialError('x'.repeat(1000))!.length).toBe(300);
expect(sanitizeTrialError('')).toBeUndefined();
});
});
+367 -1
View File
@@ -76,6 +76,18 @@ export interface EvalTestEntry {
* its body again and re-records under the same name. Set by addTest. */
attempt?: number;
// Trial identity (eval reliability policy). Stamped by addTest from the
// TRIAL_ENV variables the paid runner sets on an isolated trial shard.
/** Registry id (E2E_TIERS / LLM_JUDGE_TOUCHFILES key) this record belongs to. */
case_id?: string;
kind?: EvalCaseKind;
/** 1-based trial index within the case's panel. */
trial?: number;
panel?: PanelShape;
/** Why a failed record failed; 'contract' comes only from expectContract. */
failure_class?: TrialFailureClass;
policy_version?: number;
// E2E
transcript?: any[];
prompt?: string;
@@ -131,6 +143,329 @@ export function evalEntryOutcome(entry: unknown): 'passed' | 'failed' | 'manual-
return result.passed === true ? 'passed' : 'failed';
}
// --- Trials and panel verdicts ---
//
// Paid evals never retry. Each case's kind (E2E_KINDS) fixes its trials before
// the run; a panel verdict is computed once, by panelVerdict(), from exactly
// panel.n trial records of one run attempt. The report, collector-outcomes,
// the PR comment and pass-rates all read that one function.
export type EvalCaseKind = 'rule' | 'behavior' | 'judge';
/** assertion: an ordinary failed expectation. contract: expectContract() fired
* (fails the panel at any count). timeout: the case budget ran out.
* infra: API/CLI/runner failure before the model could be graded. */
export type TrialFailureClass = 'assertion' | 'contract' | 'timeout' | 'infra';
export type TrialOutcome = 'passed' | 'failed' | 'skipped';
export interface PanelShape { n: number; k: number }
/** Environment the paid runner sets on an isolated trial shard. */
export const TRIAL_ENV = {
caseId: 'GSTACK_EVAL_CASE_ID',
kind: 'GSTACK_EVAL_KIND',
trial: 'GSTACK_EVAL_TRIAL',
panelN: 'GSTACK_EVAL_PANEL_N',
panelK: 'GSTACK_EVAL_PANEL_K',
policyVersion: 'GSTACK_EVAL_POLICY_VERSION',
} as const;
/** Sidecar every expectContract() failure appends to (in GSTACK_EVAL_DIR), so a
* contract veto survives a test that throws before recording its entry. */
export const CONTRACT_VIOLATIONS_FILE = 'contract-violations.jsonl';
export interface TrialContext {
case_id: string;
kind: EvalCaseKind;
trial: number;
panel: PanelShape;
policy_version: number;
}
const EVAL_KINDS: readonly EvalCaseKind[] = ['rule', 'behavior', 'judge'];
const FAILURE_CLASSES: readonly TrialFailureClass[] = ['assertion', 'contract', 'timeout', 'infra'];
function positiveInt(raw: string | undefined): number | null {
if (raw === undefined || !/^[1-9][0-9]*$/.test(raw)) return null;
return Number(raw);
}
/** Trial context of this process, or null outside an isolated trial shard.
* A partial or malformed context throws: a mislabeled trial is fail-open. */
export function trialContextFromEnv(env: NodeJS.ProcessEnv = process.env): TrialContext | null {
const caseId = env[TRIAL_ENV.caseId];
if (!caseId) return null;
const kind = env[TRIAL_ENV.kind] as EvalCaseKind | undefined;
const trial = positiveInt(env[TRIAL_ENV.trial]);
const n = positiveInt(env[TRIAL_ENV.panelN]);
const k = positiveInt(env[TRIAL_ENV.panelK]);
const policy = positiveInt(env[TRIAL_ENV.policyVersion]);
if (!kind || !EVAL_KINDS.includes(kind) || trial === null || n === null || k === null || policy === null || k > n || trial > n) {
throw new Error(`Malformed trial context for ${caseId}: ${Object.values(TRIAL_ENV).map((name) => `${name}=${env[name] ?? ''}`).join(' ')}`);
}
return { case_id: caseId, kind, trial, panel: { n, k }, policy_version: policy };
}
/** Failure class of a failed record: an explicit class wins, then the exit reason. */
export function failureClassOf(entry: Pick<EvalTestEntry, 'failure_class' | 'exit_reason'>): TrialFailureClass {
if (entry.failure_class && FAILURE_CLASSES.includes(entry.failure_class)) return entry.failure_class;
return entry.exit_reason === 'timeout' ? 'timeout' : 'assertion';
}
export class ContractViolation extends Error {
constructor(message: string) {
super(`CONTRACT: ${message}`);
this.name = 'ContractViolation';
}
}
/**
* Assert a contract: an outcome the product must meet on every run. On failure
* it records failure_class 'contract' before throwing, both on the collector
* entry named `record.name` (now or when the test records it) and in the
* GSTACK_EVAL_DIR sidecar, so panelVerdict() fails the panel even at 2 of 3.
*/
export function expectContract(
condition: unknown,
message: string,
record?: { collector: EvalCollector | null; name: string },
): asserts condition {
if (condition) return;
record?.collector?.markContractViolation(record.name, message);
const evalDir = process.env.GSTACK_EVAL_DIR;
if (evalDir) {
const context = trialContextFromEnv();
fs.mkdirSync(evalDir, { recursive: true });
fs.appendFileSync(path.join(evalDir, CONTRACT_VIOLATIONS_FILE), JSON.stringify({
case_id: context?.case_id ?? record?.name ?? null,
name: record?.name ?? null,
trial: context?.trial ?? null,
message,
at: new Date().toISOString(),
}) + '\n');
}
throw new ContractViolation(message);
}
export interface PanelTrial {
trial: number;
outcome: TrialOutcome;
/** Required meaning for a failed trial; absent reads as 'assertion'. */
failure_class?: TrialFailureClass;
/** CI run attempt (github.run_attempt); absent means 1. */
attempt?: number;
exit_reason?: string;
error?: string;
execution?: 'executed' | 'reused';
}
export interface PanelVerdictInput {
case: string;
kind: EvalCaseKind;
panel: PanelShape;
trials: readonly PanelTrial[];
quarantined?: boolean;
}
export type PanelStatus = 'PASS' | 'FAIL' | 'INCOMPLETE' | 'SKIPPED';
export interface PanelVerdict {
case: string;
kind: EvalCaseKind;
panel: PanelShape;
attempt: number;
quarantined: boolean;
status: PanelStatus;
passed: number;
failed: number;
/** A failed trial carried failure_class 'contract'. */
contract: boolean;
/** PASS with at least one failed trial: shown as `PASS k/n`, never clean. */
split: boolean;
/** Whether this verdict makes the lane red. */
failsLane: boolean;
/** Whether it counts as passing coverage (never for quarantined or skipped). */
coverage: boolean;
/** Machine classification of a lane-failing verdict: INCOMPLETE (missing or
* malformed trial records), INFRA (every failed trial is infra-class), or
* VERDICT (a real red). Null when the verdict does not fail the lane. */
redClass: 'INCOMPLETE' | 'INFRA' | 'VERDICT' | null;
/** One glyph per trial index: ✓ pass, ✗ fail, – skipped, · missing. */
marks: string;
reason: string;
trials: PanelTrial[];
}
/**
* The single verdict function. `rule`/`judge` cases run panel {1,1}; `behavior`
* cases run EVAL_POLICY.panel; a quarantined case runs a full panel whose k
* keeps its kind's meaning (k = n for rule). Verdict: INCOMPLETE unless
* exactly one record per trial index 1..n; SKIPPED when every trial skipped;
* FAIL on any contract trial; otherwise PASS iff passed >= k. A quarantined
* FAIL fails the lane only on a hard break (0 of n) or a contract violation.
*/
export function panelVerdict(input: PanelVerdictInput): PanelVerdict {
const { n, k } = input.panel;
if (!Number.isInteger(n) || !Number.isInteger(k) || n < 1 || k < 1 || k > n) {
throw new Error(`${input.case}: invalid panel {n:${n}, k:${k}}`);
}
if (!EVAL_KINDS.includes(input.kind)) throw new Error(`${input.case}: unknown kind ${String(input.kind)}`);
const attempts = new Set(input.trials.map((t) => t.attempt ?? 1));
if (attempts.size > 1) {
throw new Error(`${input.case}: trials from run attempts ${[...attempts].join(', ')}; compute one verdict per attempt`);
}
const attempt = [...attempts][0] ?? 1;
const quarantined = input.quarantined === true;
const trials = [...input.trials].sort((a, b) => a.trial - b.trial);
const byIndex = new Map<number, PanelTrial>();
const problems: string[] = [];
for (const t of trials) {
if (!Number.isInteger(t.trial) || t.trial < 1 || t.trial > n) problems.push(`unexpected trial t${t.trial}`);
else if (byIndex.has(t.trial)) problems.push(`duplicate trial t${t.trial}`);
else if (t.outcome !== 'passed' && t.outcome !== 'failed' && t.outcome !== 'skipped') problems.push(`t${t.trial} has outcome ${String(t.outcome)}`);
else byIndex.set(t.trial, t);
}
for (let i = 1; i <= n; i++) if (!trials.some((t) => t.trial === i)) problems.push(`missing trial t${i}`);
const marks = Array.from({ length: n }, (_, i) => {
const t = byIndex.get(i + 1);
return !t ? '·' : t.outcome === 'passed' ? '✓' : t.outcome === 'failed' ? '✗' : '–';
}).join('');
const passed = [...byIndex.values()].filter((t) => t.outcome === 'passed').length;
const failedTrials = [...byIndex.values()].filter((t) => t.outcome === 'failed');
const skipped = [...byIndex.values()].filter((t) => t.outcome === 'skipped').length;
const contract = failedTrials.some((t) => failureClassOf(t) === 'contract');
const base = { case: input.case, kind: input.kind, panel: { n, k }, attempt, quarantined, passed, failed: failedTrials.length, contract, marks, trials };
if (problems.length === 0 && skipped === n) {
return { ...base, status: 'SKIPPED', split: false, failsLane: false, coverage: false, redClass: null, reason: 'every trial skipped (no verdict credit)' };
}
if (problems.length === 0 && skipped > 0) problems.push(`${skipped} of ${n} trials skipped`);
if (problems.length > 0) {
return { ...base, status: 'INCOMPLETE', split: false, failsLane: true, coverage: false, redClass: 'INCOMPLETE', reason: problems.join('; ') };
}
if (!contract && passed >= k) {
const split = failedTrials.length > 0;
return {
...base, status: 'PASS', split, failsLane: false, coverage: !quarantined, redClass: null,
reason: split ? `PASS ${passed}/${n}` : `${passed}/${n} passed`,
};
}
const hardBreak = passed === 0;
const failsLane = !quarantined || contract || hardBreak;
const allInfra = !contract && failedTrials.length > 0 && failedTrials.every((t) => failureClassOf(t) === 'infra');
const why = contract ? 'contract violation' : `${passed}/${n} passed, needs ${k}`;
return {
...base, status: 'FAIL', split: false, failsLane, coverage: false,
redClass: failsLane ? (allInfra ? 'INFRA' : 'VERDICT') : null,
reason: !quarantined ? why
: contract ? `${why}; quarantine never excuses a contract`
: hardBreak ? `${why}; quarantined hard break`
: `${why}; quarantined, does not fail the lane`,
};
}
// --- trial-outcomes JSONL (one line per trial; pass-rate history input) ---
export const TRIAL_OUTCOME_SCHEMA = 'gstack-trial-outcome/v1';
export const TRIAL_OUTCOMES_FILE = 'trial-outcomes.jsonl';
/** Cap on a stored `error` line (sanitized first line of the failure). */
export const TRIAL_ERROR_MAX = 300;
export interface TrialOutcomeRecord {
schema: typeof TRIAL_OUTCOME_SCHEMA;
/** Registry id. */
case: string;
file: string;
tier: string;
kind: EvalCaseKind;
trial: number;
panel: PanelShape;
/** CI run attempt (github.run_attempt); 1 locally and for pre-policy backfill. */
attempt: number;
outcome: TrialOutcome;
/** Present exactly when outcome is 'failed'. */
failure_class?: TrialFailureClass;
exit_reason?: string;
error?: string;
duration_ms: number;
cost_usd: number;
model?: string;
cli_version?: string;
/** Reuse input key of the trial's shard, when known. */
input_identity?: string;
/** EVAL_POLICY.version; 0 marks pre-policy backfill. */
policy_version: number;
quarantined: boolean;
execution: 'executed' | 'reused';
/** shard: isolated trial shard status. junit: a rule file shard's per-test
* JUnit outcome. backfill: imported pre-policy artifact record. */
source: 'shard' | 'junit' | 'backfill';
run_id?: string;
sha?: string;
lane?: string;
recorded_at?: string;
/** History series key: a hash of the case's own touchfiles (GLOBAL_TOUCHFILES excluded), stamped by the report job. */
series_identity?: string;
}
/** First line of free text, stripped of @-mentions and control characters, capped. */
export function sanitizeTrialError(text: string | undefined): string | undefined {
if (!text) return undefined;
const first = text.split('\n').map((l) => l.trim()).find((l) => l.length > 0);
if (!first) return undefined;
// eslint-disable-next-line no-control-regex
const clean = first.replace(/[\u0000-\u001f\u007f]/g, ' ').replace(/`/g, "'").replace(/@(?=[A-Za-z0-9_-])/g, '@\u200b');
return clean.length > TRIAL_ERROR_MAX ? `${clean.slice(0, TRIAL_ERROR_MAX - 1)}…` : clean;
}
function trialRecordProblems(r: any): string[] {
const problems: string[] = [];
if (!r || typeof r !== 'object' || Array.isArray(r)) return ['not an object'];
if (r.schema !== TRIAL_OUTCOME_SCHEMA) problems.push(`schema ${String(r.schema)}`);
for (const key of ['case', 'file', 'tier'] as const) if (typeof r[key] !== 'string' || r[key].length === 0) problems.push(`${key} missing`);
if (!EVAL_KINDS.includes(r.kind)) problems.push(`kind ${String(r.kind)}`);
const n = r.panel?.n, k = r.panel?.k;
if (!Number.isInteger(n) || !Number.isInteger(k) || n < 1 || k < 1 || k > n) problems.push('panel invalid');
if (!Number.isInteger(r.trial) || r.trial < 1 || (Number.isInteger(n) && r.trial > n)) problems.push('trial invalid');
if (!Number.isInteger(r.attempt) || r.attempt < 1) problems.push('attempt invalid');
if (!['passed', 'failed', 'skipped'].includes(r.outcome)) problems.push(`outcome ${String(r.outcome)}`);
if (r.outcome === 'failed' && !FAILURE_CLASSES.includes(r.failure_class)) problems.push('failed without failure_class');
if (r.outcome !== 'failed' && r.failure_class !== undefined) problems.push('failure_class on a non-failed trial');
if (typeof r.duration_ms !== 'number' || !Number.isFinite(r.duration_ms) || r.duration_ms < 0) problems.push('duration_ms invalid');
if (typeof r.cost_usd !== 'number' || !Number.isFinite(r.cost_usd) || r.cost_usd < 0) problems.push('cost_usd invalid');
if (!Number.isInteger(r.policy_version) || r.policy_version < 0) problems.push('policy_version invalid');
if (typeof r.quarantined !== 'boolean') problems.push('quarantined invalid');
if (r.execution !== 'executed' && r.execution !== 'reused') problems.push('execution invalid');
if (!['shard', 'junit', 'backfill'].includes(r.source)) problems.push('source invalid');
if (r.error !== undefined && (typeof r.error !== 'string' || r.error.length > TRIAL_ERROR_MAX)) problems.push('error invalid');
if (r.series_identity !== undefined && (typeof r.series_identity !== 'string' || !/^[\w.-]{1,64}$/.test(r.series_identity))) problems.push('series_identity invalid');
return problems;
}
/** Serialize records as JSONL; throws on any invalid record (writers fail closed). */
export function formatTrialOutcomes(records: readonly TrialOutcomeRecord[]): string {
return records.map((r) => {
const problems = trialRecordProblems(r);
if (problems.length > 0) throw new Error(`invalid trial record ${r?.case}~t${r?.trial}: ${problems.join(', ')}`);
return JSON.stringify(r);
}).join('\n') + (records.length > 0 ? '\n' : '');
}
/** Parse downloaded JSONL as data only: invalid lines are reported, never guessed. */
export function parseTrialOutcomes(text: string, opts: { maxBytes?: number } = {}): { records: TrialOutcomeRecord[]; errors: string[] } {
const maxBytes = opts.maxBytes ?? 16 * 1024 * 1024;
if (Buffer.byteLength(text) > maxBytes) return { records: [], errors: [`trial outcomes exceed ${maxBytes} bytes`] };
const records: TrialOutcomeRecord[] = [];
const errors: string[] = [];
text.split('\n').forEach((line, i) => {
if (line.trim() === '') return;
let parsed: unknown;
try { parsed = JSON.parse(line); } catch { errors.push(`line ${i + 1}: not JSON`); return; }
const problems = trialRecordProblems(parsed);
if (problems.length > 0) errors.push(`line ${i + 1}: ${problems.join(', ')}`);
else records.push(parsed as TrialOutcomeRecord);
});
return { records, errors };
}
export interface EvalResult {
schema_version: number;
version: string;
@@ -887,6 +1222,7 @@ export class EvalCollector {
private shard: string | null;
private fileNamespace?: string;
private createdAt = Date.now();
private pendingContract = new Map<string, string>();
constructor(tier: 'e2e' | 'llm-judge', evalDir?: string, fileNamespace?: string) {
if (fileNamespace !== undefined && !/^[a-z0-9]+(?:-[a-z0-9]+)*$/.test(fileNamespace)) {
@@ -903,7 +1239,29 @@ export class EvalCollector {
// names are unique by convention). Stamp the 1-based attempt so a
// pass-on-attempt-2 stays visible forever — the stream hides it.
const prior = this.tests.filter((t) => t.name === entry.name).length;
this.tests.push({ ...entry, attempt: prior + 1 });
const context = trialContextFromEnv();
const record: EvalTestEntry = { ...(context ?? {}), ...entry, attempt: prior + 1 };
const contract = this.pendingContract.get(entry.name);
if (contract !== undefined) {
this.pendingContract.delete(entry.name);
Object.assign(record, { passed: false, failure_class: 'contract', error: record.error ?? contract });
}
this.tests.push(record);
this.savePartial();
}
/** expectContract() hook: mark `name`'s latest record (or its next one) as a
* contract failure. An unmatched mark becomes its own failed record at
* finalize, so the veto is never lost. */
markContractViolation(name: string, message: string): void {
const existing = this.tests.filter((t) => t.name === name).at(-1);
if (!existing) {
this.pendingContract.set(name, message);
return;
}
existing.passed = false;
existing.failure_class = 'contract';
existing.error = existing.error ?? message;
this.savePartial();
}
@@ -959,6 +1317,14 @@ export class EvalCollector {
async finalize(): Promise<string> {
if (this.finalized) return '';
this.finalized = true;
for (const [name, message] of this.pendingContract) {
this.tests.push({
...(trialContextFromEnv() ?? {}),
name, suite: 'contract', tier: this.tier, passed: false, duration_ms: 0, cost_usd: 0,
failure_class: 'contract', error: message, attempt: 1,
});
}
this.pendingContract.clear();
const git = getGitInfo();
const version = getVersion();
+11
View File
@@ -18,3 +18,14 @@ export function installFakeImpeccable(prefix = 'gstack-fake-impeccable-'): { dir
fs.copyFileSync(DETECT_SAMPLE, path.join(dir, 'impeccable-detect-sample.json')); // the shim's documented default output, beside it
return { dir, bin };
}
/** Sample rule ids the design checklist never names: a review can carry them only from the detector's rows. */
export function detectorOnlyRuleIds(checklist: string): string[] {
const rules = JSON.parse(fs.readFileSync(DETECT_SAMPLE, 'utf-8')) as Array<{ antipattern: string }>;
return [...new Set(rules.map(rule => rule.antipattern))].filter(id => !checklist.includes(id));
}
export function carriesDetectorRows(review: string, checklist: string): boolean {
const text = review.toLowerCase();
return detectorOnlyRuleIds(checklist).some(id => new RegExp(`(?<![\\w-])${id}(?![\\w-])`).test(text));
}
+4 -3
View File
@@ -3,7 +3,8 @@
*
* Pins three contracts:
* 1. Allowlist semantics: contamination vars dropped, basics/auth/network
* kept, overrides merge last, EVALS_HERMETIC=0 is byte-identical legacy.
* kept, overrides merge last, EVALS_HERMETIC=0 is the legacy env plus the
* DISABLE_AUTOUPDATER pin.
* 2. Seed-config shape: 20-char key suffix, trusted dirs, undefined-key safe.
* 3. Dir lifecycle: /.claude suffix (extractPlanFilePath contract —
* claude-pty-runner.ts:191), sync singleton reuse, pid-aware GC.
@@ -167,7 +168,7 @@ describe('buildHermeticEnv allowlist', () => {
});
describe('EVALS_HERMETIC=0 escape hatch', () => {
test('returns byte-identical legacy env, overrides still last', () => {
test('returns the legacy env plus the updater pin, overrides still last', () => {
const base = { ...CONTAMINATED, EVALS_HERMETIC: '0' } as NodeJS.ProcessEnv;
const e = buildHermeticEnv(base, HERMETIC_VARS, { GSTACK_HEADLESS: '1' });
// Legacy spread: every base var survives, hermeticVars NOT applied.
@@ -175,7 +176,7 @@ describe('EVALS_HERMETIC=0 escape hatch', () => {
expect(e.CLAUDE_CONFIG_DIR).toBe('/Users/op/.claude');
expect(e.GSTACK_HOME).toBe('/Users/op/.gstack');
expect(e.GSTACK_HEADLESS).toBe('1');
expect(e).toEqual({ ...(base as Record<string, string>), GSTACK_HEADLESS: '1' });
expect(e).toEqual({ ...(base as Record<string, string>), DISABLE_AUTOUPDATER: '1', GSTACK_HEADLESS: '1' });
});
test('isHermeticEnabled reads at call time (ESM-hoist safety)', () => {
+9 -3
View File
@@ -21,11 +21,15 @@
* └─────────────────────────────┘
* + per-runner extraAllow (codex: OpenAI vars; gemini: Google vars)
* + CLAUDE_CONFIG_DIR=<runRoot>/.claude GSTACK_HOME=<runRoot>/gstack-home
* + DISABLE_AUTOUPDATER=1 (pinned in both branches; the scrub drops the
* workflow's copy and every PTY screen otherwise shows the updater's
* "no write permission to npm prefix" failure)
* + per-test overrides spread LAST
*
* Escape hatch: EVALS_HERMETIC=0 restores the legacy contaminated env
* byte-identically (runners must also gate --strict-mcp-config on
* isHermeticEnabled() so the escape hatch restores args too).
* plus only the DISABLE_AUTOUPDATER pin (runners must also gate
* --strict-mcp-config on isHermeticEnabled() so the escape hatch restores
* args too).
*
* isHermeticEnabled() is evaluated at CALL time, never at module load —
* ESM hoists imports above any in-file `process.env.EVALS_HERMETIC = '0'`
@@ -100,9 +104,10 @@ export function buildHermeticEnv(
opts?: HermeticEnvOpts,
): Record<string, string> {
if (!isHermeticEnabled(base)) {
// Escape hatch: byte-identical to the legacy spread.
// Escape hatch: the legacy spread plus the updater pin.
const legacy: Record<string, string> = {};
for (const [k, v] of Object.entries(base)) if (v !== undefined) legacy[k] = v;
legacy.DISABLE_AUTOUPDATER = '1';
for (const [k, v] of Object.entries(overrides ?? {})) if (v !== undefined) legacy[k] = v;
return legacy;
}
@@ -127,6 +132,7 @@ export function buildHermeticEnv(
if (allowed) out[k] = v;
}
if (!out.TERM) out.TERM = 'xterm-256color';
out.DISABLE_AUTOUPDATER = '1';
Object.assign(out, hermeticVars);
for (const [k, v] of Object.entries(overrides ?? {})) if (v !== undefined) out[k] = v;
return out;
+115 -24
View File
@@ -23,6 +23,8 @@ export interface JudgeScore {
reasoning: string;
}
export const JUDGE_SCORE_DIMENSIONS = ['clarity', 'completeness', 'actionability'] as const;
export interface JudgeRefusalEvidence {
stop_reason: 'refusal';
response_id: string | null;
@@ -102,6 +104,8 @@ export interface CallJudgeOptions {
signal?: AbortSignal;
/** Opt-in serialization contract; callers still validate the judgment locally. */
jsonSchema?: JSONOutputFormat['schema'];
/** Adaptive-thinking effort; the judge models accept no thinking token budget. */
effort?: 'low' | 'medium' | 'high';
}
export async function callJudge<T>(
@@ -127,7 +131,9 @@ export async function callJudge<T>(
model: resolvedModel,
max_tokens: maxTokens,
...(opts?.temperature !== undefined ? { temperature: opts.temperature } : {}),
...(opts?.jsonSchema === undefined ? {} : { output_config: { format: { type: 'json_schema' as const, schema: opts.jsonSchema } } }),
...(opts?.jsonSchema === undefined && opts?.effort === undefined ? {} : { output_config: {
...(opts?.jsonSchema === undefined ? {} : { format: { type: 'json_schema' as const, schema: opts.jsonSchema } }),
...(opts?.effort === undefined ? {} : { effort: opts.effort }) } }),
messages: [{ role: 'user' as const, content: prompt }],
};
const makeRequest = () => opts?.stream
@@ -196,6 +202,92 @@ export async function callJudge<T>(
}
}
/**
* Samples per judge panel: EVAL_POLICY.judge.samples, restated here so this
* helper (imported by many paid tests) does not pull the quarantine registry
* into their touchfile closure. test/judge-panel.test.ts pins the two equal.
*/
export const JUDGE_PANEL_SAMPLES = 3;
/**
* Judge panel (EVAL_POLICY.judge): every `judge`-kind entry draws a fixed number of
* independent samples of the SAME prompt concurrently, inside its unchanged
* JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean
* against the unchanged minimum; boolean fields gate on a strict majority.
* A sample that errors (refusal, truncation, non-JSON, malformed field) fails
* the whole panel and is never resampled. callJudge's 429 backoff happens
* before any model output exists, so it is transport, not a verdict retry.
*/
export async function judgePanel<T>(sample: () => Promise<T>): Promise<T[]> {
const settled = await Promise.allSettled(Array.from({ length: JUDGE_PANEL_SAMPLES }, () => sample()));
const failures = settled.flatMap((result, index) => result.status === 'rejected' ? [{ index, reason: result.reason }] : []);
if (failures.length === 0) return settled.map(result => (result as PromiseFulfilledResult<T>).value);
const first = failures[0]!;
// A refusal is an unscored panel only when EVERY sample refused; a partial
// refusal beside scored samples is an ordinary failed panel.
if (first.reason instanceof JudgeRefusalError && failures.length < settled.length) {
throw new Error(`Judge panel sample ${first.index + 1} of ${settled.length} failed beside scored samples: ${first.reason.message}`);
}
throw first.reason;
}
/** Per-dimension mean over a panel; any non-finite sample value fails the panel. */
export function judgePanelMean<K extends string>(samples: ReadonlyArray<Record<K, unknown>>, keys: readonly K[]): Record<K, number> {
if (samples.length === 0) throw new Error('Judge panel has no samples');
return Object.fromEntries(keys.map(key => {
const values = samples.map(sample => sample && typeof sample === 'object' ? sample[key] : undefined);
const bad = values.findIndex(value => typeof value !== 'number' || !Number.isFinite(value));
if (bad !== -1) throw new Error(`Judge panel sample ${bad + 1} has non-numeric ${key}: ${JSON.stringify(values[bad])}`);
return [key, (values as number[]).reduce((sum, value) => sum + value, 0) / values.length];
})) as Record<K, number>;
}
/** Strict majority of a boolean field; any non-boolean sample value fails the panel. */
export function judgePanelMajority<K extends string>(samples: ReadonlyArray<Record<K, unknown>>, key: K): boolean {
if (samples.length === 0) throw new Error('Judge panel has no samples');
const values = samples.map(sample => sample && typeof sample === 'object' ? sample[key] : undefined);
const bad = values.findIndex(value => typeof value !== 'boolean');
if (bad !== -1) throw new Error(`Judge panel sample ${bad + 1} has non-boolean ${key}: ${JSON.stringify(values[bad])}`);
return values.filter(value => value === true).length * 2 > values.length;
}
/** Sample reasoning lines, numbered, for the collector record. */
export function judgePanelReasoning(samples: ReadonlyArray<unknown>): string {
return samples.map((sample, index) => {
const reasoning = sample && typeof sample === 'object' ? (sample as { reasoning?: unknown }).reasoning : undefined;
return `[sample ${index + 1}] ${typeof reasoning === 'string' ? reasoning : ''}`;
}).join('\n');
}
const score = { type: 'integer', enum: [1, 2, 3, 4, 5] } as const;
// Structured output guarantees parseable JSON; free-form judges failed on
// unescaped quotes inside their reasoning (run 36798539821, setup block).
export const JUDGE_SCORE_SCHEMA = {
type: 'object',
properties: { clarity: score, completeness: score, actionability: score, reasoning: { type: 'string' } },
required: ['clarity', 'completeness', 'actionability', 'reasoning'],
additionalProperties: false,
};
export const OUTCOME_JUDGE_SCHEMA = {
type: 'object',
properties: {
detected: { type: 'array', items: { type: 'string' } },
missed: { type: 'array', items: { type: 'string' } },
false_positives: { type: 'integer' },
detection_rate: { type: 'integer' },
evidence_quality: score,
reasoning: { type: 'string' },
},
required: ['detected', 'missed', 'false_positives', 'detection_rate', 'evidence_quality', 'reasoning'],
additionalProperties: false,
};
export const POSTURE_SCORE_SCHEMA = {
type: 'object',
properties: { axis_a: score, axis_b: score, reasoning: { type: 'string' } },
required: ['axis_a', 'axis_b', 'reasoning'],
additionalProperties: false,
};
/**
* Score documentation quality on clarity/completeness/actionability (1-5).
*/
@@ -226,7 +318,7 @@ Respond with ONLY valid JSON in this exact format:
Here is the ${section} to evaluate:
${content}`);
${content}`, undefined, { jsonSchema: JUDGE_SCORE_SCHEMA });
}
/**
@@ -266,7 +358,7 @@ Rules:
- "detected" and "missed" arrays must only contain IDs from the ground truth: ${groundTruth.bugs.map((b: any) => b.id).join(', ')}
- detection_rate = length of detected array
- evidence_quality (1-5): Do detected bugs have screenshots, repro steps, or specific element references?
5 = excellent evidence for every bug, 1 = no evidence at all`);
5 = excellent evidence for every bug, 1 = no evidence at all`, undefined, { jsonSchema: OUTCOME_JUDGE_SCHEMA });
}
/**
@@ -320,7 +412,7 @@ Respond with ONLY valid JSON in this exact format:
Here is the output to evaluate:
${text}`, undefined, { signal });
${text}`, undefined, { signal, jsonSchema: POSTURE_SCORE_SCHEMA });
}
/**
@@ -338,6 +430,16 @@ ${text}`, undefined, { signal });
* Format spec: scripts/resolvers/preamble/generate-ask-user-format.ts
* Recommendation: <choice> because <one-line reason>
*/
export const RECOMMENDATION_JUDGE_SCHEMA = {
type: 'object',
properties: {
reason_substance: { type: 'integer', enum: [1, 2, 3, 4, 5] },
reasoning: { type: 'string' },
},
required: ['reason_substance', 'reasoning'],
additionalProperties: false,
};
export async function judgeRecommendation(askUserText: string, signal?: AbortSignal): Promise<RecommendationScore> {
signal?.throwIfAborted();
// Deterministic checks. The format spec requires:
@@ -413,7 +515,7 @@ Respond with ONLY valid JSON:
const out = await callJudge<{ reason_substance: number; reasoning: string }>(
prompt,
'claude-haiku-4-5-20251001',
{ signal },
{ signal, jsonSchema: RECOMMENDATION_JUDGE_SCHEMA },
);
// Defensive clamp: rubric is 1-5. If Haiku returns out-of-range or non-numeric,
@@ -453,9 +555,6 @@ export interface ArmJudgeScore {
*/
export const ARM_JUDGE_MODEL = CLAUDE_FRONTIER_EVAL_MODEL;
/** Bounded retry-on-malformed loop: total attempts, not extra retries. */
export const ARM_JUDGE_ATTEMPTS = 2;
/**
* Build the over-engineering rubric prompt. Exported (pure) so the free
* selftest can verify prompt construction without any API call.
@@ -528,10 +627,10 @@ export function parseArmJudgeResponse(raw: unknown): ArmJudgeScore {
*
* - Zero-diff arms are VALID scored cells: the agent built nothing, so the
* score is deterministically 0/"none" — no API call.
* - Bounded retry-on-malformed: ARM_JUDGE_ATTEMPTS total attempts. callJudge
* already retries 429s internally; this loop covers malformed/refused JSON.
* - One sample, never re-asked: a malformed or refused verdict is a failed
* sample. callJudge's transport-level 429 backoff is not a verdict retry.
* - `opts.call` is an injection seam so the free selftest can exercise the
* retry bound without spending API money. Defaults to the real callJudge.
* malformed path without spending API money. Defaults to the real callJudge.
*/
export async function armJudge(
task: string,
@@ -546,18 +645,10 @@ export async function armJudge(
};
}
const call = opts?.call ?? callJudge;
const prompt = buildArmJudgePrompt(task, diff);
let lastError: unknown;
for (let attempt = 1; attempt <= ARM_JUDGE_ATTEMPTS; attempt++) {
try {
const raw = await call<Record<string, unknown>>(prompt, ARM_JUDGE_MODEL);
return parseArmJudgeResponse(raw);
} catch (err) {
lastError = err;
}
const raw = await call<Record<string, unknown>>(buildArmJudgePrompt(task, diff), ARM_JUDGE_MODEL);
try {
return parseArmJudgeResponse(raw);
} catch (err) {
throw new Error(`armJudge: malformed verdict (never resampled) — ${err instanceof Error ? err.message : String(err)}`);
}
throw new Error(
`armJudge: no well-formed verdict after ${ARM_JUDGE_ATTEMPTS} attempts — `
+ (lastError instanceof Error ? lastError.message : String(lastError)),
);
}
+1 -1
View File
@@ -51,7 +51,7 @@ function modeField(line: string): { value: string; completed: boolean } | null {
// unfinished and unknown statuses also invalidate an earlier declaration.
const { label, status, value: rawValue } = match.groups!;
const completeStatus = !status || /^(?:done|complete|completed)$/i.test(status.trim());
const explicitMode = /^(?:the )?(?:review )?mode\b(?:\s+is\b|:)?\s*/i;
const explicitMode = /^(?:the )?(?:review )?mode\b(?:\s+is\b|:|\s*=(?!=))?\s*/i;
// An unqualified Decision field owns a review mode only when its value
// names that vocabulary. Keep unrelated decisions out of withdrawal checks;
// partial/negated mode names still own a field and therefore fail closed.
+44 -27
View File
@@ -91,6 +91,49 @@ function assignmentBody(markdown: string): string {
|| '';
}
/**
* Design-draft phase of the fixed fixture: the repo design carries every
* required section and an Assignment, and an independent Agent/Task opinion on
* RosterCheck preceded the Write that created it. The full workflow validator
* applies these same checks; the focused design-draft capture applies them alone.
*/
export function validateOfficeHoursDesignDraft(
evidence: Pick<OfficeHoursCompletionEvidence, 'designPath' | 'designContent' | 'toolCalls'>,
label = 'Office-hours design draft',
): { designPath: string; repoPath: string; firstDesignWrite: number } {
const fail = (message: string): never => { throw new Error(`${label}: ${message}`); };
if (evidence.designContent === null) fail(`repo design is missing: ${evidence.designPath}`);
const design = evidence.designContent!;
for (const [section, names] of [
['Problem Statement', ['problem statement']],
['Recommended Approach', ['recommended approach']],
['Success Criteria', ['success criteria']],
['What I noticed about how you think', ['what i noticed about how you think']],
] as const) {
if (!substantive(sectionBody(design, [...names]))) fail(`repo design lacks substantive ${section}`);
}
if (!substantive(assignmentBody(design))) fail('repo design lacks a concrete Assignment');
// A cold-read opinion before the design exists is not the required spec
// review. The fixture promises an available Agent, so require an attempt
// that names this design even when the review subsequently fails.
const designPath = evidence.designPath.replace(/\\/g, '/');
const repoPath = designPath.match(/(?:^|\/)(docs\/designs\/[^/]+\.md)$/)?.[1] ?? designPath;
const firstDesignWrite = evidence.toolCalls.findIndex(call => {
const writtenPath = String(call.input?.file_path ?? '').replace(/\\/g, '/').replace(/^\.\//, '');
return call.tool === 'Write' && (writtenPath === designPath || writtenPath === repoPath);
});
if (firstDesignWrite === -1) fail('no observed Write created the repo design');
const opinion = evidence.toolCalls.slice(0, firstDesignWrite).some(call => {
if (!['Agent', 'Task'].includes(call.tool)) return false;
const prompt = `${String(call.input?.description ?? '')}\n${String(call.input?.prompt ?? '')}`;
return /\bRosterCheck\b/i.test(prompt)
&& /\b(?:review|challenge|opinion|critique|perspective|steelman|advisor)\b|\bcold.read\b/i.test(prompt);
});
if (!opinion) fail('no independent Agent/Task opinion on RosterCheck preceded the repo design Write');
return { designPath, repoPath, firstDesignWrite };
}
export function validateOfficeHoursCompletion(evidence: OfficeHoursCompletionEvidence): OfficeHoursReviewEvidence | null {
const fail = (message: string): never => { throw new Error(`Office-hours completion: ${message}`); };
if (evidence.exitReason !== 'success') fail(`execution failed: ${evidence.exitReason}`);
@@ -111,34 +154,8 @@ export function validateOfficeHoursCompletion(evidence: OfficeHoursCompletionEvi
if (statuses.length !== 1 || statuses[0].trim().toUpperCase() !== 'APPROVED') {
fail('repo design is not marked Status: APPROVED');
}
for (const [label, names] of [
['Problem Statement', ['problem statement']],
['Recommended Approach', ['recommended approach']],
['Success Criteria', ['success criteria']],
['What I noticed about how you think', ['what i noticed about how you think']],
] as const) {
if (!substantive(sectionBody(design, [...names]))) fail(`repo design lacks substantive ${label}`);
}
if (!substantive(assignmentBody(design))) fail('repo design lacks a concrete Assignment');
const { designPath, repoPath, firstDesignWrite } = validateOfficeHoursDesignDraft(evidence, 'Office-hours completion');
if (!substantive(assignmentBody(evidence.output))) fail('REPORT.md lacks the Assignment');
// A cold-read opinion before the design exists is not the required spec
// review. The fixture promises an available Agent, so require an attempt
// that names this design even when the review subsequently fails.
const designPath = evidence.designPath.replace(/\\/g, '/');
const repoPath = designPath.match(/(?:^|\/)(docs\/designs\/[^/]+\.md)$/)?.[1] ?? designPath;
const firstDesignWrite = evidence.toolCalls.findIndex(call => {
const writtenPath = String(call.input?.file_path ?? '').replace(/\\/g, '/').replace(/^\.\//, '');
return call.tool === 'Write' && (writtenPath === designPath || writtenPath === repoPath);
});
if (firstDesignWrite === -1) fail('no observed Write created the repo design');
const opinion = evidence.toolCalls.slice(0, firstDesignWrite).some(call => {
if (!['Agent', 'Task'].includes(call.tool)) return false;
const prompt = `${String(call.input?.description ?? '')}\n${String(call.input?.prompt ?? '')}`;
return /\bRosterCheck\b/i.test(prompt)
&& /\b(?:review|challenge|opinion|critique|perspective|steelman|advisor)\b|\bcold.read\b/i.test(prompt);
});
if (!opinion) fail('no independent Agent/Task opinion on RosterCheck preceded the repo design Write');
const reviews = evidence.toolCalls.slice(firstDesignWrite + 1).filter(call => {
if (!['Agent', 'Task'].includes(call.tool)) return false;
const prompt = String(call.input?.prompt ?? '').replace(/\\/g, '/');
+79
View File
@@ -44,3 +44,82 @@ export const PERIODIC_CI_EXCLUDE: Record<string, { reason: string; tracking: str
tracking: 'TODOS.md "CI-unrunnable paid evals" (re-entry: the CLI/device is available in the CI image; review by 2026-12-28)',
},
};
/**
* Case-level exclusions for case-sharded files (`<file>#<case id>`), same
* contract as above: a case lands here only when a CI runner cannot execute it
* (it self-skips), with reason + tracking. The planner records each as an
* excluded manifest entry instead of an empty case shard, so the exact
* one-case check stays strict for every planned case. Pinned by
* test/periodic-exclude-policy.test.ts.
*/
export const CASE_CI_EXCLUDE: Record<string, { reason: string; tracking: string }> = {
'test/skill-e2e-design.test.ts#design-review-fix': {
reason: '/design-review drives the Aside browser; CI runners are Linux without Aside, so the case registers test.skip("needs Aside")',
tracking: 'TODOS.md "CI-unrunnable paid evals" (re-entry: the CLI/device is available in the CI image; review by 2026-12-28)',
},
};
/**
* Paid-eval verdict policy, pre-registered (approved 2026-09-29). Frozen before
* the census: any change after seeing census results needs Garry's
* re-approval and a fresh census, and bumps `version` (every trial record
* carries it as policy_version, so pass-rate history segments at the change).
* panel - behavior cases and quarantined cases run n independent
* trials; a behavior panel PASSES at >= k passing trials with
* no contract violation. Rule and judge cases run one trial.
* quarantine - entry below `entry.rate` per trial over >= `entry.minTrials`
* new-policy trials; exit at >= `exit.rate` over >=
* `exit.minTrials`; at most `capFraction` of each tier's
* blocking cases; an entry expires after `expiryWeeklyRuns`.
* judge - a judge case draws `samples` independent samples of one
* prompt concurrently; numeric dimensions gate on the panel
* mean against the unchanged threshold, booleans on a strict
* majority; an erroring sample fails the panel, never resampled.
* drift - one-sided Fisher exact alarm between input-identity series
* (Holm-controlled across the cases tested in one report).
* infraRedispatch - a census whose every red verdict is machine-classified
* INFRA or INCOMPLETE may be re-dispatched this many times as
* a new run; both runs are reported.
*/
export const EVAL_POLICY = {
version: 1,
panel: { n: 3, k: 2 },
quarantine: {
entry: { rate: 0.95, minTrials: 10 },
exit: { rate: 0.97, minTrials: 10 },
capFraction: 0.10,
expiryWeeklyRuns: 8,
},
judge: { samples: 3 },
drift: { fisherAlpha: 0.05, fisherMinPerSide: 6 },
infraRedispatch: 1,
} as const;
/**
* Quarantined paid cases, keyed by registry id (an E2E_TIERS key). A
* quarantined case still runs its full panel and reports in every lane, but
* its failed verdict cannot fail the lane unless the panel is a hard break
* (0 of n) or a trial violated a contract; it never counts as passing
* coverage. An entry needs the entry rule met on the current input identity,
* a written diagnosis that the failures are detector, harness or model-latency
* failures (a product defect is never quarantined), and unchanged case
* touchfiles in the change that adds it. Pinned by
* test/periodic-exclude-policy.test.ts.
* reason - the written diagnosis, with the pass-rate evidence
* failureClass - what the diagnosis found; a product defect has no class here
* tracking - issue or TODOS pointer
* owner - who removes it
* enteredAt - YYYY-MM-DD the entry landed (expiry counts weekly runs from here)
* exit - the measurable exit condition
* At most EVAL_POLICY.quarantine.capFraction of a tier's cases (gate and
* periodic are the blocking tiers) may be quarantined at once.
*/
export const CASE_QUARANTINE: Record<string, {
reason: string;
failureClass: 'detector' | 'harness' | 'model-latency';
tracking: string;
owner: string;
enteredAt: string;
exit: string;
}> = {};
+5 -1
View File
@@ -284,7 +284,11 @@ function currentCreatePreview(preview: string, r: any, config: string, cwd: stri
if(event.name!=='Write'||`${event.sessionId}:${event.toolUseId}`!==r.pendingId||event.input?.file_path!==r.expected||
Date.parse(event.timestamp)<startedAt||typeof event.input.content!=='string'||
Buffer.byteLength(event.input.content)>MAX_WRITE_INPUT_BYTES) return false;
const source=event.input.content.split(/\r?\n/), rows=preview.split('\n');
// A crop can keep the pane's file row and rule above the preview while its
// "Create file" title scrolls away. That row must name the owned path.
const header=/^ {0,3}(?![1-9]\d*(?:[ \t]|\n))(\S[^\n]*)\n[╌─━]{3,}[ \t]*\n/.exec(preview);
if(header && path.resolve(cwd,header[1]!.trim())!==r.expected) return false;
const source=event.input.content.split(/\r?\n/), rows=preview.slice(header?.[0].length ?? 0).split('\n');
const numbered:Array<{line:number;text:string}>=[];
let leading='';
for(const row of rows) {
+9 -1
View File
@@ -137,6 +137,14 @@ function deterministicPlanFloorSetup(input: PlanFloorReview): PlanFloorAssessmen
/\b(?:developer|sdk developer|user)\b/.test(question) &&
/\b(?:experiences|journey|narrative)\b/.test(question);
// plan-devex-review 0B's confirmation contract: the brief asks whether the
// narrative matches reality and every option is one of its three answers
// (accurate / some or partly wrong / way off). Any remedy option leaves it to the assessor.
const isDxNarrativeConfirmation =
/\b(?:empathy|narrative)\b/.test(header) &&
/\b(?:narrative|journey)\b[^?\n]*\bmatch\b[^?\n]*\?/.test(q.question.split(/\r?\n/)[0]!.toLowerCase()) &&
q.options.every(o => /^(?:[a-d][).:]\s*)?(?:(?:this is\s+)?accurate|(?:some|partly)\b[^,]*?\b(?:wrong|corrections?)|(?:this is\s+)?way off)\b/i.test(o.label.trim()));
const isProductTypeSetup =
header === 'product type' &&
/^is this\b/.test(question) &&
@@ -146,7 +154,7 @@ function deterministicPlanFloorSetup(input: PlanFloorReview): PlanFloorAssessmen
/^(?:mode|review mode)$/.test(header) &&
/\b(?:which|what)\b.*\breview mode\b/.test(question);
if (!isDxEmpathySetup && !isProductTypeSetup && !isReviewModeSetup) return null;
if (!isDxEmpathySetup && !isDxNarrativeConfirmation && !isProductTypeSetup && !isReviewModeSetup) return null;
return validatePlanFloorAssessment(input, {
kind: 'setup',
+6 -3
View File
@@ -54,6 +54,9 @@ export interface PlanReviewDecisionJudgment {
engReview?: EngReviewJudgment;
}
export type PlanReviewJudge = (prompt: string, model?: string, opts?: Pick<CallJudgeOptions, 'signal' | 'max_tokens' | 'jsonSchema'>) => Promise<unknown>;
// Structured outputs cannot enforce maxLength, so the reason bound the local
// validator applies is stated on the field the model writes.
const REASON_FIELD = { type: 'string', description: '1-1000 characters: under 120 words.' } as const;
// Only response structure is constrained. Identity, exact quotes, enum casing,
// uncertainty, target coverage, independence and count checks remain local.
function planReviewDecisionSchema(withPeerComparison: boolean, withEngReview = false): NonNullable<CallJudgeOptions['jsonSchema']> {
@@ -68,7 +71,7 @@ function planReviewDecisionSchema(withPeerComparison: boolean, withEngReview = f
toolUseId: { type: 'string' }, questionIndex: { type: 'integer' },
kind: { type: 'string', enum: ['finding', 'scope', 'workflow', 'backlog', 'uncertain'] },
targetIds: { type: 'array', items: { type: 'string' } },
independentDecisions: { type: 'integer' }, reason: { type: 'string' },
independentDecisions: { type: 'integer' }, reason: REASON_FIELD,
evidence: { type: 'array', items: {
type: 'object', additionalProperties: false, required: ['field', 'optionIndex', 'quote'],
properties: {
@@ -86,7 +89,7 @@ function planReviewDecisionSchema(withPeerComparison: boolean, withEngReview = f
type: 'object', additionalProperties: false,
required: ['status', 'regression', 'approvals', 'navigation', 'reason'],
properties: {
status: { type: 'string', enum: ['complete', 'missing', 'uncertain'] }, reason: { type: 'string' },
status: { type: 'string', enum: ['complete', 'missing', 'uncertain'] }, reason: REASON_FIELD,
regression: { type: 'array', items: { type: 'object', additionalProperties: false,
required: ['role', 'source', 'quote'], properties: {
role: { type: 'string', enum: ['critical', 'baseline', 'replay', 'assertions', 'approved-differences'] },
@@ -112,7 +115,7 @@ function planReviewDecisionSchema(withPeerComparison: boolean, withEngReview = f
properties: { name: { type: 'string' }, quote: { type: 'string' } },
} },
productQuote: { type: 'string' }, groundingQuote: { type: 'string' },
implicationQuote: { type: 'string' }, reason: { type: 'string' },
implicationQuote: { type: 'string' }, reason: REASON_FIELD,
},
} } : {}),
},
+11 -3
View File
@@ -152,6 +152,9 @@ export async function submitPlanSeed(session: SeedSession, seed: string, opts: {
});
if (Date.now() >= opts.deadlineAt) throw new PlanSeedTimeout('Plan seed submission exhausted the existing case budget');
session.sendKey('Enter'); // Separate input event after the acknowledged paste.
// The transcript can record end_turn before the CLI repaints, so an empty
// composer counts only when the same frame survives one more poll.
let settled = '';
await until(async () => {
const owned = read();
if (!owned || owned.pendingBytes) return false;
@@ -178,13 +181,18 @@ export async function submitPlanSeed(session: SeedSession, seed: string, opts: {
}
if (row.type === 'user') for (const c of content(row)) if (c.type === 'tool_result') pending.delete(c.tool_use_id);
}
if (!complete || pending.size || owned.status.waitingFor) return false;
const unsettled = () => { settled = ''; return false; };
if (!complete || pending.size || owned.status.waitingFor) return unsettled();
const frame = await session.currentScreen!();
if (opts.isQuestionOrPermission(frame.text)) throw new Error('Plan seed response requires an answer before skill invocation');
const input = composer(frame.text);
if (frame.rawEnd !== session.mark() || !input
|| input.line.replace(/^❯[ \u00a0]*/, '').trim() !== '') return false;
|| input.line.replace(/^❯[ \u00a0]*/, '').trim() !== '') return unsettled();
const fresh = read();
return !!fresh && !fresh.pendingBytes && fresh.rows.length === owned.rows.length && !fresh.status.waitingFor;
if (!fresh || fresh.pendingBytes || fresh.rows.length !== owned.rows.length || fresh.status.waitingFor) return unsettled();
const signature = `${frame.rawEnd}:${fresh.rows.length}:${frame.text}`;
if (signature === settled) return true;
settled = signature;
return false;
});
}
+19 -35
View File
@@ -638,42 +638,26 @@ export const engStep0Boundary: Step0BoundaryPredicate = (fp) =>
// plan-eng-review-idempotency, plan-eng-review-todos-e2e-concurrent.
/gstack-qid:\s*(?:plan-)?eng-review-/i.test(fp.promptSnippet);
/** Completed plan-wide focus and local-learnings choices remain setup, even when asked late. */
export const designReviewSetupAUQ: Step0BoundaryPredicate = (fp) => {
const call = fp.nativeCall;
if (call?.answered !== true || call.failed !== false || !call.sessionId || !call.toolUseId ||
call.questions.length !== 1 || !Array.isArray(call.unansweredQuestionIndices) || call.unansweredQuestionIndices.length ||
!Number.isFinite(Date.parse(call.answeredAt ?? '')) || fp.signature !== `${call.sessionId}:${call.toolUseId}` ||
(fp.nativeQuestionIndex !== undefined && fp.nativeQuestionIndex !== 0)) return false;
const q = call.questions[0]!;
if (q.multiSelect || q.options.length !== 2 || new Set(q.options.map(o => o.label)).size !== 2 ||
Object.keys(call.answers ?? {}).length !== 1 || q.options.filter(o => o.label === call.answers?.[q.question]).length !== 1 ||
fp.options.length !== 2 || !fp.options.every((o, i) => o.index === i + 1 && o.label === q.options[i]!.label)) return false;
/**
* The seed declares "Design: review all seven dimensions" for the pending Step 0D
* focus menu. Pick its single all-seven option only when every other option
* narrows the review and the brief approves no product action.
*/
export function pickDesignFocusAll(q: NativePlanQuestionCall['questions'][number]): number | null {
if (q.multiSelect || q.options.length < 2 || q.options.length > 4 || new Set(q.options.map(o => o.label)).size !== q.options.length) return null;
const text = q.question.trim();
const title = text.split(/\r?\n/, 1)[0]!.replace(/^D[1-9]\d*\s*[—–:-]\s*/i, '');
const sources = [...text.matchAll(/^Project\/branch\/task:\s*([^\n]+)$/gm)];
const source = sources[0]?.[1] ?? '';
// Setup never approves another product action. Quoted examples and negative
// consequences are explanatory; current imperative clauses remain decisions.
const explanatory = [text, ...q.options.map(o => o.description ?? '')].join('\n')
.replace(/`+[^`]*`+|"[^"\n]*"|“[^”\n]*”|‘[^’\n]*’/g, '')
.replace(/[✅❌*]/g, '');
if (/(?:^|[.!?;:\n]|\b(?:and|while))\s*(?:(?:also|please|then|now)\s+)*(?:approv(?:e|ing)|deploy(?:ing)?|implement(?:ing)?|ship(?:ping)?|merg(?:e|ing)|delet(?:e|ing))\b/im.test(explanatory)) return false;
if (sources.length !== 1 || (text.match(/\?/g)?.length ?? 0) !== 1 || /```|~~~|^\s*>/m.test(text) ||
!/\bplan-design-review of PLAN\.md\b/i.test(source) ||
/\b(?:historical|archived|quoted|example|foreign|other|another|previous)\b/i.test(source)) return false;
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).:]\s+/i, '')
.replace(/\s*\(recommended\)\s*$/i, ''));
if (/^(?:Learnings|Cross-project)$/i.test(q.header.trim()) &&
/^Enable cross[- ]project learnings(?: search)?\?$/i.test(title) &&
labels.some(label => /^Enable cross[- ]project learnings$/i.test(label)) &&
labels.some(label => /^Keep learnings project[- ]scoped(?: only)?$/i.test(label))) {
// Reuse the existing native cross-project premise/owned answer classifier.
return engSetupAUQ(fp);
}
return /^(?:Focus|Review focus)$/i.test(q.header.trim()) &&
/^Review all 7 (?:design )?(?:dimensions|passes),? or focus(?: on (?:specific areas|a subset))?\?$/i.test(title) &&
/^ELI10:\s*I['’]ve rated this plan (?:10(?:\.0+)?|[0-9](?:\.\d+)?)\/10 on design completeness\./mi.test(text) &&
labels.some(label => /^(?:Review )?All 7 (?:design )?(?:dimensions|passes)$/i.test(label)) &&
labels.some(label => /^(?:Only (?:the )?[1-6](?: listed)? (?:gaps|areas|dimensions|passes)|Focus on (?:specific areas|a subset))$/i.test(label));
};
.replace(/`+[^`]*`+|"[^"\n]*"|“[^”\n]*”|‘[^’\n]*’/g, '').replace(/[✅❌*]/g, '');
if (!/^Review all (?:7|seven) (?:design )?(?:dimensions|passes),? or focus(?: on [^?\n]+)?\?$/i.test(title) ||
sources.length !== 1 || !/\bplan-design-review of PLAN\.md\b/i.test(sources[0]![1]!) ||
/\b(?:historical|archived|quoted|example|foreign|other|another|previous)\b/i.test(sources[0]![1]!) ||
(text.match(/\?/g)?.length ?? 0) !== 1 || /```|~~~|^\s*>/m.test(text) ||
!/^ELI10:\s*I['’]ve rated this plan (?:10(?:\.0+)?|[0-9](?:\.\d+)?)\/10 on design completeness\./mi.test(text) ||
/(?:^|[.!?;:\n]|\b(?:and|while))\s*(?:(?:also|please|then|now)\s+)*(?:approv(?:e|ing)|deploy(?:ing)?|implement(?:ing)?|ship(?:ping)?|merg(?:e|ing)|delet(?:e|ing))\b/im.test(explanatory)) return null;
const labels = q.options.map(o => o.label.trim().replace(/^[A-Z][).:]\s+/i, '').replace(/\s*\(recommended\)\s*$/i, ''));
const all = labels.flatMap((label, i) => /^(?:Review )?All (?:7|seven) (?:design )?(?:dimensions|passes)$/i.test(label) ? [i + 1] : []);
if (all.length !== 1 || labels.some((label, i) => i + 1 !== all[0] && !/^(?:Only|Focus)\b/i.test(label))) return null;
return all[0]!;
}
+1 -1
View File
@@ -105,7 +105,7 @@ function conflictingDesignClosure(text: string): boolean {
new RegExp(`(?:^|[.!?;]\\s+|\\n)(?:If|When|Once|Unless|Assuming|Provided)\\b[^.!?\\n]*\\b${owner}\\b`, 'i').test(text);
}
function hasCompletePlanReport(expectedPlanPath: string, minimumMtime: number, maximumMtime: number,
export function hasCompletePlanReport(expectedPlanPath: string, minimumMtime: number, maximumMtime: number,
allowRunHeaderForFailure = false, requiredReview?: 'Design'): boolean {
if (!path.isAbsolute(expectedPlanPath)) return false;
try {
+6 -2
View File
@@ -221,7 +221,7 @@ export async function runPlanSkillCounting(opts: PlanSkillCountingOptions): Prom
seen: new Set(), countedCalls: new Set(), filePermission: createPlanCountPermissionGuard(), ownedFilePermissions: [],
lastMatchedNativeQuestion: undefined, transcript: { status: 'missing', calls: [], assistantMessages: [] },
boundaryFired: false, step0Count: 0, reviewCount: 0, administrativeCount: 0, isFirstAUQ: true,
lastCheckpointAt: 0, viewport: '', observedOutput: 0, lastObservationAt: -Infinity,
lastCheckpointAt: 0, viewport: '', observedOutput: 0, lastObservationAt: -Infinity, lastOutputAt: driver.now(),
};
const permissionPaths = [
...(opts.expectedPlanPath ? [opts.expectedPlanPath, path.join(fixture.cwd, 'PLAN.md')] : []),
@@ -253,7 +253,8 @@ export async function runPlanSkillCounting(opts: PlanSkillCountingOptions): Prom
poll: session => countingPoll(run, session),
tick: session => countingTick(run, session),
timeout: session => countingSnapshot(run, session, 'timeout',
`no terminal outcome within ${timeoutMs}ms total budget (including startup and ${cleanupReserveMs}ms cleanup reserve; step0=${run.step0Count}, review=${run.reviewCount})`,
`no terminal outcome within ${timeoutMs}ms total budget (including startup and ${cleanupReserveMs}ms cleanup reserve; step0=${run.step0Count}, review=${run.reviewCount}); ` +
`idleFor=${run.driver.now() - run.lastOutputAt}ms`,
run.viewport),
onError: (session, error) => countingFailure(run, session, error),
onCloseError: (session, error) => {
@@ -294,6 +295,8 @@ export interface CountingRun {
viewport: string;
observedOutput: number;
lastObservationAt: number;
/** Wall time of the last new PTY output; the timeout summary reports the idle span. */
lastOutputAt: number;
}
function remainingWork(run: CountingRun): number {
@@ -388,6 +391,7 @@ async function countingPoll(run: CountingRun, session: ClaudePtySession): Promis
if (remainingWork(run) <= 0) return false;
await session.waitForOutput(run.observedOutput, Math.min(2000, remainingWork(run)));
if (remainingWork(run) <= 0) return false;
if (session.rawOutput().length > run.observedOutput) run.lastOutputAt = run.driver.now();
const coalesceMs = session.rawOutput().length > run.observedOutput
? 250 : 250 - (run.driver.monotonic() - run.lastObservationAt);
if (coalesceMs > 0 && !await waitForWork(run, coalesceMs)) return false;
+6 -4
View File
@@ -16,7 +16,7 @@ import { isDeepStrictEqual } from 'node:util';
import { randomUUID } from 'node:crypto';
import { capturePlanCountQuestion, createPlanCountPermissionGuard, matchesNativePlanQuestion, planCountPrerequisitePick, planCountQuestionInput } from '../auq';
import { resolveClaudeBinary } from '../binary';
import { designReviewSetupAUQ } from '../boundaries';
import { pickDesignFocusAll } from '../boundaries';
import { SANCTIONED_WRITE_SUBSTRINGS, isProseAUQVisible, planCountSubmissionInput } from '../classify';
import type { ClaudePtySession } from '../launch';
import { isNumberedOptionListVisible, isPermissionDialogVisible, isPlanReadyVisible, isRejectedSlashCommand, stripPtyResidue } from '../screen';
@@ -168,6 +168,10 @@ export async function runPlanSkillFloorCheck(opts: PlanSkillFloorOptions): Promi
'Proceed directly to the requested review; skip the optional /office-hours prerequisite.',
'This actor has already declined routing setup, cross-project recall and outside reviewers.',
'Preserve the supplied product scope. For review-mode questions choose HOLD SCOPE (CEO), DX POLISH (DX), or the full BIG CHANGE review (Eng). Design: review all seven dimensions.',
// Runs 36626737820 and 36776104571: CEO read scope preservation as approving
// approach A, so the seeded premise gap never reached a question. Run 36794871032:
// CEO read this request as supplying every answer up front and asked nothing.
...(opts.skillName === 'plan-ceo-review' ? ['Preserving scope does not approve the plan\'s premise, approach or any remedy. This request answers only the routing, recall, outside-reviewer and review-mode questions named above.'] : []),
...(opts.productType === 'sdk-documentation' ? [
'Product type is confirmed: SDK quickstart documentation, with the complete journey to the first SDK call as context. If asked to classify, choose SDK + Docs when offered, otherwise Documentation. This confirms the review lens; it does not expand the plan.',
'Target persona is confirmed: a hands-on developer integrating this SDK for the first time, trying to make one successful call. Product type and persona setup are already answered; proceed to reviewing the supplied plan.',
@@ -423,11 +427,9 @@ function floorQuestion(run: FloorRun, session: ClaudePtySession, currentCalls: N
const question = pendingQuestion.questions[index]!;
const key = `${pendingQuestion.sessionId}:${pendingQuestion.toolUseId}`;
const chosen = run.setupChoices.get(key) ?? new Set<number>();
const allDesign = opts.skillName === 'plan-design-review' && designReviewSetupAUQ(fp)
? question.options.flatMap((option, i) => /^(?:Review )?All 7 (?:design )?(?:dimensions|passes)(?:\s*\(recommended\))?$/i.test(option.label.trim()) ? [i + 1] : []) : [];
const pick = pickPlanFloorMode(opts.skillName, question) ?? planCountPrerequisitePick(fp, fp)
?? (opts.skillName === 'plan-devex-review' ? pickPlanFloorProductType(question, opts.productType) : null)
?? (allDesign.length === 1 ? allDesign[0]! : null);
?? (opts.skillName === 'plan-design-review' ? pickDesignFocusAll(question) : null);
if (pick !== null) {
if (!chosen.has(index)) {
session.send(planCountQuestionInput(viewport, fp, pick));
+53 -7
View File
@@ -11,7 +11,7 @@ import { runSkillTest, SESSION_DRAIN_GRACE_MS } from './session-runner';
import { CAPTURE_MS } from './eval-budgets';
import { refreshHermeticSkillRuntime } from './hermetic-skill-runtime';
import { seedHermeticGstackHome } from './hermetic-env';
import { observeQAWrites, type QAWriteObservation } from './qa-functional-observer';
import { observeQAWrites, qaHelperUsageCommand, type QAWriteObservation } from './qa-functional-observer';
import { nativeCalls, readQACheckpointFiles, validateQACheckpoints } from './qa-checkpoint-evidence';
import { ownedPath } from './qa-functional-fixture';
import { qaEvidenceCommand, qaNativeCapture, qaProducerReceipt, type QaEvidenceContext } from './qa-evidence-producer';
@@ -214,6 +214,7 @@ function callerDeadlineCommand(command: string, context?: CallerDeadlineContext)
export function qaCallerCommandAllowed(command: string, workflowCommands: string[] = [], deadline?: CallerDeadlineContext): boolean {
const text = command.trim();
if (qaHelperUsageCommand(text)) return true;
if (callerEvidenceCommand(text, deadline)) return true;
if (callerDeadlineCommand(text, deadline)) return true;
if (/\bgstack-qa-(?:deadline|evidence)\b/.test(text) && !literalCallerCLI.test(text)
@@ -243,11 +244,13 @@ export function qaCallerCommandAllowed(command: string, workflowCommands: string
}
if (/[\n\r;&|<>`$\\(){}]/.test(text)) return false;
return /^(?:pwd|ls(?: -la)?|bun --version|date -u \+%Y-%m-%dT%H:%M:%SZ|bun (?:run test|test(?: cli\.test\.ts)?))$/.test(text)
|| /^git (?:status --(?:short|porcelain)|branch --show-current|rev-parse (?:--short )?HEAD|merge-base origin\/main HEAD|ls-files(?: --others --exclude-standard)?)$/.test(text)
|| /^git (?:status --(?:short|porcelain)|branch --show-current|rev-parse (?:--short )?HEAD|merge-base origin\/main HEAD|log origin\/main\.\.HEAD --oneline|ls-files(?: --others --exclude-standard)?)$/.test(text)
|| /^\/?[\w./-]+\/bin\/gstack-review-log --start (?:review|adversarial-review)$/.test(text)
|| /^\/?[\w./-]+\/bin\/gstack-(?:review-read|specialist-stats)$/.test(text);
}
const REVIEW_RECORD_STATUSES: Record<QaCaller, string[]> = { review: ['clean', 'issues_found'], ship: ['clean', 'issues_found', 'unavailable'] };
export function validateCallerEvidence(input: {
caller: QaCaller;
result: Pick<SkillTestResult, 'transcript' | 'exitReason'>;
@@ -300,6 +303,11 @@ export function validateCallerEvidence(input: {
if (tool.name === 'Bash' && !qaCallerCommandAllowed(String(tool.input.command), input.workflowCommands, deadline)) {
errors.push('command outside declared caller observation interface');
}
const record = tool.name === 'Bash' ? String(tool.input.command).match(/gstack-review-log '(.*)'(?: --finish \S+)?$/s)?.[1] : undefined;
const status = record?.match(/"status"\s*:\s*"([^"]*)"/)?.[1] ?? '';
if (record && /"skill"\s*:\s*"review"/.test(record) && !REVIEW_RECORD_STATUSES[input.caller].includes(status)) {
errors.push(`review record status outside the ${input.caller} vocabulary: ${status}`);
}
if (tool.name === 'Read' && /\/(?:browse|devex-review)\/SKILL\.md$|\/qa\/sections\/(?:browser-[^/]+|qa-patterns)\.md$/.test(file)) {
errors.push(`unexpected browser/DX load: ${file}`);
}
@@ -408,6 +416,9 @@ export function validateCallerEvidence(input: {
probes: checkpointProbes, requiredProbes: checkpointProbes.slice(1),
additionalTargets: expiredTargets,
files: input.checkpointFiles, reportMarkdown: input.reportMarkdown,
// An installed review-log record cites checkpoints in its findings' evidence.
citesCheckpointsOnly: command => /^\S*\/gstack-review-log '/.test(command.trim())
&& qaCallerCommandAllowed(command, input.workflowCommands, deadline),
}));
const selected = input.receipt.probes.map(id => input.probes.find(probe => probe.id === id));
if (selected.some(probe => !probe)) errors.push('receipt references an unobserved probe');
@@ -436,6 +447,11 @@ export function validateCallerEvidence(input: {
errors.push('review completion preceded handoff freshness decision');
}
}
if (input.requireCapturedEvidence && input.reportRoot) {
const evidenceFile = path.join(input.reportRoot, 'evidence.json');
const verdict = fs.existsSync(evidenceFile) ? JSON.parse(fs.readFileSync(evidenceFile, 'utf8')).verdict : undefined;
if (verdict && input.receipt.status !== verdict.status) errors.push(`receipt status ${input.receipt.status} differs from helper verdict ${verdict.status}`);
}
if (input.receipt.status === 'pass' && (input.receipt.remaining.length || selected.some(probe => probe?.status !== 'pass'))) {
errors.push('blocked, failing or incomplete coverage reported green');
}
@@ -468,8 +484,10 @@ export interface QaCallerFixture {
caller: QaCaller;
caseId: QaCallerCase;
reviewStart?: string;
startObservations: string;
gitEnvironment: Record<'GIT_OBJECT_DIRECTORY' | 'GIT_ALTERNATE_OBJECT_DIRECTORIES', string>;
journal: string;
child?: string;
mutationEvents: string[];
observerErrors: string[];
observation?: QAWriteObservation;
@@ -535,6 +553,18 @@ process.exit(exit ?? 127);
if (caseId === 'ship-exploratory-plan-checks') {
fs.writeFileSync(path.join(cwd, 'PLAN.md'), '# Scale change\n\n## Verification\n\nThe upper boundary is a required release check: run `bun scripts/probe.ts 9`; require exit 0, stdout `18\\n`, and empty stderr. Ordinary positive input is not a substitute for this check.\n');
}
const childOutputs: Partial<Record<QaCallerCase, string>> = {
'review-exploratory-small-cli': `ADVERSARIAL REVIEW (Claude subagent):
Fixtures and tests were reviewed in summary mode; none changed.
1. FIXABLE — scale.ts:3: the added \`!n\` guard treats the documented valid input 0 as invalid because Number("0") is falsy, so \`0\` now exits 2 with the rejection message instead of printing 0.
Recommendation: Fix the \`!n\` guard at scale.ts:3 because it rejects the documented lower boundary 0.
`,
'ship-exploratory-plan-checks': `{"total_items":0,"done":0,"changed":0,"partial":0,"not_done":0,"unverifiable":0,"summary":"PLAN.md has no implementation deliverables. Execution-only check retained verbatim for Step 8.1/9: run \`bun scripts/probe.ts 9\`; require exit 0, stdout \`18\\\\n\`, and empty stderr (PLAN.md, Verification). Pending execution."}
`,
};
const childText = childOutputs[caseId];
const child = childText === undefined ? undefined : path.join(root, 'children', caller === 'review' ? 'adversarial.md' : 'plan-audit.json');
if (child) { fs.mkdirSync(path.dirname(child)); fs.writeFileSync(child, childText!); }
const run = (...args: string[]) => {
const result = spawnSync('git', args, { cwd, encoding: 'utf8', timeout: 5000 });
if (result.status !== 0) throw new Error(`Caller fixture git ${args[0]} failed: ${result.stderr || result.error?.message}`);
@@ -568,8 +598,17 @@ process.exit(exit ?? 127);
GIT_ALTERNATE_OBJECT_DIRECTORIES: '',
};
fs.cpSync(fs.realpathSync(path.join(cwd, '.git/objects')), gitEnvironment.GIT_OBJECT_DIRECTORY, { recursive: true });
const startObservations = [
`git rev-parse --short HEAD: ${run('rev-parse', '--short', 'HEAD')}`,
`git log origin/main..HEAD --oneline: ${run('log', 'origin/main..HEAD', '--oneline') || '(no commits)'}`,
`git ls-files --others --exclude-standard: ${run('ls-files', '--others', '--exclude-standard') || '(none)'}`,
`git status --short: ${run('status', '--short')}`,
`git diff origin/main --stat:\n${run('diff', 'origin/main', '--stat')}`,
`git diff origin/main:\n${run('diff', 'origin/main')}`,
'reports/ and .qa-state/ contain no files yet.',
].join('\n');
let reviewStart: string | undefined;
if (caller === 'review') {
{
const start = spawnSync('bash', [path.join(QA_CALLER_ROOT, 'bin/gstack-review-log'), '--start', 'review'], {
cwd, env: { ...process.env, ...gitEnvironment, GSTACK_HOME: state, GSTACK_STATE_ROOT: state },
encoding: 'utf8', timeout: 5000,
@@ -601,7 +640,7 @@ process.exit(exit ?? 127);
const snapshot = () => callerSnapshot({ ...Object.fromEntries(productFiles.map(file => [file, fs.readFileSync(path.join(cwd, file), 'utf8')])), 'fixture.json': fs.readFileSync(fixtureInput, 'utf8') });
const probes = () => fs.readFileSync(journal, 'utf8').split('\n').filter(Boolean).map(line => JSON.parse(line) as CallerProbe);
const fixture: QaCallerFixture = {
root, cwd, state, runtime, config, caller, caseId, instructions, journal, mutationEvents, observerErrors, workflowCommands, reviewStart, gitEnvironment,
root, cwd, state, runtime, config, caller, caseId, instructions, journal, child, mutationEvents, observerErrors, workflowCommands, reviewStart, startObservations, gitEnvironment,
lateApplied: false, snapshot, probes,
observe: async () => {
if (observer || fixture.observation) throw new Error('Caller observation cannot restart mid-capture');
@@ -633,14 +672,21 @@ process.exit(exit ?? 127);
}
}
export function callerReviewRecordTemplate(fixture: Pick<QaCallerFixture, 'caller' | 'runtime'>): string {
const source = fs.readFileSync(path.join(QA_CALLER_ROOT, fixture.caller === 'review' ? 'review/SKILL.md' : 'ship/sections/review-army.md'), 'utf8');
const templates = source.match(/^~\/\.claude\/skills\/gstack\/bin\/gstack-review-log '\{"skill":"review",.*$/gm) ?? [];
if (templates.length !== 1) throw new Error(`Expected one installed /${fixture.caller} review record template, found ${templates.length}`);
return templates[0].replace('~/.claude/skills/gstack', fixture.runtime);
}
export function qaCallerSessionOptions(fixture: QaCallerFixture, runId: string): Parameters<typeof runSkillTest>[0] {
return {
prompt: `Load gstack's /${fixture.caller} supplied parent phase from caller-${fixture.caller}.md and resume it on the selected working-tree diff against origin/main. This excerpt comes from ${fixture.runtime}/${fixture.caller}/SKILL.md; resolve installed-relative references there, not from the excerpt file or product directory. That path identifies the asset base, not another entrypoint: do not read or invoke the full parent SKILL.md or rerun its preamble. Earlier preamble/branch/base setup is complete; use the existing local origin/main ref without fetch. Earlier-phase asset locators are ${fixture.runtime}/review/checklist.md and ${fixture.runtime}/qa/templates/functional-report-template.md. Read those files directly when referenced; recursive Glob does not follow the installed asset symlinks. Cross-project learnings are configured off in this owned fixture. ${fixture.reviewStart ? `The actual review-start helper already captured REVIEW_START=${fixture.reviewStart} for this unchanged core pass; retain that token. ` : ''} This fixture evaluates only the supplied parent phase, not later publication stages. Read README.md for the project contract and commands. Use diagnostic-client commands such as \`bun scripts/probe.ts <literal>\` for exploratory discoveries and their checkpoint evidence. A required \`bun run test\` is separate suite verification: report it as verification, never as a diagnostic observation or checkpoint anchor/target. Use the production evidence helper to publish each diagnostic checkpoint as \`reports/exploration-NNN.json\`, not inside a nested directory; do not transcribe its observed payload. ${fixture.caseId === 'ship-exploratory-plan-checks' ? 'The previously discovered plan is PLAN.md.' : 'No plan file was found.'} There is no remote service and no release publication is authorized. There is no interactive approver; do not invent answers or permission. Keep normal parent decision gates. Before every completion report or bookkeeping log, read HANDOFF.md and reports/HANDOFF.md if present for any concurrent collaborator update, await the results, and compare evidence with current inputs. A gstack-review-log completed:true record is a completion, not preliminary bookkeeping; a later handoff read cannot validate an earlier completion.\n\nDeadline bookkeeping additionally permits \`bun ${fixture.runtime}/bin/gstack-qa-deadline start ${fixture.cwd}/reports/deadline.json SECONDS [EARLIER_UTC]\`, \`bun ${fixture.runtime}/bin/gstack-qa-deadline status ${fixture.cwd}/reports/deadline.json\`, and \`bun ${fixture.runtime}/bin/gstack-qa-deadline run ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These are closed literal forms: SECONDS must be positive and at most 300, EARLIER_UTC is the optional caller absolute deadline: use the section clock's Hard deadline UTC, never its Runner entry UTC, reserve-start time or a clock-read time. The child is only the existing diagnostic client with zero or one literal argument. Resolve these exact helper and state paths; do not use variables, another helper, another state file, nested wrappers, scripts, operators or substitutions. Only this helper may create or change reports/deadline.json and its .qa-deadline- temporary files; never use Write/Edit/MultiEdit on those paths. Record the full outer run command in checkpoints and evidence; keep the unchanged child JSON as observed, separate from prefixed guard diagnostics. A completed expired guard-run is not a probe or a pass: retain its unused checkpoint, report not-run coverage and do not restart the deadline. Keep the 12-probe smoke limit. Required suites and explicit plan checks are outside the bounded smoke budget, not permission to reset it.\n\nFunctional evidence uses the same production helper and existing diagnostic client: \`bun ${fixture.runtime}/bin/gstack-qa-evidence capture ${fixture.cwd}/reports NNN --public --deadline ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These diagnostic receipts are declared public/synthetic, so --public is approved; a fresh three-digit ID is required each time. Explicit plan probes outside the smoke budget may replace --deadline with --timeout-ms 10000; this does not reset or bypass the smoke deadline. Publish causal intent with \`bun ${fixture.runtime}/bin/gstack-qa-evidence checkpoint ${fixture.cwd}/reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command'\`; quote arguments literally. For complex quoting, Write only capture, observationCommand, hypothesis and nextCommand to reports/intent.json; publish with the same helper: checkpoint REPORT_ROOT NNN intent.json. Materialize is supported when evidence.json is required. Sources stay inside reports. Decide to execute the next probe before publishing its checkpoint, then await successful publication and dispatch that exact probe. If you defer an optional idea or stop exploration, do not publish a checkpoint for it; descri Line truncated
prompt: `Load gstack's /${fixture.caller} supplied parent phase from caller-${fixture.caller}.md and resume it on the selected working-tree diff against origin/main. This excerpt comes from ${fixture.runtime}/${fixture.caller}/SKILL.md; resolve installed-relative references there, not from the excerpt file or product directory. That path identifies the asset base, not another entrypoint: do not read or invoke the full parent SKILL.md or rerun its preamble. Earlier preamble/branch/base setup is complete; use the existing local origin/main ref without fetch. Earlier-phase asset locators are ${fixture.runtime}/review/checklist.md and ${fixture.runtime}/qa/templates/functional-report-template.md. Read those files directly when referenced; recursive Glob does not follow the installed asset symlinks. Cross-project learnings are configured off in this owned fixture. ${fixture.reviewStart ? `The actual review-start helper already captured REVIEW_START=${fixture.reviewStart} for this unchanged core pass; retain that token. ` : ''} This fixture evaluates only the supplied parent phase, not later publication stages. The fixture owner recorded these read-only observations when this phase began; they are current until an input changes, so use them instead of re-running those commands, and recheck freshness before completion outputs:\n${fixture.startObservations}\n Read README.md for the project contract and commands. Use diagnostic-client commands such as \`bun scripts/probe.ts <literal>\` for exploratory discoveries and their checkpoint evidence. A required \`bun run test\` is separate suite verification: report it as verification, never as a diagnostic observation or checkpoint anchor/target. Use the production evidence helper to publish each diagnostic checkpoint as \`reports/exploration-NNN.json\`, not inside a nested directory; do not transcribe its observed payload. ${fixture.caseId === 'ship-exploratory-plan-checks' ? 'The previously discovered plan is PLAN.md.' : 'No plan file was found.'} There is no remote service and no release publication is authorized. There is no interactive approver; do not invent answers or permission. Keep normal parent decision gates. Before the completion report and each completed:true review record, read HANDOFF.md and reports/HANDOFF.md if present for any concurrent collaborator update, await the results, and compare evidence with current inputs. If a probe's snapshot differs from an earlier probe's, read reports/HANDOFF.md first; it names any collaborator change. A gstack-review-log completed:true record is a completion, not preliminary bookkeeping; a later handoff read cannot validate an earlier completion.\n\nDeadline bookkeeping additionally permits \`bun ${fixture.runtime}/bin/gstack-qa-deadline start ${fixture.cwd}/reports/deadline.json SECONDS [EARLIER_UTC]\`, \`bun ${fixture.runtime}/bin/gstack-qa-deadline status ${fixture.cwd}/reports/deadline.json\`, and \`bun ${fixture.runtime}/bin/gstack-qa-deadline run ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These are closed literal forms: SECONDS must be positive and at most 300, EARLIER_UTC is the optional caller absolute deadline: use the section clock's Hard deadline UTC, never its Runner entry UTC, reserve-start time or a clock-read time. The child is only the existing diagnostic client with zero or one literal argument. Resolve these exact helper and state paths; do not use variables, another helper, another state file, nested wrappers, scripts, operators or substitutions. Only this helper may create or change reports/deadline.json and its .qa-deadline- temporary files; never use Write/Edit/MultiEdit on those paths. Record the full outer run command in checkpoints and evidence; keep the unchanged child JSON as observed, separate from prefixed guard diagnostics. A completed expired guard-run is not a probe or a pass: retain its unused checkpoint, report not-run coverage and do not restart the deadline. Keep the 12-probe smoke limit. Required suites and explicit plan checks are outside the bounded smoke budget, not permission to reset it.\n\nFunctional evidence uses the same production helper and existing diagnostic client: \`bun ${fixture.runtime}/bin/gstack-qa-evidence capture ${fixture.cwd}/reports NNN --public --deadline ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These diagnostic receipts are declared public/synthetic, so --public is approved; a fresh three-digit ID is required each time. Explicit plan probes outside the smoke budget may replace --deadline with --timeout-ms 10000; this does not reset or bypass the smoke deadline. After the first capture, put causal intent in the next capture itself by adding \`--after PREV --hypothesis 'causal hypothesis'\` before \`--\` (PREV = the most recent complete capture ID); the helper publishes reports/exploration-NNN.json linking capture PREV's observation to that probe before running it, so one call is both checkpoi Line truncated
appendSystemPrompt: `Caller execution scheduling (fixture contract):
This session has at most 25 assistant turns, including required verification and final artifacts. The command boundary applies to each Bash call, not to the number of independent tool calls in an assistant turn.
After required clock and approval prerequisites settle, issue independent source Reads and read-only discovery together as separate native tool calls once their paths and inputs are known. Wait for their results before decisions that depend on them.
After required clock and approval prerequisites settle, issue independent source Reads and read-only discovery together as separate native tool calls once their paths and inputs are known. For example, after the first clock read, one response can Read the caller excerpt, README.md, every product file and the referenced checklist and templates. Wait for their results before decisions that depend on them.
The completion reserve is for required verification, affected-input revalidation and artifacts, not an earlier deadline. Keep completing required work within the actual remaining deadline; reserve entry alone is not a reason to stop. Use each native probe's snapshot to distinguish current from superseded evidence before deciding which checks still need revalidation.
Never group diagnostic probes, checkpoint publication with its next probe, or any action with the clock/status/approval result it needs. Shell composition remains forbidden outside the declared forms. Preserve every required Read, probe, verification, freshness check, approval and report field; the turn limit does not authorize skipping work or reporting incomplete work as passed.`,
Never group diagnostic probes, checkpoint publication with its next probe (a capture carrying --after is one probe, not a group), or any action with the clock/status/approval result it needs. Shell composition remains forbidden outside the declared forms. Preserve every required Read, probe, verification, freshness check, approval and report field; the turn limit does not authorize skipping work or reporting incomplete work as passed.`,
workingDirectory: fixture.cwd,
timeout: CAPTURE_MS,
completionReserveMs: CAPTURE_MS / 4,
+43 -10
View File
@@ -91,6 +91,8 @@ export function validateQACheckpoints(input: {
files: Record<string, string>;
reportMarkdown: string;
producer?: QaEvidenceContext;
/** Caller-authorized Bash that names checkpoints only as evidence citations, never as files to read or write. */
citesCheckpointsOnly?: (command: string) => boolean;
}): string[] {
const failures: string[] = [];
let disk: Record<string, string>;
@@ -111,13 +113,40 @@ export function validateQACheckpoints(input: {
else bound.push({ probe, call: matches[0] });
}
bound.sort((a, b) => a.call.start - b.call.start);
const notes: Array<{ name: string; call: Call; value: Record<string, any>; intent?: Call }> = [];
const notes: Array<{ name: string; call: Call; value: Record<string, any>; intent?: Call; merged?: boolean }> = [];
const captureId = (call: Call) => qaNativeCapture(call, input.producer)?.command.id;
const written = new Set<string>();
for (const call of calls) {
const producerCommand = call.name === 'Bash' ? qaEvidenceCommand(call.input.command, input.producer) : undefined;
let attempted = call.input.file_path;
let content = call.input.content;
let intent: Call | undefined;
const merged = producerCommand?.action === 'capture' && !!producerCommand.after;
if (merged) {
const after = producerCommand!.after!;
const name = `exploration-${producerCommand!.id}.json`;
const producer = qaProducerReceipt(call, input.producer) ?? qaProducerReceipt(call, input.producer, 'incomplete');
if (!producer || typeof disk[name] !== 'string' || qaEvidenceHash(disk[name]) !== producer.receipt.checkpointSha256) {
failures.push(`Checkpoint lacks completed native producer and intent: ${name}`);
continue;
}
let published: any;
try { published = JSON.parse(disk[name]); } catch {}
const captures = bound.filter(row => row.call.parent === call.parent && row.call.end < call.start)
.map(row => ({ row, producer: qaNativeCapture(row.call, input.producer) })).filter(row => row.producer?.command.id === after.capture);
const capture = captures.length === 1 ? captures[0] : undefined;
const read = capture && (capture.producer!.command.publicOutput || calls.some(read => read.name === 'Read' && !read.failed && read.parent === call.parent
&& read.start > capture.row.call.end && read.end > read.start && read.end < call.start
&& read.file?.path === path.join(input.reportRoot, `.qa-evidence/${after.capture}/observation.json`)
&& read.file.content === capture.producer!.captured.observationText));
if (!capture || !read || !isDeepStrictEqual(published, { observationCapture: after.capture, observationArgv: capture.producer!.captured.receipt.argv,
observed: capture.row.probe.observed, hypothesis: after.hypothesis, nextCapture: producerCommand!.id, nextArgv: producerCommand!.argv })) {
failures.push(`Checkpoint intent lacks its completed observation read: ${name}`);
continue;
}
attempted = path.join(input.reportRoot, name);
content = disk[name];
}
if (producerCommand?.action === 'checkpoint') {
const producer = qaProducerReceipt(call, input.producer);
const name = `exploration-${producerCommand.id}.json`;
@@ -154,19 +183,20 @@ export function validateQACheckpoints(input: {
attempted = path.join(input.reportRoot, name);
content = disk[name];
}
if (call.name === 'Bash' && producerCommand?.action !== 'checkpoint' && typeof call.input.command === 'string' && /exploration-\d+\.json/.test(call.input.command)) {
if (call.name === 'Bash' && producerCommand?.action !== 'checkpoint' && typeof call.input.command === 'string' && /exploration-\d+\.json/.test(call.input.command)
&& !input.citesCheckpointsOnly?.(call.input.command)) {
failures.push('Unsupported checkpoint Bash interaction');
}
if (typeof attempted !== 'string' || !path.basename(attempted).startsWith('exploration-')) continue;
if (call.name === 'Read') continue;
const name = path.basename(attempted);
if ((call.name !== 'Write' && !intent) || !checkpointName.test(name) || attempted !== path.join(input.reportRoot, name)) {
if ((call.name !== 'Write' && !intent && !merged) || !checkpointName.test(name) || attempted !== path.join(input.reportRoot, name)) {
failures.push(`Unsupported checkpoint write/path: ${attempted}`);
continue;
}
if (written.has(name)) failures.push(`Reused or overwritten checkpoint: ${name}`);
written.add(name);
if (call.failed || call.end <= call.start) { failures.push(`Checkpoint Write did not complete successfully: ${name}`); continue; }
if (!merged && (call.failed || call.end <= call.start)) { failures.push(`Checkpoint Write did not complete successfully: ${name}`); continue; }
if (typeof content !== 'string' || !Object.hasOwn(disk, name) || disk[name] !== content) {
failures.push(`Checkpoint artifact differs from captured Write: ${name}`);
continue;
@@ -177,18 +207,19 @@ export function validateQACheckpoints(input: {
}
let value: unknown;
try { value = JSON.parse(content); } catch {}
if (!object(value) || !isDeepStrictEqual(Object.keys(value).sort(), ['hypothesis', 'nextCommand', 'observationCommand', 'observed'])
|| typeof value.hypothesis !== 'string' || value.hypothesis.trim().length <= 20
|| !/[a-z]{3}/i.test(value.hypothesis) || typeof value.observationCommand !== 'string' || typeof value.nextCommand !== 'string') {
if (!object(value) || typeof value.hypothesis !== 'string' || value.hypothesis.trim().length <= 20 || !/[a-z]{3}/i.test(value.hypothesis)
|| (merged ? !isDeepStrictEqual(Object.keys(value).sort(), ['hypothesis', 'nextArgv', 'nextCapture', 'observationArgv', 'observationCapture', 'observed'])
: !isDeepStrictEqual(Object.keys(value).sort(), ['hypothesis', 'nextCommand', 'observationCommand', 'observed'])
|| typeof value.observationCommand !== 'string' || typeof value.nextCommand !== 'string')) {
failures.push(`Invalid checkpoint schema: ${name}`);
continue;
}
notes.push({ name, call, value, ...(intent ? { intent } : {}) });
notes.push({ name, call, value, ...(intent ? { intent } : {}), ...(merged ? { merged } : {}) });
}
for (const name of Object.keys(disk)) if (!written.has(name)) failures.push(`Checkpoint has no public Write: ${name}`);
const additional: Array<{ command: string; call: Call }> = [];
for (const target of input.additionalTargets ?? []) {
if (!notes.some(note => note.value.nextCommand === target.command)) continue;
if (!notes.some(note => note.value.nextCommand === target.command || note.merged && note.call.input.command === target.command && note.call.output === target.output)) continue;
const matches = calls.filter(call => call.name === 'Bash' && call.input.command === target.command
&& call.end > call.start && call.output === target.output);
if (matches.length !== 1 || bound.some(row => row.call === matches[0]) || additional.some(row => row.call === matches[0])) {
@@ -198,7 +229,9 @@ export function validateQACheckpoints(input: {
const owners = new Map<Call, typeof notes>();
for (const target of [...bound.map(row => ({ command: row.probe.command, call: row.call })), ...additional]) {
const previous = bound.filter(row => row.call.parent === target.call.parent && row.call.start < target.call.start).at(-1);
owners.set(target.call, notes.filter(note => previous && note.call.parent === target.call.parent
owners.set(target.call, notes.filter(note => previous && note.call.parent === target.call.parent && note.merged
? note.call === target.call && note.value.observationCapture === captureId(previous.call) && isDeepStrictEqual(note.value.observed, previous.probe.observed)
: previous && note.call.parent === target.call.parent && !note.merged
&& note.call.start > previous.call.end && note.call.end < target.call.start
&& (!note.intent || note.intent.start > previous.call.end)
&& note.value.observationCommand === previous.probe.command && isDeepStrictEqual(note.value.observed, previous.probe.observed)
+9 -1
View File
@@ -29,6 +29,7 @@ export interface QaEvidenceCommand {
timeoutMs?: number;
publicOutput?: boolean;
intent?: { capture: string; observationCommand: string; hypothesis: string; nextCommand: string };
after?: { capture: string; hypothesis: string };
}
export function qaEvidenceCommand(command: string, context?: QaEvidenceContext): QaEvidenceCommand | undefined {
@@ -51,13 +52,19 @@ export function qaEvidenceCommand(command: string, context?: QaEvidenceContext):
intent: { capture: tokens[5], observationCommand: tokens[6], hypothesis: tokens[7], nextCommand: tokens[8] } };
const publicOutput = tokens[5] === '--public';
const option = publicOutput ? 6 : 5;
const after = tokens[option + 2] === '--after' && /^\d{3}$/.test(tokens[option + 3] ?? '') && tokens[option + 4] === '--hypothesis'
? { capture: tokens[option + 3], hypothesis: tokens[option + 5] ?? '' } : undefined;
if (after) {
tokens.splice(option + 2, 4);
matches.splice(option + 2, 4);
}
if (tokens[2] !== 'capture' || tokens[option + 2] !== '--' || tokens.length < option + 4) return;
const deadline = tokens[option] === '--deadline' ? source(path.resolve(context.cwd, tokens[option + 1])) : undefined;
if (tokens[option] === '--deadline' && !deadline) return;
if (tokens[option] === '--timeout-ms' && (!/^[1-9]\d*$/.test(tokens[option + 1]) || Number(tokens[option + 1]) > 2_147_483_647)) return;
if (!['--deadline', '--timeout-ms'].includes(tokens[option])) return;
return { action: 'capture', id: tokens[4], publicOutput, argv: tokens.slice(option + 3), nativeCommand: command.slice(matches[option + 3].index).trim(),
...(tokens[option] === '--deadline' ? { deadline } : { timeoutMs: Number(tokens[option + 1]) }) };
...(tokens[option] === '--deadline' ? { deadline } : { timeoutMs: Number(tokens[option + 1]) }), ...(after ? { after } : {}) };
}
export type QaProducerCall = { name: string; input: Record<string, any>; output: string; failed: boolean; start: number; end: number };
@@ -75,6 +82,7 @@ export function qaProducerReceipt(call: QaProducerCall, context?: QaEvidenceCont
|| receipt.status !== status || !/^[a-f0-9]{64}$/.test(receipt.sha256)
|| receipt.exitCode !== exitCode
|| (command.id !== undefined && receipt.id !== command.id)
|| (command.after ? receipt.checkpoint !== command.id || !/^[a-f0-9]{64}$/.test(receipt.checkpointSha256) : receipt.checkpoint !== undefined)
|| (call.failed && (command.action !== 'capture' || receipt.exitCode === 0))) return;
return { command, receipt };
} catch { return; }
+12 -8
View File
@@ -10,7 +10,7 @@ import { runRecordedOfficeHoursAttempt, OFFICE_HOURS_BUN_GRACE_MS } from './offi
import { resolveEvalModel } from '../../lib/eval-model';
import { createQAFunctionalFixture, fixtureGit, ownedPath, qaFixtureActor, QA_TOOLS, type QAFamily, type QAMode } from './qa-functional-fixture';
import { observeQAWrites, type QAWriteObservation } from './qa-functional-observer';
import { qaFunctionalVerdict, verifyQANativeRegression, preserveQAArtifact, qaCaptureArtifacts } from './qa-functional-evidence';
import { QA_WEBHOOK_REQUIRED_SCENARIOS, qaFunctionalVerdict, verifyQANativeRegression, preserveQAArtifact, qaCaptureArtifacts } from './qa-functional-evidence';
import { QA_EVIDENCE_RUNTIME, qaEvidenceCommand, qaProducerReceipt, qaEvidenceHash } from './qa-evidence-producer';
import { nativeCalls } from './qa-checkpoint-evidence';
@@ -54,15 +54,15 @@ Fixture execution boundary:
- Reports belong only in existing qa-reports. Retain fixture state; the owner cleans it after preserving evidence. Authorized source/test edits use Write/Edit. Read/Glob/Grep support arbitrary read-only discovery, including directory/path inventory.
- Bash accepts separate literal commands only: no shell composition, scripts or added path operands. Read-only forms are pwd, ls, ls -la, git status --short, git status --porcelain, git branch --show-current, git diff, git diff --stat, git rev-parse HEAD, bun --version, and exactly date -u +%Y-%m-%dT%H:%M:%SZ. Native tests use bun test with optional named test/*.test.ts selectors.
- These are complete command forms, not general shell examples. For inventory inside a named directory, use Read/Glob/Grep; the listed ls forms inspect only the working directory. Do not add operands or flags beyond the declared forms, even for read-only discovery.
- The installed production helper is bin/gstack-qa-evidence (absolute owned path also accepted). Probe outputs are declared public/synthetic, so capture with: bun bin/gstack-qa-evidence capture qa-reports NNN --public --timeout-ms 10000 -- NATIVE_PROBE. The child must be one of the observation forms below. Each execution/replay gets a fresh three-digit ID. The helper does not authorize another command, interpreter, path, pipeline or redirect.
- Publish causal intent from the most recent completed native probe with bun bin/gstack-qa-evidence checkpoint qa-reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command'; quote each argument literally. For complex quoting, Write only capture, observationCommand, hypothesis and nextCommand to qa-reports/intent.json, then use bun bin/gstack-qa-evidence checkpoint qa-reports NNN intent.json. Wait for successful publication before dispatching the exact next command. Only the helper writes observed fields.
- The installed production helper is bin/gstack-qa-evidence (absolute owned path also accepted); \`bun bin/gstack-qa-evidence --help\` prints its usage, and the helper source is not part of the task. Probe outputs are declared public/synthetic, so capture with: bun bin/gstack-qa-evidence capture qa-reports NNN --public --timeout-ms 10000 -- NATIVE_PROBE. The child must be one of the observation forms below. Each execution/replay gets a fresh three-digit ID. The helper does not authorize another command, interpreter, path, pipeline or redirect.
- After the first capture, put causal intent in the next capture itself: bun bin/gstack-qa-evidence capture qa-reports NNN --public --timeout-ms 10000 --after PREV --hypothesis 'causal hypothesis' -- NATIVE_PROBE, where PREV is the capture ID of the most recent completed native probe. Before running the probe, the helper publishes qa-reports/exploration-NNN.json linking capture PREV's observation to it, so one call is both checkpoint and probe. The separate form stays valid: bun bin/gstack-qa-evidence checkpoint qa-reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command' (or an intent.json with only capture, observationCommand, hypothesis and nextCommand), then the exact next command after successful publication. Quote arguments literally. Only the helper writes observed fields.
- Write annotations.json inside qa-reports, then run bun bin/gstack-qa-evidence materialize qa-reports annotations.json to produce evidence.json before writing Markdown. Annotations have revision, runtime, cwd, evidence rows {capture,command,contract,expected,classification}, learning (selected checkpoint IDs) and limits; omit observed, which the helper supplies from captures. In each evidence row, capture is the three-digit capture ID and command is the exact full outer capture invocation, including that ID and all wrapper options, not just the native child command after --. This same full-command definition applies to observationCommand and nextCommand. Select a checkpoint whose next native command differs, not a same-command replay with a new capture ID. Retain all required safe observations and every executed probe.
- ${entry.family === 'cli' ? 'CLI observation forms: bun run probe -- balance; bun run probe -- export; bun run probe -- apply with zero to three literal arguments; bun cancel.ts. The equivalent bun run cli commands may be diagnostic but do not emit probe JSON. Arguments use ASCII letters/digits/._+- or quoted forms including spaces. The generic wrapper does NOT support wait: the only bounded wait/cancellation interface is bun cancel.ts.' : `Webhook observation form: bun run probe -- followed by one of happy, reject, duplicate, partial, concurrent-ab, concurrent-ba, cancel, dependency. ${entry.mode === 'qa-only' ? 'All eight scenarios are required coverage; a replay does not replace another scenario. ' : ''}Choose their order from observations after the happy path. bun cancel.ts is a CLI-only entrypoint, not part of this fixture.`}
Actions outside this interface are unsupported and fail acceptance; they are not implicitly approved.
Materialize qa-reports/evidence.json first, then write a concise qa-reports/report.md using the functional report structure. Link the evidence and checkpoint files rather than repeating full probe payloads in Markdown. Both artifacts are required before completion. The resulting evidence.json schema is:
{ "revision": "<full 40-character git rev-parse HEAD>", "runtime": "bun <version>", "cwd": "<working directory>", "evidence": [{"command":"<exact full outer capture invocation>","contract":"README.md","expected":"<declared expected behavior>","classification":"pass|product-defect|setup-blocked|inconclusive","observed":<complete unchanged JSON emitted by the native probe>}], "learning":[{"observationCommand":"<earlier full capture invocation>","hypothesis":"<what it taught you to challenge>","nextCommand":"<later full capture invocation>"}], "limits":["<untested or blocked coverage>"] }
Evidence rows contain ONLY complete JSON actually emitted by native probes, including failures and repeats; retain pre-repair results alongside green results. Never synthesize JSON from a tool error or raw test output. Put tests, raw CLI diagnostics, launch failures and timeouts in Markdown with their actual output and limits. The learning array is a summary: choose one completed checkpoint where an observation motivated a different later command, not the required same-command replay. Select that checkpoint ID in annotations.learning; the production helper copies its observationCommand, hypothesis and nextCommand. Both commands must name exact captured probes with different native child commands, never a combined command list or a replay distinguished only by capture ID. This selects existing exploration evidence, not another probe or a duplicate of the complete checkpoint ledger. Preserve every checkpoint and link every checkpoint in Markdown; keep every executed probe and its complete JSON in evidence, including the required replay. Missing dependencies remain setup blockers, not repairs. No browser installation or execution is needed.`;
{ "revision": "<full 40-character git rev-parse HEAD>", "runtime": "bun <version>", "cwd": "<working directory>", "evidence": [{"command":"<exact full outer capture invocation>","contract":"README.md","expected":"<declared expected behavior>","classification":"pass|product-defect|setup-blocked|inconclusive","observed":<complete unchanged JSON emitted by the native probe>}], "learning":[<rows the helper copies from the selected checkpoints>], "limits":["<untested or blocked coverage>"] }
Evidence rows contain ONLY complete JSON actually emitted by native probes, including failures and repeats; retain pre-repair results alongside green results. Never synthesize JSON from a tool error or raw test output. Put tests, raw CLI diagnostics, launch failures and timeouts in Markdown with their actual output and limits. The learning array is a summary: choose one completed checkpoint where an observation motivated a different later command, not the required same-command replay. Select that checkpoint ID in annotations.learning; the production helper copies its observationCommand, hypothesis and nextCommand, or a merged note's capture IDs, argv and hypothesis. Both commands must name exact captured probes with different native child commands, never a combined command list or a replay distinguished only by capture ID. This selects existing exploration evidence, not another probe or a duplicate of the complete checkpoint ledger. Preserve every checkpoint and link every checkpoint in Markdown; keep every executed probe and its complete JSON in evidence, including the required replay. Missing dependencies remain setup blockers, not repairs. No browser installation or execution is needed.`;
}
export async function runQAFunctionalCase(entry: { id: string; family: QAFamily; mode: QAMode }, collector: EvalCollector | null) {
@@ -103,7 +103,7 @@ export async function runQAFunctionalCase(entry: { id: string; family: QAFamily;
fixtureGit(fixture.root, ['commit', '-m', 'Bind current QA instructions to fixture']);
fixture.revision = fixtureGit(fixture.root, ['rev-parse', 'HEAD']);
if (fixtureGit(fixture.root, ['status', '--porcelain'])) throw new Error('QA fixture must start clean');
observer = await observeQAWrites(fixture.root, { evidenceProducer: true });
observer = await observeQAWrites(fixture.root, { evidenceProducer: true, atomicWriteMode: entry.mode });
await runRecordedOfficeHoursAttempt({
collector, name: entry.id, suite: 'Functional QA native E2E',
model: process.env.EVALS_MODEL ?? resolveEvalModel('capture'),
@@ -115,7 +115,8 @@ export async function runQAFunctionalCase(entry: { id: string; family: QAFamily;
workingDirectory: fixture.root, maxTurns: 40, allowedTools: QA_TOOLS, tools: QA_TOOLS,
timeout, completionReserveMs: timeout / 4,
testName: entry.id, runId, signal, env: { CLAUDE_CONFIG_DIR: fixture.config,
GIT_OPTIONAL_LOCKS: '0', QA_STATE_ROOT: path.join(fixture.root, '.qa-state') },
GIT_OPTIONAL_LOCKS: '0', QA_STATE_ROOT: path.join(fixture.root, '.qa-state'),
...(entry.family === 'webhook' ? { GSTACK_QA_REQUIRED_PROBES: JSON.stringify(QA_WEBHOOK_REQUIRED_SCENARIOS[entry.mode].map(scenario => `bun run probe -- ${scenario}`)) } : {}) },
});
return result;
},
@@ -132,7 +133,10 @@ export async function runQAFunctionalCase(entry: { id: string; family: QAFamily;
const context = { cwd: fixture.root, reportRoot: path.join(fixture.root, 'qa-reports'), executable: path.join(fixture.root, 'bin/gstack-qa-evidence') };
const calls = nativeCalls(captured.transcript, failures);
for (const action of ['capture', 'checkpoint', 'materialize']) {
if (!calls.some(call => qaProducerReceipt(call, context)?.command.action === action)) failures.push(`missing completed production ${action}`);
if (!calls.some(call => {
const command = qaProducerReceipt(call, context)?.command;
return command?.action === action || action === 'checkpoint' && !!command?.after;
})) failures.push(`missing completed production ${action}`);
}
if (!calls.some(call => {
const producer = qaProducerReceipt(call, context);
+15 -4
View File
@@ -7,6 +7,12 @@ import { readQACheckpointFiles, validateQACheckpoints } from './qa-checkpoint-ev
import { nativeCalls } from './qa-checkpoint-evidence';
import { qaNativeCapture } from './qa-evidence-producer';
/** Webhook scenarios a run must observe, per mode; the verdict and the fixture's capture nudge share this list. */
export const QA_WEBHOOK_REQUIRED_SCENARIOS: Record<QAMode, string[]> = {
'qa-only': ['happy', 'reject', 'duplicate', 'partial', 'concurrent-ab', 'concurrent-ba', 'cancel', 'dependency'],
qa: ['happy', 'cancel', 'dependency'],
};
const canonical = (value: any): string => JSON.stringify(value && typeof value === 'object'
? Array.isArray(value) ? value.map(item => JSON.parse(canonical(item)))
: Object.fromEntries(Object.keys(value).sort().map(key => [key, JSON.parse(canonical(value[key]))])) : value) ?? 'null';
@@ -15,14 +21,14 @@ const passingOutput = (text: string) => /\b[1-9]\d* pass\b/.test(text) && /\b0 f
export function qaNativeProbes(result: Pick<SkillTestResult, 'toolCalls'> & Partial<Pick<SkillTestResult, 'transcript'>>, root?: string) {
const calls = root && result.transcript ? nativeCalls(result.transcript, []) : [];
return result.toolCalls.flatMap<{ index: number; command: string; nativeCommand?: string; observed: any }>((call, index) => {
return result.toolCalls.flatMap<{ index: number; command: string; nativeCommand?: string; capture?: string; observed: any }>((call, index) => {
if (root && call.tool === 'Bash') {
const native = calls.filter(native => native.name === 'Bash' && native.input.command === call.input?.command && native.output === call.output);
const producer = native.length === 1 ? qaNativeCapture(native[0], { cwd: root, reportRoot: path.join(root, 'qa-reports'), executable: path.join(root, 'bin/gstack-qa-evidence') }) : undefined;
if (producer && /^bun (?:run probe -- |cancel\.ts$)/.test(producer.command.nativeCommand!)) {
const observed = producer.captured.observed as any;
if (observed && (Array.isArray(observed.args) || typeof observed.scenario === 'string'
|| producer.command.nativeCommand === 'bun cancel.ts' && Object.hasOwn(observed, 'exit'))) return [{ index, command: call.input.command, nativeCommand: producer.command.nativeCommand!, observed }];
|| producer.command.nativeCommand === 'bun cancel.ts' && Object.hasOwn(observed, 'exit'))) return [{ index, command: call.input.command, nativeCommand: producer.command.nativeCommand!, capture: producer.command.id!, observed }];
}
}
if (call.tool !== 'Bash' || !/^bun (?:run probe -- |cancel\.ts$)/.test(call.input?.command ?? '')) return [];
@@ -126,7 +132,7 @@ export function qaFunctionalVerdict(fixture: QAFunctionalFixture, mode: QAMode,
const cancellations = probes.filter(probe => fixture.family === 'cli' ? nativeCommand(probe) === 'bun cancel.ts' : probe.observed.scenario === 'cancel');
if (!cancellations.length) failures.push('missing cancellation observation');
if (fixture.family === 'webhook') {
for (const scenario of mode === 'qa-only' ? ['happy', 'reject', 'duplicate', 'partial', 'concurrent-ab', 'concurrent-ba'] : ['happy']) {
for (const scenario of QA_WEBHOOK_REQUIRED_SCENARIOS[mode].filter(scenario => !['cancel', 'dependency'].includes(scenario))) {
if (!probes.some(probe => probe.observed.scenario === scenario)) failures.push(`missing native ${scenario} probe`);
}
} else if (!probes.some(probe => probe.observed.args?.[0] === 'apply' && qaProbeClassification(probe.observed) === 'pass')) failures.push('missing adjacent valid CLI apply');
@@ -158,8 +164,13 @@ export function qaFunctionalVerdict(fixture: QAFunctionalFixture, mode: QAMode,
for (const row of report?.evidence ?? []) {
if (!probes.some(probe => row.command === probe.command && canonical(row.observed) === canonical(probe.observed))) failures.push('report invented an executed probe');
}
const learnedNote = (row: any) => {
try { return JSON.parse(checkpointFiles[`exploration-${row.nextCapture}.json`]).hypothesis === row.hypothesis; } catch { return false; }
};
if (!report?.learning?.some(row => typeof row.hypothesis === 'string' && row.hypothesis.trim().length > 20
&& probes.some(previous => previous.command === row.observationCommand && probes.some(next => next.index > previous.index && next.command === row.nextCommand && nativeCommand(next) !== nativeCommand(previous))))) failures.push('missing observation-to-next-hypothesis evidence');
&& probes.some(previous => (typeof row.observationCapture === 'string' ? previous.capture === row.observationCapture : previous.command === row.observationCommand)
&& probes.some(next => next.index > previous.index && nativeCommand(next) !== nativeCommand(previous)
&& (typeof row.nextCapture === 'string' ? next.capture === row.nextCapture && learnedNote(row) : next.command === row.nextCommand))))) failures.push('missing observation-to-next-hypothesis evidence');
const publicText = JSON.stringify(report) + result.output + reportMarkdown + JSON.stringify(checkpointFiles);
if (publicText.includes(QA_PRIVATE_SENTINEL)) failures.push('private sentinel leaked into published evidence');
if (mode === 'qa' && defect) {
+32 -3
View File
@@ -72,7 +72,7 @@ export interface QAWriteObservation {
limits: string[];
}
export async function observeQAWrites(root: string, options: { reportDirectory?: string; evidenceProducer?: boolean } = {}) {
export async function observeQAWrites(root: string, options: { reportDirectory?: string; evidenceProducer?: boolean; atomicTargets?: string[]; atomicWriteMode?: QAMode } = {}) {
if (process.platform !== 'linux') throw new Error('QA write observer unavailable: Linux inotify required');
if (fs.realpathSync(root) !== root) throw new Error('Observer root must be canonical');
let reportDirectory: string | undefined;
@@ -82,7 +82,10 @@ export async function observeQAWrites(root: string, options: { reportDirectory?:
reportDirectory = path.relative(root, directory);
}
const transientFile = (relative: string) => qaWriteAllowed(relative, 'qa-only')
|| (reportDirectory !== undefined && relative.startsWith(reportDirectory + path.sep));
|| (reportDirectory !== undefined && relative.startsWith(reportDirectory + path.sep))
|| (options.atomicTargets ?? []).some(target => relative.startsWith(target + '.tmp.') && /^\.tmp\.[1-9]\d*\.[0-9a-f]{12}$/.test(relative.slice(target.length)))
|| (options.atomicWriteMode !== undefined && /\.tmp\.[1-9]\d*\.[0-9a-f]{12}$/.test(relative)
&& qaWriteAllowed(relative.replace(/\.tmp\.[1-9]\d*\.[0-9a-f]{12}$/, ''), options.atomicWriteMode));
const before = qaTreeSnapshot(root);
const { dlopen, FFIType, ptr } = await import('bun:ffi');
const libc = dlopen('libc.so.6', {
@@ -95,6 +98,7 @@ export async function observeQAWrites(root: string, options: { reportDirectory?:
const events: QAWriteObservation['events'] = [];
const failures: string[] = [];
const publications = new Map<string, { temporary: string; dev: number; ino: number; bytes: string; parentDev: number; parentIno: number; mode: number }>();
const pendingLinks = new Map<string, { target: string; dev: number; ino: number; mode: number; bytes: string }>();
let stopped = false;
const observedPath = (relative: string, knownPair = false): string => {
try {
@@ -136,7 +140,7 @@ export async function observeQAWrites(root: string, options: { reportDirectory?:
|| !canonicalUTC(state.startedAt) || !canonicalUTC(state.deadlineAt)
|| Date.parse(state.deadlineAt) > Date.parse(state.startedAt) + state.budgetMs)) throw error;
if (!isDeadline && (!state || typeof state !== 'object' || Array.isArray(state))) throw error;
if (targetName.startsWith('exploration-') && Object.keys(state).sort().join(',') !== 'hypothesis,nextCommand,observationCommand,observed') throw error;
if (targetName.startsWith('exploration-') && !['hypothesis,nextCommand,observationCommand,observed', 'hypothesis,nextArgv,nextCapture,observationArgv,observationCapture,observed'].includes(Object.keys(state).sort().join(','))) throw error;
if (targetName === 'receipt.json' && (state.version !== 1 || !/^\d{3}$/.test(state.id) || !['complete', 'incomplete', 'sensitive'].includes(state.status))) throw error;
if (targetName === 'evidence.json' && (!Array.isArray(state.evidence) || !Array.isArray(state.limits))) throw error;
const final = fs.lstatSync(target);
@@ -150,6 +154,18 @@ export async function observeQAWrites(root: string, options: { reportDirectory?:
publications.set(relativeTarget, publication);
return target;
} catch (publicationError) {
if (receipt === undefined && (publicationError as NodeJS.ErrnoException).code === 'ENOENT' && temporaryName.test(basename)) {
let linking: number | undefined;
try { linking = fs.openSync(path.join(parent, basename), fs.constants.O_RDONLY | fs.constants.O_NOFOLLOW | fs.constants.O_NONBLOCK); }
catch (openError) { if ((openError as NodeJS.ErrnoException).code === 'ENOENT') return path.join(parent, basename); }
if (linking !== undefined) try {
const entry = fs.fstatSync(linking);
if (entry.isFile() && entry.nlink === 2 && entry.uid === parentStat.uid && (entry.mode & 0o777) === mode) {
pendingLinks.set(relative, { target: path.relative(root, target), dev: entry.dev, ino: entry.ino, mode, bytes: fs.readFileSync(linking, 'utf8') });
return path.join(parent, basename);
}
} finally { fs.closeSync(linking); }
}
try { return ownedPath(root, relative); } catch { throw publicationError; }
} finally { if (receipt !== undefined) fs.closeSync(receipt); }
}
@@ -241,6 +257,13 @@ export async function observeQAWrites(root: string, options: { reportDirectory?:
|| fs.readFileSync(target, 'utf8') !== publication.bytes) throw new Error('Evidence publication did not settle unchanged');
} catch (error) { failures.push(String(pathFailure(root, relative, error))); }
}
for (const [temporary, pending] of pendingLinks) {
try {
const entry = fs.lstatSync(ownedPath(root, pending.target));
if (fs.existsSync(ownedPath(root, temporary)) || entry.dev !== pending.dev || entry.ino !== pending.ino || entry.nlink !== 1
|| (entry.mode & 0o777) !== pending.mode || fs.readFileSync(ownedPath(root, pending.target), 'utf8') !== pending.bytes) throw new Error('Evidence publication did not settle unchanged');
} catch (error) { failures.push(String(pathFailure(root, temporary, error))); }
}
let after: Record<string, string> = {};
try { after = qaTreeSnapshot(root); } catch (error) { failures.push(String(error)); }
fs.closeSync(fd);
@@ -260,7 +283,13 @@ export function qaWriteVerdict(observation: QAWriteObservation, mode: QAMode): s
return failures;
}
/** Asking an installed gstack QA helper for its own usage text is read-only and always declared. */
export function qaHelperUsageCommand(command: string): boolean {
return /^bun (?:[\w./-]+\/)?bin\/gstack-qa-(?:evidence|deadline) --help$/.test(command.trim());
}
export function qaCommandAllowed(command: string, root?: string): boolean {
if (qaHelperUsageCommand(command)) return true;
const producer = root ? qaEvidenceCommand(command, { cwd: root, reportRoot: path.join(root, 'qa-reports'), executable: path.join(root, 'bin/gstack-qa-evidence') }) : undefined;
if (producer) {
try { ownedPath(root!, 'bin/gstack-qa-evidence'); } catch { return false; }
+4 -2
View File
@@ -7,7 +7,8 @@ import { isAgentRecordGone, isOurAgent, readAgentRecord } from '../../browse/src
export async function stopQaOnlyBrowser(directory: string, timeoutMs: number): Promise<void> {
if (!Number.isFinite(timeoutMs) || timeoutMs <= 100) throw new Error('QA-only browser cleanup: no settlement budget remains; retaining fixture');
const worker = Bun.spawn([process.execPath, import.meta.path, directory, String(timeoutMs - 100)], {
// An absolute deadline keeps worker startup inside its own budget, so it reports its reason before the kill.
const worker = Bun.spawn([process.execPath, import.meta.path, directory, String(Date.now() + timeoutMs - 100)], {
stdout: 'ignore', stderr: 'pipe',
});
let timedOut = false;
@@ -21,8 +22,9 @@ export async function stopQaOnlyBrowser(directory: string, timeoutMs: number): P
} finally { clearTimeout(timer); }
}
async function settleOwnedBrowser(directory: string, timeoutMs: number): Promise<void> {
async function settleOwnedBrowser(directory: string, deadlineEpochMs: number): Promise<void> {
const started = performance.now();
const timeoutMs = deadlineEpochMs - Date.now();
const deadline = started + timeoutMs;
let pending = ['identity verification'];
const fail = (message: string): never => { throw new Error(`QA-only browser cleanup: ${message}`); };
+3 -2
View File
@@ -293,7 +293,7 @@ Runner entry UTC: ${new Date(startTime).toISOString()}
Hard deadline UTC: ${new Date(deadline).toISOString()}
Completion reserve starts UTC: ${new Date(deadline - reserve).toISOString()}
Setup, CLI startup and API queueing consume this same window; it never resets.
Before source Reads and after each saved checkpoint, use Bash to run exactly \`date -u +%Y-%m-%dT%H:%M:%SZ\`. Compare that observed UTC time with the times above. When remaining time is at most ${reserve / 1000} seconds, prioritize the remaining required completion outputs and verification. No required content or gate may be skipped. If the clock read fails, report timing unavailable; do not invent remaining time or restart the deadline.`;
Before source Reads, use Bash to run exactly \`date -u +%Y-%m-%dT%H:%M:%SZ\`. After each saved checkpoint, compare the latest evidence capture's printed completedAt with the times above; run that clock read again only when no capture has completed since your last clock read. When remaining time is at most ${reserve / 1000} seconds, prioritize the remaining required completion outputs and verification. No required content or gate may be skipped. If the clock read fails, report timing unavailable; do not invent remaining time or restart the deadline.`;
systemPrompt = systemPrompt ? `${systemPrompt}\n\n${notice}` : notice;
}
@@ -696,7 +696,8 @@ Before source Reads and after each saved checkpoint, use Bash to run exactly \`d
}
// Cost from result line (exact) or estimate from chars
const turnsUsed = resultLine?.num_turns || 0;
const turnsUsed = resultLine?.num_turns
|| new Set(transcript.filter(event => event?.type === 'assistant' && !event.parent_tool_use_id).map(event => event.message?.id)).size;
const estimatedCost = resultLine?.total_cost_usd || 0;
const inputChars = prompt.length;
const outputChars = (resultLine?.result || '').length;
+13
View File
@@ -36,6 +36,19 @@ export function redactPublicValue(value: unknown, token: string, unsafe = () =>
return value;
}
/**
* Path 4 remote actor: the prompt directs Step 5a registration (the contract under
* test), so the MCP-registration question takes its register/recommended option.
* Every other gate (privacy, artifacts repo, per-remote policy) is declined or skipped.
*/
export function setupGbrainRemoteAnswer(q: { question: string; header?: string; options: Array<{ label: string }> }): string {
if (/typed tool surface|\bregister(?:s|ing)?\b[^?]*\bMCP\b|\bMCP\b[^?]*\bregist/i.test(`${q.header ?? ''}\n${q.question}`)) {
const accept = q.options.find(o => /^(?:yes|register)\b/i.test(o.label)) ?? q.options.find(o => /\(recommended\)/i.test(o.label));
if (accept) return accept.label;
}
return (q.options.find(o => /skip|decline|no thanks|local/i.test(o.label)) ?? q.options[q.options.length - 1]!).label;
}
/** Retain the prior Path 4 public projection; SDK private fields are never read. */
export function publicEvents(events: readonly unknown[]): unknown[] {
return events.flatMap((event: any) => {
+203 -22
View File
@@ -250,7 +250,7 @@ export function snapshotFixture(directory: string): Record<string, string> {
export interface SourceRequest {
tool: string; args: string[]; endpoint?: string; method?: string; cwd: string;
violation?: string;
env?: Record<string, string>; violation?: string;
pid?: number; ppid?: number; parentExecutable?: string; parentCommand?: string;
}
@@ -426,6 +426,8 @@ function sharedShellTokens(command: string): SharedShellToken[] {
return tokens;
}
const SHARED_OUTPUT_DEVICES = ['/dev/null', '/dev/stdout', '/dev/stderr', '/dev/fd/1', '/dev/fd/2'];
/** Share attempted-write checks across native, semantic and Codex standalone captures. */
export function sharedReadOnlyViolations(toolCalls: Array<{ tool: string; input: any }>, requests: SourceRequest[] = []): string[] {
const violations: string[] = [];
@@ -448,10 +450,17 @@ export function sharedReadOnlyViolations(toolCalls: Array<{ tool: string; input:
};
for (let i = 0; i < tokens.length; i++) {
const token = tokens[i].value;
if (tokens[i].operator && ['>', '>>', '&>'].includes(token) && !['/dev/null', '/dev/stdout', '/dev/stderr', '/dev/fd/1', '/dev/fd/2'].includes(tokens[i + 1]?.value))
if (tokens[i].operator && ['>', '>>', '&>'].includes(token) && !SHARED_OUTPUT_DEVICES.includes(tokens[i + 1]?.value))
violations.push('shell file output redirection');
if (tokens[i].operator && token === '>&' && !['1', '2', '-'].includes(tokens[i + 1]?.value)) violations.push('shell file output redirection');
if (isCommand(i) && /(?:^|\/)tee$/.test(token) && tokens[i + 1] && !tokens[i + 1].operator) violations.push('tee file output');
if (isCommand(i) && /(?:^|\/)tee$/.test(token)) {
// Like a redirection, tee may only duplicate to the discard/stdout devices; any other operand is a file.
const operands = sharedShellCommandTokens(tokens, i + 1).map(operand => operand.value);
const end = operands.indexOf('--');
const files = end < 0 ? operands.filter(value => !value.startsWith('-') || value === '-')
: [...operands.slice(0, end).filter(value => !value.startsWith('-') || value === '-'), ...operands.slice(end + 1)];
if (files.some(file => !SHARED_OUTPUT_DEVICES.includes(file))) violations.push('tee file output');
}
if (isCommand(i) && /(?:^|\/)curl$/.test(token)) {
// URL variables resolve only in the instrumented process. Its request
// record supplies endpoint validation; source text still reveals writes.
@@ -463,6 +472,22 @@ export function sharedReadOnlyViolations(toolCalls: Array<{ tool: string; input:
return [...new Set(violations)];
}
const SAFE_GIT_ENV = { GIT_OPTIONAL_LOCKS: '0', GIT_NO_LAZY_FETCH: '1', GIT_TERMINAL_PROMPT: '0' };
const SAFE_GIT_FLAGS = ['--no-pager', '--no-lazy-fetch', '--no-replace-objects'];
const SAFE_GIT_CONFIG = ['core.fsmonitor=false', 'log.showSignature=false', 'diff.submodule=short'];
/** A repository Git process that carries the complete safe prefix bin/gstack-safe-git applies, including its
* environment, plus patch-driver disabling for diffs; a partial literal prefix does not qualify. */
export function isGuardedGitRequest(request: SourceRequest): boolean {
const args = request.args, command = args.findIndex((arg, i) => !arg.startsWith('-') && args[i - 1] !== '-c' && args[i - 1] !== '-C');
const globals = command < 0 ? args : args.slice(0, command), rest = command < 0 ? [] : args.slice(command);
const configs = globals.flatMap((arg, i) => globals[i - 1] === '-c' ? [arg] : []);
return Object.entries(SAFE_GIT_ENV).every(([key, value]) => request.env?.[key] === value)
&& SAFE_GIT_FLAGS.every(flag => globals.includes(flag)) && SAFE_GIT_CONFIG.every(config => configs.includes(config))
&& !rest.some(arg => ['--ext-diff', '--textconv', '--output'].includes(arg) || arg.startsWith('--output='))
&& (rest[0] !== 'diff' || (rest.includes('--no-ext-diff') && rest.includes('--no-textconv')));
}
/** Claude's own workspace probes are not commands requested by the skill. */
export function isInternalClaudeGitRequest(request: SourceRequest, commands: string[]): boolean {
const hostPrefix = ['-c', 'protocol.ext.allow=never', '-c', 'submodule.recurse=false',
@@ -477,6 +502,81 @@ export function isInternalClaudeGitRequest(request: SourceRequest, commands: str
!commands.some(command => command.includes('core.safecrlf=false') || command.includes('protocol.ext.allow=never'));
}
/** Older open PRs outside the 14-day window: more than the skill's five 100-item open-metadata pages. */
export const SHARED_LIBS_OLDER_OPEN_PRS = 600;
/** One finite PR world, newest update first. Self-contained so the fixture executable uses this exact table. */
function sharedPullRequestTable(now: string, olderOpen: number) {
const hour = 3_600_000, newestOlder = Date.UTC(2025, 0, 1);
return [
{ number: 7, state: 'open', updated: now },
...Array.from({ length: olderOpen }, (_, i) => ({ number: 100 + olderOpen - 1 - i, state: 'open',
updated: new Date(newestOlder - i * hour).toISOString() })),
{ number: 42, state: 'open', updated: '2020-01-01T00:00:00Z' },
...[5, 4, 3].map((number, i) => ({ number, state: 'closed', updated: new Date(Date.UTC(2019, 5, 1) - i * hour).toISOString() })),
];
}
/** gh pr list, pulls?state= and search/issues as views of the same table; null for any other request. */
function sharedPullRequestView(table: Array<{ number: number; state: string; updated: string; title: string; body: string }>,
args: string[], endpoint: string) {
const [route, search = ''] = endpoint.replace(/^\//, '').split('?');
const params = new URLSearchParams(search);
const flag = (names: string[]) => {
for (let i = 0; i < args.length; i++) {
if (names.includes(args[i]!)) return args[i + 1];
const joined = names.find(name => name.startsWith('--') && args[i]!.startsWith(name + '='));
if (joined) return args[i]!.slice(joined.length + 1);
}
return undefined;
};
const matches = (query: string) => {
let state = '', rest = query;
const dates: Array<(pr: { updated: string }) => boolean> = [];
const words: string[] = [];
for (const term of rest.split(/\s+/).filter(Boolean)) {
const qualifier = /^(is|state|type|updated|created|repo):(.+)$/i.exec(term);
if (!qualifier) { words.push(term.replace(/^"|"$/g, '').toLowerCase()); continue; }
const [, key, value] = qualifier as unknown as [string, string, string];
if (/^(?:is|state)$/i.test(key) && /^(?:open|closed|merged)$/i.test(value)) state = value.toLowerCase();
else if (/^(?:is|type)$/i.test(key) && /^issue$/i.test(value)) return () => false;
else if (/^repo$/i.test(key) && value.toLowerCase() !== 'fixture/shared-libs') return () => false;
else if (/^(?:updated|created)$/i.test(key)) {
const range = /^(.+)\.\.(.+)$/.exec(value), op = /^(>=|<=|>|<)?(.+)$/.exec(value)!;
const at = (text: string) => Date.parse(text);
if (range) dates.push(pr => at(pr.updated) >= at(range[1]!) && at(pr.updated) <= at(range[2]!) + 86_399_999);
else dates.push(pr => { const t = at(pr.updated), v = at(op[2]!);
return op[1] === '>=' ? t >= v : op[1] === '>' ? t > v : op[1] === '<=' ? t <= v + 86_399_999 : op[1] === '<' ? t < v : t >= v && t <= v + 86_399_999; });
}
}
return (pr: { state: string; updated: string; title: string; body: string }) =>
(!state || pr.state === state) && dates.every(check => check(pr)) &&
words.every(word => (pr.title + ' ' + pr.body).toLowerCase().includes(word));
};
const page = (rows: typeof table, perPage: number, number: number) => {
const size = Math.min(100, Math.max(1, perPage || 30));
return rows.slice((Math.max(1, number || 1) - 1) * size, Math.max(1, number || 1) * size);
};
if (args[0] === 'pr' && args[1] === 'list') {
const state = (flag(['--state', '-s']) || 'open').toLowerCase();
const rows = table.filter(pr => state === 'all' || pr.state === state).filter(matches(flag(['--search', '-S']) || ''));
return rows.slice(0, Math.max(1, Number(flag(['--limit', '-L']) || 30)));
}
if (/^repos\/fixture\/shared-libs\/pulls$/.test(route!)) {
const state = (params.get('state') || 'open').toLowerCase();
const rows = table.filter(pr => state === 'all' || pr.state === state);
if (params.get('direction') === 'asc') rows.reverse();
return page(rows, Number(params.get('per_page')), Number(params.get('page')));
}
if (route === 'search/issues') {
const rows = table.filter(matches(params.get('q') || ''));
if (params.get('order') === 'asc') rows.reverse();
return { total_count: rows.length, incomplete_results: false,
items: page(rows, Number(params.get('per_page')), Number(params.get('page'))) };
}
return null;
}
export function installSourceShims(f: SharedLibsFixture, opts: {
unsupportedGit?: boolean; unavailableApi?: boolean; prCoverage?: boolean;
} = {}): void {
@@ -503,7 +603,8 @@ export function installSourceShims(f: SharedLibsFixture, opts: {
}
const common = `const fs=require('node:fs'), cp=require('node:child_process');\nconst a=process.argv.slice(2);\nconst trace=${JSON.stringify(f.trace)};\nconst parent={pid:process.pid,ppid:process.ppid};try{parent.parentExecutable=fs.readlinkSync('/proc/'+process.ppid+'/exe');parent.parentCommand=fs.readFileSync('/proc/'+process.ppid+'/cmdline','utf8').replaceAll('\\0',' ');}catch{try{const info=cp.spawnSync('ps',['-p',String(process.ppid),'-o','comm=','-o','args='],{encoding:'utf8',timeout:3_000});const line=(info.stdout||'').trim();parent.parentExecutable=line.split(/\\s+/)[0];parent.parentCommand=line;}catch{}}\n`;
fs.writeFileSync(path.join(f.bin, 'git'), `#!${nodeBin}\n${common}
fs.appendFileSync(trace,JSON.stringify({tool:'git',args:a,cwd:process.cwd(),...parent})+'\\n');
const gitEnv=Object.fromEntries(['GIT_OPTIONAL_LOCKS','GIT_NO_LAZY_FETCH','GIT_TERMINAL_PROMPT'].filter(k=>k in process.env).map(k=>[k,process.env[k]]));
fs.appendFileSync(trace,JSON.stringify({tool:'git',args:a,env:gitEnv,cwd:process.cwd(),...parent})+'\\n');
if (${!!opts.unsupportedGit} && a.some(x=>x==='--no-lazy-fetch')) { console.error('unknown option: --no-lazy-fetch'); process.exit(129); }
if(a.includes('ls-remote')) { console.log('ref: refs/heads/main\\tHEAD\\n${f.tip}\\tHEAD\\n${f.tip}\\trefs/heads/main'); process.exit(0); }
// The fixture remote is already current. Record fetch attempts without contacting a real repository.
@@ -567,24 +668,25 @@ const sources=${JSON.stringify(sources)};
const base={name:'main',sha:${JSON.stringify(f.tip)}};
const pr=(number,date,extra={})=>({number,state:'open',title:number===42?'Extract retry parsing into existing helper':'Routine documentation '+number,body:number===7?'Coordination: https://github.com/fixture/shared-libs/pull/42':'',created_at:date,updated_at:date,merged_at:null,createdAt:date,updatedAt:date,mergedAt:null,url:'https://github.com/fixture/shared-libs/pull/'+number,html_url:'https://github.com/fixture/shared-libs/pull/'+number,head:{sha:number===42?prHead:${JSON.stringify(f.tip)},ref:'feature-'+number},base:{sha:${JSON.stringify(f.tip)},ref:'main'},...extra});
const page=Number((endpoint.match(/[?&]page=(\\d+)/)||[])[1]||a[a.indexOf('-F')+1]?.match(/^page=(\\d+)/)?.[1]||1);
let out;
const view=${sharedPullRequestView.toString()};
const prTable=${!!opts.prCoverage}?(${sharedPullRequestTable.toString()})(now,${SHARED_LIBS_OLDER_OPEN_PRS}).map(row=>({...row,title:pr(row.number,row.updated).title,body:pr(row.number,row.updated).body})):[];
const toPr=row=>pr(row.number,row.updated,row.state==='closed'?{state:'closed',closed_at:row.updated}:{});
let out,listing;
if(a[0]==='auth')process.exit(0);
else if(a[0]==='repo') out={nameWithOwner:'fixture/shared-libs',defaultBranchRef:base,url:'https://github.com/fixture/shared-libs'};
else if(a[0]==='pr'&&a[1]==='list')out=${!!opts.prCoverage}?[pr(7,now),pr(42,old)]:[];
else if((listing=view(prTable,a,endpoint))!==null)out=Array.isArray(listing)?listing.map(toPr):{...listing,items:listing.items.map(row=>({...toPr(row),pull_request:{url:'https://api.github.com/repos/fixture/shared-libs/pulls/'+row.number}}))};
else if(a[0]==='pr'&&a[1]==='view')out=pr(Number(a[2])||42,Number(a[2])===7?now:old,{files:[{path:Number(a[2])===7?'docs/unrelated.md':'src/retry-worker.ts'}]});
else if(endpoint.includes('search/issues'))out={total_count:${opts.prCoverage ? 1 : 0},incomplete_results:false,items:${!!opts.prCoverage}?[pr(7,now)]:[]};
else if(endpoint.includes('/contents/')) { const p=decodeURIComponent(endpoint.split('/contents/')[1].split('?')[0]); const ref=decodeURIComponent((endpoint.match(/[?&]ref=([^&]+)/)||[])[1]||'');if(!Object.hasOwn(sources,ref))apiError(404,'unsupported or unpinned fixture revision');const source=sources[ref];if(!Object.hasOwn(source.files,p))apiError(404,'source unavailable');out={path:p,encoding:'base64',content:source.files[p],sha:source.blobs[p]}; }
else if(/\\/contents(?:\\/|\\?|$)/.test(endpoint)) { const p=decodeURIComponent(endpoint.split(/\\/contents/)[1].split('?')[0]).replace(/^\\/+|\\/+$/g,''); const ref=decodeURIComponent((endpoint.match(/[?&]ref=([^&]+)/)||[])[1]||'');if(!Object.hasOwn(sources,ref))apiError(404,'unsupported or unpinned fixture revision');const source=sources[ref];
if(Object.hasOwn(source.files,p))out={type:'file',name:p.split('/').pop(),path:p,encoding:'base64',content:source.files[p],sha:source.blobs[p]};
else { const prefix=p?p+'/':'';const names=[...new Set(Object.keys(source.files).filter(file=>file.startsWith(prefix)).map(file=>file.slice(prefix.length).split('/')[0]))].sort();if(!names.length)apiError(404,'source unavailable');out=names.map(name=>{const file=prefix+name;return Object.hasOwn(source.files,file)?{type:'file',name,path:file,sha:source.blobs[file]}:{type:'dir',name,path:file};}); } }
else if(/\\/pulls\\/42\\/files/.test(endpoint))out=page===1?Array.from({length:100},(_,i)=>({filename:'docs/coordination-'+i+'.md',status:'added',patch:'@@ -0,0 +1 @@\\n+Documentation coordination '+i+'.'})):page===2?[{filename:'src/retry-worker.ts',status:'modified',patch:${JSON.stringify(prPatch)}}]:[];
else if(/\\/pulls\\/\\d+\\/files/.test(endpoint))out=page===1?[{filename:'docs/unrelated.md',status:'modified',patch:'@@ -1 +1 @@\\n-old\\n+new'}]:[];
else if(/\\/pulls\\/42(?:\\?|$)/.test(endpoint))out=pr(42,old);
else if(endpoint.includes('/pulls')) {
if(!${!!opts.prCoverage})out=[];
else if(endpoint.includes('state=open'))out=Array.from({length:100},(_,i)=>pr((page-1)*100+i+40,old));
else out=page===1?[pr(7,now),pr(42,old)]:[];
}
else if(/\\/pulls\\/\\d+(?:\\?|$)/.test(endpoint)){const row=prTable.find(row=>row.number===Number(endpoint.match(/\\/pulls\\/(\\d+)/)[1]));if(!row)apiError(404,'Not Found');out=toPr(row);}
else if(endpoint.includes('/commits')){const isPrCommit=prHead!==${JSON.stringify(f.tip)}&&endpoint.includes(prHead);out=endpoint.includes('/commits/')?{sha:isPrCommit?prHead:${JSON.stringify(f.tip)},commit:{committer:{date:isPrCommit?old:now},message:'Fixture work'},files:Object.keys(isPrCommit?prFiles:files).map(filename=>({filename,status:'modified'}))}:[{sha:${JSON.stringify(f.tip)},commit:{committer:{date:now},message:'Fixture work'}}];}
else if(endpoint.includes('/branches/'))out={name:'main',commit:{sha:${JSON.stringify(f.tip)}}};
else out={default_branch:'main',full_name:'fixture/shared-libs',html_url:'https://github.com/fixture/shared-libs'};
else if(/^\\/?repos\\/fixture\\/shared-libs\\/?(?:\\?|$)/.test(endpoint))out={default_branch:'main',full_name:'fixture/shared-libs',html_url:'https://github.com/fixture/shared-libs'};
else apiError(404,'Not Found');
if(curl)curlResponse(out);
const qi=a.findIndex(x=>x==='--jq'||x==='-q');
if(qi>=0) {const r=cp.spawnSync('jq',['-r',a[qi+1]],{input:JSON.stringify(out),encoding:'utf8',timeout:30_000});process.stdout.write(r.stdout||'');process.stderr.write(r.stderr||'');process.exit(r.status??1);}
@@ -682,9 +784,10 @@ export function seedOpportunitySources(f: SharedLibsFixture): void {
export function standaloneInstructions(f: SharedLibsFixture, codex = false): string {
const source = codex ? path.join(SHARED_LIBS_ROOT, '.agents/skills/gstack-deslop-shared-libs') : path.join(SHARED_LIBS_ROOT, 'deslop-shared-libs');
// Resolve the installed helper to this checkout, as the hermetic runtime for the skill under test.
const text = extractSkillSections(source, [
'Scope and read-only boundary', 'Establish the reviewed source', 'Start with recent work', 'Evaluate candidates', 'Output',
]);
]).replaceAll(codex ? '~/.codex/skills/gstack' : '~/.claude/skills/gstack', SHARED_LIBS_ROOT);
const file = path.join(f.root, 'standalone-instructions.md');
fs.writeFileSync(file, text);
return file;
@@ -759,7 +862,57 @@ export interface SharedReviewStageActor {
verify(events: any[]): boolean;
}
export function reviewPrompt(f: SharedLibsFixture, instructions: string, specialistInput: string, resumed?: SharedReviewResume | Pick<SharedReviewStageActor, 'actorCommand'>): string {
export interface SharedLifecycleSeed {
token: string;
startedAt: string;
startWtree: string;
diffBase: string;
observation: string;
}
/** Pass 1's Step 3, executed once by the fixture with the real logger and Git: fetch is a no-op against the
* pinned remote, the start token precedes every read, and the saved observation is never rewritten. */
export function seedLifecycleFirstPass(f: SharedLibsFixture): SharedLifecycleSeed {
const observation = path.join(f.root, 'pass1-observation.md');
if (fs.existsSync(observation)) throw new Error('Pass 1 observation already seeded');
const env = { ...process.env, ...f.env, PATH: process.env.PATH, GSTACK_HOME: f.state };
const diffBase = fixtureGit(f, 'merge-base', 'origin/main', 'HEAD');
const token = execFileSync(path.join(SHARED_LIBS_ROOT, 'bin/gstack-review-log'), ['--start', 'review'],
{ cwd: f.repo, env, encoding: 'utf8', timeout: 30_000 }).trim();
const start = JSON.parse(fs.readFileSync(path.join(f.state, 'projects/fixture-shared-libs/.review-starts', `${token}.json`), 'utf8'));
const diff = fixtureGit(f, 'diff', '--no-ext-diff', '--no-textconv', diffBase);
const untracked = fixtureGit(f, 'ls-files', '--others', '--exclude-standard');
const tracked = fixtureGit(f, 'ls-files');
const flagged = fixtureGit(f, 'ls-files', '-v').split('\n').filter(line => line && !line.startsWith('H '));
const attributes = ['.gitattributes', '.git/info/attributes'].filter(file => fs.existsSync(path.join(f.repo, file)));
const records = execFileSync(path.join(SHARED_LIBS_ROOT, 'bin/gstack-review-read'), [], { cwd: f.repo, env, encoding: 'utf8', timeout: 30_000 });
const files = [...tracked.split('\n'), ...untracked.split('\n')].filter(Boolean).map(relative => {
const bytes = fs.readFileSync(path.join(f.repo, relative));
return `### ${relative} (${bytes.length} bytes, sha256 ${createHash('sha256').update(bytes).digest('hex')})\n\`\`\`\n${bytes.toString('utf8')}\`\`\``;
});
fs.writeFileSync(observation, `# Pass 1 Step 3 observation (fixture-captured once, in this order)
1. git fetch origin main --quiet: exit 0; origin/main is pinned at ${f.tip}.
2. DIFF_BASE=$(git merge-base origin/main HEAD) = ${diffBase}; branch ${fixtureGit(f, 'symbolic-ref', '--short', 'HEAD')}; HEAD ${fixtureGit(f, 'rev-parse', 'HEAD')}.
3. gstack-review-log --start review printed pass 1 REVIEW_START ${token} (started_at ${start.started_at}). Step 3 leaves it unused; only the final pass's token is finished.
4. Reads after the token:
- Untracked non-ignored files (git ls-files --others --exclude-standard): ${untracked || '(none)'}
- Attribute files: ${attributes.join(', ') || '(none: no .gitattributes or .git/info/attributes)'}; index flags other than H (git ls-files -v): ${flagged.join('; ') || '(none)'}
- Repository-local config (git config --local --list):\n${fixtureGit(f, 'config', '--local', '--list')}
- gstack-review-read:\n${records.trim()}
## git diff --no-ext-diff --no-textconv ${diffBase}
\`\`\`diff
${diff}
\`\`\`
## Repository files (tracked and untracked)
${files.join('\n\n')}
`, { mode: 0o600, flag: 'wx' });
return { token, startedAt: start.started_at, startWtree: start.wtree, diffBase, observation };
}
export function reviewPrompt(f: SharedLibsFixture, instructions: string, specialistInput: string, resumed?: SharedReviewResume | Pick<SharedReviewStageActor, 'actorCommand'>, seed?: SharedLifecycleSeed): string {
if (seed && !(resumed && 'actorCommand' in resumed)) throw new Error('A seeded first pass belongs to the edit-capable lifecycle replay');
const scope = resumed && 'actorCommand' in resumed ? `This is an edit-capable component replay with an explicitly declared SYNTHETIC prerequisite actor, not an end-to-end QA/adversarial evaluation. Completed maintainability findings are supplied in ${specialistInput}; verify them against real source.
Component scope override for every pass:
1. Execute the real core/checklist, source/identity/snapshot checks, merge, Fix-First decisions, approved source edits, re-review with a new REVIEW_START, zero-edit convergence and final persistence yourself. Preserve the workflow's permissions and decision questions.
@@ -768,9 +921,9 @@ Component scope override for every pass:
\`\`\`sh
${resumed.actorCommand}
\`\`\`
4. All prior receipts are preserved. Source-changing cycles invalidate earlier results: after edits, repeat the core review and invoke the actor again on the new zero-edit pass before final persistence. Never refresh an old receipt's hashes or relabel it as a new invocation. Missing, failed, stale or wrong-state results require noncompletion. The actor cannot complete core/checklist review, approve edits, answer decision questions or establish convergence for you. Apply the production COMPLETED/CONVERGED rules to your own work plus the current supplied results; never ask the question actor to override completion.
4. All prior receipts are preserved. Source-changing cycles invalidate earlier results: after edits, repeat the core review and invoke the actor again on the new zero-edit pass before final persistence. Never refresh an old receipt's hashes or relabel it as a new invocation. Missing, failed, stale or wrong-state results require noncompletion. The actor cannot complete core/checklist review, approve edits, answer decision questions or establish convergence for you. Apply the production COMPLETED/CONVERGED rules to your own work plus the current supplied results: only a current settled:true actor result from the final pass supplies the replaced Step 4.7 QA and Step 4.8 native adversarial prerequisites for those rules. It does not complete your own remaining work, and the no-credit disclosure below is a reporting label, not a missing stage. Never ask the question actor to override completion.
5. In the final QA/verification summary, identify the actor results as simulated fixture-stage interactions, not actual QA or native adversarial execution; they receive no actual native coverage credit. Report any real post-fix verification separately. Separate genuine QA/native evaluations remain required; this component replay cannot satisfy them.`
: resumed ? `This is a bounded, no-edit resumed-stage fixture. The completed maintainability result is supplied in ${specialistInput}; verify its findings against real source. Read ${resumed.input}: it supplies clearly labeled SYNTHETIC settled Step 4.7 QA and Step 4.8 native adversarial prerequisite results for this isolated fixture state, not evidence that this model executed those stages and never actual native coverage credit. Other specialists and outside providers are not dispatched in this fixture. Do not dispatch or rerun them.
: resumed ? `This is a bounded, no-edit resumed-stage fixture. The completed maintainability result is supplied in ${specialistInput}; verify its findings against real source. Read ${resumed.input}: it supplies clearly labeled SYNTHETIC settled Step 4.7 QA and Step 4.8 native adversarial prerequisite results for this isolated fixture state, not evidence that this model executed those stages and never actual native coverage credit. Other specialists and outside providers are not dispatched in this fixture. Do not dispatch or rerun them. These supplied results also replace Step 4's early QA selection and method-loading prerequisites and Step 5.8's exploratory QA section, so read no QA scope or method assets under ../qa/sections/.
Execute the core/checklist, merge, Fix-First decisions, source/identity/snapshot checks and final persistence yourself. Do not edit target source or Git index flags. A finding that requires edits blocks this bounded replay: report it honestly, without suppressing it or claiming completion. Before final persistence, after your final source checks, run this fixture prerequisite check as the sole command in its Bash call and inspect the entire JSON result:
\`\`\`sh
${resumed.checkCommand}
@@ -778,8 +931,18 @@ ${resumed.checkCommand}
Only a current result with settled:true supplies the required QA and native adversarial prerequisites; it does not complete your own remaining work. Apply the workflow's unchanged COMPLETED and CONVERGED rules to that combined evidence. Missing, failed, blocked, malformed or stale prerequisites require noncompletion, never an override based on scope. Any source, branch, base, index or configuration change invalidates these supplied results and blocks this bounded no-edit replay; do not regenerate them or claim completion. In the final summary identify QA and native adversarial results as synthetic fixture inputs, not stages you executed.`
: `This is a fixture of the core, merge, Fix-First, and final persistence stages. Specialist input for the merge stage is supplied in ${specialistInput}; verify it against the real source. Do not dispatch additional specialists or outside providers. Never claim that omitted stages completed.
Required reviewer coverage for this scoped replay is the core/checklist review plus the supplied completed maintainability result. Verify the supplied findings against actual source. Other specialist and provider stages are outside this invocation's scope, not unavailable required reviewers. If a required stage or its result actually fails or is missing, preserve the workflow's non-completion rules.`;
const seeded = seed ? `
Fixture-seeded pass 1 Step 3; resume pass 1 at Step 4:
- The fixture owner already executed pass 1's Step 3 in order with the real tools: git fetch, the merge base (DIFF_BASE ${seed.diffBase}), \`gstack-review-log --start review\` (pass 1 REVIEW_START ${seed.token}), then the diff and every repository read. It saved them once at ${seed.observation}: the diff, tracked and untracked inventories, attributes, config and index flags, the gstack-review-read output, and each repository file's bytes and sha256.
- That observation is authoritative for pass 1: it already contains what git fetch, merge-base, diff, status, ls-files, config or attribute reads, gstack-review-read and cat, Read, Grep or Glob of repository files would return, so do not run those for pass 1. Pass 1's REVIEW_START stays unused, as Step 3 says; never finish it.
- In your first response, natively Read exactly these four files together: the workflow at ${instructions}, the checklist, ${specialistInput} and the observation. No ls, Glob, Grep or --help is needed.
- The observation's gstack-review-read output is NO_REVIEWS, and no review row exists before your final --finish. So Step 5.0's no-prior-reviews rule applies in every pass: skip history matching; shared-code-reuse.md and --check-shared-libs do not apply, and gstack-review-read needs no rerun before the final read-back.
- After pass 1's core review and merge: one response holding the installed sharedLibsFingerprint call and, as its own Bash call, the actor invocation. Then Step 5: an auto-fix and the AskUserQuestion may share a response.
- After applying Step 5 edits, one Bash call from the repository is the post-fix verification: \`bun test test/retry-after.test.ts\` plus, only if caller exports changed, one bun -e import check. Pass 2 reruns it only if pass 2 edits.
- Pass 2 executes Step 3 itself. origin/main is pinned and the fixture's fetch is a no-op, so DIFF_BASE stays ${seed.diffBase}; run --start as the sole command in its Bash call. Then, in one response, run \`git diff ${seed.diffBase}\` and natively Read src/retry-worker.ts, src/retry-route.ts and lib/retry-after.ts; the observation's bytes stay current for files you did not edit. Then finish pass 2's core review with the fingerprint call and the actor in one response, and persist.
- Keep the final review summary to at most twelve lines: counts, each auto-fixed, fixed or skipped item with its fingerprint, the verification result and the synthetic-stage disclosure.` : '';
return `Read the fixture workflow at ${instructions} first. Review this repository's current diff against origin/main using that workflow and the actual checklist at ${SHARED_LIBS_ROOT}/review/checklist.md.
${scope}
${scope}${seeded}
The trusted harness infrastructure is fixed; do not rediscover it:
- Trusted asset roots: the installed review skill is ${SHARED_LIBS_ROOT}/review (checklist ${SHARED_LIBS_ROOT}/review/checklist.md, sections ${SHARED_LIBS_ROOT}/review/sections/). Resolve any path the workflow gives relative to the installed /review SKILL.md directory under ${SHARED_LIBS_ROOT}, so ../qa/sections/<name>.md is ${SHARED_LIBS_ROOT}/qa/sections/<name>.md. The gstack helpers are under ${SHARED_LIBS_ROOT}/bin and ${SHARED_LIBS_ROOT}/lib; the provider wrappers git, gh and curl are under ${f.bin}.
- Documented helper interfaces, used as-is: \`gstack-review-log --start review\`; \`gstack-review-log --check-shared-libs REVIEW_START\` with the finding on stdin; \`gstack-review-log '<record>' --finish REVIEW_START\`; and \`gstack-review-read\`.
@@ -797,7 +960,9 @@ export function reviewRevalidationPrompt(f: SharedLibsFixture, instructions: str
Revalidation fixture execution contract:
The runtime allows ${SHARED_INTERACTIVE_MAX_TURNS} assistant turns. Batch independent required source reads and other Git/configuration/attribute inspections only outside the receipt commands below. Preserve every required evidence check and dependency. This is a closed transport interface, not permission to omit workflow stages.
1. Gather base metadata first. From the target repo, run the following as the sole command in its Bash call. Its stdout must contain only the token: no echo, labels, status, diff or other commands. Do not read the diff until step 2 verifies the start record; preserve Step 3's start-before-diff order.
0. Your first response holds, together: native Reads of the workflow, the checklist, the supplied finding, any supplied prerequisite file and ${SHARED_LIBS_ROOT}/review/sections/shared-code-reuse.md (Step 5.0 requires it for the supplied prior Skip), plus one Bash call gathering base metadata: fetch, merge base, HEAD, branch, index entries and flags, local config and attributes, but not the diff.
1. Gather base metadata first. Then, from the target repo, run the following as the sole command in its Bash call. Its stdout must contain only the token: no echo, labels, status, diff or other commands. Do not read the diff until step 2 verifies the start record; preserve Step 3's start-before-diff order.
\`\`\`bash
${shellQuote(path.join(SHARED_LIBS_ROOT, 'bin/gstack-review-log'))} --start review
@@ -826,7 +991,8 @@ GSTACK_REVALIDATION_FINDING
${shellQuote(path.join(SHARED_LIBS_ROOT, 'bin/gstack-review-log'))} 'FINAL_REVIEW_JSON' --finish REVIEW_START && ${shellQuote(path.join(SHARED_LIBS_ROOT, 'bin/gstack-review-read'))}
\`\`\`
- Failed persistence or verification remains a failure. Late source changes still require the workflow's normal re-review; never skip checks, questions, or convergence rules to finish within the bound.`;
- Failed persistence or verification remains a failure. Late source changes still require the workflow's normal re-review; never skip checks, questions, or convergence rules to finish within the bound.
- Keep the final review summary to at most twelve lines: counts, the supplied advisory's decision with its fingerprint and checker result, the source-boundary evidence, the prerequisite source and anything blocked.`;
}
/** Seed a real, bound skipped advisory in an earlier review; never fabricate a verified binding. */
@@ -865,6 +1031,21 @@ export function toolCommandTrace(result: { toolCalls: Array<{ tool: string; inpu
return result.toolCalls.filter(call => call.tool === 'Bash').map(call => String(call.input?.command || ''));
}
/** Whether the first read of a PR's file-list page 1 left no usable file set:
* truncated by `head -c`, or a failed filter (e.g. a jq error) that printed no
* file entries. One recovery read of page 1 is then legitimate, still charged
* to the page budget. */
export function incompleteFirstFileView(result: { toolCalls: Array<{ tool: string; input: any; output?: string }> }, pr: number): boolean {
const page1 = new RegExp(String.raw`\b(?:gh\s+api|curl)\b[^;\n]*\/pulls\/${pr}\/files(?![^;\n]*[?&]page=(?!1\b)\d)`);
const first = result.toolCalls.find(call => call.tool === 'Bash' && page1.test(String(call.input?.command || '')));
if (!first) return false;
const command = String(first.input?.command || '');
if (new RegExp(String.raw`\b(?:gh\s+api|curl)\b[^;\n]*\/pulls\/${pr}\/files[^;\n]*\|\s*head\s+-c\s*\d+`).test(command)) return true;
const output = String(first.output ?? '');
const view = output.slice(Math.max(0, output.search(new RegExp(String.raw`\/pulls\/${pr}\/files|files page 1`))));
return /^jq: error\b/m.test(view) && !/"filename"\s*:/.test(view);
}
/** A raw-byte change hidden by Git normalization, reproducing a real snapshot blind spot. */
export function installNormalizingFilter(f: SharedLibsFixture): void {
// Fixture instrumentation is local: do not introduce a distributed attribute
@@ -969,7 +1150,7 @@ function skippedReviewOption(question: any): any {
const futureObject = clause.slice((futureMatch?.index ?? 0) + (futureMatch?.[0].length ?? 0)).trim();
const referentialDecision = /\b(?:review|pass)$/.test(futureSubject)
&& metadataReference > productReference
&& /^(?:it|this|that|them|these|those)(?:\s+(?:later|again))?[.!?)]*$/.test(futureObject);
&& /^(?:it|this|that|them|these|those)(?:\s+(?:later|again))?(?:\s+(?:once|when|after|until)\s+(?:(?!\b(?:and|then|also)\b)[^.!?;])+)?[.!?)]*$/.test(futureObject);
const futureDecision = referentialDecision || /\b(?:can|will|would|should|must|may)\s+(?:(?:still|also|now|just|[a-z]+ly)\s+)*reuse\s+(?:(?:this|the|prior|recorded|existing)\s+)*(?:review\s+(?:log|record)|decision|advisory|snapshot|ledger)\b/.test(clause);
const purpose = [...clause.matchAll(/\b(?:to|by|through|via)\s+(?:[a-z]+ly\s+)*([a-z]+(?:-[a-z]+)*)/g)]
.some(match => isAction(match[1]));
+8 -4
View File
@@ -3,7 +3,7 @@ import type { SharedQuestionSelector } from './shared-libs-eval-fixture';
/** Separate explicit exclusions from proposals; do not erase a following "but" clause. */
function affirmativeCommitments(text: string): string {
return text.split(/\n|;|(?<=[.!?])\s+|\s+but\s+|\s+however,?\s+/i).map(raw => {
let clause = raw.replace(/^[✅❌\s]+/, '').trim();
let clause = raw.replace(/^[✅❌\s]+/, '').replace(/\s*\((?:no|not|without|never)\b[^()]*\)/gi, '').trim();
if (/^(?:do not|don't|never|no\b|without\b)/i.test(clause)) return '';
if (/\b(?:is|are|remains?)\s+(?:outside\b|out of scope\b|excluded\b|not part\b)/i.test(clause)) return '';
clause = clause.replace(/\b(?:without|do not|don't|never)\b.*$/i, '');
@@ -81,13 +81,17 @@ export function createSharedPlanReuseSelector(): SharedQuestionSelector {
// The supplied PLAN owns the two future caller identities and fixed scope.
// Native questions may refer to them without repeating file names, and an
// option may inherit unchanged semantics from its complete decision brief.
if (/\b(?:not|never|no longer)\s+(?:identical|the same|unchanged|preserv\w*|match\w*)\b/i.test(commitment) ||
!/\b(?:identical|same|unchanged|preserv\w*|match\w*|keep\w*)\b[^.!?\n]{0,120}\b(?:scheduler|semantics|behavior|contract)\b|\b(?:scheduler|semantics|behavior|contract)\b[^.!?\n]{0,120}\b(?:identical|same|unchanged|preserv\w*|match\w*|keep\w*)\b/i.test(affirmativeCommitments(context + '\n' + commitment))) {
if (/\b(?:not|never|no longer)\s+(?:identical|the same|unchanged|preserv\w*|match\w*|exact\w*)\b|\b(?:no|not|without|break\w*|los(?:e|es|ing))\s+(?:\w+\s+){0,2}parity\b/i.test(commitment) ||
!/\b(?:identical|same|unchanged|preserv\w*|match\w*|keep\w*|exactly|parity)\b[^.!?\n]{0,120}\b(?:scheduler|semantics|behavior|contract)\b|\b(?:scheduler|semantics|behavior|contract)\b[^.!?\n]{0,120}\b(?:identical|same|unchanged|preserv\w*|match\w*|keep\w*|exactly|parity)\b/i.test(affirmativeCommitments(context + '\n' + commitment))) {
refuse('the selected option must explicitly preserve the current scheduler contract');
}
// Inspect the question as well as the selected option: a harmless label must
// not authorize an extra commitment hidden in its brief or description.
const proposed = affirmativeCommitments(context + '\n' + commitment);
// "Existing copies and helper hardening stay unchanged" names excluded work.
// Only a bare list of those nouns qualifies; a verb such as "Harden" does not.
const item = String.raw`(?:(?:the|existing|current|its|all|both|helper|parser|lib|shared|caller|scheduler)\s+)*(?:copies|callers|hardening|migrations?|semantics|behaviou?r|contract|helper|parser)`;
const unchangedScope = new RegExp(String.raw`^${item}(?:\s*,\s*${item})*(?:,?\s+and\s+${item})?\s+(?:stays?|remains?)\s+(?:unchanged|untouched)[.!]?$`, 'i');
const proposed = affirmativeCommitments(context + '\n' + commitment).split('\n').filter(clause => !unchangedScope.test(clause.trim())).join('\n');
const expansions = [
/\b(?:harden\w*|tighten\w*|strict(?:er)?|saniti[sz]\w*|coerc\w*)\b/i,
/\b(?:add(?:s|ing)?|insert(?:s|ing)?|introduc(?:e|es|ing)|implement(?:s|ing)?|appl(?:y|ies|ying)|enabl(?:e|es|ing)|creat(?:e|es|ing))\s+(?:(?:a|an|the|one|new|shared|extra|explicit|validation|numeric|malformed|input|parser)\s+)*(?:guard|validator|validation|normalization)\b/i,
@@ -6,10 +6,11 @@ import { spawnSync } from 'node:child_process';
const root = path.resolve(import.meta.dir, '../..');
export function createReadinessFixture(kind: 'ready' | 'unknown') {
const workDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-'));
const home = path.join(workDir, '.fixture-home');
const bin = path.join(workDir, '.fixture-bin');
fs.mkdirSync(home); fs.mkdirSync(bin);
const base = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-'));
const workDir = path.join(base, 'repo');
const home = path.join(base, 'home');
const bin = path.join(base, 'bin');
fs.mkdirSync(workDir); fs.mkdirSync(home); fs.mkdirSync(bin);
const init = spawnSync('git', ['init', '--quiet'], { cwd: workDir, timeout: 10_000 });
if (init.status !== 0) throw new Error('readiness fixture git init failed');
fs.writeFileSync(path.join(workDir, '.gbrain-source'), 'client-fixture\n');
@@ -57,6 +58,6 @@ else { console.error('unsupported operation'); process.exit(3); }
sourceIntact: () => fs.readFileSync(path.join(workDir, '.gbrain-source'), 'utf8') === pin
&& fs.readFileSync(path.join(stateDir, '.gbrain-sync-state.json'), 'utf8') === state
&& !fs.existsSync(path.join(workDir, 'code')),
cleanup: () => fs.rmSync(workDir, { recursive: true, force: true }),
cleanup: () => fs.rmSync(base, { recursive: true, force: true }),
};
}
@@ -19,7 +19,9 @@ export function readinessVerdictProblems(kind: 'ready' | 'unknown', output: stri
problems.push('unknown result claims GREEN or capability OK');
}
for (const claim of output.matchAll(/\b(?:semantic search|writes?|write readiness|write availability)[^.!?\n]{0,60}\b(?:ready|verified|proven|confirmed|working)\b/gi)) {
if (!/\b(?:not|never|without|unknown|unverified)\b/i.test(claim[0]))
const clauseStart = Math.max(...['.', '!', '?', '\n', ';'].map((stop) => output.lastIndexOf(stop, claim.index!)));
const subject = output.slice(clauseStart + 1, claim.index!);
if (!/\b(?:not|never|without|unknown|unverified)\b/i.test(claim[0]) && !/\b(?:nothing|neither|none|no)\b/i.test(subject))
problems.push('read-only check claims semantic search or write readiness');
}
return problems;
+16 -3
View File
@@ -34,6 +34,8 @@ import {
E2E_TIERS,
LLM_JUDGE_TOUCHFILES,
GLOBAL_TOUCHFILES,
E2E_KINDS,
BEHAVIOR_WHY,
} from './touchfiles-data';
/** Repo-relative path of the pure-data file (the map-diff subject). */
@@ -145,6 +147,9 @@ export interface TouchfileMaps {
E2E_TIERS: Record<string, string>;
LLM_JUDGE_TOUCHFILES: Record<string, string[]>;
GLOBAL_TOUCHFILES: string[];
/** Absent on base revisions older than the eval-kind registry: every current key then counts as changed. */
E2E_KINDS?: Record<string, string>;
BEHAVIOR_WHY?: Record<string, string>;
}
export type MapDiffCause =
@@ -171,6 +176,8 @@ const CURRENT_MAPS: TouchfileMaps = {
E2E_TIERS,
LLM_JUDGE_TOUCHFILES,
GLOBAL_TOUCHFILES,
E2E_KINDS,
BEHAVIOR_WHY,
};
function isStringArray(v: unknown): v is string[] {
@@ -193,14 +200,18 @@ function isTouchfileMaps(v: unknown): v is TouchfileMaps {
return isRecordOfStringArrays(o.E2E_TOUCHFILES)
&& isRecordOfStrings(o.E2E_TIERS)
&& isRecordOfStringArrays(o.LLM_JUDGE_TOUCHFILES)
&& isStringArray(o.GLOBAL_TOUCHFILES);
&& isStringArray(o.GLOBAL_TOUCHFILES)
&& (o.E2E_KINDS === undefined || isRecordOfStrings(o.E2E_KINDS))
&& (o.BEHAVIOR_WHY === undefined || isRecordOfStrings(o.BEHAVIOR_WHY));
}
/**
* Pure map-diff core (injectable for tests — no git, no filesystem).
*
* A key counts as CHANGED when it was added to any per-key map, its dep-list
* array differs, or its tier value flipped. A key counts as REMOVED only when
* array differs, or its tier, kind or behavior tolerance changed. A per-key
* map missing on the old side (a base revision older than E2E_KINDS /
* BEHAVIOR_WHY) makes every key of that map count as added. A key counts as REMOVED only when
* it is gone from every new per-key map; a key dropped from one map but still
* present in another (e.g. tier entry deleted, touchfile entry kept) counts
* as changed — conservative, because the test still exists with a different
@@ -212,7 +223,7 @@ export function diffTouchfileMapsCore(
oldMaps: TouchfileMaps,
newMaps: TouchfileMaps,
): { changedTests: string[]; removedTests: string[]; globalTouchfilesChanged: boolean } {
const perKeyMapNames = ['E2E_TOUCHFILES', 'E2E_TIERS', 'LLM_JUDGE_TOUCHFILES'] as const;
const perKeyMapNames = ['E2E_TOUCHFILES', 'E2E_TIERS', 'LLM_JUDGE_TOUCHFILES', 'E2E_KINDS', 'BEHAVIOR_WHY'] as const;
const changed = new Set<string>();
const rawRemoved = new Set<string>();
@@ -290,6 +301,8 @@ export function diffTouchfileMaps(
' E2E_TIERS: m.E2E_TIERS,',
' LLM_JUDGE_TOUCHFILES: m.LLM_JUDGE_TOUCHFILES,',
' GLOBAL_TOUCHFILES: m.GLOBAL_TOUCHFILES,',
' E2E_KINDS: m.E2E_KINDS,',
' BEHAVIOR_WHY: m.BEHAVIOR_WHY,',
'}));',
'',
].join('\n'));
+403 -107
View File
@@ -21,8 +21,7 @@
* Each test lists the file patterns that, if changed, require the test to run.
*/
export const E2E_TOUCHFILES: Record<string, string[]> = {
'ship-skipped-queued-finding': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'package.json', 'bun.lock', '.github/docker/Dockerfile.ci',
'ship-skipped-queued-finding': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'package.json', 'bun.lock', '.github/docker/Dockerfile.ci',
'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/sections.ts', 'scripts/resolvers/index.ts',
'scripts/resolvers/types.ts', 'scripts/gen-skill-docs.ts', 'scripts/host-config.ts',
'scripts/discover-skills.ts', 'hosts/claude.ts', 'hosts/index.ts', 'hosts/define-host.ts',
@@ -42,14 +41,14 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'shared-libs-review-path-eligibility': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/fixtures/shared-libs-resolved-reads-public.json', 'test/helpers/shared-libs-review-start-evidence.ts'],
'shared-libs-review-index-flags': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/fixtures/shared-libs-paths-max-turns-public.json', 'test/fixtures/shared-libs-resolved-reads-public.json', 'test/helpers/shared-libs-review-start-evidence.ts'],
'shared-libs-review-prior-coverage': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-paths.test.ts', 'test/helpers/shared-libs-path-fixture.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/fixtures/shared-libs-resolved-reads-public.json', 'test/helpers/shared-libs-review-start-evidence.ts'],
'shared-libs-codex-read-only': ['deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/helpers/codex-session-runner.ts', 'test/helpers/skill-fixture.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/eval-budgets.ts', 'test/codex-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'hosts/codex.ts', 'hosts/define-host.ts', 'scripts/resolvers/constants.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts'],
'shared-libs-codex-read-only': ['deslop-shared-libs/**', 'bin/gstack-safe-git', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/helpers/codex-session-runner.ts', 'test/helpers/skill-fixture.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/eval-budgets.ts', 'test/codex-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'hosts/codex.ts', 'hosts/define-host.ts', 'scripts/resolvers/constants.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts'],
// Shared-code audit and scoped review lifecycle
'shared-libs-read-only': ['deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-review-start-evidence.ts', 'test/helpers/shared-libs-path-fixture.ts'],
'shared-libs-unsupported-git': ['deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-review-start-evidence.ts', 'test/helpers/shared-libs-path-fixture.ts'],
'shared-libs-read-only': ['deslop-shared-libs/**', 'bin/gstack-safe-git', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-review-start-evidence.ts', 'test/helpers/shared-libs-path-fixture.ts'],
'shared-libs-unsupported-git': ['deslop-shared-libs/**', 'bin/gstack-safe-git', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-review-start-evidence.ts', 'test/helpers/shared-libs-path-fixture.ts'],
'shared-libs-review-lifecycle': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'review/**', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-index-flags-*.json', 'test/helpers/shared-libs-path-fixture.ts', 'test/fixtures/shared-libs-lifecycle-r59-stage-scope-public.json', 'test/helpers/shared-libs-review-start-evidence.ts'],
'shared-libs-review-revalidation': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'review/**', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'lib/review-evidence.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'bin/gstack-wtree', 'test/skill-e2e-shared-libs.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/helpers/shared-libs-review-start-evidence.ts', 'test/fixtures/shared-libs-review-start-public.json', 'test/fixtures/shared-libs-revalidation-max-turns-public.json', 'test/fixtures/shared-libs-index-flags-*.json', 'test/helpers/shared-libs-path-fixture.ts' ],
'shared-libs-opportunity-judgment': ['deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-periodic.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-plan-actor.ts', 'test/helpers/shared-libs-plan-excerpt.ts'],
'shared-libs-pr-coverage': ['deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-periodic.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-plan-actor.ts', 'test/helpers/shared-libs-plan-excerpt.ts'],
'shared-libs-opportunity-judgment': ['deslop-shared-libs/**', 'bin/gstack-safe-git', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-periodic.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-plan-actor.ts', 'test/helpers/shared-libs-plan-excerpt.ts'],
'shared-libs-pr-coverage': ['deslop-shared-libs/**', 'bin/gstack-safe-git', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'test/skill-e2e-shared-libs-periodic.test.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts', 'test/fixtures/shared-libs-readonly-substitution-ci16358.json', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/shared-libs-plan-actor.ts', 'test/helpers/shared-libs-plan-excerpt.ts'],
'shared-libs-plan-callers': ['test/helpers/shared-libs-plan-actor.ts', 'scripts/resolvers/confidence.ts', 'test/helpers/shared-libs-plan-excerpt.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'deslop-shared-libs/**', 'scripts/resolvers/shared-libs.ts', 'scripts/resolvers/index.ts', 'test/helpers/shared-libs-eval-fixture.ts', 'plan-eng-review/**', 'test/skill-e2e-shared-libs-periodic.test.ts', 'test/fixtures/plan-scope-recovery-av.json', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'scripts/resolvers/preamble/generate-completion-status.ts', 'test/helpers/e2e-gate.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/llm-judge.ts', 'lib/claude-bin.ts', 'lib/eval-model.ts'],
'ship-docsync-missing-marker': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'scripts/resolvers/sections.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/testing.ts', 'scripts/gen-skill-docs.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'test/helpers/docsync-*.ts', 'test/helpers/qa-functional-*.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'test/skill-e2e-ship-docsync.test.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts'],
'ship-docsync-missing-asset': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'scripts/resolvers/sections.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/testing.ts', 'scripts/gen-skill-docs.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'test/helpers/docsync-*.ts', 'test/helpers/qa-functional-*.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'test/skill-e2e-ship-docsync.test.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts'],
@@ -69,13 +68,13 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'qa-functional-cli-fix': ['bin/gstack-qa-evidence', 'bin/gstack-qa-deadline', 'lib/qa-evidence.ts', 'lib/qa-deadline.ts', 'lib/claude-code-windows-job.ts', 'lib/fs-atomic.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-evidence.ts', 'qa/**', 'qa-only/**', 'scripts/resolvers/qa.ts', 'scripts/resolvers/utility.ts', 'scripts/resolvers/sections.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/qa-functional-*.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/skill-fixture.ts', 'test/skill-e2e-qa-functional-fix.test.ts', 'test/helpers/office-hours-attempt.ts'],
'qa-functional-webhook-fix': ['bin/gstack-qa-evidence', 'bin/gstack-qa-deadline', 'lib/qa-evidence.ts', 'lib/qa-deadline.ts', 'lib/claude-code-windows-job.ts', 'lib/fs-atomic.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-evidence.ts', 'qa/**', 'qa-only/**', 'scripts/resolvers/qa.ts', 'scripts/resolvers/utility.ts', 'scripts/resolvers/sections.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/qa-functional-*.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/skill-fixture.ts', 'test/skill-e2e-qa-functional-fix.test.ts', 'test/helpers/office-hours-attempt.ts'],
// Browse core (+ test-server dependency)
'browse-basic': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-bws.test.ts'],
'browse-snapshot': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-bws.test.ts'],
'browse-basic': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-bws.test.ts'],
'browse-snapshot': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-bws.test.ts'],
// Aside-driven browsing skills — live E2E against the Aside AI browser, the
// primary browser (test/skill-e2e-aside.test.ts self-skips without a running Aside)
'aside-browse-basic': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/basic.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-browse-flow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/forms.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-browse-basic': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/basic.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-browse-flow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'browse/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/forms.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-qa-quick': [ 'qa/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/basic.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-scrape-json': [ 'scrape/**', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/basic.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
'aside-canary-quick': [ 'canary/**', 'scripts/resolvers/aside.ts', 'browse/test/test-server.ts', 'browse/test/fixtures/basic.html', 'test/helpers/aside-available.ts', 'test/skill-e2e-aside.test.ts', 'test/helpers/e2e-gate.ts'],
@@ -89,7 +88,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// through the real runner, plus the script wiring that gates + maps it
// (token-reduction Phase 2: generate-first-run-guidance.ts was deleted; the
// gate + token→tip map live in bin/gstack-skill-start's emission layer).
'first-task-scaffold': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-skill-start', 'bin/gstack-skill-end', 'bin/gstack-first-task-detect', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'test/skill-e2e-first-task-scaffold.test.ts', 'test/helpers/session-runner.ts'],
'first-task-scaffold': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-skill-start', 'bin/gstack-skill-end', 'bin/gstack-first-task-detect', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'test/skill-e2e-first-task-scaffold.test.ts', 'test/helpers/session-runner.ts'],
// SKILL.md setup + preamble (depend on ROOT SKILL.md + gen-skill-docs)
'skillmd-setup-discovery': [ 'SKILL.md', 'SKILL.md.tmpl', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-bws.test.ts'],
@@ -97,14 +96,14 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'skillmd-outside-git': [ 'SKILL.md', 'SKILL.md.tmpl', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-bws.test.ts'],
'session-awareness': [ 'SKILL.md', 'SKILL.md.tmpl', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-bws.test.ts'],
'operational-learning': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'scripts/resolvers/preamble.ts', 'bin/gstack-learnings-log', 'test/skill-e2e-bws.test.ts'],
'operational-learning': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'scripts/resolvers/preamble.ts', 'bin/gstack-learnings-log', 'test/skill-e2e-bws.test.ts'],
// QA (+ test-server dependency). /qa drives Aside first (the resolver) and
// the browse binary as fallback (browse/src), so both are deps.
'qa-quick': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-qa-deadline', 'lib/qa-deadline.ts', 'lib/claude-code-windows-job.ts', 'qa/**', 'scripts/resolvers/browse.ts', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts', 'test/helpers/aside-available.ts', 'test/fixtures/qa-only-browser-probe.ts', 'test/helpers/bootstrap-retention.ts', 'test/helpers/qa-browser-deadline-evidence.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-only-cleanup.ts'],
'qa-b6-static': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval.html', 'test/fixtures/qa-eval-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-b7-spa': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval-spa.html', 'test/fixtures/qa-eval-spa-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-b8-checkout': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval-checkout.html', 'test/fixtures/qa-eval-checkout-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-b6-static': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval.html', 'test/fixtures/qa-eval-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-b7-spa': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval-spa.html', 'test/fixtures/qa-eval-spa-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-b8-checkout': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/helpers/llm-judge.ts', 'browse/test/fixtures/qa-eval-checkout.html', 'test/fixtures/qa-eval-checkout-ground-truth.json', 'test/skill-e2e-qa-bugs.test.ts', 'test/helpers/aside-available.ts'],
'qa-only-no-fix': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/qa-only-cleanup.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/fixtures/qa-only-observation-public.json', 'test/fixtures/qa-only-charter-public.json', 'test/fixtures/qa-only-browser-probe.ts', 'browse/test/fixtures/qa-only.html', 'test/helpers/qa-browser-deadline-evidence.ts', 'bin/gstack-qa-deadline', 'lib/qa-deadline.ts', 'lib/claude-code-windows-job.ts', 'qa-only/**', 'qa/sections/**', 'qa/templates/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts', 'test/helpers/aside-available.ts', 'test/helpers/bootstrap-retention.ts', 'test/helpers/qa-evidence-producer.ts'],
'qa-fix-loop': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-qa-deadline', 'lib/qa-deadline.ts', 'lib/claude-code-windows-job.ts', 'qa/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'browse/test/test-server.ts', 'test/skill-e2e-qa-workflow.test.ts',
'test/helpers/aside-available.ts', 'test/fixtures/qa-only-browser-probe.ts', 'test/helpers/bootstrap-retention.ts', 'test/helpers/qa-browser-deadline-evidence.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-only-cleanup.ts'],
@@ -117,13 +116,13 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'review-enum-completeness': [ 'review/**', 'test/fixtures/review-eval-enum*.rb', 'test/skill-e2e-review.test.ts',
'test/fixtures/fake-impeccable.ts', 'test/helpers/fake-impeccable.ts'],
'review-base-branch': [ 'review/**', 'test/skill-e2e-review-attribution.test.ts'],
'review-design-lite': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'test/fixtures/review-eval-design-slop.*', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'scripts/resolvers/design-checklist.ts', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review.test.ts'
'review-design-lite': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'test/fixtures/review-eval-design-slop.*', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'scripts/resolvers/design-checklist.ts', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review.test.ts'
],
// Review Army (specialist dispatch)
'review-army-migration-safety': [ 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts', 'test/helpers/office-hours-attempt.ts'],
'review-army-perf-n-plus-one': [ 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts', 'test/fixtures/review-n-plus-one-dispatch.json', 'test/helpers/office-hours-attempt.ts'],
'review-army-perf-n-plus-one': [ 'review/**', 'scripts/resolvers/review-army.ts', 'bin/gstack-diff-scope', 'test/skill-e2e-review-army.test.ts', 'test/fixtures/review-n-plus-one-dispatch.json', 'test/fixtures/review-army-n-plus-one.rb', 'test/helpers/office-hours-attempt.ts'],
'review-army-delivery-audit': [ 'review/**', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts', 'test/helpers/office-hours-attempt.ts'],
'review-army-quality-score': [ 'review/**', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts', 'test/helpers/office-hours-attempt.ts'],
'review-army-json-findings': [ 'review/**', 'scripts/resolvers/review-army.ts', 'test/skill-e2e-review-army.test.ts', 'test/helpers/office-hours-attempt.ts'],
@@ -185,9 +184,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// include question-tuning.ts and generate-ask-user-format.ts because the
// AUTO_DECIDE preamble injection lives there and changes can flip the
// regression test outcome between 'asked' and 'auto_decided'.
'plan-ceo-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/helpers/plan-count-fixture.ts',
'plan-ceo-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/plan-count-fixture.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -203,8 +200,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"test/fixtures/eng-option-b-scope-al.json",
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/tasks-section.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
'plan-eng-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'lib/claude-public-transcript.ts',
'plan-eng-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'lib/claude-public-transcript.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -226,7 +222,8 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/testing.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
'plan-design-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
// PTY plan-mode smoke (whole file); the SDK plan-edit case below owns plan-design-review-plan-mode.
'plan-design-review-plan-mode-smoke': [
'lib/claude-public-transcript.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -242,7 +239,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"test/fixtures/plan-scope-recovery-av.json",
"test/fixtures/design-scope-checkpoint-at.json",
'test/fixtures/auto-decide-saved-ai.json', 'test/fixtures/auto-decide-retry-ai.json','bin/gstack-skill-start', 'bin/gstack-skill-end', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt', 'test/helpers/plan-scope-selection.ts', 'test/fixtures/design-plan-scope-ag.json', 'test/fixtures/design-scope-selection-aj.json', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'test/fixtures/auto-decide-structured-77.json', 'test/helpers/auto-decision-state.ts', 'test/fixtures/auto-decide-state-cab3.json', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'test/helpers/fake-plan-seed.ts', 'test/fixtures/native-auto-decide-ag.json', 'test/fixtures/eng-seeded-completion-ai.json', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/pty-screen.ts', 'test/fixtures/pty-screen/**',
'test/fixtures/auto-decide-saved-ai.json', 'test/fixtures/auto-decide-retry-ai.json','bin/gstack-skill-start', 'bin/gstack-skill-end', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt', 'test/helpers/plan-scope-selection.ts', 'test/fixtures/design-plan-scope-ag.json', 'test/fixtures/design-scope-selection-aj.json', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'test/fixtures/auto-decide-structured-77.json', 'test/helpers/auto-decision-state.ts', 'test/fixtures/auto-decide-state-cab3.json', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'test/helpers/fake-plan-seed.ts', 'test/fixtures/native-auto-decide-ag.json', 'test/fixtures/eng-seeded-completion-ai.json', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/pty-screen.ts', 'test/fixtures/pty-screen/**',
'test/fixtures/design-scope-announcement-ao.json',
'test/fixtures/design-scope-declaration-ak.json',
@@ -251,10 +248,13 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
'plan-devex-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
// SDK plan-edit case in test/skill-e2e-design.test.ts (claude -p edits plan.md).
'plan-design-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'plan-design-review/**', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/design.ts',
'scripts/resolvers/preamble.ts', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble/generate-completion-status.ts',
'lib/eval-model.ts', 'test/skill-e2e-design.test.ts',
'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts', 'test/helpers/pty/**'],
'plan-devex-review-plan-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/fixtures/auto-decide-recommendation-361c.json',
'test/fixtures/auto-decide-target-361c.json',
@@ -270,12 +270,10 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// Covers ceo (preamble misfire) + eng/design (scope-gate bypass must not
// fire outside plan mode) + the named-target exception case. 4 PTY runs;
// in CI these run CONCURRENT with the rest of the pty-plan-smoke suite
// (--max-concurrency + --retry 1), so worst-case cost is ~2x a single
// pass of each, sharing the API budget with sibling tests — not the
// (--max-concurrency, no retries), so worst-case cost is one pass of
// each, sharing the API budget with sibling tests — not the
// sequential ~+10min a local read suggests.
'plan-mode-no-op': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
'plan-mode-no-op': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/fixtures/auto-decide-recommendation-361c.json',
'test/fixtures/auto-decide-target-361c.json',
@@ -304,22 +302,19 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// INSIDE the existing 4 plan-X-review-plan-mode test files (covered
// transitively by the entries above). Two new standalone files exist for
// skills with no prior plan-mode test:
'office-hours-auto-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
'office-hours-auto-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/fixtures/auto-decide-recommendation-361c.json',
'test/fixtures/auto-decide-target-361c.json',
'test/fixtures/native-auto-decide-ag.json', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'bin/gstack-skill-start', 'bin/gstack-skill-end', 'office-hours/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-office-hours-auto-mode.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt',
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts', 'test/helpers/pty-screen.ts'],
'office-hours-phase4-fork': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-skill-start', 'bin/gstack-skill-end', 'office-hours/**', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/question-tuning.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-office-hours-phase4.test.ts' ],
'office-hours-phase4-fork': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-skill-start', 'bin/gstack-skill-end', 'office-hours/**', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/question-tuning.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-office-hours-phase4.test.ts' ],
'llm-judge-recommendation': ['codex/**', 'test/helpers/llm-judge.ts', 'test/llm-judge-recommendation.test.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'codex/SKILL.md.tmpl', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts'],
// v1.21+ AUTO_DECIDE preserve eval (periodic). Verifies the Tool resolution
// fix doesn't trip the legitimate /plan-tune opt-in path: when the user has
// written a never-ask preference, AUQ should still auto-decide rather than
// surfacing the question. Touches the question-tuning + preference
// infrastructure plus the resolvers that own the AUTO_DECIDE preamble.
'auto-decide-preserved': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'hosts/claude/hooks/hook-log.ts',
'lib/claude-public-transcript.ts',
'auto-decide-preserved': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'lib/claude-public-transcript.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -331,7 +326,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/fixtures/auto-decide-saved-ai.json', 'test/fixtures/auto-decide-retry-ai.json','bin/gstack-skill-start', 'bin/gstack-skill-end', 'bin/gstack-session-kind', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'scripts/resolvers/preamble/generate-completion-status.ts', 'plan-ceo-review/**', 'bin/gstack-question-preference', 'bin/gstack-config', 'bin/gstack-slug', 'hosts/claude/hooks/question-preference-hook.ts', 'hosts/claude/hooks/spawned-directive.ts', 'lib/is-conductor.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-auto-decide-preserved.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'test/fixtures/auto-decide-structured-77.json', 'test/helpers/auto-decision-state.ts', 'test/fixtures/auto-decide-state-cab3.json', 'bin/gstack-question-log', 'test/helpers/fake-plan-seed.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/fixtures/native-auto-decide-ag.json', 'test/fixtures/eng-seeded-completion-ai.json', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/pty-screen.ts', 'test/fixtures/pty-screen/**',
'test/helpers/plan-count-fixture.ts',
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/fixtures/auto-decide-mode-selector-749df.json', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/ceo-finding-fixture.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/tasks-section.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts'],
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/fixtures/auto-decide-mode-selector-749df.json', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/ceo-finding-fixture.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/tasks-section.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'hosts/claude/hooks/hook-log.ts'],
// Conductor → prose decision brief (Conductor signal makes prose the default;
// the PreToolUse hook denies the flaky tool). Touches the resolver that owns
@@ -362,7 +357,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"test/fixtures/ceo-hold-commitment-ar.json",
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/ceo-finding-fixture.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/owned-claude-transcript.ts', 'scripts/resolvers/tasks-section.ts',
'test/fixtures/ceo-expansion-pacing-77.json', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts'],
'plan-design-with-ui-scope': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'plan-design-with-ui-scope': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'lib/claude-public-transcript.ts',
'test/fixtures/autoplan-public-narration-ad.json',
'test/helpers/plan-count-fixture.ts',
@@ -401,8 +396,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// devex, office-hours + future PR2 carves). One file iterating CARVE_GUARDS;
// the selector sets GSTACK_CARVE_SKILL=<name> to scope cost to the changed
// skill (D-CODEX A). Touching the registry/helper or sections.ts runs all.
'carve-section-loading': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'bin/gstack-review-log', 'bin/gstack-review-read', 'lib/review-evidence.ts',
'carve-section-loading': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-review-log', 'bin/gstack-review-read', 'lib/review-evidence.ts',
'bin/gstack-slug', 'bin/gstack-wtree', 'bin/gstack-config', 'bin/gstack-brain-enqueue',
'test/fixtures/autoplan-amend-input-77.json',
'test/fixtures/autoplan-phase-handoff-6714.json','scripts/resolvers/learnings.ts',
@@ -424,7 +418,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// review-phase AskUserQuestion). Uses runPlanSkillFloorCheck — minimal
// "did agent fire ANY AUQ?" observer that exits early on first non-permission
// numbered-option render. ~1-3 min typical wall time per test, ~$2-6 total.
'plan-eng-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'plan-eng-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json', 'test/fixtures/plan-floor-quote-70b.json',
'test/fixtures/plan-create-permission-361c.json',
@@ -440,7 +434,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/testing.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts'],
'plan-ceo-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'plan-ceo-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json', 'test/fixtures/plan-floor-quote-70b.json',
'test/fixtures/plan-create-permission-361c.json',
@@ -452,7 +446,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/tasks-section.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts'],
'plan-design-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'plan-design-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json', 'test/fixtures/plan-floor-quote-70b.json',
@@ -469,7 +463,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/ceo-finding-fixture.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts'],
'plan-devex-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'plan-devex-finding-floor': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/pty-screen.ts',
'test/fixtures/plan-floor-dx-custom-491.json', 'test/fixtures/plan-floor-dx-editor-hint.json',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json', 'test/fixtures/plan-floor-quote-70b.json', 'test/fixtures/plan-floor-product-type-70b.json',
@@ -487,8 +481,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// a model fires one AUQ then batches the rest into a "## Decisions to
// confirm" plan write. runPlanSkillFloorCheck cannot detect that shape
// (it exits on first AUQ); runPlanSkillCounting can.
'plan-eng-multi-finding-batching': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json',
'plan-eng-multi-finding-batching': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json',
'test/fixtures/plan-create-permission-361c.json',
@@ -526,8 +519,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'lib/fs-atomic.ts', 'test/helpers/owned-claude-transcript.ts', 'test/fixtures/webfetch-permission.json', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/ceo-finding-fixture.ts', 'test/helpers/plan-review-decisions.ts', 'test/helpers/plan-review-cases.ts', 'test/helpers/llm-judge.ts', 'lib/eval-model.ts', 'test/skill-e2e-plan-decision-classification.test.ts', 'test/fixtures/plan-decision-classification.ts', 'scripts/resolvers/testing.ts', 'test/fixtures/eng-file-permission-repaint.json', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts'],
'plan-ceo-split-overflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json',
'plan-ceo-split-overflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'lib/claude-public-transcript.ts', 'test/fixtures/plan-create-prepublication-491.json', 'test/fixtures/plan-create-combined-permission-70b.json',
'test/fixtures/plan-create-permission-361c.json',
@@ -551,16 +543,15 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// Agent SDK. Gate-tier (deterministic stub server, fixed inputs); fires
// when the skill template, the verify helper, the artifacts-init helper,
// or the detect script changes.
'setup-gbrain-remote': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/helpers/setup-gbrain-sandbox.ts', 'test/helpers/office-hours-attempt.ts', 'test/helpers/eval-store.ts', 'test/helpers/e2e-helpers.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'lib/eval-model.ts','setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/claude-md-persist.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'bin/gstack-artifacts-init', 'bin/gstack-gbrain-detect', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-remote.test.ts',
'setup-gbrain-remote': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/setup-gbrain-sandbox.ts', 'test/helpers/office-hours-attempt.ts', 'test/helpers/eval-store.ts', 'test/helpers/e2e-helpers.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'lib/eval-model.ts','setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/claude-md-persist.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'bin/gstack-artifacts-init', 'bin/gstack-gbrain-detect', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-remote.test.ts',
'test/helpers/e2e-gate.ts'],
'setup-gbrain-bad-token': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-bad-token.test.ts',
'setup-gbrain-bad-token': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-bad-token.test.ts',
'test/helpers/setup-gbrain-sandbox.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/setup-gbrain-fixture-command.ts', 'bin/gstack-gbrain-detect', 'lib/gbrain-local-status.ts', 'lib/gbrain-exec.ts', 'bin/gstack-gbrain-lib.sh', 'bin/gstack-egress-lib.sh', 'bin/gstack-egress-receipt', 'test/helpers/office-hours-attempt.ts', 'test/helpers/e2e-gate.ts'],
// v1.34.0.0 split-engine Path 4 + Step 4.5 Yes (local PGLite for code).
// Periodic-tier per codex #12 (AgentSDK harness is non-deterministic).
// Fires when the setup-gbrain template, install/verify/init helpers, or
// the agent-sdk-runner harness changes.
'setup-gbrain-path4-local-pglite': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/claude-md-persist.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'bin/gstack-gbrain-install', 'bin/gstack-gbrain-detect', 'lib/gbrain-local-status.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-path4-local-pglite.test.ts',
'setup-gbrain-path4-local-pglite': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'setup-gbrain/sections/brain-init.md.tmpl', 'setup-gbrain/sections/claude-md-persist.md.tmpl', 'setup-gbrain/sections/manifest.json', 'test/helpers/setup-gbrain-fixture.ts', 'setup-gbrain/SKILL.md.tmpl', 'bin/gstack-gbrain-mcp-verify', 'bin/gstack-gbrain-install', 'bin/gstack-gbrain-detect', 'lib/gbrain-local-status.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-setup-gbrain-path4-local-pglite.test.ts',
'test/helpers/setup-gbrain-sandbox.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/setup-gbrain-fixture-command.ts', 'lib/gbrain-exec.ts', 'bin/gstack-gbrain-lib.sh', 'bin/gstack-egress-lib.sh', 'bin/gstack-egress-receipt', 'test/helpers/office-hours-attempt.ts', 'test/helpers/e2e-gate.ts'],
// AskUserQuestion format regression (RECOMMENDATION + Completeness: N/10)
@@ -617,7 +608,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// Expanded coverage (CT3) — 6 non-plan-review skills inherit Pros/Cons via preamble
// /plan-tune (v1 observational)
'plan-tune-inspect': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'plan-tune/**', 'scripts/question-registry.ts', 'scripts/psychographic-signals.ts', 'scripts/one-way-doors.ts', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'bin/gstack-developer-profile', 'test/skill-e2e-plan-tune.test.ts'],
'plan-tune-inspect': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'plan-tune/**', 'scripts/question-registry.ts', 'scripts/psychographic-signals.ts', 'scripts/one-way-doors.ts', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'bin/gstack-developer-profile', 'test/skill-e2e-plan-tune.test.ts'],
// /plan-tune cathedral (T16 — 5 E2E scenarios, all gate per D12)
@@ -661,7 +652,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/helpers/ship-hook-actor.ts',
'test/helpers/workflow-excerpt.ts', 'test/helpers/agent-sdk-runner.ts', 'test/skill-e2e-ship-hook-consent.test.ts',
'test/helpers/e2e-gate.ts'],
'ship-base-branch': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'bin/gstack-repo-mode', 'test/skill-e2e-review-attribution.test.ts',
'ship-base-branch': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'bin/gstack-repo-mode', 'test/skill-e2e-review-attribution.test.ts',
'scripts/resolvers/testing.ts'
],
'ship-local-workflow': [ 'ship/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-workflow.test.ts',
@@ -671,29 +662,29 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
],
// Retro
'retro': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-retro-metrics', 'retro/**', 'test/skill-e2e-retro.test.ts'],
'retro-base-branch': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-retro-metrics', 'retro/**', 'test/skill-e2e-retro.test.ts'],
'retro': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-retro-metrics', 'retro/**', 'test/skill-e2e-retro.test.ts'],
'retro-base-branch': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-retro-metrics', 'retro/**', 'test/skill-e2e-retro.test.ts'],
// CSO
'cso-full-audit': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
'cso-diff-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
'cso-infra-scope': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
'cso-full-audit': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
'cso-diff-mode': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
'cso-infra-scope': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'cso/**', 'lib/cso/**', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/skill-e2e-cso.test.ts'],
// Learnings
'learnings-show': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'learn/**', 'bin/gstack-learnings-search', 'bin/gstack-learnings-log', 'scripts/resolvers/learnings.ts', 'test/skill-e2e-learnings.test.ts'],
'learnings-show': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'learn/**', 'bin/gstack-learnings-search', 'bin/gstack-learnings-log', 'scripts/resolvers/learnings.ts', 'test/skill-e2e-learnings.test.ts'],
// Session Intelligence (timeline, context recovery, /context-save + /context-restore)
'timeline-event-flow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-timeline-log', 'bin/gstack-timeline-read', 'test/skill-e2e-session-intelligence.test.ts'],
'context-recovery-artifacts': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'scripts/resolvers/preamble.ts', 'bin/gstack-timeline-log', 'bin/gstack-slug', 'learn/**', 'test/skill-e2e-session-intelligence.test.ts'],
'context-save-writes-file': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-save/**', 'bin/gstack-slug', 'test/skill-e2e-session-intelligence.test.ts'],
'context-restore-loads-latest': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-restore/**', 'bin/gstack-slug', 'test/skill-e2e-session-intelligence.test.ts'],
'timeline-event-flow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'bin/gstack-timeline-log', 'bin/gstack-timeline-read', 'test/skill-e2e-session-intelligence.test.ts'],
'context-recovery-artifacts': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'scripts/resolvers/preamble.ts', 'bin/gstack-timeline-log', 'bin/gstack-slug', 'learn/**', 'test/skill-e2e-session-intelligence.test.ts'],
'context-save-writes-file': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-save/**', 'bin/gstack-slug', 'test/skill-e2e-session-intelligence.test.ts'],
'context-restore-loads-latest': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-restore/**', 'bin/gstack-slug', 'test/skill-e2e-session-intelligence.test.ts'],
// Context skills E2E (live-fire, Skill-tool routing path) — see
// test/skill-e2e-context-skills.test.ts. These are periodic-tier because
// each one spawns claude -p and costs ~$0.20-$0.40. Collectively they
// verify the thing the /checkpoint → /context-save rename was for.
'context-save-routing': [ 'context-save/**', 'scripts/resolvers/preamble.ts', 'test/skill-e2e-context-skills.test.ts'],
'context-save-then-restore-roundtrip': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-save/**', 'context-restore/**', 'bin/gstack-slug', 'test/skill-e2e-context-skills.test.ts'],
'context-save-then-restore-roundtrip': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'context-save/**', 'context-restore/**', 'bin/gstack-slug', 'test/skill-e2e-context-skills.test.ts'],
'context-restore-fragment-match': [ 'context-restore/**', 'test/skill-e2e-context-skills.test.ts'],
'context-restore-empty-state': [ 'context-restore/**', 'test/skill-e2e-context-skills.test.ts'],
'context-restore-list-delegates': [ 'context-restore/**', 'test/skill-e2e-context-skills.test.ts'],
@@ -724,8 +715,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/helpers/skill-fixture.ts', 'test/helpers/outside-voice-fixture.ts', 'test/helpers/outside-voice-evidence.ts', 'test/fixtures/outside-async-task-m-events.json', 'test/skill-e2e-outside-voice.test.ts',
'test/helpers/outside-voice-receipt.ts',
'test/helpers/e2e-gate.ts'],
'outside-voice-claude-code-to-codex': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'review/**', 'codex/**', 'hosts/claude.ts', 'hosts/define-host.ts',
'outside-voice-claude-code-to-codex': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'review/**', 'codex/**', 'hosts/claude.ts', 'hosts/define-host.ts',
'scripts/gen-skill-docs.ts', 'scripts/resolvers/index.ts', 'scripts/resolvers/outside-voice.ts',
'scripts/resolvers/constants.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts',
'bin/gstack-codex-probe', 'lib/outside-review-result.ts', 'test/helpers/session-runner.ts',
@@ -733,8 +723,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/helpers/skill-fixture.ts', 'test/helpers/outside-voice-fixture.ts', 'test/helpers/outside-voice-evidence.ts', 'test/fixtures/outside-async-task-m-events.json', 'test/skill-e2e-outside-voice.test.ts', 'test/helpers/codex-session-runner.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/outside-voice-receipt.ts'],
// Disabled means no extra plan review, including a native Agent fallback.
'outside-plan-disabled-no-fallback': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/fixtures/disabled-retained-record.json',
'outside-plan-disabled-no-fallback': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/fixtures/disabled-retained-record.json',
'scripts/resolvers/testing.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts',
"test/fixtures/plan-scope-recovery-av.json",
@@ -759,9 +748,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// Coverage audit (shared fixture) + triage + gates
'ship-coverage-audit': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
"test/fixtures/coverage-audit-aw.json",
'ship-coverage-audit': ['bin/gstack-state-root.sh', 'lib/state-root.ts', "test/fixtures/coverage-audit-aw.json",
"test/fixtures/coverage-checkbox-tail-av.json",
@@ -804,11 +791,9 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts", "scripts/resolvers/preamble/generate-completion-status.ts",
'test/helpers/coverage-audit.ts', 'test/helpers/office-hours-attempt.ts', 'scripts/resolvers/testing.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts'
],
'ship-triage': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'bin/gstack-repo-mode', 'test/skill-e2e-triage.test.ts',
'ship-triage': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'bin/gstack-repo-mode', 'test/skill-e2e-triage.test.ts',
'scripts/resolvers/testing.ts'
],
'ship-docsync': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/sections.ts', 'test/skill-e2e-ship-docsync.test.ts',
'scripts/resolvers/testing.ts', 'test/helpers/docsync-*.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
'ship-docsync-completion': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'test/skill-e2e-ship-docsync.test.ts', 'test/helpers/docsync-*.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'scripts/resolvers/sections.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/testing.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
'ship-docsync-current': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'test/skill-e2e-ship-docsync.test.ts', 'test/helpers/docsync-*.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'scripts/resolvers/sections.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/testing.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
'ship-docsync-failure': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/**', 'document-release/**', 'test/skill-e2e-ship-docsync.test.ts', 'test/helpers/docsync-*.ts', 'test/helpers/session-runner.ts', 'test/helpers/hermetic-env.ts', 'bin/gstack-skill-start', 'bin/gstack-session-kind', 'scripts/resolvers/sections.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/testing.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
@@ -817,8 +802,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// spawned-marked subagent. Deps name every behavior under test — the
// session-kind override, the skill-start gates, both hooks + the shared
// directive, and the AUQ prose rule — so changing any of them selects it.
'docsync-spawned': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'hosts/claude/hooks/hook-log.ts',
'ship/sections/pr-body.md',
'docsync-spawned': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'ship/sections/pr-body.md',
'document-release/**',
'ship/sections/documentation.md',
'ship/sections/documentation.md.tmpl',
@@ -829,13 +813,13 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'hosts/claude/hooks/auq-error-fallback-hook.ts',
'hosts/claude/hooks/spawned-directive.ts',
'scripts/resolvers/preamble/generate-ask-user-format.ts',
'test/skill-e2e-docsync-spawned.test.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
'test/skill-e2e-docsync-spawned.test.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts', 'hosts/claude/hooks/hook-log.ts'],
// Design
'design-consultation-core': [ 'design-consultation/**', 'lib/design-catalog.ts', 'lib/design-md.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-design.test.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-existing': [ 'design-consultation/**', 'lib/design-md.ts', 'bin/gstack-design-md.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-research': [ 'design-consultation/**', 'scripts/resolvers/aside.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/helpers/skill-fixture.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-preview': [ 'design-consultation/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-core': [ 'design-consultation/**', 'lib/design-catalog.ts', 'lib/design-md.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-design.test.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-existing': [ 'design-consultation/**', 'lib/design-md.ts', 'bin/gstack-design-md.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-research': [ 'design-consultation/**', 'scripts/resolvers/aside.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/helpers/skill-fixture.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-preview': [ 'design-consultation/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'plan-design-review-no-ui-scope': [
"test/fixtures/plan-scope-recovery-av.json",
@@ -843,16 +827,16 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts", "scripts/resolvers/preamble/generate-completion-status.ts",
'scripts/resolvers/preamble/generate-ask-user-format.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-fix': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/aside.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-catalog.ts', 'browse/src/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'scripts/resolvers/preamble/generate-ask-user-format.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-fix': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/aside.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-catalog.ts', 'browse/src/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
// Design detector (user-installed impeccable engine) through the fake engine shim: source mode on a diff and DOM mode on a served page.
'design-review-detector-shim': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'design-review-detector-shim': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-detector-shim-dom': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'browse/src/**', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'design-review-detector-shim-dom': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'browse/src/**', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-plugin-handoff': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/testing.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.html', 'test/skill-e2e-design.test.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-html-slop-gate': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-html/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/skill-e2e-design.test.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-plugin-handoff': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/testing.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/skill-e2e-design.test.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-html-slop-gate': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'design-html/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/skill-e2e-design.test.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
// /diagram (diagram-render bundle consumers). Triplet = deterministic
// functional (gate); authoring quality = LLM-judged benchmark (periodic).
@@ -860,24 +844,23 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// bin/gstack-render.ts): Aside when it is running, the browse daemon
// otherwise — so both engines are deps. Triplet = deterministic functional
// (gate); authoring quality = LLM-judged benchmark (periodic).
'diagram-triplet': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'diagram/**', 'lib/diagram-render/**', 'lib/aside-render.ts', 'bin/gstack-render.ts', 'test/helpers/aside-available.ts', 'browse/src/**', 'test/skill-e2e-diagram.test.ts', 'test/helpers/llm-judge.ts'],
'diagram-authoring-quality': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'diagram/**', 'lib/diagram-render/**', 'lib/aside-render.ts', 'bin/gstack-render.ts', 'test/helpers/aside-available.ts', 'browse/src/**', 'test/helpers/llm-judge.ts', 'test/skill-e2e-diagram.test.ts'],
'diagram-triplet': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'diagram/**', 'lib/diagram-render/**', 'lib/aside-render.ts', 'bin/gstack-render.ts', 'test/helpers/aside-available.ts', 'browse/src/**', 'test/skill-e2e-diagram.test.ts', 'test/helpers/llm-judge.ts'],
'diagram-authoring-quality': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'diagram/**', 'lib/diagram-render/**', 'lib/aside-render.ts', 'bin/gstack-render.ts', 'test/helpers/aside-available.ts', 'browse/src/**', 'test/helpers/llm-judge.ts', 'test/skill-e2e-diagram.test.ts'],
// gstack-upgrade
'gstack-upgrade-happy-path': [ 'gstack-upgrade/**', 'test/skill-e2e-workflow.test.ts', 'test/fixtures/coverage-audit-fixture.ts', 'test/helpers/coverage-audit.ts', 'test/helpers/office-hours-attempt.ts'],
// Deploy skills
'land-and-deploy-workflow': [ 'land-and-deploy/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-deploy.test.ts'],
'land-and-deploy-first-run': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'land-and-deploy/**', 'scripts/gen-skill-docs.ts', 'bin/gstack-slug', 'test/skill-e2e-deploy.test.ts'],
'land-and-deploy-review-gate': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'land-and-deploy/**', 'bin/gstack-review-read', 'test/skill-e2e-deploy.test.ts'],
'canary-workflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'canary/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'test/skill-e2e-deploy.test.ts'],
'benchmark-workflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'benchmark/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'test/skill-e2e-deploy.test.ts'],
'land-and-deploy-first-run': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'land-and-deploy/**', 'scripts/gen-skill-docs.ts', 'bin/gstack-slug', 'test/skill-e2e-deploy.test.ts'],
'land-and-deploy-review-gate': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'land-and-deploy/**', 'bin/gstack-review-read', 'test/skill-e2e-deploy.test.ts'],
'canary-workflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'canary/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'test/skill-e2e-deploy.test.ts'],
'benchmark-workflow': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'benchmark/**', 'scripts/resolvers/aside.ts', 'browse/src/**', 'test/skill-e2e-deploy.test.ts'],
'setup-deploy-workflow': [ 'setup-deploy/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-deploy.test.ts'],
// Autoplan
'autoplan-dual-voice': ['bin/gstack-state-root.sh', 'lib/state-root.ts',
'test/helpers/hermetic-env.ts',
'autoplan-dual-voice': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'test/helpers/hermetic-env.ts',
'test/helpers/autoplan-dual-voice-evidence.ts',
'test/fixtures/autoplan-dual-false-positive-6bd.json',
'test/helpers/autoplan-method-read-audit.ts',
@@ -1043,7 +1026,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// frontmatter. Touched by anything that changes resolver output, gen
// pipeline, detection helper, refresh subcommand, or the on-demand
// docs the resolver points to.
'office-hours-brain-writeback': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'office-hours/sections/**',
'office-hours-brain-writeback': ['bin/gstack-state-root.sh', 'lib/state-root.ts', 'office-hours/sections/**',
'scripts/resolvers/gbrain.ts',
'scripts/gen-skill-docs.ts',
'bin/gstack-gbrain-detect',
@@ -1097,6 +1080,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'ship/SKILL.md', 'test/helpers/e2e-gate.ts'],
'office-hours-section-loading': [ 'office-hours/**', 'bin/gstack-office-hours-review', 'lib/office-hours-review.ts', 'lib/fs-atomic.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/sections.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/carve-guards.ts', 'test/helpers/auq-sdk-capture.ts', 'test/helpers/office-hours-completion.ts', 'test/helpers/llm-judge.ts', 'test/helpers/session-runner.ts', 'test/skill-e2e-office-hours-section-loading.test.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/auq-native-capture.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/capture-parity-baseline.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/parity-harness.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/plan-skill-questions.ts', 'test/helpers/pty-screen.ts', 'test/helpers/pty-trust-dialog.ts', 'test/helpers/skill-census.ts'],
'office-hours-design-draft': [ 'office-hours/**', 'lib/office-hours-review.ts', 'scripts/resolvers/review-dashboard.ts', 'scripts/resolvers/plan-gates.ts', 'scripts/resolvers/spec-review.ts', 'scripts/resolvers/outside-voice-steps.ts', 'scripts/resolvers/review-scope.ts', 'scripts/resolvers/outside-voice.ts', 'scripts/resolvers/sections.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/carve-guards.ts', 'test/helpers/auq-sdk-capture.ts', 'test/helpers/office-hours-completion.ts', 'test/helpers/llm-judge.ts', 'test/helpers/session-runner.ts', 'test/skill-e2e-office-hours-design-draft.test.ts', 'test/helpers/agent-sdk-runner.ts', 'test/helpers/auq-native-capture.ts', 'test/helpers/auto-decision-state.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/capture-parity-baseline.ts', 'test/helpers/claude-pty-runner.ts', 'test/helpers/pty/**', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/native-auto-decide.ts', 'test/helpers/owned-claude-transcript.ts', 'test/helpers/parity-harness.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-count-transcript.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts', 'test/helpers/plan-scope-selection.ts', 'test/helpers/plan-seed-submission.ts', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/plan-skill-questions.ts', 'test/helpers/pty-screen.ts', 'test/helpers/pty-trust-dialog.ts', 'test/helpers/skill-census.ts'],
'plan-devex-peer-comparison-classification': [
@@ -1122,10 +1106,11 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
};
/**
* E2E test tiers — 'gate' blocks PRs, 'periodic' runs weekly/on-demand.
* E2E test tiers — 'gate' blocks PRs, 'periodic' runs weekly/on-demand,
* 'marathon' keeps full start-to-finish flows in a non-blocking lane only.
* Must have exactly the same keys as E2E_TOUCHFILES.
*/
export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
export const E2E_TIERS: Record<string, 'gate' | 'periodic' | 'marathon'> = {
'ship-skipped-queued-finding': 'gate',
'investigate-owned-completion': 'gate',
'investigate-owned-abort': 'gate',
@@ -1250,6 +1235,7 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
'plan-ceo-review-plan-mode': 'gate',
'plan-eng-review-plan-mode': 'periodic',
'plan-design-review-plan-mode': 'periodic',
'plan-design-review-plan-mode-smoke': 'periodic',
'plan-devex-review-plan-mode': 'gate',
'plan-mode-no-op': 'gate',
// v1.21+ auto-mode regression tests
@@ -1283,7 +1269,7 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
'plan-design-finding-floor': 'periodic', // stochastic ask-first (see plan-mode-handshake note); periodic
'plan-devex-finding-floor': 'gate',
'plan-eng-multi-finding-batching': 'periodic',
'plan-ceo-split-overflow': 'periodic',
'plan-ceo-split-overflow': 'marathon', // Full /plan-ceo-review through split overflow (504–1188 s on 2.1.251)
// Privacy gate for gstack-brain-sync — periodic (non-deterministic LLM call,
// costs ~$0.30-$0.50 per run, not needed on every commit)
@@ -1355,7 +1341,6 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
'ship-local-hook-preservation': 'gate',
'ship-coverage-audit': 'gate',
'ship-triage': 'gate',
'ship-docsync': 'gate',
'ship-docsync-missing-marker': 'gate',
'ship-docsync-missing-asset': 'gate',
'ship-docsync-launch-failure': 'gate',
@@ -1480,7 +1465,8 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic'> = {
'arm-benchmark-native-overbuild': 'periodic',
'arm-benchmark-crud-endpoint': 'periodic',
'arm-benchmark-bugfix-decoys': 'periodic',
'office-hours-section-loading': 'periodic', // Full startup design/review/approval workflow
'office-hours-section-loading': 'marathon', // Full startup design/review/approval workflow (1–3 real review rounds, ~20 min)
'office-hours-design-draft': 'periodic', // Same interview through the design-creating Write (~5 min)
'plan-decision-classification': 'periodic',
'plan-devex-peer-comparison-classification': 'periodic',
'health-reporting': 'periodic',
@@ -1563,6 +1549,9 @@ export const GLOBAL_TOUCHFILES = [
'scripts/lib/shard-engine.ts', // The shard engine test-strict-output.ts re-exports (moved there in the W2 refactor)
// Canonical paid execution and its shared time allocations affect every paid test.
'scripts/test-paid-shards.ts',
'scripts/lib/paid-cases.ts', // Case/trial shard keys, plan and report moved out of test-paid-shards.ts
'scripts/lib/paid-plan.ts',
'scripts/lib/paid-report.ts',
'scripts/test-pr-profile.ts',
'test/helpers/eval-budgets.ts',
@@ -1581,3 +1570,310 @@ export const GLOBAL_TOUCHFILES = [
// diffed per key, so a data-only edit runs just the affected tests.
// Map-diff fails CLOSED — any error on that path still runs everything.
];
/**
* Eval kind per live case (every E2E_TIERS and LLM_JUDGE_TOUCHFILES key).
* The kind fixes the trial policy before the run (EVAL_POLICY in
* periodic-exclude-data.ts):
* rule - one trial; any failed assertion fails the verdict. The default.
* behavior - a panel of independent trials, PASS at the policy majority;
* needs a BEHAVIOR_WHY entry naming the tolerated deviation.
* judge - an LLM-judge score of a static input, sampled as a panel.
* Reclassification is a reviewed diff, never a runtime switch.
*/
export const E2E_KINDS: Record<string, 'rule' | 'behavior' | 'judge'> = {
'ship-skipped-queued-finding': 'rule',
'investigate-owned-completion': 'rule',
'investigate-owned-abort': 'rule',
'investigate-owned-ending-error': 'rule',
'shared-libs-review-path-eligibility': 'rule',
'shared-libs-review-index-flags': 'rule',
'shared-libs-review-prior-coverage': 'rule',
'shared-libs-codex-read-only': 'rule',
'shared-libs-read-only': 'rule',
'shared-libs-unsupported-git': 'rule',
'shared-libs-review-lifecycle': 'rule',
'shared-libs-review-revalidation': 'rule',
'shared-libs-opportunity-judgment': 'behavior',
'shared-libs-pr-coverage': 'rule',
'shared-libs-plan-callers': 'rule',
'browse-basic': 'rule',
'browse-snapshot': 'rule',
'aside-browse-basic': 'rule',
'aside-browse-flow': 'rule',
'aside-qa-quick': 'rule',
'aside-scrape-json': 'rule',
'aside-canary-quick': 'rule',
'hermetic-canary': 'rule',
'hermetic-sentinel': 'rule',
'skillmd-setup-discovery': 'rule',
'skillmd-no-local-binary': 'rule',
'skillmd-outside-git': 'rule',
'session-awareness': 'rule',
'operational-learning': 'rule',
'first-task-scaffold': 'rule',
'qa-quick': 'rule',
'qa-b6-static': 'rule',
'qa-b7-spa': 'rule',
'qa-b8-checkout': 'rule',
'qa-only-no-fix': 'rule',
'qa-fix-loop': 'rule',
'qa-bootstrap': 'rule',
'review-exploratory-small-cli': 'rule',
'ship-exploratory-small-cli': 'rule',
'ship-exploratory-unavailable': 'rule',
'ship-exploratory-plan-checks': 'rule',
'ship-exploratory-late-input': 'rule',
'qa-functional-cli-report': 'rule',
'qa-functional-webhook-report': 'rule',
'qa-functional-cli-fix': 'rule',
'qa-functional-webhook-fix': 'rule',
'review-sql-injection': 'rule',
'review-enum-completeness': 'rule',
'review-base-branch': 'rule',
'review-design-lite': 'behavior',
'review-coverage-audit': 'rule',
'review-dashboard-via': 'rule',
'review-army-migration-safety': 'rule',
'review-army-perf-n-plus-one': 'rule',
'review-army-delivery-audit': 'rule',
'review-army-quality-score': 'rule',
'review-army-json-findings': 'rule',
'review-army-red-team': 'behavior',
'review-army-consensus': 'behavior',
'review-army-simplification': 'behavior',
'review-army-simplification-precision': 'behavior',
'office-hours-spec-review': 'rule',
'office-hours-brain-writeback': 'behavior',
'gbrain-roundtrip-local': 'rule',
'sync-gbrain-read-ready': 'rule',
'sync-gbrain-read-unknown': 'rule',
'office-hours-forcing-energy': 'behavior',
'office-hours-builder-wildness': 'behavior',
'plan-ceo-review': 'rule',
'plan-ceo-review-selective': 'rule',
'plan-ceo-review-benefits': 'rule',
'plan-ceo-review-expansion-energy': 'behavior',
'plan-eng-review': 'rule',
'plan-eng-review-artifact': 'rule',
'plan-eng-coverage-audit': 'rule',
'plan-review-report': 'rule',
'plan-ceo-review-plan-mode': 'rule',
'plan-eng-review-plan-mode': 'rule',
'plan-design-review-plan-mode': 'rule',
'plan-design-review-plan-mode-smoke': 'rule',
'ship-coverage-value': 'rule',
'review-test-value': 'rule',
'test-audit-report-only': 'rule',
'plan-devex-review-plan-mode': 'rule',
'plan-mode-no-op': 'rule',
'office-hours-auto-mode': 'rule',
'auto-decide-preserved': 'rule',
'auq-format-gate': 'rule',
'plan-ceo-mode-routing': 'rule',
'plan-design-with-ui-scope': 'rule',
'tpa-present': 'rule',
'tpa-absent-linux': 'rule',
'tpa-broken': 'rule',
'tpa-absent-darwin': 'rule',
'tpa-apple-ban': 'rule',
'ship-section-loading': 'rule',
'plan-ceo-section-loading': 'rule',
'carve-section-loading': 'rule',
'plan-eng-finding-floor': 'rule',
'plan-ceo-finding-floor': 'rule',
'plan-design-finding-floor': 'rule',
'plan-devex-finding-floor': 'rule',
'plan-eng-multi-finding-batching': 'rule',
'plan-ceo-split-overflow': 'rule',
'setup-gbrain-remote': 'rule',
'setup-gbrain-bad-token': 'rule',
'setup-gbrain-path4-local-pglite': 'rule',
'plan-ceo-review-format-mode': 'behavior',
'plan-ceo-review-format-approach': 'behavior',
'plan-eng-review-format-coverage': 'behavior',
'plan-eng-review-format-kind': 'behavior',
'office-hours-phase4-fork': 'behavior',
'llm-judge-recommendation': 'judge',
'plan-ceo-review-prosons-cadence': 'behavior',
'plan-review-prosons-format': 'behavior',
'plan-review-prosons-hardstop-neg': 'behavior',
'plan-review-prosons-neutral-neg': 'behavior',
'plan-tune-inspect': 'rule',
'codex-offered-office-hours': 'rule',
'codex-offered-ceo-review': 'rule',
'codex-offered-design-review': 'rule',
'codex-offered-eng-review': 'rule',
'timeline-event-flow': 'rule',
'context-recovery-artifacts': 'rule',
'context-save-writes-file': 'rule',
'context-restore-loads-latest': 'rule',
'context-save-routing': 'rule',
'context-save-then-restore-roundtrip': 'rule',
'context-restore-fragment-match': 'rule',
'context-restore-empty-state': 'rule',
'context-restore-list-delegates': 'rule',
'context-restore-legacy-compat': 'rule',
'context-save-list-current-branch': 'rule',
'context-save-list-all-branches': 'rule',
'ship-base-branch': 'rule',
'ship-local-workflow': 'rule',
'ship-managed-hook-refresh': 'rule',
'ship-unmanaged-hook-consent': 'rule',
'ship-local-hook-preservation': 'rule',
'ship-coverage-audit': 'rule',
'ship-triage': 'rule',
'ship-docsync-missing-marker': 'rule',
'ship-docsync-missing-asset': 'rule',
'ship-docsync-launch-failure': 'rule',
'ship-docsync-timeout-unsettled': 'rule',
'ship-docsync-late-result': 'rule',
'ship-docsync-stale-before': 'rule',
'ship-docsync-stale-after': 'rule',
'ship-docsync-recovery': 'rule',
'ship-docsync-completion': 'rule',
'ship-docsync-current': 'rule',
'ship-docsync-failure': 'rule',
'ship-docsync-store': 'rule',
'docsync-spawned': 'rule',
'retro': 'rule',
'retro-base-branch': 'rule',
'cso-full-audit': 'rule',
'cso-diff-mode': 'rule',
'cso-infra-scope': 'rule',
'learnings-show': 'rule',
'document-release': 'rule',
'codex-review': 'rule',
'codex-discover-skill': 'rule',
'codex-review-findings': 'rule',
'outside-voice-codex-to-claude-code': 'rule',
'outside-voice-claude-code-to-codex': 'rule',
'outside-plan-disabled-no-fallback': 'rule',
'codex-sol-scope-termination': 'rule',
'design-consultation-core': 'rule',
'design-consultation-existing': 'rule',
'design-consultation-research': 'rule',
'design-consultation-preview': 'rule',
'plan-design-review-no-ui-scope': 'rule',
'design-review-fix': 'rule',
'design-review-detector-shim': 'rule',
'design-review-detector-shim-dom': 'rule',
'design-review-plugin-handoff': 'rule',
'design-html-slop-gate': 'behavior',
'diagram-triplet': 'rule',
'diagram-authoring-quality': 'rule',
'gstack-upgrade-happy-path': 'rule',
'land-and-deploy-workflow': 'rule',
'land-and-deploy-first-run': 'rule',
'land-and-deploy-review-gate': 'rule',
'canary-workflow': 'rule',
'benchmark-workflow': 'rule',
'setup-deploy-workflow': 'rule',
'autoplan-dual-voice': 'rule',
'benchmark-providers-live': 'rule',
'scrape-match-path': 'behavior',
'scrape-prototype-path': 'behavior',
'skillify-happy-path': 'rule',
'skillify-provenance-refusal': 'rule',
'skillify-approval-reject': 'rule',
'journey-ideation': 'rule',
'journey-plan-eng': 'rule',
'journey-debug': 'rule',
'journey-qa': 'rule',
'journey-code-review': 'rule',
'journey-ship': 'rule',
'journey-docs': 'rule',
'journey-retro': 'rule',
'journey-design-system': 'rule',
'journey-visual-qa': 'rule',
'ios-qa-device': 'rule',
'arm-benchmark-native-overbuild': 'rule',
'arm-benchmark-crud-endpoint': 'rule',
'arm-benchmark-bugfix-decoys': 'rule',
'office-hours-section-loading': 'rule',
'office-hours-design-draft': 'rule',
'plan-decision-classification': 'rule',
'plan-devex-peer-comparison-classification': 'rule',
'health-reporting': 'rule',
'overlay-harness-claude-dedicated-tools-vs-bash': 'rule',
'overlay-harness-opus-4-7-effort-match-trivial': 'rule',
'overlay-harness-opus-4-7-literal-interpretation': 'rule',
'overlay-harness-claude-dedicated-tools-vs-bash-sonnet': 'rule',
'journey-negatives': 'rule',
'review/SKILL.md workflow': 'judge',
'setup-browser-cookies/SKILL.md workflow': 'judge',
'browse/SKILL.md reference': 'judge',
'setup block': 'judge',
'qa/SKILL.md workflow': 'judge',
'qa/SKILL.md health rubric': 'judge',
'qa/SKILL.md anti-refusal': 'judge',
'cross-skill greptile consistency': 'judge',
'ship/SKILL.md workflow': 'judge',
'document-release/SKILL.md workflow': 'judge',
'plan-ceo-review/SKILL.md modes': 'judge',
'plan-eng-review/SKILL.md sections': 'judge',
'plan-design-review/SKILL.md passes': 'judge',
'design-review/SKILL.md fix loop': 'judge',
'design-consultation/SKILL.md research': 'judge',
'land-and-deploy/SKILL.md workflow': 'judge',
'canary/SKILL.md monitoring loop': 'judge',
'benchmark/SKILL.md perf collection': 'judge',
'setup-deploy/SKILL.md platform setup': 'judge',
'retro/SKILL.md instructions': 'judge',
'qa-only/SKILL.md workflow': 'judge',
'gstack-upgrade/SKILL.md upgrade flow': 'judge',
'sync-gbrain/SKILL.md read-only readiness': 'judge',
'voice directive tone': 'judge',
};
/**
* One-line tolerance for every behavior-kind case: why an occasional
* deviation is acceptable product behavior. Keys equal the behavior ids of
* E2E_KINDS; values are non-empty.
*/
export const BEHAVIOR_WHY: Record<string, string> = {
'shared-libs-opportunity-judgment':
"Whether a candidate extraction is worth recommending is a judgment call; the read-only invariant stays a contract.",
'review-design-lite':
"How many of the seven design-lite checklist items the live review flags varies run to run; the fake-engine rows it must carry stay strict.",
'review-army-red-team':
"Whether the red-team lens surfaces on a small diff is a live model choice, not a contract.",
'review-army-consensus':
"Multi-specialist agreement on the planted SQL finding is a quality benchmark that tolerates an occasional miss.",
'review-army-simplification':
"Flagging the planted unnecessary structure is an advisory-lens quality judgment.",
'review-army-simplification-precision':
"Staying silent on a lean diff is a false-flag noise benchmark; an occasional advisory is acceptable noise.",
'office-hours-forcing-energy':
"The Q3 posture is scored by a live judge on generated prose; a single flat phrasing is tolerable.",
'office-hours-builder-wildness':
"Builder-mode creativity is scored by a live judge on generated prose; one conservative riff is tolerable.",
'office-hours-brain-writeback':
"The model's interpretation of the gbrain writeback instruction (page shape, tags) varies; no secret or safety step rides on it.",
'office-hours-phase4-fork':
"Phase 4 asks the model to invent 2-3 architectures; surfacing the fork with its reasoning is open-ended generation.",
'plan-ceo-review-expansion-energy':
"Expansion framing is scored by a live judge on generated proposals; one flat proposal set is tolerable.",
'plan-ceo-review-format-mode':
"Mode-question wording (Completeness line vs kind note) is live formatting of one AskUserQuestion.",
'plan-ceo-review-format-approach':
"Approach-menu Completeness wording is live formatting of one AskUserQuestion.",
'plan-eng-review-format-coverage':
"Coverage-issue Completeness wording is live formatting of one AskUserQuestion.",
'plan-eng-review-format-kind':
"Kind-note wording is live formatting of one AskUserQuestion.",
'plan-ceo-review-prosons-cadence':
"Pros/Cons cadence on a hard-stop question is live formatting; either the escape or the full block is accepted.",
'plan-review-prosons-format':
"The full Pros/Cons block (counts of pros and cons, labels) is live formatting of one question.",
'plan-review-prosons-hardstop-neg':
"Not using the hard-stop escape on an ordinary decision is live formatting of one question.",
'plan-review-prosons-neutral-neg':
"Avoiding neutral posture and naming a because-reason is live formatting of one question.",
'design-html-slop-gate':
"How many scan passes the one-pass slop gate takes on a fake engine's fixed output is a judgment call.",
'scrape-match-path':
"The /scrape fallback no longer prescribes the browser-skills match flow, so taking it is prompt compliance.",
'scrape-prototype-path':
"The /scrape fallback no longer prescribes the prototype flow, so taking it is prompt compliance.",
};
+30 -47
View File
@@ -1,13 +1,12 @@
/** Audited cache adapter for runWorkflowJudge only. Native/PTY evals stay fresh. */
import * as fs from 'node:fs';
import * as path from 'node:path';
import { isBuiltin } from 'node:module';
import { spawnSync } from 'node:child_process';
import { DEFAULT_JUDGE_MAX_TOKENS, resolveEvalModel } from '../../lib/eval-model';
import { JUDGE_MS } from './eval-budgets';
import type { JudgeScore } from './llm-judge';
import { JUDGE_PANEL_SAMPLES, JUDGE_SCORE_DIMENSIONS, judgePanelMean, type JudgeScore } from './llm-judge';
import { readWorkflowJudgeInput, buildWorkflowJudgePrompt, WORKFLOW_JUDGE_RESPONSE_SCHEMA, WORKFLOW_JUDGE_REASONING_WORD_LIMIT } from './workflow-judge-input';
import { buildEvalInputIdentity, lookupEvalInputCache, storeEvalInputCache,
import { buildEvalInputIdentity, lookupEvalInputCache, sourceDependencyClosure, storeEvalInputCache,
type EvalCacheValue, type EvalInputIdentity, type EvalPassingProof } from '../../scripts/eval-input-cache';
type Thresholds = { clarity: number; completeness: number; actionability: number };
@@ -19,50 +18,20 @@ export interface WorkflowCacheOptions {
structuredResponse?: boolean;
maxTokens?: number;
stream?: boolean;
effort?: 'medium';
env?: NodeJS.ProcessEnv;
}
export interface WorkflowJudgeReuse {
key: string; source: EvalPassingProof['source'];
}
/** Follow literal module imports, including installed SDK bytes, without executing them. */
/** The judge's audited closure: its runner, rubric and documents, installed SDK bytes included. */
export function workflowJudgeDependencies(root: string, documents: string[]): string[] {
const seen = new Set<string>();
const scan = new Bun.Transpiler({ loader: 'tsx' });
const visit = (file: string) => {
file = path.resolve(file);
const relative = path.relative(root, file).split(path.sep).join('/');
if (relative.startsWith('../') || path.isAbsolute(relative)) throw new Error('Dependency outside checkout');
// Root version labels collector output only; its remaining semantic fields
// are hashed separately. Installed package manifests remain byte-exact.
if (relative === 'package.json') return;
if (seen.has(relative)) return;
seen.add(relative);
const source = fs.readFileSync(file, 'utf8');
if (!/\.[cm]?[jt]sx?$/.test(file)) return;
// Entrypoint scripts carry hashbangs, which scanImports does not accept.
// Strip only for parsing; buildEvalInputIdentity still hashes the full file.
for (const entry of scan.scanImports(source.replace(/^#![^\n]*(?:\n|$)/, '\n'))) {
if (isBuiltin(entry.path) || entry.path.startsWith('bun:')) continue;
const resolved = Bun.resolveSync(entry.path, path.dirname(file));
visit(resolved);
// Package export maps/defaults affect resolution independently of code.
let directory = path.dirname(resolved);
while (directory !== root && directory.startsWith(root + path.sep)) {
const manifest = path.join(directory, 'package.json');
if (fs.existsSync(manifest)) { visit(manifest); break; }
directory = path.dirname(directory);
}
}
};
for (const file of ['test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-cache.ts',
return sourceDependencyClosure(root, ['test/skill-llm-eval.test.ts', 'test/helpers/workflow-judge-cache.ts',
'test/helpers/llm-judge.ts', 'lib/eval-model.ts', 'test/helpers/eval-budgets.ts',
'scripts/test-paid-shards.ts', 'scripts/test-strict-output.ts', 'scripts/eval-select.ts',
'scripts/test-pr-profile.ts', '.github/workflows/evals.yml',
'package.json', 'bun.lock', '.github/docker/Dockerfile.ci', ...documents]) visit(path.join(root, file));
for (const file of ['bunfig.toml', 'tsconfig.json', 'jsconfig.json'])
if (fs.existsSync(path.join(root, file))) visit(path.join(root, file));
return [...seen].sort();
'package.json', 'bun.lock', '.github/docker/Dockerfile.ci', ...documents]);
}
export function validWorkflowJudgeScore(value: EvalCacheValue, thresholds: Thresholds, structuredResponse = false): value is JudgeScore & EvalCacheValue {
@@ -71,17 +40,28 @@ export function validWorkflowJudgeScore(value: EvalCacheValue, thresholds: Thres
|| typeof value.reasoning !== 'string'
|| (structuredResponse && (!value.reasoning.trim()
|| value.reasoning.trim().split(/\s+/).length >= WORKFLOW_JUDGE_REASONING_WORD_LIMIT))) return false;
return (['clarity', 'completeness', 'actionability'] as const).every(key =>
return JUDGE_SCORE_DIMENSIONS.every(key =>
typeof value[key] === 'number' && Number.isInteger(value[key]) && value[key] >= thresholds[key] && value[key] <= 5);
}
const SAMPLE_RANGE: Thresholds = { clarity: 1, completeness: 1, actionability: 1 };
/** A complete judge panel: exactly JUDGE_PANEL_SAMPLES of valid samples whose per-dimension mean meets every threshold. */
export function validWorkflowJudgePanel(value: EvalCacheValue, thresholds: Thresholds, structuredResponse = false): value is { samples: Array<JudgeScore & EvalCacheValue> } {
if (!value || typeof value !== 'object' || Array.isArray(value) || Object.keys(value).join(',') !== 'samples'
|| !Array.isArray(value.samples) || value.samples.length !== JUDGE_PANEL_SAMPLES
|| !value.samples.every(sample => validWorkflowJudgeScore(sample, SAMPLE_RANGE, structuredResponse))) return false;
const mean = judgePanelMean(value.samples as JudgeScore[], JUDGE_SCORE_DIMENSIONS);
return JUDGE_SCORE_DIMENSIONS.every(key => mean[key] >= thresholds[key]);
}
export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): {
lookup(): { scores: JudgeScore; reuse: WorkflowJudgeReuse } | null;
lookup(): { samples: JudgeScore[]; reuse: WorkflowJudgeReuse } | null;
/** The attempt guard is rechecked after synchronous input/provenance reads. */
publish(scores: JudgeScore, isActive?: () => boolean): (() => void) | undefined;
publish(samples: JudgeScore[], isActive?: () => boolean): (() => void) | undefined;
} {
const env = opts.env ?? process.env;
const noCache = { lookup: () => null, publish: (_scores: JudgeScore) => undefined };
const noCache = { lookup: () => null, publish: (_samples: JudgeScore[]) => undefined };
const pr = Number(env.EVALS_CACHE_PR);
// Runtime ID is the immutable CI image manifest, not a mutable image tag.
// Nonstandard Node/Bun preload code or custom model endpoints need a separate
@@ -106,8 +86,10 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): {
files: workflowJudgeDependencies(opts.root, input.files.map(file => file.path)),
prompts: { [opts.testName]: prompt },
parameters: { rootPackage, thresholds: opts.thresholds, max_tokens: opts.maxTokens ?? DEFAULT_JUDGE_MAX_TOKENS, temperature: null, budget_ms: JUDGE_MS,
request: opts.stream ? 'messages.stream/user' : 'messages.create/user', retries: 1,
request: opts.stream ? 'messages.stream/user' : 'messages.create/user', retries: 0,
panel: { samples: JUDGE_PANEL_SAMPLES, numeric: 'mean', boolean: 'majority' },
...(opts.stream ? { stream: true } : {}),
...(opts.effort ? { effort: opts.effort } : {}),
...(opts.structuredResponse ? { output_config: { format: { type: 'json_schema', schema: WORKFLOW_JUDGE_RESPONSE_SCHEMA } },
response_validation: { reasoning_words_below: WORKFLOW_JUDGE_REASONING_WORD_LIMIT } } : {}) },
runtime: { image: env.EVALS_CACHE_RUNTIME_ID!, bun: Bun.version, node: process.versions.node,
@@ -126,14 +108,15 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): {
return {
lookup() {
const result = lookupEvalInputCache({ ...common, identity: before,
validateResult: value => validWorkflowJudgeScore(value, opts.thresholds, opts.structuredResponse) });
validateResult: value => validWorkflowJudgePanel(value, opts.thresholds, opts.structuredResponse) });
return result.status === 'reused'
? { scores: result.result as JudgeScore, reuse: { key: result.key, source: result.source } } : null;
? { samples: (result.result as unknown as { samples: JudgeScore[] }).samples, reuse: { key: result.key, source: result.source } } : null;
},
publish(scores, isActive = () => true) {
publish(samples, isActive = () => true) {
// Caller reaches here ONLY after its actual assertions passed. A later
// failed case in the file does not erase this independently completed case.
if (!isActive() || !validWorkflowJudgeScore(scores as unknown as EvalCacheValue, opts.thresholds, opts.structuredResponse)) return;
const panel = { samples: samples.map(({ clarity, completeness, actionability, reasoning }) => ({ clarity, completeness, actionability, reasoning })) };
if (!isActive() || !validWorkflowJudgePanel(panel as unknown as EvalCacheValue, opts.thresholds, opts.structuredResponse)) return;
const after = currentIdentity();
const runId = env.GITHUB_RUN_ID ? `${env.GITHUB_RUN_ID}/${env.GITHUB_RUN_ATTEMPT ?? '1'}` : env.EVALS_RUN_ID;
if (!after || !runId || !isActive()) return;
@@ -144,7 +127,7 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): {
cancelled: false, skipped: 0, failed: 0, passed: 1,
cases: [{ id: opts.testName, outcome: 'passed', attempt: 1 }],
source: { runId, revision: revision.stdout.trim(), completedAt: Date.now() },
result: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability, reasoning: scores.reasoning },
result: panel,
} });
// A slow synchronous write can consume the recording allowance. The
// caller withdraws this new receipt if its final deadline check fails.
+4 -2
View File
@@ -25,6 +25,8 @@ export const QA_DISCOVERY_REFERENCES = [
];
export const WORKFLOW_JUDGE_REASONING_WORD_LIMIT = 150;
/** The instructed length sits below the enforced limit: judges asked for <150 landed at 130-156 words. */
export const WORKFLOW_JUDGE_REASONING_WORD_TARGET = 120;
export const WORKFLOW_JUDGE_RESPONSE_SCHEMA = {
type: 'object',
@@ -33,7 +35,7 @@ export const WORKFLOW_JUDGE_RESPONSE_SCHEMA = {
completeness: { type: 'integer', enum: [1, 2, 3, 4, 5] },
actionability: { type: 'integer', enum: [1, 2, 3, 4, 5] },
reasoning: { type: 'string',
description: `Under ${WORKFLOW_JUDGE_REASONING_WORD_LIMIT} words with at most two decisive examples, evaluating the complete supplied workflow.` },
description: `Under ${WORKFLOW_JUDGE_REASONING_WORD_TARGET} words with at most two decisive examples, evaluating the complete supplied workflow.` },
},
required: ['clarity', 'completeness', 'actionability', 'reasoning'],
additionalProperties: false,
@@ -61,7 +63,7 @@ Clarity 4 means the target agent can determine the next permitted action on each
5 additionally means those paths are easy to locate and understand.
Score clarity 3 or lower when execution still requires guessing because of
conflicting order, undefined decisions, unclear authority or missing input/output handling.
Evaluate the whole workflow, but keep the JSON reasoning under 150 words with at most two decisive examples.
Evaluate the whole workflow, but keep the JSON reasoning under ${WORKFLOW_JUDGE_REASONING_WORD_TARGET} words with at most two decisive examples.
For a clarity defect, cite the specific file/step and explain the competing actions or missing decision.
Keep completeness and actionability independent: reader capability does not supply missing requirements.` : ''}