- ceo mode routing: a Submit review taller than the viewport, a setup tab
bundled after the mode tab, and a clip through the mode question each hung
or misread the run; the native answer is still verified after Submit.
- judgeRecommendation requests a 1-5 enum schema; a malformed Haiku reply had
scored substance 0 for a 4/5 brief. Judge failures now propagate.
- carve section-loading for design-consultation declines the optional outside
voices (a supported path) and treats DESIGN.md as the report; timeout unchanged.
The Step 0E handoff defect is not fixed (0/15 samples across four wordings,
none shipped) and is filed in TODOS.
- plan-design-review-plan-mode was registered by two files; the PTY smoke is
now plan-design-review-plan-mode-smoke, and a registry test requires one
owner per case in case-sharded files.
- plan-devex-finding-floor: the template's 0B narrative-confirmation question
is classified as setup structurally instead of timing out a Haiku assessor.
- setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration
question the test asserts; it now accepts that question and declines others.
- design-review-plugin-handoff: the fake engine cited a file absent from the
fixture repo and index.html linked a missing styles.css.
Captured-question regressions with negative controls; each case passed a
focused paid run.
Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to
claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284
census was mostly API latency: its SDK-only judges were 25% slower too. Nine
previously slow cases pass on 2.1.284 within unchanged budgets.
- ship-docsync-completion: yesterday's audit-scope result dropped the section's
status, so /ship spliced one in; the section now opens with **Status:**.
- ship-docsync-missing-asset: a missing section or old Ship-owned mode blocks
before launch.
- ship-docsync-late-result: the invocation record says prepare already saves
the candidate selection (no extra Read; budget unchanged).
- qa exploratory: await the method Reads before the first probe.
- qa-callers fixture: quote the real review-log record template; allow the
git log command plan-completion prescribes.
- qa functional observer: a receipt caught mid-link(2) is checked at stop
instead of failing with ENOENT (reproduced from CI).
Each repaired case passed a focused paid run.
- /ship design-lite: the probe is mandatory and any non-ready first line is
stated, matching /review (5 of 6 captured /review trials had skipped it).
- shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error,
so the one refetch is a legitimate recovery, charged to the same budget.
- shared-libs-review-prior-coverage: the Skip option said a future review can
"reuse it once snapshot coverage holds"; a conditional tail on the recorded
decision is not product work. Captured-text regressions and negative controls.
- review-army-perf-n-plus-one: the parent copied full checklists into agent
prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
without probing; the probe is mandatory and its first line is reported, and
the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
invocation or status vocabulary; the model ran it through bun and wrote
status "blocked". The prompt states both and the validator rejects
out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.
auto-decide-preserved: the product auto-decided HOLD SCOPE and said
"Decision: review mode = HOLD SCOPE"; the grammar knew only "is" and ":".
shared-libs-plan-callers: the recommended option said "(no hardening)" and the
actor read "hardening" as an expansion. Both replay the captured text, keep
negative controls, and passed focused paid runs.
A focused paid run still timed out at 300 s: the first three passes alone took
150 s of thinking. The case stays a named timeout red rather than cutting review depth.
- design-consultation Phase 1 asks one brief that confirms context and decides
research; the confirm-only first question scored substance 2.
- document-release defines ship-owned inputs, exact steps and the JSON result,
and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4).
- plan-design-with-ui accepts the Step 0D focus menu the same way the shared
picker does ("focus on specific ones?").
- plan-design-review plan-mode saves in three Edits instead of one final Write.
- QA functional annotations ask for the full 40-character revision.
- Outside-disabled attribution judges quoted prior-record data by its exact
timestamp or a dated, pre-existing-record sentence; four captured phrasings
replay clean and current claims still fail.
- --case can select autoplan-dual-voice by its literal test name.
The merged skeleton measured 80,166 bytes against its unchanged 80,150 cap.
Same instructions: ask separately for each addition, in turn, with no pacing
menu; lead each proposal with the felt experience, then shape, effort and impact.
Main shipped v1.91.9.0, so this wave becomes v1.91.10.0. Conflicts kept this
branch's planner-derived assertions. Main's new test-value eval gets rule
kinds for its three gate cases and its measured 332 s duration from #2998's
CI; census counts and the PR fallback floor follow the new file. The packing
test now weighs each tier's recorded durations the way the planner does.
qa-functional-webhook-report failed in two of three censuses because the
report linked .qa-evidence/NNN capture folders as "checkpoints" and never
linked exploration-NNN.json. The checkpoint receipt now prints
link: [checkpoint NNN](exploration-NNN.json), and the functional report
template says capture folders are not checkpoints.
The registered-callback fixtures mock ceo-mode-option and lacked the new
export; the split fixtures asserted the periodic tier; the registered-budget
check looked for split-overflow only in the periodic manifest.
Census 36629958451 reds:
- outside-plan-disabled-no-fallback: the model quoted the pre-existing record
as a parenthesized field list with its exact timestamp; attribution now
requires that exact timestamp and the record's own field values.
- plan-devex-peer-comparison-classification: the judge correctly returned
missing but wrote a 1069-character reason, voiding the judgment; structured
outputs cannot enforce maxLength, so the bound is stated on the field.
- plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the
periodic lane's wall clock; it now runs weekly in the marathon lane.
2.1.284 enables per-turn effort for claude-fable-5-1: in gate census
36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time,
+32% thinking tokens) and 11 cases timed out on unchanged budgets.
HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it
Defer and the assessment then judged that scope question as the rigor
decision. The actor now answers that menu Keep and assesses the next one.
Census run 36626737820: the native CEO report arrived framed and indented, so
its INPUT line never matched, and the model's exact probe plus two variable
echoes was not canonical. A column-zero line inside a frame, command
substitution, backticks, redirects, assignments, CODEX_MODE echoes and output
line-count mismatches stay rejected. The failure now records before asserting.
submitPlanSeed accepted a stale empty composer when the transcript recorded
end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the
old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E
without isPreTurnInfraFailure, so every failed case threw before recording.
The next targeted rerun (Claude Code 2.1.284) again asked eleven separate
native questions and again counted zero: its briefs named no plan and its
report declared '- **Review target (fixed):** `/abs/PLAN.md`' under
'# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one
current target field that names a PLAN.md file, whatever its list or
emphasis markup; its ledger record still supplies the cited finding and
must reproduce the brief exactly. A brief that names its plan must still
match the report title. Replays of all three captures count 9, 9 and 3;
controls reject a foreign, duplicate or missing target and an archived
title.
recordE2E sets failure_class 'infra' on a failed session whose runner
reports error_api, timeout_startup, error_output_stream or a non-zero CLI
exit with zero turns and no assistant event. A model refusal, a timeout
after model work, max turns, or an explicit caller pass/class keeps its
ordinary classification.
llm-judge-recommendation is a judge case: each fixture now draws a
3-sample judgePanel, gates reason_substance on the panel mean and the
present/commits/has_because checks on a 2-of-3 majority, thresholds
unchanged. armJudge no longer re-asks on a malformed verdict; it is a
failed sample, as the judge policy requires.
shared-libs-opportunity-judgment and review-design-lite are behavior
cases: their recommendation and checklist judgments may vary, but the
read-only invariant (commands, provider requests, fixture bytes, hooks,
state) and the deterministic fake-engine detector rows are contracts.
Both now go through expectContract, so any failure vetoes the panel.
The targeted batching rerun on Claude Code 2.1.284 left its first report
Write unanswered for 1,372 s and timed out: the viewport began at the
pane's relative file row and rule, with the 'Create file' title cropped
above, so the preview parser rejected the file row as foreign. That row
must now resolve to the owned path and is skipped before the unchanged
line-by-line preview match. Replay controls reject another file, another
directory and an edited preview row.
bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N
independent trials of one case through the CI panel runner (trial shards,
TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict();
N defaults to the case's policy panel and CI never reads it. The local
sharded path (test:gate:sharded, test:periodic:sharded) now plans the same
trial shards and exclusions as CI and exits on execution completeness plus
panel verdicts.
The planner job restores this PR's receipt store once and ships a single
filtered set with the plan: a pass or panel receipt with a same-or-newer
FAIL for its input identity is dropped, and a panel receipt ships only as
a whole PASS panel (re-verified with panelVerdict) from one run. Executors
read only that set (no per-slice cache restore or save), so every trial of
a panel sees the same receipts; a trial reuses its own record from the
panel receipt, keeping a split PASS's failed trial.
Trial identities drop the trial index (run-scoped) and bind the panel
policy. Executed shards carry their input identity; the report turns a
whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule
shard into a negative receipt, and marks a panel that mixes reused and
fresh trials INCOMPLETE. The report job merges plan, slice and report
receipts (newest per file) and saves one store per run.
Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts
whose diagnostic text drifted with program order (baseline locked, fix only).
- Slice, census and marathon artifacts carry -a<run_attempt>; reports
download them per artifact (no merge), so records never overwrite and a
re-run never replaces the first attempt's verdict.
- Planners pass --max-parallel for the capacity preflight (24/16 unchanged:
the refreshed periodic plan needs 24 slices, the gate census 12).
- PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized
failure block); the group_by(.name)|last recomputation is gone.
- Reports stamp series identities, upload trial-outcomes-* for history, and
shard logs upload always (a failed trial no longer reds its runner).
- Weekly report: headline + failure block of both lanes in the issue body,
the eval:pass-rates --gate step (fails closed without history), close the
issue on a green run, and UC-E1: when every red is machine-classified
INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency
group (redispatch_of), both runs reported.
- scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's
caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step,
keeping the history tool out of the paid runner's closure;
TrialOutcomeRecord gains the optional series_identity field.
- Slice-count plans let a registered trial spill into an ordinary lane when
its siblings hold every long lane, so panels never share a runner.
- Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY
map-diff; no new module loading) and repinned its hash.
- Detach and release floors now count trial shards (66 periodic trials in
22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s.
- Coordination fixtures supply the executor's trial records.
Both tiers, merged in run order (the later run wins). Notable: split-overflow
1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s;
multi-finding-batching 734s -> 1318s (its red path in run 36606688266).
Planner: behavior and quarantined cases become panels of isolated trial
shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact
test name; the file shard excludes them by name. Trials of one case never
share a slice, result slugs are unique, panels are validated whole, unknown
registrations throw, and the planner prints a capacity preflight.
Executor: each trial shard gets its TRIAL_ENV identity and a trial record
(outcome, failure class, cause, cost); every shard writes a JUnit report.
The slice exit now means execution completeness: a failed rule shard or a
trial without a record reds the runner, a failed trial does not.
Report: panelVerdict() decides every panel of the first run attempt (later
attempts are reported, never replacing it); rule shards keep the unchanged
fail-closed checks; collector records all count (no last-attempt wins);
census runs enforce the quarantine cap and expiry. It writes
collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge
cases), report-summary.md, and one headline + failure block with rerun
commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch.
The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green
with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red,
quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later
attempt never replacing the first.
AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.
Run 36606688266 bundled routing, learnings and the mode choice into one
native call. Its review panel was taller than the terminal, so the tab
bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD
SCOPE was never submitted ('no posture match'). With no bar on screen the
viewport must still end at the focused Submit prompt, and the accumulated
screen text supplies the one complete panel, authenticated exactly as
before. Replay controls reject another mode, an unoffered answer, an
altered question, a quoted panel, trailing output, a moved cursor and an
answered or changed call.
Run 36606688266's opportunity audit read sixteen sources one per turn and
stopped at error_max_turns; the passing run 36597762183 read the same
files in three batched commands. The skill now says turns are bounded and
asks for parallel reads or one read-only command per step.