test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon

Census 36629958451 reds:
- outside-plan-disabled-no-fallback: the model quoted the pre-existing record
  as a parenthesized field list with its exact timestamp; attribution now
  requires that exact timestamp and the record's own field values.
- plan-devex-peer-comparison-classification: the judge correctly returned
  missing but wrote a 1069-character reason, voiding the judgment; structured
  outputs cannot enforce maxLength, so the bound is stated on the field.
- plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the
  periodic lane's wall clock; it now runs weekly in the marathon lane.
This commit is contained in:
garrytan committed 2026-09-29 21:26:37 +00:00
1 parent f63e1fb7cb
commit bf4667146c
6 files changed
+52 -8

No files matched your search

+3 -1
View File
@@ -32,11 +32,13 @@ Run `bun run typecheck` and `bun run typecheck:test` before you push; both are f
- Compiled `/cso` installs called an unimported `join` when launching the assertion-witness child, breaking runtime-tested witnessing for every installed user.
- `/review` workflow ambiguities (smoke clock vs required revalidation, setup authority, plan-completion gate, findings record), `/office-hours` builder mode not loading its brainstorm section, `/sync-gbrain` Step 4 helper arguments and write path, `/plan-ceo-review` expansion framing and pacing menus, `/plan-design-review` with no designer API key, and `/deslop-shared-libs` one-file-per-turn reads.
- Eval detectors that graded wording or step order now grade outcomes: eng batching, CEO split-overflow, mode routing, section-loading stale-fill, outside-voice-disabled attribution, design focus menus, and PTY permission dialogs with cropped titles.
- Harness races and adapter gaps found by the proof runs: plan seeding accepted a stale empty input box when the CLI repainted after recording its reply, the third-party-actions recorder fixture lost every failure record, the autoplan dual-voice check could not read framed subagent reports from newer Claude Code, and the HOLD SCOPE routing check judged the skill's own defer/keep menu as its rigor decision, the outside-disabled check missed a correctly attributed quote of the pre-existing review record, and the plan-review judge was not told its reason length bound on the field it writes.
#### Changed
- Paid evals: one test file or case per machine within a 540-second slice budget, planned from recorded per-tier and per-case durations; case sharding for plan, design, review-army, shared-libs, shared-libs-paths, ship-docsync and qa-callers.
- Verdict policy: no retries; `rule` cases fail on any failed trial, `behavior` cases pass on 2 of 3 parallel trials with contract assertions still strict, `judge` entries average 3 samples against unchanged thresholds. One panel-verdict function feeds the report, PR comment, weekly issue and pass-rate history. A census whose every red is infrastructure is re-dispatched once, and both runs are reported.
- New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows (the full `/office-hours` workflow; a focused design-draft case replaces it in the weekly lane).
- New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows: the full `/office-hours` workflow (a focused design-draft case replaces it in the weekly lane) and the full `/plan-ceo-review` split-overflow run, which took 8 to 20 minutes on its own.
- The CI image stays on Claude Code 2.1.251. On 2.1.284, sessions ran 20% longer on the same model (per-turn effort) and 11 gate cases timed out on unchanged budgets; that bump needs its own budget work.
- `lib/cso/*.ts` is formatted with pinned Prettier; minified transpile output is byte-identical except three canonicalized regex flag orders.
- The duplicate dispatch-only `ship-docsync` case is removed; `ship-docsync-completion` asserts the same on the same fixture.