From bfad7fd37d9af798f6d74b1047cbe5d5b6c5dba5 Mon Sep 17 00:00:00 2001 From: garrytan Date: Tue, 29 Sep 2026 21:50:57 +0000 Subject: [PATCH] docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups --- CHANGELOG.md | 5 +++-- TODOS.md | 20 +++++++++++++++++++- 2 files changed, 22 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index de23edfbe..a4367e755 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,11 +9,11 @@ The weekly paid eval run took 2 hours 45 minutes on Sept 28, almost all of it on ### The numbers that matter -Source: the Sept 28 weekly census (run 36385945043) and the two proof censuses on this branch (runs 36597762183 and 36606688266). `bun run scripts/test-paid-shards.ts --tier periodic --slice-budget 540 --jobs 2 --list` prints the current plan. +Source: the Sept 28 weekly census (run 36385945043) and the final proof census on this branch (run 36633323521). `bun run scripts/test-paid-shards.ts --tier periodic --slice-budget 540 --jobs 2 --list` prints the current plan. | Measure | Before | After | | --- | ---: | ---: | -| Weekly periodic census wall clock | 2h 45m | 15 min for 32 of 33 machines (proof run 2) | +| Weekly periodic census wall clock | 2h 45m | 11m 42s (gate census alongside: 10m 22s) | | Longest planned slice | 160 min (one test, twice) | ~10 min | | Automatic retries on paid evals | up to 2 per file | 0 | | Product-code type errors | 103 on v1.91.8.0 (no check) | 0, required in `free-tests` | @@ -30,6 +30,7 @@ Run `bun run typecheck` and `bun run typecheck:test` before you push; both are f #### Fixed - `$B connect --supervise` respawned with a block-scoped env that no longer existed, so every restart threw and the supervisor gave up after five tries. The headed env is one helper used by connect and respawn, and the loop has behavioral tests. - Compiled `/cso` installs called an unimported `join` when launching the assertion-witness child, breaking runtime-tested witnessing for every installed user. +- `/qa` checkpoint receipts now print the report link for their `exploration-NNN.json` file; reports had been linking `.qa-evidence/NNN` capture folders as checkpoints instead. - `/review` workflow ambiguities (smoke clock vs required revalidation, setup authority, plan-completion gate, findings record), `/office-hours` builder mode not loading its brainstorm section, `/sync-gbrain` Step 4 helper arguments and write path, `/plan-ceo-review` expansion framing and pacing menus, `/plan-design-review` with no designer API key, and `/deslop-shared-libs` one-file-per-turn reads. - Eval detectors that graded wording or step order now grade outcomes: eng batching, CEO split-overflow, mode routing, section-loading stale-fill, outside-voice-disabled attribution, design focus menus, and PTY permission dialogs with cropped titles. - Harness races and adapter gaps found by the proof runs: plan seeding accepted a stale empty input box when the CLI repainted after recording its reply, the third-party-actions recorder fixture lost every failure record, the autoplan dual-voice check could not read framed subagent reports from newer Claude Code, and the HOLD SCOPE routing check judged the skill's own defer/keep menu as its rigor decision, the outside-disabled check missed a correctly attributed quote of the pre-existing review record, and the plan-review judge was not told its reason length bound on the field it writes. diff --git a/TODOS.md b/TODOS.md index ce8fae2ee..628a1d9b2 100644 --- a/TODOS.md +++ b/TODOS.md @@ -2,6 +2,24 @@ ## NEXT PRIORITY +### P1: paid-eval follow-ups from the v1.91.9.0 proof censuses (filed 2026-09-29) + +- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model: + in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251 + (+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged + budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M. +- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode` + (one ~250 s thinking block before its single write; times out at 300 s on + 2.1.251 in every recent run), `review-army-perf-n-plus-one` (290-300 s on a + 12-line diff; web search plus a conditional red-team pass), the HOLD SCOPE + routing case (no rigor decision within its 240 s window after the skill's own + defer/keep questions), and the `document-release` workflow judge (below its + floor in two of three censuses). Effort M each. +- **Let pass-rate history decide the rest** — every census on this branch had + a different handful of single-trial reds. Once `eval:pass-rates` has 10 weekly + trials per case, apply the CASE_QUARANTINE entry rule instead of chasing one + run at a time. Effort S. + ### P2/P3: impeccable interop deferrals (filed 2026-09-08, from the CEO + eng reviews of docs/designs/IMPECCABLE_INTEROP.md) Each item was weighed during the review and deferred with a reason; none blocks @@ -144,7 +162,7 @@ wave"). Each was explicitly deferred with rationale, not dropped: - **#2443 AskUserQuestion numbering redesign** — real mismatch (brief letters vs host-rendered numbers), but a prompt-behavior redesign that shifts eval baselines; needs its own PR with baseline refresh. Effort S. -- ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.8.0) +- ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.9.0) added `tsconfig.json`, `bun run typecheck` (zero product errors) and the `typecheck:test` ratchet inside the required `free-tests` check, reusing #2447's fixes where they still applied.