docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups

This commit is contained in:
garrytan committed 2026-09-29 21:50:57 +00:00
1 parent 8dd3c44fc6
commit bfad7fd37d
2 files changed
+22 -3

No files matched your search

+3 -2
View File
@@ -9,11 +9,11 @@ The weekly paid eval run took 2 hours 45 minutes on Sept 28, almost all of it on
### The numbers that matter ### The numbers that matter
Source: the Sept 28 weekly census (run 36385945043) and the two proof censuses on this branch (runs 36597762183 and 36606688266). `bun run scripts/test-paid-shards.ts --tier periodic --slice-budget 540 --jobs 2 --list` prints the current plan. Source: the Sept 28 weekly census (run 36385945043) and the final proof census on this branch (run 36633323521). `bun run scripts/test-paid-shards.ts --tier periodic --slice-budget 540 --jobs 2 --list` prints the current plan.
| Measure | Before | After | | Measure | Before | After |
| --- | ---: | ---: | | --- | ---: | ---: |
| Weekly periodic census wall clock | 2h 45m | 15 min for 32 of 33 machines (proof run 2) | | Weekly periodic census wall clock | 2h 45m | 11m 42s (gate census alongside: 10m 22s) |
| Longest planned slice | 160 min (one test, twice) | ~10 min | | Longest planned slice | 160 min (one test, twice) | ~10 min |
| Automatic retries on paid evals | up to 2 per file | 0 | | Automatic retries on paid evals | up to 2 per file | 0 |
| Product-code type errors | 103 on v1.91.8.0 (no check) | 0, required in `free-tests` | | Product-code type errors | 103 on v1.91.8.0 (no check) | 0, required in `free-tests` |
@@ -30,6 +30,7 @@ Run `bun run typecheck` and `bun run typecheck:test` before you push; both are f
#### Fixed #### Fixed
- `$B connect --supervise` respawned with a block-scoped env that no longer existed, so every restart threw and the supervisor gave up after five tries. The headed env is one helper used by connect and respawn, and the loop has behavioral tests. - `$B connect --supervise` respawned with a block-scoped env that no longer existed, so every restart threw and the supervisor gave up after five tries. The headed env is one helper used by connect and respawn, and the loop has behavioral tests.
- Compiled `/cso` installs called an unimported `join` when launching the assertion-witness child, breaking runtime-tested witnessing for every installed user. - Compiled `/cso` installs called an unimported `join` when launching the assertion-witness child, breaking runtime-tested witnessing for every installed user.
- `/qa` checkpoint receipts now print the report link for their `exploration-NNN.json` file; reports had been linking `.qa-evidence/NNN` capture folders as checkpoints instead.
- `/review` workflow ambiguities (smoke clock vs required revalidation, setup authority, plan-completion gate, findings record), `/office-hours` builder mode not loading its brainstorm section, `/sync-gbrain` Step 4 helper arguments and write path, `/plan-ceo-review` expansion framing and pacing menus, `/plan-design-review` with no designer API key, and `/deslop-shared-libs` one-file-per-turn reads. - `/review` workflow ambiguities (smoke clock vs required revalidation, setup authority, plan-completion gate, findings record), `/office-hours` builder mode not loading its brainstorm section, `/sync-gbrain` Step 4 helper arguments and write path, `/plan-ceo-review` expansion framing and pacing menus, `/plan-design-review` with no designer API key, and `/deslop-shared-libs` one-file-per-turn reads.
- Eval detectors that graded wording or step order now grade outcomes: eng batching, CEO split-overflow, mode routing, section-loading stale-fill, outside-voice-disabled attribution, design focus menus, and PTY permission dialogs with cropped titles. - Eval detectors that graded wording or step order now grade outcomes: eng batching, CEO split-overflow, mode routing, section-loading stale-fill, outside-voice-disabled attribution, design focus menus, and PTY permission dialogs with cropped titles.
- Harness races and adapter gaps found by the proof runs: plan seeding accepted a stale empty input box when the CLI repainted after recording its reply, the third-party-actions recorder fixture lost every failure record, the autoplan dual-voice check could not read framed subagent reports from newer Claude Code, and the HOLD SCOPE routing check judged the skill's own defer/keep menu as its rigor decision, the outside-disabled check missed a correctly attributed quote of the pre-existing review record, and the plan-review judge was not told its reason length bound on the field it writes. - Harness races and adapter gaps found by the proof runs: plan seeding accepted a stale empty input box when the CLI repainted after recording its reply, the third-party-actions recorder fixture lost every failure record, the autoplan dual-voice check could not read framed subagent reports from newer Claude Code, and the HOLD SCOPE routing check judged the skill's own defer/keep menu as its rigor decision, the outside-disabled check missed a correctly attributed quote of the pre-existing review record, and the plan-review judge was not told its reason length bound on the field it writes.
+19 -1
View File
@@ -2,6 +2,24 @@
## NEXT PRIORITY ## NEXT PRIORITY
### P1: paid-eval follow-ups from the v1.91.9.0 proof censuses (filed 2026-09-29)
- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model:
in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251
(+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M.
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
(one ~250 s thinking block before its single write; times out at 300 s on
2.1.251 in every recent run), `review-army-perf-n-plus-one` (290-300 s on a
12-line diff; web search plus a conditional red-team pass), the HOLD SCOPE
routing case (no rigor decision within its 240 s window after the skill's own
defer/keep questions), and the `document-release` workflow judge (below its
floor in two of three censuses). Effort M each.
- **Let pass-rate history decide the rest** — every census on this branch had
a different handful of single-trial reds. Once `eval:pass-rates` has 10 weekly
trials per case, apply the CASE_QUARANTINE entry rule instead of chasing one
run at a time. Effort S.
### P2/P3: impeccable interop deferrals (filed 2026-09-08, from the CEO + eng reviews of docs/designs/IMPECCABLE_INTEROP.md) ### P2/P3: impeccable interop deferrals (filed 2026-09-08, from the CEO + eng reviews of docs/designs/IMPECCABLE_INTEROP.md)
Each item was weighed during the review and deferred with a reason; none blocks Each item was weighed during the review and deferred with a reason; none blocks
@@ -144,7 +162,7 @@ wave"). Each was explicitly deferred with rationale, not dropped:
- **#2443 AskUserQuestion numbering redesign** — real mismatch (brief letters - **#2443 AskUserQuestion numbering redesign** — real mismatch (brief letters
vs host-rendered numbers), but a prompt-behavior redesign that shifts eval vs host-rendered numbers), but a prompt-behavior redesign that shifts eval
baselines; needs its own PR with baseline refresh. Effort S. baselines; needs its own PR with baseline refresh. Effort S.
- ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.8.0) - ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.9.0)
added `tsconfig.json`, `bun run typecheck` (zero product errors) and the added `tsconfig.json`, `bun run typecheck` (zero product errors) and the
`typecheck:test` ratchet inside the required `free-tests` check, reusing `typecheck:test` ratchet inside the required `free-tests` check, reusing
#2447's fixes where they still applied. #2447's fixes where they still applied.