Merge origin/main (v1.91.9.0, #2998) into capy/audit-fix-wave; release as v1.91.10.0

Main shipped v1.91.9.0, so this wave becomes v1.91.10.0. Conflicts kept this
branch's planner-derived assertions. Main's new test-value eval gets rule
kinds for its three gate cases and its measured 332 s duration from #2998's
CI; census counts and the PR fallback floor follow the new file. The packing
test now weighs each tier's recorded durations the way the planner does.
This commit is contained in:
garrytan committed 2026-09-29 22:00:14 +00:00
commit 58b5e3f977
58 files changed
+2453 -220

No files matched your search

+2 -2
View File
@@ -2,7 +2,7 @@
## NEXT PRIORITY
### P1: paid-eval follow-ups from the v1.91.9.0 proof censuses (filed 2026-09-29)
### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29)
- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model:
in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251
@@ -162,7 +162,7 @@ wave"). Each was explicitly deferred with rationale, not dropped:
- **#2443 AskUserQuestion numbering redesign** — real mismatch (brief letters
vs host-rendered numbers), but a prompt-behavior redesign that shifts eval
baselines; needs its own PR with baseline refresh. Effort S.
- ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.9.0)
- ~~**#2447 typecheck infra**~~ — superseded: the audit fix wave (v1.91.10.0)
added `tsconfig.json`, `bun run typecheck` (zero product errors) and the
`typecheck:test` ratchet inside the required `free-tests` check, reusing
#2447's fixes where they still applied.