ci(image): pin Claude Code 2.1.284, the version users run

Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to
claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284
census was mostly API latency: its SDK-only judges were 25% slower too. Nine
previously slow cases pass on 2.1.284 within unchanged budgets.
This commit is contained in:
garrytan committed 2026-09-30 12:14:37 +00:00
1 parent 77cce3bec4
commit 8cf87d4729
3 files changed
+13 -11

No files matched your search

+8 -6
View File
@@ -93,12 +93,14 @@ RUN curl --retry 5 --retry-delay 5 --retry-connrefused -fsSL https://bun.sh/inst
# skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162). # skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162).
# Bump deliberately, via a PR that runs the PTY gate against the new TUI. # Bump deliberately, via a PR that runs the PTY gate against the new TUI.
# test/ci-image-cli-pin.test.ts fails the free suite if this pin is removed. # test/ci-image-cli-pin.test.ts fails the free suite if this pin is removed.
# Stays on 2.1.251. 2.1.284 (tried 2026-09-29) enables per-turn effort for # 2.1.284 is the first pin that recognizes the eval model claude-fable-5-1.
# claude-fable-5-1: in gate census run 36626737820, 66 of 84 sessions ran # 2.1.251 already sent it effort "high", but ran it as an unknown model with
# longer than the same cases in run 36606688266 on 2.1.251 (session time # a generic system prompt. Users on stable (2.1.280) and latest (2.1.285) get
# +20%, thinking tokens +32%), and 11 cases timed out on unchanged budgets. # the fable-5-1 profile: its own system prompt, 64k max_tokens and per-turn
# Bumping it needs its own budget and skill-speed work. # effort, which only moves the same effort value into the conversation.
RUN npm i -g @anthropic-ai/claude-code@2.1.251 # Census 36626737820 on 2.1.284 looked slower mostly because the API was:
# its SDK-only judge evals, which never start this CLI, were 25% slower too.
RUN npm i -g @anthropic-ai/claude-code@2.1.284
# Playwright system deps (Chromium) — needed for browse E2E tests # Playwright system deps (Chromium) — needed for browse E2E tests
RUN npx playwright install-deps chromium RUN npx playwright install-deps chromium
+1 -1
View File
@@ -41,7 +41,7 @@ Run `bun run typecheck` and `bun run typecheck:test` before you push; both are f
- Paid evals: one test file or case per machine within a 540-second slice budget, planned from recorded per-tier and per-case durations; case sharding for plan, design, review-army, shared-libs, shared-libs-paths, ship-docsync and qa-callers. - Paid evals: one test file or case per machine within a 540-second slice budget, planned from recorded per-tier and per-case durations; case sharding for plan, design, review-army, shared-libs, shared-libs-paths, ship-docsync and qa-callers.
- Verdict policy: no retries; `rule` cases fail on any failed trial, `behavior` cases pass on 2 of 3 parallel trials with contract assertions still strict, `judge` entries average 3 samples against unchanged thresholds. One panel-verdict function feeds the report, PR comment, weekly issue and pass-rate history. A census whose every red is infrastructure is re-dispatched once, and both runs are reported. - Verdict policy: no retries; `rule` cases fail on any failed trial, `behavior` cases pass on 2 of 3 parallel trials with contract assertions still strict, `judge` entries average 3 samples against unchanged thresholds. One panel-verdict function feeds the report, PR comment, weekly issue and pass-rate history. A census whose every red is infrastructure is re-dispatched once, and both runs are reported.
- New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows: the full `/office-hours` workflow (a focused design-draft case replaces it in the weekly lane) and the full `/plan-ceo-review` split-overflow run, which took 8 to 20 minutes on its own. - New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows: the full `/office-hours` workflow (a focused design-draft case replaces it in the weekly lane) and the full `/plan-ceo-review` split-overflow run, which took 8 to 20 minutes on its own.
- The CI image stays on Claude Code 2.1.251. On 2.1.284, sessions ran 20% longer on the same model (per-turn effort) and 11 gate cases timed out on unchanged budgets; that bump needs its own budget work. - The CI image pins Claude Code 2.1.284, the first version that recognizes the eval model `claude-fable-5-1` and runs it with the same profile users get. Both versions send effort "high"; a census that looked slower on 2.1.284 was mostly slower API responses (its SDK-only judges, which never start the CLI, were 25% slower too), and nine previously slow cases pass on 2.1.284 within unchanged budgets.
- `lib/cso/*.ts` is formatted with pinned Prettier; minified transpile output is byte-identical except three canonicalized regex flag orders. - `lib/cso/*.ts` is formatted with pinned Prettier; minified transpile output is byte-identical except three canonicalized regex flag orders.
- The duplicate dispatch-only `ship-docsync` case is removed; `ship-docsync-completion` asserts the same on the same fixture. - The duplicate dispatch-only `ship-docsync` case is removed; `ship-docsync-completion` asserts the same on the same fixture.
+4 -4
View File
@@ -4,10 +4,10 @@
### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29) ### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29)
- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model: - **Thin budgets on slow API days** — on Claude Code 2.1.284, review-army-perf
in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251 (274 of 300 s) and the ship-docsync fault cases (250-263 of 285 s) sit at
(+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged 88-93% of their budgets; a slow-API census can time them out on either CLI
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M. version. Make those skills faster rather than raising budgets. Effort M.
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode` - **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
(one ~250 s thinking block before its single write; times out at 300 s on (one ~250 s thinking block before its single write; times out at 300 s on
2.1.251 in every recent run) and the HOLD SCOPE 2.1.251 in every recent run) and the HOLD SCOPE