ci(image): pin Claude Code 2.1.284, the version users run

Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to
claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284
census was mostly API latency: its SDK-only judges were 25% slower too. Nine
previously slow cases pass on 2.1.284 within unchanged budgets.
This commit is contained in:
garrytan committed 2026-09-30 12:14:37 +00:00
1 parent 77cce3bec4
commit 8cf87d4729
3 files changed
+13 -11

No files matched your search

+8 -6
View File
@@ -93,12 +93,14 @@ RUN curl --retry 5 --retry-delay 5 --retry-connrefused -fsSL https://bun.sh/inst
# skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162).
# Bump deliberately, via a PR that runs the PTY gate against the new TUI.
# test/ci-image-cli-pin.test.ts fails the free suite if this pin is removed.
# Stays on 2.1.251. 2.1.284 (tried 2026-09-29) enables per-turn effort for
# claude-fable-5-1: in gate census run 36626737820, 66 of 84 sessions ran
# longer than the same cases in run 36606688266 on 2.1.251 (session time
# +20%, thinking tokens +32%), and 11 cases timed out on unchanged budgets.
# Bumping it needs its own budget and skill-speed work.
RUN npm i -g @anthropic-ai/claude-code@2.1.251
# 2.1.284 is the first pin that recognizes the eval model claude-fable-5-1.
# 2.1.251 already sent it effort "high", but ran it as an unknown model with
# a generic system prompt. Users on stable (2.1.280) and latest (2.1.285) get
# the fable-5-1 profile: its own system prompt, 64k max_tokens and per-turn
# effort, which only moves the same effort value into the conversation.
# Census 36626737820 on 2.1.284 looked slower mostly because the API was:
# its SDK-only judge evals, which never start this CLI, were 25% slower too.
RUN npm i -g @anthropic-ai/claude-code@2.1.284
# Playwright system deps (Chromium) — needed for browse E2E tests
RUN npx playwright install-deps chromium