mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 01:46:55 +02:00
ci(image): pin Claude Code 2.1.284, the version users run
Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284 census was mostly API latency: its SDK-only judges were 25% slower too. Nine previously slow cases pass on 2.1.284 within unchanged budgets.
This commit is contained in:
1 parent
77cce3bec4
commit
8cf87d4729
3 files changed
+13
-11
No files matched your search
@@ -93,12 +93,14 @@ RUN curl --retry 5 --retry-delay 5 --retry-connrefused -fsSL https://bun.sh/inst
|
|||||||
# skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162).
|
# skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162).
|
||||||
# Bump deliberately, via a PR that runs the PTY gate against the new TUI.
|
# Bump deliberately, via a PR that runs the PTY gate against the new TUI.
|
||||||
# test/ci-image-cli-pin.test.ts fails the free suite if this pin is removed.
|
# test/ci-image-cli-pin.test.ts fails the free suite if this pin is removed.
|
||||||
# Stays on 2.1.251. 2.1.284 (tried 2026-09-29) enables per-turn effort for
|
# 2.1.284 is the first pin that recognizes the eval model claude-fable-5-1.
|
||||||
# claude-fable-5-1: in gate census run 36626737820, 66 of 84 sessions ran
|
# 2.1.251 already sent it effort "high", but ran it as an unknown model with
|
||||||
# longer than the same cases in run 36606688266 on 2.1.251 (session time
|
# a generic system prompt. Users on stable (2.1.280) and latest (2.1.285) get
|
||||||
# +20%, thinking tokens +32%), and 11 cases timed out on unchanged budgets.
|
# the fable-5-1 profile: its own system prompt, 64k max_tokens and per-turn
|
||||||
# Bumping it needs its own budget and skill-speed work.
|
# effort, which only moves the same effort value into the conversation.
|
||||||
RUN npm i -g @anthropic-ai/claude-code@2.1.251
|
# Census 36626737820 on 2.1.284 looked slower mostly because the API was:
|
||||||
|
# its SDK-only judge evals, which never start this CLI, were 25% slower too.
|
||||||
|
RUN npm i -g @anthropic-ai/claude-code@2.1.284
|
||||||
|
|
||||||
# Playwright system deps (Chromium) — needed for browse E2E tests
|
# Playwright system deps (Chromium) — needed for browse E2E tests
|
||||||
RUN npx playwright install-deps chromium
|
RUN npx playwright install-deps chromium
|
||||||
|
|||||||
+1
-1
@@ -41,7 +41,7 @@ Run `bun run typecheck` and `bun run typecheck:test` before you push; both are f
|
|||||||
- Paid evals: one test file or case per machine within a 540-second slice budget, planned from recorded per-tier and per-case durations; case sharding for plan, design, review-army, shared-libs, shared-libs-paths, ship-docsync and qa-callers.
|
- Paid evals: one test file or case per machine within a 540-second slice budget, planned from recorded per-tier and per-case durations; case sharding for plan, design, review-army, shared-libs, shared-libs-paths, ship-docsync and qa-callers.
|
||||||
- Verdict policy: no retries; `rule` cases fail on any failed trial, `behavior` cases pass on 2 of 3 parallel trials with contract assertions still strict, `judge` entries average 3 samples against unchanged thresholds. One panel-verdict function feeds the report, PR comment, weekly issue and pass-rate history. A census whose every red is infrastructure is re-dispatched once, and both runs are reported.
|
- Verdict policy: no retries; `rule` cases fail on any failed trial, `behavior` cases pass on 2 of 3 parallel trials with contract assertions still strict, `judge` entries average 3 samples against unchanged thresholds. One panel-verdict function feeds the report, PR comment, weekly issue and pass-rate history. A census whose every red is infrastructure is re-dispatched once, and both runs are reported.
|
||||||
- New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows: the full `/office-hours` workflow (a focused design-draft case replaces it in the weekly lane) and the full `/plan-ceo-review` split-overflow run, which took 8 to 20 minutes on its own.
|
- New non-blocking weekly `evals-marathon.yml` lane for full start-to-finish flows: the full `/office-hours` workflow (a focused design-draft case replaces it in the weekly lane) and the full `/plan-ceo-review` split-overflow run, which took 8 to 20 minutes on its own.
|
||||||
- The CI image stays on Claude Code 2.1.251. On 2.1.284, sessions ran 20% longer on the same model (per-turn effort) and 11 gate cases timed out on unchanged budgets; that bump needs its own budget work.
|
- The CI image pins Claude Code 2.1.284, the first version that recognizes the eval model `claude-fable-5-1` and runs it with the same profile users get. Both versions send effort "high"; a census that looked slower on 2.1.284 was mostly slower API responses (its SDK-only judges, which never start the CLI, were 25% slower too), and nine previously slow cases pass on 2.1.284 within unchanged budgets.
|
||||||
- `lib/cso/*.ts` is formatted with pinned Prettier; minified transpile output is byte-identical except three canonicalized regex flag orders.
|
- `lib/cso/*.ts` is formatted with pinned Prettier; minified transpile output is byte-identical except three canonicalized regex flag orders.
|
||||||
- The duplicate dispatch-only `ship-docsync` case is removed; `ship-docsync-completion` asserts the same on the same fixture.
|
- The duplicate dispatch-only `ship-docsync` case is removed; `ship-docsync-completion` asserts the same on the same fixture.
|
||||||
|
|
||||||
|
|||||||
@@ -4,10 +4,10 @@
|
|||||||
|
|
||||||
### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29)
|
### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29)
|
||||||
|
|
||||||
- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model:
|
- **Thin budgets on slow API days** — on Claude Code 2.1.284, review-army-perf
|
||||||
in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251
|
(274 of 300 s) and the ship-docsync fault cases (250-263 of 285 s) sit at
|
||||||
(+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged
|
88-93% of their budgets; a slow-API census can time them out on either CLI
|
||||||
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M.
|
version. Make those skills faster rather than raising budgets. Effort M.
|
||||||
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
|
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
|
||||||
(one ~250 s thinking block before its single write; times out at 300 s on
|
(one ~250 s thinking block before its single write; times out at 300 s on
|
||||||
2.1.251 in every recent run) and the HOLD SCOPE
|
2.1.251 in every recent run) and the HOLD SCOPE
|
||||||
|
|||||||
Reference in new issue
Block a user