mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 01:46:55 +02:00
ci(image): pin Claude Code 2.1.284, the version users run
Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284 census was mostly API latency: its SDK-only judges were 25% slower too. Nine previously slow cases pass on 2.1.284 within unchanged budgets.
This commit is contained in:
1 parent
77cce3bec4
commit
8cf87d4729
3 files changed
+13
-11
No files matched your search
@@ -4,10 +4,10 @@
|
||||
|
||||
### P1: paid-eval follow-ups from the v1.91.10.0 proof censuses (filed 2026-09-29)
|
||||
|
||||
- **Claude Code 2.1.284 bump** — it enables per-turn effort for the eval model:
|
||||
in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251
|
||||
(+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged
|
||||
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M.
|
||||
- **Thin budgets on slow API days** — on Claude Code 2.1.284, review-army-perf
|
||||
(274 of 300 s) and the ship-docsync fault cases (250-263 of 285 s) sit at
|
||||
88-93% of their budgets; a slow-API census can time them out on either CLI
|
||||
version. Make those skills faster rather than raising budgets. Effort M.
|
||||
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
|
||||
(one ~250 s thinking block before its single write; times out at 300 s on
|
||||
2.1.251 in every recent run) and the HOLD SCOPE
|
||||
|
||||
Reference in new issue
Block a user