mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
test: run seven paid evals on the current default capture model (B8)
skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned
claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned
claude-sonnet-4-6; none tests a historical model, so they now capture with
resolveEvalModel('capture'), and the free harness tests that execute these
registrations receive the same resolver. The paid re-pin run passed all of
them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep
claude-opus-4-7: six of their cases failed on the default model (three
timeouts, a missing report file, a format miss and a posture score of 3), so
per the plan's fallback they keep their pins with a TODOS entry. The pre-spend
estimate and drop threshold are in docs/test-audit-2026-09.md.
This commit is contained in:
1 parent
799f6ec36c
commit
67084a5caf
11 files changed
+55
-23
No files matched your search
@@ -0,0 +1,23 @@
|
||||
# Test audit 2026-09: evidence
|
||||
|
||||
Evidence for the test-reduction branch (plan approved via /autoplan). Sections are added by the commit they support.
|
||||
|
||||
## B8 pre-spend estimate (recorded 2026-09-29, before any B8 paid run)
|
||||
|
||||
Source: latest weekly periodic artifacts (runs 36385945043 = 09-28, 35567915613 = 09-21), per-shard eval JSON cost_usd.
|
||||
Price ratio from test/helpers/pricing.ts: claude-fable-5-1 (default capture, lib/eval-model.ts) $10/$50 per MTok in/out;
|
||||
claude-opus-4-7 $15/$75 (ratio 0.667 on both); claude-sonnet-4-6 $3/$15 (ratio 3.33 on both).
|
||||
|
||||
| Files | Old pin | Weekly $ (09-28) | Est. weekly $ on default | Delta |
|
||||
|---|---|---:|---:|---:|
|
||||
| plan, design, plan-prosons, plan-format, qa-bugs, retro, office-hours-phase4 | opus-4-7 | 15.78 | 10.52 | −5.26 |
|
||||
| office-hours, office-hours-brain-writeback | sonnet-4-6 | 0.91 | 3.03 | +2.12 |
|
||||
| auq-matrix, workflow | opus-4-7 | no result in the retained artifacts | — | ≤ 0 (ratio 0.667) |
|
||||
| **B8 total** | | 16.69 | 13.55 | **−3.14** |
|
||||
|
||||
Assumes the same token volume per case (a verbosity change moves this; the ratio applies to input and output alike).
|
||||
Wall clock: unchanged shard walls (budgets do not depend on model). Drop threshold, fixed now: B8 is dropped from this PR
|
||||
if its estimated net weekly dollars after C and B5 savings are above zero. Estimated net: −3.14 (B8) − C savings
|
||||
(five retired evals) − B5 savings (18 hollow shards, 23 census judges) < 0 → B8 proceeds to its one paid run.
|
||||
Fallback check: `git log -S claude-sonnet-4-6` on skill-e2e-office-hours and -brain-writeback shows only 636175d / #2264
|
||||
(infra hardening), no cost rationale → both re-pinned.
|
||||
Reference in new issue
Block a user