test: run seven paid evals on the current default capture model (B8)

skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned
claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned
claude-sonnet-4-6; none tests a historical model, so they now capture with
resolveEvalModel('capture'), and the free harness tests that execute these
registrations receive the same resolver. The paid re-pin run passed all of
them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep
claude-opus-4-7: six of their cases failed on the default model (three
timeouts, a missing report file, a format miss and a posture score of 3), so
per the plan's fallback they keep their pins with a TODOS entry. The pre-spend
estimate and drop threshold are in docs/test-audit-2026-09.md.
This commit is contained in:
garrytan committed 2026-09-29 09:16:19 +00:00
1 parent 799f6ec36c
commit 67084a5caf
11 files changed
+55 -23

No files matched your search

@@ -35,6 +35,7 @@
*/
import { expect, beforeAll, afterAll } from 'bun:test';
import { resolveEvalModel } from '../lib/eval-model';
import { CAPTURE_LONG_MS } from './helpers/eval-budgets';
import { execFileSync, spawnSync } from 'child_process';
import {
@@ -228,7 +229,7 @@ exit 0
collector: evalCollector,
name: '/office-hours-brain-writeback',
suite: 'Office Hours Brain Writeback E2E',
model: 'claude-sonnet-4-6',
model: resolveEvalModel('capture'),
run: (signal) => runSkillTest({
signal,
prompt: `Read office-hours/SKILL.md for the workflow.
@@ -245,7 +246,7 @@ This is a test of the brain-writeback path. Do NOT skip the gbrain save step und
timeout: CAPTURE_LONG_MS,
testName: 'office-hours-brain-writeback',
runId,
model: 'claude-sonnet-4-6',
model: resolveEvalModel('capture'),
env: childEnv,
}),
validate: (result) => {