test: run seven paid evals on the current default capture model (B8)

skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned
claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned
claude-sonnet-4-6; none tests a historical model, so they now capture with
resolveEvalModel('capture'), and the free harness tests that execute these
registrations receive the same resolver. The paid re-pin run passed all of
them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep
claude-opus-4-7: six of their cases failed on the default model (three
timeouts, a missing report file, a format miss and a posture score of 3), so
per the plan's fallback they keep their pins with a TODOS entry. The pre-spend
estimate and drop threshold are in docs/test-audit-2026-09.md.
This commit is contained in:
garrytan committed 2026-09-29 09:16:19 +00:00
1 parent 799f6ec36c
commit 67084a5caf
11 files changed
+55 -23

No files matched your search

+3 -1
View File
@@ -1,5 +1,6 @@
/** Free recording fixtures; every runner and judge below is synthetic. */
import { describe, expect, spyOn, test } from 'bun:test';
import { resolveEvalModel } from '../lib/eval-model';
import * as fs from 'node:fs';
import * as os from 'node:os';
import * as path from 'node:path';
@@ -649,7 +650,8 @@ describe('Plan format actual capture and judge lifecycle', () => {
expect(output).not.toContain('Unhandled error between tests');
const starts = events.filter(event => event.kind === 'start');
expect(starts.map(({ timeout, maxTurns, model }) => ({ timeout, maxTurns, model })))
.toEqual([1, 2].map(() => ({ timeout: 300, maxTurns: 10, model: 'claude-opus-4-7' })));
.toEqual([1, 2].map(() => ({ timeout: 300, maxTurns: 10,
model: file.includes('plan-prosons') ? 'claude-opus-4-7' : resolveEvalModel('capture') })));
expect(events.filter(event => event.kind === 'ready').map(event => event.fixtureExists)).toEqual([true, true]);
expect(entries).toHaveLength(2);
expect(entries.map(entry => entry.attempt)).toEqual([1, 2]);