Commit Graph
2 Commits
Author SHA1 Message Date
Garry TanandClaude Fable 5 c9e9a653cf feat(evals): arm benchmark runs each fixture's functional oracle — correctness before LOC
The plan's metric order is diff-quality FIRST, but cells never ran the
fixtures' own run-tests.js, so a refusal, a broken implementation, and
working code were indistinguishable in aggregates (Codex adversarial catch).
Tasks with an oracle declare checkCmd; every cell records checks=pass|fail|none
in the report line and eval store. Selftest pins the oracle declarations and
that the planted bug fails its own check pre-fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-29 05:57:56 +00:00
Garry TanandClaude Fable 5 2ee61ab455 refactor(evals): arm-benchmark selftest runs FREE on every PR
The selftest lived inside the paid skill-e2e-* file, so fixture-integrity
and plumbing pins executed weekly at best — a broken fixture would ship past
every gating check and be discovered when the periodic run burned money on a
dead instrument. Harness extracted to test/helpers/arm-benchmark-harness.ts,
selftest to test/arm-benchmark-selftest.test.ts (free suite). Touchfiles:
harness added to the three benchmark dep lists; the auq-repetition-cut-ab
tier comment now states the MANUAL re-run obligation honestly (periodic runs
force EVALS_ALL, so dep lists cannot auto-trigger it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-29 05:42:57 +00:00