mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-31 10:20:42 +02:00
skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s) reported PASS on every periodic run while asserting nothing. Deleting them would remove the periodic-tier selector surface they exist to register (diff-based selection for spec/ changes), so they become test.todo — reported as todo/skip, never pass — with the v1.1 implementation specs kept in-file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
36 lines
1.5 KiB
TypeScript
36 lines
1.5 KiB
TypeScript
/**
|
|
* /spec LLM-judge eval (periodic, paid).
|
|
*
|
|
* Asserts: when /spec runs against a fixture vague request, the agent
|
|
* produces a spec body that scores >= 8/10 against an LLM judge using
|
|
* the contributor's 14 Quality Standards as the rubric.
|
|
*
|
|
* Cost: ~$0.15/run. Periodic — runs weekly via cron or on demand via
|
|
* `EVALS=1 EVALS_TIER=periodic bun run test:evals`.
|
|
*
|
|
* TODO (v1.1): expand fixture set to cover bug / feature / refactor / audit
|
|
* framings + project-level prompts (no concrete file mapping, exercises the
|
|
* Phase 3 fallback path).
|
|
*/
|
|
|
|
import { describe, test } from 'bun:test';
|
|
|
|
const evalsEnabled = !!process.env.EVALS;
|
|
const describeEval = evalsEnabled ? describe : describe.skip;
|
|
|
|
describeEval('/spec LLM-judge eval (periodic)', () => {
|
|
// test.todo, not expect(true): the placeholder reported PASS on every
|
|
// run while asserting nothing — a lying green with a 300s budget. The
|
|
// file stays as the periodic-tier selector surface for spec/ changes.
|
|
//
|
|
// Expected v1.1 implementation:
|
|
// 1. Pick fixture prompt from test/fixtures/spec/vague-bug.md
|
|
// 2. Spawn `claude -p` with /spec loaded, send the prompt + role-play
|
|
// five Phase 1 answers (from test/fixtures/spec/vague-bug-answers.json)
|
|
// 3. Capture final spec body
|
|
// 4. Dispatch to Claude judge with prompt encoding the 14 Quality
|
|
// Standards from spec/SKILL.md.tmpl
|
|
// 5. Assert numeric score >= 8
|
|
test.todo('spec body scores >= 8/10 against 14-standard rubric on fixture request');
|
|
});
|