fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures

- design-consultation Phase 1 asks one brief that confirms context and decides
  research; the confirm-only first question scored substance 2.
- document-release defines ship-owned inputs, exact steps and the JSON result,
  and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4).
- plan-design-with-ui accepts the Step 0D focus menu the same way the shared
  picker does ("focus on specific ones?").
- plan-design-review plan-mode saves in three Edits instead of one final Write.
- QA functional annotations ask for the full 40-character revision.
- Outside-disabled attribution judges quoted prior-record data by its exact
  timestamp or a dated, pre-existing-record sentence; four captured phrasings
  replay clean and current claims still fail.
- --case can select autoplan-dual-voice by its literal test name.
This commit is contained in:
garrytan committed 2026-09-29 22:36:28 +00:00
1 parent aba80c8fb6
commit a111225e78
17 files changed
+180 -92

No files matched your search

+4 -3
View File
@@ -89,7 +89,7 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
expect(opts.signal.aborted).toBe(false);
// Bind the complete actual compact-delivery prompt, not selected snippets.
expect(new Bun.CryptoHasher('sha256').update(opts.prompt).digest('hex'))
.toBe('2fa957ab9d56850a1629a845d6fe0ee5a1cb7c0843ab6555b621971d270604cb');
.toBe('7f2dbb6b2671588e7fbc83cfecc2869c59c388b45b627ac9d481c4d917d1e76f');
expect(opts.testName).toBe(id); expect(opts.maxTurns).toBe(15); expect(opts.timeout).toBe(CAPTURE_MS);
for (const key of ['model', 'tools', 'allowedTools', 'appendSystemPrompt', 'env']) expect(opts).not.toHaveProperty(key);
expect(opts.prompt).toContain('Review the plan in ./plan.md');
@@ -98,7 +98,8 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
expect(opts.prompt).toContain('preserve the unresolved-decisions pass');
expect(opts.prompt).toContain('interaction state table, empty states, responsive behavior');
expect(opts.prompt).toContain('full required review report');
expect(opts.prompt).toContain('Write before publishing a completed walkthrough');
expect(opts.prompt).toContain('(or one Write) before publishing a completed walkthrough');
expect(opts.prompt).toContain('Save as you go in three Edits: after passes 1-3, apply their decisions to plan.md');
expect(opts.prompt).toContain('Read plan.md back to verify the saved changes');
expect(opts.prompt).toContain('Then return a brief, concrete summary');
expect(opts.prompt).toContain('execute every required pass and lazy-section Read');
@@ -109,7 +110,7 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
expect(opts.prompt).toContain('concise score rationales and 10/10 explanations');
const ordered = ['Read every lazy section', 'Review all 7 design passes',
'EDIT plan.md', 'Keep the saved review compact',
'Persist that complete plan and review with Write', 'Read plan.md back',
'Persist that complete plan and review with those Edits', 'Read plan.md back',
'Then return a brief, concrete summary'].map(text => opts.prompt.indexOf(text));
expect(ordered.every(index => index >= 0)).toBe(true);
expect(ordered).toEqual([...ordered].sort((a, b) => a - b));