feat(evals): trial planner, slice exit split and panel-verdict report

Planner: behavior and quarantined cases become panels of isolated trial
shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact
test name; the file shard excludes them by name. Trials of one case never
share a slice, result slugs are unique, panels are validated whole, unknown
registrations throw, and the planner prints a capacity preflight.

Executor: each trial shard gets its TRIAL_ENV identity and a trial record
(outcome, failure class, cause, cost); every shard writes a JUnit report.
The slice exit now means execution completeness: a failed rule shard or a
trial without a record reds the runner, a failed trial does not.

Report: panelVerdict() decides every panel of the first run attempt (later
attempts are reported, never replacing it); rule shards keep the unchanged
fail-closed checks; collector records all count (no last-attempt wins);
census runs enforce the quarantine cap and expiry. It writes
collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge
cases), report-summary.md, and one headline + failure block with rerun
commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch.

The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green
with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red,
quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later
attempt never replacing the first.
This commit is contained in:
garrytan committed 2026-09-29 19:23:55 +00:00
1 parent b1f5bc0032
commit 8eb55b6ebf
6 files changed
+1225 -178

No files matched your search

+2 -1
View File
@@ -248,7 +248,8 @@ describe('dependency-free CI planner and report execution', () => {
const red = run(['--report', reportDir], tier);
expect(red.status).toBe(1);
expect(red.stderr).toContain(`${failed.outcomes[0].files[0]}: failed`);
expect(red.stdout).toContain('3 executed, 0 reused; 1 passed, 2 failed, 0 manual accepted (unscored; no score-cache credit) (6 attempt records from 1 collectors)');
// Paid evals never retry: every record counts, a later pass never hides an earlier failure.
expect(red.stdout).toContain('6 executed, 0 reused; 2 passed, 4 failed, 0 manual accepted (unscored; no score-cache credit) (6 attempt records from 1 collectors');
expect(red.stdout).toContain('3 cases with multiple attempts this run:');
expect(red.stdout).not.toMatch(/passed only on retry|not blocking/);