diff --git a/AGENTS.md b/AGENTS.md index a68337ffd..c847ad649 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -148,7 +148,8 @@ When fixing failures or preparing `/ship`, follow this order: public events in free regressions, including negative controls, before paying for another agent run. Check behavior and acknowledgments; match exact prose only when that prose is the contract. Do not lower thresholds, increase model - budgets, skip cases, or rejudge a failure to manufacture a pass. + budgets, skip cases, or rejudge a failure to manufacture a pass. A + pre-registered fixed panel is not rejudging. For policy or validation repairs, exercise the actual registered callback with representative native input and assert that it uses the helper’s result. When renderer or parser failures recur at the same boundary, verify the @@ -208,10 +209,16 @@ When fixing failures or preparing `/ship`, follow this order: result and pending permission state; diagnose a blocked actor before waiting through its deadline. Preserve cancellation separately from a test verdict. Skipped or unstarted cases - do not satisfy coverage; preserve every attempt. Retries follow the approved - policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so - only files whose every case budget is CAPTURE tier or shorter keep one retry; - never add retries to pass a longer case. + do not satisfy coverage; preserve every attempt. Paid evals never retry. Each + case's kind (`E2E_KINDS`) fixes its trials before the run: `rule` one trial; + `behavior` a panel of 3 independent trials, PASS at >= 2 with no contract + violation; `judge` 3 samples on one output, gated on the mean against the + unchanged threshold. Never add trials, samples or dispatches after seeing a + result, never change a kind to change a verdict without pass-rate evidence, + and report every trial. Quarantine follows `CASE_QUARANTINE`'s entry and exit + rules only (`EVAL_POLICY`, `docs/TESTING_INTERNALS.md`). A census whose every + red is machine-classified INFRA or INCOMPLETE may be re-dispatched once as a + new run; report both runs. 7. Prove all known repairs with focused tests, including affected paid cases. Rerun a failed case only after a concrete repair or a demonstrated launch correction. Run the remaining required selected evaluations on the integrated @@ -245,6 +252,7 @@ bun run test # complete free suite via the strict shard runner (no A bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY) bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals bun run eval:bg:release # fresh complete gate + periodic live coverage +bun run eval:pass-rates # per-case trial pass rates (Wilson), drift and quarantine alarms (--case, --gate) bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free) bun run test:windows # curated Windows-safe subset (runs on windows-latest) bun run build # generate docs + compile binaries diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index db9f74c8e..07496edbb 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -239,10 +239,29 @@ Complete start-to-finish flows belong to the `marathon` tier (`describeE2ETier('marathon')`), which runs only in the non-blocking `evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge. -Retries: a timed-out attempt is a verdict. Only files whose every case budget is -CAPTURE tier or shorter (`RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`) -keep one automatic retry for fast-failing flakes; every other paid file runs once. -Case budgets themselves never change with this rule. +Verdicts: paid evals never retry. Each case's kind in `E2E_KINDS` +(`test/helpers/touchfiles-data.ts`) fixes its trials before the run, from the +constants in `EVAL_POLICY` (`test/helpers/periodic-exclude-data.ts`): + +- `rule` (the default): one trial; any failed assertion fails the case. Use it + when nothing stochastic decides the verdict, or when the verdict checks a + contract the product must meet every run (no writes in plan mode, a question + before a decision, a skill-mandated step, no leaked secret). +- `behavior`: a panel of 3 independent trials run as parallel case shards, + PASS at 2 or more with no contract violation (`expectContract()`). Use it only + when a live model choice decides the verdict and an occasional deviation is + acceptable product behavior; the one-line reason goes in `BEHAVIOR_WHY`. +- `judge`: an LLM judge scoring a fixed input; 3 samples of the same prompt, + gated on the per-dimension mean (booleans on a majority) against the + unchanged threshold. An erroring sample fails the panel and is never resampled. + +A timed-out, crashed or infrastructure-failed trial counts as a failed trial and +is reported with its class; a missing trial makes the case INCOMPLETE, which +fails the lane. A 2-of-3 pass is reported as `PASS 2/3` with the failed trial's +cause, never as a clean pass. Case budgets and thresholds never change with +this policy. Quarantine (`CASE_QUARANTINE`) and history are described in +`docs/TESTING_INTERNALS.md`; `bun run eval:pass-rates --case ` shows a +case's per-trial pass rate with its Wilson interval. CI enables verified first-attempt reuse for 16 workflow quality judges for 24 hours within the same PR. The cookie workflow's custom input and the other 11 @@ -415,7 +434,7 @@ When E2E tests run, they produce machine-readable artifacts in `~/.gstack-dev/`: bun run eval:list # list all eval runs (turns, duration, cost per run) bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary bun run eval:summary # aggregate stats + per-test efficiency averages across runs -bun run eval:flake-rank # rank tests by flake signal: retried passes first, then failure rate (--json, --dir, --since-days) +bun run eval:pass-rates # per-case trial pass rates + Wilson intervals from recent weekly runs (--case, --runs, --dir, --backfill, --json, --gate); eval:flake-rank is an alias ``` **Detached runs for agents and long suites.** When an agent (or you, for a run @@ -463,7 +482,9 @@ Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`: - **Completeness** — Are all commands, flags, and usage patterns documented? - **Actionability** — Can the agent execute tasks using only the information in the doc? -Each dimension is scored 1-5. Threshold: every dimension must score **≥ 4**. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher. +Each dimension is scored 1-5 by a panel of 3 samples of the same prompt, drawn +concurrently; each dimension's panel mean must meet that judge's threshold (≥ 4 +for most dimensions; see each case). An erroring sample fails the panel. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher. ```bash # Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals @@ -483,6 +504,24 @@ fails, add the named path to the named key and check selection with `bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. The rule is a lower bound: a fixture path the test builds at runtime is not visible to it, so add such paths to the key by hand. +### Add a paid eval + +1. **Test file.** Write the case in a paid test file, registered with a literal + name (`testIfSelected('', ...)`), grading the outcome (files, git + state, native questions, exit status) rather than wording, unless the step + itself is the contract. Wrap contract assertions in `expectContract()`. +2. **Touchfiles.** Add `'': [...]` to `E2E_TOUCHFILES`; `bun test + test/touchfiles.test.ts` names any missing closure path. +3. **Tier.** Add it to `E2E_TIERS`: `gate` for cheap contracts every PR needs, + `periodic` for long or model-quality cases, `marathon` for complete flows. +4. **Kind.** Add it to `E2E_KINDS` (`rule` unless a live model choice may + acceptably deviate; then `behavior` plus a `BEHAVIOR_WHY` line). + `bun test test/eval-kinds.test.ts` prints the literal to add. +5. **PR profile.** If a PR should run it, add it to `scripts/test-pr-profile.ts` + and check `bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. +6. **Try the panel locally.** `bun run scripts/test-paid-shards.ts --tier + --case --trials 3` runs the same panel CI runs, before you push. + ### CI A GitHub Action (`.github/workflows/skill-docs.yml`) generates all hosts on pushes to main and on PRs, then rejects tracked differences and nonignored untracked output. Generation errors also fail the job. Optional ignored host caches are not compared against Git. diff --git a/docs/TESTING_INTERNALS.md b/docs/TESTING_INTERNALS.md index 36d89ff3c..0cb64ac28 100644 --- a/docs/TESTING_INTERNALS.md +++ b/docs/TESTING_INTERNALS.md @@ -252,24 +252,20 @@ processes × `EVALS_CONCURRENCY` within-shard, per-shard `GSTACK_EVAL_DIR`, full-stream spooling to per-shard log files (path printed at START and on failure), never-started/timed-out taxonomy, and parent-computed diff selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a -child that can't parse it recomputes locally with one warning). Retries follow -one rule (`retriesForFiles`, `RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`): -a timed-out attempt is a verdict, so a file keeps one Bun retry only when every -case budget is CAPTURE tier plus its recording grace or shorter (registered rows -derive it from `caseMs`, `SHORT_CASE_RETRY_FILES` lists the rest); every other -file, including the former `retries: 2` matrix rows, runs once. `--list` prints -each shard's retries. Files in `CASE_SHARDED_FILES` run one registered case per +child that can't parse it recomputes locally with one warning). Paid evals +never retry; each case's kind fixes its trials before the run (see "Eval verdict +policy" below). Files in `CASE_SHARDED_FILES` run one registered case per process (`#`, an exact `--test-name-pattern`, exactly one executed case), so a long file of short cases spreads across runners and each case gets its own SDK semaphore. -Flake telemetry rides the store: every recorded test carries its 1-based -`attempt` (a pass-on-attempt-2 stays visible forever — bun's own stream hides -it), runs list `flaky_retries`, the report warns on passed-only-on-retry -tests, and `bun run eval:flake-rank` ranks the series (retried passes first, -then failure rate; 60-day recency bound on eval files; the free lane's flake -ledger is folded in from `flakeLedgerPath()` — override with -`GSTACK_FLAKE_LEDGER`, the same env var the CI free lane sets before -uploading the ledger as the `flake-ledger` artifact). Census integrity is +Trial telemetry rides the store: every recorded test carries its 1-based +`attempt` plus, on an isolated trial shard, its `case_id`, `kind`, `trial`, +`panel` and `policy_version`, and each lane's report uploads one +`trial-outcomes` JSONL line per trial. `bun run eval:pass-rates` +(`eval:flake-rank` is an alias) turns that history into per-case pass rates +(see "Pass-rate history" below; the free lane's flake ledger is folded in from +`flakeLedgerPath()` — override with `GSTACK_FLAKE_LEDGER`, the same env var the +CI free lane sets before uploading the ledger as the `flake-ledger` artifact). Census integrity is enforced from the free suite: every `E2E_TOUCHFILES` / `LLM_JUDGE_TOUCHFILES` key must name a living paid test (`test/touchfiles.test.ts`'s reverse invariant), and `git show :path` fixtures are banned — vendor the bytes @@ -313,7 +309,7 @@ CI supplies the scoped cache/runtime configuration; local runs are fresh by default. Cached scores must pass current assertions; reused records retain their original source and time and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts -to PR-profile E2E shards that run with zero retries (so the pass is structurally a +to PR-profile E2E shards (paid evals never retry, so a pass is structurally a first attempt): the identity hashes the test's import closure, every tracked file matched by the touchfiles of every case the file registers plus the global touchfiles, the runner/workflow/setup actions, the child's environment pins, the @@ -377,6 +373,118 @@ the runner parent and handed to shard children as `GSTACK_CLAUDE_CLI_VERSION` (never spawned on a test thread), so a TUI-drift flake hunt is a grep, not archaeology. +**Eval verdict policy** (`EVAL_POLICY` version 1 in +`test/helpers/periodic-exclude-data.ts`, pre-registered 2026-09-29). Paid evals +never retry. Each live case has exactly one kind in `E2E_KINDS` +(`test/helpers/touchfiles-data.ts`; `test/eval-kinds.test.ts` enforces coverage), +and the kind fixes its trials before the run: + +- `rule` (default): one trial; any failed assertion fails the verdict. For + cases where nothing stochastic decides the verdict, or where it checks a + contract the product must meet every run. +- `behavior`: a panel of `n = 3` independent trials, launched together as + isolated case shards on different slices (key `#~t`). All three + always run: no early stop and no conditional extra trial. PASS when at least + `k = 2` pass and no trial violated a contract (`expectContract()` stamps + `failure_class: 'contract'`). Each behavior case names its tolerated deviation + in `BEHAVIOR_WHY` and must have a literal registration so it can run alone. +- `judge`: an LLM judge scoring a fixed input. `judgePanel()` + (`test/helpers/llm-judge.ts`) draws 3 samples of the same prompt concurrently + inside the unchanged `JUDGE_MS`; numeric dimensions gate on the per-dimension + mean against the unchanged threshold (no dimension compensates for another), + booleans on a strict majority. A sample that errors (refusal, truncation, + non-JSON, a malformed field) fails the panel and is never resampled; a + refusal counts as an unscored panel only when every sample refused. + `callJudge`'s 429 backoff happens before any model output and is transport, + not a verdict retry. The workflow-judge cache stores whole panels only. + +`panelVerdict()` (`test/helpers/eval-store.ts`) is the single verdict +function the report, `collector-outcomes.json`, the PR comment and pass-rates +all use. A timed-out, crashed or infrastructure-failed trial is a failed trial +recorded with its class; a missing or duplicate trial record makes the verdict +INCOMPLETE, which fails the lane; a 2/3 PASS is shown as `PASS 2/3` with the +failed trial's cause. A manual re-run adds trials under a new run attempt and +never replaces the first attempt's verdict. A red census is never rerun on +unchanged inputs: each red is diagnosed as product, test/detector, harness or +infra and resolved by a concrete repair and a fresh census, or listed as a named +red. The one exception: a census whose every red verdict is machine-classified +INFRA or INCOMPLETE (missing slice artifact, runner loss, API error before the +first model turn) may be re-dispatched once as a new run, and both runs are +reported. Changing any `EVAL_POLICY` constant after seeing census results needs +Garry's re-approval, a `version` bump and a fresh census; +`test/periodic-exclude-policy.test.ts` pins the approved values. + +**Quarantine** (`CASE_QUARANTINE`, same file). An entry needs: a per-trial rate +below 95% over at least 10 post-policy trials of the case's current input +identity (pre-policy backfill may justify only an initial entry, labeled as +such); a written diagnosis in `reason` whose `failureClass` is `detector`, +`harness` or `model-latency` (a product defect is fixed or listed as a named +red, never quarantined); unchanged case touchfiles in the change that adds it; +and an owner, tracking pointer, `enteredAt` date and measurable `exit`. A +quarantined case still runs its full panel and reports in every lane but cannot +fail it, except on a hard break (0 of n) or a contract violation, and it never +counts as passing coverage. At most 10% of a blocking tier (gate, periodic) may +be quarantined. The weekly report fails when an entry passes its exit rule (at +least 97% over at least 10 trials) without being removed, when an entry is 8 +weekly runs old, or when a tier is over its cap. + +**Pass-rate history** (`bun run eval:pass-rates`, `scripts/eval-flake-rank.ts`). +It reads the `trial-outcomes` artifact of the last N completed +`evals-periodic.yml` runs on the current branch and `main` (flags: `--case`, +`--runs N`, `--branch`, `--dir`, `--backfill`, `--json`, `--gate`) and prints +per-case per-trial pass rates with 95% Wilson intervals. A series is one case +under one input identity, the hash of its own touchfiles minus +`GLOBAL_TOUCHFILES` (harness edits do not restart it), per model, Claude CLI +version and policy version; a change starts a new series and older ones stay +visible. Labels: INCONCLUSIVE below 10 trials, BROKEN when the latest run is +0/n after a prior interval at or above 95%, FLAKY when failures leave the +interval straddling 95%, FAILING when the whole interval is below it, PASSING +otherwise. `--backfill` imports legacy slice artifacts as pre-policy trials +(first attempt only; a record that names no registry id is listed as +unattributed, never guessed); they are display-only. `--gate` (the weekly +report) fails with ACTION REQUIRED, on post-policy trials of the current series +only, when a non-quarantined blocking case meets the entry rule (proposing an +entry), when a `rule` case does (rule case behaving like behavior: fix or +reclassify), when a blocking case's current identity is significantly below its +previous one (one-sided Fisher exact, α = 0.05, at least 6 trials each side, +Holm-controlled across the cases tested), and on the quarantine rules above. +History that cannot be fetched fails the gate closed. + +**The arithmetic.** With per-trial pass rate p, the chance a single case goes +red (a false red while the product works, the catch rate once it has +regressed): + +| p | 1 trial | 2-of-3 panel | +|---|---|---| +| 0.99 | 1.0% | 0.03% | +| 0.95 | 5.0% | 0.72% | +| 0.90 | 10.0% | 2.8% | +| 0.70 | 30.0% | 21.6% | +| 0.30 | 70.0% | 78.4% | + +The panel removes most false reds at healthy rates, but it catches a 0.95 → 0.70 +regression in one run only 21.6% of the time (a single trial 30%, retry-until-green +3%), so drift detection is the history rule's job, not the per-run verdict's. +The Fisher alarm is weak at the minimum sample (5.4% power for 0.95 → 0.70 at +6 trials a side), and ten straight passes still leave a 72% Wilson lower bound: +after this policy lands, every series starts INCONCLUSIVE. + +A lane is all green with probability Π p_rule × Π P(≥2 of 3 | p_behavior) × +Π p_judge. For the current registry (PR gate worst case: 107 rule cases and 24 +judges; weekly census: 190 rule, 22 behavior and 25 judge verdicts), with rule +and judge verdicts at p_rule: + +| p_rule | full PR gate | weekly, behavior p = 0.90 | 0.95 | 0.97 | +|---|---|---|---|---| +| 0.99 | 26.8% | 6.2% | 9.8% | 10.9% | +| 0.995 | 51.9% | 18.2% | 29.0% | 32.1% | +| 0.999 | 87.7% | 43.2% | 68.7% | 76.1% | + +The rule term dominates: a green lane on a working product needs rule cases to +be near-deterministic (0.999), which is why failing detectors are converted to +outcome checks and product defects are fixed or named, and why each census +reports its expected lane false-red from the measured rates. + **Timeout policy.** Paid tests use the tiers in `test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG); `test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall diff --git a/scripts/eval-flake-rank.ts b/scripts/eval-flake-rank.ts index 3e6028a10..363fd66d8 100644 --- a/scripts/eval-flake-rank.ts +++ b/scripts/eval-flake-rank.ts @@ -1,29 +1,55 @@ #!/usr/bin/env bun /** - * eval-flake-rank — the flake-telemetry dial (WS1). + * eval-pass-rates (alias: eval-flake-rank) — per-case trial pass rates. * - * Aggregates per-test series across every FINALIZED eval-store run on this - * machine (default: ~/.gstack/projects//evals/, shard dirs included) - * plus the free suite's flake ledger, and ranks tests by flake signal: - * retried passes first (a test that needs attempt 2 to go green is the - * definition of a flake), then failure rate. + * Reads trial records (one JSONL line per trial: case, kind, trial, outcome, + * exit_reason, duration, cost, model, CLI version, series identity, run id, + * sha, policy_version) from the last N completed `evals-periodic.yml` runs on + * the current branch and `main` (downloading only each run's small + * `trial-outcomes` artifact through `gh`), plus any local eval dirs, and + * prints per-case per-trial pass rates with 95% Wilson intervals. * - * This is the readable dial behind two policies: - * - a flaky pass never blocks a merge, but it is recorded and RANKED here; - * - the required-check promotion (WS16) needs weeks of clean flake-rank, - * not vibes. + * A series is one case under one input identity: the case's own touchfiles + * minus GLOBAL_TOUCHFILES (`caseSeriesIdentities`), grouped by model and CLI + * version, per policy_version. A new identity starts a new series; earlier + * series stay visible. Only post-policy trials of the current series feed the + * labels and alarms. Legacy eval-store records (`--backfill`, `--dir`) are + * imported as pre-policy trials (first attempt only; a missing attempt means + * 1) and are display-only. + * + * Labels: INCONCLUSIVE (below the entry rule's minimum trials), BROKEN (latest run 0/n + * after a prior interval at or above the entry rate), FLAKY (failures and an + * interval straddling the entry rate), FAILING (interval below the entry + * rate), PASSING (otherwise). + * + * The weekly gate (`--gate`) exits non-zero with ACTION REQUIRED when a + * non-quarantined case meets the quarantine entry rule, a rule case behaves + * like a behavior case, a blocking case's current-identity rate is + * significantly below its previous identity (one-sided Fisher exact, + * Holm-controlled across cases), or a CASE_QUARANTINE entry has met its exit + * rule, expired, or pushed its tier over the cap. History that cannot be + * fetched fails the gate closed. * * Usage: - * bun run eval:flake-rank # project eval dir - * bun run eval:flake-rank --dir # e.g. downloaded CI artifacts - * bun run eval:flake-rank --json # machine-readable + * bun run eval:pass-rates # last 10 weekly runs, this branch + main + * bun run eval:pass-rates --case --runs 20 + * bun run eval:pass-rates --dir # local eval dirs / downloaded artifacts (repeatable) + * bun run eval:pass-rates --backfill # also import legacy slice artifacts, labeled pre-policy + * bun run eval:pass-rates --json | --gate */ import * as fs from 'node:fs'; +import * as os from 'node:os'; import * as path from 'node:path'; -import { getProjectEvalDir, isPartialEval, isFinalizedEvalResultFile, type EvalResult } from '../test/helpers/eval-store'; -import { evalEntryOutcome } from '../test/helpers/eval-store'; +import { spawnSync } from 'node:child_process'; +import { createHash } from 'node:crypto'; +import { isPartialEval, isFinalizedEvalResultFile, evalEntryOutcome, failureClassOf, parseTrialOutcomes, sanitizeTrialError, + TRIAL_OUTCOME_SCHEMA, type EvalCaseKind, type EvalResult, type TrialOutcomeRecord } from '../test/helpers/eval-store'; import { flakeLedgerPath, type FlakeLedgerEntry } from './test-free-shards'; +import { E2E_KINDS, E2E_TIERS, E2E_TOUCHFILES, GLOBAL_TOUCHFILES, LLM_JUDGE_TOUCHFILES } from '../test/helpers/touchfiles-data'; +import { CASE_QUARANTINE, EVAL_POLICY } from '../test/helpers/periodic-exclude-data'; +import { matchGlob } from '../test/helpers/test-selection'; +import { CASE_TEST_NAMES } from './test-paid-shards'; interface TestSeries { name: string; @@ -111,33 +137,561 @@ function readFreeLedger(): FlakeLedgerEntry[] { return out; } +// --- Trial records --- + +/** + * A trial record as pass-rates reads it: eval-store's trial-outcomes schema + * plus the series identity the report job stamps (caseSeriesIdentities). + * policy_version 0 marks a pre-policy (backfilled) record. + */ +export type TrialRecord = TrialOutcomeRecord & { series_identity?: string }; + +/** Per-file cap for downloaded artifacts: pass-rates parses data only, never executes it. */ +export const TRIAL_OUTCOMES_MAX_BYTES = 8 * 1024 * 1024; + +/** Every `trial-outcomes*.jsonl` file under a directory, size-capped, schema-validated by eval-store. */ +export function readTrialOutcomeDir(dir: string): { records: TrialRecord[]; errors: string[] } { + const records: TrialRecord[] = []; + const errors: string[] = []; + if (!fs.existsSync(dir)) return { records, errors }; + for (const name of fs.readdirSync(dir, { recursive: true }) as string[]) { + if (!/(^|\/)trial-outcomes[^/]*\.jsonl$/.test(name)) continue; + const full = path.join(dir, name); + const parsed = parseTrialOutcomes(fs.readFileSync(full, 'utf8'), { maxBytes: TRIAL_OUTCOMES_MAX_BYTES }); + records.push(...parsed.records.map(record => ({ + ...record, series_identity: typeof (record as TrialRecord).series_identity === 'string' + ? (record as TrialRecord).series_identity!.slice(0, 64) : undefined }))); + errors.push(...parsed.errors.map(error => `${full}: ${error}`)); + } + return { records, errors }; +} + +// --- Registry attribution and series identity --- + +export interface Registry { + kinds: Record; + tiers: Record; + touchfiles: Record; + judgeTouchfiles: Record; + globals: readonly string[]; + testNames: Record; +} + +export const LIVE_REGISTRY: Registry = { + kinds: E2E_KINDS, tiers: E2E_TIERS, touchfiles: E2E_TOUCHFILES, judgeTouchfiles: LLM_JUDGE_TOUCHFILES, + globals: GLOBAL_TOUCHFILES, testNames: CASE_TEST_NAMES, +}; + +/** A case's tier: its E2E_TIERS value, or 'judge' for an LLM-judge entry. */ +export function caseTier(id: string, registry: Registry = LIVE_REGISTRY): string { + return registry.tiers[id] ?? (id in registry.judgeTouchfiles ? 'judge' : 'unknown'); +} + +/** + * Attribute a legacy eval-store record to a registry id: the case-shard slug + * suffix (`--`), the recorded name or its exact slug (`/qa b6-static` + * is `qa-b6-static`), a CASE_TEST_NAMES label, or the only id its shard file + * registers. Anything else is unattributed (null). + */ +export function attributeLegacyRecord(name: string, shard: string | undefined, registry: Registry = LIVE_REGISTRY): string | null { + const known = (id: string) => id in registry.kinds; + const [slugFile, slugCase] = (shard ?? '').split('--'); + if (slugCase && known(slugCase)) return slugCase; + if (known(name)) return name; + const slug = name.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, ''); + if (known(slug)) return slug; + const labeled = Object.entries(registry.testNames).find(([, label]) => label === name)?.[0]; + if (labeled && known(labeled)) return labeled; + if (slugFile) { + const file = `test/${slugFile}.test.ts`; + const owners = Object.keys(registry.touchfiles).filter(id => registry.touchfiles[id]!.includes(file)); + if (owners.length === 1 && known(owners[0]!)) return owners[0]!; + } + return null; +} + +/** + * Series identity per case: a hash of the git blob ids of the files matching + * the case's own touchfiles, excluding GLOBAL_TOUCHFILES (harness edits are + * markers, not new series). The report job stamps this on every trial record. + */ +export function caseSeriesIdentities(ids: string[], root: string, registry: Registry = LIVE_REGISTRY): Record { + const listed = spawnSync('git', ['ls-files', '-s'], { cwd: root, encoding: 'utf8', timeout: 20_000, maxBuffer: 64 * 1024 * 1024 }); + if (listed.status !== 0) throw new Error(`git ls-files failed: ${listed.stderr}`); + const blobs = listed.stdout.split('\n').filter(Boolean).map(line => { + const [meta, file] = line.split('\t'); + return { file: file!, blob: meta!.split(' ')[1]! }; + }).filter(entry => !registry.globals.some(pattern => matchGlob(entry.file, pattern))); + return Object.fromEntries(ids.map(id => { + const patterns = registry.touchfiles[id] ?? registry.judgeTouchfiles[id] ?? []; + const lines = blobs.filter(entry => patterns.some(pattern => matchGlob(entry.file, pattern))) + .map(entry => `${entry.file} ${entry.blob}`).sort(); + return [id, createHash('sha256').update(`${id}\n${lines.join('\n')}`).digest('hex').slice(0, 16)]; + })); +} + +/** + * Import legacy eval-store result files as pre-policy trials (policy_version + * 0, source 'backfill'): first attempt only (a missing attempt means 1), + * attributed by registry id, never guessed. A manual-review acceptance carries + * no automated verdict: it is counted and shown, never scored. Without a CI + * run, each local result file is its own run. + */ +export function backfillEvalFiles(files: string[], run?: { run_id: string; sha?: string; timestamp?: string }, + registry: Registry = LIVE_REGISTRY): { records: TrialRecord[]; unattributed: string[]; manualReviews: string[] } { + const records: TrialRecord[] = []; + const unattributed = new Set(); + const manualReviews: string[] = []; + for (const file of files) { + let result: EvalResult & { shard?: string; claude_cli_version?: string }; + try { result = JSON.parse(fs.readFileSync(file, 'utf8')); } catch { continue; } + if (isPartialEval(result, file) || !Array.isArray(result.tests)) continue; + const seen = new Set(); + for (const entry of result.tests) { + if ((entry.attempt ?? 1) !== 1 || seen.has(entry.name)) continue; + seen.add(entry.name); + const id = attributeLegacyRecord(entry.name, result.shard, registry); + if (!id) { unattributed.add(entry.name); continue; } + const outcome = evalEntryOutcome(entry); + if (outcome === 'manual-review') { manualReviews.push(id); continue; } + records.push({ + schema: TRIAL_OUTCOME_SCHEMA, case: id, + file: result.shard ? `test/${result.shard.split('--')[0]}.test.ts` : 'unknown', + tier: caseTier(id, registry), kind: registry.kinds[id]!, trial: 1, panel: { n: 1, k: 1 }, attempt: 1, + outcome, ...(outcome === 'failed' ? { failure_class: failureClassOf(entry) } : {}), + exit_reason: entry.exit_reason, error: sanitizeTrialError(entry.error), + duration_ms: Math.max(0, entry.duration_ms || 0), cost_usd: Math.max(0, entry.cost_usd || 0), + model: entry.model, cli_version: result.claude_cli_version, policy_version: 0, quarantined: false, + execution: entry.execution === 'reused' ? 'reused' : 'executed', source: 'backfill', + run_id: run?.run_id ?? `local:${file}`, sha: run?.sha ?? result.git_sha, recorded_at: run?.timestamp ?? result.timestamp, + }); + } + } + return { records, unattributed: [...unattributed].sort(), manualReviews }; +} + +// --- Statistics --- + +/** 95% Wilson score interval for k successes in n trials. */ +export function wilsonInterval(k: number, n: number, z = 1.96): { lo: number; hi: number } { + if (n <= 0) return { lo: 0, hi: 1 }; + const p = k / n, z2 = z * z, denom = 1 + z2 / n; + const center = (p + z2 / (2 * n)) / denom; + const half = (z * Math.sqrt(p * (1 - p) / n + z2 / (4 * n * n))) / denom; + return { lo: Math.max(0, center - half), hi: k === n ? 1 : Math.min(1, center + half) }; +} + +function logChoose(n: number, k: number): number { + let sum = 0; + for (let i = 1; i <= k; i++) sum += Math.log(n - k + i) - Math.log(i); + return sum; +} + +/** + * One-sided Fisher exact p-value that the CURRENT pass rate is below the + * PREVIOUS one: P(X <= curPass) under the hypergeometric null with the + * observed margins. + */ +export function fisherOneSidedLower(curPass: number, curN: number, prevPass: number, prevN: number): number { + const passes = curPass + prevPass, total = curN + prevN; + const denom = logChoose(total, passes); + let p = 0; + for (let x = Math.max(0, passes - prevN); x <= curPass; x++) p += Math.exp(logChoose(curN, x) + logChoose(prevN, passes - x) - denom); + return Math.min(1, p); +} + +/** Holm step-down: the indices whose p-values are rejected at family-wise alpha. */ +export function holmRejections(pValues: number[], alpha: number): Set { + const order = pValues.map((p, index) => ({ p, index })).sort((a, b) => a.p - b.p); + const rejected = new Set(); + for (let rank = 0; rank < order.length; rank++) { + if (order[rank]!.p > alpha / (order.length - rank)) break; + rejected.add(order[rank]!.index); + } + return rejected; +} + +// --- Analysis --- + +export type PassRateLabel = 'INCONCLUSIVE' | 'BROKEN' | 'FLAKY' | 'FAILING' | 'PASSING'; +export type AlarmKind = 'drift' | 'rule-as-behavior' | 'regression' | 'quarantine-exit' | 'quarantine-expired' + | 'quarantine-cap' | 'quarantine-invalid'; + +/** The EVAL_POLICY fields pass-rates reads (structural, so tests can vary them). */ +export interface PassRatePolicy { + version: number; + quarantine: { entry: { rate: number; minTrials: number }; exit: { rate: number; minTrials: number }; capFraction: number; expiryWeeklyRuns: number }; + drift: { fisherAlpha: number; fisherMinPerSide: number }; +} + +export type QuarantineEntry = (typeof CASE_QUARANTINE)[string]; + +/** Tiers whose cases block a lane; quarantine applies only to them. */ +export const BLOCKING_TIERS: readonly string[] = ['gate', 'periodic']; +const QUARANTINE_FAILURE_CLASSES: readonly string[] = ['detector', 'harness', 'model-latency']; + +export interface SeriesStats { + key: string; + identity: string; + model: string; + cli: string; + policyVersion: number; + passes: number; + /** Scored trials: passed + failed (skipped trials carry no verdict). */ + trials: number; + infra: number; + interval: { lo: number; hi: number }; + firstSeen: string; + lastSeen: string; + runs: string[]; +} + +export interface CasePassRate { + case: string; + kind: EvalCaseKind; + tier: string; + quarantined: boolean; + label: PassRateLabel; + /** Manual-review acceptances: visible, never scored. */ + manualReviews: number; + current: SeriesStats | null; + previous: SeriesStats | null; + prePolicy: SeriesStats | null; + series: SeriesStats[]; + latestRun: { runId: string; passes: number; trials: number } | null; +} + +export interface Alarm { kind: AlarmKind; case: string; message: string } + +export interface PassRateReport { + policyVersion: number; + cases: CasePassRate[]; + alarms: Alarm[]; + postPolicyTrials: number; + prePolicyTrials: number; + unattributed: string[]; + errors: string[]; +} + +export interface AnalyzeOptions { + registry?: Registry; + quarantine?: Record; + policy?: PassRatePolicy; + /** Completed weekly-run timestamps in the window, for quarantine expiry. */ + weeklyRuns?: string[]; + now?: number; + unattributed?: string[]; + errors?: string[]; + manualReviews?: string[]; +} + +const at = (record: TrialRecord) => record.recorded_at ?? ''; +const runOf = (record: TrialRecord) => `${record.run_id ?? record.sha ?? 'local'}#${record.attempt}`; + +function seriesStats(key: string, records: TrialRecord[]): SeriesStats { + const scored = records.filter(record => record.outcome !== 'skipped'); + const passes = scored.filter(record => record.outcome === 'passed').length; + const times = records.map(at).sort(); + const first = records[0]!; + return { + key, identity: first.series_identity ?? 'unknown', model: first.model ?? 'unknown', cli: first.cli_version ?? 'unknown', + policyVersion: first.policy_version, passes, trials: scored.length, + infra: scored.filter(record => record.outcome === 'failed' && record.failure_class === 'infra').length, + interval: wilsonInterval(passes, scored.length), firstSeen: times[0] ?? '', lastSeen: times[times.length - 1] ?? '', + runs: [...new Set(records.map(runOf))], + }; +} + +/** Weekly runs completed after an entry's enteredAt; offline, whole weeks elapsed. */ +export function quarantineRunsSince(enteredAt: string, weeklyRuns: string[] | undefined, now: number): number { + const entered = Date.parse(enteredAt); + if (!Number.isFinite(entered)) return Number.POSITIVE_INFINITY; + if (weeklyRuns && weeklyRuns.length) return weeklyRuns.filter(time => Date.parse(time) > entered).length; + return Math.floor((now - entered) / (7 * 86_400_000)); +} + +/** + * Static CASE_QUARANTINE problems, shared by the free policy test and the + * weekly gate: an id that is not a blocking-tier E2E case, a missing field, + * a failure class outside detector / harness / model-latency (a product + * defect is fixed or named, never quarantined), a malformed or future date, + * and a tier over its cap. + */ +export function quarantinePolicyProblems(quarantine: Record, + registry: Registry = LIVE_REGISTRY, policy: PassRatePolicy = EVAL_POLICY, now = Date.now()): Alarm[] { + const problems: Alarm[] = []; + const invalid = (id: string, message: string) => problems.push({ kind: 'quarantine-invalid', case: id, message: `${id}: ${message}` }); + const perTier = new Map(); + for (const [id, entry] of Object.entries(quarantine)) { + const tier = registry.tiers[id]; + if (!tier || !(id in registry.kinds)) { invalid(id, 'CASE_QUARANTINE names no registered E2E case'); continue; } + if (!BLOCKING_TIERS.includes(tier)) invalid(id, `tier ${tier} is not blocking; only ${BLOCKING_TIERS.join(' and ')} cases are quarantined`); + for (const field of ['reason', 'failureClass', 'tracking', 'owner', 'enteredAt', 'exit'] as const) { + if (typeof entry[field] !== 'string' || !entry[field].trim()) invalid(id, `missing ${field}`); + } + if (typeof entry.reason === 'string' && entry.reason.trim().length < 40) invalid(id, 'reason must be a written diagnosis (at least 40 characters)'); + if (!QUARANTINE_FAILURE_CLASSES.includes(entry.failureClass)) { + invalid(id, `failureClass ${JSON.stringify(entry.failureClass)} is not ${QUARANTINE_FAILURE_CLASSES.join(', ')}; a product defect is fixed or named as a red, never quarantined`); + } + const entered = Date.parse(entry.enteredAt); + if (!/^\d{4}-\d{2}-\d{2}$/.test(entry.enteredAt ?? '') || !Number.isFinite(entered)) invalid(id, 'enteredAt must be YYYY-MM-DD'); + else if (entered > now) invalid(id, 'enteredAt is in the future'); + perTier.set(tier, (perTier.get(tier) ?? 0) + 1); + } + for (const [tier, count] of perTier) { + const size = Object.values(registry.tiers).filter(value => value === tier).length; + const cap = Math.floor(size * policy.quarantine.capFraction); + if (count > cap) problems.push({ kind: 'quarantine-cap', case: tier, + message: `${count} quarantined ${tier} cases exceed the ${pct(policy.quarantine.capFraction)} cap (${cap} of ${size})` }); + } + return problems; +} + +export function analyzePassRates(records: TrialRecord[], options: AnalyzeOptions = {}): PassRateReport { + const registry = options.registry ?? LIVE_REGISTRY; + const quarantine = options.quarantine ?? CASE_QUARANTINE; + const policy = options.policy ?? EVAL_POLICY; + const now = options.now ?? Date.now(); + const byCase = new Map(); + for (const record of records) { + const list = byCase.get(record.case) ?? []; + list.push(record); + byCase.set(record.case, list); + } + for (const id of options.manualReviews ?? []) if (!byCase.has(id)) byCase.set(id, []); + const cases: CasePassRate[] = []; + for (const [id, list] of [...byCase].sort(([a], [b]) => a.localeCompare(b))) { + list.sort((a, b) => at(a).localeCompare(at(b)) || runOf(a).localeCompare(runOf(b)) || a.trial - b.trial); + const groups = new Map(); + for (const record of list) { + const key = record.policy_version === 0 ? 'pre-policy' + : [record.series_identity ?? 'unknown', record.model ?? 'unknown', record.cli_version ?? 'unknown', `v${record.policy_version}`].join('|'); + const group = groups.get(key) ?? []; + group.push(record); + groups.set(key, group); + } + const series = [...groups].map(([key, group]) => seriesStats(key, group)) + .sort((a, b) => a.lastSeen.localeCompare(b.lastSeen)); + const post = series.filter(entry => entry.policyVersion !== 0); + const current = post[post.length - 1] ?? null; + const previous = post[post.length - 2] ?? null; + const scored = current ? groups.get(current.key)!.filter(record => record.outcome !== 'skipped') : []; + const latestRun = scored.length ? runOf(scored[scored.length - 1]!) : null; + const latest = scored.filter(record => runOf(record) === latestRun); + const prior = scored.filter(record => runOf(record) !== latestRun); + const priorPasses = prior.filter(record => record.outcome === 'passed').length; + const entryRate = policy.quarantine.entry.rate; + let label: PassRateLabel; + if (latest.length > 0 && latest.every(record => record.outcome === 'failed') + && prior.length > 0 && wilsonInterval(priorPasses, prior.length).lo >= entryRate) label = 'BROKEN'; + else if (!current || current.trials < policy.quarantine.entry.minTrials) label = 'INCONCLUSIVE'; + else if (current.interval.hi < entryRate) label = 'FAILING'; + else if (current.passes < current.trials && current.interval.lo < entryRate) label = 'FLAKY'; + else label = 'PASSING'; + cases.push({ + case: id, kind: registry.kinds[id] ?? list[0]!.kind, tier: caseTier(id, registry), + quarantined: id in quarantine, label, current, previous, + manualReviews: (options.manualReviews ?? []).filter(name => name === id).length, + prePolicy: series.find(entry => entry.policyVersion === 0) ?? null, series, + latestRun: latestRun ? { runId: latestRun, passes: latest.filter(record => record.outcome === 'passed').length, trials: latest.length } : null, + }); + } + + const alarms: Alarm[] = []; + const rate = (stats: SeriesStats) => stats.passes / stats.trials; + for (const entry of cases) { + const current = entry.current; + if (!current) continue; + const below = current.trials >= policy.quarantine.entry.minTrials && rate(current) < policy.quarantine.entry.rate; + if (below && !entry.quarantined && BLOCKING_TIERS.includes(entry.tier)) alarms.push({ kind: 'drift', case: entry.case, + message: `${entry.case} passes ${current.passes}/${current.trials} (below ${pct(policy.quarantine.entry.rate)} over >= ${policy.quarantine.entry.minTrials} trials): fix it, or propose a CASE_QUARANTINE entry with a written diagnosis (product defects are never quarantined)` }); + if (below && entry.kind === 'rule') alarms.push({ kind: 'rule-as-behavior', case: entry.case, + message: `${entry.case}: rule case behaving like behavior (${current.passes}/${current.trials}): fix or reclassify` }); + if (entry.quarantined && current.trials >= policy.quarantine.exit.minTrials && rate(current) >= policy.quarantine.exit.rate) { + alarms.push({ kind: 'quarantine-exit', case: entry.case, + message: `${entry.case} passes ${current.passes}/${current.trials} (>= ${pct(policy.quarantine.exit.rate)}): remove its CASE_QUARANTINE entry` }); + } + } + const tested = cases.filter(entry => BLOCKING_TIERS.includes(entry.tier) && entry.current && entry.previous + && entry.current.trials >= policy.drift.fisherMinPerSide && entry.previous.trials >= policy.drift.fisherMinPerSide); + const pValues = tested.map(entry => fisherOneSidedLower(entry.current!.passes, entry.current!.trials, entry.previous!.passes, entry.previous!.trials)); + for (const index of holmRejections(pValues, policy.drift.fisherAlpha)) { + const entry = tested[index]!; + alarms.push({ kind: 'regression', case: entry.case, + message: `${entry.case}: current identity ${entry.current!.passes}/${entry.current!.trials} is significantly below the previous ${entry.previous!.passes}/${entry.previous!.trials} (one-sided Fisher p=${pValues[index]!.toFixed(4)}, Holm over ${tested.length} cases)` }); + } + for (const [id, entry] of Object.entries(quarantine)) { + const runs = quarantineRunsSince(entry.enteredAt, options.weeklyRuns, now); + if (runs >= policy.quarantine.expiryWeeklyRuns) alarms.push({ kind: 'quarantine-expired', case: id, + message: `${id}: entered ${runs} weekly runs ago (limit ${policy.quarantine.expiryWeeklyRuns}): fix it, name it as a red, or re-diagnose with fresh evidence` }); + } + alarms.push(...quarantinePolicyProblems(quarantine, registry, policy, now)); + + const post = records.filter(record => record.policy_version !== 0).length; + return { policyVersion: policy.version, cases, alarms, postPolicyTrials: post, prePolicyTrials: records.length - post, + unattributed: options.unattributed ?? [], errors: options.errors ?? [] }; +} + +function pct(value: number): string { return `${Math.round(value * 1000) / 10}%`; } + +function formatStats(stats: SeriesStats | null): string { + if (!stats) return '-'; + return `${stats.passes}/${stats.trials} [${pct(stats.interval.lo)}–${pct(stats.interval.hi)}]${stats.infra ? ` (${stats.infra} infra)` : ''}`; +} + +export function formatPassRates(report: PassRateReport, options: { caseFilter?: string } = {}): string { + const lines: string[] = []; + lines.push(`pass-rates: policy v${report.policyVersion}, ${report.postPolicyTrials} post-policy trial(s), ${report.prePolicyTrials} pre-policy (display only)`); + if (report.postPolicyTrials === 0) lines.push(' no post-policy trials yet: every series starts INCONCLUSIVE'); + const cases = report.cases.filter(entry => !options.caseFilter || entry.case === options.caseFilter); + lines.push(' label kind tier current series pre-policy manual case'); + for (const entry of cases) { + const group = entry.current ? ` ${entry.current.model} / ${entry.current.cli}` : ''; + const reset = entry.previous ? ' (baseline reset)' : ''; + lines.push(` ${entry.label.padEnd(12)} ${entry.kind.padEnd(8)} ${entry.tier.padEnd(8)} ${formatStats(entry.current).padEnd(29)} ` + + `${formatStats(entry.prePolicy).padEnd(18)} ${String(entry.manualReviews).padStart(6)} ${entry.case}${entry.quarantined ? ' [quarantined]' : ''}${group}${reset}`); + } + if (report.unattributed.length) lines.push(` unattributed records (${report.unattributed.length}, never guessed): ${report.unattributed.slice(0, 20).join(', ')}`); + if (report.errors.length) lines.push(` rejected ${report.errors.length} invalid trial line(s): ${report.errors.slice(0, 5).join('; ')}`); + if (report.alarms.length) { + lines.push(`ACTION REQUIRED (${report.alarms.length}):`); + for (const alarm of report.alarms) lines.push(` [${alarm.kind}] ${alarm.message}`); + } + return lines.join('\n'); +} + +// --- GitHub history --- + +export interface WeeklyRun { id: number; attempt: number; sha: string; branch: string; createdAt: string } +export interface RunArtifact { id: number; name: string; size: number } + +/** The GitHub calls pass-rates makes; injectable so the free tests never touch the network. */ +export interface HistoryFetcher { + listRuns(repo: string, workflow: string, branch: string, limit: number): WeeklyRun[]; + listArtifacts(repo: string, runId: number): RunArtifact[]; + downloadZip(repo: string, artifactId: number, destination: string): void; +} + +function gh(args: string[]): Buffer { + const result = spawnSync('gh', args, { timeout: 300_000, maxBuffer: 256 * 1024 * 1024 }); + if (result.status !== 0) throw new Error(`gh ${args.slice(0, 2).join(' ')} failed: ${String(result.stderr || result.error || '').trim()}`); + return result.stdout; +} + +const jsonLines = (buffer: Buffer): T[] => buffer.toString('utf8').split('\n').filter(Boolean).map(line => JSON.parse(line) as T); + +export const GH_HISTORY: HistoryFetcher = { + listRuns: (repo, workflow, branch, limit) => jsonLines(gh(['api', + `repos/${repo}/actions/workflows/${workflow}/runs?branch=${encodeURIComponent(branch)}&status=completed&per_page=${limit}`, + '--jq', '.workflow_runs[] | {id, attempt: .run_attempt, sha: .head_sha, branch: .head_branch, createdAt: .created_at}'])), + listArtifacts: (repo, runId) => jsonLines(gh(['api', `repos/${repo}/actions/runs/${runId}/artifacts?per_page=100`, + '--paginate', '--jq', '.artifacts[] | select(.expired | not) | {id, name, size: .size_in_bytes}'])), + downloadZip: (repo, artifactId, destination) => fs.writeFileSync(destination, gh(['api', `repos/${repo}/actions/artifacts/${artifactId}/zip`])), +}; + +/** The last `limit` completed runs of `workflow` on each branch, newest first, deduplicated. */ +export function listWeeklyRuns(opts: { repo: string; workflow: string; branches: string[]; limit: number; fetcher?: HistoryFetcher }): WeeklyRun[] { + const fetcher = opts.fetcher ?? GH_HISTORY; + const runs = new Map(); + for (const branch of opts.branches) for (const run of fetcher.listRuns(opts.repo, opts.workflow, branch, opts.limit)) runs.set(run.id, run); + return [...runs.values()].sort((a, b) => b.createdAt.localeCompare(a.createdAt)); +} + +/** + * Download the artifacts of one run whose names match into a per-run cache + * directory (reused on later calls) and return the extracted directories. + * Oversized or oddly named artifacts are skipped: downloads are data only. + */ +export function downloadRunArtifacts(opts: { repo: string; run: WeeklyRun; match: (name: string) => boolean; cacheDir: string; + fetcher?: HistoryFetcher; maxBytes?: number }): string[] { + const fetcher = opts.fetcher ?? GH_HISTORY; + const dirs: string[] = []; + for (const artifact of fetcher.listArtifacts(opts.repo, opts.run.id)) { + if (!opts.match(artifact.name) || !/^[A-Za-z0-9._-]+$/.test(artifact.name)) continue; + if (artifact.size > (opts.maxBytes ?? TRIAL_OUTCOMES_MAX_BYTES)) continue; + const dir = path.join(opts.cacheDir, `${opts.run.id}`, artifact.name); + if (!fs.existsSync(path.join(dir, '.complete'))) { + fs.rmSync(dir, { recursive: true, force: true }); + fs.mkdirSync(dir, { recursive: true }); + const zip = path.join(dir, 'artifact.zip'); + fetcher.downloadZip(opts.repo, artifact.id, zip); + const unzip = spawnSync('unzip', ['-o', '-q', zip, '-d', dir], { timeout: 120_000 }); + if (unzip.status !== 0) throw new Error(`unzip failed for ${artifact.name}: ${String(unzip.stderr || unzip.error || '')}`); + fs.rmSync(zip, { force: true }); + fs.writeFileSync(path.join(dir, '.complete'), ''); + } + dirs.push(dir); + } + return dirs; +} + +function gitOutput(args: string[]): string | null { + const result = spawnSync('git', args, { encoding: 'utf8', timeout: 5_000 }); + return result.status === 0 ? result.stdout.trim() : null; +} + +function repoSlug(): string { + const url = gitOutput(['remote', 'get-url', 'origin']) ?? ''; + return url.match(/[:/]([^/:]+\/[^/]+?)(?:\.git)?$/)?.[1] ?? 'garrytan/gstack'; +} + if (import.meta.main) { const argv = process.argv.slice(2); - const dirFlag = argv.indexOf('--dir'); - const dir = dirFlag !== -1 ? argv[dirFlag + 1] : getProjectEvalDir(); + const flag = (name: string) => { const index = argv.indexOf(name); return index === -1 ? undefined : argv[index + 1]; }; + const dirs = argv.flatMap((arg, index) => arg === '--dir' && argv[index + 1] ? [argv[index + 1]!] : []); const asJson = argv.includes('--json'); - const sinceFlag = argv.indexOf('--since-days'); - const sinceDays = sinceFlag !== -1 ? Number(argv[sinceFlag + 1]) || 60 : 60; + const gate = argv.includes('--gate'); + const backfill = argv.includes('--backfill'); + const caseFilter = flag('--case'); + const runsLimit = Number(flag('--runs')) || 10; + const sinceDays = Number(flag('--since-days')) || 60; + const repo = flag('--repo') ?? repoSlug(); + const workflow = flag('--workflow') ?? 'evals-periodic.yml'; + const branch = flag('--branch') ?? gitOutput(['rev-parse', '--abbrev-ref', 'HEAD']) ?? 'main'; - const files = collectEvalFiles(dir, sinceDays); - const series = [...aggregate(files).values()] - .sort((a, b) => b.retriedPasses - a.retriedPasses || (b.fails / Math.max(1, b.runs)) - (a.fails / Math.max(1, a.runs))); - const ledger = readFreeLedger(); + const records: TrialRecord[] = []; + const unattributed = new Set(); + const errors: string[] = []; + let historyError: string | null = null; + let weeklyRuns: string[] | undefined; - if (asJson) { - console.log(JSON.stringify({ dir, runsScanned: files.length, tests: series, freeLedger: ledger }, null, 2)); + const manualReviews: string[] = []; + const importDir = (dir: string, run: { run_id: string; sha?: string; timestamp?: string } | undefined, legacyDays: number) => { + const trials = readTrialOutcomeDir(dir); + records.push(...trials.records); + errors.push(...trials.errors); + const legacy = backfillEvalFiles(collectEvalFiles(dir, legacyDays), run); + records.push(...legacy.records); + manualReviews.push(...legacy.manualReviews); + legacy.unattributed.forEach(name => unattributed.add(name)); + }; + + if (dirs.length) { + for (const dir of dirs) importDir(dir, undefined, sinceDays); } else { - console.log(`flake-rank: ${files.length} finalized run file(s) under ${dir}`); - const flaky = series.filter((s) => s.retriedPasses > 0 || s.fails > 0 || s.manualAccepted > 0); - if (flaky.length === 0) { - console.log(' no retried passes and no failures recorded — clean series'); - } else { - console.log(' retries fails/runs manual avg-dur test'); - for (const s of flaky.slice(0, 30)) { - console.log(` ${String(s.retriedPasses).padStart(7)} ${String(s.fails).padStart(5)}/${String(s.runs).padEnd(6)} ` - + `${String(s.manualAccepted).padStart(6)} ${Math.round(s.totalDurationMs / s.totalAttempts / 1000).toString().padStart(5)}s ${s.name}`); + try { + const runs = listWeeklyRuns({ repo, workflow, branches: [...new Set([branch, 'main'])], limit: runsLimit }); + weeklyRuns = runs.map(run => run.createdAt); + const cacheDir = path.join(os.homedir(), '.gstack', 'eval-pass-rates-cache', repo.replace('/', '-')); + const match = backfill + ? (name: string) => name.startsWith('trial-outcomes') || /^(paid-slice-\d+|gate-census-\d+)$/.test(name) + : (name: string) => name.startsWith('trial-outcomes'); + for (const run of runs) { + const dirsForRun = downloadRunArtifacts({ repo, run, match, cacheDir, maxBytes: backfill ? 64 * 1024 * 1024 : undefined }); + for (const dir of dirsForRun) importDir(dir, { run_id: `${run.id}`, sha: run.sha, timestamp: run.createdAt }, 3650); } + } catch (error) { + historyError = error instanceof Error ? error.message : String(error); } + } + + const report = analyzePassRates(records, { weeklyRuns, unattributed: [...unattributed].sort(), errors, manualReviews }); + const ledger = readFreeLedger(); + if (asJson) { + console.log(JSON.stringify({ repo, workflow, branch, dirs, historyError, ...report, freeLedger: ledger }, null, 2)); + } else { + if (historyError) console.log(`pass-rates: history unavailable (${historyError}); every label below is INCONCLUSIVE`); + console.log(formatPassRates(report, { caseFilter })); if (ledger.length > 0) { const byFile = new Map(); for (const e of ledger) byFile.set(e.file, (byFile.get(e.file) ?? 0) + 1); @@ -147,4 +701,5 @@ if (import.meta.main) { } } } + if (gate && (historyError || report.alarms.length)) process.exit(1); } diff --git a/scripts/typecheck-test-baseline.json b/scripts/typecheck-test-baseline.json index c309682d1..612942cf7 100644 --- a/scripts/typecheck-test-baseline.json +++ b/scripts/typecheck-test-baseline.json @@ -261,7 +261,6 @@ "test/helpers/shared-libs-eval-fixture.ts\tTS7006\tParameter 'candidate' implicitly has an 'any' type.": 2, "test/helpers/shared-libs-path-fixture.ts\tTS2352\tConversion of type '{ root: string; repo: string; state: string; env: { GSTACK_HOME: string; }; }' to type 'SharedLibsFixture' may be a mistake because neither type sufficiently overlaps with the other. If this was intentional, convert the expression to 'unknown' first. Type '{ root: string; repo: string; state: string; env: { GSTACK_HOME: string; }; }' is missing the following properties from type 'SharedLibsFixture': bin, trace, hookTrace, tip": 1, "test/helpers/shared-libs-plan-actor.ts\tTS18046\t'questions' is of type 'unknown'.": 1, - "test/helpers/workflow-judge-cache.ts\tTS2352\tConversion of type 'string | number | boolean | EvalCacheValue[] | { [key: string]: EvalCacheValue; } | null' to type 'JudgeScore' may be a mistake because neither type sufficiently overlaps with the other. If this was intentional, convert the expression to 'unknown' first. Type '{ [key: string]: EvalCacheValue; }' is missing the following properties from type 'JudgeScore': clarity, completeness, actionability, reasoning": 1, "test/impeccable-fixtures.test.ts\tTS2769\tNo overload matches this call. The last overload gave the following error. Argument of type 'string | undefined' is not assignable to parameter of type 'string'. Type 'undefined' is not assignable to type 'string'.": 1, "test/llm-judge-abort.test.ts\tTS2769\tNo overload matches this call. The last overload gave the following error. Argument of type '{ passed: boolean; }' is not assignable to parameter of type 'undefined'.": 3, "test/llm-judge-frontier.test.ts\tTS2769\tNo overload matches this call. The last overload gave the following error. Argument of type '{ score: number; reason: string; }' is not assignable to parameter of type 'undefined'.": 1, diff --git a/test/cookie-workflow-judge-input.test.ts b/test/cookie-workflow-judge-input.test.ts index cc174995e..70e2b489c 100644 --- a/test/cookie-workflow-judge-input.test.ts +++ b/test/cookie-workflow-judge-input.test.ts @@ -11,7 +11,7 @@ import { selectTests } from './helpers/test-selection'; import { E2E_TOUCHFILES, LLM_JUDGE_TOUCHFILES } from './helpers/touchfiles-data'; import { selectPrProfile } from '../scripts/test-pr-profile'; import { JUDGE_MS } from './helpers/eval-budgets'; -import { JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS } from './helpers/llm-judge'; +import { JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS, judgePanel, judgePanelMean, judgePanelReasoning, JUDGE_SCORE_DIMENSIONS, JUDGE_PANEL_SAMPLES } from './helpers/llm-judge'; import { COOKIE_MANUAL_REVIEW_FILE, getCookieWorkflowManualReview, isManualReviewEntry } from './helpers/cookie-workflow-manual-review'; const ROOT = resolve(import.meta.dir, '..'); @@ -70,7 +70,7 @@ function actualCookieCallback(root: string, overrides: { const records: EvalTestEntry[] = []; const attempts = new Map(); let callback: () => Promise = async () => { throw new Error('Judge callback was not registered'); }; - new Function('describeIfSelected', 'testIfSelected', 'ROOT', 'buildCookieWorkflowJudgeInput', 'resolveEvalModel', 'callJudge', 'COOKIE_WORKFLOW_JUDGE', 'JUDGE_MS', 'WORKFLOW_JUDGE_TEST_MS', 'WORKFLOW_JUDGE_RECORD_MS', 'evalCollector', 'expect', 'console', 'readWorkflowJudgeInput', 'buildWorkflowJudgePrompt', 'prepareWorkflowJudgeCache', 'workflowJudgeAttempts', 'performance', 'setTimeout', 'clearTimeout', 'JudgeRefusalError', 'getCookieWorkflowManualReview', 'DEFAULT_JUDGE_MAX_TOKENS', 'WORKFLOW_JUDGE_RESPONSE_SCHEMA', 'validWorkflowJudgeScore', registration)( + new Function('describeIfSelected', 'testIfSelected', 'ROOT', 'buildCookieWorkflowJudgeInput', 'resolveEvalModel', 'callJudge', 'COOKIE_WORKFLOW_JUDGE', 'JUDGE_MS', 'WORKFLOW_JUDGE_TEST_MS', 'WORKFLOW_JUDGE_RECORD_MS', 'evalCollector', 'expect', 'console', 'readWorkflowJudgeInput', 'buildWorkflowJudgePrompt', 'prepareWorkflowJudgeCache', 'workflowJudgeAttempts', 'performance', 'setTimeout', 'clearTimeout', 'JudgeRefusalError', 'getCookieWorkflowManualReview', 'DEFAULT_JUDGE_MAX_TOKENS', 'WORKFLOW_JUDGE_RESPONSE_SCHEMA', 'validWorkflowJudgeScore', 'judgePanel', 'judgePanelMean', 'judgePanelReasoning', 'JUDGE_SCORE_DIMENSIONS', registration)( (_suite: string, names: string[], run: () => void) => { expect(names).toEqual([NAME]); run(); }, (name: string, run: () => Promise, budget: number) => { expect(name).toBe(NAME); expect(budget).toBe(JUDGE_MS + 10_000); callback = run; }, root, buildCookieWorkflowJudgeInput, (_kind: string, explicit?: string) => explicit ?? 'fixture-model', @@ -85,7 +85,7 @@ function actualCookieCallback(root: string, overrides: { attempts, overrides.clock ? { now: overrides.clock } : performance, overrides.setTimer ?? setTimeout, overrides.clearTimer ?? clearTimeout, JudgeRefusalError, getCookieWorkflowManualReview, DEFAULT_JUDGE_MAX_TOKENS, - WORKFLOW_JUDGE_RESPONSE_SCHEMA, validWorkflowJudgeScore, + WORKFLOW_JUDGE_RESPONSE_SCHEMA, validWorkflowJudgeScore, judgePanel, judgePanelMean, judgePanelReasoning, JUDGE_SCORE_DIMENSIONS, ); return { run: () => callback(), requests, records, attempts }; } @@ -96,7 +96,7 @@ describe('cookie workflow judge input', () => { approveFixture(root); const h = actualCookieCallback(root, { judge: async () => { throw refusal(); } }); await h.run(); - expect(h.requests).toHaveLength(1); + expect(h.requests).toHaveLength(JUDGE_PANEL_SAMPLES); expect(h.records).toHaveLength(1); expect(h.records[0]).toMatchObject({ passed: false, execution: 'executed', exit_reason: 'provider_refusal' }); expect(isManualReviewEntry(h.records[0])).toBe(true); @@ -161,7 +161,7 @@ describe('cookie workflow judge input', () => { const root = fixture(); approveFixture(root); let calls = 0; const h = actualCookieCallback(root, { judge: async () => { - if (++calls === 1) return { ...passingScore, clarity: 1 }; + if (++calls <= JUDGE_PANEL_SAMPLES) return { ...passingScore, clarity: 1 }; throw refusal(); } }); await expect(h.run()).rejects.toThrow(); @@ -275,7 +275,7 @@ describe('cookie workflow judge input', () => { let scores = passingScore; const h = actualCookieCallback(root, { judge: async () => scores }); await h.run(); - expect(h.requests).toHaveLength(1); + expect(h.requests).toHaveLength(JUDGE_PANEL_SAMPLES); expect(h.requests[0].prompt).toBe(input.prompt); expect(h.requests[0].model).toBe(COOKIE_WORKFLOW_JUDGE.model); expect(h.requests[0].signal).toBeInstanceOf(AbortSignal); @@ -283,7 +283,7 @@ describe('cookie workflow judge input', () => { expect(existsSync(join(root, 'cache'))).toBe(false); const fresh = actualCookieCallback(root); await fresh.run(); - expect(fresh.requests).toHaveLength(1); + expect(fresh.requests).toHaveLength(JUDGE_PANEL_SAMPLES); for (const dimension of ['clarity', 'completeness', 'actionability'] as const) { scores = { ...COOKIE_WORKFLOW_JUDGE.thresholds, [dimension]: COOKIE_WORKFLOW_JUDGE.thresholds[dimension] - 1, reasoning: 'Synthetic failing fixture score' }; await expect(h.run()).rejects.toThrow(); diff --git a/test/eval-flake-rank.test.ts b/test/eval-flake-rank.test.ts index 198a64d54..f3d815c72 100644 --- a/test/eval-flake-rank.test.ts +++ b/test/eval-flake-rank.test.ts @@ -13,6 +13,13 @@ import * as path from 'node:path'; import { spawnSync } from 'node:child_process'; import { aggregate, collectEvalFiles } from '../scripts/eval-flake-rank'; import { manualReviewFixture } from './helpers/manual-judge-review-fixture'; +import { + analyzePassRates, attributeLegacyRecord, backfillEvalFiles, caseSeriesIdentities, downloadRunArtifacts, fisherOneSidedLower, + formatPassRates, holmRejections, listWeeklyRuns, quarantinePolicyProblems, quarantineRunsSince, readTrialOutcomeDir, + wilsonInterval, type HistoryFetcher, type PassRatePolicy, type QuarantineEntry, type Registry, type TrialRecord, +} from '../scripts/eval-flake-rank'; +import { EVAL_POLICY } from './helpers/periodic-exclude-data'; +import { TRIAL_OUTCOME_SCHEMA, formatTrialOutcomes } from './helpers/eval-store'; const entry = (name: string, passed: boolean, attempt: number) => ({ name, suite: 's', tier: 'e2e', passed, attempt, duration_ms: 1000, cost_usd: 0.1, @@ -39,9 +46,10 @@ describe('eval-flake-rank aggregate', () => { const display = spawnSync(process.execPath, [path.resolve(import.meta.dir, '../scripts/eval-flake-rank.ts'), '--dir', dir], { encoding: 'utf8', timeout: 10_000 }); expect(display.status, display.stderr).toBe(0); - expect(display.stdout).toContain('fails/runs manual'); - expect(display.stdout).toContain('0/1'); - expect(display.stdout).toContain(manual.name); + // pass-rates view: the prior automated pass is the one scored pre-policy + // trial; the manual acceptance is counted in its own column, never scored. + expect(display.stdout).toContain('pre-policy manual case'); + expect(display.stdout).toMatch(new RegExp(`1/1 \\[[^\\]]+\\]\\s+1 ${manual.name.replace(/[.*+?^${}()|[\]\\/]/g, '\\$&')}`)); fs.writeFileSync(path.join(dir, 'invalid-retry.json'), run([ { ...ordinary, attempt: 1 }, { ...manual, attempt: 2 }, ])); @@ -97,3 +105,300 @@ describe('eval-flake-rank aggregate', () => { fs.rmSync(dir, { recursive: true, force: true }); }); }); + +// --- pass-rates --- + +const registry: Registry = { + kinds: { 'rule-a': 'rule', 'beh-b': 'behavior', 'gate-c': 'rule', 'mar-d': 'rule', 'judge one': 'judge', + ...Object.fromEntries(Array.from({ length: 18 }, (_, i) => [`filler-${i}`, 'rule'])) }, + tiers: { 'rule-a': 'periodic', 'beh-b': 'periodic', 'gate-c': 'gate', 'mar-d': 'marathon', + ...Object.fromEntries(Array.from({ length: 18 }, (_, i) => [`filler-${i}`, i < 9 ? 'gate' : 'periodic'])) }, + touchfiles: { 'rule-a': ['test/skill-e2e-a.test.ts', 'a/**'], 'beh-b': ['test/skill-e2e-b.test.ts', 'b/**'], + 'gate-c': ['test/skill-e2e-shared.test.ts'], 'mar-d': ['test/skill-e2e-shared.test.ts'] }, + judgeTouchfiles: { 'judge one': ['j/SKILL.md'] }, + globals: ['harness/**'], + testNames: { 'gate-c': '/gate c labeled' }, +}; + +let clock = 0; +function trial(id: string, outcome: 'passed' | 'failed' | 'skipped', extra: Partial = {}): TrialRecord { + clock += 1; + return { + schema: TRIAL_OUTCOME_SCHEMA, case: id, file: 'test/x.test.ts', tier: registry.tiers[id] ?? 'judge', + kind: registry.kinds[id]!, trial: 1, panel: { n: 1, k: 1 }, attempt: 1, outcome, + ...(outcome === 'failed' ? { failure_class: 'assertion' as const } : {}), + duration_ms: 1, cost_usd: 0, model: 'model-x', cli_version: '2.1.284', policy_version: 1, quarantined: false, + execution: 'executed', source: 'shard', run_id: `run-${clock}`, recorded_at: new Date(Date.UTC(2026, 9, 1) + clock * 60_000).toISOString(), + series_identity: 'id-1', ...extra, + }; +} +const many = (id: string, passes: number, fails: number, extra: Partial = {}) => + [...Array.from({ length: passes }, () => trial(id, 'passed', extra)), ...Array.from({ length: fails }, () => trial(id, 'failed', extra))]; +const analyze = (records: TrialRecord[], quarantine: Record = {}, extra = {}) => + analyzePassRates(records, { registry, quarantine, now: Date.UTC(2026, 9, 2), ...extra }); +const qEntry = (overrides: Partial = {}): QuarantineEntry => ({ + reason: 'Detector graded the posture wording; 7 of 10 fresh trials failed only the regex, transcripts attached.', + failureClass: 'detector', tracking: 'TODOS.md "x"', owner: 'garrytan', enteredAt: '2026-09-29', + exit: '>= 97% over >= 10 trials on the current identity', ...overrides, +}); + +describe('pass-rates statistics', () => { + test('Wilson bounds match the documented policy arithmetic', () => { + expect(wilsonInterval(10, 10).lo).toBeCloseTo(0.7225, 4); + expect(wilsonInterval(6, 6).lo).toBeCloseTo(0.6097, 4); + expect(wilsonInterval(125, 125).lo).toBeCloseTo(0.9702, 4); + expect(wilsonInterval(10, 10).hi).toBe(1); + expect(wilsonInterval(0, 0)).toEqual({ lo: 0, hi: 1 }); + const mid = wilsonInterval(7, 10); + expect(mid.lo).toBeGreaterThan(0.39); expect(mid.hi).toBeLessThan(0.9); + }); + + test('one-sided Fisher exact matches a known table and is one-sided', () => { + expect(fisherOneSidedLower(4, 6, 6, 6)).toBeCloseTo(0.227272727, 8); + expect(fisherOneSidedLower(6, 6, 4, 6)).toBe(1); + expect(fisherOneSidedLower(0, 10, 10, 10)).toBeLessThan(1e-4); + }); + + test('Holm rejects step-down and stops at the first non-rejection', () => { + expect([...holmRejections([0.001, 0.02, 0.04], 0.05)].sort()).toEqual([0, 1, 2]); + expect([...holmRejections([0.001, 0.03, 0.04], 0.05)]).toEqual([0]); + expect([...holmRejections([0.03, 0.04], 0.05)]).toEqual([]); + expect([...holmRejections([], 0.05)]).toEqual([]); + }); +}); + +describe('pass-rates labels', () => { + test('thin history is INCONCLUSIVE, and after this PR every series starts there', () => { + const report = analyze(many('rule-a', 9, 0)); + expect(report.cases[0]).toMatchObject({ case: 'rule-a', label: 'INCONCLUSIVE' }); + expect(formatPassRates(analyze([]))).toContain('every series starts INCONCLUSIVE'); + }); + + test('PASSING, FLAKY and FAILING come from the interval against the entry rate', () => { + expect(analyze(many('rule-a', 12, 0)).cases[0]!.label).toBe('PASSING'); + expect(analyze(many('beh-b', 10, 1)).cases[0]!.label).toBe('FLAKY'); + expect(analyze(many('beh-b', 2, 10)).cases[0]!.label).toBe('FAILING'); + }); + + test('BROKEN: the latest run is 0/n after a prior interval at or above the entry rate', () => { + const prior = many('beh-b', 80, 0, { run_id: 'old' }); + const latest = [1, 2, 3].map(n => trial('beh-b', 'failed', { run_id: 'new', trial: n, panel: { n: 3, k: 2 } })); + expect(analyze([...prior, ...latest]).cases[0]!.label).toBe('BROKEN'); + }); + + test('skipped trials carry no verdict; infra failures count as failed trials', () => { + const stats = analyze([...many('rule-a', 10, 0), trial('rule-a', 'skipped'), + trial('rule-a', 'failed', { failure_class: 'infra' })]).cases[0]!.current!; + expect(stats).toMatchObject({ passes: 10, trials: 11, infra: 1 }); + }); + + test('a new identity, model or CLI starts a new series; earlier series stay visible', () => { + const report = analyze([...many('rule-a', 10, 0), ...many('rule-a', 3, 0, { series_identity: 'id-2' }), + ...many('rule-a', 2, 0, { series_identity: 'id-2', cli_version: '2.1.285' })]); + const c = report.cases[0]!; + expect(c.series).toHaveLength(3); + expect(c.current).toMatchObject({ identity: 'id-2', cli: '2.1.285', trials: 2 }); + expect(c.previous).toMatchObject({ identity: 'id-2', cli: '2.1.284', trials: 3 }); + expect(c.label).toBe('INCONCLUSIVE'); + }); +}); + +describe('pass-rates alarms count post-policy trials of the current series only', () => { + test('backfilled pre-policy failures are displayed but never alarm', () => { + const report = analyze(many('rule-a', 2, 20, { policy_version: 0, source: 'backfill' })); + expect(report.alarms).toEqual([]); + expect(report.cases[0]!.prePolicy).toMatchObject({ passes: 2, trials: 22 }); + expect(report.cases[0]!.label).toBe('INCONCLUSIVE'); + }); + + test('drift proposes quarantine for a blocking case below the entry rule; a rule case is flagged as behaving like behavior', () => { + const kinds = analyze([...many('rule-a', 8, 2), ...many('beh-b', 8, 2), ...many('mar-d', 0, 10)]).alarms.map(a => `${a.kind}:${a.case}`); + expect([...kinds].sort()).toEqual(['drift:beh-b', 'drift:rule-a', 'rule-as-behavior:mar-d', 'rule-as-behavior:rule-a']); + expect(analyze(many('rule-a', 19, 1)).alarms).toEqual([]); + }); + + test('the Fisher regression alarm needs the minimum trials on both sides', () => { + const old = many('gate-c', 6, 0, { series_identity: 'old' }); + const fresh = many('gate-c', 0, 6, { series_identity: 'new' }); + expect(analyze([...old, ...fresh]).alarms.map(a => a.kind)).toContain('regression'); + expect(analyze([...old, ...fresh.slice(0, 5)]).alarms.map(a => a.kind)).not.toContain('regression'); + }); + + test('quarantine exit, expiry and cap', () => { + const exit = analyze(many('beh-b', 10, 0), { 'beh-b': qEntry() }).alarms.map(a => a.kind); + expect(exit).toContain('quarantine-exit'); + expect(exit).not.toContain('drift'); + const weekly = Array.from({ length: 8 }, (_, i) => new Date(Date.UTC(2026, 8, 30) + i * 7 * 86_400_000).toISOString()); + expect(analyze([], { 'beh-b': qEntry() }, { weeklyRuns: weekly }).alarms.map(a => a.kind)).toContain('quarantine-expired'); + expect(analyze([], { 'beh-b': qEntry() }, { weeklyRuns: weekly.slice(0, 7) }).alarms.map(a => a.kind)).not.toContain('quarantine-expired'); + expect(quarantineRunsSince('2026-09-01', undefined, Date.UTC(2026, 9, 27))).toBe(8); + expect(quarantineRunsSince('not a date', undefined, 0)).toBe(Number.POSITIVE_INFINITY); + }); +}); + +describe('quarantine policy', () => { + const policy: PassRatePolicy = EVAL_POLICY; + const now = Date.UTC(2026, 9, 2); + test('a valid entry has no problems', () => { + expect(quarantinePolicyProblems({ 'beh-b': qEntry() }, registry, policy, now)).toEqual([]); + }); + + test('a product defect, a missing diagnosis or field, a bad date or a non-blocking case is rejected', () => { + const problems = (quarantine: Record) => quarantinePolicyProblems(quarantine, registry, policy, now).map(p => p.message); + expect(problems({ 'beh-b': qEntry({ failureClass: 'product' as QuarantineEntry['failureClass'] }) }).join()).toContain('never quarantined'); + expect(problems({ 'beh-b': qEntry({ reason: 'flaky' }) }).join()).toContain('written diagnosis'); + expect(problems({ 'beh-b': qEntry({ owner: ' ' }) }).join()).toContain('missing owner'); + expect(problems({ 'beh-b': qEntry({ enteredAt: '09/29/2026' }) }).join()).toContain('YYYY-MM-DD'); + expect(problems({ 'beh-b': qEntry({ enteredAt: '2027-01-01' }) }).join()).toContain('future'); + expect(problems({ 'mar-d': qEntry() }).join()).toContain('not blocking'); + expect(problems({ 'judge one': qEntry() }).join()).toContain('no registered E2E case'); + expect(problems({ ghost: qEntry() }).join()).toContain('no registered E2E case'); + }); + + test('at most 10% of a tier may be quarantined', () => { + // 11 periodic cases in the fixture registry: the cap is 1. + expect(quarantinePolicyProblems({ 'beh-b': qEntry() }, registry, policy, now)).toEqual([]); + const over = quarantinePolicyProblems({ 'beh-b': qEntry(), 'rule-a': qEntry() }, registry, policy, now); + expect(over.map(p => p.kind)).toEqual(['quarantine-cap']); + expect(over[0]!.message).toContain('2 quarantined periodic cases exceed the 10% cap (1 of 11)'); + }); +}); + +describe('pass-rates inputs', () => { + test('trial-outcomes JSONL is schema-validated; invalid lines are reported, never guessed', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'passrates-')); + const valid = trial('rule-a', 'passed'); + fs.mkdirSync(path.join(dir, 'nested')); + fs.writeFileSync(path.join(dir, 'nested', 'trial-outcomes.jsonl'), formatTrialOutcomes([valid]) + '{"schema":"other"}\nnot json\n'); + fs.writeFileSync(path.join(dir, 'unrelated.jsonl'), formatTrialOutcomes([valid])); + const read = readTrialOutcomeDir(dir); + expect(read.records).toHaveLength(1); + expect(read.records[0]).toMatchObject({ case: 'rule-a', series_identity: 'id-1' }); + expect(read.errors).toHaveLength(2); + fs.rmSync(dir, { recursive: true, force: true }); + }); + + test('legacy records attribute by shard suffix, id, label or single-owner file, else stay unattributed', () => { + expect(attributeLegacyRecord('/anything', 'skill-e2e-b--beh-b', registry)).toBe('beh-b'); + expect(attributeLegacyRecord('rule-a', 'skill-e2e-zzz', registry)).toBe('rule-a'); + expect(attributeLegacyRecord('/Rule a', 'skill-e2e-zzz', registry)).toBe('rule-a'); + expect(attributeLegacyRecord('/rule a extra', 'skill-e2e-zzz', registry)).toBeNull(); + expect(attributeLegacyRecord('/gate c labeled', undefined, registry)).toBe('gate-c'); + expect(attributeLegacyRecord('/a display name', 'skill-e2e-a', registry)).toBe('rule-a'); + expect(attributeLegacyRecord('/shared display', 'skill-e2e-shared', registry)).toBeNull(); + }); + + test('backfill keeps only first attempts, defaults a missing attempt to 1, and labels records pre-policy', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'passrates-backfill-')); + fs.writeFileSync(path.join(dir, 'run.json'), run([ + { ...entry_('rule-a', false, 1), exit_reason: 'timeout' }, entry_('rule-a', true, 2), + { name: 'beh-b', suite: 's', tier: 'e2e', passed: true, duration_ms: 1, cost_usd: 0 }, + entry_('/unknown display', true, 1), + ], { shard: 'skill-e2e-zzz' })); + const { records, unattributed } = backfillEvalFiles(collectEvalFiles(dir), { run_id: '42', sha: 'abc' }, registry); + expect(records.map(r => [r.case, r.outcome, r.failure_class, r.policy_version, r.source, r.run_id])) + .toEqual([['rule-a', 'failed', 'timeout', 0, 'backfill', '42'], ['beh-b', 'passed', undefined, 0, 'backfill', '42']]); + expect(formatTrialOutcomes(records)).toContain(TRIAL_OUTCOME_SCHEMA); + expect(unattributed).toEqual(['/unknown display']); + fs.rmSync(dir, { recursive: true, force: true }); + }); + + test('series identity follows the case\'s own touchfiles, not GLOBAL_TOUCHFILES', () => { + const root = fs.mkdtempSync(path.join(os.tmpdir(), 'passrates-series-')); + const git = (...args: string[]) => spawnSync('git', args, { cwd: root, encoding: 'utf8', timeout: 10_000 }); + for (const [file, body] of [['a/x.ts', '1'], ['b/y.ts', '1'], ['harness/run.ts', '1'], ['test/skill-e2e-a.test.ts', '1']] as const) { + fs.mkdirSync(path.join(root, path.dirname(file)), { recursive: true }); + fs.writeFileSync(path.join(root, file), body); + } + const snapshot = () => { expect(git('add', '-A').status).toBe(0); return caseSeriesIdentities(['rule-a', 'beh-b'], root, registry); }; + expect(git('init', '-q').status).toBe(0); + const first = snapshot(); + expect(first['rule-a']).not.toBe(first['beh-b']); + fs.writeFileSync(path.join(root, 'harness/run.ts'), '2'); + expect(snapshot()).toEqual(first); + fs.writeFileSync(path.join(root, 'a/x.ts'), '2'); + const next = snapshot(); + expect(next['rule-a']).not.toBe(first['rule-a']); + expect(next['beh-b']).toBe(first['beh-b']); + fs.rmSync(root, { recursive: true, force: true }); + }); +}); + +describe('pass-rates history fetch (injected, no network)', () => { + function storedZip(files: Record): Buffer { + const locals: Buffer[] = [], centrals: Buffer[] = []; + let offset = 0; + for (const [name, text] of Object.entries(files)) { + const data = Buffer.from(text), fileName = Buffer.from(name), crc = Bun.hash.crc32(data) >>> 0; + const local = Buffer.alloc(30); local.writeUInt32LE(0x04034b50, 0); local.writeUInt16LE(20, 4); + local.writeUInt32LE(crc, 14); local.writeUInt32LE(data.length, 18); local.writeUInt32LE(data.length, 22); local.writeUInt16LE(fileName.length, 26); + const central = Buffer.alloc(46); central.writeUInt32LE(0x02014b50, 0); central.writeUInt16LE(20, 4); central.writeUInt16LE(20, 6); + central.writeUInt32LE(crc, 16); central.writeUInt32LE(data.length, 20); central.writeUInt32LE(data.length, 24); + central.writeUInt16LE(fileName.length, 28); central.writeUInt32LE(offset, 42); + locals.push(local, fileName, data); centrals.push(central, fileName); + offset += 30 + fileName.length + data.length; + } + const size = centrals.reduce((sum, b) => sum + b.length, 0); + const end = Buffer.alloc(22); end.writeUInt32LE(0x06054b50, 0); end.writeUInt16LE(Object.keys(files).length, 8); + end.writeUInt16LE(Object.keys(files).length, 10); end.writeUInt32LE(size, 12); end.writeUInt32LE(offset, 16); + return Buffer.concat([...locals, ...centrals, end]); + } + + test('lists runs per branch, deduplicated and newest first', () => { + const fetcher: HistoryFetcher = { + listRuns: (_repo, _workflow, branch) => branch === 'main' + ? [{ id: 1, attempt: 1, sha: 'a', branch, createdAt: '2026-09-01T00:00:00Z' }, { id: 3, attempt: 1, sha: 'c', branch, createdAt: '2026-09-15T00:00:00Z' }] + : [{ id: 3, attempt: 1, sha: 'c', branch, createdAt: '2026-09-15T00:00:00Z' }, { id: 2, attempt: 2, sha: 'b', branch, createdAt: '2026-09-08T00:00:00Z' }], + listArtifacts: () => [], downloadZip: () => { throw new Error('unused'); }, + }; + expect(listWeeklyRuns({ repo: 'o/r', workflow: 'evals-periodic.yml', branches: ['feature', 'main'], limit: 10, fetcher }).map(r => r.id)).toEqual([3, 2, 1]); + }); + + test('downloads only matching, bounded artifacts once, and caches them', () => { + const cacheDir = fs.mkdtempSync(path.join(os.tmpdir(), 'passrates-cache-')); + const downloads: number[] = []; + const fetcher: HistoryFetcher = { + listRuns: () => [], + listArtifacts: () => [{ id: 10, name: 'trial-outcomes-gate', size: 100 }, { id: 11, name: 'paid-slice-1', size: 100 }, + { id: 12, name: 'trial-outcomes-huge', size: 10 ** 9 }, { id: 13, name: 'trial-outcomes/../escape', size: 1 }], + downloadZip: (_repo, id, destination) => { downloads.push(id); fs.writeFileSync(destination, storedZip({ 'trial-outcomes.jsonl': formatTrialOutcomes([trial('rule-a', 'passed')]) })); }, + }; + const options = { repo: 'o/r', run: { id: 7, attempt: 1, sha: 's', branch: 'main', createdAt: '' }, cacheDir, fetcher, + match: (name: string) => name.startsWith('trial-outcomes') }; + const dirs = downloadRunArtifacts(options); + expect(downloads).toEqual([10]); + expect(dirs).toHaveLength(1); + expect(readTrialOutcomeDir(dirs[0]!).records.map(r => r.case)).toEqual(['rule-a']); + expect(downloadRunArtifacts(options)).toEqual(dirs); + expect(downloads).toEqual([10]); + fs.rmSync(cacheDir, { recursive: true, force: true }); + }); +}); + +describe('pass-rates CLI', () => { + const cli = (args: string[]) => spawnSync(process.execPath, [path.resolve(import.meta.dir, '../scripts/eval-flake-rank.ts'), ...args], + { encoding: 'utf8', timeout: 20_000 }); + + test('--dir prints per-case pass rates; --gate fails only on ACTION REQUIRED', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'passrates-cli-')); + const id = 'plan-ceo-review-format-mode'; + const records = Array.from({ length: 12 }, (_, i) => ({ ...trial(id, i < 11 ? 'passed' : 'failed'), kind: 'behavior' as const, tier: 'periodic' })); + fs.writeFileSync(path.join(dir, 'trial-outcomes.jsonl'), formatTrialOutcomes(records)); + const shown = cli(['--dir', dir, '--case', id]); + expect(shown.status, shown.stderr).toBe(0); + expect(shown.stdout).toContain(`11/12 [`); + expect(shown.stdout).toMatch(new RegExp(`FLAKY\\s+behavior\\s+periodic.*${id}`)); + expect(shown.stdout).toContain('ACTION REQUIRED'); + expect(shown.stdout).toContain(`[drift] ${id} passes 11/12`); + expect(cli(['--dir', dir, '--gate']).status).toBe(1); + fs.writeFileSync(path.join(dir, 'trial-outcomes.jsonl'), formatTrialOutcomes(records.slice(0, 11))); + const clean = cli(['--dir', dir, '--gate', '--json']); + expect(clean.status, clean.stdout).toBe(0); + expect(JSON.parse(clean.stdout).cases[0]).toMatchObject({ case: id, label: 'PASSING', current: { passes: 11, trials: 11 } }); + fs.rmSync(dir, { recursive: true, force: true }); + }); +}); + +function entry_(name: string, passed: boolean, attempt: number) { + return { name, suite: 's', tier: 'e2e', passed, attempt, duration_ms: 1000, cost_usd: 0.1 }; +} diff --git a/test/eval-kinds.test.ts b/test/eval-kinds.test.ts new file mode 100644 index 000000000..ea6480f6c --- /dev/null +++ b/test/eval-kinds.test.ts @@ -0,0 +1,107 @@ +/** + * Eval kind registry (E2E_KINDS / BEHAVIOR_WHY in touchfiles-data.ts). The + * kind fixes a case's trial policy before the run, so the registry must cover + * every live case exactly once, every behavior case must name its tolerated + * deviation, and a behavior case must be isolatable as its own trial shard. + * A kind edit must re-select the case in the PR lane (map-diff). + */ +import { describe, expect, test } from 'bun:test'; +import * as fs from 'node:fs'; +import * as path from 'node:path'; + +import { BEHAVIOR_WHY, E2E_KINDS, E2E_TIERS, E2E_TOUCHFILES, LLM_JUDGE_TOUCHFILES } from './helpers/touchfiles-data'; +import { diffTouchfileMapsCore, type TouchfileMaps } from './helpers/test-selection'; +import { CASE_TEST_NAMES, fileCaseRegistration } from '../scripts/test-paid-shards'; +import { isPaidTestFile } from './helpers/paid-test-set'; + +const ROOT = path.resolve(import.meta.dir, '..'); +const KIND_RULE = "Pick the kind by what can make the verdict differ between two runs of the same commit: 'rule' when nothing " + + "stochastic decides it or it checks a contract the product must meet every run (the default); 'behavior' when a live " + + "model choice decides it and a sub-100% per-trial rate is acceptable (add a BEHAVIOR_WHY line); 'judge' when the only " + + 'stochastic step is an LLM judge scoring a fixed input.'; + +const liveIds = [...Object.keys(E2E_TIERS), ...Object.keys(LLM_JUDGE_TOUCHFILES)]; +const behaviorIds = Object.keys(E2E_KINDS).filter(id => E2E_KINDS[id] === 'behavior').sort(); + +describe('E2E_KINDS registry', () => { + test('every live case has exactly one kind and no kind names a dead case', () => { + const missing = liveIds.filter(id => !(id in E2E_KINDS)); + expect(missing.length, missing.length ? `add to E2E_KINDS:\n${missing.map(id => ` '${id}': 'rule', // `).join('\n')}\n${KIND_RULE}` : '').toBe(0); + const unknown = Object.keys(E2E_KINDS).filter(id => !liveIds.includes(id)); + expect(unknown, `E2E_KINDS names ids that are neither E2E_TIERS nor LLM_JUDGE_TOUCHFILES keys`).toEqual([]); + expect(new Set(liveIds).size).toBe(liveIds.length); + }); + + test('kinds are rule, behavior or judge; every LLM-judge entry is judge-kind', () => { + for (const [id, kind] of Object.entries(E2E_KINDS)) expect(['rule', 'behavior', 'judge'], id).toContain(kind); + for (const id of Object.keys(LLM_JUDGE_TOUCHFILES)) expect(E2E_KINDS[id], `${id}: a workflow judge scores a fixed input`).toBe('judge'); + }); + + test('BEHAVIOR_WHY names the tolerance of exactly the behavior cases', () => { + expect(Object.keys(BEHAVIOR_WHY).sort()).toEqual(behaviorIds); + for (const id of behaviorIds) { + expect(BEHAVIOR_WHY[id]!.trim().length, `${id}: BEHAVIOR_WHY must say why an occasional deviation is acceptable`).toBeGreaterThanOrEqual(30); + } + }); + + test('a behavior case is an isolatable trial shard: known literal registration and an exact Bun test name', () => { + for (const id of behaviorIds) { + const files = E2E_TOUCHFILES[id]!.filter(file => /^test\/[^/]+\.test\.ts$/.test(file) && isPaidTestFile(file)); + expect(files.length, `${id}: no paid test file registers it`).toBeGreaterThan(0); + for (const file of files) { + const source = fs.readFileSync(path.join(ROOT, file), 'utf8'); + expect(fileCaseRegistration(file, source).known, `${id}: ${file} has a computed registration; behavior needs a literal one`).toBe(true); + const name = CASE_TEST_NAMES[id] ?? id; + const literal = new RegExp(`\\b(?:test(?:\\.serial|\\.concurrent)?|testIfSelected|testConcurrentIfSelected)\\(\\s*(['"\`])${name.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}\\1`); + expect(literal.test(source), `${id}: ${file} must register the Bun test named '${name}'`).toBe(true); + } + } + }); + + test('the classification is the reviewed one: rule by default, 22 behavior, 25 judge', () => { + const counts = Object.values(E2E_KINDS).reduce>((acc, kind) => ({ ...acc, [kind]: (acc[kind] ?? 0) + 1 }), {}); + expect(counts).toEqual({ rule: liveIds.length - 22 - 25, behavior: 22, judge: 25 }); + // Contract-shaped cases stay rule: ask-before-decide, plan-mode no-writes, + // mandated steps, secrets, and the batching floor never ride a majority. + for (const id of ['plan-ceo-mode-routing', 'plan-eng-multi-finding-batching', 'plan-design-review-plan-mode', + 'plan-eng-review-plan-mode', 'plan-ceo-section-loading', 'setup-gbrain-bad-token', 'qa-only-no-fix', 'review-sql-injection']) { + expect(E2E_KINDS[id], id).toBe('rule'); + } + }); +}); + +describe('kind edits re-select their case (map-diff)', () => { + const base = (): TouchfileMaps => ({ + E2E_TOUCHFILES: { alpha: ['a/**'], beta: ['b/**'] }, + E2E_TIERS: { alpha: 'gate', beta: 'periodic' }, + LLM_JUDGE_TOUCHFILES: { 'judge one': ['j/SKILL.md'] }, + GLOBAL_TOUCHFILES: [], + E2E_KINDS: { alpha: 'rule', beta: 'rule', 'judge one': 'judge' }, + BEHAVIOR_WHY: {}, + }); + + test('a rule -> behavior flip selects exactly that case', () => { + const next = base(); + next.E2E_KINDS = { ...next.E2E_KINDS, beta: 'behavior' }; + next.BEHAVIOR_WHY = { beta: 'tolerated deviation' }; + expect(diffTouchfileMapsCore(base(), next).changedTests).toEqual(['beta']); + }); + + test('a BEHAVIOR_WHY edit alone selects its case', () => { + const old = base(); old.E2E_KINDS!.beta = 'behavior'; old.BEHAVIOR_WHY = { beta: 'one' }; + const next = base(); next.E2E_KINDS!.beta = 'behavior'; next.BEHAVIOR_WHY = { beta: 'two' }; + expect(diffTouchfileMapsCore(old, next).changedTests).toEqual(['beta']); + }); + + test('a base revision without the kind maps selects every key', () => { + const old = base(); delete old.E2E_KINDS; delete old.BEHAVIOR_WHY; + expect(diffTouchfileMapsCore(old, base()).changedTests).toEqual(['alpha', 'beta', 'judge one']); + }); + + test('dropping a kind entry while the case lives on counts as changed, not removed', () => { + const next = base(); delete next.E2E_KINDS!.alpha; + const result = diffTouchfileMapsCore(base(), next); + expect(result.changedTests).toEqual(['alpha']); + expect(result.removedTests).toEqual([]); + }); +}); diff --git a/test/helpers/llm-judge.ts b/test/helpers/llm-judge.ts index 4a0f72a1c..da7811b47 100644 --- a/test/helpers/llm-judge.ts +++ b/test/helpers/llm-judge.ts @@ -23,6 +23,8 @@ export interface JudgeScore { reasoning: string; } +export const JUDGE_SCORE_DIMENSIONS = ['clarity', 'completeness', 'actionability'] as const; + export interface JudgeRefusalEvidence { stop_reason: 'refusal'; response_id: string | null; @@ -196,6 +198,63 @@ export async function callJudge( } } +/** + * Samples per judge panel: EVAL_POLICY.judge.samples, restated here so this + * helper (imported by many paid tests) does not pull the quarantine registry + * into their touchfile closure. test/judge-panel.test.ts pins the two equal. + */ +export const JUDGE_PANEL_SAMPLES = 3; + +/** + * Judge panel (EVAL_POLICY.judge): every `judge`-kind entry draws a fixed number of + * independent samples of the SAME prompt concurrently, inside its unchanged + * JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean + * against the unchanged minimum; boolean fields gate on a strict majority. + * A sample that errors (refusal, truncation, non-JSON, malformed field) fails + * the whole panel and is never resampled. callJudge's 429 backoff happens + * before any model output exists, so it is transport, not a verdict retry. + */ +export async function judgePanel(sample: () => Promise): Promise { + const settled = await Promise.allSettled(Array.from({ length: JUDGE_PANEL_SAMPLES }, () => sample())); + const failures = settled.flatMap((result, index) => result.status === 'rejected' ? [{ index, reason: result.reason }] : []); + if (failures.length === 0) return settled.map(result => (result as PromiseFulfilledResult).value); + const first = failures[0]!; + // A refusal is an unscored panel only when EVERY sample refused; a partial + // refusal beside scored samples is an ordinary failed panel. + if (first.reason instanceof JudgeRefusalError && failures.length < settled.length) { + throw new Error(`Judge panel sample ${first.index + 1} of ${settled.length} failed beside scored samples: ${first.reason.message}`); + } + throw first.reason; +} + +/** Per-dimension mean over a panel; any non-finite sample value fails the panel. */ +export function judgePanelMean(samples: ReadonlyArray>, keys: readonly K[]): Record { + if (samples.length === 0) throw new Error('Judge panel has no samples'); + return Object.fromEntries(keys.map(key => { + const values = samples.map(sample => sample && typeof sample === 'object' ? sample[key] : undefined); + const bad = values.findIndex(value => typeof value !== 'number' || !Number.isFinite(value)); + if (bad !== -1) throw new Error(`Judge panel sample ${bad + 1} has non-numeric ${key}: ${JSON.stringify(values[bad])}`); + return [key, (values as number[]).reduce((sum, value) => sum + value, 0) / values.length]; + })) as Record; +} + +/** Strict majority of a boolean field; any non-boolean sample value fails the panel. */ +export function judgePanelMajority(samples: ReadonlyArray>, key: K): boolean { + if (samples.length === 0) throw new Error('Judge panel has no samples'); + const values = samples.map(sample => sample && typeof sample === 'object' ? sample[key] : undefined); + const bad = values.findIndex(value => typeof value !== 'boolean'); + if (bad !== -1) throw new Error(`Judge panel sample ${bad + 1} has non-boolean ${key}: ${JSON.stringify(values[bad])}`); + return values.filter(value => value === true).length * 2 > values.length; +} + +/** Sample reasoning lines, numbered, for the collector record. */ +export function judgePanelReasoning(samples: ReadonlyArray): string { + return samples.map((sample, index) => { + const reasoning = sample && typeof sample === 'object' ? (sample as { reasoning?: unknown }).reasoning : undefined; + return `[sample ${index + 1}] ${typeof reasoning === 'string' ? reasoning : ''}`; + }).join('\n'); +} + /** * Score documentation quality on clarity/completeness/actionability (1-5). */ diff --git a/test/helpers/periodic-exclude-data.ts b/test/helpers/periodic-exclude-data.ts index 5959f44d7..638c66d18 100644 --- a/test/helpers/periodic-exclude-data.ts +++ b/test/helpers/periodic-exclude-data.ts @@ -72,7 +72,12 @@ export const CASE_CI_EXCLUDE: Record= `exit.rate` over >= * `exit.minTrials`; at most `capFraction` of each tier's * blocking cases; an entry expires after `expiryWeeklyRuns`. - * drift - one-sided Fisher exact alarm between input-identity series. + * judge - a judge case draws `samples` independent samples of one + * prompt concurrently; numeric dimensions gate on the panel + * mean against the unchanged threshold, booleans on a strict + * majority; an erroring sample fails the panel, never resampled. + * drift - one-sided Fisher exact alarm between input-identity series + * (Holm-controlled across the cases tested in one report). * infraRedispatch - a census whose every red verdict is machine-classified * INFRA or INCOMPLETE may be re-dispatched this many times as * a new run; both runs are reported. @@ -86,6 +91,7 @@ export const EVAL_POLICY = { capFraction: 0.10, expiryWeeklyRuns: 8, }, + judge: { samples: 3 }, drift: { fisherAlpha: 0.05, fisherMinPerSide: 6 }, infraRedispatch: 1, } as const; @@ -100,14 +106,18 @@ export const EVAL_POLICY = { * failures (a product defect is never quarantined), and unchanged case * touchfiles in the change that adds it. Pinned by * test/periodic-exclude-policy.test.ts. - * reason - the written diagnosis - * tracking - issue or TODOS pointer - * owner - who removes it - * enteredAt - ISO date the entry landed (expiry counts weekly runs from here) - * exit - the measurable exit condition + * reason - the written diagnosis, with the pass-rate evidence + * failureClass - what the diagnosis found; a product defect has no class here + * tracking - issue or TODOS pointer + * owner - who removes it + * enteredAt - YYYY-MM-DD the entry landed (expiry counts weekly runs from here) + * exit - the measurable exit condition + * At most EVAL_POLICY.quarantine.capFraction of a tier's cases (gate and + * periodic are the blocking tiers) may be quarantined at once. */ export const CASE_QUARANTINE: Record; LLM_JUDGE_TOUCHFILES: Record; GLOBAL_TOUCHFILES: string[]; + /** Absent on base revisions older than the eval-kind registry: every current key then counts as changed. */ + E2E_KINDS?: Record; + BEHAVIOR_WHY?: Record; } export type MapDiffCause = @@ -171,6 +176,8 @@ const CURRENT_MAPS: TouchfileMaps = { E2E_TIERS, LLM_JUDGE_TOUCHFILES, GLOBAL_TOUCHFILES, + E2E_KINDS, + BEHAVIOR_WHY, }; function isStringArray(v: unknown): v is string[] { @@ -193,14 +200,18 @@ function isTouchfileMaps(v: unknown): v is TouchfileMaps { return isRecordOfStringArrays(o.E2E_TOUCHFILES) && isRecordOfStrings(o.E2E_TIERS) && isRecordOfStringArrays(o.LLM_JUDGE_TOUCHFILES) - && isStringArray(o.GLOBAL_TOUCHFILES); + && isStringArray(o.GLOBAL_TOUCHFILES) + && (o.E2E_KINDS === undefined || isRecordOfStrings(o.E2E_KINDS)) + && (o.BEHAVIOR_WHY === undefined || isRecordOfStrings(o.BEHAVIOR_WHY)); } /** * Pure map-diff core (injectable for tests — no git, no filesystem). * * A key counts as CHANGED when it was added to any per-key map, its dep-list - * array differs, or its tier value flipped. A key counts as REMOVED only when + * array differs, or its tier, kind or behavior tolerance changed. A per-key + * map missing on the old side (a base revision older than E2E_KINDS / + * BEHAVIOR_WHY) makes every key of that map count as added. A key counts as REMOVED only when * it is gone from every new per-key map; a key dropped from one map but still * present in another (e.g. tier entry deleted, touchfile entry kept) counts * as changed — conservative, because the test still exists with a different @@ -212,7 +223,7 @@ export function diffTouchfileMapsCore( oldMaps: TouchfileMaps, newMaps: TouchfileMaps, ): { changedTests: string[]; removedTests: string[]; globalTouchfilesChanged: boolean } { - const perKeyMapNames = ['E2E_TOUCHFILES', 'E2E_TIERS', 'LLM_JUDGE_TOUCHFILES'] as const; + const perKeyMapNames = ['E2E_TOUCHFILES', 'E2E_TIERS', 'LLM_JUDGE_TOUCHFILES', 'E2E_KINDS', 'BEHAVIOR_WHY'] as const; const changed = new Set(); const rawRemoved = new Set(); @@ -290,6 +301,8 @@ export function diffTouchfileMaps( ' E2E_TIERS: m.E2E_TIERS,', ' LLM_JUDGE_TOUCHFILES: m.LLM_JUDGE_TOUCHFILES,', ' GLOBAL_TOUCHFILES: m.GLOBAL_TOUCHFILES,', + ' E2E_KINDS: m.E2E_KINDS,', + ' BEHAVIOR_WHY: m.BEHAVIOR_WHY,', '}));', '', ].join('\n')); diff --git a/test/helpers/touchfiles-data.ts b/test/helpers/touchfiles-data.ts index 18fc17f77..860afc9f2 100644 --- a/test/helpers/touchfiles-data.ts +++ b/test/helpers/touchfiles-data.ts @@ -1597,7 +1597,7 @@ export const E2E_KINDS: Record = { 'shared-libs-unsupported-git': 'rule', 'shared-libs-review-lifecycle': 'rule', 'shared-libs-review-revalidation': 'rule', - 'shared-libs-opportunity-judgment': 'rule', + 'shared-libs-opportunity-judgment': 'behavior', 'shared-libs-pr-coverage': 'rule', 'shared-libs-plan-callers': 'rule', 'browse-basic': 'rule', @@ -1634,7 +1634,7 @@ export const E2E_KINDS: Record = { 'review-sql-injection': 'rule', 'review-enum-completeness': 'rule', 'review-base-branch': 'rule', - 'review-design-lite': 'rule', + 'review-design-lite': 'behavior', 'review-coverage-audit': 'rule', 'review-dashboard-via': 'rule', 'review-army-migration-safety': 'rule', @@ -1642,21 +1642,21 @@ export const E2E_KINDS: Record = { 'review-army-delivery-audit': 'rule', 'review-army-quality-score': 'rule', 'review-army-json-findings': 'rule', - 'review-army-red-team': 'rule', - 'review-army-consensus': 'rule', - 'review-army-simplification': 'rule', - 'review-army-simplification-precision': 'rule', + 'review-army-red-team': 'behavior', + 'review-army-consensus': 'behavior', + 'review-army-simplification': 'behavior', + 'review-army-simplification-precision': 'behavior', 'office-hours-spec-review': 'rule', - 'office-hours-brain-writeback': 'rule', + 'office-hours-brain-writeback': 'behavior', 'gbrain-roundtrip-local': 'rule', 'sync-gbrain-read-ready': 'rule', 'sync-gbrain-read-unknown': 'rule', - 'office-hours-forcing-energy': 'rule', - 'office-hours-builder-wildness': 'rule', + 'office-hours-forcing-energy': 'behavior', + 'office-hours-builder-wildness': 'behavior', 'plan-ceo-review': 'rule', 'plan-ceo-review-selective': 'rule', 'plan-ceo-review-benefits': 'rule', - 'plan-ceo-review-expansion-energy': 'rule', + 'plan-ceo-review-expansion-energy': 'behavior', 'plan-eng-review': 'rule', 'plan-eng-review-artifact': 'rule', 'plan-eng-coverage-audit': 'rule', @@ -1688,16 +1688,16 @@ export const E2E_KINDS: Record = { 'setup-gbrain-remote': 'rule', 'setup-gbrain-bad-token': 'rule', 'setup-gbrain-path4-local-pglite': 'rule', - 'plan-ceo-review-format-mode': 'rule', - 'plan-ceo-review-format-approach': 'rule', - 'plan-eng-review-format-coverage': 'rule', - 'plan-eng-review-format-kind': 'rule', - 'office-hours-phase4-fork': 'rule', - 'llm-judge-recommendation': 'rule', - 'plan-ceo-review-prosons-cadence': 'rule', - 'plan-review-prosons-format': 'rule', - 'plan-review-prosons-hardstop-neg': 'rule', - 'plan-review-prosons-neutral-neg': 'rule', + 'plan-ceo-review-format-mode': 'behavior', + 'plan-ceo-review-format-approach': 'behavior', + 'plan-eng-review-format-coverage': 'behavior', + 'plan-eng-review-format-kind': 'behavior', + 'office-hours-phase4-fork': 'behavior', + 'llm-judge-recommendation': 'judge', + 'plan-ceo-review-prosons-cadence': 'behavior', + 'plan-review-prosons-format': 'behavior', + 'plan-review-prosons-hardstop-neg': 'behavior', + 'plan-review-prosons-neutral-neg': 'behavior', 'plan-tune-inspect': 'rule', 'codex-offered-office-hours': 'rule', 'codex-offered-ceo-review': 'rule', @@ -1758,7 +1758,7 @@ export const E2E_KINDS: Record = { 'design-review-detector-shim': 'rule', 'design-review-detector-shim-dom': 'rule', 'design-review-plugin-handoff': 'rule', - 'design-html-slop-gate': 'rule', + 'design-html-slop-gate': 'behavior', 'diagram-triplet': 'rule', 'diagram-authoring-quality': 'rule', 'gstack-upgrade-happy-path': 'rule', @@ -1770,8 +1770,8 @@ export const E2E_KINDS: Record = { 'setup-deploy-workflow': 'rule', 'autoplan-dual-voice': 'rule', 'benchmark-providers-live': 'rule', - 'scrape-match-path': 'rule', - 'scrape-prototype-path': 'rule', + 'scrape-match-path': 'behavior', + 'scrape-prototype-path': 'behavior', 'skillify-happy-path': 'rule', 'skillify-provenance-refusal': 'rule', 'skillify-approval-reject': 'rule', @@ -1799,30 +1799,30 @@ export const E2E_KINDS: Record = { 'overlay-harness-opus-4-7-literal-interpretation': 'rule', 'overlay-harness-claude-dedicated-tools-vs-bash-sonnet': 'rule', 'journey-negatives': 'rule', - 'review/SKILL.md workflow': 'rule', - 'setup-browser-cookies/SKILL.md workflow': 'rule', - 'browse/SKILL.md reference': 'rule', - 'setup block': 'rule', - 'qa/SKILL.md workflow': 'rule', - 'qa/SKILL.md health rubric': 'rule', - 'qa/SKILL.md anti-refusal': 'rule', - 'cross-skill greptile consistency': 'rule', - 'ship/SKILL.md workflow': 'rule', - 'document-release/SKILL.md workflow': 'rule', - 'plan-ceo-review/SKILL.md modes': 'rule', - 'plan-eng-review/SKILL.md sections': 'rule', - 'plan-design-review/SKILL.md passes': 'rule', - 'design-review/SKILL.md fix loop': 'rule', - 'design-consultation/SKILL.md research': 'rule', - 'land-and-deploy/SKILL.md workflow': 'rule', - 'canary/SKILL.md monitoring loop': 'rule', - 'benchmark/SKILL.md perf collection': 'rule', - 'setup-deploy/SKILL.md platform setup': 'rule', - 'retro/SKILL.md instructions': 'rule', - 'qa-only/SKILL.md workflow': 'rule', - 'gstack-upgrade/SKILL.md upgrade flow': 'rule', - 'sync-gbrain/SKILL.md read-only readiness': 'rule', - 'voice directive tone': 'rule', + 'review/SKILL.md workflow': 'judge', + 'setup-browser-cookies/SKILL.md workflow': 'judge', + 'browse/SKILL.md reference': 'judge', + 'setup block': 'judge', + 'qa/SKILL.md workflow': 'judge', + 'qa/SKILL.md health rubric': 'judge', + 'qa/SKILL.md anti-refusal': 'judge', + 'cross-skill greptile consistency': 'judge', + 'ship/SKILL.md workflow': 'judge', + 'document-release/SKILL.md workflow': 'judge', + 'plan-ceo-review/SKILL.md modes': 'judge', + 'plan-eng-review/SKILL.md sections': 'judge', + 'plan-design-review/SKILL.md passes': 'judge', + 'design-review/SKILL.md fix loop': 'judge', + 'design-consultation/SKILL.md research': 'judge', + 'land-and-deploy/SKILL.md workflow': 'judge', + 'canary/SKILL.md monitoring loop': 'judge', + 'benchmark/SKILL.md perf collection': 'judge', + 'setup-deploy/SKILL.md platform setup': 'judge', + 'retro/SKILL.md instructions': 'judge', + 'qa-only/SKILL.md workflow': 'judge', + 'gstack-upgrade/SKILL.md upgrade flow': 'judge', + 'sync-gbrain/SKILL.md read-only readiness': 'judge', + 'voice directive tone': 'judge', }; /** @@ -1830,4 +1830,49 @@ export const E2E_KINDS: Record = { * deviation is acceptable product behavior. Keys equal the behavior ids of * E2E_KINDS; values are non-empty. */ -export const BEHAVIOR_WHY: Record = {}; +export const BEHAVIOR_WHY: Record = { + 'shared-libs-opportunity-judgment': + "Whether a candidate extraction is worth recommending is a judgment call; the read-only invariant stays a contract.", + 'review-design-lite': + "How many of the seven design-lite checklist items the live review flags varies run to run; the fake-engine rows it must carry stay strict.", + 'review-army-red-team': + "Whether the red-team lens surfaces on a small diff is a live model choice, not a contract.", + 'review-army-consensus': + "Multi-specialist agreement on the planted SQL finding is a quality benchmark that tolerates an occasional miss.", + 'review-army-simplification': + "Flagging the planted unnecessary structure is an advisory-lens quality judgment.", + 'review-army-simplification-precision': + "Staying silent on a lean diff is a false-flag noise benchmark; an occasional advisory is acceptable noise.", + 'office-hours-forcing-energy': + "The Q3 posture is scored by a live judge on generated prose; a single flat phrasing is tolerable.", + 'office-hours-builder-wildness': + "Builder-mode creativity is scored by a live judge on generated prose; one conservative riff is tolerable.", + 'office-hours-brain-writeback': + "The model's interpretation of the gbrain writeback instruction (page shape, tags) varies; no secret or safety step rides on it.", + 'office-hours-phase4-fork': + "Phase 4 asks the model to invent 2-3 architectures; surfacing the fork with its reasoning is open-ended generation.", + 'plan-ceo-review-expansion-energy': + "Expansion framing is scored by a live judge on generated proposals; one flat proposal set is tolerable.", + 'plan-ceo-review-format-mode': + "Mode-question wording (Completeness line vs kind note) is live formatting of one AskUserQuestion.", + 'plan-ceo-review-format-approach': + "Approach-menu Completeness wording is live formatting of one AskUserQuestion.", + 'plan-eng-review-format-coverage': + "Coverage-issue Completeness wording is live formatting of one AskUserQuestion.", + 'plan-eng-review-format-kind': + "Kind-note wording is live formatting of one AskUserQuestion.", + 'plan-ceo-review-prosons-cadence': + "Pros/Cons cadence on a hard-stop question is live formatting; either the escape or the full block is accepted.", + 'plan-review-prosons-format': + "The full Pros/Cons block (counts of pros and cons, labels) is live formatting of one question.", + 'plan-review-prosons-hardstop-neg': + "Not using the hard-stop escape on an ordinary decision is live formatting of one question.", + 'plan-review-prosons-neutral-neg': + "Avoiding neutral posture and naming a because-reason is live formatting of one question.", + 'design-html-slop-gate': + "How many scan passes the one-pass slop gate takes on a fake engine's fixed output is a judgment call.", + 'scrape-match-path': + "The /scrape fallback no longer prescribes the browser-skills match flow, so taking it is prompt compliance.", + 'scrape-prototype-path': + "The /scrape fallback no longer prescribes the prototype flow, so taking it is prompt compliance.", +}; diff --git a/test/helpers/workflow-judge-cache.ts b/test/helpers/workflow-judge-cache.ts index 0b421567a..ee9607546 100644 --- a/test/helpers/workflow-judge-cache.ts +++ b/test/helpers/workflow-judge-cache.ts @@ -4,7 +4,7 @@ import * as path from 'node:path'; import { spawnSync } from 'node:child_process'; import { DEFAULT_JUDGE_MAX_TOKENS, resolveEvalModel } from '../../lib/eval-model'; import { JUDGE_MS } from './eval-budgets'; -import type { JudgeScore } from './llm-judge'; +import { JUDGE_PANEL_SAMPLES, JUDGE_SCORE_DIMENSIONS, judgePanelMean, type JudgeScore } from './llm-judge'; import { readWorkflowJudgeInput, buildWorkflowJudgePrompt, WORKFLOW_JUDGE_RESPONSE_SCHEMA, WORKFLOW_JUDGE_REASONING_WORD_LIMIT } from './workflow-judge-input'; import { buildEvalInputIdentity, lookupEvalInputCache, sourceDependencyClosure, storeEvalInputCache, type EvalCacheValue, type EvalInputIdentity, type EvalPassingProof } from '../../scripts/eval-input-cache'; @@ -39,17 +39,28 @@ export function validWorkflowJudgeScore(value: EvalCacheValue, thresholds: Thres || typeof value.reasoning !== 'string' || (structuredResponse && (!value.reasoning.trim() || value.reasoning.trim().split(/\s+/).length >= WORKFLOW_JUDGE_REASONING_WORD_LIMIT))) return false; - return (['clarity', 'completeness', 'actionability'] as const).every(key => + return JUDGE_SCORE_DIMENSIONS.every(key => typeof value[key] === 'number' && Number.isInteger(value[key]) && value[key] >= thresholds[key] && value[key] <= 5); } +const SAMPLE_RANGE: Thresholds = { clarity: 1, completeness: 1, actionability: 1 }; + +/** A complete judge panel: exactly JUDGE_PANEL_SAMPLES of valid samples whose per-dimension mean meets every threshold. */ +export function validWorkflowJudgePanel(value: EvalCacheValue, thresholds: Thresholds, structuredResponse = false): value is { samples: Array } { + if (!value || typeof value !== 'object' || Array.isArray(value) || Object.keys(value).join(',') !== 'samples' + || !Array.isArray(value.samples) || value.samples.length !== JUDGE_PANEL_SAMPLES + || !value.samples.every(sample => validWorkflowJudgeScore(sample, SAMPLE_RANGE, structuredResponse))) return false; + const mean = judgePanelMean(value.samples as JudgeScore[], JUDGE_SCORE_DIMENSIONS); + return JUDGE_SCORE_DIMENSIONS.every(key => mean[key] >= thresholds[key]); +} + export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): { - lookup(): { scores: JudgeScore; reuse: WorkflowJudgeReuse } | null; + lookup(): { samples: JudgeScore[]; reuse: WorkflowJudgeReuse } | null; /** The attempt guard is rechecked after synchronous input/provenance reads. */ - publish(scores: JudgeScore, isActive?: () => boolean): (() => void) | undefined; + publish(samples: JudgeScore[], isActive?: () => boolean): (() => void) | undefined; } { const env = opts.env ?? process.env; - const noCache = { lookup: () => null, publish: (_scores: JudgeScore) => undefined }; + const noCache = { lookup: () => null, publish: (_samples: JudgeScore[]) => undefined }; const pr = Number(env.EVALS_CACHE_PR); // Runtime ID is the immutable CI image manifest, not a mutable image tag. // Nonstandard Node/Bun preload code or custom model endpoints need a separate @@ -74,7 +85,8 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): { files: workflowJudgeDependencies(opts.root, input.files.map(file => file.path)), prompts: { [opts.testName]: prompt }, parameters: { rootPackage, thresholds: opts.thresholds, max_tokens: opts.maxTokens ?? DEFAULT_JUDGE_MAX_TOKENS, temperature: null, budget_ms: JUDGE_MS, - request: opts.stream ? 'messages.stream/user' : 'messages.create/user', retries: 1, + request: opts.stream ? 'messages.stream/user' : 'messages.create/user', retries: 0, + panel: { samples: JUDGE_PANEL_SAMPLES, numeric: 'mean', boolean: 'majority' }, ...(opts.stream ? { stream: true } : {}), ...(opts.structuredResponse ? { output_config: { format: { type: 'json_schema', schema: WORKFLOW_JUDGE_RESPONSE_SCHEMA } }, response_validation: { reasoning_words_below: WORKFLOW_JUDGE_REASONING_WORD_LIMIT } } : {}) }, @@ -94,14 +106,15 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): { return { lookup() { const result = lookupEvalInputCache({ ...common, identity: before, - validateResult: value => validWorkflowJudgeScore(value, opts.thresholds, opts.structuredResponse) }); + validateResult: value => validWorkflowJudgePanel(value, opts.thresholds, opts.structuredResponse) }); return result.status === 'reused' - ? { scores: result.result as JudgeScore, reuse: { key: result.key, source: result.source } } : null; + ? { samples: (result.result as unknown as { samples: JudgeScore[] }).samples, reuse: { key: result.key, source: result.source } } : null; }, - publish(scores, isActive = () => true) { + publish(samples, isActive = () => true) { // Caller reaches here ONLY after its actual assertions passed. A later // failed case in the file does not erase this independently completed case. - if (!isActive() || !validWorkflowJudgeScore(scores as unknown as EvalCacheValue, opts.thresholds, opts.structuredResponse)) return; + const panel = { samples: samples.map(({ clarity, completeness, actionability, reasoning }) => ({ clarity, completeness, actionability, reasoning })) }; + if (!isActive() || !validWorkflowJudgePanel(panel as unknown as EvalCacheValue, opts.thresholds, opts.structuredResponse)) return; const after = currentIdentity(); const runId = env.GITHUB_RUN_ID ? `${env.GITHUB_RUN_ID}/${env.GITHUB_RUN_ATTEMPT ?? '1'}` : env.EVALS_RUN_ID; if (!after || !runId || !isActive()) return; @@ -112,7 +125,7 @@ export function prepareWorkflowJudgeCache(opts: WorkflowCacheOptions): { cancelled: false, skipped: 0, failed: 0, passed: 1, cases: [{ id: opts.testName, outcome: 'passed', attempt: 1 }], source: { runId, revision: revision.stdout.trim(), completedAt: Date.now() }, - result: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability, reasoning: scores.reasoning }, + result: panel, } }); // A slow synchronous write can consume the recording allowance. The // caller withdraws this new receipt if its final deadline check fails. diff --git a/test/periodic-exclude-policy.test.ts b/test/periodic-exclude-policy.test.ts index 8e4946507..d1a3eeb95 100644 --- a/test/periodic-exclude-policy.test.ts +++ b/test/periodic-exclude-policy.test.ts @@ -9,10 +9,11 @@ import { describe, expect, test } from 'bun:test'; import * as fs from 'node:fs'; import * as path from 'node:path'; -import { CASE_CI_EXCLUDE, PERIODIC_CI_EXCLUDE } from './helpers/periodic-exclude-data'; +import { CASE_CI_EXCLUDE, CASE_QUARANTINE, EVAL_POLICY, PERIODIC_CI_EXCLUDE } from './helpers/periodic-exclude-data'; import { E2E_TOUCHFILES } from './helpers/touchfiles'; +import { quarantinePolicyProblems } from '../scripts/eval-flake-rank'; import { isPaidTestFile } from './helpers/paid-test-set'; -import { buildRunManifest, CASE_SHARDED_FILES, expandCaseShards, partitionCaseExclusions, selectPaidTestFiles, shardCaseId, shardFile } from '../scripts/test-paid-shards'; +import { buildRunManifest, CASE_SHARDED_FILES, expandCaseShards, fileCaseRegistration, partitionCaseExclusions, selectPaidTestFiles, shardCaseId, shardFile } from '../scripts/test-paid-shards'; const ROOT = path.resolve(__dirname, '..'); @@ -72,3 +73,30 @@ describe('periodic exclude policy', () => { } }); }); + +describe('eval verdict policy (pre-registered)', () => { + test('EVAL_POLICY carries exactly the approved constants; a change needs re-approval and a version bump', () => { + expect(EVAL_POLICY).toEqual({ + version: 1, + panel: { n: 3, k: 2 }, + quarantine: { entry: { rate: 0.95, minTrials: 10 }, exit: { rate: 0.97, minTrials: 10 }, capFraction: 0.10, expiryWeeklyRuns: 8 }, + judge: { samples: 3 }, + drift: { fisherAlpha: 0.05, fisherMinPerSide: 6 }, + infraRedispatch: 1, + }); + }); + + test('every CASE_QUARANTINE entry is a diagnosed, dated, non-product blocking case within the tier cap', () => { + expect(quarantinePolicyProblems(CASE_QUARANTINE).map(problem => problem.message)).toEqual([]); + }); + + test('a quarantined case runs as isolated trial shards: its files register it literally', () => { + for (const id of Object.keys(CASE_QUARANTINE)) { + const files = (E2E_TOUCHFILES[id] ?? []).filter(file => /^test\/[^/]+\.test\.ts$/.test(file) && isPaidTestFile(file)); + expect(files.length, `${id}: no paid file registers it`).toBeGreaterThan(0); + for (const file of files) { + expect(fileCaseRegistration(file, fs.readFileSync(path.join(ROOT, file), 'utf8')).known, `${id}: ${file} registration must be statically known`).toBe(true); + } + } + }); +}); diff --git a/test/skill-llm-eval.test.ts b/test/skill-llm-eval.test.ts index 41b5ab13c..af9d1774c 100644 --- a/test/skill-llm-eval.test.ts +++ b/test/skill-llm-eval.test.ts @@ -14,7 +14,7 @@ import { afterAll, expect } from 'bun:test'; import { JUDGE_MS } from './helpers/eval-budgets'; import * as fs from 'fs'; import * as path from 'path'; -import { callJudge, judge, JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS } from './helpers/llm-judge'; +import { callJudge, judge, JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS, judgePanel, judgePanelMean, judgePanelMajority, judgePanelReasoning, JUDGE_SCORE_DIMENSIONS } from './helpers/llm-judge'; import { ENG_REVIEW_EXCERPT } from './helpers/workflow-excerpt'; import type { JudgeScore } from './helpers/llm-judge'; import { readWorkflowJudgeInput, buildWorkflowJudgePrompt, QA_DISCOVERY_REFERENCES, WORKFLOW_JUDGE_RESPONSE_SCHEMA, type WorkflowJudgeInput } from './helpers/workflow-judge-input'; @@ -100,8 +100,9 @@ describeIfSelected('LLM-as-judge quality evals', [ // rewrites the pin). const section = sliceBrowseSection('## Snapshot Flags'); - const scores = await judge('browse skill reference (flags + commands)', section); - console.log('Browse SKILL.md scores:', JSON.stringify(scores, null, 2)); + const samples = await judgePanel(() => judge('browse skill reference (flags + commands)', section)); + const scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS); + console.log('Browse SKILL.md panel:', JSON.stringify({ mean: scores, samples }, null, 2)); const baselinesPath = path.join(ROOT, 'test', 'fixtures', 'eval-baselines.json'); const baselines = JSON.parse(fs.readFileSync(baselinesPath, 'utf-8')); @@ -120,9 +121,9 @@ describeIfSelected('LLM-as-judge quality evals', [ tier: 'llm-judge', passed: scores.clarity >= 3 && scores.completeness >= 4 && scores.actionability >= 4 && regressions.length === 0, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability }, - judge_reasoning: regressions.length ? `${scores.reasoning} | ${regressions.join('; ')}` : scores.reasoning, + judge_reasoning: regressions.length ? `${judgePanelReasoning(samples)} | ${regressions.join('; ')}` : judgePanelReasoning(samples), }); expect(scores.clarity).toBeGreaterThanOrEqual(3); @@ -144,8 +145,9 @@ describeIfSelected('LLM-as-judge quality evals', [ if (setupStart < 0 || setupEnd < 0) throw new Error('browse/SKILL.md: setup block not found — regenerate with: bun run gen:skill-docs'); const section = content.slice(setupStart, setupEnd); - const scores = await judge('setup/binary discovery instructions', section); - console.log('Setup block scores:', JSON.stringify(scores, null, 2)); + const samples = await judgePanel(() => judge('setup/binary discovery instructions', section)); + const scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS); + console.log('Setup block panel:', JSON.stringify({ mean: scores, samples }, null, 2)); evalCollector?.addTest({ name: 'setup block', @@ -153,9 +155,9 @@ describeIfSelected('LLM-as-judge quality evals', [ tier: 'llm-judge', passed: scores.actionability >= 3 && scores.clarity >= 3, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability }, - judge_reasoning: scores.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); // Setup block is intentionally minimal (binary discovery only). @@ -203,7 +205,7 @@ describeIfSelected('QA skill quality evals', ['qa/SKILL.md workflow', 'qa/SKILL. startMarker: '# /qa: Test', endMarker: null, references: ['qa/templates/functional-report-template.md'] }).text; - const scores = await callJudge(`You are evaluating the quality of a QA testing workflow document for an AI coding agent. + const samples = await judgePanel(() => callJudge(`You are evaluating the quality of a QA testing workflow document for an AI coding agent. The agent reads this source-file bundle to select browser, native functional or mixed surfaces, explore with bounded probes, reproduce and diagnose defects, add a regression @@ -222,8 +224,9 @@ Respond with ONLY valid JSON: Here is the QA workflow to evaluate: -${section}`); - console.log('QA workflow scores:', JSON.stringify(scores, null, 2)); +${section}`)); + const scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS); + console.log('QA workflow panel:', JSON.stringify({ mean: scores, samples }, null, 2)); evalCollector?.addTest({ name: 'qa/SKILL.md workflow', @@ -231,9 +234,9 @@ ${section}`); tier: 'llm-judge', passed: scores.clarity >= 3 && scores.completeness >= 3 && scores.actionability >= 4, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability }, - judge_reasoning: scores.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); expect(scores.clarity).toBeGreaterThanOrEqual(3); @@ -247,7 +250,7 @@ ${section}`); const t0 = Date.now(); const section = sliceQaPatterns('## Health Score Rubric'); - const scores = await callJudge(`You are evaluating a health score rubric that an AI agent must follow to compute a numeric QA score. + const samples = await judgePanel(() => callJudge(`You are evaluating a health score rubric that an AI agent must follow to compute a numeric QA score. The agent uses this rubric after QA testing a website. It needs to: 1. Understand each scoring category and what counts as a deduction @@ -264,8 +267,9 @@ Respond with ONLY valid JSON: Here is the rubric to evaluate: -${section}`); - console.log('QA health rubric scores:', JSON.stringify(scores, null, 2)); +${section}`)); + const scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS); + console.log('QA health rubric panel:', JSON.stringify({ mean: scores, samples }, null, 2)); evalCollector?.addTest({ name: 'qa/SKILL.md health rubric', @@ -273,9 +277,9 @@ ${section}`); tier: 'llm-judge', passed: scores.clarity >= 3 && scores.completeness >= 3 && scores.actionability >= 4, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability }, - judge_reasoning: scores.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); expect(scores.clarity).toBeGreaterThanOrEqual(3); @@ -294,7 +298,7 @@ ${section}`); const diffAwareSection = sliceQaPatterns('### Diff-aware', '### Full'); const rulesSection = sliceQaPatterns('## Important Rules'); - const result = await callJudge<{ would_browse: boolean; fallback_behavior: string; confidence: number; reasoning: string }>(`You are evaluating whether a QA testing skill document would cause an AI agent to USE THE BROWSER or REFUSE to use the browser in a specific scenario. + const samples = await judgePanel(() => callJudge<{ would_browse: boolean; fallback_behavior: string; confidence: number; reasoning: string }>(`You are evaluating whether a QA testing skill document would cause an AI agent to USE THE BROWSER or REFUSE to use the browser in a specific scenario. SCENARIO: A user runs /qa (a browser-based QA testing skill). The branch diff shows ONLY prompt template files and config file changes — no routes, views, controllers, components, or CSS were changed. The changes are "purely backend" with no obvious UI surface. @@ -318,9 +322,10 @@ Respond with ONLY valid JSON: Rules: - would_browse should be true if the document instructs the agent to always use the browser regardless of diff content - would_browse should be false if the document allows the agent to skip browser testing for non-UI changes -- confidence: 5 = document is unambiguous, 1 = document is unclear or contradictory`); +- confidence: 5 = document is unambiguous, 1 = document is unclear or contradictory`)); + const result = { would_browse: judgePanelMajority(samples, 'would_browse'), ...judgePanelMean(samples, ['confidence'] as const) }; - console.log('QA anti-refusal result:', JSON.stringify(result, null, 2)); + console.log('QA anti-refusal panel:', JSON.stringify({ result, samples }, null, 2)); evalCollector?.addTest({ name: 'qa/SKILL.md anti-refusal', @@ -328,9 +333,9 @@ Rules: tier: 'llm-judge', passed: result.would_browse === true && result.confidence >= 4, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { would_browse: result.would_browse ? 1 : 0, confidence: result.confidence }, - judge_reasoning: result.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); expect(result.would_browse).toBe(true); @@ -362,7 +367,7 @@ describeIfSelected('Cross-skill consistency evals', ['cross-skill greptile consi extractGrepLines(retroContent, 'retro/SKILL.md'), ].join('\n\n'); - const result = await callJudge<{ consistent: boolean; issues: string[]; score: number; reasoning: string }>(`You are evaluating whether multiple skill configuration files implement the same data architecture consistently. + const samples = await judgePanel(() => callJudge<{ consistent: boolean; issues: string[]; score: number; reasoning: string }>(`You are evaluating whether multiple skill configuration files implement the same data architecture consistently. INTENDED ARCHITECTURE: - greptile-history has TWO paths: per-project (~/.gstack/projects/{slug}/greptile-history.md) and global (~/.gstack/greptile-history.md) @@ -383,9 +388,10 @@ Evaluate consistency. Respond with ONLY valid JSON: "reasoning": "brief explanation" } -score (1-5): 5 = perfectly consistent, 1 = contradictory`); +score (1-5): 5 = perfectly consistent, 1 = contradictory`)); + const result = { consistent: judgePanelMajority(samples, 'consistent'), ...judgePanelMean(samples, ['score'] as const) }; - console.log('Cross-skill consistency:', JSON.stringify(result, null, 2)); + console.log('Cross-skill consistency panel:', JSON.stringify({ result, samples }, null, 2)); evalCollector?.addTest({ name: 'cross-skill greptile consistency', @@ -393,9 +399,9 @@ score (1-5): 5 = perfectly consistent, 1 = contradictory`); tier: 'llm-judge', passed: result.consistent && result.score >= 4, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { consistency_score: result.score }, - judge_reasoning: result.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); expect(result.consistent).toBe(true); @@ -439,7 +445,8 @@ async function runWorkflowJudge(opts: { const workDeadline = started + JUDGE_MS; let stage: 'input' | 'judge' | 'validation' | 'recording' = 'input'; let finalized = false; - let scores: JudgeScore | undefined; + let samples: JudgeScore[] | undefined; + let scores: Record | undefined; let manualReview: ManualJudgeReview | undefined; let customInputMetadata: { prompt: string; model: string } | undefined; let reused: ReturnType['lookup']> = null; @@ -458,19 +465,19 @@ async function runWorkflowJudge(opts: { evalCollector?.addTest({ name: opts.testName, suite: opts.suite, tier: 'llm-judge', passed, attempt, duration_ms: Math.max(0, performance.now() - started), - cost_usd: reused || !scores ? 0 : 0.02, + cost_usd: reused || !samples ? 0 : 0.02 * samples.length, execution: reused ? 'reused' : 'executed', ...customInputMetadata, ...(manualReview ? { manual_review: manualReview } : {}), ...(reused ? { reused_from: { input_key: reused.reuse.key, run_id: reused.reuse.source.runId, revision: reused.reuse.source.revision, completed_at: new Date(reused.reuse.source.completedAt).toISOString() } } : {}), - ...(scores ? { judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability }, - judge_reasoning: scores.reasoning } : {}), + ...(scores ? { judge_scores: { clarity: scores.clarity, completeness: scores.completeness, actionability: scores.actionability } } : {}), + ...(samples ? { judge_reasoning: judgePanelReasoning(samples) } : {}), ...(passed ? {} : { exit_reason: error instanceof JudgeRefusalError ? 'provider_refusal' : error instanceof Error && error.name === 'WorkflowJudgeDeadline' ? 'timeout' : error instanceof Error && error.name === 'WorkflowJudgeSuperseded' ? 'cancelled' : stage === 'validation' ? 'validation_failed' : 'harness_error', - error: `${error instanceof Error ? error.message : String(error)}${scores ? '' : error instanceof JudgeRefusalError + error: `${error instanceof Error ? error.message : String(error)}${samples ? '' : error instanceof JudgeRefusalError ? '\nNo automated score; provider refusal usage retained when manually accepted; cost unavailable.' : '\nNo completed model response; cost and usage unavailable.'}` }), }); @@ -508,11 +515,11 @@ async function runWorkflowJudge(opts: { checkActive(); stage = 'judge'; const maxTokens = opts.maxTokens ?? DEFAULT_JUDGE_MAX_TOKENS; - let result: JudgeScore; + let result: JudgeScore[]; try { - result = reused?.scores ?? await callJudge(prompt, opts.model, { signal: controller.signal, max_tokens: maxTokens, + result = reused?.samples ?? await judgePanel(() => callJudge(prompt, opts.model, { signal: controller.signal, max_tokens: maxTokens, ...(opts.stream ? { stream: true } : {}), - ...(opts.structuredResponse ? { jsonSchema: WORKFLOW_JUDGE_RESPONSE_SCHEMA } : {}) }); + ...(opts.structuredResponse ? { jsonSchema: WORKFLOW_JUDGE_RESPONSE_SCHEMA } : {}) })); } catch (error) { checkActive(); if (error instanceof JudgeRefusalError && customInputMetadata) { @@ -529,20 +536,21 @@ async function runWorkflowJudge(opts: { throw error; } checkActive(); - scores = result; + samples = result; console.log(`[workflow-judge] ${opts.testName}: ${reused ? `reused ${reused.reuse.source.runId} @ ${reused.reuse.source.revision} (${new Date(reused.reuse.source.completedAt).toISOString()})` : 'executed'}`); - console.log(`${opts.testName} scores:`, JSON.stringify(scores, null, 2)); stage = 'validation'; - if (opts.structuredResponse && !validWorkflowJudgeScore(scores as unknown as EvalCacheValue, { clarity: 1, completeness: 1, actionability: 1 }, true)) { + if (opts.structuredResponse && !samples.every(sample => validWorkflowJudgeScore(sample as unknown as EvalCacheValue, { clarity: 1, completeness: 1, actionability: 1 }, true))) { throw new Error('Structured workflow judge violated the response schema'); } + scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS); + console.log(`${opts.testName} panel:`, JSON.stringify({ mean: scores, samples }, null, 2)); expect(scores.clarity).toBeGreaterThanOrEqual(thresholds.clarity); expect(scores.completeness).toBeGreaterThanOrEqual(thresholds.completeness); expect(scores.actionability).toBeGreaterThanOrEqual(thresholds.actionability); checkActive(); stage = 'recording'; arm(); - const discardReceipt = reused ? undefined : cache.publish(scores, active); + const discardReceipt = reused ? undefined : cache.publish(samples, active); try { checkActive(); finish(true); } catch (error) { discardReceipt?.(); throw error; } }; @@ -792,7 +800,7 @@ describeIfSelected('Voice directive eval', ['voice directive tone'], () => { const voiceEnd = content.indexOf('\n## ', voiceStart + 1); const voiceSection = content.slice(voiceStart, voiceEnd > 0 ? voiceEnd : voiceStart + 3000); - const result = await callJudge<{ + const samples = await judgePanel(() => callJudge<{ directness: number; concreteness: number; avoids_corporate: number; @@ -812,9 +820,10 @@ Return JSON only: {"directness": N, "concreteness": N, "avoids_corporate": N, "avoids_ai_vocabulary": N, "connects_user_outcomes": N, "reasoning": "..."} THE VOICE DIRECTIVE: -${voiceSection}`); +${voiceSection}`)); + const result = judgePanelMean(samples, ['directness', 'concreteness', 'avoids_corporate', 'avoids_ai_vocabulary', 'connects_user_outcomes'] as const); - console.log('Voice directive scores:', JSON.stringify(result, null, 2)); + console.log('Voice directive panel:', JSON.stringify({ mean: result, samples }, null, 2)); evalCollector?.addTest({ name: 'voice directive tone', @@ -823,7 +832,7 @@ ${voiceSection}`); passed: result.directness >= 4 && result.concreteness >= 4 && result.avoids_corporate >= 4 && result.avoids_ai_vocabulary >= 4 && result.connects_user_outcomes >= 4, duration_ms: Date.now() - t0, - cost_usd: 0.02, + cost_usd: 0.02 * samples.length, judge_scores: { directness: result.directness, concreteness: result.concreteness, @@ -831,7 +840,7 @@ ${voiceSection}`); avoids_ai_vocabulary: result.avoids_ai_vocabulary, connects_user_outcomes: result.connects_user_outcomes, }, - judge_reasoning: result.reasoning, + judge_reasoning: judgePanelReasoning(samples), }); expect(result.directness).toBeGreaterThanOrEqual(4); diff --git a/test/workflow-judge-cache.test.ts b/test/workflow-judge-cache.test.ts index 79feebcdd..1d55fad33 100644 --- a/test/workflow-judge-cache.test.ts +++ b/test/workflow-judge-cache.test.ts @@ -1,18 +1,22 @@ -import { afterEach, expect, spyOn, test } from 'bun:test'; +import { afterEach, describe, expect, spyOn, test } from 'bun:test'; import { Messages } from '@anthropic-ai/sdk/resources/messages'; -import { callJudge, JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS } from './helpers/llm-judge'; +import { callJudge, JudgeRefusalError, DEFAULT_JUDGE_MAX_TOKENS, judgePanel, judgePanelMajority, judgePanelMean, judgePanelReasoning, JUDGE_SCORE_DIMENSIONS, JUDGE_PANEL_SAMPLES } from './helpers/llm-judge'; +import { EVAL_POLICY } from './helpers/periodic-exclude-data'; import { getCookieWorkflowManualReview } from './helpers/cookie-workflow-manual-review'; import { resolveEvalModel } from '../lib/eval-model'; import * as fs from 'node:fs'; import * as os from 'node:os'; import * as path from 'node:path'; import { execFileSync } from 'node:child_process'; -import { prepareWorkflowJudgeCache, validWorkflowJudgeScore, workflowJudgeDependencies, type WorkflowCacheOptions } from './helpers/workflow-judge-cache'; +import { prepareWorkflowJudgeCache, validWorkflowJudgePanel, validWorkflowJudgeScore, workflowJudgeDependencies, type WorkflowCacheOptions } from './helpers/workflow-judge-cache'; import { readWorkflowJudgeInput, buildWorkflowJudgePrompt, QA_DISCOVERY_REFERENCES, WORKFLOW_JUDGE_RESPONSE_SCHEMA } from './helpers/workflow-judge-input'; const roots: string[] = []; afterEach(() => { for (const root of roots.splice(0)) fs.rmSync(root, { recursive: true, force: true }); }); const scores = { clarity: 4, completeness: 5, actionability: 4, reasoning: 'Concrete steps' }; +const SAMPLES = JUDGE_PANEL_SAMPLES; +const panelOf = (sample: typeof scores) => Array.from({ length: SAMPLES }, () => sample); +const panel = panelOf(scores); function fixture() { const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-judge-cache-')); roots.push(root); const files = { @@ -53,9 +57,9 @@ function fixture() { } test('the audited adapter reuses only the exact completed score and original provenance', () => { - const f = fixture(); const first = f.cache(); expect(first.lookup()).toBeNull(); first.publish(scores); + const f = fixture(); const first = f.cache(); expect(first.lookup()).toBeNull(); first.publish(panel); expect(f.entries()).toHaveLength(1); - const reused = f.cache().lookup(); expect(reused?.scores).toEqual(scores); + const reused = f.cache().lookup(); expect(reused?.samples).toEqual(panel); expect(reused?.reuse.source.runId).toBe('free-cache-test'); expect(reused?.reuse.source.revision).toMatch(/^[a-f0-9]{40}$/); expect(reused?.reuse.source.completedAt).toBeLessThanOrEqual(Date.now()); @@ -73,11 +77,11 @@ test('the dependency closure includes actual installed SDK bytes and local trans }); test('release-label changes preserve reuse; other package semantics invalidate it', () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); const file = path.join(f.root, 'package.json'); const original = JSON.parse(fs.readFileSync(file, 'utf8')); fs.writeFileSync(file, JSON.stringify({ ...original, version: '2.0.0' }, null, 2)); - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); for (const change of [{ scripts: { 'test:gate': 'changed command' } }, { dependencies: { 'some-sdk': '2.0.0' } }]) { fs.writeFileSync(file, JSON.stringify({ ...original, ...change, version: '2.0.0' })); expect(f.cache().lookup()).toBeNull(); @@ -89,7 +93,7 @@ for (const file of ['test/helpers/nested.ts', 'node_modules/@anthropic-ai/sdk/in 'scripts/test-paid-shards.ts', 'scripts/test-strict-output.ts', 'scripts/eval-select.ts', 'scripts/test-pr-profile.ts', '.github/workflows/evals.yml']) { test(`changes in ${file} require new evaluation`, () => { - const f = fixture(); f.cache().publish(scores); const target = path.join(f.root, file); + const f = fixture(); f.cache().publish(panel); const target = path.join(f.root, file); fs.appendFileSync(target, file.endsWith('.json') ? ' ' : '\n// changed'); f.refreshPrompt(); expect(f.cache().lookup()).toBeNull(); }); @@ -98,8 +102,8 @@ for (const file of ['test/helpers/nested.ts', 'node_modules/@anthropic-ai/sdk/in test('changed sources during an attempt and mismatched actual prompt cannot publish', () => { const f = fixture(); const before = f.cache(); fs.appendFileSync(path.join(f.root, 'example/sections/review.md'), 'new finding'); - before.publish(scores); expect(f.entries()).toHaveLength(0); - f.refreshPrompt(); f.opts.prompt += ' hidden new request'; f.cache().publish(scores); + before.publish(panel); expect(f.entries()).toHaveLength(0); + f.refreshPrompt(); f.opts.prompt += ' hidden new request'; f.cache().publish(panel); expect(f.entries()).toHaveLength(0); }); @@ -108,52 +112,52 @@ for (const [key, value] of Object.entries({ EVALS_FRESH: '1', EVALS_TIER: 'perio EVALS_CACHE_REPOSITORY: '', NODE_OPTIONS: '--require=unknown', BUN_OPTIONS: '--preload=unknown', ANTHROPIC_BASE_URL: 'https://custom-provider.example.test' })) { test(`${key}=${value} is fresh or ineligible`, () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); f.opts.env = { ...f.env, [key]: value }; const cache = f.cache(); - expect(cache.lookup()).toBeNull(); cache.publish(scores); expect(f.entries()).toHaveLength(1); + expect(cache.lookup()).toBeNull(); cache.publish(panel); expect(f.entries()).toHaveLength(1); }); } test('runtime/model/threshold changes miss, and retries never reuse or publish', () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); for (const overrides of [{ GSTACK_EVAL_MODEL_JUDGE: 'different-model' }, { EVALS_CACHE_RUNTIME_ID: 'c'.repeat(64) }]) { f.opts.env = { ...f.env, ...overrides }; expect(f.cache().lookup()).toBeNull(); } f.opts.env = f.env; f.opts.thresholds.clarity = 5; expect(f.cache().lookup()).toBeNull(); f.opts.thresholds.clarity = 4; f.opts.attempt = 2; const retry = f.cache(); - expect(retry.lookup()).toBeNull(); retry.publish(scores); expect(f.entries()).toHaveLength(1); + expect(retry.lookup()).toBeNull(); retry.publish(panel); expect(f.entries()).toHaveLength(1); }); test('frontier reader calibration cannot reuse a score from the unspecified-reader rubric', () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); const original = f.opts.prompt; f.opts.agentCapability = 'frontier'; f.refreshPrompt(); expect(f.opts.prompt).not.toBe(original); expect(f.cache().lookup()).toBeNull(); - f.cache().publish(scores); + f.cache().publish(panel); expect(f.entries()).toHaveLength(2); - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); delete f.opts.agentCapability; f.refreshPrompt(); expect(f.opts.prompt).toBe(original); - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); }); test('a pinned workflow judge model overrides the global model and changes the cache identity', () => { const f = fixture(); f.opts.model = 'claude-sonnet-4-6'; - f.cache().publish(scores); + f.cache().publish(panel); expect(f.entries()).toHaveLength(1); f.opts.env = { ...f.env, GSTACK_EVAL_MODEL_JUDGE: 'different-global-model' }; - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); f.opts.model = 'claude-opus-4-7'; expect(f.cache().lookup()).toBeNull(); }); test('failed assertions, missing provenance, and missing imported dependencies cannot supply a receipt', () => { - const f = fixture(); f.cache().publish({ ...scores, clarity: 3 }); expect(f.entries()).toHaveLength(0); - f.opts.env = { ...f.env, EVALS_RUN_ID: '' }; f.cache().publish(scores); expect(f.entries()).toHaveLength(0); + const f = fixture(); f.cache().publish(panelOf({ ...scores, clarity: 3 })); expect(f.entries()).toHaveLength(0); + f.opts.env = { ...f.env, EVALS_RUN_ID: '' }; f.cache().publish(panel); expect(f.entries()).toHaveLength(0); f.opts.env = f.env; fs.unlinkSync(path.join(f.root, 'test/helpers/nested.ts')); - f.cache().publish(scores); expect(f.entries()).toHaveLength(0); + f.cache().publish(panel); expect(f.entries()).toHaveLength(0); }); test('cached payload schema remains small and cannot carry operational fields', () => { @@ -167,8 +171,9 @@ test('workflow registration preserves model work and reserves only terminal-reco const source = fs.readFileSync(path.join(import.meta.dir, 'skill-llm-eval.test.ts'), 'utf8'); const body = source.split('async function runWorkflowJudge')[1]!.split('// Block 1:')[0]!; const stages = ['workflowJudgeAttempts.set', 'readWorkflowJudgeInput(', 'cache.lookup()', - 'callJudge(prompt, opts.model, { signal: controller.signal, max_tokens: maxTokens,', - 'expect(scores.clarity)', 'expect(scores.completeness)', 'expect(scores.actionability)', 'cache.publish(scores, active)'] + 'judgePanel(() => callJudge(prompt, opts.model, { signal: controller.signal, max_tokens: maxTokens,', + 'scores = judgePanelMean(samples, JUDGE_SCORE_DIMENSIONS);', + 'expect(scores.clarity)', 'expect(scores.completeness)', 'expect(scores.actionability)', 'cache.publish(samples, active)'] .map(stage => body.indexOf(stage)); expect(stages.every(position => position >= 0)).toBe(true); expect(stages).toEqual([...stages].sort((a, b) => a - b)); @@ -203,6 +208,7 @@ function actualCallback(f: ReturnType, overrides: { 'evalCollector', 'expect', 'console', 'performance', 'JUDGE_MS', 'WORKFLOW_JUDGE_RECORD_MS', 'setTimeout', 'clearTimeout', 'JudgeRefusalError', 'getCookieWorkflowManualReview', 'DEFAULT_JUDGE_MAX_TOKENS', 'resolveEvalModel', 'WORKFLOW_JUDGE_RESPONSE_SCHEMA', 'validWorkflowJudgeScore', + 'judgePanel', 'judgePanelMean', 'judgePanelReasoning', 'JUDGE_SCORE_DIMENSIONS', `${javascript}\nreturn runWorkflowJudge;`)( f.root, overrides.read ?? readWorkflowJudgeInput, buildWorkflowJudgePrompt, (options: WorkflowCacheOptions) => (overrides.prepare ?? prepareWorkflowJudgeCache)({ ...options, env: f.env }), @@ -213,7 +219,8 @@ function actualCallback(f: ReturnType, overrides: { overrides.clock ? { now: overrides.clock } : performance, overrides.budget ?? 120_000, overrides.allowance ?? 5_000, overrides.setTimer ?? setTimeout, overrides.clearTimer ?? clearTimeout, JudgeRefusalError, getCookieWorkflowManualReview, DEFAULT_JUDGE_MAX_TOKENS, resolveEvalModel, - WORKFLOW_JUDGE_RESPONSE_SCHEMA, validWorkflowJudgeScore); + WORKFLOW_JUDGE_RESPONSE_SCHEMA, validWorkflowJudgeScore, + judgePanel, judgePanelMean, judgePanelReasoning, JUDGE_SCORE_DIMENSIONS); return { run, records, signals, prompts, attempts, options: { ...f.opts, suite: 'Cache regression' } }; } @@ -223,7 +230,7 @@ test('the actual workflow callback preserves the pinned model and frontier rubri const actual = actualCallback(f, { judge: async (_prompt, model) => { models.push(model); return scores; } }); await actual.run({ ...actual.options, model: 'claude-sonnet-4-6', agentCapability: 'frontier', readInput: () => readWorkflowJudgeInput(f.opts) }); - expect(models).toEqual(['claude-sonnet-4-6']); + expect(models).toEqual(Array(SAMPLES).fill('claude-sonnet-4-6')); expect(actual.prompts[0]).toContain('GPT-5.6 Sol-level capability or stronger'); expect(actual.records[0]).toMatchObject({ passed: true, model: 'claude-sonnet-4-6', prompt: actual.prompts[0] }); }); @@ -239,7 +246,7 @@ test.each(['ship', 'review'])('the registered %s callback sends the frontier rub endMarker: f.opts.endMarker, references: [] }; const passing = actualCallback(f, { judge: async () => ({ ...scores, clarity: 3 }) }); await passing.run(options); - expect(passing.prompts).toHaveLength(1); + expect(passing.prompts).toHaveLength(SAMPLES); expect(passing.prompts[0]).toContain('GPT-5.6 Sol-level capability or stronger'); expect(passing.records[0]).toMatchObject({ passed: true, execution: 'executed', judge_scores: { clarity: 3 } }); const failing = actualCallback(f, { judge: async () => ({ ...scores, clarity: 2 }) }); @@ -254,8 +261,8 @@ test('the actual workflow callback executes once, reuses with provenance, and pr const f = fixture(); const first = actualCallback(f); const options = { ...f.opts, suite: 'Cache regression' }; await first.run(options); - expect(first.prompts).toEqual([f.opts.prompt]); expect(f.entries()).toHaveLength(1); - expect(first.records[0]).toMatchObject({ passed: true, execution: 'executed', cost_usd: 0.02 }); + expect(first.prompts).toEqual(Array(SAMPLES).fill(f.opts.prompt)); expect(f.entries()).toHaveLength(1); + expect(first.records[0]).toMatchObject({ passed: true, execution: 'executed', cost_usd: 0.02 * SAMPLES }); expect(first.records[0]).not.toHaveProperty('prompt'); expect(first.records[0]).not.toHaveProperty('model'); const reused = actualCallback(f, { judge: async () => ({ ...scores, clarity: 1 }) }); @@ -312,13 +319,13 @@ test('a superseding attempt cancels its predecessor before either can record a s }); test('a failed input read consumes attempt one and prevents a retry from borrowing or publishing a receipt', async () => { - const f = fixture(); f.cache().publish(scores); const receipt = fs.readFileSync(path.join(f.env.EVALS_CACHE_DIR, f.entries()[0]), 'utf8'); + const f = fixture(); f.cache().publish(panel); const receipt = fs.readFileSync(path.join(f.env.EVALS_CACHE_DIR, f.entries()[0]), 'utf8'); let reads = 0; const h = actualCallback(f, { read: options => { if (++reads === 1) throw new Error('Missing workflow fixture'); return readWorkflowJudgeInput(options); } }); await expect(h.run(h.options)).rejects.toThrow('Missing workflow fixture'); expect(h.records[0]).toMatchObject({ passed: false, exit_reason: 'harness_error' }); await h.run(h.options); - expect(h.prompts).toHaveLength(1); + expect(h.prompts).toHaveLength(SAMPLES); expect(h.records.map(record => record.execution)).toEqual(['executed', 'executed']); expect(h.attempts.get(f.opts.testName).attempt).toBe(2); expect(fs.readFileSync(path.join(f.env.EVALS_CACHE_DIR, f.entries()[0]), 'utf8')).toBe(receipt); @@ -333,14 +340,14 @@ test('monotonic expiry after a synchronous preparation or late model response re await expect(h.run(h.options)).rejects.toThrow('deadline'); expect(h.records).toHaveLength(1); expect(h.records[0]).toMatchObject({ passed: false, exit_reason: 'timeout', duration_ms: 21 }); - expect(h.prompts).toHaveLength(phase === 'preparation' ? 0 : 1); + expect(h.prompts).toHaveLength(phase === 'preparation' ? 0 : SAMPLES); expect(f.entries()).toHaveLength(0); } }); test('publication rechecks after input scanning and withdraws a receipt if recording expires', async () => { const f = fixture(); let checks = 0; - f.cache().publish(scores, () => ++checks < 2); + f.cache().publish(panel, () => ++checks < 2); expect(checks).toBe(2); expect(f.entries()).toHaveLength(0); let now = 0; const h = actualCallback(f, { budget: 20, allowance: 5, clock: () => now, @@ -373,7 +380,7 @@ test('the actual workflow callback preserves the complete public API body; cance try { const h = actualCallback(f, { judge: (prompt, model, options) => callJudge(prompt, model, options) }); await h.run(h.options); - expect(create).toHaveBeenCalledTimes(1); + expect(create).toHaveBeenCalledTimes(SAMPLES); expect(create.mock.calls[0]).toEqual([{ model: resolveEvalModel('judge'), max_tokens: 8192, messages: [{ role: 'user', content: f.opts.prompt }], @@ -412,40 +419,40 @@ test('Ship sends its authorized 64k cap and compact response contract through th f.opts.structuredResponse = true; f.opts.maxTokens = 65_536; f.opts.stream = true; - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); } finally { stream.mockRestore(); } }); test('changing response serialization misses the cache even when prompt and model match', () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); f.opts.structuredResponse = true; expect(f.cache().lookup()).toBeNull(); - f.cache().publish(scores); + f.cache().publish(panel); expect(f.entries()).toHaveLength(2); - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); const description = WORKFLOW_JUDGE_RESPONSE_SCHEMA.properties.reasoning.description; try { WORKFLOW_JUDGE_RESPONSE_SCHEMA.properties.reasoning.description += ' Changed response contract.'; expect(f.cache().lookup()).toBeNull(); } finally { WORKFLOW_JUDGE_RESPONSE_SCHEMA.properties.reasoning.description = description; } - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); f.opts.structuredResponse = false; - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); }); test('the actual cap and streaming transport independently affect workflow cache identity', () => { - const f = fixture(); f.cache().publish(scores); + const f = fixture(); f.cache().publish(panel); f.opts.maxTokens = 65_536; expect(f.cache().lookup()).toBeNull(); - f.cache().publish(scores); + f.cache().publish(panel); f.opts.stream = true; expect(f.cache().lookup()).toBeNull(); - f.cache().publish(scores); + f.cache().publish(panel); expect(f.entries()).toHaveLength(3); - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); delete f.opts.maxTokens; delete f.opts.stream; - expect(f.cache().lookup()?.scores).toEqual(scores); + expect(f.cache().lookup()?.samples).toEqual(panel); }); test('the structured callback rejects incomplete, schema-invalid and below-threshold answers without cache credit', async () => { @@ -473,3 +480,96 @@ test('the structured callback rejects incomplete, schema-invalid and below-thres expect(validWorkflowJudgeScore({ ...scores, reasoning: Array(149).fill('word').join(' ') }, { clarity: 1, completeness: 1, actionability: 1 }, true)).toBe(true); } finally { stream.mockRestore(); diagnostics.mockRestore(); } }); + +// --- Judge panel policy (EVAL_POLICY.judge): fixed concurrent samples, per-dimension +// mean and boolean majority against unchanged thresholds, an erroring sample fails +// the whole panel and is never resampled. The provider is always a stub. +const panelScore = (clarity: number, completeness = 4, actionability = 4) => ({ clarity, completeness, actionability, reasoning: `c${clarity}` }); +const panelThresholds = { clarity: 3, completeness: 3, actionability: 4 }; +const panelRefusal = () => new JudgeRefusalError({ id: 'msg_1', _request_id: 'req_1', model: 'm', usage: { input_tokens: 1, output_tokens: 0 }, content: [] }); + +describe('judge panel', () => { + test('the pre-registered panel is three samples, and the helper restates EVAL_POLICY exactly', () => { + expect(EVAL_POLICY.judge.samples).toBe(3); + expect(JUDGE_PANEL_SAMPLES).toBe(EVAL_POLICY.judge.samples); + }); + + test('draws every sample concurrently before any resolves', async () => { + let started = 0; + const releases: Array<() => void> = []; + const panel = judgePanel(() => new Promise(resolve => { started += 1; releases.push(() => resolve(started)); })); + await Promise.resolve(); + expect(started).toBe(SAMPLES); + releases.forEach(release => release()); + expect(await panel).toHaveLength(SAMPLES); + }); + + test('an erroring sample fails the panel and is never resampled', async () => { + let calls = 0; + const panel = judgePanel(async () => { + calls += 1; + if (calls === 2) throw new Error('Judge returned non-JSON: nope'); + return panelScore(5); + }); + await expect(panel).rejects.toThrow('non-JSON'); + expect(calls).toBe(SAMPLES); + }); + + test('a refusal on every sample stays a provider refusal; a partial refusal is an ordinary failure', async () => { + await expect(judgePanel(async () => { throw panelRefusal(); })).rejects.toBeInstanceOf(JudgeRefusalError); + let calls = 0; + const partial = judgePanel(async () => { if (++calls === 1) throw panelRefusal(); return panelScore(4); }); + const error = await partial.then(() => null, (reason: unknown) => reason); + expect(error).toBeInstanceOf(Error); + expect(error).not.toBeInstanceOf(JudgeRefusalError); + expect(String(error)).toContain(`sample 1 of ${SAMPLES} failed beside scored samples`); + }); + + test('numeric dimensions gate on the per-dimension mean; one low sample can be outvoted, a low mean cannot', () => { + const outvoted = judgePanelMean([panelScore(2), panelScore(4), panelScore(4)], JUDGE_SCORE_DIMENSIONS); + expect(outvoted.clarity).toBeCloseTo(10 / 3); + expect(outvoted.clarity).toBeGreaterThanOrEqual(panelThresholds.clarity); + const low = judgePanelMean([panelScore(2), panelScore(2), panelScore(4)], JUDGE_SCORE_DIMENSIONS); + expect(low.clarity).toBeLessThan(panelThresholds.clarity); + // No compensation across dimensions: each is averaged on its own. + expect(judgePanelMean([panelScore(5, 1), panelScore(5, 1), panelScore(5, 1)], JUDGE_SCORE_DIMENSIONS).completeness).toBe(1); + }); + + test('malformed sample fields fail the panel instead of averaging to NaN', () => { + expect(() => judgePanelMean([panelScore(4), { ...panelScore(4), clarity: '4' as unknown as number }, panelScore(4)], JUDGE_SCORE_DIMENSIONS)).toThrow('sample 2 has non-numeric clarity'); + expect(() => judgePanelMean([panelScore(4), null as unknown as ReturnType], JUDGE_SCORE_DIMENSIONS)).toThrow('sample 2'); + expect(() => judgePanelMean([], JUDGE_SCORE_DIMENSIONS)).toThrow('no samples'); + }); + + test('boolean fields gate on a strict majority', () => { + const vote = (...values: boolean[]) => judgePanelMajority(values.map(value => ({ ok: value })), 'ok'); + expect(vote(true, true, false)).toBe(true); + expect(vote(true, false, false)).toBe(false); + expect(vote(true, false)).toBe(false); + expect(() => judgePanelMajority([{ ok: true }, { ok: 'yes' }], 'ok')).toThrow('sample 2 has non-boolean ok'); + }); + + test('reasoning keeps every sample, numbered, even for malformed samples', () => { + expect(judgePanelReasoning([panelScore(4), null, { reasoning: 7 }])).toBe('[sample 1] c4\n[sample 2] \n[sample 3] '); + }); + + test('the cache stores and validates only a complete panel against the mean', () => { + expect(validWorkflowJudgePanel({ samples: [panelScore(2), panelScore(4), panelScore(4)] }, panelThresholds)).toBe(true); + expect(validWorkflowJudgePanel({ samples: [panelScore(2), panelScore(2), panelScore(4)] }, panelThresholds)).toBe(false); + expect(validWorkflowJudgePanel({ samples: [panelScore(4), panelScore(4)] }, panelThresholds)).toBe(false); + expect(validWorkflowJudgePanel({ samples: [panelScore(4), panelScore(4), panelScore(4), panelScore(4)] }, panelThresholds)).toBe(false); + expect(validWorkflowJudgePanel({ samples: [panelScore(4), panelScore(4), { ...panelScore(4), clarity: 6 }] }, panelThresholds)).toBe(false); + expect(validWorkflowJudgePanel({ samples: [panelScore(4), panelScore(4), panelScore(4)], prompt: 'x' }, panelThresholds)).toBe(false); + expect(validWorkflowJudgePanel(panelScore(4), panelThresholds)).toBe(false); + }); + + test('every judge in the quality file samples through the panel, never a lone call', () => { + const source = fs.readFileSync(path.join(import.meta.dir, 'skill-llm-eval.test.ts'), 'utf8'); + const calls = [...source.matchAll(/\b(?:callJudge<[^>(]*(?:<[^>]*>[^>(]*)*>|judge)\(/g)]; + expect(calls.length).toBeGreaterThanOrEqual(8); + for (const call of calls) { + expect(source.slice(Math.max(0, call.index! - 25), call.index), `unpaneled judge call at offset ${call.index}`).toMatch(/judgePanel\(\(\) => $/); + } + expect(source).not.toMatch(/\bscores\.reasoning\b|\bresult\.reasoning\b/); + }); +});