docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic

AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.
This commit is contained in:
garrytan committed 2026-09-29 19:17:12 +00:00
1 parent 8ee9f887ed
commit ccb5f3c07c
3 files changed
+182 -27

No files matched your search

+13 -5
View File
@@ -148,7 +148,8 @@ When fixing failures or preparing `/ship`, follow this order:
public events in free regressions, including negative controls, before paying
for another agent run. Check behavior and acknowledgments; match exact prose
only when that prose is the contract. Do not lower thresholds, increase model
budgets, skip cases, or rejudge a failure to manufacture a pass.
budgets, skip cases, or rejudge a failure to manufacture a pass. A
pre-registered fixed panel is not rejudging.
For policy or validation repairs, exercise the actual registered callback with
representative native input and assert that it uses the helper’s result.
When renderer or parser failures recur at the same boundary, verify the
@@ -208,10 +209,16 @@ When fixing failures or preparing `/ship`, follow this order:
result and pending permission state; diagnose a blocked actor before waiting
through its deadline. Preserve cancellation separately from a test verdict.
Skipped or unstarted cases
do not satisfy coverage; preserve every attempt. Retries follow the approved
policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so
only files whose every case budget is CAPTURE tier or shorter keep one retry;
never add retries to pass a longer case.
do not satisfy coverage; preserve every attempt. Paid evals never retry. Each
case's kind (`E2E_KINDS`) fixes its trials before the run: `rule` one trial;
`behavior` a panel of 3 independent trials, PASS at >= 2 with no contract
violation; `judge` 3 samples on one output, gated on the mean against the
unchanged threshold. Never add trials, samples or dispatches after seeing a
result, never change a kind to change a verdict without pass-rate evidence,
and report every trial. Quarantine follows `CASE_QUARANTINE`'s entry and exit
rules only (`EVAL_POLICY`, `docs/TESTING_INTERNALS.md`). A census whose every
red is machine-classified INFRA or INCOMPLETE may be re-dispatched once as a
new run; report both runs.
7. Prove all known repairs with focused tests, including affected paid cases.
Rerun a failed case only after a concrete repair or a demonstrated launch
correction. Run the remaining required selected evaluations on the integrated
@@ -245,6 +252,7 @@ bun run test # complete free suite via the strict shard runner (no A
bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY)
bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals
bun run eval:bg:release # fresh complete gate + periodic live coverage
bun run eval:pass-rates # per-case trial pass rates (Wilson), drift and quarantine alarms (--case, --gate)
bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free)
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
bun run build # generate docs + compile binaries