docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic

AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.
This commit is contained in:
garrytan committed 2026-09-29 19:17:12 +00:00
1 parent 8ee9f887ed
commit ccb5f3c07c
3 files changed
+182 -27

No files matched your search

+13 -5
View File
@@ -148,7 +148,8 @@ When fixing failures or preparing `/ship`, follow this order:
public events in free regressions, including negative controls, before paying public events in free regressions, including negative controls, before paying
for another agent run. Check behavior and acknowledgments; match exact prose for another agent run. Check behavior and acknowledgments; match exact prose
only when that prose is the contract. Do not lower thresholds, increase model only when that prose is the contract. Do not lower thresholds, increase model
budgets, skip cases, or rejudge a failure to manufacture a pass. budgets, skip cases, or rejudge a failure to manufacture a pass. A
pre-registered fixed panel is not rejudging.
For policy or validation repairs, exercise the actual registered callback with For policy or validation repairs, exercise the actual registered callback with
representative native input and assert that it uses the helper’s result. representative native input and assert that it uses the helper’s result.
When renderer or parser failures recur at the same boundary, verify the When renderer or parser failures recur at the same boundary, verify the
@@ -208,10 +209,16 @@ When fixing failures or preparing `/ship`, follow this order:
result and pending permission state; diagnose a blocked actor before waiting result and pending permission state; diagnose a blocked actor before waiting
through its deadline. Preserve cancellation separately from a test verdict. through its deadline. Preserve cancellation separately from a test verdict.
Skipped or unstarted cases Skipped or unstarted cases
do not satisfy coverage; preserve every attempt. Retries follow the approved do not satisfy coverage; preserve every attempt. Paid evals never retry. Each
policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so case's kind (`E2E_KINDS`) fixes its trials before the run: `rule` one trial;
only files whose every case budget is CAPTURE tier or shorter keep one retry; `behavior` a panel of 3 independent trials, PASS at >= 2 with no contract
never add retries to pass a longer case. violation; `judge` 3 samples on one output, gated on the mean against the
unchanged threshold. Never add trials, samples or dispatches after seeing a
result, never change a kind to change a verdict without pass-rate evidence,
and report every trial. Quarantine follows `CASE_QUARANTINE`'s entry and exit
rules only (`EVAL_POLICY`, `docs/TESTING_INTERNALS.md`). A census whose every
red is machine-classified INFRA or INCOMPLETE may be re-dispatched once as a
new run; report both runs.
7. Prove all known repairs with focused tests, including affected paid cases. 7. Prove all known repairs with focused tests, including affected paid cases.
Rerun a failed case only after a concrete repair or a demonstrated launch Rerun a failed case only after a concrete repair or a demonstrated launch
correction. Run the remaining required selected evaluations on the integrated correction. Run the remaining required selected evaluations on the integrated
@@ -245,6 +252,7 @@ bun run test # complete free suite via the strict shard runner (no A
bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY) bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY)
bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals
bun run eval:bg:release # fresh complete gate + periodic live coverage bun run eval:bg:release # fresh complete gate + periodic live coverage
bun run eval:pass-rates # per-case trial pass rates (Wilson), drift and quarantine alarms (--case, --gate)
bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free) bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free)
bun run test:windows # curated Windows-safe subset (runs on windows-latest) bun run test:windows # curated Windows-safe subset (runs on windows-latest)
bun run build # generate docs + compile binaries bun run build # generate docs + compile binaries
+45 -6
View File
@@ -239,10 +239,29 @@ Complete start-to-finish flows belong to the `marathon` tier
(`describeE2ETier('marathon')`), which runs only in the non-blocking (`describeE2ETier('marathon')`), which runs only in the non-blocking
`evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge. `evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge.
Retries: a timed-out attempt is a verdict. Only files whose every case budget is Verdicts: paid evals never retry. Each case's kind in `E2E_KINDS`
CAPTURE tier or shorter (`RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`) (`test/helpers/touchfiles-data.ts`) fixes its trials before the run, from the
keep one automatic retry for fast-failing flakes; every other paid file runs once. constants in `EVAL_POLICY` (`test/helpers/periodic-exclude-data.ts`):
Case budgets themselves never change with this rule.
- `rule` (the default): one trial; any failed assertion fails the case. Use it
when nothing stochastic decides the verdict, or when the verdict checks a
contract the product must meet every run (no writes in plan mode, a question
before a decision, a skill-mandated step, no leaked secret).
- `behavior`: a panel of 3 independent trials run as parallel case shards,
PASS at 2 or more with no contract violation (`expectContract()`). Use it only
when a live model choice decides the verdict and an occasional deviation is
acceptable product behavior; the one-line reason goes in `BEHAVIOR_WHY`.
- `judge`: an LLM judge scoring a fixed input; 3 samples of the same prompt,
gated on the per-dimension mean (booleans on a majority) against the
unchanged threshold. An erroring sample fails the panel and is never resampled.
A timed-out, crashed or infrastructure-failed trial counts as a failed trial and
is reported with its class; a missing trial makes the case INCOMPLETE, which
fails the lane. A 2-of-3 pass is reported as `PASS 2/3` with the failed trial's
cause, never as a clean pass. Case budgets and thresholds never change with
this policy. Quarantine (`CASE_QUARANTINE`) and history are described in
`docs/TESTING_INTERNALS.md`; `bun run eval:pass-rates --case <id>` shows a
case's per-trial pass rate with its Wilson interval.
CI enables verified first-attempt reuse for 16 workflow quality judges for CI enables verified first-attempt reuse for 16 workflow quality judges for
24 hours within the same PR. The cookie workflow's custom input and the other 11 24 hours within the same PR. The cookie workflow's custom input and the other 11
@@ -415,7 +434,7 @@ When E2E tests run, they produce machine-readable artifacts in `~/.gstack-dev/`:
bun run eval:list # list all eval runs (turns, duration, cost per run) bun run eval:list # list all eval runs (turns, duration, cost per run)
bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary
bun run eval:summary # aggregate stats + per-test efficiency averages across runs bun run eval:summary # aggregate stats + per-test efficiency averages across runs
bun run eval:flake-rank # rank tests by flake signal: retried passes first, then failure rate (--json, --dir, --since-days) bun run eval:pass-rates # per-case trial pass rates + Wilson intervals from recent weekly runs (--case, --runs, --dir, --backfill, --json, --gate); eval:flake-rank is an alias
``` ```
**Detached runs for agents and long suites.** When an agent (or you, for a run **Detached runs for agents and long suites.** When an agent (or you, for a run
@@ -463,7 +482,9 @@ Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`:
- **Completeness** — Are all commands, flags, and usage patterns documented? - **Completeness** — Are all commands, flags, and usage patterns documented?
- **Actionability** — Can the agent execute tasks using only the information in the doc? - **Actionability** — Can the agent execute tasks using only the information in the doc?
Each dimension is scored 1-5. Threshold: every dimension must score **≥ 4**. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher. Each dimension is scored 1-5 by a panel of 3 samples of the same prompt, drawn
concurrently; each dimension's panel mean must meet that judge's threshold (≥ 4
for most dimensions; see each case). An erroring sample fails the panel. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher.
```bash ```bash
# Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals # Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals
@@ -483,6 +504,24 @@ fails, add the named path to the named key and check selection with
`bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. The rule is a lower bound: a fixture `bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. The rule is a lower bound: a fixture
path the test builds at runtime is not visible to it, so add such paths to the key by hand. path the test builds at runtime is not visible to it, so add such paths to the key by hand.
### Add a paid eval
1. **Test file.** Write the case in a paid test file, registered with a literal
name (`testIfSelected('<case-id>', ...)`), grading the outcome (files, git
state, native questions, exit status) rather than wording, unless the step
itself is the contract. Wrap contract assertions in `expectContract()`.
2. **Touchfiles.** Add `'<case-id>': [...]` to `E2E_TOUCHFILES`; `bun test
test/touchfiles.test.ts` names any missing closure path.
3. **Tier.** Add it to `E2E_TIERS`: `gate` for cheap contracts every PR needs,
`periodic` for long or model-quality cases, `marathon` for complete flows.
4. **Kind.** Add it to `E2E_KINDS` (`rule` unless a live model choice may
acceptably deviate; then `behavior` plus a `BEHAVIOR_WHY` line).
`bun test test/eval-kinds.test.ts` prints the literal to add.
5. **PR profile.** If a PR should run it, add it to `scripts/test-pr-profile.ts`
and check `bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`.
6. **Try the panel locally.** `bun run scripts/test-paid-shards.ts --tier <tier>
--case <case-id> --trials 3` runs the same panel CI runs, before you push.
### CI ### CI
A GitHub Action (`.github/workflows/skill-docs.yml`) generates all hosts on pushes to main and on PRs, then rejects tracked differences and nonignored untracked output. Generation errors also fail the job. Optional ignored host caches are not compared against Git. A GitHub Action (`.github/workflows/skill-docs.yml`) generates all hosts on pushes to main and on PRs, then rejects tracked differences and nonignored untracked output. Generation errors also fail the job. Optional ignored host caches are not compared against Git.
+124 -16
View File
@@ -252,24 +252,20 @@ processes × `EVALS_CONCURRENCY` within-shard, per-shard `GSTACK_EVAL_DIR`,
full-stream spooling to per-shard log files (path printed at START and on full-stream spooling to per-shard log files (path printed at START and on
failure), never-started/timed-out taxonomy, and parent-computed diff failure), never-started/timed-out taxonomy, and parent-computed diff
selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a
child that can't parse it recomputes locally with one warning). Retries follow child that can't parse it recomputes locally with one warning). Paid evals
one rule (`retriesForFiles`, `RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`): never retry; each case's kind fixes its trials before the run (see "Eval verdict
a timed-out attempt is a verdict, so a file keeps one Bun retry only when every policy" below). Files in `CASE_SHARDED_FILES` run one registered case per
case budget is CAPTURE tier plus its recording grace or shorter (registered rows
derive it from `caseMs`, `SHORT_CASE_RETRY_FILES` lists the rest); every other
file, including the former `retries: 2` matrix rows, runs once. `--list` prints
each shard's retries. Files in `CASE_SHARDED_FILES` run one registered case per
process (`<file>#<case id>`, an exact `--test-name-pattern`, exactly one executed process (`<file>#<case id>`, an exact `--test-name-pattern`, exactly one executed
case), so a long file of short cases spreads across runners and each case gets case), so a long file of short cases spreads across runners and each case gets
its own SDK semaphore. its own SDK semaphore.
Flake telemetry rides the store: every recorded test carries its 1-based Trial telemetry rides the store: every recorded test carries its 1-based
`attempt` (a pass-on-attempt-2 stays visible forever — bun's own stream hides `attempt` plus, on an isolated trial shard, its `case_id`, `kind`, `trial`,
it), runs list `flaky_retries`, the report warns on passed-only-on-retry `panel` and `policy_version`, and each lane's report uploads one
tests, and `bun run eval:flake-rank` ranks the series (retried passes first, `trial-outcomes` JSONL line per trial. `bun run eval:pass-rates`
then failure rate; 60-day recency bound on eval files; the free lane's flake (`eval:flake-rank` is an alias) turns that history into per-case pass rates
ledger is folded in from `flakeLedgerPath()` — override with (see "Pass-rate history" below; the free lane's flake ledger is folded in from
`GSTACK_FLAKE_LEDGER`, the same env var the CI free lane sets before `flakeLedgerPath()` — override with `GSTACK_FLAKE_LEDGER`, the same env var the
uploading the ledger as the `flake-ledger` artifact). Census integrity is CI free lane sets before uploading the ledger as the `flake-ledger` artifact). Census integrity is
enforced from the free suite: every `E2E_TOUCHFILES` / `LLM_JUDGE_TOUCHFILES` enforced from the free suite: every `E2E_TOUCHFILES` / `LLM_JUDGE_TOUCHFILES`
key must name a living paid test (`test/touchfiles.test.ts`'s reverse key must name a living paid test (`test/touchfiles.test.ts`'s reverse
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
@@ -313,7 +309,7 @@ CI supplies the scoped cache/runtime configuration; local runs are fresh by
default. Cached scores must default. Cached scores must
pass current assertions; reused records retain their original source and time pass current assertions; reused records retain their original source and time
and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts
to PR-profile E2E shards that run with zero retries (so the pass is structurally a to PR-profile E2E shards (paid evals never retry, so a pass is structurally a
first attempt): the identity hashes the test's import closure, every tracked file first attempt): the identity hashes the test's import closure, every tracked file
matched by the touchfiles of every case the file registers plus the global matched by the touchfiles of every case the file registers plus the global
touchfiles, the runner/workflow/setup actions, the child's environment pins, the touchfiles, the runner/workflow/setup actions, the child's environment pins, the
@@ -377,6 +373,118 @@ the runner parent and handed to shard children as `GSTACK_CLAUDE_CLI_VERSION`
(never spawned on a test thread), so a TUI-drift flake hunt is a grep, not (never spawned on a test thread), so a TUI-drift flake hunt is a grep, not
archaeology. archaeology.
**Eval verdict policy** (`EVAL_POLICY` version 1 in
`test/helpers/periodic-exclude-data.ts`, pre-registered 2026-09-29). Paid evals
never retry. Each live case has exactly one kind in `E2E_KINDS`
(`test/helpers/touchfiles-data.ts`; `test/eval-kinds.test.ts` enforces coverage),
and the kind fixes its trials before the run:
- `rule` (default): one trial; any failed assertion fails the verdict. For
cases where nothing stochastic decides the verdict, or where it checks a
contract the product must meet every run.
- `behavior`: a panel of `n = 3` independent trials, launched together as
isolated case shards on different slices (key `<file>#<id>~t<N>`). All three
always run: no early stop and no conditional extra trial. PASS when at least
`k = 2` pass and no trial violated a contract (`expectContract()` stamps
`failure_class: 'contract'`). Each behavior case names its tolerated deviation
in `BEHAVIOR_WHY` and must have a literal registration so it can run alone.
- `judge`: an LLM judge scoring a fixed input. `judgePanel()`
(`test/helpers/llm-judge.ts`) draws 3 samples of the same prompt concurrently
inside the unchanged `JUDGE_MS`; numeric dimensions gate on the per-dimension
mean against the unchanged threshold (no dimension compensates for another),
booleans on a strict majority. A sample that errors (refusal, truncation,
non-JSON, a malformed field) fails the panel and is never resampled; a
refusal counts as an unscored panel only when every sample refused.
`callJudge`'s 429 backoff happens before any model output and is transport,
not a verdict retry. The workflow-judge cache stores whole panels only.
`panelVerdict()` (`test/helpers/eval-store.ts`) is the single verdict
function the report, `collector-outcomes.json`, the PR comment and pass-rates
all use. A timed-out, crashed or infrastructure-failed trial is a failed trial
recorded with its class; a missing or duplicate trial record makes the verdict
INCOMPLETE, which fails the lane; a 2/3 PASS is shown as `PASS 2/3` with the
failed trial's cause. A manual re-run adds trials under a new run attempt and
never replaces the first attempt's verdict. A red census is never rerun on
unchanged inputs: each red is diagnosed as product, test/detector, harness or
infra and resolved by a concrete repair and a fresh census, or listed as a named
red. The one exception: a census whose every red verdict is machine-classified
INFRA or INCOMPLETE (missing slice artifact, runner loss, API error before the
first model turn) may be re-dispatched once as a new run, and both runs are
reported. Changing any `EVAL_POLICY` constant after seeing census results needs
Garry's re-approval, a `version` bump and a fresh census;
`test/periodic-exclude-policy.test.ts` pins the approved values.
**Quarantine** (`CASE_QUARANTINE`, same file). An entry needs: a per-trial rate
below 95% over at least 10 post-policy trials of the case's current input
identity (pre-policy backfill may justify only an initial entry, labeled as
such); a written diagnosis in `reason` whose `failureClass` is `detector`,
`harness` or `model-latency` (a product defect is fixed or listed as a named
red, never quarantined); unchanged case touchfiles in the change that adds it;
and an owner, tracking pointer, `enteredAt` date and measurable `exit`. A
quarantined case still runs its full panel and reports in every lane but cannot
fail it, except on a hard break (0 of n) or a contract violation, and it never
counts as passing coverage. At most 10% of a blocking tier (gate, periodic) may
be quarantined. The weekly report fails when an entry passes its exit rule (at
least 97% over at least 10 trials) without being removed, when an entry is 8
weekly runs old, or when a tier is over its cap.
**Pass-rate history** (`bun run eval:pass-rates`, `scripts/eval-flake-rank.ts`).
It reads the `trial-outcomes` artifact of the last N completed
`evals-periodic.yml` runs on the current branch and `main` (flags: `--case`,
`--runs N`, `--branch`, `--dir`, `--backfill`, `--json`, `--gate`) and prints
per-case per-trial pass rates with 95% Wilson intervals. A series is one case
under one input identity, the hash of its own touchfiles minus
`GLOBAL_TOUCHFILES` (harness edits do not restart it), per model, Claude CLI
version and policy version; a change starts a new series and older ones stay
visible. Labels: INCONCLUSIVE below 10 trials, BROKEN when the latest run is
0/n after a prior interval at or above 95%, FLAKY when failures leave the
interval straddling 95%, FAILING when the whole interval is below it, PASSING
otherwise. `--backfill` imports legacy slice artifacts as pre-policy trials
(first attempt only; a record that names no registry id is listed as
unattributed, never guessed); they are display-only. `--gate` (the weekly
report) fails with ACTION REQUIRED, on post-policy trials of the current series
only, when a non-quarantined blocking case meets the entry rule (proposing an
entry), when a `rule` case does (rule case behaving like behavior: fix or
reclassify), when a blocking case's current identity is significantly below its
previous one (one-sided Fisher exact, α = 0.05, at least 6 trials each side,
Holm-controlled across the cases tested), and on the quarantine rules above.
History that cannot be fetched fails the gate closed.
**The arithmetic.** With per-trial pass rate p, the chance a single case goes
red (a false red while the product works, the catch rate once it has
regressed):
| p | 1 trial | 2-of-3 panel |
|---|---|---|
| 0.99 | 1.0% | 0.03% |
| 0.95 | 5.0% | 0.72% |
| 0.90 | 10.0% | 2.8% |
| 0.70 | 30.0% | 21.6% |
| 0.30 | 70.0% | 78.4% |
The panel removes most false reds at healthy rates, but it catches a 0.95 → 0.70
regression in one run only 21.6% of the time (a single trial 30%, retry-until-green
3%), so drift detection is the history rule's job, not the per-run verdict's.
The Fisher alarm is weak at the minimum sample (5.4% power for 0.95 → 0.70 at
6 trials a side), and ten straight passes still leave a 72% Wilson lower bound:
after this policy lands, every series starts INCONCLUSIVE.
A lane is all green with probability Π p_rule × Π P(≥2 of 3 | p_behavior) ×
Π p_judge. For the current registry (PR gate worst case: 107 rule cases and 24
judges; weekly census: 190 rule, 22 behavior and 25 judge verdicts), with rule
and judge verdicts at p_rule:
| p_rule | full PR gate | weekly, behavior p = 0.90 | 0.95 | 0.97 |
|---|---|---|---|---|
| 0.99 | 26.8% | 6.2% | 9.8% | 10.9% |
| 0.995 | 51.9% | 18.2% | 29.0% | 32.1% |
| 0.999 | 87.7% | 43.2% | 68.7% | 76.1% |
The rule term dominates: a green lane on a working product needs rule cases to
be near-deterministic (0.999), which is why failing detectors are converted to
outcome checks and product defects are fixed or named, and why each census
reports its expected lane false-red from the measured rates.
**Timeout policy.** Paid tests use the tiers in **Timeout policy.** Paid tests use the tiers in
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG); `test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall `test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall