mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic
AGENTS.md replaces the retry rule with the approved policy text (no retries; kind fixes trials; no added trials, samples or dispatches after a result; quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval' checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the current 191 rule / 22 behavior / 25 judge registry.
This commit is contained in:
1 parent
8ee9f887ed
commit
ccb5f3c07c
3 files changed
+182
-27
No files matched your search
@@ -148,7 +148,8 @@ When fixing failures or preparing `/ship`, follow this order:
|
|||||||
public events in free regressions, including negative controls, before paying
|
public events in free regressions, including negative controls, before paying
|
||||||
for another agent run. Check behavior and acknowledgments; match exact prose
|
for another agent run. Check behavior and acknowledgments; match exact prose
|
||||||
only when that prose is the contract. Do not lower thresholds, increase model
|
only when that prose is the contract. Do not lower thresholds, increase model
|
||||||
budgets, skip cases, or rejudge a failure to manufacture a pass.
|
budgets, skip cases, or rejudge a failure to manufacture a pass. A
|
||||||
|
pre-registered fixed panel is not rejudging.
|
||||||
For policy or validation repairs, exercise the actual registered callback with
|
For policy or validation repairs, exercise the actual registered callback with
|
||||||
representative native input and assert that it uses the helper’s result.
|
representative native input and assert that it uses the helper’s result.
|
||||||
When renderer or parser failures recur at the same boundary, verify the
|
When renderer or parser failures recur at the same boundary, verify the
|
||||||
@@ -208,10 +209,16 @@ When fixing failures or preparing `/ship`, follow this order:
|
|||||||
result and pending permission state; diagnose a blocked actor before waiting
|
result and pending permission state; diagnose a blocked actor before waiting
|
||||||
through its deadline. Preserve cancellation separately from a test verdict.
|
through its deadline. Preserve cancellation separately from a test verdict.
|
||||||
Skipped or unstarted cases
|
Skipped or unstarted cases
|
||||||
do not satisfy coverage; preserve every attempt. Retries follow the approved
|
do not satisfy coverage; preserve every attempt. Paid evals never retry. Each
|
||||||
policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so
|
case's kind (`E2E_KINDS`) fixes its trials before the run: `rule` one trial;
|
||||||
only files whose every case budget is CAPTURE tier or shorter keep one retry;
|
`behavior` a panel of 3 independent trials, PASS at >= 2 with no contract
|
||||||
never add retries to pass a longer case.
|
violation; `judge` 3 samples on one output, gated on the mean against the
|
||||||
|
unchanged threshold. Never add trials, samples or dispatches after seeing a
|
||||||
|
result, never change a kind to change a verdict without pass-rate evidence,
|
||||||
|
and report every trial. Quarantine follows `CASE_QUARANTINE`'s entry and exit
|
||||||
|
rules only (`EVAL_POLICY`, `docs/TESTING_INTERNALS.md`). A census whose every
|
||||||
|
red is machine-classified INFRA or INCOMPLETE may be re-dispatched once as a
|
||||||
|
new run; report both runs.
|
||||||
7. Prove all known repairs with focused tests, including affected paid cases.
|
7. Prove all known repairs with focused tests, including affected paid cases.
|
||||||
Rerun a failed case only after a concrete repair or a demonstrated launch
|
Rerun a failed case only after a concrete repair or a demonstrated launch
|
||||||
correction. Run the remaining required selected evaluations on the integrated
|
correction. Run the remaining required selected evaluations on the integrated
|
||||||
@@ -245,6 +252,7 @@ bun run test # complete free suite via the strict shard runner (no A
|
|||||||
bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY)
|
bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY)
|
||||||
bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals
|
bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals
|
||||||
bun run eval:bg:release # fresh complete gate + periodic live coverage
|
bun run eval:bg:release # fresh complete gate + periodic live coverage
|
||||||
|
bun run eval:pass-rates # per-case trial pass rates (Wilson), drift and quarantine alarms (--case, --gate)
|
||||||
bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free)
|
bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free)
|
||||||
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
|
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
|
||||||
bun run build # generate docs + compile binaries
|
bun run build # generate docs + compile binaries
|
||||||
|
|||||||
+45
-6
@@ -239,10 +239,29 @@ Complete start-to-finish flows belong to the `marathon` tier
|
|||||||
(`describeE2ETier('marathon')`), which runs only in the non-blocking
|
(`describeE2ETier('marathon')`), which runs only in the non-blocking
|
||||||
`evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge.
|
`evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge.
|
||||||
|
|
||||||
Retries: a timed-out attempt is a verdict. Only files whose every case budget is
|
Verdicts: paid evals never retry. Each case's kind in `E2E_KINDS`
|
||||||
CAPTURE tier or shorter (`RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`)
|
(`test/helpers/touchfiles-data.ts`) fixes its trials before the run, from the
|
||||||
keep one automatic retry for fast-failing flakes; every other paid file runs once.
|
constants in `EVAL_POLICY` (`test/helpers/periodic-exclude-data.ts`):
|
||||||
Case budgets themselves never change with this rule.
|
|
||||||
|
- `rule` (the default): one trial; any failed assertion fails the case. Use it
|
||||||
|
when nothing stochastic decides the verdict, or when the verdict checks a
|
||||||
|
contract the product must meet every run (no writes in plan mode, a question
|
||||||
|
before a decision, a skill-mandated step, no leaked secret).
|
||||||
|
- `behavior`: a panel of 3 independent trials run as parallel case shards,
|
||||||
|
PASS at 2 or more with no contract violation (`expectContract()`). Use it only
|
||||||
|
when a live model choice decides the verdict and an occasional deviation is
|
||||||
|
acceptable product behavior; the one-line reason goes in `BEHAVIOR_WHY`.
|
||||||
|
- `judge`: an LLM judge scoring a fixed input; 3 samples of the same prompt,
|
||||||
|
gated on the per-dimension mean (booleans on a majority) against the
|
||||||
|
unchanged threshold. An erroring sample fails the panel and is never resampled.
|
||||||
|
|
||||||
|
A timed-out, crashed or infrastructure-failed trial counts as a failed trial and
|
||||||
|
is reported with its class; a missing trial makes the case INCOMPLETE, which
|
||||||
|
fails the lane. A 2-of-3 pass is reported as `PASS 2/3` with the failed trial's
|
||||||
|
cause, never as a clean pass. Case budgets and thresholds never change with
|
||||||
|
this policy. Quarantine (`CASE_QUARANTINE`) and history are described in
|
||||||
|
`docs/TESTING_INTERNALS.md`; `bun run eval:pass-rates --case <id>` shows a
|
||||||
|
case's per-trial pass rate with its Wilson interval.
|
||||||
|
|
||||||
CI enables verified first-attempt reuse for 16 workflow quality judges for
|
CI enables verified first-attempt reuse for 16 workflow quality judges for
|
||||||
24 hours within the same PR. The cookie workflow's custom input and the other 11
|
24 hours within the same PR. The cookie workflow's custom input and the other 11
|
||||||
@@ -415,7 +434,7 @@ When E2E tests run, they produce machine-readable artifacts in `~/.gstack-dev/`:
|
|||||||
bun run eval:list # list all eval runs (turns, duration, cost per run)
|
bun run eval:list # list all eval runs (turns, duration, cost per run)
|
||||||
bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary
|
bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary
|
||||||
bun run eval:summary # aggregate stats + per-test efficiency averages across runs
|
bun run eval:summary # aggregate stats + per-test efficiency averages across runs
|
||||||
bun run eval:flake-rank # rank tests by flake signal: retried passes first, then failure rate (--json, --dir, --since-days)
|
bun run eval:pass-rates # per-case trial pass rates + Wilson intervals from recent weekly runs (--case, --runs, --dir, --backfill, --json, --gate); eval:flake-rank is an alias
|
||||||
```
|
```
|
||||||
|
|
||||||
**Detached runs for agents and long suites.** When an agent (or you, for a run
|
**Detached runs for agents and long suites.** When an agent (or you, for a run
|
||||||
@@ -463,7 +482,9 @@ Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`:
|
|||||||
- **Completeness** — Are all commands, flags, and usage patterns documented?
|
- **Completeness** — Are all commands, flags, and usage patterns documented?
|
||||||
- **Actionability** — Can the agent execute tasks using only the information in the doc?
|
- **Actionability** — Can the agent execute tasks using only the information in the doc?
|
||||||
|
|
||||||
Each dimension is scored 1-5. Threshold: every dimension must score **≥ 4**. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher.
|
Each dimension is scored 1-5 by a panel of 3 samples of the same prompt, drawn
|
||||||
|
concurrently; each dimension's panel mean must meet that judge's threshold (≥ 4
|
||||||
|
for most dimensions; see each case). An erroring sample fails the panel. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals
|
# Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals
|
||||||
@@ -483,6 +504,24 @@ fails, add the named path to the named key and check selection with
|
|||||||
`bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. The rule is a lower bound: a fixture
|
`bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`. The rule is a lower bound: a fixture
|
||||||
path the test builds at runtime is not visible to it, so add such paths to the key by hand.
|
path the test builds at runtime is not visible to it, so add such paths to the key by hand.
|
||||||
|
|
||||||
|
### Add a paid eval
|
||||||
|
|
||||||
|
1. **Test file.** Write the case in a paid test file, registered with a literal
|
||||||
|
name (`testIfSelected('<case-id>', ...)`), grading the outcome (files, git
|
||||||
|
state, native questions, exit status) rather than wording, unless the step
|
||||||
|
itself is the contract. Wrap contract assertions in `expectContract()`.
|
||||||
|
2. **Touchfiles.** Add `'<case-id>': [...]` to `E2E_TOUCHFILES`; `bun test
|
||||||
|
test/touchfiles.test.ts` names any missing closure path.
|
||||||
|
3. **Tier.** Add it to `E2E_TIERS`: `gate` for cheap contracts every PR needs,
|
||||||
|
`periodic` for long or model-quality cases, `marathon` for complete flows.
|
||||||
|
4. **Kind.** Add it to `E2E_KINDS` (`rule` unless a live model choice may
|
||||||
|
acceptably deviate; then `behavior` plus a `BEHAVIOR_WHY` line).
|
||||||
|
`bun test test/eval-kinds.test.ts` prints the literal to add.
|
||||||
|
5. **PR profile.** If a PR should run it, add it to `scripts/test-pr-profile.ts`
|
||||||
|
and check `bun run scripts/test-paid-shards.ts --tier gate --profile pr --list`.
|
||||||
|
6. **Try the panel locally.** `bun run scripts/test-paid-shards.ts --tier <tier>
|
||||||
|
--case <case-id> --trials 3` runs the same panel CI runs, before you push.
|
||||||
|
|
||||||
### CI
|
### CI
|
||||||
|
|
||||||
A GitHub Action (`.github/workflows/skill-docs.yml`) generates all hosts on pushes to main and on PRs, then rejects tracked differences and nonignored untracked output. Generation errors also fail the job. Optional ignored host caches are not compared against Git.
|
A GitHub Action (`.github/workflows/skill-docs.yml`) generates all hosts on pushes to main and on PRs, then rejects tracked differences and nonignored untracked output. Generation errors also fail the job. Optional ignored host caches are not compared against Git.
|
||||||
|
|||||||
+124
-16
@@ -252,24 +252,20 @@ processes × `EVALS_CONCURRENCY` within-shard, per-shard `GSTACK_EVAL_DIR`,
|
|||||||
full-stream spooling to per-shard log files (path printed at START and on
|
full-stream spooling to per-shard log files (path printed at START and on
|
||||||
failure), never-started/timed-out taxonomy, and parent-computed diff
|
failure), never-started/timed-out taxonomy, and parent-computed diff
|
||||||
selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a
|
selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a
|
||||||
child that can't parse it recomputes locally with one warning). Retries follow
|
child that can't parse it recomputes locally with one warning). Paid evals
|
||||||
one rule (`retriesForFiles`, `RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`):
|
never retry; each case's kind fixes its trials before the run (see "Eval verdict
|
||||||
a timed-out attempt is a verdict, so a file keeps one Bun retry only when every
|
policy" below). Files in `CASE_SHARDED_FILES` run one registered case per
|
||||||
case budget is CAPTURE tier plus its recording grace or shorter (registered rows
|
|
||||||
derive it from `caseMs`, `SHORT_CASE_RETRY_FILES` lists the rest); every other
|
|
||||||
file, including the former `retries: 2` matrix rows, runs once. `--list` prints
|
|
||||||
each shard's retries. Files in `CASE_SHARDED_FILES` run one registered case per
|
|
||||||
process (`<file>#<case id>`, an exact `--test-name-pattern`, exactly one executed
|
process (`<file>#<case id>`, an exact `--test-name-pattern`, exactly one executed
|
||||||
case), so a long file of short cases spreads across runners and each case gets
|
case), so a long file of short cases spreads across runners and each case gets
|
||||||
its own SDK semaphore.
|
its own SDK semaphore.
|
||||||
Flake telemetry rides the store: every recorded test carries its 1-based
|
Trial telemetry rides the store: every recorded test carries its 1-based
|
||||||
`attempt` (a pass-on-attempt-2 stays visible forever — bun's own stream hides
|
`attempt` plus, on an isolated trial shard, its `case_id`, `kind`, `trial`,
|
||||||
it), runs list `flaky_retries`, the report warns on passed-only-on-retry
|
`panel` and `policy_version`, and each lane's report uploads one
|
||||||
tests, and `bun run eval:flake-rank` ranks the series (retried passes first,
|
`trial-outcomes` JSONL line per trial. `bun run eval:pass-rates`
|
||||||
then failure rate; 60-day recency bound on eval files; the free lane's flake
|
(`eval:flake-rank` is an alias) turns that history into per-case pass rates
|
||||||
ledger is folded in from `flakeLedgerPath()` — override with
|
(see "Pass-rate history" below; the free lane's flake ledger is folded in from
|
||||||
`GSTACK_FLAKE_LEDGER`, the same env var the CI free lane sets before
|
`flakeLedgerPath()` — override with `GSTACK_FLAKE_LEDGER`, the same env var the
|
||||||
uploading the ledger as the `flake-ledger` artifact). Census integrity is
|
CI free lane sets before uploading the ledger as the `flake-ledger` artifact). Census integrity is
|
||||||
enforced from the free suite: every `E2E_TOUCHFILES` / `LLM_JUDGE_TOUCHFILES`
|
enforced from the free suite: every `E2E_TOUCHFILES` / `LLM_JUDGE_TOUCHFILES`
|
||||||
key must name a living paid test (`test/touchfiles.test.ts`'s reverse
|
key must name a living paid test (`test/touchfiles.test.ts`'s reverse
|
||||||
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
|
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
|
||||||
@@ -313,7 +309,7 @@ CI supplies the scoped cache/runtime configuration; local runs are fresh by
|
|||||||
default. Cached scores must
|
default. Cached scores must
|
||||||
pass current assertions; reused records retain their original source and time
|
pass current assertions; reused records retain their original source and time
|
||||||
and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts
|
and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts
|
||||||
to PR-profile E2E shards that run with zero retries (so the pass is structurally a
|
to PR-profile E2E shards (paid evals never retry, so a pass is structurally a
|
||||||
first attempt): the identity hashes the test's import closure, every tracked file
|
first attempt): the identity hashes the test's import closure, every tracked file
|
||||||
matched by the touchfiles of every case the file registers plus the global
|
matched by the touchfiles of every case the file registers plus the global
|
||||||
touchfiles, the runner/workflow/setup actions, the child's environment pins, the
|
touchfiles, the runner/workflow/setup actions, the child's environment pins, the
|
||||||
@@ -377,6 +373,118 @@ the runner parent and handed to shard children as `GSTACK_CLAUDE_CLI_VERSION`
|
|||||||
(never spawned on a test thread), so a TUI-drift flake hunt is a grep, not
|
(never spawned on a test thread), so a TUI-drift flake hunt is a grep, not
|
||||||
archaeology.
|
archaeology.
|
||||||
|
|
||||||
|
**Eval verdict policy** (`EVAL_POLICY` version 1 in
|
||||||
|
`test/helpers/periodic-exclude-data.ts`, pre-registered 2026-09-29). Paid evals
|
||||||
|
never retry. Each live case has exactly one kind in `E2E_KINDS`
|
||||||
|
(`test/helpers/touchfiles-data.ts`; `test/eval-kinds.test.ts` enforces coverage),
|
||||||
|
and the kind fixes its trials before the run:
|
||||||
|
|
||||||
|
- `rule` (default): one trial; any failed assertion fails the verdict. For
|
||||||
|
cases where nothing stochastic decides the verdict, or where it checks a
|
||||||
|
contract the product must meet every run.
|
||||||
|
- `behavior`: a panel of `n = 3` independent trials, launched together as
|
||||||
|
isolated case shards on different slices (key `<file>#<id>~t<N>`). All three
|
||||||
|
always run: no early stop and no conditional extra trial. PASS when at least
|
||||||
|
`k = 2` pass and no trial violated a contract (`expectContract()` stamps
|
||||||
|
`failure_class: 'contract'`). Each behavior case names its tolerated deviation
|
||||||
|
in `BEHAVIOR_WHY` and must have a literal registration so it can run alone.
|
||||||
|
- `judge`: an LLM judge scoring a fixed input. `judgePanel()`
|
||||||
|
(`test/helpers/llm-judge.ts`) draws 3 samples of the same prompt concurrently
|
||||||
|
inside the unchanged `JUDGE_MS`; numeric dimensions gate on the per-dimension
|
||||||
|
mean against the unchanged threshold (no dimension compensates for another),
|
||||||
|
booleans on a strict majority. A sample that errors (refusal, truncation,
|
||||||
|
non-JSON, a malformed field) fails the panel and is never resampled; a
|
||||||
|
refusal counts as an unscored panel only when every sample refused.
|
||||||
|
`callJudge`'s 429 backoff happens before any model output and is transport,
|
||||||
|
not a verdict retry. The workflow-judge cache stores whole panels only.
|
||||||
|
|
||||||
|
`panelVerdict()` (`test/helpers/eval-store.ts`) is the single verdict
|
||||||
|
function the report, `collector-outcomes.json`, the PR comment and pass-rates
|
||||||
|
all use. A timed-out, crashed or infrastructure-failed trial is a failed trial
|
||||||
|
recorded with its class; a missing or duplicate trial record makes the verdict
|
||||||
|
INCOMPLETE, which fails the lane; a 2/3 PASS is shown as `PASS 2/3` with the
|
||||||
|
failed trial's cause. A manual re-run adds trials under a new run attempt and
|
||||||
|
never replaces the first attempt's verdict. A red census is never rerun on
|
||||||
|
unchanged inputs: each red is diagnosed as product, test/detector, harness or
|
||||||
|
infra and resolved by a concrete repair and a fresh census, or listed as a named
|
||||||
|
red. The one exception: a census whose every red verdict is machine-classified
|
||||||
|
INFRA or INCOMPLETE (missing slice artifact, runner loss, API error before the
|
||||||
|
first model turn) may be re-dispatched once as a new run, and both runs are
|
||||||
|
reported. Changing any `EVAL_POLICY` constant after seeing census results needs
|
||||||
|
Garry's re-approval, a `version` bump and a fresh census;
|
||||||
|
`test/periodic-exclude-policy.test.ts` pins the approved values.
|
||||||
|
|
||||||
|
**Quarantine** (`CASE_QUARANTINE`, same file). An entry needs: a per-trial rate
|
||||||
|
below 95% over at least 10 post-policy trials of the case's current input
|
||||||
|
identity (pre-policy backfill may justify only an initial entry, labeled as
|
||||||
|
such); a written diagnosis in `reason` whose `failureClass` is `detector`,
|
||||||
|
`harness` or `model-latency` (a product defect is fixed or listed as a named
|
||||||
|
red, never quarantined); unchanged case touchfiles in the change that adds it;
|
||||||
|
and an owner, tracking pointer, `enteredAt` date and measurable `exit`. A
|
||||||
|
quarantined case still runs its full panel and reports in every lane but cannot
|
||||||
|
fail it, except on a hard break (0 of n) or a contract violation, and it never
|
||||||
|
counts as passing coverage. At most 10% of a blocking tier (gate, periodic) may
|
||||||
|
be quarantined. The weekly report fails when an entry passes its exit rule (at
|
||||||
|
least 97% over at least 10 trials) without being removed, when an entry is 8
|
||||||
|
weekly runs old, or when a tier is over its cap.
|
||||||
|
|
||||||
|
**Pass-rate history** (`bun run eval:pass-rates`, `scripts/eval-flake-rank.ts`).
|
||||||
|
It reads the `trial-outcomes` artifact of the last N completed
|
||||||
|
`evals-periodic.yml` runs on the current branch and `main` (flags: `--case`,
|
||||||
|
`--runs N`, `--branch`, `--dir`, `--backfill`, `--json`, `--gate`) and prints
|
||||||
|
per-case per-trial pass rates with 95% Wilson intervals. A series is one case
|
||||||
|
under one input identity, the hash of its own touchfiles minus
|
||||||
|
`GLOBAL_TOUCHFILES` (harness edits do not restart it), per model, Claude CLI
|
||||||
|
version and policy version; a change starts a new series and older ones stay
|
||||||
|
visible. Labels: INCONCLUSIVE below 10 trials, BROKEN when the latest run is
|
||||||
|
0/n after a prior interval at or above 95%, FLAKY when failures leave the
|
||||||
|
interval straddling 95%, FAILING when the whole interval is below it, PASSING
|
||||||
|
otherwise. `--backfill` imports legacy slice artifacts as pre-policy trials
|
||||||
|
(first attempt only; a record that names no registry id is listed as
|
||||||
|
unattributed, never guessed); they are display-only. `--gate` (the weekly
|
||||||
|
report) fails with ACTION REQUIRED, on post-policy trials of the current series
|
||||||
|
only, when a non-quarantined blocking case meets the entry rule (proposing an
|
||||||
|
entry), when a `rule` case does (rule case behaving like behavior: fix or
|
||||||
|
reclassify), when a blocking case's current identity is significantly below its
|
||||||
|
previous one (one-sided Fisher exact, α = 0.05, at least 6 trials each side,
|
||||||
|
Holm-controlled across the cases tested), and on the quarantine rules above.
|
||||||
|
History that cannot be fetched fails the gate closed.
|
||||||
|
|
||||||
|
**The arithmetic.** With per-trial pass rate p, the chance a single case goes
|
||||||
|
red (a false red while the product works, the catch rate once it has
|
||||||
|
regressed):
|
||||||
|
|
||||||
|
| p | 1 trial | 2-of-3 panel |
|
||||||
|
|---|---|---|
|
||||||
|
| 0.99 | 1.0% | 0.03% |
|
||||||
|
| 0.95 | 5.0% | 0.72% |
|
||||||
|
| 0.90 | 10.0% | 2.8% |
|
||||||
|
| 0.70 | 30.0% | 21.6% |
|
||||||
|
| 0.30 | 70.0% | 78.4% |
|
||||||
|
|
||||||
|
The panel removes most false reds at healthy rates, but it catches a 0.95 → 0.70
|
||||||
|
regression in one run only 21.6% of the time (a single trial 30%, retry-until-green
|
||||||
|
3%), so drift detection is the history rule's job, not the per-run verdict's.
|
||||||
|
The Fisher alarm is weak at the minimum sample (5.4% power for 0.95 → 0.70 at
|
||||||
|
6 trials a side), and ten straight passes still leave a 72% Wilson lower bound:
|
||||||
|
after this policy lands, every series starts INCONCLUSIVE.
|
||||||
|
|
||||||
|
A lane is all green with probability Π p_rule × Π P(≥2 of 3 | p_behavior) ×
|
||||||
|
Π p_judge. For the current registry (PR gate worst case: 107 rule cases and 24
|
||||||
|
judges; weekly census: 190 rule, 22 behavior and 25 judge verdicts), with rule
|
||||||
|
and judge verdicts at p_rule:
|
||||||
|
|
||||||
|
| p_rule | full PR gate | weekly, behavior p = 0.90 | 0.95 | 0.97 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 0.99 | 26.8% | 6.2% | 9.8% | 10.9% |
|
||||||
|
| 0.995 | 51.9% | 18.2% | 29.0% | 32.1% |
|
||||||
|
| 0.999 | 87.7% | 43.2% | 68.7% | 76.1% |
|
||||||
|
|
||||||
|
The rule term dominates: a green lane on a working product needs rule cases to
|
||||||
|
be near-deterministic (0.999), which is why failing detectors are converted to
|
||||||
|
outcome checks and product defects are fixed or named, and why each census
|
||||||
|
reports its expected lane false-red from the measured rates.
|
||||||
|
|
||||||
**Timeout policy.** Paid tests use the tiers in
|
**Timeout policy.** Paid tests use the tiers in
|
||||||
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
|
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
|
||||||
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall
|
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall
|
||||||
|
|||||||
Reference in new issue
Block a user