From 0e9345aadee693144de8b11073eba050b2c50f68 Mon Sep 17 00:00:00 2001 From: garrytan Date: Tue, 29 Sep 2026 16:19:52 +0000 Subject: [PATCH] docs: blocking lane budget, marathon lane, retry policy and E2E reuse --- AGENTS.md | 6 ++- CONTRIBUTING.md | 27 +++++++++- docs/TESTING_INTERNALS.md | 106 ++++++++++++++++++++++++-------------- 3 files changed, 98 insertions(+), 41 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 8865ab565..a68337ffd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -208,7 +208,10 @@ When fixing failures or preparing `/ship`, follow this order: result and pending permission state; diagnose a blocked actor before waiting through its deadline. Preserve cancellation separately from a test verdict. Skipped or unstarted cases - do not satisfy coverage; preserve configured retries and every attempt. + do not satisfy coverage; preserve every attempt. Retries follow the approved + policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so + only files whose every case budget is CAPTURE tier or shorter keep one retry; + never add retries to pass a longer case. 7. Prove all known repairs with focused tests, including affected paid cases. Rerun a failed case only after a concrete repair or a demonstrated launch correction. Run the remaining required selected evaluations on the integrated @@ -242,6 +245,7 @@ bun run test # complete free suite via the strict shard runner (no A bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY) bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals bun run eval:bg:release # fresh complete gate + periodic live coverage +bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free) bun run test:windows # curated Windows-safe subset (runs on windows-latest) bun run build # generate docs + compile binaries bun run gen:skill-docs # regenerate SKILL.md files from templates diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 5722e33ba..db9f74c8e 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -228,9 +228,32 @@ gate and periodic censuses run fresh weekly and on manual dispatch of `evals-periodic.yml`; `bun run eval:bg:release` runs both locally. Some broad behavioral failures will therefore be found after the PR gate. +Blocking paid lanes (the PR gate and the weekly periodic + gate census) aim to +finish in about 12 minutes including setup. The planner packs recorded wall +times (`scripts/paid-test-durations.json`, per tier) into as many ~9-minute +runners as the work needs, one file or a tightly packed group each; files whose +cases are short but whose total is long run one case per runner. Matrix size and +job timeout come from that plan. Preview it for free with +`bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2`. +Complete start-to-finish flows belong to the `marathon` tier +(`describeE2ETier('marathon')`), which runs only in the non-blocking +`evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge. + +Retries: a timed-out attempt is a verdict. Only files whose every case budget is +CAPTURE tier or shorter (`RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`) +keep one automatic retry for fast-failing flakes; every other paid file runs once. +Case budgets themselves never change with this rule. + CI enables verified first-attempt reuse for 16 workflow quality judges for -24 hours within the same PR. The cookie workflow's custom input, the other 11 -quality cases and all dynamic agent cases stay fresh. Local runs stay fresh unless +24 hours within the same PR. The cookie workflow's custom input and the other 11 +quality cases stay fresh. PR-profile E2E shards that run once (no retry, so the +pass is provably a first attempt) reuse a pass from the same PR when every +consumed input is byte-identical: the test's import closure, every tracked file +its registered cases' touchfiles and the global touchfiles match, the runner and +workflow, the child's EVALS_/GSTACK_/CLAUDE_/ANTHROPIC_ environment (secret +presence only), the CI image and Claude CLI version (`scripts/e2e-shard-reuse.ts`). +A computed case registration or a touchfile pattern matching nothing keeps the +shard fresh. The weekly census, marathon and release lanes never reuse. Local runs stay fresh unless the complete scoped cache and runtime configuration is supplied. The key includes complete prompt bytes, generated inputs, fixtures, runner/rubric code, installed dependencies, model settings and runtime. The current assertions validate a reused score again. Records retain the original diff --git a/docs/TESTING_INTERNALS.md b/docs/TESTING_INTERNALS.md index 046caf712..36d89ff3c 100644 --- a/docs/TESTING_INTERNALS.md +++ b/docs/TESTING_INTERNALS.md @@ -252,8 +252,16 @@ processes × `EVALS_CONCURRENCY` within-shard, per-shard `GSTACK_EVAL_DIR`, full-stream spooling to per-shard log files (path printed at START and on failure), never-started/timed-out taxonomy, and parent-computed diff selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a -child that can't parse it recomputes locally with one warning). Retry parity -lives in `RETRY_OVERRIDES` (literals; old matrix rows' earned `retries: 2`). +child that can't parse it recomputes locally with one warning). Retries follow +one rule (`retriesForFiles`, `RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`): +a timed-out attempt is a verdict, so a file keeps one Bun retry only when every +case budget is CAPTURE tier plus its recording grace or shorter (registered rows +derive it from `caseMs`, `SHORT_CASE_RETRY_FILES` lists the rest); every other +file, including the former `retries: 2` matrix rows, runs once. `--list` prints +each shard's retries. Files in `CASE_SHARDED_FILES` run one registered case per +process (`#`, an exact `--test-name-pattern`, exactly one executed +case), so a long file of short cases spreads across runners and each case gets +its own SDK semaphore. Flake telemetry rides the store: every recorded test carries its 1-based `attempt` (a pass-on-attempt-2 stays visible forever — bun's own stream hides it), runs list `flaky_retries`, the report warns on passed-only-on-retry @@ -304,8 +312,17 @@ does not match the cache adapter and stays fresh, as do the other 11 quality cas CI supplies the scoped cache/runtime configuration; local runs are fresh by default. Cached scores must pass current assertions; reused records retain their original source and time -and cannot renew the receipt. Dynamic live-agent runs are currently ineligible. -`EVALS_FRESH=1`, periodic and release validation bypass both lookup and publishing. +and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts +to PR-profile E2E shards that run with zero retries (so the pass is structurally a +first attempt): the identity hashes the test's import closure, every tracked file +matched by the touchfiles of every case the file registers plus the global +touchfiles, the runner/workflow/setup actions, the child's environment pins, the +CI image and Claude CLI version, and the shard's case ids, pattern, wall and +concurrency. A computed registration, an unmatched touchfile pattern, a retrying +file, a preload option or a custom endpoint makes the shard ineligible. A reused +shard reports `reused` with its source run and writes `execution: "reused"` +collector records; the report rejects reused outcomes outside the fast PR profile. +`EVALS_FRESH=1`, periodic, marathon and release validation bypass both lookup and publishing. **Free test timing and isolation.** `test:quick` is an explicitly partial measured subset for edit feedback. `test` remains complete local acceptance with its @@ -318,18 +335,24 @@ the entire lane, not separately to every machine. Refresh the full timing list with `bun run test:ubicloud --record-durations`. Profiling records failures faithfully and is separate from final release acceptance. -**CI planner/executor/report.** `--emit-plan --slices K` computes -selection + the slice plan ONCE (killing per-slice selector divergence); +**CI planner/executor/report.** `--emit-plan --slice-budget S --jobs J` +(CI) or `--slices K` (local) computes selection + the slice plan ONCE (killing +per-slice selector divergence); `--plan --slice i` executors consume the manifest and write slice-result artifacts; `--report ` reconciles them FAIL-CLOSED (a slice whose artifact never landed, or a planned shard nobody reported, is a -failure). Slices start from the supervision baseline (registered long files -spread by budget, the rest round-robin), then are re-packed by the recorded -wall times in `scripts/paid-test-durations.json`: a file moves or swaps out of -the heaviest slice only if no slice's worst-case wall (`paidShardWallUpperBoundMs` -for 1–4 workers) rises above the baseline's maximum, so CI timeout coverage is -never weakened. Refresh the seed from a downloaded report directory with -`--report --write-durations`. Under `EVALS_ALL` the hollow-shard guard marks exit-0 shards with +failure). Budget mode (`packBySliceBudget`) places shards longest-recorded-first +into the fullest slice whose estimated wall on J FIFO workers stays within S +seconds, else a new slice; a shard with no recorded wall weighs the whole budget +(its own runner), a shard longer than the budget runs alone, and overlays keep +one final one-at-a-time slice. The manifest's `plan` records each slice estimate +and `ciTimeoutMinutes` (every slice's supervised worst case plus 20 minutes +setup); CI derives the matrix (`[range(1; .sliceCount + 1)]`) and job timeout +from it, and executors refuse an `EVALS_JOBS` other than the planned J and run +their shards longest first. `--slices K` keeps the supervised round-robin +baseline re-packed by recorded times for local runs. Durations are recorded per +tier (a file's gate and periodic cases differ); refresh one tier from a +downloaded report directory with `--report --write-durations`. Under `EVALS_ALL` the hollow-shard guard marks exit-0 shards with ZERO executed tests `passed-empty` (a failure) — census-health, not just test runs. evals.yml runs the sliced gate lane per PR — the ONLY paid lane since the legacy 17-row matrix (22.6 min/$21 per PR serialized ahead of the @@ -337,7 +360,12 @@ slices) was deleted after demonstrated parity; its `KNOWN_MATRIX_GAPS`/`KNOWN_TIER_UNSET` ratchets retired with it and `test/evals-workflow-wiring.test.ts` pins the surviving wiring (slice-count agreement, tier consistency, the shared register-skills composite with its -fail-fast verification loop). evals-periodic.yml runs ALL +fail-fast verification loop). Tier `marathon` (complete start-to-finish flows) +is selected positively: a file enters the marathon plan only when it declares +`describeE2ETier('marathon')` or registers a marathon-tier case, and the gate and +periodic planners exclude marathon-only files; `evals-marathon.yml` runs them +weekly and on dispatch, fresh, one file per runner, with its own fail-closed +report and tracking issue, and nothing requires it. evals-periodic.yml runs ALL periodic-tier files weekly (the coverage contract) minus the reasoned exclusions in `test/helpers/periodic-exclude-data.ts` (reason + tracking required per entry; removal re-activates the file), plus a weekly @@ -356,28 +384,31 @@ minus overhead and ratchets raw literals. Budget above the wall is fiction. No paid test may exceed the ordinary tiers. `FINDING_RETRY_BUDGETS` also registers the CEO split-overflow and Eng -multi-finding batching files. Each retains its 25-minute case deadline and one -retry in a 52-minute shard wall, including two minutes for cleanup. No per-case budget grows. Overlay wrappers +multi-finding batching files. Each retains its 25-minute case deadline and, as a +case past the retry cap, runs once in a 27-minute shard wall including two +minutes for cleanup. No per-case budget grows. Overlay wrappers have a 1,830-second minimum shard wall and run without Bun retries; see the [overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget. -The quality file reserves 7,180 seconds for all 28 cases and their existing -retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow +The quality file reserves its whole-file wall for every case and its one retry +(judge cases are under the retry cap), plus cleanup. Each still has 120 seconds of model work. Its 17 workflow judges own their deadline and abort signal, with five seconds for terminal recording inside a ten-second Bun grace; the other 11 retain their existing 120-second Bun timeout. Late responses cannot create records or cache passes. -The ship documentation file reserves 10,920 seconds for five 600-second cases and -eight 300-second fault cases, each with one retry, plus cleanup. The standalone +The ship documentation file reserves 5,520 seconds for five 600-second cases and +eight 300-second fault cases, run once (a 600-second case is past the retry cap), plus cleanup. The standalone documentation child retains its 600-second case. The five review/ship explorer cases reserve 3,270 seconds including their existing retry and finalization grace. These are whole-file supervision limits, not additional model work per case. -The shared-library path file reserves 3,720 seconds for its three serial -600-second cases, each with one retry, plus 120 seconds for cleanup. Its +The shared-library path file reserves 1,920 seconds for its three serial +600-second cases, run once, plus 120 seconds for cleanup; in CI each case runs as +its own shard with a 720-second wall (a registered file's case shard supervises +`caseMs` times its allowed attempts plus the reserve). Its registered budget keeps the file in its own shard and binds the expected wall to both the saved plan and the execution receipt; missing or stale budget -records fail reconciliation. Case deadlines, model budgets and retries do not grow. +records fail reconciliation. Case deadlines and model budgets do not grow. `resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver. Each registered finding file and each overlay wrapper requires its @@ -389,27 +420,26 @@ Planner entries and execution results record the effective wall, its source and policy identifier. Custom drivers must resolve each job instead of passing their ordinary 1800-second default as an explicit cap; their outer controller/detach wall must also cover the allocated work and cleanup. -The current paid census has 104 files: 46 gate-tier and 70 periodic-tier. +The paid census counts are printed by `--list` for each tier. `eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR wrapper covers a full-gate fallback at its default two workers. The broad gate wrapper reserves 49320 seconds, and release reserves 116700 seconds for both -tiers. Legacy monolithic +tiers; free tests recompute each floor from the live shard census, case shards +included. Legacy monolithic `eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not promise every registered retry; use the sharded periodic path for this policy. -Periodic CI plans `--slices 7`. When overlays are selected, the seventh is -reserved for their serial wrappers; registered finding files are distributed -across the remaining ordinary slices by their supervised walls. Each slice job -has a 360-minute cap. Reconciliation rejects missing, duplicated or misplaced -registered work and absent budget records. The weekly gate census has a -352-minute cap across seven single-worker slices with at most four running at -once. Its longest current work wall is 272 minutes. PR slices retain seven -two-worker slices with a 265-minute cap for their 212-minute work wall plus -setup. Free supervision tests -verify these bounds against the complete current census, configured retries, -and setup reserve. Ordinary paid tiers and the default 1800-second -shard wall remain unchanged; the registered and overlay policies above supply -exceptions, and unregistered over-ceiling tests still fail policy checks. +CI plans with `--slice-budget 540 --jobs 2` for the PR gate, the periodic census +and the weekly gate census (the gate census also `--skip-judges`), and +`--slice-budget 1 --jobs 1` (one file per runner) for marathon. The live plans +must fit their workflow's `max-parallel` so every slice starts at once, and +`ciTimeoutMinutes` must stay within 360; `test/evals-workflow-wiring.test.ts` +recomputes both from the complete census. Reconciliation rejects missing, +duplicated or misplaced registered work, absent budget records, case shards that +did not execute exactly their case, and reused results outside the PR profile. +Ordinary paid tiers and the default 1800-second shard wall remain unchanged; the +registered and overlay policies above supply exceptions, and unregistered +over-ceiling tests still fail policy checks. Session timeouts are two-phase: a silent API dies at the startup grace (90s local / 300s CI floor, distinct exit reason `timeout_startup`) and the work