mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 01:46:55 +02:00
docs: blocking lane budget, marathon lane, retry policy and E2E reuse
This commit is contained in:
1 parent
0023d011a3
commit
0e9345aade
3 files changed
+98
-41
No files matched your search
@@ -208,7 +208,10 @@ When fixing failures or preparing `/ship`, follow this order:
|
||||
result and pending permission state; diagnose a blocked actor before waiting
|
||||
through its deadline. Preserve cancellation separately from a test verdict.
|
||||
Skipped or unstarted cases
|
||||
do not satisfy coverage; preserve configured retries and every attempt.
|
||||
do not satisfy coverage; preserve every attempt. Retries follow the approved
|
||||
policy in `test/helpers/eval-budgets.ts`: a timed-out attempt is a verdict, so
|
||||
only files whose every case budget is CAPTURE tier or shorter keep one retry;
|
||||
never add retries to pass a longer case.
|
||||
7. Prove all known repairs with focused tests, including affected paid cases.
|
||||
Rerun a failed case only after a concrete repair or a demonstrated launch
|
||||
correction. Run the remaining required selected evaluations on the integrated
|
||||
@@ -242,6 +245,7 @@ bun run test # complete free suite via the strict shard runner (no A
|
||||
bun run test:ubicloud # same suite on an ephemeral 16-vCPU Ubicloud VM (needs UBICLOUD_API_KEY)
|
||||
bun run eval:bg:pr # changed fast live probes + selected judges, with explicit deferrals
|
||||
bun run eval:bg:release # fresh complete gate + periodic live coverage
|
||||
bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2 # CI slice plan preview (free)
|
||||
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
|
||||
bun run build # generate docs + compile binaries
|
||||
bun run gen:skill-docs # regenerate SKILL.md files from templates
|
||||
|
||||
+25
-2
@@ -228,9 +228,32 @@ gate and periodic censuses run fresh weekly and on manual
|
||||
dispatch of `evals-periodic.yml`; `bun run eval:bg:release` runs both locally.
|
||||
Some broad behavioral failures will therefore be found after the PR gate.
|
||||
|
||||
Blocking paid lanes (the PR gate and the weekly periodic + gate census) aim to
|
||||
finish in about 12 minutes including setup. The planner packs recorded wall
|
||||
times (`scripts/paid-test-durations.json`, per tier) into as many ~9-minute
|
||||
runners as the work needs, one file or a tightly packed group each; files whose
|
||||
cases are short but whose total is long run one case per runner. Matrix size and
|
||||
job timeout come from that plan. Preview it for free with
|
||||
`bun run scripts/test-paid-shards.ts --tier periodic --list --slice-budget 540 --jobs 2`.
|
||||
Complete start-to-finish flows belong to the `marathon` tier
|
||||
(`describeE2ETier('marathon')`), which runs only in the non-blocking
|
||||
`evals-marathon.yml` lane (weekly and on dispatch) and never gates a merge.
|
||||
|
||||
Retries: a timed-out attempt is a verdict. Only files whose every case budget is
|
||||
CAPTURE tier or shorter (`RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`)
|
||||
keep one automatic retry for fast-failing flakes; every other paid file runs once.
|
||||
Case budgets themselves never change with this rule.
|
||||
|
||||
CI enables verified first-attempt reuse for 16 workflow quality judges for
|
||||
24 hours within the same PR. The cookie workflow's custom input, the other 11
|
||||
quality cases and all dynamic agent cases stay fresh. Local runs stay fresh unless
|
||||
24 hours within the same PR. The cookie workflow's custom input and the other 11
|
||||
quality cases stay fresh. PR-profile E2E shards that run once (no retry, so the
|
||||
pass is provably a first attempt) reuse a pass from the same PR when every
|
||||
consumed input is byte-identical: the test's import closure, every tracked file
|
||||
its registered cases' touchfiles and the global touchfiles match, the runner and
|
||||
workflow, the child's EVALS_/GSTACK_/CLAUDE_/ANTHROPIC_ environment (secret
|
||||
presence only), the CI image and Claude CLI version (`scripts/e2e-shard-reuse.ts`).
|
||||
A computed case registration or a touchfile pattern matching nothing keeps the
|
||||
shard fresh. The weekly census, marathon and release lanes never reuse. Local runs stay fresh unless
|
||||
the complete scoped cache and runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
|
||||
fixtures, runner/rubric code, installed dependencies, model settings and runtime.
|
||||
The current assertions validate a reused score again. Records retain the original
|
||||
|
||||
+68
-38
@@ -252,8 +252,16 @@ processes × `EVALS_CONCURRENCY` within-shard, per-shard `GSTACK_EVAL_DIR`,
|
||||
full-stream spooling to per-shard log files (path printed at START and on
|
||||
failure), never-started/timed-out taxonomy, and parent-computed diff
|
||||
selection propagated to children via `EVALS_SELECTION_JSON` (fail-open: a
|
||||
child that can't parse it recomputes locally with one warning). Retry parity
|
||||
lives in `RETRY_OVERRIDES` (literals; old matrix rows' earned `retries: 2`).
|
||||
child that can't parse it recomputes locally with one warning). Retries follow
|
||||
one rule (`retriesForFiles`, `RETRY_MAX_CASE_MS` in `test/helpers/eval-budgets.ts`):
|
||||
a timed-out attempt is a verdict, so a file keeps one Bun retry only when every
|
||||
case budget is CAPTURE tier plus its recording grace or shorter (registered rows
|
||||
derive it from `caseMs`, `SHORT_CASE_RETRY_FILES` lists the rest); every other
|
||||
file, including the former `retries: 2` matrix rows, runs once. `--list` prints
|
||||
each shard's retries. Files in `CASE_SHARDED_FILES` run one registered case per
|
||||
process (`<file>#<case id>`, an exact `--test-name-pattern`, exactly one executed
|
||||
case), so a long file of short cases spreads across runners and each case gets
|
||||
its own SDK semaphore.
|
||||
Flake telemetry rides the store: every recorded test carries its 1-based
|
||||
`attempt` (a pass-on-attempt-2 stays visible forever — bun's own stream hides
|
||||
it), runs list `flaky_retries`, the report warns on passed-only-on-retry
|
||||
@@ -304,8 +312,17 @@ does not match the cache adapter and stays fresh, as do the other 11 quality cas
|
||||
CI supplies the scoped cache/runtime configuration; local runs are fresh by
|
||||
default. Cached scores must
|
||||
pass current assertions; reused records retain their original source and time
|
||||
and cannot renew the receipt. Dynamic live-agent runs are currently ineligible.
|
||||
`EVALS_FRESH=1`, periodic and release validation bypass both lookup and publishing.
|
||||
and cannot renew the receipt. `scripts/e2e-shard-reuse.ts` extends the same receipts
|
||||
to PR-profile E2E shards that run with zero retries (so the pass is structurally a
|
||||
first attempt): the identity hashes the test's import closure, every tracked file
|
||||
matched by the touchfiles of every case the file registers plus the global
|
||||
touchfiles, the runner/workflow/setup actions, the child's environment pins, the
|
||||
CI image and Claude CLI version, and the shard's case ids, pattern, wall and
|
||||
concurrency. A computed registration, an unmatched touchfile pattern, a retrying
|
||||
file, a preload option or a custom endpoint makes the shard ineligible. A reused
|
||||
shard reports `reused` with its source run and writes `execution: "reused"`
|
||||
collector records; the report rejects reused outcomes outside the fast PR profile.
|
||||
`EVALS_FRESH=1`, periodic, marathon and release validation bypass both lookup and publishing.
|
||||
|
||||
**Free test timing and isolation.** `test:quick` is an explicitly partial measured
|
||||
subset for edit feedback. `test` remains complete local acceptance with its
|
||||
@@ -318,18 +335,24 @@ the entire lane, not separately to every machine. Refresh the full timing list
|
||||
with `bun run test:ubicloud --record-durations`. Profiling records failures faithfully
|
||||
and is separate from final release acceptance.
|
||||
|
||||
**CI planner/executor/report.** `--emit-plan <path> --slices K` computes
|
||||
selection + the slice plan ONCE (killing per-slice selector divergence);
|
||||
**CI planner/executor/report.** `--emit-plan <path> --slice-budget S --jobs J`
|
||||
(CI) or `--slices K` (local) computes selection + the slice plan ONCE (killing
|
||||
per-slice selector divergence);
|
||||
`--plan <path> --slice i` executors consume the manifest and write
|
||||
slice-result artifacts; `--report <dir>` reconciles them FAIL-CLOSED (a slice
|
||||
whose artifact never landed, or a planned shard nobody reported, is a
|
||||
failure). Slices start from the supervision baseline (registered long files
|
||||
spread by budget, the rest round-robin), then are re-packed by the recorded
|
||||
wall times in `scripts/paid-test-durations.json`: a file moves or swaps out of
|
||||
the heaviest slice only if no slice's worst-case wall (`paidShardWallUpperBoundMs`
|
||||
for 1–4 workers) rises above the baseline's maximum, so CI timeout coverage is
|
||||
never weakened. Refresh the seed from a downloaded report directory with
|
||||
`--report <dir> --write-durations`. Under `EVALS_ALL` the hollow-shard guard marks exit-0 shards with
|
||||
failure). Budget mode (`packBySliceBudget`) places shards longest-recorded-first
|
||||
into the fullest slice whose estimated wall on J FIFO workers stays within S
|
||||
seconds, else a new slice; a shard with no recorded wall weighs the whole budget
|
||||
(its own runner), a shard longer than the budget runs alone, and overlays keep
|
||||
one final one-at-a-time slice. The manifest's `plan` records each slice estimate
|
||||
and `ciTimeoutMinutes` (every slice's supervised worst case plus 20 minutes
|
||||
setup); CI derives the matrix (`[range(1; .sliceCount + 1)]`) and job timeout
|
||||
from it, and executors refuse an `EVALS_JOBS` other than the planned J and run
|
||||
their shards longest first. `--slices K` keeps the supervised round-robin
|
||||
baseline re-packed by recorded times for local runs. Durations are recorded per
|
||||
tier (a file's gate and periodic cases differ); refresh one tier from a
|
||||
downloaded report directory with `--report <dir> --write-durations`. Under `EVALS_ALL` the hollow-shard guard marks exit-0 shards with
|
||||
ZERO executed tests `passed-empty` (a failure) — census-health, not just
|
||||
test runs. evals.yml runs the sliced gate lane per PR — the ONLY paid lane
|
||||
since the legacy 17-row matrix (22.6 min/$21 per PR serialized ahead of the
|
||||
@@ -337,7 +360,12 @@ slices) was deleted after demonstrated parity; its
|
||||
`KNOWN_MATRIX_GAPS`/`KNOWN_TIER_UNSET` ratchets retired with it and
|
||||
`test/evals-workflow-wiring.test.ts` pins the surviving wiring (slice-count
|
||||
agreement, tier consistency, the shared register-skills composite with its
|
||||
fail-fast verification loop). evals-periodic.yml runs ALL
|
||||
fail-fast verification loop). Tier `marathon` (complete start-to-finish flows)
|
||||
is selected positively: a file enters the marathon plan only when it declares
|
||||
`describeE2ETier('marathon')` or registers a marathon-tier case, and the gate and
|
||||
periodic planners exclude marathon-only files; `evals-marathon.yml` runs them
|
||||
weekly and on dispatch, fresh, one file per runner, with its own fail-closed
|
||||
report and tracking issue, and nothing requires it. evals-periodic.yml runs ALL
|
||||
periodic-tier files weekly (the coverage contract) minus the reasoned
|
||||
exclusions in `test/helpers/periodic-exclude-data.ts` (reason + tracking
|
||||
required per entry; removal re-activates the file), plus a weekly
|
||||
@@ -356,28 +384,31 @@ minus overhead and ratchets raw literals. Budget above the wall is fiction.
|
||||
No paid test may exceed the ordinary tiers.
|
||||
|
||||
`FINDING_RETRY_BUDGETS` also registers the CEO split-overflow and Eng
|
||||
multi-finding batching files. Each retains its 25-minute case deadline and one
|
||||
retry in a 52-minute shard wall, including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
multi-finding batching files. Each retains its 25-minute case deadline and, as a
|
||||
case past the retry cap, runs once in a 27-minute shard wall including two
|
||||
minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
have a 1,830-second minimum shard wall and run without Bun retries; see the
|
||||
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
|
||||
|
||||
The quality file reserves 7,180 seconds for all 28 cases and their existing
|
||||
retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
|
||||
The quality file reserves its whole-file wall for every case and its one retry
|
||||
(judge cases are under the retry cap), plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
|
||||
judges own their deadline and abort signal, with five seconds for terminal
|
||||
recording inside a ten-second Bun grace; the other 11 retain their existing
|
||||
120-second Bun timeout. Late responses cannot create records or cache passes.
|
||||
|
||||
The ship documentation file reserves 10,920 seconds for five 600-second cases and
|
||||
eight 300-second fault cases, each with one retry, plus cleanup. The standalone
|
||||
The ship documentation file reserves 5,520 seconds for five 600-second cases and
|
||||
eight 300-second fault cases, run once (a 600-second case is past the retry cap), plus cleanup. The standalone
|
||||
documentation child retains its 600-second case. The five review/ship explorer
|
||||
cases reserve 3,270 seconds including their existing retry and finalization grace.
|
||||
These are whole-file supervision limits, not additional model work per case.
|
||||
|
||||
The shared-library path file reserves 3,720 seconds for its three serial
|
||||
600-second cases, each with one retry, plus 120 seconds for cleanup. Its
|
||||
The shared-library path file reserves 1,920 seconds for its three serial
|
||||
600-second cases, run once, plus 120 seconds for cleanup; in CI each case runs as
|
||||
its own shard with a 720-second wall (a registered file's case shard supervises
|
||||
`caseMs` times its allowed attempts plus the reserve). Its
|
||||
registered budget keeps the file in its own shard and binds the expected wall
|
||||
to both the saved plan and the execution receipt; missing or stale budget
|
||||
records fail reconciliation. Case deadlines, model budgets and retries do not grow.
|
||||
records fail reconciliation. Case deadlines and model budgets do not grow.
|
||||
|
||||
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
|
||||
Each registered finding file and each overlay wrapper requires its
|
||||
@@ -389,27 +420,26 @@ Planner entries and execution results record the effective wall,
|
||||
its source and policy identifier. Custom drivers must resolve each job instead
|
||||
of passing their ordinary 1800-second default as an explicit cap;
|
||||
their outer controller/detach wall must also cover the allocated work and cleanup.
|
||||
The current paid census has 104 files: 46 gate-tier and 70 periodic-tier.
|
||||
The paid census counts are printed by `--list` for each tier.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
|
||||
wrapper covers a full-gate fallback at its default two workers. The broad gate
|
||||
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
|
||||
tiers. Legacy monolithic
|
||||
tiers; free tests recompute each floor from the live shard census, case shards
|
||||
included. Legacy monolithic
|
||||
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
|
||||
promise every registered retry; use the sharded periodic path for this policy.
|
||||
|
||||
Periodic CI plans `--slices 7`. When overlays are selected, the seventh is
|
||||
reserved for their serial wrappers; registered finding files are distributed
|
||||
across the remaining ordinary slices by their supervised walls. Each slice job
|
||||
has a 360-minute cap. Reconciliation rejects missing, duplicated or misplaced
|
||||
registered work and absent budget records. The weekly gate census has a
|
||||
352-minute cap across seven single-worker slices with at most four running at
|
||||
once. Its longest current work wall is 272 minutes. PR slices retain seven
|
||||
two-worker slices with a 265-minute cap for their 212-minute work wall plus
|
||||
setup. Free supervision tests
|
||||
verify these bounds against the complete current census, configured retries,
|
||||
and setup reserve. Ordinary paid tiers and the default 1800-second
|
||||
shard wall remain unchanged; the registered and overlay policies above supply
|
||||
exceptions, and unregistered over-ceiling tests still fail policy checks.
|
||||
CI plans with `--slice-budget 540 --jobs 2` for the PR gate, the periodic census
|
||||
and the weekly gate census (the gate census also `--skip-judges`), and
|
||||
`--slice-budget 1 --jobs 1` (one file per runner) for marathon. The live plans
|
||||
must fit their workflow's `max-parallel` so every slice starts at once, and
|
||||
`ciTimeoutMinutes` must stay within 360; `test/evals-workflow-wiring.test.ts`
|
||||
recomputes both from the complete census. Reconciliation rejects missing,
|
||||
duplicated or misplaced registered work, absent budget records, case shards that
|
||||
did not execute exactly their case, and reused results outside the PR profile.
|
||||
Ordinary paid tiers and the default 1800-second shard wall remain unchanged; the
|
||||
registered and overlay policies above supply exceptions, and unregistered
|
||||
over-ceiling tests still fail policy checks.
|
||||
|
||||
Session timeouts are two-phase: a silent API dies at the startup grace (90s
|
||||
local / 300s CI floor, distinct exit reason `timeout_startup`) and the work
|
||||
|
||||
Reference in new issue
Block a user