- Slice, census and marathon artifacts carry -a<run_attempt>; reports
download them per artifact (no merge), so records never overwrite and a
re-run never replaces the first attempt's verdict.
- Planners pass --max-parallel for the capacity preflight (24/16 unchanged:
the refreshed periodic plan needs 24 slices, the gate census 12).
- PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized
failure block); the group_by(.name)|last recomputation is gone.
- Reports stamp series identities, upload trial-outcomes-* for history, and
shard logs upload always (a failed trial no longer reds its runner).
- Weekly report: headline + failure block of both lanes in the issue body,
the eval:pass-rates --gate step (fails closed without history), close the
issue on a green run, and UC-E1: when every red is machine-classified
INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency
group (redispatch_of), both runs reported.
- Planner budget mode (--slice-budget S --jobs J): recorded per-tier wall
times pack into as many ~9-minute executors as the work needs; the plan
records per-slice estimates and the CI job timeout (supervised worst case
+ 20 min). evals.yml and evals-periodic.yml derive matrix size and
timeout-minutes from it; max-parallel covers every slice at once.
- Case shards: plan/design/review-army/shared-libs(-paths) run one registered
case per process (<file>#<case id>, exact name pattern, exactly one case).
- Retry rule: a timed-out attempt is a verdict. Only files whose every case
budget is CAPTURE tier or shorter keep one retry; walls shrink to match.
- Marathon tier: positive selection, excluded from gate/periodic planners,
run by the new evals-marathon.yml (weekly + dispatch, fresh, own report).
- PR-lane E2E reuse of verified first-attempt passes on identical inputs;
the report rejects reuse outside the fast PR profile.
- Duration seed from census run 36385945043, per tier and per case shard.