test: retire the finding-count cluster and trim its helpers (C)

- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
  skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
  budget, never on skill behavior; delete them, their touchfile/tier ids,
  AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
  (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
  closure, and delete the free replay tests whose assertions exercised only
  that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
  only as input for a live subject keep their assertions: the multiSelect
  default moved to plan-review-decisions, runner PTY tests use inline caller
  policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
  (its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
  and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
garrytan committed 2026-09-29 06:08:49 +00:00
1 parent 53e7f3212f
commit 6415690a18
317 files changed
+315 -54111

No files matched your search

+12 -30
View File
@@ -41,7 +41,7 @@ Seeded planning sessions also receive an isolated runtime home through
to the working tree under test. Explicit per-test home overrides remain intact.
Autoplan resolves each review skill from its own installed host registry.
**Interactive planning evidence.** Finding-count and autoplan-chain drivers use
**Interactive planning evidence.** Native plan-review count drivers use
`observeScreen: true` and await `currentScreen()` before choosing an input. The
existing xterm dependency interprets cursor moves and erases; old menus in the
raw stream cannot establish a current prompt. Snapshots preserve
@@ -300,28 +300,11 @@ archaeology.
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall
minus overhead and ratchets raw literals. Budget above the wall is fiction.
The registered four-phase exception is `AUTOPLAN_CHAIN_BUDGET` for
`test/skill-e2e-autoplan-chain.test.ts`: 80 minutes of work (four `PTY_LONG`
allocations), an 84-minute session watchdog, an 85-minute Bun test deadline,
and a 172-minute supervised shard wall. The unchanged retry count of one
permits two 85-minute attempts plus two minutes for cleanup. This is a
**specified allocation for the stronger four-phase contract**, not a measured
calibration or statistical upper bound. The historical 900-second failures
remain failures. Models, fixtures, phase assertions and production review
caller timeouts are unchanged; this explicitly changes eval latency/cost policy.
No paid test may exceed the ordinary tiers.
The Autoplan chain explicitly enables native `PreToolUse` approval for edits to
its owned temporary review artifacts. Approval starts with the `/autoplan`
command and requires the exact parent session, prior successful file history,
and a current request digest. Other recorder callers remain observational.
A rejected artifact edit fails the test instead of falling through to terminal
permission input. Approval itself supplies no edit success or phase credit:
the native tool result and all four completed review phases are still required.
`FINDING_RETRY_BUDGETS` also registers six finding files. Each retains its
25-minute case deadline and one retry: the two-case CEO finding-count file has
a 102-minute shard wall, and the five single-case files have 52-minute walls,
including two minutes for cleanup. No per-case budget grows. Overlay wrappers
`FINDING_RETRY_BUDGETS` also registers the CEO split-overflow and Eng
multi-finding batching files. Each retains its 25-minute case deadline and one
retry in a 52-minute shard wall, including two minutes for cleanup. No per-case budget grows. Overlay wrappers
have a 1,830-second minimum shard wall and run without Bun retries; see the
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
@@ -332,27 +315,26 @@ recording inside a ten-second Bun grace; the other 11 retain their existing
120-second Bun timeout. Late responses cannot create records or cache passes.
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
Autoplan, each registered finding file, and each overlay wrapper require their
Each registered finding file and each overlay wrapper requires its
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
jobs are rejected so ordinary files retain their configured retries. An explicit
CLI `--timeout`, `EVALS_SHARD_TIMEOUT_MS`, or API `timeoutMs` still wins for these
policies, including a lower cap; overlay overrides below their minimum are rejected.
Planner entries and execution results record the effective wall,
its source and policy identifier. Custom drivers must resolve each job instead
of passing their ordinary 1800-second default as an explicit Autoplan cap;
of passing their ordinary 1800-second default as an explicit cap;
their outer controller/detach wall must also cover the allocated work and cleanup.
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
wrapper covers a full-gate fallback at its default two workers. The broad gate
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
tiers. Legacy monolithic
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
promise every registered retry; use the sharded periodic path for this policy.
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
When overlays are selected, the seventh is reserved for their serial wrappers;
registered finding files are distributed across the remaining ordinary slices
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
Periodic CI plans `--slices 7`. When overlays are selected, the seventh is
reserved for their serial wrappers; registered finding files are distributed
across the remaining ordinary slices by their supervised walls. Each slice job
has a 358-minute cap. Reconciliation rejects missing, duplicated or misplaced
registered work and absent budget records. The weekly gate census has a
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
verify these bounds against the complete current census, configured retries,