test: retire the finding-count cluster and trim its helpers (C)

- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
  skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
  budget, never on skill behavior; delete them, their touchfile/tier ids,
  AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
  (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
  closure, and delete the free replay tests whose assertions exercised only
  that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
  only as input for a live subject keep their assertions: the multiSelect
  default moved to plan-review-decisions, runner PTY tests use inline caller
  policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
  (its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
  and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
garrytan committed 2026-09-29 06:08:49 +00:00
1 parent 53e7f3212f
commit 6415690a18
317 files changed
+315 -54111

No files matched your search

+20
View File
@@ -825,6 +825,19 @@ audit trail lives in Aside.
## Test infrastructure
### P3: No paid eval runs the full /autoplan chain
**What:** `skill-e2e-autoplan-chain` was retired (it never reached a product
verdict: launch failures, then 85-minute budget overruns). Phase order is still
enforced by `autoplan/bin/phase-publication-hook.ts` and pinned by the free
`test/autoplan-publication-guard.test.ts`, and `skill-e2e-autoplan-dual-voice`
covers CEO Phase 1 dispatch. Nothing proves a live model completes
CEO → Design → DX → Eng or reads the required phase sections
(`CARVE_GUARDS.autoplan` is `behavioral: 'none'`).
**Re-entry:** a chain eval that fits the ordinary PTY tiers, for example one that
runs the no-UI, no-DX path (CEO then Eng) and asserts the section reads.
### P3: CI-unrunnable paid evals
**What:** Seven paid files cannot execute in the CI image (no `codex` CLI, no
@@ -1972,6 +1985,13 @@ plus a TTL so abandoned PTYs eventually exit.
**Priority:** P2.
**Effort:** S (CC: ~30 min once fixture exists). Captured from v1.21.1.0 plan-eng-review D2.
**Status (2026-09):** The four `skill-e2e-plan-*-finding-count` evals were retired
after eight red weekly runs whose failures were harness and budget, not skill
behavior. The `*-finding-floor` evals assert at least one AskUserQuestion, not one
per finding, so this contract has no paid coverage today. Re-entry test: a
qid-keyed per-finding count on a multi-finding fixture with `QUESTION_TUNING: true`
(the `<gstack-qid:…>` markers only appear with tuning on).
---
## P3: Honor env vars in gstack-config (so QUESTION_TUNING/EXPLAIN_LEVEL actually isolate tests)