mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
test: retire the finding-count cluster and trim its helpers (C)
- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
budget, never on skill behavior; delete them, their touchfile/tier ids,
AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
(11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
closure, and delete the free replay tests whose assertions exercised only
that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
only as input for a live subject keep their assertions: the multiSelect
default moved to plan-review-decisions, runner PTY tests use inline caller
policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
(its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
1 parent
53e7f3212f
commit
6415690a18
317 files changed
+315
-54111
No files matched your search
@@ -825,6 +825,19 @@ audit trail lives in Aside.
|
||||
|
||||
## Test infrastructure
|
||||
|
||||
### P3: No paid eval runs the full /autoplan chain
|
||||
|
||||
**What:** `skill-e2e-autoplan-chain` was retired (it never reached a product
|
||||
verdict: launch failures, then 85-minute budget overruns). Phase order is still
|
||||
enforced by `autoplan/bin/phase-publication-hook.ts` and pinned by the free
|
||||
`test/autoplan-publication-guard.test.ts`, and `skill-e2e-autoplan-dual-voice`
|
||||
covers CEO Phase 1 dispatch. Nothing proves a live model completes
|
||||
CEO → Design → DX → Eng or reads the required phase sections
|
||||
(`CARVE_GUARDS.autoplan` is `behavioral: 'none'`).
|
||||
|
||||
**Re-entry:** a chain eval that fits the ordinary PTY tiers, for example one that
|
||||
runs the no-UI, no-DX path (CEO then Eng) and asserts the section reads.
|
||||
|
||||
### P3: CI-unrunnable paid evals
|
||||
|
||||
**What:** Seven paid files cannot execute in the CI image (no `codex` CLI, no
|
||||
@@ -1972,6 +1985,13 @@ plus a TTL so abandoned PTYs eventually exit.
|
||||
**Priority:** P2.
|
||||
**Effort:** S (CC: ~30 min once fixture exists). Captured from v1.21.1.0 plan-eng-review D2.
|
||||
|
||||
**Status (2026-09):** The four `skill-e2e-plan-*-finding-count` evals were retired
|
||||
after eight red weekly runs whose failures were harness and budget, not skill
|
||||
behavior. The `*-finding-floor` evals assert at least one AskUserQuestion, not one
|
||||
per finding, so this contract has no paid coverage today. Re-entry test: a
|
||||
qid-keyed per-finding count on a multi-finding fixture with `QUESTION_TUNING: true`
|
||||
(the `<gstack-qid:…>` markers only appear with tuning on).
|
||||
|
||||
---
|
||||
|
||||
## P3: Honor env vars in gstack-config (so QUESTION_TUNING/EXPLAIN_LEVEL actually isolate tests)
|
||||
|
||||
Reference in new issue
Block a user