mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 01:46:55 +02:00
test: retire the finding-count cluster and trim its helpers (C)
- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
budget, never on skill behavior; delete them, their touchfile/tier ids,
AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
(11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
closure, and delete the free replay tests whose assertions exercised only
that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
only as input for a live subject keep their assertions: the multiSelect
default moved to plan-review-decisions, runner PTY tests use inline caller
policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
(its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
1 parent
53e7f3212f
commit
6415690a18
317 files changed
+315
-54111
No files matched your search
+12
-30
@@ -41,7 +41,7 @@ Seeded planning sessions also receive an isolated runtime home through
|
||||
to the working tree under test. Explicit per-test home overrides remain intact.
|
||||
Autoplan resolves each review skill from its own installed host registry.
|
||||
|
||||
**Interactive planning evidence.** Finding-count and autoplan-chain drivers use
|
||||
**Interactive planning evidence.** Native plan-review count drivers use
|
||||
`observeScreen: true` and await `currentScreen()` before choosing an input. The
|
||||
existing xterm dependency interprets cursor moves and erases; old menus in the
|
||||
raw stream cannot establish a current prompt. Snapshots preserve
|
||||
@@ -300,28 +300,11 @@ archaeology.
|
||||
`test/helpers/eval-budgets.ts` (JUDGE/CAPTURE/CAPTURE_LONG/PTY/PTY_LONG);
|
||||
`test/eval-budgets-policy.test.ts` pins that every tier fits the shard wall
|
||||
minus overhead and ratchets raw literals. Budget above the wall is fiction.
|
||||
The registered four-phase exception is `AUTOPLAN_CHAIN_BUDGET` for
|
||||
`test/skill-e2e-autoplan-chain.test.ts`: 80 minutes of work (four `PTY_LONG`
|
||||
allocations), an 84-minute session watchdog, an 85-minute Bun test deadline,
|
||||
and a 172-minute supervised shard wall. The unchanged retry count of one
|
||||
permits two 85-minute attempts plus two minutes for cleanup. This is a
|
||||
**specified allocation for the stronger four-phase contract**, not a measured
|
||||
calibration or statistical upper bound. The historical 900-second failures
|
||||
remain failures. Models, fixtures, phase assertions and production review
|
||||
caller timeouts are unchanged; this explicitly changes eval latency/cost policy.
|
||||
No paid test may exceed the ordinary tiers.
|
||||
|
||||
The Autoplan chain explicitly enables native `PreToolUse` approval for edits to
|
||||
its owned temporary review artifacts. Approval starts with the `/autoplan`
|
||||
command and requires the exact parent session, prior successful file history,
|
||||
and a current request digest. Other recorder callers remain observational.
|
||||
A rejected artifact edit fails the test instead of falling through to terminal
|
||||
permission input. Approval itself supplies no edit success or phase credit:
|
||||
the native tool result and all four completed review phases are still required.
|
||||
|
||||
`FINDING_RETRY_BUDGETS` also registers six finding files. Each retains its
|
||||
25-minute case deadline and one retry: the two-case CEO finding-count file has
|
||||
a 102-minute shard wall, and the five single-case files have 52-minute walls,
|
||||
including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
`FINDING_RETRY_BUDGETS` also registers the CEO split-overflow and Eng
|
||||
multi-finding batching files. Each retains its 25-minute case deadline and one
|
||||
retry in a 52-minute shard wall, including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
have a 1,830-second minimum shard wall and run without Bun retries; see the
|
||||
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
|
||||
|
||||
@@ -332,27 +315,26 @@ recording inside a ten-second Bun grace; the other 11 retain their existing
|
||||
120-second Bun timeout. Late responses cannot create records or cache passes.
|
||||
|
||||
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
|
||||
Autoplan, each registered finding file, and each overlay wrapper require their
|
||||
Each registered finding file and each overlay wrapper requires its
|
||||
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
|
||||
jobs are rejected so ordinary files retain their configured retries. An explicit
|
||||
CLI `--timeout`, `EVALS_SHARD_TIMEOUT_MS`, or API `timeoutMs` still wins for these
|
||||
policies, including a lower cap; overlay overrides below their minimum are rejected.
|
||||
Planner entries and execution results record the effective wall,
|
||||
its source and policy identifier. Custom drivers must resolve each job instead
|
||||
of passing their ordinary 1800-second default as an explicit Autoplan cap;
|
||||
of passing their ordinary 1800-second default as an explicit cap;
|
||||
their outer controller/detach wall must also cover the allocated work and cleanup.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
|
||||
wrapper covers a full-gate fallback at its default two workers. The broad gate
|
||||
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
|
||||
tiers. Legacy monolithic
|
||||
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
|
||||
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
|
||||
promise every registered retry; use the sharded periodic path for this policy.
|
||||
|
||||
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
|
||||
When overlays are selected, the seventh is reserved for their serial wrappers;
|
||||
registered finding files are distributed across the remaining ordinary slices
|
||||
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
|
||||
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
|
||||
Periodic CI plans `--slices 7`. When overlays are selected, the seventh is
|
||||
reserved for their serial wrappers; registered finding files are distributed
|
||||
across the remaining ordinary slices by their supervised walls. Each slice job
|
||||
has a 358-minute cap. Reconciliation rejects missing, duplicated or misplaced
|
||||
registered work and absent budget records. The weekly gate census has a
|
||||
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
|
||||
verify these bounds against the complete current census, configured retries,
|
||||
|
||||
Reference in new issue
Block a user