test: retire the finding-count cluster and trim its helpers (C)

- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
  skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
  budget, never on skill behavior; delete them, their touchfile/tier ids,
  AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
  (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
  closure, and delete the free replay tests whose assertions exercised only
  that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
  only as input for a live subject keep their assertions: the multiSelect
  default moved to plan-review-decisions, runner PTY tests use inline caller
  policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
  (its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
  and per-finding cadence coverage with their re-entry tests.
This commit is contained in:
garrytan committed 2026-09-29 06:08:49 +00:00
1 parent 53e7f3212f
commit 6415690a18
317 files changed
+315 -54111

No files matched your search

+6 -6
View File
@@ -4,7 +4,7 @@ name: Periodic Evals
# tests can't rot invisibly — the class where the autoplan-dual-voice E2E was
# silently broken for months until a lucky local diff selected it. Engine:
# scripts/test-paid-shards.ts (the same runner local eval:bg:periodic uses):
# one planner manifest, 6 ordinary slices plus overlay and Autoplan slices, and a FAIL-CLOSED report — a slice
# one planner manifest, 6 ordinary slices plus an overlay slice, and a FAIL-CLOSED report — a slice
# whose artifact never landed is a failure, not an absence. The gate-census
# job is the weekly EVALS_ALL backstop for the gate tier (PR lanes are
# diff-billed, so without it the full gate census might never execute
@@ -96,7 +96,7 @@ jobs:
- name: Emit run manifest (ALL periodic tests minus reasoned excludes)
env:
EVALS_ALL: "1"
run: EVALS_TIER=periodic bun --no-install run scripts/test-paid-shards.ts --tier periodic --emit-plan /tmp/paid-plan/manifest.json --slices 8 --autoplan-slice
run: EVALS_TIER=periodic bun --no-install run scripts/test-paid-shards.ts --tier periodic --emit-plan /tmp/paid-plan/manifest.json --slices 7
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
@@ -118,8 +118,8 @@ jobs:
eval-slices:
runs-on: ubicloud-standard-8
needs: [build-image, plan-slices]
# Eight slices retain every registered case and retry. The complete
# census needs at most 338 minutes per slice, plus 20 minutes setup/upload.
# Seven slices retain every registered case and retry. The complete
# census needs at most 251 minutes per slice, plus 20 minutes setup/upload.
timeout-minutes: 358
permissions:
contents: read
@@ -133,7 +133,7 @@ jobs:
strategy:
fail-fast: false
matrix:
slice: [1, 2, 3, 4, 5, 6, 7, 8]
slice: [1, 2, 3, 4, 5, 6, 7]
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
@@ -171,7 +171,7 @@ jobs:
name: paid-plan
path: /tmp/paid-plan
- name: Run slice ${{ matrix.slice }}/8
- name: Run slice ${{ matrix.slice }}/7
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}