v1.91.8.0 test: smaller suite, derived paid selection, retired never-green evals

Release metadata for the test-reduction branch: VERSION 1.91.8.0 (1.91.7.0 is
claimed by #2983), CHANGELOG with the measured before/after table and a
contributor section, durations re-recorded on Ubicloud standard-16 (857 files,
0 failures), the agents digest, CONTRIBUTING's after-measurement row, the B8
fallback TODOS entry, and the after metrics, kept-vs-plan notes, B8 run and
census estimate in docs/test-audit-2026-09.md.
This commit is contained in:
garrytan committed 2026-09-29 09:18:02 +00:00
1 parent 13e4f37201
commit a2a239b1b3
9 files changed
+961 -1059

No files matched your search

+39
View File
@@ -1,5 +1,44 @@
# Changelog
## [1.91.8.0] - 2026-09-29
The test suite is smaller and every remaining test maps to a product contract: 227 fewer test files, about 90,000 fewer lines of tests, helpers and fixtures, and the weekly paid lane drops the five evals that were red eight runs straight. Free tests that only replayed one captured failure are folded into their detector's owner test, and paid eval selection is derived from each eval's own imports instead of hand-copied lists.
| Measure | Before (v1.91.6.0) | After |
| --- | ---: | ---: |
| Tracked test files | 1,184 | 957 |
| Test-file lines (all `*.test.ts`) | 279,640 | 244,504 |
| `test/` TypeScript lines (tests + helpers) | 274,208 | 227,713 |
| `test/helpers` lines | 51,390 | 39,427 |
| `test/fixtures` bytes | 16.3 MB | 9.7 MB |
| Free suite files / passing tests (Ubicloud standard-16) | 1,065 / 27,331 | 857 / 20,302 |
| Free suite serial seconds (recorded durations, same machine class) | 1,888 | 1,738 |
| Paid files / gate-lane files / periodic files | 119 / 58 / 100 | 100 / 42 / 69 |
| Weekly gate-census files | 58 | 41 (LLM judges run in the periodic and PR lanes) |
| Weekly periodic shard-minutes spent on files this release removes (09-21 run) | 235 of 462 | 0 |
### Removed
- The never-green finding-count cluster: `skill-e2e-autoplan-chain` and `skill-e2e-plan-{ceo,eng,design,devex}-finding-count`, whose weekly failures were harness and budget failures, never skill behavior (triage in `docs/test-audit-2026-09.md`). No paid eval now proves a live model completes the full `/autoplan` chain or asks one question per finding; both gaps have TODOS entries with re-entry tests. The dedicated eighth periodic slice and `AUTOPLAN_CHAIN_BUDGET` go with them.
- Paid files that asserted nothing or could not pass: `skill-llm-eval-spec`, `skill-e2e-spec-execute`, `gemini-e2e` (no Gemini CLI in CI), `skill-e2e-ship-idempotency`, `skill-e2e-conductor-prose`, `codex-e2e-plan-format`, `skill-e2e-brain-privacy-gate`, `skill-e2e-opus-47` (its negative routing controls moved into `skill-routing-e2e`) and two duplicate overlay wrappers; `test:gemini` scripts removed.
- Free tests of dead eval code, product tests that exercised copies of the product, and test-infrastructure dead code.
### Changed
- Tests that faked the product now drive it: the design `serve()` server, terminal-agent `/internal/grant` and `/internal/revoke` bearer auth, `/health` liveness, and brain-sync consent before egress.
- Per-incident replay files are folded verbatim into twelve detector owner tests (listed in `docs/TEST_PORTFOLIO.md`), keeping every captured case.
- Paid touchfiles are derived: `test/touchfiles.test.ts` checks that each case's key covers its eval's static helper/fixture imports and the fixture paths it names, and free `*.test.ts` files are no longer touchfiles, so editing a free test no longer selects paid evals.
- The paid planner skips a file for a tier lane when every E2E id it registers belongs to the other tier (the hollow shards), and the weekly gate census skips the LLM judges.
- Seven paid evals that pinned `claude-opus-4-7` or `claude-sonnet-4-6` now capture with the default model from `resolveEvalModel`; all passed on it. Four more (`skill-e2e-design`, `-office-hours-phase4`, `-plan-prosons`, `-plan`) keep `claude-opus-4-7` because six of their cases failed on the default model; TODOS tracks re-pinning them.
- memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral make no model calls and now run in the free suite; CI-unrunnable Codex, Aside, outside-voice and iOS-device files are excluded from the weekly lane with a tracked re-entry condition.
- The plan-count history PTY test waits for its startup marker instead of a fixed 8-second sleep.
### For contributors
- When a paid eval fails, fix the product or harness and add the captured case as one row in the detector's owner test; `test/test-of-test-ratchet.test.ts` fails on any new test file that imports only `test/` code and names the owner test to extend. `CONTRIBUTING.md` "Test tiers" has an example.
- Deleted `test/helpers` modules and where their live cases went:
- `autoplan-setup-question`, `ceo-approach-pick`, `ceo-completion-handoff`, `ceo-payment-findings`, `design-artifact-question`, `design-count-fixture`, `design-count-outside`, `design-count-review`, `devex-count-fixture`, `devex-seed-coverage`, `eng-count-question-policy`: consumed only by the retired finding-count evals; runner tests that used them as caller policies now use inline policies, and the omitted-`multiSelect` default moved to `test/plan-review-decisions.test.ts`.
- `autoplan-phase-order`, `pty-current-screen`: never wired; the settings-overwrite card assertion moved to `test/helpers/claude-pty-runner.unit.test.ts`.
- `ceo-paired-fixture`, `design-ui-scope`, `plan-skill-completion`, `required-reads`, `transcript-section-logger`, `eng-finding-fixture`, `eng-completion-handoff`, `eng-retained-corpus`, `captured-paths`, `gemini-session-runner`: no live cases.
- `test/helpers/resolve-repo-path.ts` resolves specifiers and path literals for both the ratchet and the touchfile closure check. The full evidence (inventories, selection proof, security mapping, retained false positives) is in `docs/test-audit-2026-09.md`.
## [1.91.6.0] - 2026-09-28
PR eval slices are balanced by how long each eval actually takes, so the slowest slice no longer carries most of the run.
+6
View File
@@ -230,6 +230,12 @@ Historical measurements from 2026-09-21:
| Local complete free suite | All 993 files, six workers | 4m 35s |
| Complete Linux CI | All 993 files, 20 isolated runners | 1m 40s across test steps; 3m 7s including setup and aggregation |
After the 2026-09-29 test audit ([evidence](docs/test-audit-2026-09.md)):
| Run | Coverage | Elapsed |
|---|---|---|
| Complete free suite, `bun run test:ubicloud` (standard-16) | All 857 files, 20,302 passing tests | 136 seconds on the VM; 1,738 seconds of recorded serial test time |
The [Linux CI run](https://github.com/garrytan/gstack/actions/runs/35642667809)
on `25030d68` included one recorded successful retry. Its slowest test step was 77 seconds;
staggered starts made the complete test span longer. Typical PR paid-gate timing
+10
View File
@@ -842,6 +842,16 @@ and `test/dx-selected-navigation-ap.test.ts`. One shared table run once against
only after `engFirstReviewAUQ` checks native completion once at entry; today each branch gates it
separately, so the change alters a paid verdict and needs its own paid run.
### P3: Re-pin the four remaining claude-opus-4-7 paid files
**What:** The 2026-09 audit moved seven paid evals to the default capture model (`resolveEvalModel('capture')`).
`skill-e2e-design`, `skill-e2e-office-hours-phase4`, `skill-e2e-plan-prosons` and `skill-e2e-plan` keep
`claude-opus-4-7` because six cases failed on the default model in one run (plan-design-review-plan-mode timeout,
office-hours-phase4-fork format, plan-review-prosons-neutral-neg missing output, plan-ceo-review-selective and
plan-eng-review 600 s timeouts, plan-ceo-review-expansion-energy posture score 3). They measure an old model.
**Re-entry:** fix the prompt, budget or rubric so each case passes on the default model in one run, then drop the pin.
### P3: Retire the unused CEO payment seeder
**What:** `seedCeoPaymentProject` and `pickSuppliedCeoPlanStart` in `test/helpers/ceo-finding-fixture.ts`
+1 -1
View File
@@ -1 +1 @@
1.91.6.0
1.91.8.0
+1 -1
View File
@@ -1,4 +1,4 @@
# gstack digest v1.91.6.0 — regenerate/re-copy after upgrading gstack
# gstack digest v1.91.8.0 — regenerate/re-copy after upgrading gstack
Behavioral rules from gstack (https://github.com/garrytan/gstack), compressed
for agent hosts without a full skill install. The full skills add workflows,
+3 -1
View File
@@ -141,12 +141,14 @@ test/plan-count-design-ui-recovery.test.ts
test/plan-count-native-input.test.ts
test/plan-count-empty-review.test.ts
test/plan-count-owned-permission.test.ts
test/plan-count-quoted-frame-ak.test.ts
test/plan-count-file-permission.test.ts
test/plan-count-truncated-question.test.ts
test/plan-count-preview-footer.test.ts
test/eng-test-plan-edit-approval.test.ts
```
The quoted-frame selector was folded into `test/plan-count-file-permission.test.ts` in the 2026-09 audit.
The publication/watchdog pair is `test/autoplan-publication-guard.test.ts` and
`test/cso-watchdog.test.ts`. The live pair is
`test/skill-e2e-auq-consistency.test.ts` and
+55 -2
View File
@@ -262,7 +262,7 @@ import/literal chain, the key, the verify command and CONTRIBUTING.md#paid-test-
|---|---|---|
| E | Selection proof above: no lost case for the four sample edits under either profile; growth only from real static dependencies | kept |
| B5 | Gate lane 52 → 42 files, weekly gate census 52 → 41 (judges skipped), periodic 77 → 69; PR-profile selection for the sample edits byte-identical before and after | kept |
| B8 | Pre-spend estimate net −$3.14/week (below); paid run result below | see paid validation |
| B8 | Paid run: gate 16/16 pass; periodic 28 pass, 6 fail (all in four files). Fallback taken: those four files keep claude-opus-4-7; seven files re-pinned. Estimated B8 delta after the fallback: +$0.69/week (opus −$1.43, sonnet +$2.12), below zero once C and B5 savings are counted | kept (seven files) |
### B8 pre-spend estimate (recorded 2026-09-29, before any B8 paid run)
@@ -289,7 +289,20 @@ Fallback check: `git log -S claude-sonnet-4-6` on skill-e2e-office-hours and -br
- B2 union judge "browse/SKILL.md reference": PASS (clarity 4, completeness 4, actionability 4), $0.02. Fallback not
taken; the three original browse judges are deleted.
- B6 folded journey negatives in `skill-routing-e2e`: 3/3 unrouted, $0.36. Fallback not taken; `skill-e2e-opus-47` deleted.
- B8 re-pin run and the full gate census: recorded in the release commit.
- B8 re-pin run (commit B8 tree, `EVALS_ALL=1`, `EVALS_TIER=gate` then `periodic`, 11 files, detached, about $32 logged
capture cost): gate 16 pass / 0 fail; periodic 28 pass / 6 fail / 33 skip. Failures, all passing in the 09-14, 09-21
and 09-28 weekly runs on the old pins, so attributed to the default model:
`plan-design-review-plan-mode` (timeout at 300 s, no turns recorded), `office-hours-phase4-fork` (no two-alternative
fork), `plan-review-prosons-neutral-neg` (output file not written), `plan-ceo-review-selective` and `plan-eng-review`
(600 s timeouts), `plan-ceo-review-expansion-energy` (surface-framing score 3 < 4). Fallback taken: skill-e2e-design,
-office-hours-phase4, -plan-prosons and -plan keep claude-opus-4-7 (TODOS entry); the other seven files stay re-pinned.
- PR-profile list on the final diff (`--tier gate --profile pr --list`, no EVALS_ALL): unknown dependencies (deleted
helpers, fixtures and workflow edits) restore every gate case: 86 of 192 tests, 38 of 42 shards. Recorded as data.
- Gate census pre-spend estimate (recorded before running): 41 planned files (judges skipped). The 21 files with
per-file cost in the retained weekly artifacts total about $22; the other 20 have no retained cost, so about $40–45
in all at the same average. Wall clock with 8 local workers: about 1–2 hours. The census is the one full paid run
this PR spends on; the B8 run above already covered the re-pinned files.
- Full gate census: results in the final report and PR body.
## Before metrics (65bfb0c)
@@ -306,6 +319,46 @@ test/fixtures bytes: 16289330 total
all test-file LOC: 279898
- free tests: 27,331 passed, 0 failed (1065 shards)
## After metrics (release commit, same counting script as before)
| Measure | Before (65bfb0c) | After |
|---|---:|---:|
| Tracked `*.test.ts` files | 1,184 | 957 |
| `test/*.test.ts` files | 979 | 755 |
| `test/` TypeScript lines | 274,208 | 227,713 |
| `test/helpers` lines | 51,390 | 39,427 |
| `test/fixtures` bytes | 16,289,330 | 9,695,914 |
| All `*.test.ts` lines | 279,640 | 244,504 |
| Free suite (Ubicloud standard-16, `--record-durations`) | 1,065 files, 27,331 passing, 142 s wall, 1,888.2 s serial | 857 files, 20,302 passing, 136 s wall, 1,737.7 s serial |
| Paid files / gate lane / periodic lane | 119 / 58 / 100 | 100 / 42 / 69 |
| Weekly gate census planned files | 58 | 41 |
| 09-21 weekly periodic shard-minutes on files this branch removes | 235 of 462 | 0 |
`git diff --numstat 65bfb0c..release`: production, CI and scripts 21 files (+90/−215); docs 6 (+574/−78 before the
release docs sweep); tests 383 (+9,375/−44,379); test helpers 47 (+786/−12,749); fixtures 199 (−43,667).
## Kept vs plan
- Kept `AUTOPLAN_PREFLIGHT_BUDGET_BYTES` (G): `skill-preflight-budget.test.ts` enforces it on real generated output.
- Deleted `plan-tune-cathedral-fixture.test.ts` beyond the plan (B3): it only replayed the renamed file's fixture.
- `eng-finding-fixture.test.ts`: the plan named four prompt-builder tests; only two existed, and C deleted them with
the paid file they read.
- C0 agreement rule: harness and budget were treated as one non-product group; every artifact of the five files was
harness or budget, none product.
- C kept seven of the eight production-touching files; `ceo-current-decision-record` went because its template read
only fed the retired counter. Three helpers were restored for kept tests (`autoplan-method-read-audit.ts`,
`autoplan-preconfigured-fixture.ts`, `readPendingAutoplanArtifact`).
- `CARVE_GUARDS.autoplan` became `behavioral: 'none'` (the retired chain was its only section-read proof).
- D folds incident files verbatim into owner `describe` blocks rather than rewriting them into value tables, so no
incident control can be dropped; the native-completion negative table is deferred (TODOS) because collapsing it
changes `engFirstReviewAUQ` gating on a paid verdict.
- E stops the closure walk at global touchfile modules and excludes the selection modules; helpers imported by a
paid file now select every case that file registers (for example the cookie judge helpers select all judges).
- H edited only `plan-count-history`: `eng-semantic-terminal`'s sleeping cases and `design-artifact-question` went in C.
- B5 has no CLI file selector to bypass the skip; running a file directly with `bun test` bypasses it.
- Fixes to earlier commits: the B commit's census, judge-count, touchfile-count and selection literals were stale
(nine free failures found by a full local run) and were fixed inside that commit before C.
## Retained false positives (lane reports §4)
### Lane 1
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "gstack",
"version": "1.91.6",
"version": "1.91.8",
"description": "Garry's Stack — Claude Code skills + fast headless browser. One repo, one install, entire AI engineering workflow.",
"license": "MIT",
"type": "module",
File diff suppressed because it is too large. Load diff