v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998)

This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 14:35:00 -07:00
1 parent 943105f109
commit 96764e80a6
56 files changed
+2445 -216

No files matched your search

+1
View File
@@ -49,6 +49,7 @@ gstack/
├── ship/ # Ship workflow skill
├── review/ # PR review skill (checklist.md is hand-written; design-checklist.md is GENERATED from lib/design-catalog.ts)
├── deslop-shared-libs/ # Recommendations-only audit for worthwhile shared-code extractions
├── test-audit/ # Report-first sweep for low-value tests (test value bar, audit mode)
├── plan-ceo-review/ # /plan-ceo-review skill
├── plan-eng-review/ # /plan-eng-review skill
├── autoplan/ # /autoplan skill (auto-review pipeline: CEO → design → DX → eng, eng always last)
+1 -1
View File
@@ -389,7 +389,7 @@ Planner entries and execution results record the effective wall,
its source and policy identifier. Custom drivers must resolve each job instead
of passing their ordinary 1800-second default as an explicit cap;
their outer controller/detach wall must also cover the allocated work and cleanup.
The current paid census has 104 files: 46 gate-tier and 70 periodic-tier.
The current paid census has 105 files: 47 gate-tier and 71 periodic-tier.
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
wrapper covers a full-gate fallback at its default two workers. The broad gate
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
+23
View File
@@ -39,6 +39,7 @@ Detailed guides for every gstack skill — philosophy, workflow, and examples.
| [`/context-restore`](#context-restore) | **Restore State** | Resume from a saved context, even across Conductor workspace handoffs. |
| [`/health`](#health) | **Code Quality Dashboard** | Wraps type checker, linter, tests, dead code detection. Computes a weighted 0-10 score; tracks trends over time. |
| [`/deslop-shared-libs`](#deslop-shared-libs) | **Shared Code Reviewer** | Find worthwhile shared-code extractions in recent work. Recommendations only. |
| [`/test-audit`](#test-audit) | **Test Auditor** | Sweep existing tests for low-value, implementation-coupled or duplicate tests. Report-only unless you approve a batch. |
| [`/landing-report`](#landing-report) | **Ship Queue Dashboard** | Read-only snapshot of the workspace-aware ship queue. Which version slots are claimed, which sibling workspaces have WIP. |
| [`/benchmark-models`](#benchmark-models) | **Model Benchmark** | Side-by-side cross-model benchmark for skills (Claude vs GPT vs Gemini). Latency, tokens, cost, optional LLM-judged quality. |
| | | |
@@ -782,6 +783,28 @@ checks do not run the history audit. Optional extractions are advisory and requi
approval; they do not block a clean review or reduce its score. Actual defects
keep their normal fix handling.
## `/test-audit`
Find existing tests that cost more than they protect. `/review`, `/ship`, `/qa` and
`/plan-eng-review` apply the same [test value bar](test-value-bar.md) to tests in a
diff; `/test-audit` sweeps the tests that already exist.
```text
You: /test-audit
You: /test-audit test/ --max-candidates 5
You: /test-audit --since origin/main
```
A mechanical pre-filter shortlists assertion-free probes, source greps, export-list
copies and near-duplicate files before any model reading. Each candidate gets a
retirement card (what it detects, non-test callers with the search command, the
stronger remaining proof, history, what retiring it unlocks, and the validation
command). Contract tests such as SKILL.md goldens and prompt-byte checks are
retained. The report and a JSON sidecar land in `~/.gstack/projects/<slug>/`.
Nothing is edited unless you approve a batch; spawned sessions stay report-only.
Tests marked `gstack:test-value keep reason="..."` are skipped and listed in the
report's appendix.
## `/benchmark`
This is my **performance engineer mode**.
+223
View File
@@ -0,0 +1,223 @@
# Test value bar
gstack's test workflows share one rule: a test earns its place by protecting
behavior that a real regression would break. Test count is not a goal. The bar is
embedded in `/plan-eng-review`, `/review`, `/qa`, `/qa-only` and `/ship`, so every
new or changed test in a diff meets it without a separate step. `/test-audit`
applies the same bar to tests that already exist.
The source of truth is `scripts/resolvers/test-value.ts`. It was adapted from
OpenClaw's `test-audit` skill (openclaw/openclaw@a214e76,
`.agents/skills/test-audit/SKILL.md`), generalized to any test runner.
## Glossary
- **Authoring gate**: four questions every new or changed test must answer. What
behavior, invariant or contract does it protect? What credible regression makes it
fail? Why does existing coverage not already catch that? Does it need a production
seam no production caller needs? A missing answer means extend an existing test or
drop the proposal.
- **Value bar**: the authoring gate plus the rule that a test which breaks under a
behavior-preserving refactor asserts implementation, unless its exact output is the
declared contract (goldens, prompt bytes, wire formats).
- **Retention bar**: keep a test that independently enforces a public API, protocol,
config, migration, storage, security, platform, default, prompt-byte,
generated-output (SKILL.md golden), package, release or architecture contract; call
order when order is observable; source inspection when it is the cheapest
independent guard. Static or slow is not a reason to delete. Anything reachable
from the package entrypoint is never retired.
- **Value card**: the gate's four answers as one line,
`Value: protects=<...>; fails_when=<...>; why_new=<...>; seam=none`. Each field is
at most 160 UTF-8 bytes in that line (clamped to 157 bytes plus `...`); JSON keeps
full values and header comments wrap instead of truncating.
- **Retirement card**: the evidence a deletion needs, complete before any edit:
`test`, `detects`, `non_test_callers`, `search_command`, `stronger_proof`,
`history`, `unlocks`, `validation`.
- **Weak path**: a changed path whose only tests are weak: a ★ test (smoke,
existence, trivial assertion), a new test that fails the gate, or a test written in
this `/ship` run that the rating pass has not rated. Reasons:
`star_one | gate_failed | unrated`.
- **X and Y**: X = paths with a ★★ or ★★★ test / total paths (value-weighted; the
`/ship` gate uses X). Y = paths with any test / total paths. Total paths is the
diff's codepath trace, capped at 30; zero paths skips the gate.
- **Base control**: running a new regression test against the base branch in a
temporary worktree to prove the behavior existed and the test is valid.
- **Grep-only**: caller evidence from a text search. It cannot see re-exports,
dynamic dispatch or generated code, so production code is never removed on grep
evidence alone.
## Worked X/Y example
A diff touches 10 paths. Four have ★★ or ★★★ tests, three have only a ★ smoke test,
and three have no test.
- X = 4 / 10 = 40% value-weighted. The gate uses this number.
- Y = 7 / 10 = 70% including weakly covered paths.
- `gaps` = 3 (no test); `weak_gaps` = the 3 ★-only paths with reason `star_one`.
The PR body shows `Coverage: 40% value-weighted (70% including 3 weakly covered paths)`.
## Cards
A good value card:
```text
Value: protects=refundPayment rejects an empty reason; fails_when=the reason guard is removed or inverted; why_new=billing.test.ts covers processPayment only; seam=none
```
A rejected proposal:
```text
Rejected (covered_elsewhere): "checkout renders"; checkout.e2e.ts:15 covers it, so extend that test.
```
Rejection codes: `duplicate_protects`, `needs_seam`, `incomplete_card`,
`no_credible_regression`, `covered_elsewhere`, `implementation_coupled`.
## Header comments written by `/ship`
TypeScript:
```ts
// Generated by /ship coverage audit
// Value: protects=refundPayment rejects an empty reason;
// fails_when=the reason guard is removed or inverted;
// why_new=billing.test.ts covers processPayment only; seam=none
test('refundPayment rejects an empty reason', () => {
expect(() => refundPayment('pay_1', '')).toThrow('Reason required');
});
```
Python:
```python
# Generated by /ship coverage audit
# Value: protects=refund rejects an empty reason;
# fails_when=the reason guard is removed or inverted;
# why_new=test_billing.py covers process_payment only; seam=none
def test_refund_rejects_empty_reason():
with pytest.raises(ValueError, match="Reason required"):
refund_payment("pay_1", "")
```
A file type with no known comment syntax gets its card in the PR body's Test value
details instead.
## Sample PR body block
```markdown
## Test Coverage
<coverage diagram>
Tests: 41 → 44 (+3 new)
Coverage: 58% value-weighted (81% including 4 weakly covered paths)
Test value: 3 tests written, 2 rejected by the authoring gate, 1 existing test extended, 4 paths weakly covered (weak = ★, gate-failing or unrated).
Regression proof — fails at HEAD: yes · passes at base: yes · passes after fix: yes
```
## Sample `/test-audit` report section
```markdown
### Owner: src/billing/refund.ts (1 candidate, production -12 LOC, test -40 LOC)
- test: test/refund-exports.test.ts "exports the refund helpers"
- detects: a renamed export, not a behavior change
- non_test_callers: 0 for `normalizeRefundReason` (grep-only)
- search_command: git grep -n -F -w -e 'normalizeRefundReason' -- . ':!test/' ...
- stronger_proof: test/refund.test.ts covers refund reasons through refundPayment
- history: added in a1b2c3d to keep the helper exported during a refactor
- unlocks: delete the test-only export `normalizeRefundReason`
- validation: bun test test/refund.test.ts; bunx tsc --noEmit
- verdict: retire
Retained: test/fixtures/golden/ship-SKILL.md check (generated-output contract).
Suppressed: test/legacy-api.test.ts (reason="public API snapshot for v1 clients").
```
## Overrides
Projects tune `/ship` in their CLAUDE.md `## Test Coverage` section. Every key is
optional; absent keys use the defaults.
```markdown
## Test Coverage
Minimum: 60%
Target: 80%
Generation cap: 5
Base control: auto
Base control budget: 90
Star rating: auto
```
- `Generation cap:` tests per generation pass (default 5; 2 passes max).
- `Base control:` `auto` runs regression tests at the base branch; `off` skips it.
- `Base control budget:` seconds per base-control run (default 90; 3 minutes total).
- `Star rating:` `off` makes the gate use `coverage_pct` (any test); weak paths are
still listed.
A per-test pragma, in the file's comment syntax, makes `/review` and `/test-audit`
skip a deliberate test and report the reason:
```ts
// gstack:test-value keep reason="pins the v1 wire format for external clients"
```
## Messages
Each degraded-mode message names the problem, its consequence and the fix.
### rating-unavailable
`rating unavailable`: the read-only rating dispatch failed or timed out, so the
coverage gate is skipped for this run. Re-run Step 7 of `/ship` to re-rate the tests.
### value-weighted-coverage-unavailable
The coverage audit returned no usable `coverage_pct_value`, usually because the
installed skill is older than this change. The gate used `coverage_pct` (any test)
for this run. Run `/gstack-upgrade`.
### inconsistent-coverage-inputs
`coverage_pct_value` was above `coverage_pct`, which cannot happen when both are
computed from the same paths. It was clamped to `coverage_pct`. Re-run Step 7 if the
numbers look wrong.
### malformed-key-ignored
A Step 7 JSON key had the wrong type (for example `weak_gaps` not an array), so it
counts as empty. The likely cause is an outdated installed skill; run `/gstack-upgrade`.
### all-generated-tests-rejected
Every test written in a generation pass failed the machine checks (incomplete card,
duplicate `protects`, or a seam with no non-test caller). The rejected files were
removed and the gate proceeds with the unchanged value-weighted coverage. See
`tests_rejected` for each reason.
### base-control-unavailable
The regression test could not run at the base branch: `ecosystem` (not a Node/Bun
project), `no base remote`, `base not fetched`, `worktree add failed`, `budget`, or
`collection error` (the test could not load at base). The fails-at-HEAD result still
stands. To check by hand:
```bash
git worktree add --detach /tmp/base-check origin/<base>
cp <test> /tmp/base-check/<test> # plus any new test-only fixtures
(cd /tmp/base-check && <test command for that file>)
git worktree remove --force /tmp/base-check
```
Then report `passes at base: manual`.
### caller-check-unavailable
The non-test caller search failed, or the symbol is not a plain identifier
(`unsupported symbol`). The finding stays INFORMATIONAL and nothing is proposed for
deletion. Run the search by hand to complete the evidence.
### unknown-test-value-bar-mode
A template used `{{TEST_VALUE_BAR:<mode>}}` with a mode other than
`plan|ship|qa|audit`. Fix the placeholder or add the mode in
`scripts/resolvers/test-value.ts`.