9.0 KiB
Test value bar
gstack's test workflows share one rule: a test earns its place by protecting
behavior that a real regression would break. Test count is not a goal. The bar is
embedded in /plan-eng-review, /review, /qa, /qa-only and /ship, so every
new or changed test in a diff meets it without a separate step. /test-audit
applies the same bar to tests that already exist.
The source of truth is scripts/resolvers/test-value.ts. It was adapted from
OpenClaw's test-audit skill (openclaw/openclaw@a214e76,
.agents/skills/test-audit/SKILL.md), generalized to any test runner.
Glossary
- Authoring gate: four questions every new or changed test must answer. What behavior, invariant or contract does it protect? What credible regression makes it fail? Why does existing coverage not already catch that? Does it need a production seam no production caller needs? A missing answer means extend an existing test or drop the proposal.
- Value bar: the authoring gate plus the rule that a test which breaks under a behavior-preserving refactor asserts implementation, unless its exact output is the declared contract (goldens, prompt bytes, wire formats).
- Retention bar: keep a test that independently enforces a public API, protocol, config, migration, storage, security, platform, default, prompt-byte, generated-output (SKILL.md golden), package, release or architecture contract; call order when order is observable; source inspection when it is the cheapest independent guard. Static or slow is not a reason to delete. Anything reachable from the package entrypoint is never retired.
- Value card: the gate's four answers as one line,
Value: protects=<...>; fails_when=<...>; why_new=<...>; seam=none. Each field is at most 160 UTF-8 bytes in that line (clamped to 157 bytes plus...); JSON keeps full values and header comments wrap instead of truncating. - Retirement card: the evidence a deletion needs, complete before any edit:
test,detects,non_test_callers,search_command,stronger_proof,history,unlocks,validation. - Weak path: a changed path whose only tests are weak: a ★ test (smoke,
existence, trivial assertion), a new test that fails the gate, or a test written in
this
/shiprun that the rating pass has not rated. Reasons:star_one | gate_failed | unrated. - X and Y: X = paths with a ★★ or ★★★ test / total paths (value-weighted; the
/shipgate uses X). Y = paths with any test / total paths. Total paths is the diff's codepath trace, capped at 30; zero paths skips the gate. - Base control: running a new regression test against the base branch in a temporary worktree to prove the behavior existed and the test is valid.
- Grep-only: caller evidence from a text search. It cannot see re-exports, dynamic dispatch or generated code, so production code is never removed on grep evidence alone.
Worked X/Y example
A diff touches 10 paths. Four have ★★ or ★★★ tests, three have only a ★ smoke test, and three have no test.
- X = 4 / 10 = 40% value-weighted. The gate uses this number.
- Y = 7 / 10 = 70% including weakly covered paths.
gaps= 3 (no test);weak_gaps= the 3 ★-only paths with reasonstar_one.
The PR body shows Coverage: 40% value-weighted (70% including 3 weakly covered paths).
Cards
A good value card:
Value: protects=refundPayment rejects an empty reason; fails_when=the reason guard is removed or inverted; why_new=billing.test.ts covers processPayment only; seam=none
A rejected proposal:
Rejected (covered_elsewhere): "checkout renders"; checkout.e2e.ts:15 covers it, so extend that test.
Rejection codes: duplicate_protects, needs_seam, incomplete_card,
no_credible_regression, covered_elsewhere, implementation_coupled.
Header comments written by /ship
TypeScript:
// Generated by /ship coverage audit
// Value: protects=refundPayment rejects an empty reason;
// fails_when=the reason guard is removed or inverted;
// why_new=billing.test.ts covers processPayment only; seam=none
test('refundPayment rejects an empty reason', () => {
expect(() => refundPayment('pay_1', '')).toThrow('Reason required');
});
Python:
# Generated by /ship coverage audit
# Value: protects=refund rejects an empty reason;
# fails_when=the reason guard is removed or inverted;
# why_new=test_billing.py covers process_payment only; seam=none
def test_refund_rejects_empty_reason():
with pytest.raises(ValueError, match="Reason required"):
refund_payment("pay_1", "")
A file type with no known comment syntax gets its card in the PR body's Test value details instead.
Sample PR body block
## Test Coverage
<coverage diagram>
Tests: 41 → 44 (+3 new)
Coverage: 58% value-weighted (81% including 4 weakly covered paths)
Test value: 3 tests written, 2 rejected by the authoring gate, 1 existing test extended, 4 paths weakly covered (weak = ★, gate-failing or unrated).
Regression proof — fails at HEAD: yes · passes at base: yes · passes after fix: yes
Sample /test-audit report section
### Owner: src/billing/refund.ts (1 candidate, production -12 LOC, test -40 LOC)
- test: test/refund-exports.test.ts "exports the refund helpers"
- detects: a renamed export, not a behavior change
- non_test_callers: 0 for `normalizeRefundReason` (grep-only)
- search_command: git grep -n -F -w -e 'normalizeRefundReason' -- . ':!test/' ...
- stronger_proof: test/refund.test.ts covers refund reasons through refundPayment
- history: added in a1b2c3d to keep the helper exported during a refactor
- unlocks: delete the test-only export `normalizeRefundReason`
- validation: bun test test/refund.test.ts; bunx tsc --noEmit
- verdict: retire
Retained: test/fixtures/golden/ship-SKILL.md check (generated-output contract).
Suppressed: test/legacy-api.test.ts (reason="public API snapshot for v1 clients").
Overrides
Projects tune /ship in their CLAUDE.md ## Test Coverage section. Every key is
optional; absent keys use the defaults.
## Test Coverage
Minimum: 60%
Target: 80%
Generation cap: 5
Base control: auto
Base control budget: 90
Star rating: auto
Generation cap:tests per generation pass (default 5; 2 passes max).Base control:autoruns regression tests at the base branch;offskips it.Base control budget:seconds per base-control run (default 90; 3 minutes total).Star rating:offmakes the gate usecoverage_pct(any test); weak paths are still listed.
A per-test pragma, in the file's comment syntax, makes /review and /test-audit
skip a deliberate test and report the reason:
// gstack:test-value keep reason="pins the v1 wire format for external clients"
Messages
Each degraded-mode message names the problem, its consequence and the fix.
rating-unavailable
rating unavailable: the read-only rating dispatch failed or timed out, so the
coverage gate is skipped for this run. Re-run Step 7 of /ship to re-rate the tests.
value-weighted-coverage-unavailable
The coverage audit returned no usable coverage_pct_value, usually because the
installed skill is older than this change. The gate used coverage_pct (any test)
for this run. Run /gstack-upgrade.
inconsistent-coverage-inputs
coverage_pct_value was above coverage_pct, which cannot happen when both are
computed from the same paths. It was clamped to coverage_pct. Re-run Step 7 if the
numbers look wrong.
malformed-key-ignored
A Step 7 JSON key had the wrong type (for example weak_gaps not an array), so it
counts as empty. The likely cause is an outdated installed skill; run /gstack-upgrade.
all-generated-tests-rejected
Every test written in a generation pass failed the machine checks (incomplete card,
duplicate protects, or a seam with no non-test caller). The rejected files were
removed and the gate proceeds with the unchanged value-weighted coverage. See
tests_rejected for each reason.
base-control-unavailable
The regression test could not run at the base branch: ecosystem (not a Node/Bun
project), no base remote, base not fetched, worktree add failed, budget, or
collection error (the test could not load at base). The fails-at-HEAD result still
stands. To check by hand:
git worktree add --detach /tmp/base-check origin/<base>
cp <test> /tmp/base-check/<test> # plus any new test-only fixtures
(cd /tmp/base-check && <test command for that file>)
git worktree remove --force /tmp/base-check
Then report passes at base: manual.
caller-check-unavailable
The non-test caller search failed, or the symbol is not a plain identifier
(unsupported symbol). The finding stays INFORMATIONAL and nothing is proposed for
deletion. Run the search by hand to complete the evidence.
unknown-test-value-bar-mode
A template used {{TEST_VALUE_BAR:<mode>}} with a mode other than
plan|ship|qa|audit. Fix the placeholder or add the mode in
scripts/resolvers/test-value.ts.