v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998)

This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 14:35:00 -07:00
1 parent 943105f109
commit 96764e80a6
56 files changed
+2445 -216

No files matched your search

+65
View File
@@ -80,3 +80,68 @@ red tests. Unit tests suit logic; real integration/E2E tests protect boundaries
- New public methods/functions with zero test coverage
- Changed methods where existing tests only cover the old behavior, not the new branch
- Utility functions called from multiple places but tested only indirectly
- A path whose only tests are weak (★ smoke/existence/trivial, or a new test failing the
authoring gate below) stays a coverage gap at its existing severity; a low-value test
never closes a gap
### Low-value or implementation-coupled tests
Scope: tests and test-only production seams added or changed in the diff. Findings are
INFORMATIONAL, never CRITICAL, and never an auto-delete: recommend rewrite at the owning
boundary, extend an existing test, or retire with a complete retirement card
(`test, detects, non_test_callers, search_command, stronger_proof, history, unlocks,
validation`). End each finding's `fix` with `Repo-wide sweep: run /test-audit.`
A new or changed test passes the authoring gate only when all four answers exist (read
its `Value: protects=...; fails_when=...; why_new=...; seam=...` header comment when the
diff has one):
1. What observable behavior, invariant or independent contract does it protect?
2. What credible regression makes it fail?
3. Why does existing coverage not already catch that? Prefer adding a row to an existing table-driven test or shared fixture over a near-duplicate.
4. Does it need a production seam (export, flag, wrapper, injection hook) that no production caller needs? If yes, test at the real boundary instead.
Patterns:
- assertion-free coverage probes
- self-comparisons and identity copies
- copied fixtures, inventories or export lists
- exact source, import or string greps that are not a declared contract
- private predicate or call-shape tests duplicated at a real boundary
- duplicate invocations of the same contract
- per-caller replays of a shared helper's tests
- tests whose only purpose is keeping a test-only export, global or wrapper alive
- production code whose only callers are tests
Retention bar (never flag): a test that independently enforces a public API, protocol,
config, migration, storage, security, platform, default, prompt-byte, generated-output
(SKILL.md golden), package, release or architecture contract; call order when order is
observable; source inspection when it is the cheapest independent guard; anything
reachable from the package entrypoint (`package.json` exports/main, index re-exports).
Static or slow is not a reason to delete.
Evidence for each finding goes in `evidence` with the fields `detects` (what failure the
test can actually detect), `non_test_callers`, `search_command` and `stronger_proof`.
For the last two patterns, run the caller check for each symbol the diff adds or exports,
only when the symbol matches `^[A-Za-z_][A-Za-z0-9_]*$` (otherwise record "caller check
unavailable: unsupported symbol"), with the symbol quoted, never interpolated unquoted:
```bash
git grep -n -F -w -e '<symbol>' -- . ':!test/' ':!tests/' ':!spec/' ':!**/__tests__/**' ':!**/*.test.*' ':!**/*.spec.*' ':!**/*_test.*' ':!**/test_*.py'
```
Record the command, exclusions and hit count, and mark the evidence "grep-only" (it cannot
see re-exports, dynamic dispatch or generated code). If the search fails, record: caller
check unavailable: <command>. The finding stays INFORMATIONAL and nothing is proposed for
deletion; run the search by hand to complete the evidence. (see
~/.claude/skills/gstack/docs/test-value-bar.md#caller-check-unavailable)
Skip a test carrying `gstack:test-value keep reason="<why>"` (any comment syntax). Report
the count and reasons of skipped tests as one INFORMATIONAL line.
Rejection vocabulary, when a finding names why a new test fails the gate:
`duplicate_protects`, `needs_seam`, `incomplete_card`, `no_credible_regression`,
`covered_elsewhere`, `implementation_coupled`.
### Regression test without red proof
- The diff adds a test named or commented as a regression, and neither the commit history
nor the PR body shows it failing before the fix. A regression test that never
demonstrably failed proves the mock, not the fix. INFORMATIONAL; ask for the
fails-at-HEAD / passes-after-fix record.