mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998)
This commit is contained in:
1 parent
943105f109
commit
96764e80a6
56 files changed
+2445
-216
No files matched your search
@@ -49,6 +49,7 @@ gstack/
|
||||
├── ship/ # Ship workflow skill
|
||||
├── review/ # PR review skill (checklist.md is hand-written; design-checklist.md is GENERATED from lib/design-catalog.ts)
|
||||
├── deslop-shared-libs/ # Recommendations-only audit for worthwhile shared-code extractions
|
||||
├── test-audit/ # Report-first sweep for low-value tests (test value bar, audit mode)
|
||||
├── plan-ceo-review/ # /plan-ceo-review skill
|
||||
├── plan-eng-review/ # /plan-eng-review skill
|
||||
├── autoplan/ # /autoplan skill (auto-review pipeline: CEO → design → DX → eng, eng always last)
|
||||
|
||||
@@ -389,7 +389,7 @@ Planner entries and execution results record the effective wall,
|
||||
its source and policy identifier. Custom drivers must resolve each job instead
|
||||
of passing their ordinary 1800-second default as an explicit cap;
|
||||
their outer controller/detach wall must also cover the allocated work and cleanup.
|
||||
The current paid census has 104 files: 46 gate-tier and 70 periodic-tier.
|
||||
The current paid census has 105 files: 47 gate-tier and 71 periodic-tier.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
|
||||
wrapper covers a full-gate fallback at its default two workers. The broad gate
|
||||
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
|
||||
|
||||
@@ -39,6 +39,7 @@ Detailed guides for every gstack skill — philosophy, workflow, and examples.
|
||||
| [`/context-restore`](#context-restore) | **Restore State** | Resume from a saved context, even across Conductor workspace handoffs. |
|
||||
| [`/health`](#health) | **Code Quality Dashboard** | Wraps type checker, linter, tests, dead code detection. Computes a weighted 0-10 score; tracks trends over time. |
|
||||
| [`/deslop-shared-libs`](#deslop-shared-libs) | **Shared Code Reviewer** | Find worthwhile shared-code extractions in recent work. Recommendations only. |
|
||||
| [`/test-audit`](#test-audit) | **Test Auditor** | Sweep existing tests for low-value, implementation-coupled or duplicate tests. Report-only unless you approve a batch. |
|
||||
| [`/landing-report`](#landing-report) | **Ship Queue Dashboard** | Read-only snapshot of the workspace-aware ship queue. Which version slots are claimed, which sibling workspaces have WIP. |
|
||||
| [`/benchmark-models`](#benchmark-models) | **Model Benchmark** | Side-by-side cross-model benchmark for skills (Claude vs GPT vs Gemini). Latency, tokens, cost, optional LLM-judged quality. |
|
||||
| | | |
|
||||
@@ -782,6 +783,28 @@ checks do not run the history audit. Optional extractions are advisory and requi
|
||||
approval; they do not block a clean review or reduce its score. Actual defects
|
||||
keep their normal fix handling.
|
||||
|
||||
## `/test-audit`
|
||||
|
||||
Find existing tests that cost more than they protect. `/review`, `/ship`, `/qa` and
|
||||
`/plan-eng-review` apply the same [test value bar](test-value-bar.md) to tests in a
|
||||
diff; `/test-audit` sweeps the tests that already exist.
|
||||
|
||||
```text
|
||||
You: /test-audit
|
||||
You: /test-audit test/ --max-candidates 5
|
||||
You: /test-audit --since origin/main
|
||||
```
|
||||
|
||||
A mechanical pre-filter shortlists assertion-free probes, source greps, export-list
|
||||
copies and near-duplicate files before any model reading. Each candidate gets a
|
||||
retirement card (what it detects, non-test callers with the search command, the
|
||||
stronger remaining proof, history, what retiring it unlocks, and the validation
|
||||
command). Contract tests such as SKILL.md goldens and prompt-byte checks are
|
||||
retained. The report and a JSON sidecar land in `~/.gstack/projects/<slug>/`.
|
||||
Nothing is edited unless you approve a batch; spawned sessions stay report-only.
|
||||
Tests marked `gstack:test-value keep reason="..."` are skipped and listed in the
|
||||
report's appendix.
|
||||
|
||||
## `/benchmark`
|
||||
|
||||
This is my **performance engineer mode**.
|
||||
|
||||
@@ -0,0 +1,223 @@
|
||||
# Test value bar
|
||||
|
||||
gstack's test workflows share one rule: a test earns its place by protecting
|
||||
behavior that a real regression would break. Test count is not a goal. The bar is
|
||||
embedded in `/plan-eng-review`, `/review`, `/qa`, `/qa-only` and `/ship`, so every
|
||||
new or changed test in a diff meets it without a separate step. `/test-audit`
|
||||
applies the same bar to tests that already exist.
|
||||
|
||||
The source of truth is `scripts/resolvers/test-value.ts`. It was adapted from
|
||||
OpenClaw's `test-audit` skill (openclaw/openclaw@a214e76,
|
||||
`.agents/skills/test-audit/SKILL.md`), generalized to any test runner.
|
||||
|
||||
## Glossary
|
||||
|
||||
- **Authoring gate**: four questions every new or changed test must answer. What
|
||||
behavior, invariant or contract does it protect? What credible regression makes it
|
||||
fail? Why does existing coverage not already catch that? Does it need a production
|
||||
seam no production caller needs? A missing answer means extend an existing test or
|
||||
drop the proposal.
|
||||
- **Value bar**: the authoring gate plus the rule that a test which breaks under a
|
||||
behavior-preserving refactor asserts implementation, unless its exact output is the
|
||||
declared contract (goldens, prompt bytes, wire formats).
|
||||
- **Retention bar**: keep a test that independently enforces a public API, protocol,
|
||||
config, migration, storage, security, platform, default, prompt-byte,
|
||||
generated-output (SKILL.md golden), package, release or architecture contract; call
|
||||
order when order is observable; source inspection when it is the cheapest
|
||||
independent guard. Static or slow is not a reason to delete. Anything reachable
|
||||
from the package entrypoint is never retired.
|
||||
- **Value card**: the gate's four answers as one line,
|
||||
`Value: protects=<...>; fails_when=<...>; why_new=<...>; seam=none`. Each field is
|
||||
at most 160 UTF-8 bytes in that line (clamped to 157 bytes plus `...`); JSON keeps
|
||||
full values and header comments wrap instead of truncating.
|
||||
- **Retirement card**: the evidence a deletion needs, complete before any edit:
|
||||
`test`, `detects`, `non_test_callers`, `search_command`, `stronger_proof`,
|
||||
`history`, `unlocks`, `validation`.
|
||||
- **Weak path**: a changed path whose only tests are weak: a ★ test (smoke,
|
||||
existence, trivial assertion), a new test that fails the gate, or a test written in
|
||||
this `/ship` run that the rating pass has not rated. Reasons:
|
||||
`star_one | gate_failed | unrated`.
|
||||
- **X and Y**: X = paths with a ★★ or ★★★ test / total paths (value-weighted; the
|
||||
`/ship` gate uses X). Y = paths with any test / total paths. Total paths is the
|
||||
diff's codepath trace, capped at 30; zero paths skips the gate.
|
||||
- **Base control**: running a new regression test against the base branch in a
|
||||
temporary worktree to prove the behavior existed and the test is valid.
|
||||
- **Grep-only**: caller evidence from a text search. It cannot see re-exports,
|
||||
dynamic dispatch or generated code, so production code is never removed on grep
|
||||
evidence alone.
|
||||
|
||||
## Worked X/Y example
|
||||
|
||||
A diff touches 10 paths. Four have ★★ or ★★★ tests, three have only a ★ smoke test,
|
||||
and three have no test.
|
||||
|
||||
- X = 4 / 10 = 40% value-weighted. The gate uses this number.
|
||||
- Y = 7 / 10 = 70% including weakly covered paths.
|
||||
- `gaps` = 3 (no test); `weak_gaps` = the 3 ★-only paths with reason `star_one`.
|
||||
|
||||
The PR body shows `Coverage: 40% value-weighted (70% including 3 weakly covered paths)`.
|
||||
|
||||
## Cards
|
||||
|
||||
A good value card:
|
||||
|
||||
```text
|
||||
Value: protects=refundPayment rejects an empty reason; fails_when=the reason guard is removed or inverted; why_new=billing.test.ts covers processPayment only; seam=none
|
||||
```
|
||||
|
||||
A rejected proposal:
|
||||
|
||||
```text
|
||||
Rejected (covered_elsewhere): "checkout renders"; checkout.e2e.ts:15 covers it, so extend that test.
|
||||
```
|
||||
|
||||
Rejection codes: `duplicate_protects`, `needs_seam`, `incomplete_card`,
|
||||
`no_credible_regression`, `covered_elsewhere`, `implementation_coupled`.
|
||||
|
||||
## Header comments written by `/ship`
|
||||
|
||||
TypeScript:
|
||||
|
||||
```ts
|
||||
// Generated by /ship coverage audit
|
||||
// Value: protects=refundPayment rejects an empty reason;
|
||||
// fails_when=the reason guard is removed or inverted;
|
||||
// why_new=billing.test.ts covers processPayment only; seam=none
|
||||
test('refundPayment rejects an empty reason', () => {
|
||||
expect(() => refundPayment('pay_1', '')).toThrow('Reason required');
|
||||
});
|
||||
```
|
||||
|
||||
Python:
|
||||
|
||||
```python
|
||||
# Generated by /ship coverage audit
|
||||
# Value: protects=refund rejects an empty reason;
|
||||
# fails_when=the reason guard is removed or inverted;
|
||||
# why_new=test_billing.py covers process_payment only; seam=none
|
||||
def test_refund_rejects_empty_reason():
|
||||
with pytest.raises(ValueError, match="Reason required"):
|
||||
refund_payment("pay_1", "")
|
||||
```
|
||||
|
||||
A file type with no known comment syntax gets its card in the PR body's Test value
|
||||
details instead.
|
||||
|
||||
## Sample PR body block
|
||||
|
||||
```markdown
|
||||
## Test Coverage
|
||||
<coverage diagram>
|
||||
Tests: 41 → 44 (+3 new)
|
||||
Coverage: 58% value-weighted (81% including 4 weakly covered paths)
|
||||
Test value: 3 tests written, 2 rejected by the authoring gate, 1 existing test extended, 4 paths weakly covered (weak = ★, gate-failing or unrated).
|
||||
Regression proof — fails at HEAD: yes · passes at base: yes · passes after fix: yes
|
||||
```
|
||||
|
||||
## Sample `/test-audit` report section
|
||||
|
||||
```markdown
|
||||
### Owner: src/billing/refund.ts (1 candidate, production -12 LOC, test -40 LOC)
|
||||
|
||||
- test: test/refund-exports.test.ts "exports the refund helpers"
|
||||
- detects: a renamed export, not a behavior change
|
||||
- non_test_callers: 0 for `normalizeRefundReason` (grep-only)
|
||||
- search_command: git grep -n -F -w -e 'normalizeRefundReason' -- . ':!test/' ...
|
||||
- stronger_proof: test/refund.test.ts covers refund reasons through refundPayment
|
||||
- history: added in a1b2c3d to keep the helper exported during a refactor
|
||||
- unlocks: delete the test-only export `normalizeRefundReason`
|
||||
- validation: bun test test/refund.test.ts; bunx tsc --noEmit
|
||||
- verdict: retire
|
||||
|
||||
Retained: test/fixtures/golden/ship-SKILL.md check (generated-output contract).
|
||||
Suppressed: test/legacy-api.test.ts (reason="public API snapshot for v1 clients").
|
||||
```
|
||||
|
||||
## Overrides
|
||||
|
||||
Projects tune `/ship` in their CLAUDE.md `## Test Coverage` section. Every key is
|
||||
optional; absent keys use the defaults.
|
||||
|
||||
```markdown
|
||||
## Test Coverage
|
||||
Minimum: 60%
|
||||
Target: 80%
|
||||
Generation cap: 5
|
||||
Base control: auto
|
||||
Base control budget: 90
|
||||
Star rating: auto
|
||||
```
|
||||
|
||||
- `Generation cap:` tests per generation pass (default 5; 2 passes max).
|
||||
- `Base control:` `auto` runs regression tests at the base branch; `off` skips it.
|
||||
- `Base control budget:` seconds per base-control run (default 90; 3 minutes total).
|
||||
- `Star rating:` `off` makes the gate use `coverage_pct` (any test); weak paths are
|
||||
still listed.
|
||||
|
||||
A per-test pragma, in the file's comment syntax, makes `/review` and `/test-audit`
|
||||
skip a deliberate test and report the reason:
|
||||
|
||||
```ts
|
||||
// gstack:test-value keep reason="pins the v1 wire format for external clients"
|
||||
```
|
||||
|
||||
## Messages
|
||||
|
||||
Each degraded-mode message names the problem, its consequence and the fix.
|
||||
|
||||
### rating-unavailable
|
||||
|
||||
`rating unavailable`: the read-only rating dispatch failed or timed out, so the
|
||||
coverage gate is skipped for this run. Re-run Step 7 of `/ship` to re-rate the tests.
|
||||
|
||||
### value-weighted-coverage-unavailable
|
||||
|
||||
The coverage audit returned no usable `coverage_pct_value`, usually because the
|
||||
installed skill is older than this change. The gate used `coverage_pct` (any test)
|
||||
for this run. Run `/gstack-upgrade`.
|
||||
|
||||
### inconsistent-coverage-inputs
|
||||
|
||||
`coverage_pct_value` was above `coverage_pct`, which cannot happen when both are
|
||||
computed from the same paths. It was clamped to `coverage_pct`. Re-run Step 7 if the
|
||||
numbers look wrong.
|
||||
|
||||
### malformed-key-ignored
|
||||
|
||||
A Step 7 JSON key had the wrong type (for example `weak_gaps` not an array), so it
|
||||
counts as empty. The likely cause is an outdated installed skill; run `/gstack-upgrade`.
|
||||
|
||||
### all-generated-tests-rejected
|
||||
|
||||
Every test written in a generation pass failed the machine checks (incomplete card,
|
||||
duplicate `protects`, or a seam with no non-test caller). The rejected files were
|
||||
removed and the gate proceeds with the unchanged value-weighted coverage. See
|
||||
`tests_rejected` for each reason.
|
||||
|
||||
### base-control-unavailable
|
||||
|
||||
The regression test could not run at the base branch: `ecosystem` (not a Node/Bun
|
||||
project), `no base remote`, `base not fetched`, `worktree add failed`, `budget`, or
|
||||
`collection error` (the test could not load at base). The fails-at-HEAD result still
|
||||
stands. To check by hand:
|
||||
|
||||
```bash
|
||||
git worktree add --detach /tmp/base-check origin/<base>
|
||||
cp <test> /tmp/base-check/<test> # plus any new test-only fixtures
|
||||
(cd /tmp/base-check && <test command for that file>)
|
||||
git worktree remove --force /tmp/base-check
|
||||
```
|
||||
|
||||
Then report `passes at base: manual`.
|
||||
|
||||
### caller-check-unavailable
|
||||
|
||||
The non-test caller search failed, or the symbol is not a plain identifier
|
||||
(`unsupported symbol`). The finding stays INFORMATIONAL and nothing is proposed for
|
||||
deletion. Run the search by hand to complete the evidence.
|
||||
|
||||
### unknown-test-value-bar-mode
|
||||
|
||||
A template used `{{TEST_VALUE_BAR:<mode>}}` with a mode other than
|
||||
`plan|ship|qa|audit`. Fix the placeholder or add the mode in
|
||||
`scripts/resolvers/test-value.ts`.
|
||||
Reference in new issue
Block a user