mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
310 lines
17 KiB
Markdown
310 lines
17 KiB
Markdown
<!-- AUTO-GENERATED from test-coverage.md.tmpl — do not edit directly -->
|
|
<!-- Regenerate: bun run gen:skill-docs -->
|
|
## Step 7: Test Coverage Audit
|
|
|
|
### Shared subagent dispatch
|
|
|
|
For Steps 7, 8 and 10, use the Agent tool with `run_in_background: false`.
|
|
Omitting the flag runs the subagent in the background. The explicit flag waits
|
|
for a result while keeping a fresh context. Do not invoke the target as a Skill
|
|
or run it inline instead. Inline work is allowed only under that section's
|
|
documented fallback, after a failed subagent has stopped.
|
|
|
|
Dispatch the audit through Agent with `subagent_type: "general-purpose"` and
|
|
`run_in_background: false`, using the shared foreground-dispatch rule above.
|
|
Wait for its LAST-line JSON before applying the coverage gate.
|
|
|
|
**Generation allowance:** Maximum 2 generation passes total per invocation.
|
|
Count each generation-authorized attempt before dispatch/inline execution, including
|
|
the initial audit, failures and zero-test results. Re-entry never resets it.
|
|
Two passes already used means no further generation; read-only reassessment uses no pass.
|
|
|
|
**Subagent prompt:** Supply `<base>`, Step 4's framework/bootstrap decision,
|
|
permitted paths/commands, remaining gaps, passes used and generation allowance.
|
|
No allowance means audit only; missing permission is not approval. Preserve the
|
|
30-path/20-test/2-minute per-test caps.
|
|
|
|
````text
|
|
You are running a ship-workflow test coverage audit. Run `git diff origin/<base>` to include uncommitted tracked changes; also read relevant non-ignored untracked source/tests. Do not commit or push. Perform only this audit; return unresolved user decisions to the parent instead of asking or advancing to another workflow step.
|
|
|
|
Generation: <allowed|audit-only>; passes used: <N> of 2. Audit-only overrides every generation instruction below.
|
|
|
|
100% coverage is the goal — every untested path is a path where bugs hide and vibe coding becomes yolo coding. Evaluate what was ACTUALLY coded (from the diff), not what was planned.
|
|
|
|
### Test Framework Detection
|
|
|
|
Before analyzing coverage, detect the project's test framework:
|
|
|
|
1. **Read CLAUDE.md** — look for a `## Testing` section with test command and framework name. If found, use that as the authoritative source.
|
|
2. **If CLAUDE.md has no testing section, auto-detect:**
|
|
|
|
```bash
|
|
setopt +o nomatch 2>/dev/null || true # zsh compat
|
|
# Detect project runtime (markers are evidence, not commands to run blind)
|
|
[ -f manage.py ] && echo "RUNTIME:python FRAMEWORK:django"
|
|
{ [ -f pyproject.toml ] || [ -f pytest.ini ] || [ -f tox.ini ] || [ -f setup.cfg ] || [ -f requirements.txt ]; } && echo "RUNTIME:python"
|
|
{ [ -f Gemfile ] || [ -f Rakefile ] || [ -f .rspec ]; } && echo "RUNTIME:ruby"
|
|
[ -f package.json ] && echo "RUNTIME:node"
|
|
[ -f go.mod ] && echo "RUNTIME:go"
|
|
[ -f Cargo.toml ] && echo "RUNTIME:rust"
|
|
[ -f pom.xml ] && echo "RUNTIME:jvm BUILD:maven"
|
|
{ [ -f build.gradle ] || [ -f build.gradle.kts ]; } && echo "RUNTIME:jvm BUILD:gradle"
|
|
# Check for existing test infrastructure — config files, scripts, AND test files
|
|
ls jest.config.* vitest.config.* playwright.config.* cypress.config.* .rspec pytest.ini tox.ini phpunit.xml 2>/dev/null
|
|
[ -f package.json ] && grep -q '"test"[[:space:]]*:' package.json && echo "SCRIPT:package.json test"
|
|
[ -f Makefile ] && grep -qE '^(test|check):' Makefile && echo "TARGET:make test"
|
|
git ls-files | grep -cE '(^|/)(tests?|spec|__tests__)/|(^|/)tests?\.py$|(^|/)test_[^/]+\.py$|_test\.(go|py|rb|ts|js|exs)$|\.(test|spec)\.[jt]sx?$|_spec\.rb$|Test\.(java|kt)$' | sed 's/^/TESTFILES:/'
|
|
```
|
|
|
|
3. **If no framework detected:** use the bootstrap decision already made in Step 4; report diagram-only coverage if setup was declined. Do not restart bootstrap from this audit.
|
|
|
|
**0. Before/after test count:**
|
|
|
|
```bash
|
|
# Count test files before any generation
|
|
git ls-files 2>/dev/null | grep -E '(\.test\.|\.spec\.|_test\.|_spec\.)' | wc -l
|
|
```
|
|
|
|
Store this number for the PR body.
|
|
|
|
**1. Trace every codepath changed** using `git diff origin/<base>`:
|
|
|
|
Read every changed file. For each one, trace how data flows through the code — don't just list functions, actually follow the execution:
|
|
|
|
1. **Read the diff.** For each changed file, read the full file (not just the diff hunk) to understand context.
|
|
Definition: a **targeted audit** reviews named concrete source/test files or a
|
|
branch diff. A **prototype** is existing runnable code referenced by the plan,
|
|
not a proposed future component.
|
|
|
|
When grounded in concrete source and test files, read them in a dedicated tool
|
|
call before drawing the diagram. Finish this source read before tracing data
|
|
flow in audit item 2 below; map user flows afterward. Do not mix diff, grep,
|
|
package/config, git, or commentary into that read; use separate calls for
|
|
context. Base the diagram on that read.
|
|
2. **Trace data flow.** Starting from each entry point (route handler, exported function, event listener, component render), follow the data through every branch:
|
|
- Where does input come from? (request params, props, database, API call)
|
|
- What transforms it? (validation, mapping, computation)
|
|
- Where does it go? (database write, API response, rendered output, side effect)
|
|
- What can go wrong at each step? (null/undefined, invalid input, network failure, empty collection)
|
|
3. **Diagram the execution.** For each changed file, draw an ASCII diagram showing:
|
|
- Every function/method that was added or modified
|
|
- Every conditional branch (if/else, switch, ternary, guard clause, early return)
|
|
- Every error path (try/catch, rescue, error boundary, fallback)
|
|
- Every call to another function (trace into it — does IT have untested branches?)
|
|
- Every edge: what happens with null input? Empty array? Invalid type?
|
|
|
|
This is the critical step — you're building a map of every line of code that can execute differently based on input. Every branch in this diagram needs a test.
|
|
|
|
**2. Map user flows, interactions, and error states:**
|
|
|
|
Code coverage isn't enough — you need to cover how real users interact with the changed code. For each changed feature, think through:
|
|
|
|
- **User flows:** What sequence of actions does a user take that touches this code? Map the full journey (e.g., "user clicks 'Pay' → form validates → API call → success/failure screen"). Each step in the journey needs a test.
|
|
- **Interaction edge cases:** What happens when the user does something unexpected?
|
|
- Double-click/rapid resubmit
|
|
- Navigate away mid-operation (back button, close tab, click another link)
|
|
- Submit with stale data (page sat open for 30 minutes, session expired)
|
|
- Slow connection (API takes 10 seconds — what does the user see?)
|
|
- Concurrent actions (two tabs, same form)
|
|
- **Error states the user can see:** For every error the code handles, what does the user actually experience?
|
|
- Is there a clear error message or a silent failure?
|
|
- Can the user recover (retry, go back, fix input) or are they stuck?
|
|
- What happens with no network? With a 500 from the API? With invalid data from the server?
|
|
- **Empty/zero/boundary states:** What does the UI show with zero results? With 10,000 results? With a single character input? With maximum-length input?
|
|
|
|
Add these to your diagram alongside the code branches. A user flow with no test is just as much a gap as an untested if/else.
|
|
|
|
**3. Check each branch against existing tests:**
|
|
|
|
Go through your diagram branch by branch — both code paths AND user flows. For each one, search for a test that exercises it:
|
|
- Function `processPayment()` → look for `billing.test.ts`, `billing.spec.ts`, `test/billing_test.rb`
|
|
- An if/else → look for tests covering BOTH the true AND false path
|
|
- An error handler → look for a test that triggers that specific error condition
|
|
- A call to `helperFn()` that has its own branches → those branches need tests too
|
|
- A user flow → look for an integration or E2E test that walks through the journey
|
|
- An interaction edge case → look for a test that simulates the unexpected action
|
|
|
|
Quality scoring rubric:
|
|
- ★★★ Tests behavior with edge cases AND error paths
|
|
- ★★ Tests correct behavior, happy path only
|
|
- ★ Smoke test / existence check / trivial assertion (e.g., "it renders", "it doesn't throw")
|
|
|
|
### E2E Test Decision Matrix
|
|
|
|
When checking each branch, also determine whether a unit test or E2E/integration test is the right tool:
|
|
|
|
**RECOMMEND E2E (mark as [→E2E] in the diagram):**
|
|
- Common user flow spanning 3+ components/services (e.g., signup → verify email → first login)
|
|
- Integration point where mocking hides real failures (e.g., API → queue → worker → DB)
|
|
- Auth/payment/data-destruction flows — too important to trust unit tests alone
|
|
|
|
**RECOMMEND EVAL (mark as [→EVAL] in the diagram):**
|
|
- Critical LLM call that needs a quality eval (e.g., prompt change → test output still meets quality bar)
|
|
- Changes to prompt templates, system instructions, or tool definitions
|
|
|
|
**STICK WITH UNIT TESTS:**
|
|
- Pure function with clear inputs/outputs
|
|
- Internal helper with no side effects
|
|
- Edge case of a single function (null input, empty array)
|
|
- Obscure/rare flow that isn't customer-facing
|
|
|
|
### REGRESSION RULE (mandatory)
|
|
|
|
**IRON RULE:** When the coverage audit identifies a REGRESSION — code that previously worked but the diff broke — a regression test is written immediately. No AskUserQuestion. No skipping. Regressions are the highest-priority test because they prove something broke.
|
|
|
|
A regression is when:
|
|
- The diff modifies existing behavior (not new code)
|
|
- The existing test suite (if any) doesn't cover the changed path
|
|
- The change introduces a new failure mode for existing callers
|
|
|
|
When uncertain whether a change is a regression, err on the side of writing the test.
|
|
|
|
**4. Output ASCII coverage diagram:**
|
|
|
|
For targeted audits, start Test review output with the coverage diagram. In full
|
|
plan reviews, put it inside the normal Test review section. Required outputs
|
|
keep the final terminal report order.
|
|
|
|
Include BOTH code paths and user flows in the same diagram. Mark E2E-worthy and eval-worthy paths:
|
|
|
|
```
|
|
CODE PATHS USER FLOWS
|
|
[+] src/services/billing.ts [+] Payment checkout
|
|
├── processPayment() ├── [★★★ TESTED] Complete purchase — checkout.e2e.ts:15
|
|
│ ├── [★★★ TESTED] happy + declined + timeout ├── [GAP] [→E2E] Double-click submit
|
|
│ ├── [GAP] Network timeout └── [GAP] Navigate away mid-payment
|
|
│ └── [GAP] Invalid currency
|
|
└── refundPayment() [+] Error states
|
|
├── [★★ TESTED] Full refund — :89 ├── [★★ TESTED] Card declined message
|
|
└── [★ TESTED] Partial (non-throw only) — :101 └── [GAP] Network timeout UX
|
|
|
|
LLM integration: [GAP] [→EVAL] Prompt template change — needs eval test
|
|
|
|
COVERAGE: 5/13 paths tested (38%) | Code paths: 3/5 (60%) | User flows: 2/8 (25%)
|
|
QUALITY: ★★★:2 ★★:2 ★:1 | GAPS: 8 (2 E2E, 1 eval)
|
|
```
|
|
|
|
Legend: ★★★ behavior + edge + error | ★★ happy path | ★ smoke check
|
|
[→E2E] = needs integration test | [→EVAL] = needs LLM eval
|
|
|
|
Avoid bare `[ ]` or `[x]` in diagrams unless the block includes
|
|
`Legend: [x] tested | [ ] no test`. Prefer `[GAP]`, `[★★ TESTED]`,
|
|
`[→E2E]`, `[→EVAL]`; keep user-flow markers off code-path rows.
|
|
|
|
**Fast path:** All paths covered → "Step 7: All new code paths have test coverage ✓" Continue.
|
|
|
|
**5. Generate tests for uncovered paths:**
|
|
|
|
If test framework detected (or bootstrapped in Step 4):
|
|
- Prioritize error handlers and edge cases first (happy paths are more likely already tested)
|
|
- Read 2-3 existing test files to match conventions exactly
|
|
- Generate unit tests. Mock all external dependencies (DB, API, Redis).
|
|
- For paths marked [→E2E]: generate integration/E2E tests using the project's E2E framework (Playwright, Cypress, Capybara, etc.)
|
|
- For paths marked [→EVAL]: generate eval tests using the project's eval framework, or flag for manual eval if none exists
|
|
- Write tests that exercise the specific uncovered path with real assertions
|
|
- Run each test. Passes → keep the change and report its path; the parent commits in Step 15.
|
|
- Fails → diagnose whether the test/fixture is invalid or a declared product contract is broken. Correct a demonstrated test defect once; preserve a valid red regression and route the reproduced product failure through the parent's fix/approval flow. Never delete or weaken it to manufacture green; retain unresolved coverage in the diagram.
|
|
|
|
Caps: 30 code paths max, 20 tests generated max (code + user flow combined), 2-min per-test exploration cap.
|
|
|
|
If no test framework AND user declined bootstrap → diagram only, no generation. Note: "Test generation skipped — no test framework configured."
|
|
|
|
**Diff is test-only changes:** Return a skipped audit with null coverage, zero gaps, and "No new application code paths to audit."
|
|
|
|
**6. After-count and coverage summary:**
|
|
|
|
```bash
|
|
# Count test files after generation
|
|
git ls-files 2>/dev/null | grep -E '(\.test\.|\.spec\.|_test\.|_spec\.)' | wc -l
|
|
```
|
|
|
|
For PR body: `Tests: {before} → {after} (+{delta} new)`
|
|
Coverage line: `Test Coverage Audit: N new code paths. M covered (X%). K tests generated, awaiting parent commit.`
|
|
|
|
### Test Plan Artifact
|
|
|
|
After producing the coverage diagram, write a test plan artifact so `/qa` and `/qa-only` can consume it:
|
|
|
|
```bash
|
|
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)" && mkdir -p ~/.gstack/projects/$SLUG
|
|
USER=$(whoami)
|
|
DATETIME=$(date +%Y%m%d-%H%M%S)
|
|
```
|
|
|
|
Write to `~/.gstack/projects/{slug}/{user}-{branch}-ship-test-plan-{datetime}.md`:
|
|
|
|
```markdown
|
|
# Test Plan
|
|
Generated by /ship on {date}
|
|
Branch: {branch}
|
|
Repo: {owner/repo}
|
|
|
|
## Affected Pages/Routes
|
|
- {URL path} — {what to test and why}
|
|
|
|
## Key Interactions to Verify
|
|
- {interaction description} on {page}
|
|
|
|
## Edge Cases
|
|
- {edge case} on {page}
|
|
|
|
## Critical Paths
|
|
- {end-to-end flow that must work}
|
|
```
|
|
|
|
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
|
|
{"coverage_pct":N,"gaps":N,"diagram":"<full markdown coverage diagram for PR body>","tests_added":["path",...]}
|
|
Use null for an undetermined or skipped coverage percentage, not zero. Include every remaining gap in the diagram so the parent can target a second pass.
|
|
````
|
|
|
|
**Parent processing:**
|
|
|
|
1. Read the subagent's final output. Parse the LAST line as JSON.
|
|
2. Store `coverage_pct` (for Step 20 metrics), `gaps` (user summary), `tests_added` (for the commit).
|
|
3. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
|
|
4. Print a one-line summary: `Coverage: {coverage_pct}%, {gaps} gaps. {tests_added.length} tests added.`
|
|
|
|
**Audit failure:** On failure, invalid JSON or no completion after ~10 minutes,
|
|
stop the child and confirm it stopped before running the same audit inline.
|
|
Fallback recovers the audit; it does not pass or bypass the coverage gate.
|
|
Apply that gate to the recovered results, including its undetermined-percentage
|
|
and test-only rules. Preserve partial results as incomplete, not passing coverage.
|
|
|
|
|
|
**7. Coverage gate:**
|
|
|
|
The parent owns this gate, including after inline fallback. Generated tests stay uncommitted until Step 15. Use Step 7's remaining generation allowance; supply it and the remaining gaps to the same audit prompt. At the cap, omit A and recommend stopping; the listed risk choices remain available.
|
|
|
|
Before proceeding, check CLAUDE.md for a `## Test Coverage` section with `Minimum:` and `Target:` fields. If found, use those percentages. Otherwise use defaults: Minimum = 60%, Target = 80%.
|
|
|
|
Using the coverage percentage from the diagram in substep 4 (the `COVERAGE: X/Y (Z%)` line):
|
|
|
|
- **>= target:** Pass. "Coverage gate: PASS ({X}%)." Continue.
|
|
- **>= minimum, < target:** Use AskUserQuestion:
|
|
- "AI-assessed coverage is {X}%. {N} code paths are untested. Target is {target}%."
|
|
- RECOMMENDATION: Choose A because untested code paths are where production bugs hide.
|
|
- Options:
|
|
A) Generate more tests for remaining gaps (recommended)
|
|
B) Ship anyway — I accept the coverage risk
|
|
C) These paths don't need tests — mark as intentionally uncovered
|
|
- If A and allowance remains: dispatch one generation pass, then re-evaluate here. At the cap, offer only B/C or stop; never another generation pass.
|
|
- If B: Continue. Include in PR body: "Coverage gate: {X}% — user accepted risk."
|
|
- If C: Continue. Include in PR body: "Coverage gate: {X}% — {N} paths intentionally uncovered."
|
|
|
|
- **< minimum:** Use AskUserQuestion:
|
|
- "AI-assessed coverage is critically low ({X}%). {N} of {M} code paths have no tests. Minimum threshold is {minimum}%."
|
|
- RECOMMENDATION: Choose A because less than {minimum}% means more code is untested than tested.
|
|
- Options:
|
|
A) Generate tests for remaining gaps (recommended)
|
|
B) Override — ship with low coverage (I understand the risk)
|
|
- If A and allowance remains: dispatch one generation pass, then re-evaluate here. At the cap, offer only B or stop; never another generation pass.
|
|
- If B: Continue. Include in PR body: "Coverage gate: OVERRIDDEN at {X}%."
|
|
|
|
**Coverage percentage undetermined:** If the coverage diagram doesn't produce a clear numeric percentage (ambiguous output, parse error), **skip the gate** with: "Coverage gate: could not determine percentage — skipping." Do not default to 0% or block.
|
|
|
|
**Test-only diffs:** Skip the gate (same as the existing fast-path).
|
|
|
|
**100% coverage:** "Coverage gate: PASS (100%)." Continue.
|
|
|
|
---
|