Files
gstack/ship/sections/test-coverage.md.tmpl
T

97 lines
6.0 KiB
Cheetah

## Step 7: Test Coverage Audit
### Shared subagent dispatch
For Steps 7, 8 and 10, use the Agent tool with `run_in_background: false`.
Omitting the flag runs the subagent in the background. The explicit flag waits
for a result while keeping a fresh context. Do not invoke the target as a Skill
or run it inline instead. Inline work is allowed only under that section's
documented fallback, after a failed subagent has stopped.
Dispatch the audit through Agent with `subagent_type: "general-purpose"` and
`run_in_background: false`, using the shared foreground-dispatch rule above.
Wait for its LAST-line JSON before applying the coverage gate.
**Generation allowance:** Maximum 2 generation passes total per invocation.
Count each generation-authorized attempt before dispatch/inline execution, including
the initial audit, failures and zero-test results. Re-entry never resets it.
Two passes already used means no further generation; read-only reassessment uses no pass.
**Subagent prompt:** Supply `<base>`, Step 4's framework/bootstrap decision,
permitted paths/commands, remaining gaps, passes used and generation allowance,
plus the CLAUDE.md `## Test Coverage` values the gate below reads (`Generation cap:`,
`Base control:`, `Base control budget:`). No allowance means audit only; missing
permission is not approval. Preserve the 30-path/5-tests-per-pass/2-minute per-test caps.
**Before the first dispatch,** sweep base-control worktrees a previous interrupted run left behind:
```bash
git worktree prune
find "${TMPDIR:-/tmp}" -maxdepth 1 -name 'gstack-base-control.*' -mmin +10 2>/dev/null | while IFS= read -r d; do git worktree remove --force "$d/wt" >/dev/null 2>&1; rm -rf "$d"; done
```
````text
You are running a ship-workflow test coverage audit. Run `git diff origin/<base>` to include uncommitted tracked changes; also read relevant non-ignored untracked source/tests. Do not commit or push. Perform only this audit; return unresolved user decisions to the parent instead of asking or advancing to another workflow step.
Generation: <allowed|audit-only>; passes used: <N> of 2. Audit-only overrides every generation instruction below.
{{TEST_COVERAGE_AUDIT_SHIP}}
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
{"coverage_pct":N,"gaps":N,"diagram":"<full markdown coverage diagram for PR body>","tests_added":["path",...],"coverage_pct_value":N,"weak_gaps":[{"path":"...","existing_test":"...","reason":"star_one|gate_failed|unrated"}],"tests_extended":["path",...],"tests_rejected":[{"path_or_gap":"...","reason_code":"...","reason":"..."}],"regression_proof":{"red_at_head":N,"base_green":N,"base_unavailable":N}}
`coverage_pct` is Y (paths with any test), `coverage_pct_value` is X (paths with a ★★/★★★ test), `gaps` counts only paths with no test. Use null for an undetermined or skipped coverage percentage, not zero. Include every remaining gap in the diagram so the parent can target a second pass.
````
**Parent processing:**
1. Read the subagent's final output. Parse the LAST line as JSON.
2. Store `coverage_pct`, `coverage_pct_value`, `gaps`, `weak_gaps`, `tests_added`,
`tests_extended`, `tests_rejected` and `regression_proof`. A missing new key counts
as empty; say so in the summary (an older installed prompt must not fail the gate).
A key with the wrong type (for example `weak_gaps` not an array) is ignored the same
way and printed as: {{TEST_VALUE_MESSAGE:malformedKey}}
3. **Machine checks** on every test written in this run (`tests_added` and
`tests_extended`): a value-card header with four non-empty fields (else
`incomplete_card`); `protects` unique across the run after casefolding and stripping
punctuation and repeated whitespace (a later duplicate is `duplicate_protects`); seam
`none`, or a named seam with at least one non-test caller (N = 0 or an unavailable
caller check is `needs_seam`). Move each failure to `tests_rejected` with its
`reason_code`, then remove it before anything else reads the diff: an untracked new
file is deleted; for a tracked file, revert only this run's hunk with Edit, never
the whole file.
```bash
while IFS= read -r f; do
[ -n "$f" ] || continue
if git ls-files --error-unmatch -- "$f" >/dev/null 2>&1; then echo "REVERT_HUNK: $f"; else rm -f -- "$f" && echo "REMOVED: $f"; fi
done <<'REJECTED'
<one rejected test path per line>
REJECTED
```
No `tests_rejected` path may remain on disk as a new file. If every test written in
a pass is rejected, print {{TEST_VALUE_MESSAGE:allRejected}}
4. **Rating dispatch.** When this run wrote tests that survived the machine checks,
dispatch one read-only Agent (`subagent_type: "general-purpose"`,
`run_in_background: false`) with no generation permission; it uses no generation
pass. Give it the diagram and the surviving test paths. It rates each against the
★ rubric and the test value bar and returns a LAST-line JSON
`{"coverage_pct_value":N,"weak_gaps":[...]}` recomputed with its ratings; use those
two values. Until rated, this run's tests count as weak (`unrated`). If it fails,
times out or returns invalid JSON, the gate is skipped for this run ("rating
unavailable").
5. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
6. Print a one-line summary: `Coverage: {X}% value-weighted ({Y}% including {W} weakly covered paths), {gaps} gaps. {tests_added.length} tests added.`
Bindings for the PR body's Test value line: K = `tests_added.length`,
R = `tests_rejected.length`, E = `tests_extended.length`, W = `weak_gaps.length`.
**Audit failure:** On failure, invalid JSON or no completion after ~10 minutes,
stop the child and confirm it stopped before running the same audit inline.
Fallback recovers the audit; it does not pass or bypass the coverage gate.
Apply that gate to the recovered results, including its undetermined-percentage
and test-only rules. Preserve partial results as incomplete, not passing coverage.
{{TEST_COVERAGE_GATE_SHIP}}
---