mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
97 lines
6.0 KiB
Cheetah
97 lines
6.0 KiB
Cheetah
## Step 7: Test Coverage Audit
|
|
|
|
### Shared subagent dispatch
|
|
|
|
For Steps 7, 8 and 10, use the Agent tool with `run_in_background: false`.
|
|
Omitting the flag runs the subagent in the background. The explicit flag waits
|
|
for a result while keeping a fresh context. Do not invoke the target as a Skill
|
|
or run it inline instead. Inline work is allowed only under that section's
|
|
documented fallback, after a failed subagent has stopped.
|
|
|
|
Dispatch the audit through Agent with `subagent_type: "general-purpose"` and
|
|
`run_in_background: false`, using the shared foreground-dispatch rule above.
|
|
Wait for its LAST-line JSON before applying the coverage gate.
|
|
|
|
**Generation allowance:** Maximum 2 generation passes total per invocation.
|
|
Count each generation-authorized attempt before dispatch/inline execution, including
|
|
the initial audit, failures and zero-test results. Re-entry never resets it.
|
|
Two passes already used means no further generation; read-only reassessment uses no pass.
|
|
|
|
**Subagent prompt:** Supply `<base>`, Step 4's framework/bootstrap decision,
|
|
permitted paths/commands, remaining gaps, passes used and generation allowance,
|
|
plus the CLAUDE.md `## Test Coverage` values the gate below reads (`Generation cap:`,
|
|
`Base control:`, `Base control budget:`). No allowance means audit only; missing
|
|
permission is not approval. Preserve the 30-path/5-tests-per-pass/2-minute per-test caps.
|
|
|
|
**Before the first dispatch,** sweep base-control worktrees a previous interrupted run left behind:
|
|
|
|
```bash
|
|
git worktree prune
|
|
find "${TMPDIR:-/tmp}" -maxdepth 1 -name 'gstack-base-control.*' -mmin +10 2>/dev/null | while IFS= read -r d; do git worktree remove --force "$d/wt" >/dev/null 2>&1; rm -rf "$d"; done
|
|
```
|
|
|
|
````text
|
|
You are running a ship-workflow test coverage audit. Run `git diff origin/<base>` to include uncommitted tracked changes; also read relevant non-ignored untracked source/tests. Do not commit or push. Perform only this audit; return unresolved user decisions to the parent instead of asking or advancing to another workflow step.
|
|
|
|
Generation: <allowed|audit-only>; passes used: <N> of 2. Audit-only overrides every generation instruction below.
|
|
|
|
{{TEST_COVERAGE_AUDIT_SHIP}}
|
|
|
|
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
|
|
{"coverage_pct":N,"gaps":N,"diagram":"<full markdown coverage diagram for PR body>","tests_added":["path",...],"coverage_pct_value":N,"weak_gaps":[{"path":"...","existing_test":"...","reason":"star_one|gate_failed|unrated"}],"tests_extended":["path",...],"tests_rejected":[{"path_or_gap":"...","reason_code":"...","reason":"..."}],"regression_proof":{"red_at_head":N,"base_green":N,"base_unavailable":N}}
|
|
`coverage_pct` is Y (paths with any test), `coverage_pct_value` is X (paths with a ★★/★★★ test), `gaps` counts only paths with no test. Use null for an undetermined or skipped coverage percentage, not zero. Include every remaining gap in the diagram so the parent can target a second pass.
|
|
````
|
|
|
|
**Parent processing:**
|
|
|
|
1. Read the subagent's final output. Parse the LAST line as JSON.
|
|
2. Store `coverage_pct`, `coverage_pct_value`, `gaps`, `weak_gaps`, `tests_added`,
|
|
`tests_extended`, `tests_rejected` and `regression_proof`. A missing new key counts
|
|
as empty; say so in the summary (an older installed prompt must not fail the gate).
|
|
A key with the wrong type (for example `weak_gaps` not an array) is ignored the same
|
|
way and printed as: {{TEST_VALUE_MESSAGE:malformedKey}}
|
|
3. **Machine checks** on every test written in this run (`tests_added` and
|
|
`tests_extended`): a value-card header with four non-empty fields (else
|
|
`incomplete_card`); `protects` unique across the run after casefolding and stripping
|
|
punctuation and repeated whitespace (a later duplicate is `duplicate_protects`); seam
|
|
`none`, or a named seam with at least one non-test caller (N = 0 or an unavailable
|
|
caller check is `needs_seam`). Move each failure to `tests_rejected` with its
|
|
`reason_code`, then remove it before anything else reads the diff: an untracked new
|
|
file is deleted; for a tracked file, revert only this run's hunk with Edit, never
|
|
the whole file.
|
|
|
|
```bash
|
|
while IFS= read -r f; do
|
|
[ -n "$f" ] || continue
|
|
if git ls-files --error-unmatch -- "$f" >/dev/null 2>&1; then echo "REVERT_HUNK: $f"; else rm -f -- "$f" && echo "REMOVED: $f"; fi
|
|
done <<'REJECTED'
|
|
<one rejected test path per line>
|
|
REJECTED
|
|
```
|
|
|
|
No `tests_rejected` path may remain on disk as a new file. If every test written in
|
|
a pass is rejected, print {{TEST_VALUE_MESSAGE:allRejected}}
|
|
4. **Rating dispatch.** When this run wrote tests that survived the machine checks,
|
|
dispatch one read-only Agent (`subagent_type: "general-purpose"`,
|
|
`run_in_background: false`) with no generation permission; it uses no generation
|
|
pass. Give it the diagram and the surviving test paths. It rates each against the
|
|
★ rubric and the test value bar and returns a LAST-line JSON
|
|
`{"coverage_pct_value":N,"weak_gaps":[...]}` recomputed with its ratings; use those
|
|
two values. Until rated, this run's tests count as weak (`unrated`). If it fails,
|
|
times out or returns invalid JSON, the gate is skipped for this run ("rating
|
|
unavailable").
|
|
5. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
|
|
6. Print a one-line summary: `Coverage: {X}% value-weighted ({Y}% including {W} weakly covered paths), {gaps} gaps. {tests_added.length} tests added.`
|
|
Bindings for the PR body's Test value line: K = `tests_added.length`,
|
|
R = `tests_rejected.length`, E = `tests_extended.length`, W = `weak_gaps.length`.
|
|
|
|
**Audit failure:** On failure, invalid JSON or no completion after ~10 minutes,
|
|
stop the child and confirm it stopped before running the same audit inline.
|
|
Fallback recovers the audit; it does not pass or bypass the coverage gate.
|
|
Apply that gate to the recovered results, including its undetermined-percentage
|
|
and test-only rules. Preserve partial results as incomplete, not passing coverage.
|
|
|
|
{{TEST_COVERAGE_GATE_SHIP}}
|
|
|
|
---
|