mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 18:36:54 +02:00
v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998)
This commit is contained in:
1 parent
943105f109
commit
96764e80a6
56 files changed
+2445
-216
No files matched your search
@@ -18,9 +18,17 @@ the initial audit, failures and zero-test results. Re-entry never resets it.
|
||||
Two passes already used means no further generation; read-only reassessment uses no pass.
|
||||
|
||||
**Subagent prompt:** Supply `<base>`, Step 4's framework/bootstrap decision,
|
||||
permitted paths/commands, remaining gaps, passes used and generation allowance.
|
||||
No allowance means audit only; missing permission is not approval. Preserve the
|
||||
30-path/20-test/2-minute per-test caps.
|
||||
permitted paths/commands, remaining gaps, passes used and generation allowance,
|
||||
plus the CLAUDE.md `## Test Coverage` values the gate below reads (`Generation cap:`,
|
||||
`Base control:`, `Base control budget:`). No allowance means audit only; missing
|
||||
permission is not approval. Preserve the 30-path/5-tests-per-pass/2-minute per-test caps.
|
||||
|
||||
**Before the first dispatch,** sweep base-control worktrees a previous interrupted run left behind:
|
||||
|
||||
```bash
|
||||
git worktree prune
|
||||
find "${TMPDIR:-/tmp}" -maxdepth 1 -name 'gstack-base-control.*' -mmin +10 2>/dev/null | while IFS= read -r d; do git worktree remove --force "$d/wt" >/dev/null 2>&1; rm -rf "$d"; done
|
||||
```
|
||||
|
||||
````text
|
||||
You are running a ship-workflow test coverage audit. Run `git diff origin/<base>` to include uncommitted tracked changes; also read relevant non-ignored untracked source/tests. Do not commit or push. Perform only this audit; return unresolved user decisions to the parent instead of asking or advancing to another workflow step.
|
||||
@@ -30,16 +38,52 @@ Generation: <allowed|audit-only>; passes used: <N> of 2. Audit-only overrides ev
|
||||
{{TEST_COVERAGE_AUDIT_SHIP}}
|
||||
|
||||
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
|
||||
{"coverage_pct":N,"gaps":N,"diagram":"<full markdown coverage diagram for PR body>","tests_added":["path",...]}
|
||||
Use null for an undetermined or skipped coverage percentage, not zero. Include every remaining gap in the diagram so the parent can target a second pass.
|
||||
{"coverage_pct":N,"gaps":N,"diagram":"<full markdown coverage diagram for PR body>","tests_added":["path",...],"coverage_pct_value":N,"weak_gaps":[{"path":"...","existing_test":"...","reason":"star_one|gate_failed|unrated"}],"tests_extended":["path",...],"tests_rejected":[{"path_or_gap":"...","reason_code":"...","reason":"..."}],"regression_proof":{"red_at_head":N,"base_green":N,"base_unavailable":N}}
|
||||
`coverage_pct` is Y (paths with any test), `coverage_pct_value` is X (paths with a ★★/★★★ test), `gaps` counts only paths with no test. Use null for an undetermined or skipped coverage percentage, not zero. Include every remaining gap in the diagram so the parent can target a second pass.
|
||||
````
|
||||
|
||||
**Parent processing:**
|
||||
|
||||
1. Read the subagent's final output. Parse the LAST line as JSON.
|
||||
2. Store `coverage_pct` (for Step 20 metrics), `gaps` (user summary), `tests_added` (for the commit).
|
||||
3. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
|
||||
4. Print a one-line summary: `Coverage: {coverage_pct}%, {gaps} gaps. {tests_added.length} tests added.`
|
||||
2. Store `coverage_pct`, `coverage_pct_value`, `gaps`, `weak_gaps`, `tests_added`,
|
||||
`tests_extended`, `tests_rejected` and `regression_proof`. A missing new key counts
|
||||
as empty; say so in the summary (an older installed prompt must not fail the gate).
|
||||
A key with the wrong type (for example `weak_gaps` not an array) is ignored the same
|
||||
way and printed as: {{TEST_VALUE_MESSAGE:malformedKey}}
|
||||
3. **Machine checks** on every test written in this run (`tests_added` and
|
||||
`tests_extended`): a value-card header with four non-empty fields (else
|
||||
`incomplete_card`); `protects` unique across the run after casefolding and stripping
|
||||
punctuation and repeated whitespace (a later duplicate is `duplicate_protects`); seam
|
||||
`none`, or a named seam with at least one non-test caller (N = 0 or an unavailable
|
||||
caller check is `needs_seam`). Move each failure to `tests_rejected` with its
|
||||
`reason_code`, then remove it before anything else reads the diff: an untracked new
|
||||
file is deleted; for a tracked file, revert only this run's hunk with Edit, never
|
||||
the whole file.
|
||||
|
||||
```bash
|
||||
while IFS= read -r f; do
|
||||
[ -n "$f" ] || continue
|
||||
if git ls-files --error-unmatch -- "$f" >/dev/null 2>&1; then echo "REVERT_HUNK: $f"; else rm -f -- "$f" && echo "REMOVED: $f"; fi
|
||||
done <<'REJECTED'
|
||||
<one rejected test path per line>
|
||||
REJECTED
|
||||
```
|
||||
|
||||
No `tests_rejected` path may remain on disk as a new file. If every test written in
|
||||
a pass is rejected, print {{TEST_VALUE_MESSAGE:allRejected}}
|
||||
4. **Rating dispatch.** When this run wrote tests that survived the machine checks,
|
||||
dispatch one read-only Agent (`subagent_type: "general-purpose"`,
|
||||
`run_in_background: false`) with no generation permission; it uses no generation
|
||||
pass. Give it the diagram and the surviving test paths. It rates each against the
|
||||
★ rubric and the test value bar and returns a LAST-line JSON
|
||||
`{"coverage_pct_value":N,"weak_gaps":[...]}` recomputed with its ratings; use those
|
||||
two values. Until rated, this run's tests count as weak (`unrated`). If it fails,
|
||||
times out or returns invalid JSON, the gate is skipped for this run ("rating
|
||||
unavailable").
|
||||
5. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
|
||||
6. Print a one-line summary: `Coverage: {X}% value-weighted ({Y}% including {W} weakly covered paths), {gaps} gaps. {tests_added.length} tests added.`
|
||||
Bindings for the PR body's Test value line: K = `tests_added.length`,
|
||||
R = `tests_rejected.length`, E = `tests_extended.length`, W = `weak_gaps.length`.
|
||||
|
||||
**Audit failure:** On failure, invalid JSON or no completion after ~10 minutes,
|
||||
stop the child and confirm it stopped before running the same audit inline.
|
||||
|
||||
Reference in new issue
Block a user