fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture

- review-army-perf-n-plus-one: the parent copied full checklists into agent
  prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
  without probing; the probe is mandatory and its first line is reported, and
  the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
  invocation or status vocabulary; the model ran it through bun and wrote
  status "blocked". The prompt states both and the validator rejects
  out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.
This commit is contained in:
garrytan committed 2026-09-29 22:49:12 +00:00
1 parent a18e6cf655
commit a06d22e52a
15 files changed
+155 -62

No files matched your search

+1 -1
View File
@@ -693,7 +693,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
and use existing knowledge.
and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
+1 -1
View File
@@ -173,7 +173,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
and use existing knowledge.
and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
+1 -1
View File
@@ -15,7 +15,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)
If `SCOPE_FRONTEND=false`, skip the entire design review silently.
**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
```bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
+7 -9
View File
@@ -78,7 +78,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -90,7 +90,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -109,10 +109,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -181,6 +178,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log entry in Step 5.8. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -239,13 +237,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
1. The red-team checklist from `~/.claude/skills/gstack/review/specialists/red-team.md`
2. The merged specialist findings from Step 4.6 (so it knows what was already caught)
1. The red-team checklist path `~/.claude/skills/gstack/review/specialists/red-team.md` (it reads the file)
2. The merged specialist findings from Step 4.6, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."