fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture

- review-army-perf-n-plus-one: the parent copied full checklists into agent
  prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
  without probing; the probe is mandatory and its first line is reported, and
  the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
  invocation or status vocabulary; the model ran it through bun and wrote
  status "blocked". The prompt states both and the validator rejects
  out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.
This commit is contained in:
garrytan committed 2026-09-29 22:49:12 +00:00
1 parent a18e6cf655
commit a06d22e52a
15 files changed
+155 -62

No files matched your search

+1 -1
View File
@@ -68,7 +68,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)
If \`SCOPE_FRONTEND=false\`, skip the entire design review silently.
**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
\`\`\`bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
+7 -9
View File
@@ -95,7 +95,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -107,7 +107,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
\`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -126,10 +126,7 @@ If no findings: output \`NO FINDINGS\` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use \`subagent_type: "general-purpose"\`
@@ -203,6 +200,7 @@ Only specialist findings enter this header and \`quality_score\`; core findings
Use the merged NON-advisory specialist findings for both counts and score:
\`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))\`
Cap at 10 and retain for ${persistRef}. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and \`test_stub\` bodies are log and Fix-First data.
Validated \`"advisory": true\` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -264,13 +262,13 @@ function generateRedTeam(ctx: TemplateContext): string {
If activated, dispatch one more subagent via the Agent tool (pass \`run_in_background: false\` — foreground; subagents default to background since ${CC_BACKGROUND_DEFAULT_SINCE}).
The Red Team subagent receives:
1. The red-team checklist from \`${ctx.paths.skillRoot}/review/specialists/red-team.md\`
2. The merged specialist findings from Step ${stepMerge} (so it knows what was already caught)
1. The red-team checklist path \`${ctx.paths.skillRoot}/review/specialists/red-team.md\` (it reads the file)
2. The merged specialist findings from Step ${stepMerge}, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."