mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture
- review-army-perf-n-plus-one: the parent copied full checklists into agent prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now. - review-design-lite: 5 of 6 captured trials reported the detector absent without probing; the probe is mandatory and its first line is reported, and the contract credits only fake-engine rule ids the checklist never names. - review-exploratory-small-cli: the fixture never gave review-log's direct invocation or status vocabulary; the model ran it through bun and wrote status "blocked". The prompt states both and the validator rejects out-of-vocabulary review statuses. Each case passed a focused paid run after repair.
This commit is contained in:
1 parent
a18e6cf655
commit
a06d22e52a
15 files changed
+155
-62
No files matched your search
@@ -68,7 +68,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)
|
||||
|
||||
If \`SCOPE_FRONTEND=false\`, skip the entire design review silently.
|
||||
|
||||
**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
|
||||
**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
|
||||
|
||||
\`\`\`bash
|
||||
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
|
||||
|
||||
@@ -95,7 +95,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
|
||||
|
||||
Construct the prompt for each specialist. The prompt includes:
|
||||
|
||||
1. The specialist's checklist content (you already read the file above)
|
||||
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
|
||||
2. Stack context: "This is a {STACK} project."
|
||||
3. Past learnings for this domain (if any exist):
|
||||
|
||||
@@ -107,7 +107,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
|
||||
|
||||
4. Instructions:
|
||||
|
||||
"You are a specialist code reviewer. Read the checklist below, then run
|
||||
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
|
||||
\`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\` to get the full diff. Apply the checklist against the diff.
|
||||
|
||||
For each finding, output a JSON object on its own line:
|
||||
@@ -126,10 +126,7 @@ If no findings: output \`NO FINDINGS\` and nothing else.
|
||||
Do not output anything else — no preamble, no summary, no commentary.
|
||||
|
||||
Stack context: {STACK}
|
||||
Past learnings: {learnings or 'none'}
|
||||
|
||||
CHECKLIST:
|
||||
{checklist content}"
|
||||
Past learnings: {learnings or 'none'}"
|
||||
|
||||
**Subagent configuration:**
|
||||
- Use \`subagent_type: "general-purpose"\`
|
||||
@@ -203,6 +200,7 @@ Only specialist findings enter this header and \`quality_score\`; core findings
|
||||
Use the merged NON-advisory specialist findings for both counts and score:
|
||||
\`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))\`
|
||||
Cap at 10 and retain for ${persistRef}. These are not final unresolved-defect totals.
|
||||
Print only this block: the stage 6 activity object and \`test_stub\` bodies are log and Fix-First data.
|
||||
Validated \`"advisory": true\` findings from any source are excluded from score,
|
||||
header, unresolved-defect totals and clean-status blockers. Show them separately;
|
||||
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
|
||||
@@ -264,13 +262,13 @@ function generateRedTeam(ctx: TemplateContext): string {
|
||||
If activated, dispatch one more subagent via the Agent tool (pass \`run_in_background: false\` — foreground; subagents default to background since ${CC_BACKGROUND_DEFAULT_SINCE}).
|
||||
|
||||
The Red Team subagent receives:
|
||||
1. The red-team checklist from \`${ctx.paths.skillRoot}/review/specialists/red-team.md\`
|
||||
2. The merged specialist findings from Step ${stepMerge} (so it knows what was already caught)
|
||||
1. The red-team checklist path \`${ctx.paths.skillRoot}/review/specialists/red-team.md\` (it reads the file)
|
||||
2. The merged specialist findings from Step ${stepMerge}, one line each (so it knows what was already caught)
|
||||
3. The git diff command
|
||||
|
||||
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
|
||||
who found the following issues: {merged findings summary}. Your job is to find what they
|
||||
MISSED. Read the checklist, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
|
||||
MISSED. Read the checklist at {red-team checklist path}, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
|
||||
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
|
||||
concerns, integration boundary issues, and failure modes that specialist checklists
|
||||
don't cover."
|
||||
|
||||
Reference in new issue
Block a user