diff --git a/TODOS.md b/TODOS.md
index 8f5959eac..7d20a689b 100644
--- a/TODOS.md
+++ b/TODOS.md
@@ -10,8 +10,7 @@
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M.
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
(one ~250 s thinking block before its single write; times out at 300 s on
- 2.1.251 in every recent run), `review-army-perf-n-plus-one` (290-300 s on a
- 12-line diff; web search plus a conditional red-team pass), the HOLD SCOPE
+ 2.1.251 in every recent run) and the HOLD SCOPE
routing case when its next brief happens not to name the mode (see the
handoff item below). Effort M each.
- **`/plan-ceo-review` skips its Step 0E mode handoff** — in 4 of 4 asked-mode
@@ -20,6 +19,11 @@
decisions: …` chat. The routing case still passes on other posture text; a
wording change moving the handoff ahead of the question log did not change
the behavior in two paid runs, so it was not shipped. Effort M.
+- **`/ship` design-lite probe wording** — `scripts/resolvers/design.ts` still
+ says "Probe for a design detector the user installed", the wording that let
+ `/review` skip its probe in 5 of 6 captured trials before this release made it
+ mandatory. The ship union ratio is at 1.3966 of 1.397, so the same sentence
+ needs a trim elsewhere first. Effort S.
- **Let pass-rate history decide the rest** — every census on this branch had
a different handful of single-trial reds. Once `eval:pass-rates` has 10 weekly
trials per case, apply the CASE_QUARANTINE entry rule instead of chasing one
diff --git a/review/SKILL.md b/review/SKILL.md
index 01757a13f..d43064512 100644
--- a/review/SKILL.md
+++ b/review/SKILL.md
@@ -693,7 +693,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
-and use existing knowledge.
+and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
diff --git a/review/SKILL.md.tmpl b/review/SKILL.md.tmpl
index 95b5423b2..9aae36eb3 100644
--- a/review/SKILL.md.tmpl
+++ b/review/SKILL.md.tmpl
@@ -173,7 +173,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
-and use existing knowledge.
+and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
diff --git a/review/design-checklist.md b/review/design-checklist.md
index fab8a7637..8c78dfff4 100644
--- a/review/design-checklist.md
+++ b/review/design-checklist.md
@@ -15,7 +15,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope 2>/dev/null)
If `SCOPE_FRONTEND=false`, skip the entire design review silently.
-**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
+**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
```bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
diff --git a/review/sections/review-army.md b/review/sections/review-army.md
index 137133909..2cc0465ec 100644
--- a/review/sections/review-army.md
+++ b/review/sections/review-army.md
@@ -78,7 +78,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
-1. The specialist's checklist content (you already read the file above)
+1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -90,7 +90,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
-"You are a specialist code reviewer. Read the checklist below, then run
+"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -109,10 +109,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
-Past learnings: {learnings or 'none'}
-
-CHECKLIST:
-{checklist content}"
+Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -181,6 +178,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log entry in Step 5.8. These are not final unresolved-defect totals.
+Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -239,13 +237,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
-1. The red-team checklist from `~/.claude/skills/gstack/review/specialists/red-team.md`
-2. The merged specialist findings from Step 4.6 (so it knows what was already caught)
+1. The red-team checklist path `~/.claude/skills/gstack/review/specialists/red-team.md` (it reads the file)
+2. The merged specialist findings from Step 4.6, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
-MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
+MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
diff --git a/scripts/resolvers/design-checklist.ts b/scripts/resolvers/design-checklist.ts
index 6c5509d25..9268704f8 100644
--- a/scripts/resolvers/design-checklist.ts
+++ b/scripts/resolvers/design-checklist.ts
@@ -68,7 +68,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope 2>/dev/null)
If \`SCOPE_FRONTEND=false\`, skip the entire design review silently.
-**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
+**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
\`\`\`bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
diff --git a/scripts/resolvers/review-army.ts b/scripts/resolvers/review-army.ts
index 707c3ddd9..3af34dcb8 100644
--- a/scripts/resolvers/review-army.ts
+++ b/scripts/resolvers/review-army.ts
@@ -95,7 +95,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
-1. The specialist's checklist content (you already read the file above)
+1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -107,7 +107,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
-"You are a specialist code reviewer. Read the checklist below, then run
+"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
\`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"\` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -126,10 +126,7 @@ If no findings: output \`NO FINDINGS\` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
-Past learnings: {learnings or 'none'}
-
-CHECKLIST:
-{checklist content}"
+Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use \`subagent_type: "general-purpose"\`
@@ -203,6 +200,7 @@ Only specialist findings enter this header and \`quality_score\`; core findings
Use the merged NON-advisory specialist findings for both counts and score:
\`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))\`
Cap at 10 and retain for ${persistRef}. These are not final unresolved-defect totals.
+Print only this block: the stage 6 activity object and \`test_stub\` bodies are log and Fix-First data.
Validated \`"advisory": true\` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -264,13 +262,13 @@ function generateRedTeam(ctx: TemplateContext): string {
If activated, dispatch one more subagent via the Agent tool (pass \`run_in_background: false\` — foreground; subagents default to background since ${CC_BACKGROUND_DEFAULT_SINCE}).
The Red Team subagent receives:
-1. The red-team checklist from \`${ctx.paths.skillRoot}/review/specialists/red-team.md\`
-2. The merged specialist findings from Step ${stepMerge} (so it knows what was already caught)
+1. The red-team checklist path \`${ctx.paths.skillRoot}/review/specialists/red-team.md\` (it reads the file)
+2. The merged specialist findings from Step ${stepMerge}, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
-MISSED. Read the checklist, run \`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
+MISSED. Read the checklist at {red-team checklist path}, run \`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
diff --git a/ship/sections/review-army.md b/ship/sections/review-army.md
index 03c7f2627..4e29932c0 100644
--- a/ship/sections/review-army.md
+++ b/ship/sections/review-army.md
@@ -303,7 +303,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
-1. The specialist's checklist content (you already read the file above)
+1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -315,7 +315,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
-"You are a specialist code reviewer. Read the checklist below, then run
+"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -334,10 +334,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
-Past learnings: {learnings or 'none'}
-
-CHECKLIST:
-{checklist content}"
+Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -406,6 +403,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log persist. These are not final unresolved-defect totals.
+Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -464,13 +462,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
-1. The red-team checklist from `~/.claude/skills/gstack/review/specialists/red-team.md`
-2. The merged specialist findings from Step 9.2 (so it knows what was already caught)
+1. The red-team checklist path `~/.claude/skills/gstack/review/specialists/red-team.md` (it reads the file)
+2. The merged specialist findings from Step 9.2, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
-MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
+MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
diff --git a/test/fixtures/golden/factory-ship-SKILL.md b/test/fixtures/golden/factory-ship-SKILL.md
index 1fb0bb370..e66d6ccda 100644
--- a/test/fixtures/golden/factory-ship-SKILL.md
+++ b/test/fixtures/golden/factory-ship-SKILL.md
@@ -2163,7 +2163,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
-1. The specialist's checklist content (you already read the file above)
+1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -2175,7 +2175,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
-"You are a specialist code reviewer. Read the checklist below, then run
+"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -2194,10 +2194,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
-Past learnings: {learnings or 'none'}
-
-CHECKLIST:
-{checklist content}"
+Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -2266,6 +2263,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log persist. These are not final unresolved-defect totals.
+Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -2324,13 +2322,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
-1. The red-team checklist from `$GSTACK_ROOT/review/specialists/red-team.md`
-2. The merged specialist findings from Step 9.2 (so it knows what was already caught)
+1. The red-team checklist path `$GSTACK_ROOT/review/specialists/red-team.md` (it reads the file)
+2. The merged specialist findings from Step 9.2, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
-MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
+MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/ HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
diff --git a/test/fixtures/review-design-lite-reports-ci-36633323521.json b/test/fixtures/review-design-lite-reports-ci-36633323521.json
new file mode 100644
index 000000000..db116ab94
--- /dev/null
+++ b/test/fixtures/review-design-lite-reports-ci-36633323521.json
@@ -0,0 +1,35 @@
+{
+ "provenance": "review-output.md bodies written by the review-design-lite Write tool call in Periodic Evals native captures (native-captures-ci--1-eval-slices-N). scanRan records whether the transcript ran gstack-design-detect.ts scan. Only the 36633323521 t1 capture probed and scanned; the others never ran the probe yet the legacy contract accepted t3 and 36629958451 t1 because they mention 'detector' or the checklist-named [ai-color-palette]. The local focused run t2 (aba80c8 plus the first checklist repair) inferred absence from a file listing and never probed.",
+ "reports": [
+ {
+ "run": "36633323521",
+ "trial": "t1",
+ "scanRan": true,
+ "report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote; git-native fallback). Diff: `landing.html` (+41), `styles.css` (+86/−1).\nScope: `SCOPE_FRONTEND=true` → design checklist applied. No `DESIGN.md` → universal principles.\nDetector: `IMPECCABLE_READY`, exit 2, 6 hits (low-contrast ×3, skipped-heading, ai-color-palette, marketing-buzzword). Detector `file` field names `test/fixtures/review-eval-design-slop.html`, which does not exist in this repo — hits are credited only where they match code I read in `landing.html`/`styles.css`. No ignored rules in config.\nFixes are listed by tier but **not applied** — this run was requested as a written report.\n\n---\n\n## Pre-Landing Review: 4 issues (0 critical, 4 informational)\n\nPass 1 (SQL, races, LLM trust boundary, shell injection, enum completeness): nothing applicable — static HTML/CSS, no data or execution paths.\n\n**AUTO-FIX:**\n- [INFORMATIONAL] (confidence: 9/10) `styles.css:1`, `landing.html:10,17` — Eval annotations (`/* Planted design anti-patterns for E2E eval — 7 issues */`, ``) ship with the page.\n Fix: delete the fixture comments before landing.\n- [INFORMATIONAL] (confidence: 8/10) `landing.html:14` — `