fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture

- review-army-perf-n-plus-one: the parent copied full checklists into agent
  prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
  without probing; the probe is mandatory and its first line is reported, and
  the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
  invocation or status vocabulary; the model ran it through bun and wrote
  status "blocked". The prompt states both and the validator rejects
  out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.
This commit is contained in:
garrytan committed 2026-09-29 22:49:12 +00:00
1 parent a18e6cf655
commit a06d22e52a
15 files changed
+155 -62

No files matched your search

+6 -2
View File
@@ -10,8 +10,7 @@
budgets. The CI image stays on 2.1.251 until those cases get faster. Effort M.
- **Recurring reds to repair, not rerun** — `plan-design-review-plan-mode`
(one ~250 s thinking block before its single write; times out at 300 s on
2.1.251 in every recent run), `review-army-perf-n-plus-one` (290-300 s on a
12-line diff; web search plus a conditional red-team pass), the HOLD SCOPE
2.1.251 in every recent run) and the HOLD SCOPE
routing case when its next brief happens not to name the mode (see the
handoff item below). Effort M each.
- **`/plan-ceo-review` skips its Step 0E mode handoff** — in 4 of 4 asked-mode
@@ -20,6 +19,11 @@
decisions: …` chat. The routing case still passes on other posture text; a
wording change moving the handoff ahead of the question log did not change
the behavior in two paid runs, so it was not shipped. Effort M.
- **`/ship` design-lite probe wording** — `scripts/resolvers/design.ts` still
says "Probe for a design detector the user installed", the wording that let
`/review` skip its probe in 5 of 6 captured trials before this release made it
mandatory. The ship union ratio is at 1.3966 of 1.397, so the same sentence
needs a trim elsewhere first. Effort S.
- **Let pass-rate history decide the rest** — every census on this branch had
a different handful of single-trial reds. Once `eval:pass-rates` has 10 weekly
trials per case, apply the CASE_QUARANTINE entry rule instead of chasing one
+1 -1
View File
@@ -693,7 +693,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
and use existing knowledge.
and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
+1 -1
View File
@@ -173,7 +173,7 @@ _aside_exec "Search the web for {framework} {version} {pattern} current best pra
```
Without Aside `READY`, use WebSearch if available; with neither, disclose the gap
and use existing knowledge.
and use existing knowledge. Don't wait on research: run it alongside independent work, such as specialist dispatch.
### Shared-code opportunities (core pass)
+1 -1
View File
@@ -15,7 +15,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)
If `SCOPE_FRONTEND=false`, skip the entire design review silently.
**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On `IMPECCABLE_READY`, scan the changed frontend files before reading them yourself:
```bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
+7 -9
View File
@@ -78,7 +78,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -90,7 +90,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -109,10 +109,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -181,6 +178,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log entry in Step 5.8. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -239,13 +237,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
1. The red-team checklist from `~/.claude/skills/gstack/review/specialists/red-team.md`
2. The merged specialist findings from Step 4.6 (so it knows what was already caught)
1. The red-team checklist path `~/.claude/skills/gstack/review/specialists/red-team.md` (it reads the file)
2. The merged specialist findings from Step 4.6, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
+1 -1
View File
@@ -68,7 +68,7 @@ source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)
If \`SCOPE_FRONTEND=false\`, skip the entire design review silently.
**0. Mechanical pass first.** Probe for a design detector the user installed (this pass never offers to install one; the design skills ask, once) and, on \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
**0. Mechanical pass first.** Always run the probe below for a design detector the user installed. It searches the environment and install caches, which no file listing shows, so never assume or report a detector absent without its output; state its first line in the design review. This pass never offers to install one (the design skills ask, once). On \`${SENTINEL.READY}\`, scan the changed frontend files before reading them yourself:
\`\`\`bash
bun --no-env-file run ~/.claude/skills/gstack/bin/gstack-design-detect.ts probe --host claude
+7 -9
View File
@@ -95,7 +95,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -107,7 +107,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
\`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -126,10 +126,7 @@ If no findings: output \`NO FINDINGS\` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use \`subagent_type: "general-purpose"\`
@@ -203,6 +200,7 @@ Only specialist findings enter this header and \`quality_score\`; core findings
Use the merged NON-advisory specialist findings for both counts and score:
\`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))\`
Cap at 10 and retain for ${persistRef}. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and \`test_stub\` bodies are log and Fix-First data.
Validated \`"advisory": true\` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -264,13 +262,13 @@ function generateRedTeam(ctx: TemplateContext): string {
If activated, dispatch one more subagent via the Agent tool (pass \`run_in_background: false\` — foreground; subagents default to background since ${CC_BACKGROUND_DEFAULT_SINCE}).
The Red Team subagent receives:
1. The red-team checklist from \`${ctx.paths.skillRoot}/review/specialists/red-team.md\`
2. The merged specialist findings from Step ${stepMerge} (so it knows what was already caught)
1. The red-team checklist path \`${ctx.paths.skillRoot}/review/specialists/red-team.md\` (it reads the file)
2. The merged specialist findings from Step ${stepMerge}, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run \`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"\`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
+7 -9
View File
@@ -303,7 +303,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -315,7 +315,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -334,10 +334,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -406,6 +403,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log persist. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -464,13 +462,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
1. The red-team checklist from `~/.claude/skills/gstack/review/specialists/red-team.md`
2. The merged specialist findings from Step 9.2 (so it knows what was already caught)
1. The red-team checklist path `~/.claude/skills/gstack/review/specialists/red-team.md` (it reads the file)
2. The merged specialist findings from Step 9.2, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
+7 -9
View File
@@ -2163,7 +2163,7 @@ so they run in parallel. Each subagent has fresh context — no prior review bia
Construct the prompt for each specialist. The prompt includes:
1. The specialist's checklist content (you already read the file above)
1. The specialist's checklist path from the selection above (the subagent reads it; never paste its content)
2. Stack context: "This is a {STACK} project."
3. Past learnings for this domain (if any exist):
@@ -2175,7 +2175,7 @@ If learnings are found, include them: "Past learnings for this domain: {learning
4. Instructions:
"You are a specialist code reviewer. Read the checklist below, then run
"You are a specialist code reviewer. Read the checklist at {checklist path}, then run
`DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"` to get the full diff. Apply the checklist against the diff.
For each finding, output a JSON object on its own line:
@@ -2194,10 +2194,7 @@ If no findings: output `NO FINDINGS` and nothing else.
Do not output anything else — no preamble, no summary, no commentary.
Stack context: {STACK}
Past learnings: {learnings or 'none'}
CHECKLIST:
{checklist content}"
Past learnings: {learnings or 'none'}"
**Subagent configuration:**
- Use `subagent_type: "general-purpose"`
@@ -2266,6 +2263,7 @@ Only specialist findings enter this header and `quality_score`; core findings do
Use the merged NON-advisory specialist findings for both counts and score:
`quality_score = max(0, 10 - (critical_count * 2 + informational_count * 0.5))`
Cap at 10 and retain for the review-log persist. These are not final unresolved-defect totals.
Print only this block: the stage 6 activity object and `test_stub` bodies are log and Fix-First data.
Validated `"advisory": true` findings from any source are excluded from score,
header, unresolved-defect totals and clean-status blockers. Show them separately;
they remain ASK-only, never auto-applied. Real defects follow normal Fix-First.
@@ -2324,13 +2322,13 @@ completion. Advice never permits edits while readers are active or replaces a re
If activated, dispatch one more subagent via the Agent tool (pass `run_in_background: false` — foreground; subagents default to background since Claude Code v2.1.198).
The Red Team subagent receives:
1. The red-team checklist from `$GSTACK_ROOT/review/specialists/red-team.md`
2. The merged specialist findings from Step 9.2 (so it knows what was already caught)
1. The red-team checklist path `$GSTACK_ROOT/review/specialists/red-team.md` (it reads the file)
2. The merged specialist findings from Step 9.2, one line each (so it knows what was already caught)
3. The git diff command
Prompt: "You are a red team reviewer. The code has already been reviewed by N specialists
who found the following issues: {merged findings summary}. Your job is to find what they
MISSED. Read the checklist, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
MISSED. Read the checklist at {red-team checklist path}, run `DIFF_BASE=$(git merge-base origin/<base> HEAD) && git diff "$DIFF_BASE"`, and look for gaps.
Output findings as JSON objects (same schema as the specialists). Focus on cross-cutting
concerns, integration boundary issues, and failure modes that specialist checklists
don't cover."
@@ -0,0 +1,35 @@
{
"provenance": "review-output.md bodies written by the review-design-lite Write tool call in Periodic Evals native captures (native-captures-ci-<run>-1-eval-slices-N). scanRan records whether the transcript ran gstack-design-detect.ts scan. Only the 36633323521 t1 capture probed and scanned; the others never ran the probe yet the legacy contract accepted t3 and 36629958451 t1 because they mention 'detector' or the checklist-named [ai-color-palette]. The local focused run t2 (aba80c8 plus the first checklist repair) inferred absence from a file listing and never probed.",
"reports": [
{
"run": "36633323521",
"trial": "t1",
"scanRan": true,
"report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote; git-native fallback). Diff: `landing.html` (+41), `styles.css` (+86/−1).\nScope: `SCOPE_FRONTEND=true` → design checklist applied. No `DESIGN.md` → universal principles.\nDetector: `IMPECCABLE_READY`, exit 2, 6 hits (low-contrast ×3, skipped-heading, ai-color-palette, marketing-buzzword). Detector `file` field names `test/fixtures/review-eval-design-slop.html`, which does not exist in this repo — hits are credited only where they match code I read in `landing.html`/`styles.css`. No ignored rules in config.\nFixes are listed by tier but **not applied** — this run was requested as a written report.\n\n---\n\n## Pre-Landing Review: 4 issues (0 critical, 4 informational)\n\nPass 1 (SQL, races, LLM trust boundary, shell injection, enum completeness): nothing applicable — static HTML/CSS, no data or execution paths.\n\n**AUTO-FIX:**\n- [INFORMATIONAL] (confidence: 9/10) `styles.css:1`, `landing.html:10,17` — Eval annotations (`/* Planted design anti-patterns for E2E eval — 7 issues */`, `<!-- Issue 6: ... -->`) ship with the page.\n Fix: delete the fixture comments before landing.\n- [INFORMATIONAL] (confidence: 8/10) `landing.html:14` — `<button>` has no `type`; inside a future `<form>` it defaults to `submit`, and it currently does nothing at all (no handler, no link).\n Fix: `<a class=\"button\" href=\"/signup\">` if it navigates, or `<button type=\"button\">` with a handler.\n\n**NEEDS INPUT:**\n- [INFORMATIONAL] (confidence: 8/10) `landing.html:11,18,36` — Landmark regions are plain `<div>`s (`.hero`, `.features`, `.footer`); no `<main>`, `<header>`, `<footer>`, `<section>`. Screen readers get no page structure.\n Recommended fix: `<header class=\"hero\">`, `<main><section class=\"features\">`, `<footer class=\"footer\">`.\n- [INFORMATIONAL] (confidence: 7/10) `landing.html:38` — `<a href=\"#\">Terms of Service</a>` is a dead placeholder link (Completeness Gap — the real URL is a one-line change).\n Recommended fix: point at the actual terms page, or drop the link until one exists.\n\nShared-code opportunities: none — two files, no repeated authored blocks with ≥2 callers.\nTODOS.md: absent, skipped. Docs staleness: no root docs describe this page, skipped.\n\n---\n\n## Design Review: 19 issues (5 auto-fixable, 11 need input, 3 possible)\n\n**AUTO-FIX** (mechanical CSS, HIGH confidence):\n- [HIGH] `styles.css:6` — Blacklisted font: `font-family: 'Papyrus', sans-serif`. Also a fantasy face falling back to a sans stack, so the two renderings look nothing alike.\n Fix: pick a real text face (self-hosted or system stack) and give it a matching generic fallback.\n- [HIGH] [tiny-text] `styles.css:7` — Body text `font-size: 14px` (base was 16px on main; this PR regresses it).\n Fix: `font-size: 16px` (or `1rem`).\n- [HIGH] [tiny-text] `styles.css:71` — `.small-link { font-size: 11px }` — legal link at 11px is unreadable on mobile.\n Fix: `font-size: 1rem` (or 0.875rem minimum for secondary text).\n- [HIGH] `styles.css:61` — `button { outline: none }` with no replacement focus indicator — keyboard users lose the focus ring on the only CTA.\n Fix: remove `outline: none`; add `button:focus-visible { outline: 2px solid currentColor; outline-offset: 2px; }`.\n- [HIGH] `styles.css:77-78` — `!important` ×2 (`color: red !important; margin-left: 10px !important`). `.override` is a single-class selector with nothing to override.\n Fix: drop both `!important`s; the rule already wins on specificity.\n\n**NEEDS INPUT** (design judgment):\n- [MEDIUM] [ai-color-palette] `styles.css:14` — `linear-gradient(135deg, #6366f1, #8b5cf6)` hero, `#6366f1` button (`:62`), `#ede9fe` icon circles (`:51`), `#1e1b4b` footer (`:84`). Indigo→violet is the canonical AI palette. (detector + checklist)\n Recommended fix: one solid brand color the product actually owns; kill the gradient.\n- [MEDIUM] Generic hero copy `landing.html:12-13,37` — \"Welcome to Our Platform\", \"Your all-in-one solution for everything you need\", \"Unlock the power of our platform today\". Three of the five phrases the checklist greps for, verbatim.\n Recommended fix: say what the product does, for whom, in one concrete sentence.\n- [MEDIUM] [marketing-buzzword] `landing.html:32` — \"streamline your workflow effortlessly\" hits two buzzwords in one clause; `:37` \"Unlock\". (detector + checklist)\n Recommended fix: replace with the specific thing the feature does.\n- [MEDIUM] `landing.html:14` — \"Get Started\" is the only CTA on the page.\n Recommended fix: name the outcome (\"Start a free project\", \"Book a demo\").\n- [MEDIUM] Centered everything — `text-align: center` on `.hero`, `.hero h1`, `.hero p`, `.features`, `.feature-card`, `.footer` (`styles.css:15,21,26,36,42,82`): 6 of 7 text containers, well past the 60% threshold.\n Recommended fix: left-align feature copy and footer; center only Line truncated
},
{
"run": "36633323521",
"trial": "t2",
"scanRan": false,
"report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote; git-native fallback). Diff: `landing.html` (+41), `styles.css` (+86/−1).\nGreptile step skipped (no PR/remote). No `DESIGN.md` — universal design principles apply.\nFrontend files changed → design checklist applied. No source files were modified in this run (findings only); AUTO-FIX items are listed with the exact change to apply.\n\n---\n\n## Pre-Landing Review: 4 issues (0 critical, 4 informational)\n\nPass 1 (SQL / races / LLM trust / shell / enums): nothing applicable — static HTML+CSS, no data layer, no scripts.\n\n**AUTO-FIX:**\n- [styles.css:1] `/* Planted design anti-patterns for E2E eval — 7 issues */` and the `Issue N:` comments (styles.css:3-4, 12, 30, 57, 74; landing.html:10, 17) narrate the defects rather than the code → remove before landing.\n- [landing.html:11-38] Page is built from `<div>`s only — no `<main>`/`<header>`/`<section>`/`<footer>` landmarks → `<div class=\"hero\">` → `<header class=\"hero\">`, `<div class=\"features\">` → `<section class=\"features\">`, `<div class=\"footer\">` → `<footer class=\"footer\">`.\n\n**NEEDS INPUT:**\n- [landing.html:38] `<a href=\"#\" class=\"small-link\">Terms of Service</a>` is a dead placeholder link that navigates to page top.\n Recommended fix: point at the real ToS URL or drop the link until one exists.\n- [styles.css:75, 83] `.override { color: red }` on the `#1e1b4b` footer ≈ 4.0:1 contrast, below WCAG AA 4.5:1 for 14px text. (Medium confidence, verify this is actually an issue — computed, not measured.) `white` on `#6366f1` for the button is ≈ 4.5:1, borderline.\n Recommended fix: pick a footer accent from the palette with ≥4.5:1 against `#1e1b4b`, or drop the red entirely (see design finding on `.override`).\n\n---\n\n## Design Review: 20 issues (5 auto-fixable, 12 need input, 3 possible)\n\n**AUTO-FIX:**\n- [styles.css:6] `font-family: 'Papyrus', sans-serif` — blacklisted font [HIGH] → replace with a real typeface (e.g. a self-hosted humanist sans); keep `sans-serif` fallback.\n- [styles.css:7] `body { font-size: 14px }` — body text under 16px [HIGH, tiny-text]. Note the diff also *removed* the previous `body { font-size: 16px; }` → restore `font-size: 16px`.\n- [styles.css:59] `button { outline: none; }` with no replacement focus indicator — removes keyboard focus visibility [HIGH] → delete the line and add `button:focus-visible { outline: 2px solid currentColor; outline-offset: 2px; }`.\n- [styles.css:75-76] `color: red !important; margin-left: 10px !important;` — `!important` in new CSS [HIGH]. `.override` is a single-class selector with nothing competing; the flags do nothing → remove both `!important`s.\n- [styles.css:69] `.small-link { font-size: 11px }` — text under 16px [HIGH, tiny-text] → `font-size: 16px` (or ≥14px with a documented reason for de-emphasized legal links).\n\n**NEEDS INPUT:**\n- [styles.css:14] `linear-gradient(135deg, #6366f1, #8b5cf6)` — the canonical indigo→violet AI gradient [MEDIUM, ai-color-palette]. The whole palette is derived from it: button `#6366f1` (:60), icon bg `#ede9fe` (:50), footer `#1e1b4b` (:84).\n Recommended fix: one solid brand color the product owns; if a gradient is kept, make it subtle and not blue-to-purple.\n- [landing.html:12-13] `Welcome to Our Platform` / `Your all-in-one solution for everything you need` — generic hero copy [MEDIUM]. Says nothing about what the product does.\n Recommended fix: lead with the concrete outcome for a named audience.\n- [landing.html:36] `Unlock the power of our platform today` — generic CTA copy [MEDIUM]; `unlock` is also a listed buzzword.\n Recommended fix: state what the click does.\n- [landing.html:31] `streamline your workflow effortlessly` — two buzzwords in one sentence [MEDIUM, marketing-buzzword]. Also `amazing feature that will change your life` (:21), `sets us apart from the competition` (:26).\n Recommended fix: replace with what each feature literally does.\n- [landing.html:14] `Get Started` is the only CTA on the page [MEDIUM].\n Recommended fix: name the outcome (e.g. \"Create a free workspace\", \"See the demo\").\n- [landing.html:20, 25, 30] `&#9733;` ★ `&#9889;` ⚡ `&#9881;` ⚙ — unicode symbols/emoji used as icons [MEDIUM]. Render inconsistently across platforms and can't be sized/colored reliably.\n Recommended fix: inline SVG icons, or drop the icons.\n- [styles.css:15, 21, 26, 35, 41, 82] `text-align: center` on every text container (hero, h1, p, features, cards, footer) — 100% centered [MEDIUM].\n Recommended fix: left-align body copy and feature descriptions; center only the hero heading if at all.\n- [landing.html:12 → 22] Heading hierarchy skips `h1` → `h3`; no `h2` in the file [HIGH, typography].\n Recommended fix: feature titles become `<h2>` (or add an `<h2>` section heading and keep `<h3>`s).\n- [styles.css:59-67, 69-72] `button` and `.smaLine truncated
},
{
"run": "36633323521",
"trial": "t3",
"scanRan": false,
"report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote → git-native fallback). Diff: `landing.html` (+41), `styles.css` (+86/-1).\nGreptile: skipped (no PR). Design detector: not installed, skipped. DESIGN.md: none → universal principles.\n\n**Fixes NOT applied.** `styles.css:1` declares the file a planted E2E fixture (\"Planted design anti-patterns for E2E eval — 7 issues\"), so editing it would alter the fixture. Every AUTO-FIX below has its exact replacement listed; say the word and I'll apply them.\n\n---\n\n## Pre-Landing Review: 5 issues (0 critical, 5 informational)\n\nPass 1 (SQL / races / LLM trust / shell / enum): nothing applicable — static HTML + CSS, no data paths.\n\n**AUTO-FIXED:** (none applied — see note above; these are the AUTO-FIX-tier items)\n- [INFORMATIONAL] (confidence: 9/10) `styles.css:1` — Comment announces the stylesheet is a planted test fixture with \"7 issues\"; must not ship on a real landing page. → Delete the header comment (and the per-issue `/* Issue N: ... */` comments at :3, :4, :12, :30, :46, :55, :71).\n- [INFORMATIONAL] (confidence: 8/10) `styles.css:71-74` — `.override` is named for its mechanism, not its meaning, and only exists to carry `!important`. → Rename to what it is (e.g. `.footer-tagline`) once the `!important`s are removed (design finding D2).\n\n**NEEDS INPUT:**\n- [INFORMATIONAL] (confidence: 8/10) `landing.html:38` + `styles.css:83` — Footer `<a class=\"small-link\">` gets the UA default link colour (`#0000EE`); `.footer { color: white }` does not cascade onto `<a>`. Blue on `#1e1b4b` is ~1.7:1 contrast — the Terms of Service link is effectively invisible.\n Recommended fix: `.footer a { color: inherit; }` (plus a visible underline / `text-underline-offset`).\n- [INFORMATIONAL] (confidence: 7/10) `styles.css:72` + `styles.css:80` — `color: red` on the `#1e1b4b` footer is ~4.0:1, below WCAG AA 4.5:1 for body-size text (and it's 14px, see D3).\n Recommended fix: drop the red override; let the paragraph inherit `white`, or pick a palette tint that clears 4.5:1.\n- [INFORMATIONAL] (confidence: 7/10) `landing.html:14`, `landing.html:38` — Completeness gap: `<button>Get Started</button>` has no `type`, no handler and no destination; `<a href=\"#\">Terms of Service</a>` is a dead placeholder link. Both CTAs are inert.\n Recommended fix: make the primary CTA an `<a href=\"/signup\">` styled as a button (or give the button a `type` + handler); point Terms at the real URL.\n\n---\n\n## Design Review: 18 issues (4 auto-fixable, 11 need input, 3 possible)\n\n**AUTO-FIXED:** (none applied — see note above; these are the AUTO-FIX-tier items)\n- D1 [HIGH] (confidence: 10/10) `styles.css:56` — `button { outline: none; }` with no replacement focus indicator; keyboard users lose the focus ring on the only CTA. → Remove `outline: none`; add `button:focus-visible { outline: 2px solid currentColor; outline-offset: 2px; }`.\n- D2 [HIGH] (confidence: 10/10) `styles.css:72-73` — `!important` ×2 in `.override` (`color: red !important; margin-left: 10px !important;`). Nothing competes with `.override`'s specificity; the escape hatch is unneeded. → Delete both `!important`s.\n- D3 [HIGH] [tiny-text] (confidence: 10/10) `styles.css:7` — `body { font-size: 14px; }` — base body text under 16px (and this diff *removes* the previous `body { font-size: 16px; }`). → `font-size: 16px` (or `1rem`).\n- D4 [HIGH] [tiny-text] (confidence: 9/10) `styles.css:67` — `.small-link { font-size: 11px; }` — 11px link text is below any readable body floor. → Bump to ≥ 14px for a legal footer link, 16px preferred.\n\n**NEEDS INPUT:**\n- D5 [HIGH] Blacklisted font (confidence: 10/10) `styles.css:6` — `font-family: 'Papyrus', sans-serif;` — Papyrus is on the blacklist, and the fallback is a bare generic.\n Recommended fix: pick a real typeface with a proper stack (e.g. a self-hosted or system-available serif/sans, then generic fallback). Avoid the overused-default list too (Inter, Roboto, Poppins…).\n- D6 [MEDIUM] [ai-color-palette] (confidence: 9/10) `styles.css:14` — `linear-gradient(135deg, #6366f1, #8b5cf6)` is the canonical indigo→violet AI gradient. The whole palette follows it: button `#6366f1` (:57), icon circle `#ede9fe` (:49), footer `#1e1b4b` (:80).\n Recommended fix: one solid brand colour the palette owns for the hero and CTA; drop the gradient.\n- D7 [MEDIUM] Generic hero copy (confidence: 10/10) `landing.html:12-13` — \"Welcome to Our Platform\" / \"Your all-in-one solution for everything you need\" — two of the checklist's literal grep strings, plus `<title>Our Platform</title>` (:7).\n Recommended fix: say what the product does and for whom in the h1; make the subhead a concrete claim.\n- D8 [MEDIUM] [marketing-buzzword] (confidence: 9/10) `landing.html:32`, `landing.html:37` — \"streamline your workflow effortlessly\" (streamline + effortless) and \"Unlock the poweLine truncated
},
{
"run": "36629958451",
"trial": "t1",
"scanRan": false,
"report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote; git-native fallback). Diff: `landing.html` (+41), `styles.css` (+86/−1).\nGreptile step skipped (no PR). No `DESIGN.md` → universal design principles. No `TODOS.md`.\nNote: the source is annotated \"Planted design anti-patterns for E2E eval\" — findings were **written, not applied**. Every AUTO-FIX row below is mechanical and can be applied on request.\n\n---\n\n## Pre-Landing Review: 8 issues (0 critical, 8 informational)\n\nPass 1 (SQL / Race / LLM trust / Shell / Enum): not applicable — static HTML+CSS only.\n\n**AUTO-FIXED (eligible — not applied):**\n- [INFORMATIONAL] (confidence: 9/10) styles.css:5-7 — Diff **deletes** the pre-existing `body { font-size: 16px; }` and replaces it with `14px`; this is a regression on a shared stylesheet, not just a new rule. → Restore `font-size: 16px` (or `1rem`).\n- [INFORMATIONAL] (confidence: 9/10) styles.css:1-3,12,30,45,56,70 — `/* Planted design anti-patterns for E2E eval — 7 issues */` and `/* Issue N: ... */` comments, plus `<!-- Issue 6/7 -->` in landing.html:10,17, ship to users via view-source. → Remove eval scaffolding comments before landing.\n- [INFORMATIONAL] (confidence: 8/10) landing.html:14 — `<button>` with no `type`; defaults to `submit` if ever placed in a form, and has no handler or `href` today. → `<a class=\"button\" href=\"/signup\">` or `<button type=\"button\">` wired to an action.\n\n**NEEDS INPUT:**\n- [INFORMATIONAL] (confidence: 9/10) landing.html:19-33 — Completeness gap: placeholder copy shipped (\"Feature One/Two/Three\", \"A short description of this amazing feature that will change your life\"). Page is not launch-ready.\n Recommended fix: replace with real feature names and one concrete sentence each describing what the product does.\n- [INFORMATIONAL] (confidence: 9/10) landing.html:38 — `<a href=\"#\">Terms of Service</a>` is a dead link; legal link pointing to `#` is a completeness gap, not a stub.\n Recommended fix: point at the real `/terms` URL.\n- [INFORMATIONAL] (confidence: 8/10) landing.html:11,18,36 — Non-semantic `<div class=\"hero|features|footer\">`; no landmarks (`<main>`, `<section>`, `<footer>`), so screen readers get a flat page.\n Recommended fix: `<main><section class=\"hero\">…</section><section class=\"features\">…</section></main><footer>…</footer>`.\n- [INFORMATIONAL] (confidence: 8/10) landing.html:38 + styles.css:83-88 — `.footer` sets `color: white` but `<a>` does not inherit color; default link blue (`#0000ee`) on `#1e1b4b` is ~1.7:1 contrast, and the text is 11px. Effectively unreadable.\n Recommended fix: `.footer a { color: inherit; text-underline-offset: 0.15em; }` and remove the 11px size (see design review).\n- [INFORMATIONAL] (confidence: 6/10) styles.css:76-79 — `margin-left: 10px !important` on a `text-align: center` paragraph knocks the centered text 5px off axis for no stated reason. Medium confidence, verify this is actually an issue.\n Recommended fix: delete the margin (and the `!important`, see below).\n\n---\n\n## Design Review: 20 issues (6 auto-fixable, 12 need input, 2 possible)\n\nFrontend scope: `landing.html`, `styles.css` (both read in full). Mechanical detector (`gstack-design-detect`) not installed on this host — checklist pass only.\n\n**AUTO-FIXED (eligible — not applied):**\n- [HIGH] styles.css:7 — [tiny-text] Body text `font-size: 14px` (<16px). → `font-size: 16px`. (Same line as the code-review regression above — one fix.)\n- [HIGH] styles.css:66 — `outline: none` on `button` with no replacement focus indicator; keyboard users lose focus visibility entirely. → Delete the line, or `button:focus-visible { outline: 2px solid currentColor; outline-offset: 2px; }`.\n- [HIGH] styles.css:78 — `color: red !important` — specificity escape hatch. → Drop `!important`; `.footer .override { color: … }` already wins.\n- [HIGH] styles.css:79 — `margin-left: 10px !important`. → Remove (see code review: it also breaks centering).\n- [HIGH] styles.css:81-84 — `.small-link { font-size: 11px }` — text under 16px on an interactive legal link. → `font-size: 1rem` (or `0.875rem` minimum for footer meta); combined with `padding: 4px 8px` the touch target is ~19px tall — bump padding to reach 44px.\n- [HIGH] styles.css:6 — **Blacklisted font: `Papyrus`** as the primary body face, falling back to generic `sans-serif`. → Pick a real typeface (a non-`[overused-font]` face, self-hosted or via `@font-face`) with a sane fallback stack.\n\n**NEEDS INPUT (design judgment):**\n- [MEDIUM] styles.css:14 — [ai-color-palette] `linear-gradient(135deg, #6366f1, #8b5cf6)` — the exact indigo→violet range the checklist names as the AI default palette; `#6366f1` reused on `button:67`, `#ede9fe` on `.icon-circle:50`, `#1e1b4b` on `.footer:86`. The whole palette is Tailwind indigo/violet.\n Recommended fix: choose one brand coloLine truncated
},
{
"run": "local-focused-aba80c8",
"trial": "t2",
"scanRan": false,
"report": "# /review — feature/add-landing-page vs main\n\nBase branch: `main` (no `origin` remote; local merge-base used). Diff: `landing.html` (+41), `styles.css` (+86/−1).\nGreptile: skipped (no PR). DESIGN.md: none — universal principles. Design detector: not installed on this host — mechanical pass skipped, checklist pass only.\n\nNote: `styles.css:1` describes these files as planted eval fixtures, so AUTO-FIX items are listed but **not applied** to avoid altering the fixture. Approve and they can be applied in one pass.\n\n---\n\n## Pre-Landing Review: 5 issues (0 critical, 5 informational)\n\nPass 1 (SQL, races, LLM trust, shell, enums): no applicable code — static HTML/CSS only. Nothing to flag.\n\n**AUTO-FIX (pending approval):**\n- `[INFORMATIONAL] (confidence: 9/10) styles.css:7` — Base font size regressed from 16px (removed at old line 1) to 14px → restore `font-size: 16px`. (Same root cause as design tiny-text finding below.)\n- `[INFORMATIONAL] (confidence: 9/10) styles.css:76-79` — `.override` exists only to force `color`/`margin-left` with `!important`; nothing else targets `.footer p`, so no specificity conflict exists → drop `!important`, and drop the `margin-left: 10px` (it nudges a centered `<p>` 10px off-center inside `text-align: center`).\n\n**NEEDS INPUT:**\n- `[INFORMATIONAL] (confidence: 8/10) landing.html:38` — `<a href=\"#\">Terms of Service</a>` is a placeholder link that scrolls to top; ships a dead legal link.\n Recommended fix: point at the real terms URL, or remove until one exists.\n- `[INFORMATIONAL] (confidence: 7/10) landing.html:11,18,36` — Page is `<div>`-only: no `<header>`/`<main>`/`<section>`/`<footer>`, so screen readers get no landmarks.\n Recommended fix: `.hero` → `<header>` or `<section>` inside `<main>`, `.features` → `<section aria-labelledby=…>`, `.footer` → `<footer>`.\n- `[INFORMATIONAL] (confidence: 7/10) landing.html:10,17` / `styles.css:1,3,4,12,30,46,59,75` — \"Issue N: …\" / \"Planted … for E2E eval\" comments describe the review harness, not the code; they would ship to production.\n Recommended fix: strip before landing (or confirm these files are test-only fixtures).\n\nShared-code opportunities: none. Single page, no repeated authored logic; the three `.feature-card` blocks are markup, not code worth templating at this size.\nTODOS.md: absent, skipped. Doc staleness: no root docs describe this page, skipped.\n\n---\n\n## Design Review: 18 issues (5 auto-fixable, 10 need input, 3 possible)\n\n**AUTO-FIX (pending approval — HIGH confidence, mechanical CSS):**\n- `[styles.css:61]` `outline: none` on `button` with no replacement — removes the keyboard focus ring entirely → delete the line and add `button:focus-visible, a:focus-visible { outline: 2px solid currentColor; outline-offset: 2px; }`.\n- `[styles.css:77]` `color: red !important` → `color: red` (no competing rule; see code review above).\n- `[styles.css:78]` `margin-left: 10px !important` → remove.\n- `[styles.css:7]` [tiny-text] `body { font-size: 14px }` — body text under 16px → `16px`.\n- `[styles.css:71]` [tiny-text] `.small-link { font-size: 11px }` — 11px is illegible on most displays and below WCAG-comfortable size for a legal link → `min 14px`, ideally `1rem`.\n\n**NEEDS INPUT (design judgment):**\n- `[styles.css:6]` **Blacklisted font:** `font-family: 'Papyrus', sans-serif` as the site-wide body/display face.\n Recommended fix: choose a real typeface for the page's voice (and note the fallback is the browser's generic sans, so most visitors get an unstyled default anyway).\n- `[styles.css:14]` [ai-color-palette] `linear-gradient(135deg, #6366f1, #8b5cf6)` — the exact indigo→violet gradient the checklist calls out, echoed by `button` `#6366f1` (`:62`), `.icon-circle` `#ede9fe` (`:51`) and `.footer` `#1e1b4b` (`:84`). Entire palette is Tailwind indigo/violet defaults.\n Recommended fix: pick a palette the brand owns; use one solid color for the hero and CTA.\n- `[landing.html:12-13]` **Generic hero copy:** \"Welcome to Our Platform\" / \"Your all-in-one solution for everything you need\". Also `<title>Our Platform</title>` (`:7`).\n Recommended fix: lead with what the product does and for whom; the `<title>` should name the product.\n- `[landing.html:14]` \"Get Started\" is the only CTA on the page, and the button is a bare `<button>` with no handler or form — it does nothing.\n Recommended fix: name the outcome (\"Start a free trial\", \"Book a demo\") and wire it to a link/action.\n- `[landing.html:32,37]` [marketing-buzzword] \"streamline your workflow effortlessly\", \"Unlock the power of our platform today\". Also filler feature copy: \"amazing feature that will change your life\", \"sets us apart from the competition\" (`:22,27`).\n Recommended fix: replace with concrete claims (what it does, measurable outcome).\n- `[landing.html:20,25,30]` Emoji/symbol glyphs (★ ⚡ ⚙ via `&#9733; &#9889; &#9881;`) used Line truncated
}
]
}
+11
View File
@@ -18,3 +18,14 @@ export function installFakeImpeccable(prefix = 'gstack-fake-impeccable-'): { dir
fs.copyFileSync(DETECT_SAMPLE, path.join(dir, 'impeccable-detect-sample.json')); // the shim's documented default output, beside it
return { dir, bin };
}
/** Sample rule ids the design checklist never names: a review can carry them only from the detector's rows. */
export function detectorOnlyRuleIds(checklist: string): string[] {
const rules = JSON.parse(fs.readFileSync(DETECT_SAMPLE, 'utf-8')) as Array<{ antipattern: string }>;
return [...new Set(rules.map(rule => rule.antipattern))].filter(id => !checklist.includes(id));
}
export function carriesDetectorRows(review: string, checklist: string): boolean {
const text = review.toLowerCase();
return detectorOnlyRuleIds(checklist).some(id => new RegExp(`(?<![\\w-])${id}(?![\\w-])`).test(text));
}
+8 -1
View File
@@ -248,6 +248,8 @@ export function qaCallerCommandAllowed(command: string, workflowCommands: string
|| /^\/?[\w./-]+\/bin\/gstack-(?:review-read|specialist-stats)$/.test(text);
}
const REVIEW_RECORD_STATUSES: Record<QaCaller, string[]> = { review: ['clean', 'issues_found'], ship: ['clean', 'issues_found', 'unavailable'] };
export function validateCallerEvidence(input: {
caller: QaCaller;
result: Pick<SkillTestResult, 'transcript' | 'exitReason'>;
@@ -300,6 +302,11 @@ export function validateCallerEvidence(input: {
if (tool.name === 'Bash' && !qaCallerCommandAllowed(String(tool.input.command), input.workflowCommands, deadline)) {
errors.push('command outside declared caller observation interface');
}
const record = tool.name === 'Bash' ? String(tool.input.command).match(/gstack-review-log '(.*)'(?: --finish \S+)?$/s)?.[1] : undefined;
const status = record?.match(/"status"\s*:\s*"([^"]*)"/)?.[1] ?? '';
if (record && /"skill"\s*:\s*"review"/.test(record) && !REVIEW_RECORD_STATUSES[input.caller].includes(status)) {
errors.push(`review record status outside the ${input.caller} vocabulary: ${status}`);
}
if (tool.name === 'Read' && /\/(?:browse|devex-review)\/SKILL\.md$|\/qa\/sections\/(?:browser-[^/]+|qa-patterns)\.md$/.test(file)) {
errors.push(`unexpected browser/DX load: ${file}`);
}
@@ -633,7 +640,7 @@ process.exit(exit ?? 127);
export function qaCallerSessionOptions(fixture: QaCallerFixture, runId: string): Parameters<typeof runSkillTest>[0] {
return {
prompt: `Load gstack's /${fixture.caller} supplied parent phase from caller-${fixture.caller}.md and resume it on the selected working-tree diff against origin/main. This excerpt comes from ${fixture.runtime}/${fixture.caller}/SKILL.md; resolve installed-relative references there, not from the excerpt file or product directory. That path identifies the asset base, not another entrypoint: do not read or invoke the full parent SKILL.md or rerun its preamble. Earlier preamble/branch/base setup is complete; use the existing local origin/main ref without fetch. Earlier-phase asset locators are ${fixture.runtime}/review/checklist.md and ${fixture.runtime}/qa/templates/functional-report-template.md. Read those files directly when referenced; recursive Glob does not follow the installed asset symlinks. Cross-project learnings are configured off in this owned fixture. ${fixture.reviewStart ? `The actual review-start helper already captured REVIEW_START=${fixture.reviewStart} for this unchanged core pass; retain that token. ` : ''} This fixture evaluates only the supplied parent phase, not later publication stages. Read README.md for the project contract and commands. Use diagnostic-client commands such as \`bun scripts/probe.ts <literal>\` for exploratory discoveries and their checkpoint evidence. A required \`bun run test\` is separate suite verification: report it as verification, never as a diagnostic observation or checkpoint anchor/target. Use the production evidence helper to publish each diagnostic checkpoint as \`reports/exploration-NNN.json\`, not inside a nested directory; do not transcribe its observed payload. ${fixture.caseId === 'ship-exploratory-plan-checks' ? 'The previously discovered plan is PLAN.md.' : 'No plan file was found.'} There is no remote service and no release publication is authorized. There is no interactive approver; do not invent answers or permission. Keep normal parent decision gates. Before every completion report or bookkeeping log, read HANDOFF.md and reports/HANDOFF.md if present for any concurrent collaborator update, await the results, and compare evidence with current inputs. A gstack-review-log completed:true record is a completion, not preliminary bookkeeping; a later handoff read cannot validate an earlier completion.\n\nDeadline bookkeeping additionally permits \`bun ${fixture.runtime}/bin/gstack-qa-deadline start ${fixture.cwd}/reports/deadline.json SECONDS [EARLIER_UTC]\`, \`bun ${fixture.runtime}/bin/gstack-qa-deadline status ${fixture.cwd}/reports/deadline.json\`, and \`bun ${fixture.runtime}/bin/gstack-qa-deadline run ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These are closed literal forms: SECONDS must be positive and at most 300, EARLIER_UTC is the optional caller absolute deadline: use the section clock's Hard deadline UTC, never its Runner entry UTC, reserve-start time or a clock-read time. The child is only the existing diagnostic client with zero or one literal argument. Resolve these exact helper and state paths; do not use variables, another helper, another state file, nested wrappers, scripts, operators or substitutions. Only this helper may create or change reports/deadline.json and its .qa-deadline- temporary files; never use Write/Edit/MultiEdit on those paths. Record the full outer run command in checkpoints and evidence; keep the unchanged child JSON as observed, separate from prefixed guard diagnostics. A completed expired guard-run is not a probe or a pass: retain its unused checkpoint, report not-run coverage and do not restart the deadline. Keep the 12-probe smoke limit. Required suites and explicit plan checks are outside the bounded smoke budget, not permission to reset it.\n\nFunctional evidence uses the same production helper and existing diagnostic client: \`bun ${fixture.runtime}/bin/gstack-qa-evidence capture ${fixture.cwd}/reports NNN --public --deadline ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These diagnostic receipts are declared public/synthetic, so --public is approved; a fresh three-digit ID is required each time. Explicit plan probes outside the smoke budget may replace --deadline with --timeout-ms 10000; this does not reset or bypass the smoke deadline. Publish causal intent with \`bun ${fixture.runtime}/bin/gstack-qa-evidence checkpoint ${fixture.cwd}/reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command'\`; quote arguments literally. For complex quoting, Write only capture, observationCommand, hypothesis and nextCommand to reports/intent.json; publish with the same helper: checkpoint REPORT_ROOT NNN intent.json. Materialize is supported when evidence.json is required. Sources stay inside reports. Decide to execute the next probe before publishing its checkpoint, then await successful publication and dispatch that exact probe. If you defer an optional idea or stop exploration, do not publish a checkpoint for it; descri Line truncated
prompt: `Load gstack's /${fixture.caller} supplied parent phase from caller-${fixture.caller}.md and resume it on the selected working-tree diff against origin/main. This excerpt comes from ${fixture.runtime}/${fixture.caller}/SKILL.md; resolve installed-relative references there, not from the excerpt file or product directory. That path identifies the asset base, not another entrypoint: do not read or invoke the full parent SKILL.md or rerun its preamble. Earlier preamble/branch/base setup is complete; use the existing local origin/main ref without fetch. Earlier-phase asset locators are ${fixture.runtime}/review/checklist.md and ${fixture.runtime}/qa/templates/functional-report-template.md. Read those files directly when referenced; recursive Glob does not follow the installed asset symlinks. Cross-project learnings are configured off in this owned fixture. ${fixture.reviewStart ? `The actual review-start helper already captured REVIEW_START=${fixture.reviewStart} for this unchanged core pass; retain that token. ` : ''} This fixture evaluates only the supplied parent phase, not later publication stages. Read README.md for the project contract and commands. Use diagnostic-client commands such as \`bun scripts/probe.ts <literal>\` for exploratory discoveries and their checkpoint evidence. A required \`bun run test\` is separate suite verification: report it as verification, never as a diagnostic observation or checkpoint anchor/target. Use the production evidence helper to publish each diagnostic checkpoint as \`reports/exploration-NNN.json\`, not inside a nested directory; do not transcribe its observed payload. ${fixture.caseId === 'ship-exploratory-plan-checks' ? 'The previously discovered plan is PLAN.md.' : 'No plan file was found.'} There is no remote service and no release publication is authorized. There is no interactive approver; do not invent answers or permission. Keep normal parent decision gates. Before every completion report or bookkeeping log, read HANDOFF.md and reports/HANDOFF.md if present for any concurrent collaborator update, await the results, and compare evidence with current inputs. A gstack-review-log completed:true record is a completion, not preliminary bookkeeping; a later handoff read cannot validate an earlier completion.\n\nDeadline bookkeeping additionally permits \`bun ${fixture.runtime}/bin/gstack-qa-deadline start ${fixture.cwd}/reports/deadline.json SECONDS [EARLIER_UTC]\`, \`bun ${fixture.runtime}/bin/gstack-qa-deadline status ${fixture.cwd}/reports/deadline.json\`, and \`bun ${fixture.runtime}/bin/gstack-qa-deadline run ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These are closed literal forms: SECONDS must be positive and at most 300, EARLIER_UTC is the optional caller absolute deadline: use the section clock's Hard deadline UTC, never its Runner entry UTC, reserve-start time or a clock-read time. The child is only the existing diagnostic client with zero or one literal argument. Resolve these exact helper and state paths; do not use variables, another helper, another state file, nested wrappers, scripts, operators or substitutions. Only this helper may create or change reports/deadline.json and its .qa-deadline- temporary files; never use Write/Edit/MultiEdit on those paths. Record the full outer run command in checkpoints and evidence; keep the unchanged child JSON as observed, separate from prefixed guard diagnostics. A completed expired guard-run is not a probe or a pass: retain its unused checkpoint, report not-run coverage and do not restart the deadline. Keep the 12-probe smoke limit. Required suites and explicit plan checks are outside the bounded smoke budget, not permission to reset it.\n\nFunctional evidence uses the same production helper and existing diagnostic client: \`bun ${fixture.runtime}/bin/gstack-qa-evidence capture ${fixture.cwd}/reports NNN --public --deadline ${fixture.cwd}/reports/deadline.json -- bun scripts/probe.ts [literal]\`. These diagnostic receipts are declared public/synthetic, so --public is approved; a fresh three-digit ID is required each time. Explicit plan probes outside the smoke budget may replace --deadline with --timeout-ms 10000; this does not reset or bypass the smoke deadline. Publish causal intent with \`bun ${fixture.runtime}/bin/gstack-qa-evidence checkpoint ${fixture.cwd}/reports NNN CAPTURE_ID 'full prior capture command' 'causal hypothesis' 'full next capture command'\`; quote arguments literally. For complex quoting, Write only capture, observationCommand, hypothesis and nextCommand to reports/intent.json; publish with the same helper: checkpoint REPORT_ROOT NNN intent.json. Materialize is supported when evidence.json is required. Sources stay inside reports. Decide to execute the next probe before publishing its checkpoint, then await successful publication and dispatch that exact probe. If you defer an optional idea or stop exploration, do not publish a checkpoint for it; descri Line truncated
appendSystemPrompt: `Caller execution scheduling (fixture contract):
This session has at most 25 assistant turns, including required verification and final artifacts. The command boundary applies to each Bash call, not to the number of independent tool calls in an assistant turn.
After required clock and approval prerequisites settle, issue independent source Reads and read-only discovery together as separate native tool calls once their paths and inputs are known. Wait for their results before decisions that depend on them.
+23
View File
@@ -118,6 +118,29 @@ describe('caller native-event observer controls', () => {
expect(validateCallerEvidence(observed)).toContain('review completion preceded handoff freshness decision');
});
test('captured gate-census-5 review record: bun-wrapped helper and receipt status are outside the interface', () => {
// ci-36629958451-1-gate-census-5 review-exploratory-small-cli, native events 68 and 70 (fixture paths shortened).
const record = (status: string) => `/runtime/bin/gstack-review-log '{"skill":"review","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","commit":"'"$(git rev-parse --short HEAD)"'","branch":"caller-change","status":"${status}","completed":false,"converged":false,"critical":1,"informational":0,"findings":[{"path":"scale.ts","line":3,"category":"functional-contract","severity":"CRITICAL","fingerprint":"scale.ts:3:functional-contract","action":"ask-pending"}]}' --finish d45404cf-ba19-4bc6-a505-9801e9322f54`;
const errorsFor = (command: string, caller: 'review' | 'ship' = 'review') => {
const observed = { ...evidence(), caller };
observed.result.transcript.push(...nativeCall('record', 'Bash', { command }, 'Saved'));
return validateCallerEvidence(observed);
};
expect(errorsFor(`bun ${record('blocked')}`)).toContain('command outside declared caller observation interface');
expect(errorsFor(record('blocked'))).toEqual(['review record status outside the review vocabulary: blocked']);
expect(errorsFor(record('issues_found'))).toEqual([]);
expect(errorsFor(record('unavailable'))).toEqual(['review record status outside the review vocabulary: unavailable']);
expect(errorsFor(record('unavailable'), 'ship').filter(error => error.includes('vocabulary'))).toEqual([]);
expect(errorsFor(generatedReviewRecord('/runtime/bin/gstack-review-log', 'native-token').replace('"status":"clean"', '"status":"blocked"'))).toEqual([]);
const prompt = (id: 'review-exploratory-small-cli' | 'ship-exploratory-small-cli') => {
const fixture = createQaCallerFixture(id, { installRuntime: false });
try { return qaCallerSessionOptions(fixture, 'free-control').prompt; } finally { fs.rmSync(fixture.root, { recursive: true, force: true }); }
};
expect(prompt('review-exploratory-small-cli')).toContain("/bin/gstack-review-log '<JSON>' --finish <token>`, never through bun");
expect(prompt('review-exploratory-small-cli')).toContain('otherwise issues_found; a review stopped at a gate records completed:false');
expect(prompt('ship-exploratory-small-cli')).toContain('otherwise issues_found (unavailable for missing dispatched reviewer output)');
});
test('a later unchanged handoff reread does not invalidate an already completed freshness decision', () => {
const observed = evidence();
observed.result.transcript.push(
+36 -16
View File
@@ -11,12 +11,17 @@ const CASES = [
['review-enum-completeness', 300, 15],
['review-design-lite', 400, 35],
] as const;
for (const [id, workMs, maxTurns] of CASES) {
test.each(['success', 'timeout'])(`${id} records late results before finalization: %s`, scenario => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'review-finalization-'));
const script = path.join(dir, 'registration.test.ts');
const facts = path.join(dir, 'events.jsonl');
fs.writeFileSync(script, `
const SYNTHETIC_REPORT = 'SQL injection. Returned enum status critical. Papyrus font family;14px font-size;outline focus;!important;purple gradient;generic hero copy;3-column feature grid;detector [low-contrast] x3.';
const DESIGN_CAPTURES = JSON.parse(fs.readFileSync(path.join(ROOT, 'test/fixtures/review-design-lite-reports-ci-36633323521.json'), 'utf8')) as {
reports: Array<{ run: string; trial: string; scanRan: boolean; report: string }>;
};
function runRegistration(id: string, scenario: string, report: string) {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'review-finalization-'));
const script = path.join(dir, 'registration.test.ts');
const facts = path.join(dir, 'events.jsonl');
const evalDir = path.join(dir, 'eval');
fs.writeFileSync(script, `
import { describe, expect, mock, test } from 'bun:test';
import * as fs from 'node:fs';
import * as path from 'node:path';
@@ -50,8 +55,7 @@ mock.module(path.join(root, 'test/helpers/session-runner.ts'), () => ({
const target = selected === 'review-enum-completeness'
? opts.prompt.match(/Write your review findings once to (\\S+)/)[1]
: path.join(opts.workingDirectory, 'review-output.md');
fs.writeFileSync(target,
'SQL injection. Returned enum status critical. Papyrus font family;14px font-size;outline focus;!important;purple gradient;generic hero copy;3-column feature grid;impeccable detector [ai-color-palette].');
fs.writeFileSync(target, ${JSON.stringify(report)});
}
return { attemptId: id, exitReason: timeout ? 'timeout' : 'success', duration: opts.timeout,
model: 'free-fixture-model', toolCalls: [], browseErrors: [], output: '', transcript: [],
@@ -60,15 +64,22 @@ mock.module(path.join(root, 'test/helpers/session-runner.ts'), () => ({
}));
await import(path.join(root, ${JSON.stringify(PAID_FILE)}));
`);
const retries = retriesForFiles([PAID_FILE]);
expect(retries).toBe(0);
const child = Bun.spawnSync([process.execPath, ...buildPaidShardArgs([script], resolvePaidShardTimeoutMs([PAID_FILE]), 2, retries)], {
cwd: ROOT, timeout: 15_000, stdout: 'pipe', stderr: 'pipe',
env: { ...process.env, EVALS: '', EVALS_ALL: '', TMPDIR: dir, TMP: dir, TEMP: dir, GSTACK_EVAL_DIR: evalDir },
});
const violations = path.join(evalDir, 'contract-violations.jsonl');
return { dir, facts, exitCode: child.exitCode, output: child.stdout.toString() + child.stderr.toString(),
violations: fs.existsSync(violations) ? fs.readFileSync(violations, 'utf8') : '' };
}
for (const [id, workMs, maxTurns] of CASES) {
test.each(['success', 'timeout'])(`${id} records late results before finalization: %s`, scenario => {
const { dir, facts, exitCode, output } = runRegistration(id, scenario, SYNTHETIC_REPORT);
try {
const retries = retriesForFiles([PAID_FILE]);
expect(retries).toBe(0);
const child = Bun.spawnSync([process.execPath, ...buildPaidShardArgs([script], resolvePaidShardTimeoutMs([PAID_FILE]), 2, retries)], {
cwd: ROOT, timeout: 15_000, stdout: 'pipe', stderr: 'pipe',
env: { ...process.env, EVALS: '', EVALS_ALL: '', TMPDIR: dir, TMP: dir, TEMP: dir },
});
const output = child.stdout.toString() + child.stderr.toString();
expect(child.exitCode, output).toBe(scenario === 'success' ? 0 : 1);
expect(exitCode, output).toBe(scenario === 'success' ? 0 : 1);
expect(output).not.toContain('Unhandled error between tests');
const events = fs.readFileSync(facts, 'utf8').trim().split('\n').map(line => JSON.parse(line));
const starts = events.filter(event => event.kind === 'start');
@@ -88,3 +99,12 @@ await import(path.join(root, ${JSON.stringify(PAID_FILE)}));
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
}
test.each(DESIGN_CAPTURES.reports.map(capture => [`${capture.run} ${capture.trial}`, capture] as const))(
'review-design-lite credits detector rows in captured report %s only when the scan ran', (_label, capture) => {
const { dir, exitCode, output, violations } = runRegistration('review-design-lite', 'success', capture.report);
try {
expect(exitCode, output).toBe(capture.scanRan ? 0 : 1);
expect(violations.includes('the review omitted the mechanical detector rows'), output).toBe(!capture.scanRan);
} finally { fs.rmSync(dir, { recursive: true, force: true }); }
});
+4 -3
View File
@@ -13,7 +13,7 @@ import { spawnSync } from 'child_process';
import * as fs from 'fs';
import * as path from 'path';
import * as os from 'os';
import { installFakeImpeccable } from './helpers/fake-impeccable';
import { carriesDetectorRows, installFakeImpeccable } from './helpers/fake-impeccable';
const evalCollector = createEvalCollector('e2e-review');
// Capture cleanup and recording must finish before Bun starts its retry.
@@ -287,8 +287,9 @@ Important: The design checklist should catch issues like blacklisted fonts, smal
if (review.includes('welcome to') || review.includes('all-in-one') || review.includes('generic') || review.includes('hero copy') || review.includes('ai slop')) detected++;
// Issue 7: 3-column feature grid — LOW
if (review.includes('3-column') || review.includes('three-column') || review.includes('feature grid') || review.includes('icon') || review.includes('circle')) detected++;
// Signal 8: the mechanical pass (fake impeccable engine via IMPECCABLE_BIN) surfaced a detector row
const detectorSeen = review.includes('detector') || review.includes('[ai-color-palette]') || review.includes('[low-contrast]') || review.includes('impeccable');
// Signal 8: the mechanical pass (fake impeccable engine via IMPECCABLE_BIN) surfaced a detector row.
// Only rule ids the checklist never names count; claiming the detector is absent is not a row.
const detectorSeen = carriesDetectorRows(review, fs.readFileSync(path.join(designDir, 'review-design-checklist.md'), 'utf-8'));
console.log(`Design review detected ${detected}/7 planted checklist signals; detector rows surfaced: ${detectorSeen}`);
expect(detected).toBeGreaterThanOrEqual(4); // the LLM-checklist bar, unchanged by the detector