diff --git a/office-hours/SKILL.md b/office-hours/SKILL.md index f966d7c63..550dc195a 100644 --- a/office-hours/SKILL.md +++ b/office-hours/SKILL.md @@ -612,7 +612,7 @@ sections. Read a section in full before doing its step; do not work from memory. | When | Read this section | |------|-------------------| | running the startup-mode diagnostic (Phase 2A: operating principles, pushback patterns, and the six forcing questions) | `sections/phase-2a-startup-diagnostic.md` | -| running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions) | `sections/phase-2b-builder-brainstorm.md` | +| giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions) | `sections/phase-2b-builder-brainstorm.md` | | writing the design doc and running the tiered relationship handoff (Phases 5-6, after the conversation and alternatives are done) | `sections/design-and-handoff.md` | --- @@ -628,8 +628,9 @@ Use this mode when the user is building a startup or doing intrapreneurship. ## Phase 2B: Builder Mode — Design Partner Use this mode when the user is building for fun, learning, hacking on open source, at a hackathon, or doing research. +The section below applies to every builder-mode reply, including a direct request for ideas or unlocks that skips the generative questions. -> **STOP.** Before running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions), Read `~/.claude/skills/gstack/office-hours/sections/phase-2b-builder-brainstorm.md` and execute it +> **STOP.** Before giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions), Read `~/.claude/skills/gstack/office-hours/sections/phase-2b-builder-brainstorm.md` and execute it > in full. Do not work from memory — that section is the source of truth for this step. **If the vibe shifts mid-session** — the user starts in builder mode but says "actually I think this could be a real company" or mentions customers, revenue, fundraising — upgrade to Startup mode naturally. Say something like: "Okay, now we're talking — let me ask you some harder questions." Then switch to the Phase 2A questions. diff --git a/office-hours/SKILL.md.tmpl b/office-hours/SKILL.md.tmpl index 898a77b77..9b6e57f03 100644 --- a/office-hours/SKILL.md.tmpl +++ b/office-hours/SKILL.md.tmpl @@ -135,6 +135,7 @@ Use this mode when the user is building a startup or doing intrapreneurship. ## Phase 2B: Builder Mode — Design Partner Use this mode when the user is building for fun, learning, hacking on open source, at a hackathon, or doing research. +The section below applies to every builder-mode reply, including a direct request for ideas or unlocks that skips the generative questions. {{SECTION:phase-2b-builder-brainstorm}} diff --git a/office-hours/sections/manifest.json b/office-hours/sections/manifest.json index fa6df5264..fec89a51f 100644 --- a/office-hours/sections/manifest.json +++ b/office-hours/sections/manifest.json @@ -14,7 +14,7 @@ "id": "phase-2b-builder-brainstorm", "file": "phase-2b-builder-brainstorm.md", "title": "Phase 2B builder-mode brainstorm", - "trigger": "running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions)" + "trigger": "giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions)" }, { "id": "design-and-handoff", diff --git a/qa/sections/exploratory.md b/qa/sections/exploratory.md index fb29af889..961713199 100644 --- a/qa/sections/exploratory.md +++ b/qa/sections/exploratory.md @@ -24,7 +24,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co For /review and /ship, no plan/server is required. Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces). -Explicit plan checks remain required beyond this smoke budget. +Explicit plan checks and revalidation remain required beyond this smoke budget. For /qa and /qa-only: - Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900. - Functional Full, Quick and Regression have no default total timer. @@ -66,7 +66,7 @@ Never batch probes. 4. Replay the exact failing command/request from the same initial fixture state via steps 2–3 (same native command, fresh capture ID) before repair, then minimize via those gates. Expiry leaves confirmation/minimization incomplete. Another input or a regression test is not that replay. -5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence. +5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence. ## 3. Parent handoff diff --git a/review/SKILL.md b/review/SKILL.md index c77b48537..01757a13f 100644 --- a/review/SKILL.md +++ b/review/SKILL.md @@ -823,9 +823,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser - Required: plan commands/assertions, listed separately. Other ideas are optional, untested. **3. Run smoke and plan checks.** -Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit. -Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. -Use finite command timeouts, capped at the caller's remaining time if it has a deadline. +Follow the shared Probe loop for smoke checks and replays until the smoke limit. +Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. +Use finite command timeouts, capped at the caller's remaining time if it has a deadline. /review sets none; only an invoker-supplied EARLIER_UTC counts. Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run. **4. Check freshness before reporting.** @@ -843,7 +843,8 @@ Return verified defects to Fix-First: `path`, `line`, `category`, `fingerprint: path:line:category`, replay, `test_stub`. Use checklist severity; unmatched functional failures are `functional-contract`, `CRITICAL`. Setup/permission blockers are not defects. Test creation needs user approval. -Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it. +Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval. +After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it. **5. Prepare one provisional QA section.** Read QA's `templates/functional-report-template.md`. Title it @@ -1066,8 +1067,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap - Use Step 4.6's `specialists` object unchanged, including its empty small-diff map. If this host omits Review Army, use `specialists: {}` without claiming specialist coverage. -- Build `findings` from final-pass core, specialist, verified exploratory QA - findings and invocation actions. Retain `fingerprint`, `severity` +- Build `findings` from the final-pass findings Step 5 combined (core, specialist, + Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA + findings) and invocation actions. Retain `fingerprint`, `severity` (`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`, `helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`, never supplied/model hashes. diff --git a/review/SKILL.md.tmpl b/review/SKILL.md.tmpl index a9c4f3167..95b5423b2 100644 --- a/review/SKILL.md.tmpl +++ b/review/SKILL.md.tmpl @@ -386,8 +386,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap - Use Step 4.6's `specialists` object unchanged, including its empty small-diff map. If this host omits Review Army, use `specialists: {}` without claiming specialist coverage. -- Build `findings` from final-pass core, specialist, verified exploratory QA - findings and invocation actions. Retain `fingerprint`, `severity` +- Build `findings` from the final-pass findings Step 5 combined (core, specialist, + Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA + findings) and invocation actions. Retain `fingerprint`, `severity` (`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`, `helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`, never supplied/model hashes. diff --git a/review/sections/plan-completion.md b/review/sections/plan-completion.md index d06f7bc75..42ae4ece1 100644 --- a/review/sections/plan-completion.md +++ b/review/sections/plan-completion.md @@ -27,8 +27,8 @@ done 3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check. **Error handling:** -- No plan file found → skip with "No plan file detected — skipping." -- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping." +- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below. +- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified. ### Actionable Item Extraction @@ -192,13 +192,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla - **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report. - **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection. -- **HIGH-impact discrepancies** trigger AskUserQuestion: +- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion: - Show the investigation findings - Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped - A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review. - B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification. -This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion). +This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion). +Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger +this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements. When continuing after the audit (no HIGH-impact gate, or option B/C), emit the single final Scope Check using Step 1.5's provisional notes and this plan context: diff --git a/scripts/resolvers/qa.ts b/scripts/resolvers/qa.ts index ace9b8773..55841180e 100644 --- a/scripts/resolvers/qa.ts +++ b/scripts/resolvers/qa.ts @@ -71,7 +71,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co ${reportOnly ? '' : `For /review and /ship, no plan/server is required. Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces). -Explicit plan checks remain required beyond this smoke budget.`} +Explicit plan checks and revalidation remain required beyond this smoke budget.`} For /qa and /qa-only: - Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900. - Functional Full, Quick and Regression have no default total timer. @@ -126,7 +126,7 @@ ${reportOnly ? ` For guarded text, copy the complete span between the guard's Another input or a regression test is not that replay. ${reportOnly ? `5. If the user or another process changes source, commands or fixtures, review the affected contracts and return to step 2 for each affected revalidation. Do not make product changes yourself. - Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`} + Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`} ## 3. Parent handoff @@ -255,9 +255,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser - Required: plan commands/assertions, listed separately. Other ideas are optional, untested. **3. Run smoke and plan checks.** -Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit. -Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. -Use finite command timeouts, capped at the caller's remaining time if it has a deadline. +Follow the shared Probe loop for smoke checks and replays until the smoke limit. +Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. +Use finite command timeouts, capped at the caller's remaining time if it has a deadline.${ship ? '' : ' /review sets none; only an invoker-supplied EARLIER_UTC counts.'} Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run. **4. Check freshness before reporting.** @@ -275,7 +275,8 @@ Return verified defects to Fix-First: \`path\`, \`line\`, \`category\`, \`fingerprint: path:line:category\`, replay, \`test_stub\`. Use checklist severity; unmatched functional failures are \`functional-contract\`, \`CRITICAL\`. Setup/permission blockers are not defects. Test creation needs user approval. -${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : 'Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.'} +${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : `Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval. +After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.`} ${ship ? `Read QA's \`templates/functional-report-template.md\`: PR section \`## Exploratory QA\`, fields as subsections. Link every checkpoint; no second report. Separate browser results; diff --git a/scripts/resolvers/review.ts b/scripts/resolvers/review.ts index e62db45b8..6034b9eb9 100644 --- a/scripts/resolvers/review.ts +++ b/scripts/resolvers/review.ts @@ -1483,8 +1483,9 @@ done 3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check. **Error handling:** -- No plan file found → skip with "No plan file detected — skipping." -${ship ? '- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.' : '- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping."'}`; +${ship ? `- No plan file found → skip with "No plan file detected — skipping." +- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.` : `- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below. +- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified.`}`; } // ─── Plan Completion Audit ──────────────────────────────────────────── @@ -1723,13 +1724,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla - **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report. - **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection. -- **HIGH-impact discrepancies** trigger AskUserQuestion: +- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion: - Show the investigation findings - Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped - A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review. - B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification. -This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion). +This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion). +Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger +this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements. When continuing after the audit (no HIGH-impact gate, or option B/C), emit the single final Scope Check using Step 1.5's provisional notes and this plan context: diff --git a/ship/sections/review-army.md b/ship/sections/review-army.md index 4e65646bf..03c7f2627 100644 --- a/ship/sections/review-army.md +++ b/ship/sections/review-army.md @@ -502,8 +502,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F - Required: plan commands/assertions, listed separately. Other ideas are optional, untested. **3. Run smoke and plan checks.** -Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit. -Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. +Follow the shared Probe loop for smoke checks and replays until the smoke limit. +Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. Use finite command timeouts, capped at the caller's remaining time if it has a deadline. Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run. diff --git a/sync-gbrain/SKILL.md b/sync-gbrain/SKILL.md index bf60cfae3..a3455f9dc 100644 --- a/sync-gbrain/SKILL.md +++ b/sync-gbrain/SKILL.md @@ -672,6 +672,10 @@ Capability check (per /plan-eng-review §6): bun run ~/.claude/skills/gstack/bin/gstack-gbrain-read-capability.ts ``` +`` are the same flags this /sync-gbrain invocation passed to Step 2, +unchanged (empty for a plain run). The helper needs no other input: run it once +and use its JSON result; do not inspect its source or the gbrain CLI first. + The helper reports JSON `status: ready` only after the successful code sync's source and real worktree match `.gbrain-source`, the source registration points to that worktree, and a bounded, source-scoped list/get returns the same page. @@ -746,16 +750,17 @@ sync code walk for them requires an explicit `--allow-reclone` opt-in. ``` -Use the Read + Edit tools. The find-and-replace target is the entire region -from `` through +Read CLAUDE.md once and compute its new content. The replacement target is +the entire region from `` through ``. If those markers are missing, search for `## GBrain Search Guidance (configured by /sync-gbrain)` heading and replace from there to the next `## ` or EOF. If no heading exists, append the entire block at the end of CLAUDE.md. -**Atomic write:** write the new CLAUDE.md content to a tmp file alongside it -(e.g., `CLAUDE.md.sync-gbrain.tmp`) then `mv` to atomic-rename, so a crash -mid-write never leaves the file half-modified. +**Atomic write (the only write path; do not Edit CLAUDE.md in place):** Write +the complete new content to `CLAUDE.md.sync-gbrain.tmp` beside it, then `mv` it +over CLAUDE.md, so a crash mid-write never leaves the file half-modified. Verify +the block count in the same Bash call as the `mv`, then go to Step 5. **If `status=unknown`** — preserve the existing guidance block, if any, and report the helper's reason as WARN with advice to retry `/sync-gbrain` or the diff --git a/sync-gbrain/SKILL.md.tmpl b/sync-gbrain/SKILL.md.tmpl index fe3bec81c..0398939c6 100644 --- a/sync-gbrain/SKILL.md.tmpl +++ b/sync-gbrain/SKILL.md.tmpl @@ -323,6 +323,10 @@ Capability check (per /plan-eng-review §6): bun run ~/.claude/skills/gstack/bin/gstack-gbrain-read-capability.ts ``` +`` are the same flags this /sync-gbrain invocation passed to Step 2, +unchanged (empty for a plain run). The helper needs no other input: run it once +and use its JSON result; do not inspect its source or the gbrain CLI first. + The helper reports JSON `status: ready` only after the successful code sync's source and real worktree match `.gbrain-source`, the source registration points to that worktree, and a bounded, source-scoped list/get returns the same page. @@ -397,16 +401,17 @@ sync code walk for them requires an explicit `--allow-reclone` opt-in. ``` -Use the Read + Edit tools. The find-and-replace target is the entire region -from `` through +Read CLAUDE.md once and compute its new content. The replacement target is +the entire region from `` through ``. If those markers are missing, search for `## GBrain Search Guidance (configured by /sync-gbrain)` heading and replace from there to the next `## ` or EOF. If no heading exists, append the entire block at the end of CLAUDE.md. -**Atomic write:** write the new CLAUDE.md content to a tmp file alongside it -(e.g., `CLAUDE.md.sync-gbrain.tmp`) then `mv` to atomic-rename, so a crash -mid-write never leaves the file half-modified. +**Atomic write (the only write path; do not Edit CLAUDE.md in place):** Write +the complete new content to `CLAUDE.md.sync-gbrain.tmp` beside it, then `mv` it +over CLAUDE.md, so a crash mid-write never leaves the file half-modified. Verify +the block count in the same Bash call as the `mv`, then go to Step 5. **If `status=unknown`** — preserve the existing guidance block, if any, and report the helper's reason as WARN with advice to retry `/sync-gbrain` or the diff --git a/test/fixtures/golden/codex-ship-SKILL.md b/test/fixtures/golden/codex-ship-SKILL.md index b53ff9102..aae295f39 100644 --- a/test/fixtures/golden/codex-ship-SKILL.md +++ b/test/fixtures/golden/codex-ship-SKILL.md @@ -1996,8 +1996,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F - Required: plan commands/assertions, listed separately. Other ideas are optional, untested. **3. Run smoke and plan checks.** -Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit. -Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. +Follow the shared Probe loop for smoke checks and replays until the smoke limit. +Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. Use finite command timeouts, capped at the caller's remaining time if it has a deadline. Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run. diff --git a/test/fixtures/golden/factory-ship-SKILL.md b/test/fixtures/golden/factory-ship-SKILL.md index b7d48c1a5..104cb00cc 100644 --- a/test/fixtures/golden/factory-ship-SKILL.md +++ b/test/fixtures/golden/factory-ship-SKILL.md @@ -2255,8 +2255,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F - Required: plan commands/assertions, listed separately. Other ideas are optional, untested. **3. Run smoke and plan checks.** -Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit. -Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. +Follow the shared Probe loop for smoke checks and replays until the smoke limit. +Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock. Use finite command timeouts, capped at the caller's remaining time if it has a deadline. Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run. diff --git a/test/helpers/sync-gbrain-readiness-fixture.ts b/test/helpers/sync-gbrain-readiness-fixture.ts index 3ef477ee8..b818b50b8 100644 --- a/test/helpers/sync-gbrain-readiness-fixture.ts +++ b/test/helpers/sync-gbrain-readiness-fixture.ts @@ -6,10 +6,11 @@ import { spawnSync } from 'node:child_process'; const root = path.resolve(import.meta.dir, '../..'); export function createReadinessFixture(kind: 'ready' | 'unknown') { - const workDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-')); - const home = path.join(workDir, '.fixture-home'); - const bin = path.join(workDir, '.fixture-bin'); - fs.mkdirSync(home); fs.mkdirSync(bin); + const base = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-')); + const workDir = path.join(base, 'repo'); + const home = path.join(base, 'home'); + const bin = path.join(base, 'bin'); + fs.mkdirSync(workDir); fs.mkdirSync(home); fs.mkdirSync(bin); const init = spawnSync('git', ['init', '--quiet'], { cwd: workDir, timeout: 10_000 }); if (init.status !== 0) throw new Error('readiness fixture git init failed'); fs.writeFileSync(path.join(workDir, '.gbrain-source'), 'client-fixture\n'); @@ -57,6 +58,6 @@ else { console.error('unsupported operation'); process.exit(3); } sourceIntact: () => fs.readFileSync(path.join(workDir, '.gbrain-source'), 'utf8') === pin && fs.readFileSync(path.join(stateDir, '.gbrain-sync-state.json'), 'utf8') === state && !fs.existsSync(path.join(workDir, 'code')), - cleanup: () => fs.rmSync(workDir, { recursive: true, force: true }), + cleanup: () => fs.rmSync(base, { recursive: true, force: true }), }; } diff --git a/test/qa-browser-preservation.test.ts b/test/qa-browser-preservation.test.ts index 788e66c89..0cd24f262 100644 --- a/test/qa-browser-preservation.test.ts +++ b/test/qa-browser-preservation.test.ts @@ -149,7 +149,7 @@ describe('compact QA browser recipes retain native operations', () => { ]) expect(loop).toContain(contract); for (const field of ['observationCommand', 'observed', 'hypothesis', 'nextCommand']) expect(loop).toContain(`${field}:`); expect(loop).toContain('Functional Full, Quick and Regression have no default total timer'); - if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks remain required beyond this smoke budget'); + if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget'); } }); diff --git a/test/qa-caller-freshness-order.test.ts b/test/qa-caller-freshness-order.test.ts index f85101850..0a5b667b1 100644 --- a/test/qa-caller-freshness-order.test.ts +++ b/test/qa-caller-freshness-order.test.ts @@ -70,7 +70,7 @@ describe('review and ship completion freshness contracts', () => { expect(gate).toContain('Reporting reserves cannot stop required revalidation within the caller\'s deadline'); expect(body).toContain('Await clock/guard results before acting'); expect(body).toContain('Smoke: 5 minutes/12 probes'); - expect(body).toContain('Then run required plan checks, even after smoke expires'); + expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires'); expect(body).toContain('no smoke guard; never reset the clock'); expect(body).toContain('Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline'); }); diff --git a/test/qa-exploratory-callers.test.ts b/test/qa-exploratory-callers.test.ts index 564bf0c88..592b5dd05 100644 --- a/test/qa-exploratory-callers.test.ts +++ b/test/qa-exploratory-callers.test.ts @@ -608,8 +608,8 @@ describe('generated actual parent paths', () => { expect(load).toContain('Templates cannot replace them'); const flat = parent.replace(/\s+/g, ' '); expect(flat).toContain('Only the parent runs report-only discovery'); - expect(flat).toContain('Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit'); - expect(flat).toContain('Then run required plan checks, even after smoke expires'); + expect(flat).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit'); + expect(flat).toContain('Then run required plan checks and revalidation, even after smoke expires'); expect(flat).toContain('using the same procedure but no smoke guard; never reset the clock'); expect(flat).toContain("Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline"); expect(flat).toContain('When the caller\'s deadline expires, mark unfinished checks not-run'); @@ -640,7 +640,7 @@ describe('generated actual parent paths', () => { expect(body).toContain('Stop after 5 minutes or 12 probes, whichever comes first'); expect(body).toContain('G enforces the deadline'); expect(body).toContain('Never reset D/bypass G'); - expect(body).toContain('Explicit plan checks remain required beyond this smoke budget'); + expect(body).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget'); expect(body).toContain('leaves /review incomplete'); expect(body).toContain('/ship blocked unless the user explicitly accepts that named risk'); }); diff --git a/test/qa-probe-gates.test.ts b/test/qa-probe-gates.test.ts index bd2f88869..c959a920e 100644 --- a/test/qa-probe-gates.test.ts +++ b/test/qa-probe-gates.test.ts @@ -40,8 +40,8 @@ function assertBoundsAndLayout(text: string) { function assertPlanExecution(text: string, shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude })) { const step = text.slice(text.indexOf('**3. Run smoke and plan checks.**'), text.indexOf('**4. Check freshness before reporting.**')).replace(/\s+/g, ' '); for (const contract of [ - 'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit', - 'Then run required plan checks, even after smoke expires', + 'Follow the shared Probe loop for smoke checks and replays until the smoke limit', + 'Then run required plan checks and revalidation, even after smoke expires', 'using the same procedure but no smoke guard; never reset the clock', "Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline", 'When the caller\'s deadline expires, mark unfinished checks not-run', @@ -261,17 +261,19 @@ describe('QA probe entry and checkpoint gates', () => { const text = generateQAReview({ host: 'claude', skillName, tmplPath: '', paths: HOST_PATHS.claude }).replace(/\s+/g, ' '); assertPlanExecution(text); for (const [before, after] of [ - ['Then run required plan checks, even after smoke expires', 'Skip plan checks when smoke expired'], + ['Then run required plan checks and revalidation, even after smoke expires', 'Skip plan checks when smoke expired'], ['no smoke guard; never reset the clock', 'restart and use the smoke guard'], ['same procedure', 'Start a new checkpoint sequence'], ['at the caller\'s remaining time', 'with no caller cap'], ['When the caller\'s deadline expires, mark unfinished checks not-run', 'If that deadline expired, mark the check passed'], ['Await clock/guard results before acting', 'Ignore clock results'], + ['plan checks and revalidation, even', 'plan checks, even'], + ['smoke checks and replays until', 'smoke checks, replays and revalidation until'], ]) { expect(text).toContain(before); expect(() => assertPlanExecution(text.replace(before, after))).toThrow(); } - const smoke = 'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.'; + const smoke = 'Follow the shared Probe loop for smoke checks and replays until the smoke limit.'; expect(() => assertPlanExecution(text.replace(smoke, '').replace('**4. Check', smoke + '\n**4. Check'))).toThrow(); const shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude }); for (const contract of ['First demonstrate success: output AND durable effects', diff --git a/test/review-workflow-clarity.test.ts b/test/review-workflow-clarity.test.ts index fac2da981..8d4bf7e8c 100644 --- a/test/review-workflow-clarity.test.ts +++ b/test/review-workflow-clarity.test.ts @@ -32,7 +32,7 @@ test('review audits deliverables before deferring behavioral plan checks to the expect(audit.indexOf('Inspect the validator and its hooks')).toBeLessThan(audit.indexOf('If found and verified safe above, invoke it')); expect(audit).not.toContain('For each extracted plan item, run the verification dispatch'); const qa = generateQAReview(ctx); - expect(qa).toContain('Then run required plan checks, even after smoke expires'); + expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires'); expect(qa).toContain('Report clean/completed only when all required checks pass on current inputs'); }); @@ -249,7 +249,7 @@ test('caller QA defines execution, evidence ownership and report adaptation befo for (const contract of [ 'Only the parent runs report-only discovery', 'Never overwrite another run', - 'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit', + 'Follow the shared Probe loop for smoke checks and replays until the smoke limit', 'using the same procedure but no smoke guard; never reset the clock', 'Read agent/user updates and await results without batching them with reporting/logging', 'Compare each probe\'s recorded source, tests, contracts, commands and fixtures (or input fingerprint) with current inputs, even without updates', @@ -378,7 +378,7 @@ test('review finalization ownership: the plan audit retains its high-impact gate expect(plan).not.toContain('never blocks the review'); expect(plan).toContain('{{PLAN_COMPLETION_AUDIT_REVIEW}}'); const audit = readFileSync(join(root, 'review/sections/plan-completion.md'), 'utf8'); - const gate = audit.indexOf('**HIGH-impact discrepancies** trigger AskUserQuestion'); + const gate = audit.indexOf('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion'); expect(gate).toBeGreaterThan(-1); expect(gate).toBeLessThan(audit.indexOf('When continuing after the audit (no HIGH-impact gate, or option B/C)')); expect(audit).toContain('then it gates via AskUserQuestion'); @@ -564,3 +564,25 @@ test('ship review clarity: parent settlement gate precedes classification and ca expect(gate).toContain('After terminal failure, independent evidence may support fixes'); expect(gate).toContain('missing dispatched output still blocks continuation, even with a QA exception'); }); + +test('review resolves the judged smoke-clock, setup-authority, plan-gate and findings-source ambiguities', () => { + const ctx = { skillName: 'review', tmplPath: 'review/SKILL.md.tmpl', host: 'claude', paths: HOST_PATHS.claude } as TemplateContext; + const qa = generateQAReview(ctx).replace(/\s+/g, ' '); + expect(qa).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit'); + expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard'); + expect(qa).toContain('/review sets none; only an invoker-supplied EARLIER_UTC counts'); + expect(qa).toContain('Report-only /review never runs setup, installs or cookie import, even after approval'); + expect(qa).toContain('After a grant, recheck readiness and run affected checks; otherwise they stay blocked'); + expect(qa).not.toContain('Ask for setup/permission'); + expect(generateQAReview({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).not.toContain('/review sets none'); + const shared = generateQAExploratory({ ...ctx, skillName: 'qa' }).replace(/\s+/g, ' '); + expect(shared).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget'); + const audit = generatePlanCompletionAuditReview(ctx).replace(/\s+/g, ' '); + expect(audit).toContain('No plan file found → say "No plan file detected." and use the Fallback Intent Sources below'); + expect(audit).not.toContain('skip with "No plan file detected — skipping."'); + expect(audit).toContain('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion'); + expect(audit).toContain('Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger this question'); + expect(generatePlanCompletionAuditShip({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).toContain('skip with "No plan file detected — skipping."'); + const persist = skill.slice(skill.indexOf('### 2. Fill the record')).replace(/\s+/g, ' '); + expect(persist).toContain('findings Step 5 combined (core, specialist, Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA findings)'); +}); diff --git a/test/ship-workflow-clarity.test.ts b/test/ship-workflow-clarity.test.ts index 3b225c71c..8a20a9dd5 100644 --- a/test/ship-workflow-clarity.test.ts +++ b/test/ship-workflow-clarity.test.ts @@ -14,7 +14,7 @@ test('Ship initializes and applies its smoke guard independently of required pla expect(body).toContain('Run the shared preflight; start its smoke guard once. Guard every smoke probe.'); expect(body.indexOf('start its smoke guard once')).toBeLessThan(body.indexOf('**3. Run smoke and plan checks.**')); expect(body).toContain('Required even for small diffs or missing plans/servers'); - expect(body).toContain('Then run required plan checks, even after smoke expires'); + expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires'); expect(body).toContain('using the same procedure but no smoke guard; never reset the clock'); } });