Merge remote-tracking branch 'origin/capy/fixwave-baseline-repairs' into capy/audit-fix-wave

This commit is contained in:
garrytan committed 2026-09-29 16:43:05 +00:00
commit bfd1783fbf
21 files changed
+107 -61

No files matched your search

+3 -2
View File
@@ -612,7 +612,7 @@ sections. Read a section in full before doing its step; do not work from memory.
| When | Read this section |
|------|-------------------|
| running the startup-mode diagnostic (Phase 2A: operating principles, pushback patterns, and the six forcing questions) | `sections/phase-2a-startup-diagnostic.md` |
| running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions) | `sections/phase-2b-builder-brainstorm.md` |
| giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions) | `sections/phase-2b-builder-brainstorm.md` |
| writing the design doc and running the tiered relationship handoff (Phases 5-6, after the conversation and alternatives are done) | `sections/design-and-handoff.md` |
---
@@ -628,8 +628,9 @@ Use this mode when the user is building a startup or doing intrapreneurship.
## Phase 2B: Builder Mode — Design Partner
Use this mode when the user is building for fun, learning, hacking on open source, at a hackathon, or doing research.
The section below applies to every builder-mode reply, including a direct request for ideas or unlocks that skips the generative questions.
> **STOP.** Before running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions), Read `~/.claude/skills/gstack/office-hours/sections/phase-2b-builder-brainstorm.md` and execute it
> **STOP.** Before giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions), Read `~/.claude/skills/gstack/office-hours/sections/phase-2b-builder-brainstorm.md` and execute it
> in full. Do not work from memory — that section is the source of truth for this step.
**If the vibe shifts mid-session** — the user starts in builder mode but says "actually I think this could be a real company" or mentions customers, revenue, fundraising — upgrade to Startup mode naturally. Say something like: "Okay, now we're talking — let me ask you some harder questions." Then switch to the Phase 2A questions.
+1
View File
@@ -135,6 +135,7 @@ Use this mode when the user is building a startup or doing intrapreneurship.
## Phase 2B: Builder Mode — Design Partner
Use this mode when the user is building for fun, learning, hacking on open source, at a hackathon, or doing research.
The section below applies to every builder-mode reply, including a direct request for ideas or unlocks that skips the generative questions.
{{SECTION:phase-2b-builder-brainstorm}}
+1 -1
View File
@@ -14,7 +14,7 @@
"id": "phase-2b-builder-brainstorm",
"file": "phase-2b-builder-brainstorm.md",
"title": "Phase 2B builder-mode brainstorm",
"trigger": "running the builder-mode brainstorm (Phase 2B: operating principles, the wild exemplar, and the generative questions)"
"trigger": "giving any builder-mode response (Phase 2B: brainstorm questions and every suggestion, adjacent unlock or riff; holds the operating principles, the wild exemplar, the response posture and the generative questions)"
},
{
"id": "design-and-handoff",
+2 -2
View File
@@ -24,7 +24,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co
For /review and /ship, no plan/server is required.
Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces).
Explicit plan checks remain required beyond this smoke budget.
Explicit plan checks and revalidation remain required beyond this smoke budget.
For /qa and /qa-only:
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
- Functional Full, Quick and Regression have no default total timer.
@@ -66,7 +66,7 @@ Never batch probes.
4. Replay the exact failing command/request from the same initial fixture state via steps 2–3 (same native command, fresh capture ID)
before repair, then minimize via those gates. Expiry leaves confirmation/minimization incomplete.
Another input or a regression test is not that replay.
5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.
5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.
## 3. Parent handoff
+8 -6
View File
@@ -823,9 +823,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
**3. Run smoke and plan checks.**
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline. /review sets none; only an invoker-supplied EARLIER_UTC counts.
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
**4. Check freshness before reporting.**
@@ -843,7 +843,8 @@ Return verified defects to Fix-First: `path`, `line`, `category`,
`fingerprint: path:line:category`, replay, `test_stub`. Use checklist severity;
unmatched functional failures are `functional-contract`, `CRITICAL`.
Setup/permission blockers are not defects. Test creation needs user approval.
Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.
Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval.
After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.
**5. Prepare one provisional QA section.**
Read QA's `templates/functional-report-template.md`. Title it
@@ -1066,8 +1067,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap
- Use Step 4.6's `specialists` object unchanged, including its empty small-diff map.
If this host omits Review Army, use `specialists: {}` without claiming specialist coverage.
- Build `findings` from final-pass core, specialist, verified exploratory QA
findings and invocation actions. Retain `fingerprint`, `severity`
- Build `findings` from the final-pass findings Step 5 combined (core, specialist,
Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA
findings) and invocation actions. Retain `fingerprint`, `severity`
(`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`,
`helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`,
never supplied/model hashes.
+3 -2
View File
@@ -386,8 +386,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap
- Use Step 4.6's `specialists` object unchanged, including its empty small-diff map.
If this host omits Review Army, use `specialists: {}` without claiming specialist coverage.
- Build `findings` from final-pass core, specialist, verified exploratory QA
findings and invocation actions. Retain `fingerprint`, `severity`
- Build `findings` from the final-pass findings Step 5 combined (core, specialist,
Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA
findings) and invocation actions. Retain `fingerprint`, `severity`
(`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`,
`helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`,
never supplied/model hashes.
+6 -4
View File
@@ -27,8 +27,8 @@ done
3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check.
**Error handling:**
- No plan file found → skip with "No plan file detected — skipping."
- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping."
- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below.
- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified.
### Actionable Item Extraction
@@ -192,13 +192,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla
- **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report.
- **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection.
- **HIGH-impact discrepancies** trigger AskUserQuestion:
- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion:
- Show the investigation findings
- Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped
- A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review.
- B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification.
This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion).
This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion).
Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger
this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements.
When continuing after the audit (no HIGH-impact gate, or option B/C), emit the
single final Scope Check using Step 1.5's provisional notes and this plan context:
+7 -6
View File
@@ -71,7 +71,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co
${reportOnly ? '' : `For /review and /ship, no plan/server is required.
Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces).
Explicit plan checks remain required beyond this smoke budget.`}
Explicit plan checks and revalidation remain required beyond this smoke budget.`}
For /qa and /qa-only:
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
- Functional Full, Quick and Regression have no default total timer.
@@ -126,7 +126,7 @@ ${reportOnly ? ` For guarded text, copy the complete span between the guard's
Another input or a regression test is not that replay.
${reportOnly ? `5. If the user or another process changes source, commands or fixtures, review the affected
contracts and return to step 2 for each affected revalidation. Do not make product changes yourself.
Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`}
Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`}
## 3. Parent handoff
@@ -255,9 +255,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
**3. Run smoke and plan checks.**
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.${ship ? '' : ' /review sets none; only an invoker-supplied EARLIER_UTC counts.'}
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
**4. Check freshness before reporting.**
@@ -275,7 +275,8 @@ Return verified defects to Fix-First: \`path\`, \`line\`, \`category\`,
\`fingerprint: path:line:category\`, replay, \`test_stub\`. Use checklist severity;
unmatched functional failures are \`functional-contract\`, \`CRITICAL\`.
Setup/permission blockers are not defects. Test creation needs user approval.
${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : 'Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.'}
${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : `Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval.
After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.`}
${ship ? `Read QA's \`templates/functional-report-template.md\`: PR section \`## Exploratory QA\`,
fields as subsections. Link every checkpoint; no second report. Separate browser results;
+7 -4
View File
@@ -1483,8 +1483,9 @@ done
3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check.
**Error handling:**
- No plan file found → skip with "No plan file detected — skipping."
${ship ? '- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.' : '- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping."'}`;
${ship ? `- No plan file found → skip with "No plan file detected — skipping."
- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.` : `- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below.
- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified.`}`;
}
// ─── Plan Completion Audit ────────────────────────────────────────────
@@ -1723,13 +1724,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla
- **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report.
- **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection.
- **HIGH-impact discrepancies** trigger AskUserQuestion:
- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion:
- Show the investigation findings
- Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped
- A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review.
- B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification.
This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion).
This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion).
Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger
this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements.
When continuing after the audit (no HIGH-impact gate, or option B/C), emit the
single final Scope Check using Step 1.5's provisional notes and this plan context:
+2 -2
View File
@@ -502,8 +502,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
**3. Run smoke and plan checks.**
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
+10 -5
View File
@@ -672,6 +672,10 @@ Capability check (per /plan-eng-review §6):
bun run ~/.claude/skills/gstack/bin/gstack-gbrain-read-capability.ts <user-args>
```
`<user-args>` are the same flags this /sync-gbrain invocation passed to Step 2,
unchanged (empty for a plain run). The helper needs no other input: run it once
and use its JSON result; do not inspect its source or the gbrain CLI first.
The helper reports JSON `status: ready` only after the successful code sync's
source and real worktree match `.gbrain-source`, the source registration points
to that worktree, and a bounded, source-scoped list/get returns the same page.
@@ -746,16 +750,17 @@ sync code walk for them requires an explicit `--allow-reclone` opt-in.
<!-- gstack-gbrain-search-guidance:end -->
```
Use the Read + Edit tools. The find-and-replace target is the entire region
from `<!-- gstack-gbrain-search-guidance:start -->` through
Read CLAUDE.md once and compute its new content. The replacement target is
the entire region from `<!-- gstack-gbrain-search-guidance:start -->` through
`<!-- gstack-gbrain-search-guidance:end -->`. If those markers are missing,
search for `## GBrain Search Guidance (configured by /sync-gbrain)` heading
and replace from there to the next `## ` or EOF. If no heading exists, append
the entire block at the end of CLAUDE.md.
**Atomic write:** write the new CLAUDE.md content to a tmp file alongside it
(e.g., `CLAUDE.md.sync-gbrain.tmp`) then `mv` to atomic-rename, so a crash
mid-write never leaves the file half-modified.
**Atomic write (the only write path; do not Edit CLAUDE.md in place):** Write
the complete new content to `CLAUDE.md.sync-gbrain.tmp` beside it, then `mv` it
over CLAUDE.md, so a crash mid-write never leaves the file half-modified. Verify
the block count in the same Bash call as the `mv`, then go to Step 5.
**If `status=unknown`** — preserve the existing guidance block, if any, and
report the helper's reason as WARN with advice to retry `/sync-gbrain` or the
+10 -5
View File
@@ -323,6 +323,10 @@ Capability check (per /plan-eng-review §6):
bun run ~/.claude/skills/gstack/bin/gstack-gbrain-read-capability.ts <user-args>
```
`<user-args>` are the same flags this /sync-gbrain invocation passed to Step 2,
unchanged (empty for a plain run). The helper needs no other input: run it once
and use its JSON result; do not inspect its source or the gbrain CLI first.
The helper reports JSON `status: ready` only after the successful code sync's
source and real worktree match `.gbrain-source`, the source registration points
to that worktree, and a bounded, source-scoped list/get returns the same page.
@@ -397,16 +401,17 @@ sync code walk for them requires an explicit `--allow-reclone` opt-in.
<!-- gstack-gbrain-search-guidance:end -->
```
Use the Read + Edit tools. The find-and-replace target is the entire region
from `<!-- gstack-gbrain-search-guidance:start -->` through
Read CLAUDE.md once and compute its new content. The replacement target is
the entire region from `<!-- gstack-gbrain-search-guidance:start -->` through
`<!-- gstack-gbrain-search-guidance:end -->`. If those markers are missing,
search for `## GBrain Search Guidance (configured by /sync-gbrain)` heading
and replace from there to the next `## ` or EOF. If no heading exists, append
the entire block at the end of CLAUDE.md.
**Atomic write:** write the new CLAUDE.md content to a tmp file alongside it
(e.g., `CLAUDE.md.sync-gbrain.tmp`) then `mv` to atomic-rename, so a crash
mid-write never leaves the file half-modified.
**Atomic write (the only write path; do not Edit CLAUDE.md in place):** Write
the complete new content to `CLAUDE.md.sync-gbrain.tmp` beside it, then `mv` it
over CLAUDE.md, so a crash mid-write never leaves the file half-modified. Verify
the block count in the same Bash call as the `mv`, then go to Step 5.
**If `status=unknown`** — preserve the existing guidance block, if any, and
report the helper's reason as WARN with advice to retry `/sync-gbrain` or the
+2 -2
View File
@@ -1996,8 +1996,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
**3. Run smoke and plan checks.**
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
+2 -2
View File
@@ -2255,8 +2255,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
**3. Run smoke and plan checks.**
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
@@ -6,10 +6,11 @@ import { spawnSync } from 'node:child_process';
const root = path.resolve(import.meta.dir, '../..');
export function createReadinessFixture(kind: 'ready' | 'unknown') {
const workDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-'));
const home = path.join(workDir, '.fixture-home');
const bin = path.join(workDir, '.fixture-bin');
fs.mkdirSync(home); fs.mkdirSync(bin);
const base = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-ready-'));
const workDir = path.join(base, 'repo');
const home = path.join(base, 'home');
const bin = path.join(base, 'bin');
fs.mkdirSync(workDir); fs.mkdirSync(home); fs.mkdirSync(bin);
const init = spawnSync('git', ['init', '--quiet'], { cwd: workDir, timeout: 10_000 });
if (init.status !== 0) throw new Error('readiness fixture git init failed');
fs.writeFileSync(path.join(workDir, '.gbrain-source'), 'client-fixture\n');
@@ -57,6 +58,6 @@ else { console.error('unsupported operation'); process.exit(3); }
sourceIntact: () => fs.readFileSync(path.join(workDir, '.gbrain-source'), 'utf8') === pin
&& fs.readFileSync(path.join(stateDir, '.gbrain-sync-state.json'), 'utf8') === state
&& !fs.existsSync(path.join(workDir, 'code')),
cleanup: () => fs.rmSync(workDir, { recursive: true, force: true }),
cleanup: () => fs.rmSync(base, { recursive: true, force: true }),
};
}
+1 -1
View File
@@ -149,7 +149,7 @@ describe('compact QA browser recipes retain native operations', () => {
]) expect(loop).toContain(contract);
for (const field of ['observationCommand', 'observed', 'hypothesis', 'nextCommand']) expect(loop).toContain(`${field}:`);
expect(loop).toContain('Functional Full, Quick and Regression have no default total timer');
if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks remain required beyond this smoke budget');
if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
}
});
+1 -1
View File
@@ -70,7 +70,7 @@ describe('review and ship completion freshness contracts', () => {
expect(gate).toContain('Reporting reserves cannot stop required revalidation within the caller\'s deadline');
expect(body).toContain('Await clock/guard results before acting');
expect(body).toContain('Smoke: 5 minutes/12 probes');
expect(body).toContain('Then run required plan checks, even after smoke expires');
expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires');
expect(body).toContain('no smoke guard; never reset the clock');
expect(body).toContain('Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline');
});
+3 -3
View File
@@ -608,8 +608,8 @@ describe('generated actual parent paths', () => {
expect(load).toContain('Templates cannot replace them');
const flat = parent.replace(/\s+/g, ' ');
expect(flat).toContain('Only the parent runs report-only discovery');
expect(flat).toContain('Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit');
expect(flat).toContain('Then run required plan checks, even after smoke expires');
expect(flat).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit');
expect(flat).toContain('Then run required plan checks and revalidation, even after smoke expires');
expect(flat).toContain('using the same procedure but no smoke guard; never reset the clock');
expect(flat).toContain("Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline");
expect(flat).toContain('When the caller\'s deadline expires, mark unfinished checks not-run');
@@ -640,7 +640,7 @@ describe('generated actual parent paths', () => {
expect(body).toContain('Stop after 5 minutes or 12 probes, whichever comes first');
expect(body).toContain('G enforces the deadline');
expect(body).toContain('Never reset D/bypass G');
expect(body).toContain('Explicit plan checks remain required beyond this smoke budget');
expect(body).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
expect(body).toContain('leaves /review incomplete');
expect(body).toContain('/ship blocked unless the user explicitly accepts that named risk');
});
+6 -4
View File
@@ -40,8 +40,8 @@ function assertBoundsAndLayout(text: string) {
function assertPlanExecution(text: string, shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude })) {
const step = text.slice(text.indexOf('**3. Run smoke and plan checks.**'), text.indexOf('**4. Check freshness before reporting.**')).replace(/\s+/g, ' ');
for (const contract of [
'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit',
'Then run required plan checks, even after smoke expires',
'Follow the shared Probe loop for smoke checks and replays until the smoke limit',
'Then run required plan checks and revalidation, even after smoke expires',
'using the same procedure but no smoke guard; never reset the clock',
"Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline",
'When the caller\'s deadline expires, mark unfinished checks not-run',
@@ -261,17 +261,19 @@ describe('QA probe entry and checkpoint gates', () => {
const text = generateQAReview({ host: 'claude', skillName, tmplPath: '', paths: HOST_PATHS.claude }).replace(/\s+/g, ' ');
assertPlanExecution(text);
for (const [before, after] of [
['Then run required plan checks, even after smoke expires', 'Skip plan checks when smoke expired'],
['Then run required plan checks and revalidation, even after smoke expires', 'Skip plan checks when smoke expired'],
['no smoke guard; never reset the clock', 'restart and use the smoke guard'],
['same procedure', 'Start a new checkpoint sequence'],
['at the caller\'s remaining time', 'with no caller cap'],
['When the caller\'s deadline expires, mark unfinished checks not-run', 'If that deadline expired, mark the check passed'],
['Await clock/guard results before acting', 'Ignore clock results'],
['plan checks and revalidation, even', 'plan checks, even'],
['smoke checks and replays until', 'smoke checks, replays and revalidation until'],
]) {
expect(text).toContain(before);
expect(() => assertPlanExecution(text.replace(before, after))).toThrow();
}
const smoke = 'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.';
const smoke = 'Follow the shared Probe loop for smoke checks and replays until the smoke limit.';
expect(() => assertPlanExecution(text.replace(smoke, '').replace('**4. Check', smoke + '\n**4. Check'))).toThrow();
const shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude });
for (const contract of ['First demonstrate success: output AND durable effects',
+25 -3
View File
@@ -32,7 +32,7 @@ test('review audits deliverables before deferring behavioral plan checks to the
expect(audit.indexOf('Inspect the validator and its hooks')).toBeLessThan(audit.indexOf('If found and verified safe above, invoke it'));
expect(audit).not.toContain('For each extracted plan item, run the verification dispatch');
const qa = generateQAReview(ctx);
expect(qa).toContain('Then run required plan checks, even after smoke expires');
expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires');
expect(qa).toContain('Report clean/completed only when all required checks pass on current inputs');
});
@@ -249,7 +249,7 @@ test('caller QA defines execution, evidence ownership and report adaptation befo
for (const contract of [
'Only the parent runs report-only discovery',
'Never overwrite another run',
'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit',
'Follow the shared Probe loop for smoke checks and replays until the smoke limit',
'using the same procedure but no smoke guard; never reset the clock',
'Read agent/user updates and await results without batching them with reporting/logging',
'Compare each probe\'s recorded source, tests, contracts, commands and fixtures (or input fingerprint) with current inputs, even without updates',
@@ -378,7 +378,7 @@ test('review finalization ownership: the plan audit retains its high-impact gate
expect(plan).not.toContain('never blocks the review');
expect(plan).toContain('{{PLAN_COMPLETION_AUDIT_REVIEW}}');
const audit = readFileSync(join(root, 'review/sections/plan-completion.md'), 'utf8');
const gate = audit.indexOf('**HIGH-impact discrepancies** trigger AskUserQuestion');
const gate = audit.indexOf('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion');
expect(gate).toBeGreaterThan(-1);
expect(gate).toBeLessThan(audit.indexOf('When continuing after the audit (no HIGH-impact gate, or option B/C)'));
expect(audit).toContain('then it gates via AskUserQuestion');
@@ -564,3 +564,25 @@ test('ship review clarity: parent settlement gate precedes classification and ca
expect(gate).toContain('After terminal failure, independent evidence may support fixes');
expect(gate).toContain('missing dispatched output still blocks continuation, even with a QA exception');
});
test('review resolves the judged smoke-clock, setup-authority, plan-gate and findings-source ambiguities', () => {
const ctx = { skillName: 'review', tmplPath: 'review/SKILL.md.tmpl', host: 'claude', paths: HOST_PATHS.claude } as TemplateContext;
const qa = generateQAReview(ctx).replace(/\s+/g, ' ');
expect(qa).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit');
expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard');
expect(qa).toContain('/review sets none; only an invoker-supplied EARLIER_UTC counts');
expect(qa).toContain('Report-only /review never runs setup, installs or cookie import, even after approval');
expect(qa).toContain('After a grant, recheck readiness and run affected checks; otherwise they stay blocked');
expect(qa).not.toContain('Ask for setup/permission');
expect(generateQAReview({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).not.toContain('/review sets none');
const shared = generateQAExploratory({ ...ctx, skillName: 'qa' }).replace(/\s+/g, ' ');
expect(shared).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
const audit = generatePlanCompletionAuditReview(ctx).replace(/\s+/g, ' ');
expect(audit).toContain('No plan file found → say "No plan file detected." and use the Fallback Intent Sources below');
expect(audit).not.toContain('skip with "No plan file detected — skipping."');
expect(audit).toContain('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion');
expect(audit).toContain('Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger this question');
expect(generatePlanCompletionAuditShip({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).toContain('skip with "No plan file detected — skipping."');
const persist = skill.slice(skill.indexOf('### 2. Fill the record')).replace(/\s+/g, ' ');
expect(persist).toContain('findings Step 5 combined (core, specialist, Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA findings)');
});
+1 -1
View File
@@ -14,7 +14,7 @@ test('Ship initializes and applies its smoke guard independently of required pla
expect(body).toContain('Run the shared preflight; start its smoke guard once. Guard every smoke probe.');
expect(body.indexOf('start its smoke guard once')).toBeLessThan(body.indexOf('**3. Run smoke and plan checks.**'));
expect(body).toContain('Required even for small diffs or missing plans/servers');
expect(body).toContain('Then run required plan checks, even after smoke expires');
expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires');
expect(body).toContain('using the same procedure but no smoke guard; never reset the clock');
}
});