mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
fix(review): resolve the judged revalidation, setup-authority, plan-gate and findings-record ambiguities
The census review workflow judge scored clarity/actionability 3 on both attempts: smoke-clock limits appeared to forbid post-repair revalidation, the caller deadline was undefined, 'ask for setup' conflicted with the report-only browser rule, fallback-sourced HIGH discrepancies had no gate decision, and the Step 5.8 record omitted adversarial findings.
This commit is contained in:
1 parent
05ffcaf9e2
commit
175b12933d
15 files changed
+76
-43
No files matched your search
@@ -24,7 +24,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co
|
||||
|
||||
For /review and /ship, no plan/server is required.
|
||||
Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces).
|
||||
Explicit plan checks remain required beyond this smoke budget.
|
||||
Explicit plan checks and revalidation remain required beyond this smoke budget.
|
||||
For /qa and /qa-only:
|
||||
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
|
||||
- Functional Full, Quick and Regression have no default total timer.
|
||||
@@ -66,7 +66,7 @@ Never batch probes.
|
||||
4. Replay the exact failing command/request from the same initial fixture state via steps 2–3 (same native command, fresh capture ID)
|
||||
before repair, then minimize via those gates. Expiry leaves confirmation/minimization incomplete.
|
||||
Another input or a regression test is not that replay.
|
||||
5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.
|
||||
5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.
|
||||
|
||||
## 3. Parent handoff
|
||||
|
||||
|
||||
+8
-6
@@ -823,9 +823,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser
|
||||
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
|
||||
|
||||
**3. Run smoke and plan checks.**
|
||||
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
|
||||
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
|
||||
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
|
||||
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline. /review sets none; only an invoker-supplied EARLIER_UTC counts.
|
||||
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
|
||||
|
||||
**4. Check freshness before reporting.**
|
||||
@@ -843,7 +843,8 @@ Return verified defects to Fix-First: `path`, `line`, `category`,
|
||||
`fingerprint: path:line:category`, replay, `test_stub`. Use checklist severity;
|
||||
unmatched functional failures are `functional-contract`, `CRITICAL`.
|
||||
Setup/permission blockers are not defects. Test creation needs user approval.
|
||||
Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.
|
||||
Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval.
|
||||
After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.
|
||||
|
||||
**5. Prepare one provisional QA section.**
|
||||
Read QA's `templates/functional-report-template.md`. Title it
|
||||
@@ -1066,8 +1067,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap
|
||||
|
||||
- Use Step 4.6's `specialists` object unchanged, including its empty small-diff map.
|
||||
If this host omits Review Army, use `specialists: {}` without claiming specialist coverage.
|
||||
- Build `findings` from final-pass core, specialist, verified exploratory QA
|
||||
findings and invocation actions. Retain `fingerprint`, `severity`
|
||||
- Build `findings` from the final-pass findings Step 5 combined (core, specialist,
|
||||
Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA
|
||||
findings) and invocation actions. Retain `fingerprint`, `severity`
|
||||
(`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`,
|
||||
`helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`,
|
||||
never supplied/model hashes.
|
||||
|
||||
@@ -386,8 +386,9 @@ for the native result, or vice versa. Step 4.8's structured-review gate still ap
|
||||
|
||||
- Use Step 4.6's `specialists` object unchanged, including its empty small-diff map.
|
||||
If this host omits Review Army, use `specialists: {}` without claiming specialist coverage.
|
||||
- Build `findings` from final-pass core, specialist, verified exploratory QA
|
||||
findings and invocation actions. Retain `fingerprint`, `severity`
|
||||
- Build `findings` from the final-pass findings Step 5 combined (core, specialist,
|
||||
Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA
|
||||
findings) and invocation actions. Retain `fingerprint`, `severity`
|
||||
(`CRITICAL|INFORMATIONAL`), `action`, and any `advisory`, `evidence_paths`,
|
||||
`helper_target`. Recheck source after fixes. The logger uses `sharedLibsFingerprint`,
|
||||
never supplied/model hashes.
|
||||
|
||||
@@ -27,8 +27,8 @@ done
|
||||
3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check.
|
||||
|
||||
**Error handling:**
|
||||
- No plan file found → skip with "No plan file detected — skipping."
|
||||
- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping."
|
||||
- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below.
|
||||
- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified.
|
||||
|
||||
### Actionable Item Extraction
|
||||
|
||||
@@ -192,13 +192,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla
|
||||
|
||||
- **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report.
|
||||
- **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection.
|
||||
- **HIGH-impact discrepancies** trigger AskUserQuestion:
|
||||
- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion:
|
||||
- Show the investigation findings
|
||||
- Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped
|
||||
- A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review.
|
||||
- B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification.
|
||||
|
||||
This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion).
|
||||
This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion).
|
||||
Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger
|
||||
this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements.
|
||||
|
||||
When continuing after the audit (no HIGH-impact gate, or option B/C), emit the
|
||||
single final Scope Check using Step 1.5's provisional notes and this plan context:
|
||||
|
||||
@@ -71,7 +71,7 @@ Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit co
|
||||
|
||||
${reportOnly ? '' : `For /review and /ship, no plan/server is required.
|
||||
Stop after 5 minutes or 12 probes, whichever comes first (SECONDS=300 across surfaces).
|
||||
Explicit plan checks remain required beyond this smoke budget.`}
|
||||
Explicit plan checks and revalidation remain required beyond this smoke budget.`}
|
||||
For /qa and /qa-only:
|
||||
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
|
||||
- Functional Full, Quick and Regression have no default total timer.
|
||||
@@ -126,7 +126,7 @@ ${reportOnly ? ` For guarded text, copy the complete span between the guard's
|
||||
Another input or a regression test is not that replay.
|
||||
${reportOnly ? `5. If the user or another process changes source, commands or fixtures, review the affected
|
||||
contracts and return to step 2 for each affected revalidation. Do not make product changes yourself.
|
||||
Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, repeat affected review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`}
|
||||
Keep the original limits/notes; update outcomes only from fresh evidence.` : `5. After source/commands/fixtures change, re-review and return to step 2 for each affected revalidation. Keep limits/notes; status requires fresh evidence.`}
|
||||
|
||||
## 3. Parent handoff
|
||||
|
||||
@@ -255,9 +255,9 @@ Never install, import cookies or bootstrap tests. Functional-only skips browser
|
||||
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
|
||||
|
||||
**3. Run smoke and plan checks.**
|
||||
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
|
||||
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
|
||||
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
|
||||
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.${ship ? '' : ' /review sets none; only an invoker-supplied EARLIER_UTC counts.'}
|
||||
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
|
||||
|
||||
**4. Check freshness before reporting.**
|
||||
@@ -275,7 +275,8 @@ Return verified defects to Fix-First: \`path\`, \`line\`, \`category\`,
|
||||
\`fingerprint: path:line:category\`, replay, \`test_stub\`. Use checklist severity;
|
||||
unmatched functional failures are \`functional-contract\`, \`CRITICAL\`.
|
||||
Setup/permission blockers are not defects. Test creation needs user approval.
|
||||
${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : 'Ask for setup/permission, never secrets. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.'}
|
||||
${ship ? 'Step 9.4 asks: permission/repair or explicit named-risk acceptance; otherwise blocked.' : `Ask only for a named permission or setup the user performs, never secrets. Report-only /review never runs setup, installs or cookie import, even after approval.
|
||||
After a grant, recheck readiness and run affected checks; otherwise they stay blocked. Unresolved coverage makes Step 5.8 incomplete; a ship waiver cannot complete it.`}
|
||||
|
||||
${ship ? `Read QA's \`templates/functional-report-template.md\`: PR section \`## Exploratory QA\`,
|
||||
fields as subsections. Link every checkpoint; no second report. Separate browser results;
|
||||
|
||||
@@ -1483,8 +1483,9 @@ done
|
||||
3. **Validation:** For search results, read the first 20 lines and verify the project, feature and current branch. A mismatch means "no plan file found." Conversation-supplied paths bypass this search-result check.
|
||||
|
||||
**Error handling:**
|
||||
- No plan file found → skip with "No plan file detected — skipping."
|
||||
${ship ? '- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.' : '- Plan file found but unreadable (permissions, encoding) → skip with "Plan file found but unreadable — skipping."'}`;
|
||||
${ship ? `- No plan file found → skip with "No plan file detected — skipping."
|
||||
- Plan file found but unreadable (permissions, encoding) → return an audit error to the parent. Do not report no plan or successful zero counts; the parent applies its audit-failure recovery and skip/stop decision.` : `- No plan file found → say "No plan file detected." and use the Fallback Intent Sources below.
|
||||
- Plan file found but unreadable (permissions, encoding) → say "Plan file found but unreadable." and use the Fallback Intent Sources below; never report plan items as verified.`}`;
|
||||
}
|
||||
|
||||
// ─── Plan Completion Audit ────────────────────────────────────────────
|
||||
@@ -1723,13 +1724,15 @@ The plan completion results augment the existing Scope Drift Detection. If a pla
|
||||
|
||||
- **NOT DONE items** become additional evidence for **MISSING REQUIREMENTS** in the scope drift report.
|
||||
- **Items in the diff that don't match any plan item** become evidence for **SCOPE CREEP** detection.
|
||||
- **HIGH-impact discrepancies** trigger AskUserQuestion:
|
||||
- **HIGH-impact plan-file discrepancies** trigger AskUserQuestion:
|
||||
- Show the investigation findings
|
||||
- Options: A) Stop this review for implementation, B) Continue this review with P1 TODOs, C) Record the items as intentionally dropped
|
||||
- A ends this invocation before code review or implementation. List the missing work; after implementation, start a fresh /review.
|
||||
- B queues the approved TODO changes for Step 5, not this read-only audit. B/C continue to the final Scope Check and Step 2. None of these choices authorizes shipping or waives required verification.
|
||||
|
||||
This is **INFORMATIONAL** unless HIGH-impact discrepancies are found (then it gates via AskUserQuestion).
|
||||
This is **INFORMATIONAL** unless HIGH-impact plan-file discrepancies are found (then it gates via AskUserQuestion).
|
||||
Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger
|
||||
this question, whatever their IMPACT: report them in the Scope Check as lower-confidence missing requirements.
|
||||
|
||||
When continuing after the audit (no HIGH-impact gate, or option B/C), emit the
|
||||
single final Scope Check using Step 1.5's provisional notes and this plan context:
|
||||
|
||||
@@ -502,8 +502,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
|
||||
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
|
||||
|
||||
**3. Run smoke and plan checks.**
|
||||
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
|
||||
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
|
||||
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
|
||||
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
|
||||
|
||||
|
||||
+2
-2
@@ -1996,8 +1996,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
|
||||
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
|
||||
|
||||
**3. Run smoke and plan checks.**
|
||||
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
|
||||
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
|
||||
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
|
||||
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
|
||||
|
||||
|
||||
+2
-2
@@ -2255,8 +2255,8 @@ Run the shared preflight; start its smoke guard once. Guard every smoke probe. F
|
||||
- Required: plan commands/assertions, listed separately. Other ideas are optional, untested.
|
||||
|
||||
**3. Run smoke and plan checks.**
|
||||
Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.
|
||||
Then run required plan checks, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Follow the shared Probe loop for smoke checks and replays until the smoke limit.
|
||||
Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard; never reset the clock.
|
||||
Use finite command timeouts, capped at the caller's remaining time if it has a deadline.
|
||||
Await clock/guard results before acting. When the caller's deadline expires, mark unfinished checks not-run.
|
||||
|
||||
|
||||
@@ -149,7 +149,7 @@ describe('compact QA browser recipes retain native operations', () => {
|
||||
]) expect(loop).toContain(contract);
|
||||
for (const field of ['observationCommand', 'observed', 'hypothesis', 'nextCommand']) expect(loop).toContain(`${field}:`);
|
||||
expect(loop).toContain('Functional Full, Quick and Regression have no default total timer');
|
||||
if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks remain required beyond this smoke budget');
|
||||
if (skillName !== 'qa-only') expect(loop).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
@@ -70,7 +70,7 @@ describe('review and ship completion freshness contracts', () => {
|
||||
expect(gate).toContain('Reporting reserves cannot stop required revalidation within the caller\'s deadline');
|
||||
expect(body).toContain('Await clock/guard results before acting');
|
||||
expect(body).toContain('Smoke: 5 minutes/12 probes');
|
||||
expect(body).toContain('Then run required plan checks, even after smoke expires');
|
||||
expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires');
|
||||
expect(body).toContain('no smoke guard; never reset the clock');
|
||||
expect(body).toContain('Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline');
|
||||
});
|
||||
|
||||
@@ -608,8 +608,8 @@ describe('generated actual parent paths', () => {
|
||||
expect(load).toContain('Templates cannot replace them');
|
||||
const flat = parent.replace(/\s+/g, ' ');
|
||||
expect(flat).toContain('Only the parent runs report-only discovery');
|
||||
expect(flat).toContain('Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit');
|
||||
expect(flat).toContain('Then run required plan checks, even after smoke expires');
|
||||
expect(flat).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit');
|
||||
expect(flat).toContain('Then run required plan checks and revalidation, even after smoke expires');
|
||||
expect(flat).toContain('using the same procedure but no smoke guard; never reset the clock');
|
||||
expect(flat).toContain("Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline");
|
||||
expect(flat).toContain('When the caller\'s deadline expires, mark unfinished checks not-run');
|
||||
@@ -640,7 +640,7 @@ describe('generated actual parent paths', () => {
|
||||
expect(body).toContain('Stop after 5 minutes or 12 probes, whichever comes first');
|
||||
expect(body).toContain('G enforces the deadline');
|
||||
expect(body).toContain('Never reset D/bypass G');
|
||||
expect(body).toContain('Explicit plan checks remain required beyond this smoke budget');
|
||||
expect(body).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
|
||||
expect(body).toContain('leaves /review incomplete');
|
||||
expect(body).toContain('/ship blocked unless the user explicitly accepts that named risk');
|
||||
});
|
||||
|
||||
@@ -40,8 +40,8 @@ function assertBoundsAndLayout(text: string) {
|
||||
function assertPlanExecution(text: string, shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude })) {
|
||||
const step = text.slice(text.indexOf('**3. Run smoke and plan checks.**'), text.indexOf('**4. Check freshness before reporting.**')).replace(/\s+/g, ' ');
|
||||
for (const contract of [
|
||||
'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit',
|
||||
'Then run required plan checks, even after smoke expires',
|
||||
'Follow the shared Probe loop for smoke checks and replays until the smoke limit',
|
||||
'Then run required plan checks and revalidation, even after smoke expires',
|
||||
'using the same procedure but no smoke guard; never reset the clock',
|
||||
"Use finite command timeouts, capped at the caller\'s remaining time if it has a deadline",
|
||||
'When the caller\'s deadline expires, mark unfinished checks not-run',
|
||||
@@ -261,17 +261,19 @@ describe('QA probe entry and checkpoint gates', () => {
|
||||
const text = generateQAReview({ host: 'claude', skillName, tmplPath: '', paths: HOST_PATHS.claude }).replace(/\s+/g, ' ');
|
||||
assertPlanExecution(text);
|
||||
for (const [before, after] of [
|
||||
['Then run required plan checks, even after smoke expires', 'Skip plan checks when smoke expired'],
|
||||
['Then run required plan checks and revalidation, even after smoke expires', 'Skip plan checks when smoke expired'],
|
||||
['no smoke guard; never reset the clock', 'restart and use the smoke guard'],
|
||||
['same procedure', 'Start a new checkpoint sequence'],
|
||||
['at the caller\'s remaining time', 'with no caller cap'],
|
||||
['When the caller\'s deadline expires, mark unfinished checks not-run', 'If that deadline expired, mark the check passed'],
|
||||
['Await clock/guard results before acting', 'Ignore clock results'],
|
||||
['plan checks and revalidation, even', 'plan checks, even'],
|
||||
['smoke checks and replays until', 'smoke checks, replays and revalidation until'],
|
||||
]) {
|
||||
expect(text).toContain(before);
|
||||
expect(() => assertPlanExecution(text.replace(before, after))).toThrow();
|
||||
}
|
||||
const smoke = 'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit.';
|
||||
const smoke = 'Follow the shared Probe loop for smoke checks and replays until the smoke limit.';
|
||||
expect(() => assertPlanExecution(text.replace(smoke, '').replace('**4. Check', smoke + '\n**4. Check'))).toThrow();
|
||||
const shared = generateQAExploratory({ host: 'claude', skillName: 'qa', tmplPath: '', paths: HOST_PATHS.claude });
|
||||
for (const contract of ['First demonstrate success: output AND durable effects',
|
||||
|
||||
@@ -32,7 +32,7 @@ test('review audits deliverables before deferring behavioral plan checks to the
|
||||
expect(audit.indexOf('Inspect the validator and its hooks')).toBeLessThan(audit.indexOf('If found and verified safe above, invoke it'));
|
||||
expect(audit).not.toContain('For each extracted plan item, run the verification dispatch');
|
||||
const qa = generateQAReview(ctx);
|
||||
expect(qa).toContain('Then run required plan checks, even after smoke expires');
|
||||
expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires');
|
||||
expect(qa).toContain('Report clean/completed only when all required checks pass on current inputs');
|
||||
});
|
||||
|
||||
@@ -249,7 +249,7 @@ test('caller QA defines execution, evidence ownership and report adaptation befo
|
||||
for (const contract of [
|
||||
'Only the parent runs report-only discovery',
|
||||
'Never overwrite another run',
|
||||
'Follow the shared Probe loop for smoke checks, replays and revalidation until the smoke limit',
|
||||
'Follow the shared Probe loop for smoke checks and replays until the smoke limit',
|
||||
'using the same procedure but no smoke guard; never reset the clock',
|
||||
'Read agent/user updates and await results without batching them with reporting/logging',
|
||||
'Compare each probe\'s recorded source, tests, contracts, commands and fixtures (or input fingerprint) with current inputs, even without updates',
|
||||
@@ -378,7 +378,7 @@ test('review finalization ownership: the plan audit retains its high-impact gate
|
||||
expect(plan).not.toContain('never blocks the review');
|
||||
expect(plan).toContain('{{PLAN_COMPLETION_AUDIT_REVIEW}}');
|
||||
const audit = readFileSync(join(root, 'review/sections/plan-completion.md'), 'utf8');
|
||||
const gate = audit.indexOf('**HIGH-impact discrepancies** trigger AskUserQuestion');
|
||||
const gate = audit.indexOf('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion');
|
||||
expect(gate).toBeGreaterThan(-1);
|
||||
expect(gate).toBeLessThan(audit.indexOf('When continuing after the audit (no HIGH-impact gate, or option B/C)'));
|
||||
expect(audit).toContain('then it gates via AskUserQuestion');
|
||||
@@ -564,3 +564,25 @@ test('ship review clarity: parent settlement gate precedes classification and ca
|
||||
expect(gate).toContain('After terminal failure, independent evidence may support fixes');
|
||||
expect(gate).toContain('missing dispatched output still blocks continuation, even with a QA exception');
|
||||
});
|
||||
|
||||
test('review resolves the judged smoke-clock, setup-authority, plan-gate and findings-source ambiguities', () => {
|
||||
const ctx = { skillName: 'review', tmplPath: 'review/SKILL.md.tmpl', host: 'claude', paths: HOST_PATHS.claude } as TemplateContext;
|
||||
const qa = generateQAReview(ctx).replace(/\s+/g, ' ');
|
||||
expect(qa).toContain('Follow the shared Probe loop for smoke checks and replays until the smoke limit');
|
||||
expect(qa).toContain('Then run required plan checks and revalidation, even after smoke expires, using the same procedure but no smoke guard');
|
||||
expect(qa).toContain('/review sets none; only an invoker-supplied EARLIER_UTC counts');
|
||||
expect(qa).toContain('Report-only /review never runs setup, installs or cookie import, even after approval');
|
||||
expect(qa).toContain('After a grant, recheck readiness and run affected checks; otherwise they stay blocked');
|
||||
expect(qa).not.toContain('Ask for setup/permission');
|
||||
expect(generateQAReview({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).not.toContain('/review sets none');
|
||||
const shared = generateQAExploratory({ ...ctx, skillName: 'qa' }).replace(/\s+/g, ' ');
|
||||
expect(shared).toContain('Explicit plan checks and revalidation remain required beyond this smoke budget');
|
||||
const audit = generatePlanCompletionAuditReview(ctx).replace(/\s+/g, ' ');
|
||||
expect(audit).toContain('No plan file found → say "No plan file detected." and use the Fallback Intent Sources below');
|
||||
expect(audit).not.toContain('skip with "No plan file detected — skipping."');
|
||||
expect(audit).toContain('**HIGH-impact plan-file discrepancies** trigger AskUserQuestion');
|
||||
expect(audit).toContain('Discrepancies derived only from fallback sources (commit messages, TODOS.md, PR description) never trigger this question');
|
||||
expect(generatePlanCompletionAuditShip({ ...ctx, skillName: 'ship', tmplPath: 'ship/SKILL.md.tmpl' })).toContain('skip with "No plan file detected — skipping."');
|
||||
const persist = skill.slice(skill.indexOf('### 2. Fill the record')).replace(/\s+/g, ' ');
|
||||
expect(persist).toContain('findings Step 5 combined (core, specialist, Step 4.8 adversarial, VALID & ACTIONABLE Greptile and verified exploratory QA findings)');
|
||||
});
|
||||
@@ -14,7 +14,7 @@ test('Ship initializes and applies its smoke guard independently of required pla
|
||||
expect(body).toContain('Run the shared preflight; start its smoke guard once. Guard every smoke probe.');
|
||||
expect(body.indexOf('start its smoke guard once')).toBeLessThan(body.indexOf('**3. Run smoke and plan checks.**'));
|
||||
expect(body).toContain('Required even for small diffs or missing plans/servers');
|
||||
expect(body).toContain('Then run required plan checks, even after smoke expires');
|
||||
expect(body).toContain('Then run required plan checks and revalidation, even after smoke expires');
|
||||
expect(body).toContain('using the same procedure but no smoke guard; never reset the clock');
|
||||
}
|
||||
});
|
||||
|
||||
Reference in new issue
Block a user