mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-13 17:50:22 +02:00
v1.62.0.0 feat: plan-mode auto-select at the review scope gate (#2533)
* fix(evals): align plan-eng/design plan-mode + finding-floor smokes to their declared periodic tier The #2077 demotion of these four stochastic tests to 'periodic' was inert: E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate', so they kept running in the blocking gate lane and never in the weekly lane. Flip the four self-gates to 'periodic' (headers/describe labels updated), add a free static tier-alignment invariant test (dep-list filename mapping; unmapped self-gated files are reported, never silently skipped), and name the two plan-mode test files in their own touchfiles dep lists so the invariant binds for them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(pty-runner): scope-gate question/auto-select detectors + observation flags Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5 collapsed-form handling): isScopeGateQuestionVisible requires the question text PLUS option-body text (native AskUserQuestion renders numbered options, prose fallback renders lettered — the option body appears in both; narration doesn't), and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the selected-B token. runPlanSkillObservation gains scopeGateQuestionObserved / scopeGateAutoSelectObserved high-water flags (attached at every return path) so paid smokes can assert gate behavior across the whole run instead of the lossy 2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven questions, and the gate could fire inside the 3s pre-target window. Unit fixtures pin clean/native/collapsed positives, narration negatives, and the verbatim template announcement string (template rewording fails here first, before the paid smokes degrade to vacuous asserts). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(plan-eng/design-review): auto-select B in plan mode at the scope gate In plan mode the scope gate's "What should I review? A/B/C" question is pure friction: there is no branch diff and the target is the plan being drafted. Both gates gain an ordered exceptions block, checked BEFORE asking: 1. Plan mode → auto-select B: review the active plan (in context or pasted), announce it in one line ("Scope gate: plan mode — auto-selected B (reviewing <target>)") so the user can interrupt; an explicitly different user-named target still wins; no plan drafted yet → ask as normal. 2. User-named target (outside plan mode): explicit-only — a path, a pasted doc, or the literal words "branch diff". A passing mention is not naming; when in doubt, ask. Outside plan mode with no explicitly-named target, nothing changes. Plan-mode is checked FIRST because the PTY harness seeds drafts as pasted user messages (claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic. Pinning: seeded plan-mode smokes assert no gate render + announcement rendered (eng test 2; new design seeded test); plan-mode-no-op extends to eng/design (bypass must not misfire outside plan mode; first question must be the gate) plus a named-target case proving the pasted target is consumed; a drift-guard asserts the two hand-duplicated exceptions blocks stay identical modulo the two variant slots and carry the announcement string the detectors pin. Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): skip the scope gate when following loaded review skills autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim, and its section skip list omitted the scope gate — so autoplan ingested a hard-STOP AskUserQuestion that contradicts its every-question-auto-decides contract. One skip-list line fixes it; a static toContain pin in skill-validation keeps the entry load-bearing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: file scope-gate resolver-extraction TODO (eng-review D5 follow-up) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): positional floor exclusion, flag builder, outcome union, token tracking Review-army + adversarial findings on the scope-gate observability work, all verified before fixing: - Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate exclusion scanned only the 1500-byte tail, so an early gate render satisfied the floor vacuously once ~1.5KB of output accumulated (found independently by 4 review passes; predicate reproduced). Acceptance now scans only content APPENDED after the first gate render (positional anchor), and the LLM-judge 'waiting' shortcut no longer fires while the gate menu is the pending render. - High-water flags are built once and spread at every return path — the hand-spread pattern had already drifted (judge-waiting return omitted two flags), which made must-stay-false asserts vacuous on those paths. - isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases) and quoted-occurrence rejection (a model verbatim-quoting the announcement while declining must not trip must-stay-FALSE asserts). Fixtures added for both directions. - PlanSkillObservation outcome union gains 'wrote_findings_before_asking' (returned at runtime via classifyVisible but missing from the type). - trackTokens/tokensObserved: cumulative-buffer token high-water for consumption asserts (the 2KB evidence tail is lossy and the plan-file fallback is unreachable outside plan mode). - New scope-gate-floor unit pins (from the ship coverage audit): both gate render forms trip acceptance and exclusion; a genuine finding AUQ is not excluded; tail-scoping semantics pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(evals): harden no-op asserts, close tier-invariant fail-open holes, pin gate question strings - no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the outcome==='asked' conditional let a silent-bypass plan_ready run sail through); eng/design cases force --disallowedTools so the pinned prose shape is contractual rather than hoping native AUQ renders match; the named-target case uses trackTokens for consumption and lists wrote_findings_before_asking in its diagnostic throw branch. - tier-alignment invariant: both quote styles matched; zero-self-gate, mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED instead of silently skipped (the fail-open holes three reviewers found). - drift-guard: the generated gate menus must carry the exact question/option strings the PTY question detector anchors on — free CI fails before the paid smokes can go vacuous on a menu reword. - touchfiles: corrected the no-op cost note for CI concurrency + retry semantics. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): register plan-eng/design-review skills in PTY eval containers The extended plan-mode-no-op smoke invokes /plan-eng-review and /plan-design-review, but the fresh CI containers registered only office-hours and plan-ceo-review — both new runs would return 'Unknown command' and fail every PR's gate job (Codex structured review P1, verified against evals.yml). Registration loops, the dangling-target fail-fast list, and the frontmatter checks (now a loop over the same skill list, so the lists can't drift) all cover the two skills. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(plan-eng/design-review): harden scope-gate exceptions against injection and ambiguity Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex cross-confirmation), applied to both gate templates + regen: - Host-anchored mode signal: only the host's own system messages (plan-mode reminder or active plan file path) arm the auto-select; plan-shaped text inside pasted documents, tool results, or fetched pages does NOT count — injected content can't disarm the consent gate or nominate the target. - Multiple plan candidates: the host-referenced plan file wins; still ambiguous means ask. - The DIFFERENT-target override carries the passing-mention guard. - Plan mode + explicitly named target + no drafted plan resolves to the named target instead of a contradictory re-ask. - The numbered ask-path rules are qualified ('When no exception above applied:') so they no longer restate an unconditional MUST-ask that contradicts the exceptions. - 'Whenever this gate does ask — in any mode — it is a hard STOP.' - Shared preamble: 'any AskUserQuestion the skill fires is the workflow operating within plan mode' (was 'the first AskUserQuestion is the workflow entering plan mode', which framed the opposite of the bypass); regenerates every skill. - Ceilings ratcheted with attribution: plan-eng union ratio 1.10, investigate 1.10 (the ~250B shared-preamble reword lands the closest-to-ceiling skill at 1.092). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.62.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): active-render gate veto in the floor check + honest periodic-wiring docs Codex re-review P2s on the fix wave, both verified: - A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting, no further output) was vetoed by the blanket tail exclusion until timeout. The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last cursor menu, so only a pending GATE menu vetoes; the judge fallback shares the same check. Residual (documented): prose gate + prose finding inside one tail — floors run the native-menu path in practice. - The four demoted periodic tests are not in evals-periodic.yml's explicit matrix (a named instance of the pre-existing periodic-orphans TODO), so they run locally/manually until the PTY-capable periodic job lands. CHANGELOG claim softened accordingly; TODO filed with the wiring recipe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.62.0.0 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply codex doc-review fixes for v1.62.0.0 - CLAUDE.md: scope the tier-alignment invariant claim (mapped files enforced, unmapped files reported) - docs/skills.md: document the plan-mode auto-select scope gate for /plan-eng-review and /plan-design-review - evals.yml: fix stale comment (PTY smokes register four skills, not two) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh ship golden baselines for the plan-mode preamble reword The generate-completion-status.ts wording change ('any AskUserQuestion the skill fires…') intentionally regenerates every SKILL.md; the byte-compare goldens carry the generator's output and refresh with it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): custom-hooks-path detection false-negatives on git worktrees The pre-push guard's HOOKS_IN_GIT_DIR check compared the hooks dir against --absolute-git-dir, which in a linked worktree is .git/worktrees/<name> while hooks resolve to the COMMON .git/hooks — so every Conductor worktree read as a 'custom hooks path' and the consented guard install was skipped. Match against the resolved --git-common-dir too (with a /nonexistent fallback so a failed resolution can't collapse the case pattern into match-everything). Verified live: this worktree now reports yes (was no), and the main checkout still reports yes. Goldens refreshed (--host all). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: changelog bullet for the worktree hooks-detection fix Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): give the plan-ceo plan-mode smoke real budget headroom Measured 2026-08-11: a clean isolated pass took 295.7s against the 300s inner budget (4s of margin) and the same test timed out at ~308s three times under concurrent eval load — a budget-edge flake in the gate lane, not a behavior regression (it passed isolated on both this branch and main). Inner budget 300s -> 420s, outer bun timeout 360s -> 480s, and the test file is now named in its own touchfiles dep list so the tier-alignment invariant binds for it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): 300s budget floor for the two 90s design-consultation SDK tests Root cause of PR #2533's e2e-design CI failure: design-consultation-preview failed 3 attempts at 0 turns/$0.00/93s — the session was up but the model's first completion queued past the 90s inner budget under concurrent API load (11 matrix jobs; the sibling research test booted its first tool at 4s, so this is API-side queuing, not CPU boot contention). The test was selected only because touchfiles.ts is a global touchfile; the tested behavior is untouched by this branch. 90s budgets cannot absorb one slow first completion. Both 90s tests in the file move to the repo's saturated-runner standard (300s inner / 360s outer, matching review-dashboard-via and retro-base-branch). Deliberately NOT re-arming the runner's inner timer on first stream event: an audit found ~100 outer bun-timeout literals sized inner+30-60s that a re-arm would silently break — the structural options are written up in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
94993f7401
commit
d078622b73
@@ -0,0 +1,91 @@
|
||||
/**
|
||||
* Tier-alignment invariant (free, static).
|
||||
*
|
||||
* Kills the "inert demotion" defect class: E2E_TIERS declares a test's tier,
|
||||
* but the paid test files also self-gate on `process.env.EVALS_TIER === '<tier>'`.
|
||||
* When the two disagree, the touchfiles declaration is dead metadata — the
|
||||
* #2077 demotion of the plan-mode/finding-floor smokes to 'periodic' was inert
|
||||
* for months because the files still gated on 'gate' and ran in the blocking
|
||||
* lane on every gate run.
|
||||
*
|
||||
* Mapping rule (test filenames do NOT map mechanically to tier keys): for each
|
||||
* `test/skill-e2e-*.test.ts` with an EVALS_TIER self-gate, search the
|
||||
* E2E_TOUCHFILES / LLM_JUDGE_TOUCHFILES dep lists for the exact file path. If
|
||||
* found under key K, the file's self-gate tier must equal E2E_TIERS[K]. Files
|
||||
* not named in any dep list are REPORTED as unmapped (a nudge to add them to
|
||||
* their eval's dep list), never silently skipped.
|
||||
*/
|
||||
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { readdirSync, readFileSync } from 'fs';
|
||||
import * as path from 'path';
|
||||
import { E2E_TOUCHFILES, E2E_TIERS, LLM_JUDGE_TOUCHFILES } from './helpers/touchfiles';
|
||||
|
||||
const TEST_DIR = import.meta.dir;
|
||||
// Both quote styles — a mechanical refactor to double quotes must not
|
||||
// silently drop a file from the invariant (fail-open is the defect class
|
||||
// this test exists to kill).
|
||||
const SELF_GATE_RE = /EVALS_TIER\s*===\s*['"](gate|periodic)['"]/g;
|
||||
|
||||
describe('E2E tier alignment (touchfiles declaration vs test self-gate)', () => {
|
||||
const testFiles = readdirSync(TEST_DIR)
|
||||
.filter((f) => f.startsWith('skill-e2e-') && f.endsWith('.test.ts'))
|
||||
.sort();
|
||||
|
||||
const allDeps: Record<string, string[]> = { ...E2E_TOUCHFILES, ...LLM_JUDGE_TOUCHFILES };
|
||||
|
||||
test('every self-gated test file named in a dep list matches its declared tier', () => {
|
||||
const misaligned: string[] = [];
|
||||
const reported: string[] = [];
|
||||
|
||||
for (const file of testFiles) {
|
||||
const content = readFileSync(path.join(TEST_DIR, file), 'utf-8');
|
||||
const tiers = new Set<string>();
|
||||
for (const m of content.matchAll(SELF_GATE_RE)) tiers.add(m[1]);
|
||||
const repoPath = `test/${file}`;
|
||||
if (tiers.size === 0) {
|
||||
// Every skill-e2e file is expected to self-gate; zero matches means
|
||||
// either a genuinely ungated file or a gate shape the regex can't
|
||||
// see — both worth a visible report, never a silent skip.
|
||||
reported.push(`${repoPath}: no detectable EVALS_TIER self-gate`);
|
||||
continue;
|
||||
}
|
||||
if (tiers.size > 1) {
|
||||
reported.push(`${repoPath}: mixed-tier self-gates (${[...tiers].join(', ')}) — not tier-checked`);
|
||||
continue;
|
||||
}
|
||||
const selfTier = [...tiers][0];
|
||||
|
||||
const owningKeys = Object.keys(allDeps).filter((k) => allDeps[k].includes(repoPath));
|
||||
if (owningKeys.length === 0) {
|
||||
reported.push(`${repoPath} (self-gates '${selfTier}'): not named in any touchfiles dep list`);
|
||||
continue;
|
||||
}
|
||||
for (const k of owningKeys) {
|
||||
const declared = E2E_TIERS[k];
|
||||
if (!declared) {
|
||||
// A dep-list key with no E2E_TIERS entry (e.g. an LLM-judge key)
|
||||
// can't tier-check this file — report instead of silently passing.
|
||||
reported.push(`${repoPath}: matched key '${k}' which has no E2E_TIERS entry`);
|
||||
continue;
|
||||
}
|
||||
if (declared !== selfTier) {
|
||||
misaligned.push(
|
||||
`${repoPath}: self-gates on '${selfTier}' but E2E_TIERS['${k}'] declares '${declared}' — the declaration is inert`,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Reported, not asserted: coverage holes the invariant can see but not
|
||||
// arbitrate. Add the test file to its eval's dep list (or a tier entry
|
||||
// for the key) to bring it under the invariant.
|
||||
if (reported.length > 0) {
|
||||
console.warn(
|
||||
`[tier-alignment] ${reported.length} file(s) outside the invariant:\n ` + reported.join('\n '),
|
||||
);
|
||||
}
|
||||
|
||||
expect(misaligned).toEqual([]);
|
||||
});
|
||||
});
|
||||
+8
-2
@@ -150,7 +150,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
|
||||
|
||||
## Skill Invocation During Plan Mode
|
||||
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
|
||||
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
|
||||
|
||||
@@ -1272,9 +1272,15 @@ _HOOK_INSTALLED="no"
|
||||
# committed hook and write a machine-local wrapper into the working tree.
|
||||
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
|
||||
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
|
||||
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
|
||||
# resolve to the COMMON .git/hooks, so match against the common dir too or
|
||||
# every Conductor worktree false-negatives as a "custom hooks path". The
|
||||
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
|
||||
# (match-everything) when resolution fails.
|
||||
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
|
||||
_HOOKS_IN_GIT_DIR="no"
|
||||
case "$_HOOKS_DIR" in
|
||||
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
esac
|
||||
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
|
||||
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
|
||||
|
||||
+8
-2
@@ -136,7 +136,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
|
||||
|
||||
## Skill Invocation During Plan Mode
|
||||
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
|
||||
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
|
||||
|
||||
@@ -2455,9 +2455,15 @@ _HOOK_INSTALLED="no"
|
||||
# committed hook and write a machine-local wrapper into the working tree.
|
||||
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
|
||||
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
|
||||
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
|
||||
# resolve to the COMMON .git/hooks, so match against the common dir too or
|
||||
# every Conductor worktree false-negatives as a "custom hooks path". The
|
||||
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
|
||||
# (match-everything) when resolution fails.
|
||||
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
|
||||
_HOOKS_IN_GIT_DIR="no"
|
||||
case "$_HOOKS_DIR" in
|
||||
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
esac
|
||||
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
|
||||
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
|
||||
|
||||
+8
-2
@@ -138,7 +138,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
|
||||
|
||||
## Skill Invocation During Plan Mode
|
||||
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
|
||||
|
||||
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
|
||||
|
||||
@@ -2861,9 +2861,15 @@ _HOOK_INSTALLED="no"
|
||||
# committed hook and write a machine-local wrapper into the working tree.
|
||||
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
|
||||
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
|
||||
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
|
||||
# resolve to the COMMON .git/hooks, so match against the common dir too or
|
||||
# every Conductor worktree false-negatives as a "custom hooks path". The
|
||||
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
|
||||
# (match-everything) when resolution fails.
|
||||
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
|
||||
_HOOKS_IN_GIT_DIR="no"
|
||||
case "$_HOOKS_DIR" in
|
||||
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
|
||||
esac
|
||||
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
|
||||
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
|
||||
|
||||
@@ -3244,6 +3244,66 @@ describe('EXIT PLAN MODE GATE placement', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('scope-gate exceptions drift-guard', () => {
|
||||
// The plan-mode auto-select-B exceptions block is hand-duplicated in the
|
||||
// plan-eng-review and plan-design-review templates (matching the gate
|
||||
// around it, which predates this block). The two copies must stay
|
||||
// byte-identical modulo exactly two known variant slots:
|
||||
// 1. the plan-mode bullet's action tail (Design Doc Check vs pre-review
|
||||
// audit + mockups),
|
||||
// 2. the named-target vocabulary ("a path, a doc" vs "a path, a page, a doc").
|
||||
// A future edit to one copy that silently misses the other fails here
|
||||
// instead of drifting. The real fix (shared {{SCOPE_GATE}} resolver) is a
|
||||
// filed TODO — this guard is the stopgap that makes the duplication safe.
|
||||
const START_MARKER = '**Exceptions — check in this order, BEFORE asking:**';
|
||||
const END_MARKER = 'in any mode — it is a hard STOP.';
|
||||
|
||||
function extractExceptionsBlock(skill: string): string {
|
||||
const md = fs.readFileSync(path.join(ROOT, skill, 'SKILL.md'), 'utf-8');
|
||||
const start = md.indexOf(START_MARKER);
|
||||
expect(start, `${skill}/SKILL.md: exceptions block start marker present`).toBeGreaterThan(-1);
|
||||
const end = md.indexOf(END_MARKER, start);
|
||||
expect(end, `${skill}/SKILL.md: exceptions block end marker present`).toBeGreaterThan(start);
|
||||
return md.slice(start, end + END_MARKER.length);
|
||||
}
|
||||
|
||||
const normalizeVariantSlots = (block: string) =>
|
||||
block
|
||||
.replace('Then run the Design Doc Check and Step 0 against that plan.', '<ACTION_TAIL>')
|
||||
.replace('Then run the pre-review audit, mockups, and Step 0 against that plan.', '<ACTION_TAIL>')
|
||||
.replace('a path, a page, a doc they pasted,', 'a path, a doc they pasted,');
|
||||
|
||||
test('eng and design exceptions blocks are identical modulo the two variant slots', () => {
|
||||
const eng = normalizeVariantSlots(extractExceptionsBlock('plan-eng-review'));
|
||||
const design = normalizeVariantSlots(extractExceptionsBlock('plan-design-review'));
|
||||
expect(eng).toBe(design);
|
||||
// The action tail must actually have been normalized in both (guards
|
||||
// against a rewording that bypasses the normalizer and vacuously passes).
|
||||
expect(eng).toContain('<ACTION_TAIL>');
|
||||
});
|
||||
|
||||
test('exceptions block carries the announcement string the PTY detectors pin', () => {
|
||||
for (const skill of ['plan-eng-review', 'plan-design-review']) {
|
||||
const block = extractExceptionsBlock(skill);
|
||||
expect(block, `${skill}: verbatim announcement`).toContain(
|
||||
'Scope gate: plan mode — auto-selected B (reviewing <target>).',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('gate menu carries the question strings the PTY question detector pins', () => {
|
||||
// isScopeGateQuestionVisible (claude-pty-runner.ts) anchors on the
|
||||
// question text + option A's body. If the menu is reworded without
|
||||
// updating the detector, the paid smokes' must-stay-false assertions go
|
||||
// vacuous — this free pin fails first.
|
||||
for (const skill of ['plan-eng-review', 'plan-design-review']) {
|
||||
const md = fs.readFileSync(path.join(ROOT, skill, 'SKILL.md'), 'utf-8');
|
||||
expect(md, `${skill}: gate question text`).toContain('What should I review?');
|
||||
expect(md, `${skill}: option A body text`).toContain('The current branch diff');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('GSTACK REVIEW REPORT mandatory unresolved-decisions status', () => {
|
||||
// Report text rides in PLAN_FILE_REVIEW_REPORT → every report consumer gets it.
|
||||
// devex-review is a report consumer but NOT a gate consumer, so the two target
|
||||
|
||||
@@ -164,7 +164,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
},
|
||||
behavioral: 'plan',
|
||||
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
|
||||
maxSkeletonBytes: 67_000,
|
||||
// +~1 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
|
||||
maxSkeletonBytes: 68_000,
|
||||
minUnionBytes: 70_000,
|
||||
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
|
||||
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
|
||||
@@ -172,7 +173,10 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
||||
// prose, replacing the smaller opt-in question) land this at ~6.6% over the
|
||||
// v1.53.0.0 baseline. Headroom for those intentional additions.
|
||||
maxSizeRatio: 1.08,
|
||||
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
|
||||
// hardening: host-anchored mode signal, precedence, passing-mention
|
||||
// guards) and the plan-mode preamble reword land the union at 1.092.
|
||||
maxSizeRatio: 1.10,
|
||||
},
|
||||
'plan-design-review': {
|
||||
skill: 'plan-design-review',
|
||||
@@ -189,7 +193,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
||||
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
|
||||
// always-loaded AskUserQuestion Format section.
|
||||
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
|
||||
maxSkeletonBytes: 88_000,
|
||||
// +~1.3 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
|
||||
maxSkeletonBytes: 89_000,
|
||||
minUnionBytes: 70_000,
|
||||
mustContain: ['design', 'visual'],
|
||||
maxSizeRatio: 1.07,
|
||||
|
||||
@@ -0,0 +1,143 @@
|
||||
/**
|
||||
* Scope-gate floor-exclusion regression pins (free, static).
|
||||
*
|
||||
* runPlanSkillFloorCheck's acceptance condition changed with the plan-mode
|
||||
* auto-select-B work: a render only satisfies the finding floor when
|
||||
*
|
||||
* (isNumberedOptionListVisible(visible) || isProseAUQVisible(visible))
|
||||
* && !isPermissionDialogVisible(tail)
|
||||
* && !isScopeGateQuestionVisible(tail) // <- new exclusion
|
||||
*
|
||||
* where tail = visible.slice(-TAIL_SCAN_BYTES). The composition lives inline
|
||||
* in the paid PTY loop, so these tests pin the load-bearing behavior of each
|
||||
* detector on the exact render shapes the floor passes them:
|
||||
*
|
||||
* 1. Both scope-gate render forms (native numbered UI, prose lettered
|
||||
* fallback) trip the acceptance detectors — WITHOUT the exclusion the
|
||||
* gate would trivially satisfy the floor inside the 3s pre-target
|
||||
* window. The exclusion must catch both forms.
|
||||
* 2. A genuine finding-driven AskUserQuestion must NOT trip the exclusion,
|
||||
* or the floor becomes unsatisfiable.
|
||||
* 3. The exclusion is TAIL-scoped by design: an early gate render that has
|
||||
* scrolled past TAIL_SCAN_BYTES must not suppress a later real finding
|
||||
* AskUserQuestion.
|
||||
*
|
||||
* Also closes the untested OR-branch of isScopeGateAutoSelectVisible: the
|
||||
* fully-collapsed hyphen-less 'autoselectedb' form.
|
||||
*/
|
||||
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import {
|
||||
TAIL_SCAN_BYTES,
|
||||
isNumberedOptionListVisible,
|
||||
isProseAUQVisible,
|
||||
isPermissionDialogVisible,
|
||||
isScopeGateQuestionVisible,
|
||||
isScopeGateAutoSelectVisible,
|
||||
parseNumberedOptions,
|
||||
} from './claude-pty-runner';
|
||||
|
||||
// The gate's native AskUserQuestion render (numbered options + cursor) —
|
||||
// what fires inside the floor check's 3s window before the seed arrives.
|
||||
const GATE_NATIVE_RENDER = `
|
||||
What should I review?
|
||||
|
||||
❯ 1. The current branch diff — the work in progress on this branch.
|
||||
2. A plan or design doc I'll paste or point you to.
|
||||
3. A specific file, directory, or path.
|
||||
`;
|
||||
|
||||
// The gate's prose fallback render (lettered options under --disallowedTools).
|
||||
const GATE_PROSE_RENDER = `
|
||||
What should I review?
|
||||
A) The current branch diff — the work in progress on this branch.
|
||||
B) A plan or design doc I'll paste or point you to.
|
||||
C) A specific file, directory, or path.
|
||||
Recommendation: A when a branch diff exists, otherwise B.
|
||||
`;
|
||||
|
||||
// A genuine finding-driven AskUserQuestion — the render the floor MEASURES.
|
||||
const FINDING_AUQ_RENDER = `
|
||||
Finding 1: the plan reimplements test sharding that Bun provides natively.
|
||||
|
||||
❯ 1. Use Bun's native --shard flag (recommended)
|
||||
2. Keep the custom scheduler as planned
|
||||
3. Defer this decision to implementation
|
||||
`;
|
||||
|
||||
describe('floor-check scope-gate exclusion (acceptance-condition regression)', () => {
|
||||
test('native gate render trips the acceptance detector — the exclusion is load-bearing', () => {
|
||||
// Pre-exclusion, this render satisfied the floor by itself.
|
||||
expect(isNumberedOptionListVisible(GATE_NATIVE_RENDER)).toBe(true);
|
||||
expect(isPermissionDialogVisible(GATE_NATIVE_RENDER)).toBe(false);
|
||||
// The new exclusion catches it.
|
||||
expect(isScopeGateQuestionVisible(GATE_NATIVE_RENDER)).toBe(true);
|
||||
});
|
||||
|
||||
test('prose gate render trips the prose-AUQ arm — the exclusion catches that form too', () => {
|
||||
expect(isProseAUQVisible(GATE_PROSE_RENDER)).toBe(true);
|
||||
expect(isPermissionDialogVisible(GATE_PROSE_RENDER)).toBe(false);
|
||||
expect(isScopeGateQuestionVisible(GATE_PROSE_RENDER)).toBe(true);
|
||||
});
|
||||
|
||||
test('a genuine finding AskUserQuestion is NOT excluded — the floor stays satisfiable', () => {
|
||||
expect(isNumberedOptionListVisible(FINDING_AUQ_RENDER)).toBe(true);
|
||||
expect(isPermissionDialogVisible(FINDING_AUQ_RENDER)).toBe(false);
|
||||
expect(isScopeGateQuestionVisible(FINDING_AUQ_RENDER)).toBe(false);
|
||||
});
|
||||
|
||||
test('tail-scoping: an early gate render scrolled out of the tail does not suppress a later finding AUQ', () => {
|
||||
// Gate render, then >TAIL_SCAN_BYTES of review output, then the real
|
||||
// finding AskUserQuestion — the shape the TAIL-scoped exclusion exists for.
|
||||
const filler = 'Reading the plan and auditing the design system.\n'.repeat(
|
||||
Math.ceil(TAIL_SCAN_BYTES / 48) + 4,
|
||||
);
|
||||
const visible = GATE_NATIVE_RENDER + filler + FINDING_AUQ_RENDER;
|
||||
const tail = visible.slice(-TAIL_SCAN_BYTES);
|
||||
|
||||
// Full buffer still remembers the gate (scrollback)…
|
||||
expect(isScopeGateQuestionVisible(visible)).toBe(true);
|
||||
// …but the floor's exclusion looks only at the tail, which is clean:
|
||||
expect(isScopeGateQuestionVisible(tail)).toBe(false);
|
||||
// and the acceptance arm (full-buffer scan) sees the finding AUQ.
|
||||
expect(isNumberedOptionListVisible(visible)).toBe(true);
|
||||
expect(isPermissionDialogVisible(tail)).toBe(false);
|
||||
});
|
||||
|
||||
test('a gate render inside the tail IS suppressed (no false floor pass)', () => {
|
||||
const tail = GATE_NATIVE_RENDER.slice(-TAIL_SCAN_BYTES);
|
||||
expect(isScopeGateQuestionVisible(tail)).toBe(true);
|
||||
});
|
||||
|
||||
test('active-render veto: a finding AUQ close after the gate is NOT vetoed (codex P2 re-review)', () => {
|
||||
// The finding menu renders <TAIL_SCAN_BYTES after the gate, then the
|
||||
// model waits (no further output). A blanket tail veto would suppress
|
||||
// this until timeout; the active-render veto anchors on the LAST cursor
|
||||
// menu, which is the finding AUQ, so the floor is satisfiable.
|
||||
const visible = GATE_NATIVE_RENDER + '\nAuditing the plan…\n' + FINDING_AUQ_RENDER;
|
||||
const activeMenu = parseNumberedOptions(visible);
|
||||
expect(activeMenu.length).toBeGreaterThan(0);
|
||||
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
|
||||
expect(gateIsActiveRender).toBe(false);
|
||||
});
|
||||
|
||||
test('active-render veto: the gate as the pending menu IS vetoed', () => {
|
||||
const visible = 'booting…\n' + GATE_NATIVE_RENDER;
|
||||
const activeMenu = parseNumberedOptions(visible);
|
||||
expect(activeMenu.length).toBeGreaterThan(0);
|
||||
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
|
||||
expect(gateIsActiveRender).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('isScopeGateAutoSelectVisible collapsed hyphen-less branch', () => {
|
||||
test("matches the fully-collapsed 'autoselectedb' form (hyphen lost in TTY reflow)", () => {
|
||||
const sample = 'Scopegate:planmode—autoselectedB(reviewingPLAN.md).';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('hyphen-less token without the announcement prefix stays false', () => {
|
||||
const sample = 'The agent autoselectedB from the menu without announcing a scope gate decision.';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -618,6 +618,55 @@ export function isProseAUQVisible(visible: string): boolean {
|
||||
return false;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Scope-gate render detectors (plan-eng-review / plan-design-review)
|
||||
// ---------------------------------------------------------------------------
|
||||
//
|
||||
// Both anchor on the RENDER SHAPE, not bare keywords, so model narration
|
||||
// about the gate ("normally I'd ask what should I review…") stays false.
|
||||
// Matching is whitespace-squished + lowercased because stripAnsi collapses
|
||||
// TTY cursor-positioning escapes unpredictably (the same failure mode the
|
||||
// Pattern-4/5 collapsed-form handling above exists for).
|
||||
|
||||
/**
|
||||
* True when the scope-gate QUESTION is actually rendered: the question text
|
||||
* plus option A's body text. Option-body anchoring (not `A)`/`B)` markers)
|
||||
* because native AskUserQuestion renders NUMBERED options in the TTY while
|
||||
* the --disallowedTools prose fallback renders lettered ones — the option
|
||||
* body appears in both renders; narration rarely quotes both the question
|
||||
* and an option body.
|
||||
*/
|
||||
export function isScopeGateQuestionVisible(visible: string): boolean {
|
||||
const squished = visible.replace(/\s+/g, '').toLowerCase();
|
||||
return squished.includes('whatshouldireview') && squished.includes('currentbranchdiff');
|
||||
}
|
||||
|
||||
/**
|
||||
* True when the plan-mode auto-select announcement is rendered:
|
||||
* "Scope gate: plan mode — auto-selected B (reviewing <target>)."
|
||||
* Requires BOTH the announcement prefix and an auto-select-B token so
|
||||
* narration ("in plan mode I'd auto-select B") stays false. The token is
|
||||
* tense-tolerant (selected/selecting/selects) because the smokes assert
|
||||
* must-be-TRUE on it — a semantically-perfect paraphrase must not fail a
|
||||
* paid run — while the prefix stays exact so paraphrase narration without
|
||||
* the announcement frame stays false. A prefix immediately preceded by a
|
||||
* quote character is a QUOTATION (e.g. the model explaining why it is NOT
|
||||
* announcing), not a render — the announcement line itself never renders
|
||||
* quoted.
|
||||
*/
|
||||
export function isScopeGateAutoSelectVisible(visible: string): boolean {
|
||||
const squished = visible.replace(/\s+/g, '').toLowerCase();
|
||||
const QUOTES = ['"', "'", '`', '“', '‘'];
|
||||
const re = /scopegate:planmode/g;
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(squished)) !== null) {
|
||||
const before = m.index > 0 ? squished[m.index - 1]! : '';
|
||||
if (QUOTES.includes(before)) continue; // quoted occurrence — narration, keep scanning
|
||||
if (/auto-?select(?:ed|ing|s)?b/.test(squished.slice(m.index))) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a rendered numbered-option list out of the visible TTY text.
|
||||
*
|
||||
@@ -1476,10 +1525,20 @@ export interface PlanSkillObservation {
|
||||
* "Ready to execute" confirmation
|
||||
* - 'silent_write' — a Write/Edit landed BEFORE any prompt, to a path
|
||||
* outside the sanctioned plan/project directories
|
||||
* - 'wrote_findings_before_asking' — strictPlanWrites only (seeded runs):
|
||||
* the plan file was rewritten with findings before any
|
||||
* AskUserQuestion render (the May-2026 transcript bug)
|
||||
* - 'exited' — claude process died before any of the above
|
||||
* - 'timeout' — none of the above within budget
|
||||
*/
|
||||
outcome: 'asked' | 'auto_decided' | 'plan_ready' | 'silent_write' | 'exited' | 'timeout';
|
||||
outcome:
|
||||
| 'asked'
|
||||
| 'auto_decided'
|
||||
| 'plan_ready'
|
||||
| 'silent_write'
|
||||
| 'wrote_findings_before_asking'
|
||||
| 'exited'
|
||||
| 'timeout';
|
||||
/** Human-readable summary. */
|
||||
summary: string;
|
||||
/** Visible terminal text since the slash command was sent (last 2KB). */
|
||||
@@ -1516,6 +1575,28 @@ export interface PlanSkillObservation {
|
||||
* Haiku judge fallback rather than the regex detector.
|
||||
*/
|
||||
waitingEverObserved?: boolean;
|
||||
/**
|
||||
* High-water-mark flag: did the scope-gate QUESTION ("What should I
|
||||
* review?" plus option-body text) ever render during the run? Same
|
||||
* lossy-2KB-evidence rationale as proseAUQEverObserved. The plan-mode
|
||||
* smokes assert this stays false (gate bypassed via auto-select B); the
|
||||
* no-op regression asserts it fires outside plan mode.
|
||||
*/
|
||||
scopeGateQuestionObserved?: boolean;
|
||||
/**
|
||||
* High-water-mark flag: did the plan-mode auto-select announcement
|
||||
* ("Scope gate: plan mode — auto-selected B …") ever render? The
|
||||
* plan-mode smokes assert true; the no-op regression asserts false.
|
||||
*/
|
||||
scopeGateAutoSelectObserved?: boolean;
|
||||
/**
|
||||
* High-water map for opts.trackTokens: token → did it EVER appear in the
|
||||
* cumulative visible buffer? Consumption asserts (e.g. "the pasted target's
|
||||
* distinctive token shows up in the review output") must not depend on the
|
||||
* lossy 2KB evidence tail — plan-file fallbacks are unreachable outside
|
||||
* plan mode (extractPlanFilePath only matches plan-mode save renders).
|
||||
*/
|
||||
tokensObserved?: Record<string, boolean>;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1576,6 +1657,10 @@ export async function runPlanSkillObservation(opts: {
|
||||
/** Override the spawned model. Defaults via launchClaudePty's chain
|
||||
* (opts.model ?? EVALS_MODEL ?? 'claude-sonnet-4-6'). */
|
||||
model?: string;
|
||||
/** Literal tokens to track as high-water marks over the CUMULATIVE visible
|
||||
* buffer (case-sensitive). Results land in obs.tokensObserved. Use for
|
||||
* consumption asserts that must survive the 2KB evidence tail. */
|
||||
trackTokens?: string[];
|
||||
}): Promise<PlanSkillObservation> {
|
||||
const startedAt = Date.now();
|
||||
const session = await launchClaudePty({
|
||||
@@ -1619,6 +1704,21 @@ export async function runPlanSkillObservation(opts: {
|
||||
// even if the current state is 'working'.
|
||||
let proseAUQEverObserved = false;
|
||||
let waitingEverObserved = false;
|
||||
let scopeGateQuestionObserved = false;
|
||||
let scopeGateAutoSelectObserved = false;
|
||||
const tokensObserved: Record<string, boolean> = {};
|
||||
for (const t of opts.trackTokens ?? []) tokensObserved[t] = false;
|
||||
// Single source for the high-water flags at EVERY return site. Hand-
|
||||
// spreading them per-site already drifted once (the judge-waiting return
|
||||
// omitted the prose/waiting flags); a site that forgets a must-stay-false
|
||||
// flag makes `obs.flag ?? false` negative assertions pass vacuously.
|
||||
const highWaterFlags = () => ({
|
||||
proseAUQEverObserved,
|
||||
waitingEverObserved,
|
||||
scopeGateQuestionObserved,
|
||||
scopeGateAutoSelectObserved,
|
||||
...(opts.trackTokens?.length ? { tokensObserved } : {}),
|
||||
});
|
||||
const JUDGE_AFTER_MS = 60_000;
|
||||
const JUDGE_INTERVAL_MS = 30_000;
|
||||
while (Date.now() - start < budgetMs) {
|
||||
@@ -1631,6 +1731,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
summary: `claude exited (code=${session.exitCode()}) before reaching a terminal outcome`,
|
||||
evidence: visible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
}
|
||||
if (visible.includes('Unknown command:')) {
|
||||
@@ -1639,6 +1740,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
summary: `claude rejected /${opts.skillName} as unknown command (skill not registered in this cwd)`,
|
||||
evidence: visible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1652,6 +1754,18 @@ export async function runPlanSkillObservation(opts: {
|
||||
tag: 'prose-auq-surfaced',
|
||||
});
|
||||
}
|
||||
// Scope-gate render tracking (same high-water shape). Full-run
|
||||
// detection matters because the 2KB evidence tail usually scrolls
|
||||
// past the gate render before the outcome fires.
|
||||
if (!scopeGateQuestionObserved && isScopeGateQuestionVisible(visible)) {
|
||||
scopeGateQuestionObserved = true;
|
||||
}
|
||||
if (!scopeGateAutoSelectObserved && isScopeGateAutoSelectVisible(visible)) {
|
||||
scopeGateAutoSelectObserved = true;
|
||||
}
|
||||
for (const t of opts.trackTokens ?? []) {
|
||||
if (!tokensObserved[t] && visible.includes(t)) tokensObserved[t] = true;
|
||||
}
|
||||
|
||||
const classified = classifyVisible(visible, {
|
||||
strictPlanWrites: !!opts.initialPlanContent,
|
||||
@@ -1661,8 +1775,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
...classified,
|
||||
evidence: visible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
proseAUQEverObserved,
|
||||
waitingEverObserved,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
// Capture the plan file path on any outcome where one may have been
|
||||
// written. Gating only on 'plan_ready' missed two cases: (1) the
|
||||
@@ -1693,6 +1806,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
summary: `LLM judge: ${lastJudgeVerdict.reasoning} (state=waiting after ${Math.round(elapsed / 1000)}s)`,
|
||||
evidence: visible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -1714,8 +1828,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
: ''),
|
||||
evidence: finalVisible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
proseAUQEverObserved,
|
||||
waitingEverObserved,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
}
|
||||
return {
|
||||
@@ -1727,8 +1840,7 @@ export async function runPlanSkillObservation(opts: {
|
||||
: ''),
|
||||
evidence: finalVisible.slice(-2000),
|
||||
elapsedMs: Date.now() - startedAt,
|
||||
proseAUQEverObserved,
|
||||
waitingEverObserved,
|
||||
...highWaterFlags(),
|
||||
};
|
||||
} finally {
|
||||
await session.close();
|
||||
@@ -2099,11 +2211,23 @@ export async function runPlanSkillFloorCheck(opts: {
|
||||
const start = Date.now();
|
||||
let lastJudgeAt = 0;
|
||||
let lastJudgeVerdict: PtyStateVerdict | null = null;
|
||||
// Positional anchor for the scope-gate exclusion. The visible buffer is
|
||||
// append-only (old renders never leave scrollback), so a gate question
|
||||
// rendered in the 3s pre-target window would keep satisfying the
|
||||
// full-buffer acceptance checks forever while a tail-only exclusion
|
||||
// stops seeing it after ~TAIL_SCAN_BYTES of output — a vacuous
|
||||
// auq_observed (found independently by 4 review passes). Once the gate
|
||||
// render is seen, acceptance only counts AUQ renders in content APPENDED
|
||||
// after that point.
|
||||
let gateSeenIdx = -1;
|
||||
const JUDGE_AFTER_MS = 60_000;
|
||||
const JUDGE_INTERVAL_MS = 30_000;
|
||||
while (Date.now() - start < timeoutMs) {
|
||||
await Bun.sleep(2000);
|
||||
const visible = session.visibleSince(since);
|
||||
if (gateSeenIdx === -1 && isScopeGateQuestionVisible(visible)) {
|
||||
gateSeenIdx = visible.length;
|
||||
}
|
||||
|
||||
if (session.exited()) {
|
||||
return {
|
||||
@@ -2129,10 +2253,34 @@ export async function runPlanSkillFloorCheck(opts: {
|
||||
// OR via prose-rendered options under --disallowedTools when no MCP
|
||||
// variant is callable (isProseAUQVisible). Both surface the question
|
||||
// to the user; the bug we're catching is "fired zero AUQs."
|
||||
//
|
||||
// Scope-gate renders do NOT count: the gate's "What should I review?"
|
||||
// can fire inside the 3s pre-target window and would trivially satisfy
|
||||
// the floor, but the floor measures FINDING-driven questions. Once a
|
||||
// gate render has been seen, acceptance scans only the content APPENDED
|
||||
// after it (positional anchor above) — the buffer is append-only, so a
|
||||
// whole-buffer acceptance would keep matching the stale gate render
|
||||
// forever.
|
||||
//
|
||||
// The gate veto is ACTIVE-RENDER-aware, not blanket-tail: when a
|
||||
// numbered menu is up, parseNumberedOptions anchors on the LAST cursor
|
||||
// line, so we veto only when the pending menu IS the gate — a finding
|
||||
// AUQ that renders within TAIL_SCAN_BYTES of the gate (model waiting,
|
||||
// no further output) still satisfies the floor. Prose renders have no
|
||||
// cursor anchor, so the prose path falls back to the tail check
|
||||
// (accepted residual: prose gate + prose finding inside one tail can
|
||||
// suppress until timeout; floors run the native-menu path in practice).
|
||||
const tail = visible.slice(-TAIL_SCAN_BYTES);
|
||||
const acceptWindow = gateSeenIdx === -1 ? visible : visible.slice(gateSeenIdx);
|
||||
const activeMenu = parseNumberedOptions(visible);
|
||||
const gateIsActiveRender =
|
||||
activeMenu.length > 0
|
||||
? activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label))
|
||||
: isScopeGateQuestionVisible(tail);
|
||||
if (
|
||||
(isNumberedOptionListVisible(visible) || isProseAUQVisible(visible)) &&
|
||||
!isPermissionDialogVisible(tail)
|
||||
(isNumberedOptionListVisible(acceptWindow) || isProseAUQVisible(acceptWindow)) &&
|
||||
!isPermissionDialogVisible(tail) &&
|
||||
!gateIsActiveRender
|
||||
) {
|
||||
return {
|
||||
auqObserved: true,
|
||||
@@ -2154,7 +2302,11 @@ export async function runPlanSkillFloorCheck(opts: {
|
||||
lastJudgeAt = Date.now();
|
||||
logPtySnapshot(visible, { testName: opts.skillName, elapsedMs: elapsed, tag: 'floor-judge-tick' });
|
||||
lastJudgeVerdict = judgePtyState(visible, { testName: opts.skillName });
|
||||
if (lastJudgeVerdict.state === 'waiting') {
|
||||
// The judge can't tell a scope-gate question from a finding question,
|
||||
// so a 'waiting' verdict while the gate menu is the pending render
|
||||
// must NOT satisfy the floor — same active-render exclusion as the
|
||||
// regex path.
|
||||
if (lastJudgeVerdict.state === 'waiting' && !gateIsActiveRender) {
|
||||
return {
|
||||
auqObserved: true,
|
||||
outcome: 'auq_observed',
|
||||
|
||||
@@ -28,6 +28,8 @@ import {
|
||||
isPermissionDialogVisible,
|
||||
isNumberedOptionListVisible,
|
||||
isProseAUQVisible,
|
||||
isScopeGateQuestionVisible,
|
||||
isScopeGateAutoSelectVisible,
|
||||
isPlanReadyVisible,
|
||||
parseNumberedOptions,
|
||||
classifyVisible,
|
||||
@@ -194,6 +196,113 @@ describe('isNumberedOptionListVisible', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('scope-gate render detectors', () => {
|
||||
// The verbatim announcement string from the plan-eng/plan-design SKILL.md
|
||||
// templates. If the template rewording drifts, THIS fixture fails first —
|
||||
// before the paid plan-mode smokes silently degrade to vacuous asserts.
|
||||
const TEMPLATE_ANNOUNCEMENT =
|
||||
'Scope gate: plan mode — auto-selected B (reviewing <target>).';
|
||||
|
||||
describe('isScopeGateQuestionVisible', () => {
|
||||
test('matches the clean prose gate render (question + option bodies)', () => {
|
||||
const sample = `
|
||||
What should I review?
|
||||
A) The current branch diff — the work in progress on this branch.
|
||||
B) A plan or design doc I'll paste or point you to.
|
||||
C) A specific file, directory, or path.
|
||||
Recommendation: A when a branch diff exists, otherwise B.
|
||||
`;
|
||||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('matches the native numbered render (no lettered markers)', () => {
|
||||
const sample = `
|
||||
What should I review?
|
||||
|
||||
❯ 1. The current branch diff — the work in progress on this branch.
|
||||
2. A plan or design doc I'll paste or point you to.
|
||||
3. A specific file, directory, or path.
|
||||
`;
|
||||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('matches the PTY-collapsed render (stripAnsi squished spaces)', () => {
|
||||
const sample = 'WhatshouldIreview?A)Thecurrentbranchdiff—theworkinprogress';
|
||||
expect(isScopeGateQuestionVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('stays false on narration quoting only the question', () => {
|
||||
const sample =
|
||||
"Normally I'd ask 'What should I review?' but plan mode is active, so I'm proceeding.";
|
||||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||||
});
|
||||
|
||||
test('stays false on unrelated review prose', () => {
|
||||
const sample = 'I will review the current branch diff and report findings.';
|
||||
expect(isScopeGateQuestionVisible(sample)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('isScopeGateAutoSelectVisible', () => {
|
||||
test('matches the verbatim template announcement', () => {
|
||||
expect(isScopeGateAutoSelectVisible(TEMPLATE_ANNOUNCEMENT)).toBe(true);
|
||||
});
|
||||
|
||||
test('matches a real announcement with a concrete target', () => {
|
||||
const sample =
|
||||
'Scope gate: plan mode — auto-selected B (reviewing ~/.claude/plans/my-feature.md). Running the Design Doc Check next.';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('matches the PTY-collapsed announcement', () => {
|
||||
const sample = 'Scopegate:planmode—auto-selectedB(reviewingPLAN.md).';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('stays false on narration about the behavior', () => {
|
||||
const sample = "In plan mode I'd auto-select B and review the active plan.";
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||||
});
|
||||
|
||||
test('stays false on a VERBATIM QUOTE of the announcement (negation narration)', () => {
|
||||
// The exact announcement line sits quoted in the skill context, so a
|
||||
// model explaining why it is NOT firing it can reproduce it byte-exact
|
||||
// inside quotes — that must not trip a must-stay-false assert.
|
||||
const sample =
|
||||
'Not in plan mode, so I won\'t announce "Scope gate: plan mode — auto-selected B (reviewing <target>)." and will ask instead.';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||||
});
|
||||
|
||||
test('a later real render still matches after an earlier quoted mention', () => {
|
||||
const sample =
|
||||
'Earlier I said I would render "Scope gate: plan mode — auto-selected B (…)" and now:\n' +
|
||||
'Scope gate: plan mode — auto-selected B (reviewing PLAN.md).';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||||
});
|
||||
|
||||
test('matches tense paraphrases WITH the announcement prefix (auto-selecting / auto-selects)', () => {
|
||||
expect(
|
||||
isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selecting B (reviewing the drafted plan).'),
|
||||
).toBe(true);
|
||||
expect(isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selects B.')).toBe(true);
|
||||
});
|
||||
|
||||
test('stays false on tense paraphrases WITHOUT the announcement prefix', () => {
|
||||
expect(isScopeGateAutoSelectVisible('Auto-selecting B since we are in plan mode.')).toBe(false);
|
||||
});
|
||||
|
||||
test('stays false on AUTO_DECIDE preamble output', () => {
|
||||
const sample = 'Auto-decided scope question → B (your preference). Change with /plan-tune.';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||||
});
|
||||
|
||||
test('stays false on a bare "selected B" without the announcement prefix', () => {
|
||||
const sample = 'I selected B as the review target.';
|
||||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe('isProseAUQVisible', () => {
|
||||
test('matches 4 lettered options A) B) C) D) at line starts (plan-eng prose AUQ shape)', () => {
|
||||
const sample = `
|
||||
|
||||
@@ -234,7 +234,10 @@ const MONOLITH_INVARIANTS: ParityInvariant[] = [
|
||||
// cross-session decision-memory nudge) lands this skill just over the strict 1.05;
|
||||
// headroom for the shared preamble additions (matches the carved-skill overrides).
|
||||
// v1.2.0 activation lift adds the first-run-guidance section on top.
|
||||
maxSizeRatio: 1.09,
|
||||
// 1.09 → 1.10: the plan-mode preamble reword (scope-gate auto-select-B
|
||||
// change) adds ~250 B to every skill's shared preamble; investigate was
|
||||
// the closest to its ceiling (landed 1.092).
|
||||
maxSizeRatio: 1.10,
|
||||
minBytes: 30_000,
|
||||
},
|
||||
{
|
||||
|
||||
@@ -98,11 +98,17 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
|
||||
// include question-tuning.ts and generate-ask-user-format.ts because the
|
||||
// AUTO_DECIDE preamble injection lives there and changes can flip the
|
||||
// regression test outcome between 'asked' and 'auto_decided'.
|
||||
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
|
||||
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
|
||||
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
|
||||
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-ceo-plan-mode.test.ts'],
|
||||
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-eng-plan-mode.test.ts'],
|
||||
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts'],
|
||||
'plan-devex-review-plan-mode': ['plan-devex-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
|
||||
'plan-mode-no-op': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts'],
|
||||
// Covers ceo (preamble misfire) + eng/design (scope-gate bypass must not
|
||||
// fire outside plan mode) + the named-target exception case. 4 PTY runs;
|
||||
// in CI these run CONCURRENT with the rest of the pty-plan-smoke suite
|
||||
// (--max-concurrency + --retry 2), so worst-case cost is ~3x a single
|
||||
// pass of each, sharing the API budget with sibling tests — not the
|
||||
// sequential ~+10min a local read suggests.
|
||||
'plan-mode-no-op': ['plan-ceo-review/**', 'plan-eng-review/**', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-mode-no-op.test.ts'],
|
||||
|
||||
// v1.21+ AskUserQuestion-blocked regression tests — Conductor launches
|
||||
// claude with `--disallowedTools AskUserQuestion --permission-mode default`
|
||||
|
||||
@@ -176,7 +176,13 @@ Include: color trends, typography patterns, and layout conventions you observed.
|
||||
Do NOT generate a full DESIGN.md — just research notes.`,
|
||||
workingDirectory: researchDir,
|
||||
maxTurns: 8,
|
||||
timeout: 90_000,
|
||||
// 300s, not 90s: saturated-runner class (same as review-dashboard-via /
|
||||
// retro-base-branch). PR #2533 CI observed the sibling preview test at
|
||||
// 0 turns/$0.00 for 93s x3 attempts — session up, first completion
|
||||
// queued past the budget under concurrent API load. 90s budgets cannot
|
||||
// absorb one slow first completion; 300s is the repo's standard floor
|
||||
// for CI SDK tests. Outer timeout below rises to 360s for headroom.
|
||||
timeout: 300_000,
|
||||
testName: 'design-consultation-research',
|
||||
runId,
|
||||
});
|
||||
@@ -206,7 +212,7 @@ Do NOT generate a full DESIGN.md — just research notes.`,
|
||||
}
|
||||
|
||||
try { fs.rmSync(researchDir, { recursive: true, force: true }); } catch {}
|
||||
}, 120_000);
|
||||
}, 360_000);
|
||||
|
||||
testConcurrentIfSelected('design-consultation-existing', async () => {
|
||||
// Pre-create a minimal DESIGN.md (independent of core test)
|
||||
@@ -274,7 +280,9 @@ Write a single HTML file to ${previewDir}/design-preview.html that shows:
|
||||
Do NOT write DESIGN.md — only the preview HTML.`,
|
||||
workingDirectory: previewDir,
|
||||
maxTurns: 8,
|
||||
timeout: 90_000,
|
||||
// 300s, not 90s: this is the test that failed 3x at 0 turns/$0.00/93s
|
||||
// on PR #2533 CI — see the research test's comment for the class.
|
||||
timeout: 300_000,
|
||||
testName: 'design-consultation-preview',
|
||||
runId,
|
||||
});
|
||||
@@ -303,7 +311,7 @@ Do NOT write DESIGN.md — only the preview HTML.`,
|
||||
}
|
||||
|
||||
try { fs.rmSync(previewDir, { recursive: true, force: true }); } catch {}
|
||||
}, 120_000);
|
||||
}, 360_000);
|
||||
});
|
||||
|
||||
// --- Plan Design Review E2E (plan-mode) ---
|
||||
|
||||
@@ -47,7 +47,12 @@ describeE2E('plan-ceo-review plan-mode smoke (gate)', () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName: 'plan-ceo-review',
|
||||
inPlanMode: true,
|
||||
timeoutMs: 300_000,
|
||||
// 420s, not 300s: measured 2026-08-11, a clean isolated pass took
|
||||
// 295.7s (80s on a quiet main run) — 4s under the old budget — and the
|
||||
// same run timed out at ~308s three times under concurrent eval load.
|
||||
// Same runner-contention class as review-dashboard-via/retro-base-
|
||||
// branch; headroom instead of a budget-edge flake in the gate lane.
|
||||
timeoutMs: 420_000,
|
||||
env: { QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' },
|
||||
});
|
||||
|
||||
@@ -72,5 +77,5 @@ describeE2E('plan-ceo-review plan-mode smoke (gate)', () => {
|
||||
);
|
||||
}
|
||||
assertReportAtBottomIfPlanWritten(obs);
|
||||
}, 360_000);
|
||||
}, 480_000);
|
||||
});
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* /plan-design-review AskUserQuestion floor regression (gate, paid, real-PTY).
|
||||
* /plan-design-review AskUserQuestion floor regression (periodic, paid, real-PTY).
|
||||
*
|
||||
* See test/skill-e2e-plan-eng-finding-floor.test.ts for the contract.
|
||||
*/
|
||||
@@ -8,10 +8,10 @@ import { describe, test } from 'bun:test';
|
||||
import { runPlanSkillFloorCheck } from './helpers/claude-pty-runner';
|
||||
import { FORCING_FLOOR_DESIGN } from './fixtures/forcing-finding-seeds';
|
||||
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
|
||||
const describeE2E = shouldRun ? describe : describe.skip;
|
||||
|
||||
describeE2E('/plan-design-review AskUserQuestion floor (gate)', () => {
|
||||
describeE2E('/plan-design-review AskUserQuestion floor (periodic)', () => {
|
||||
test(
|
||||
'seeded forcing finding causes the agent to fire at least one AskUserQuestion',
|
||||
async () => {
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* plan-design-review plan-mode smoke (gate, paid, real-PTY).
|
||||
* plan-design-review plan-mode smoke (periodic, paid, real-PTY).
|
||||
*
|
||||
* See test/skill-e2e-plan-ceo-plan-mode.test.ts for the shared assertion
|
||||
* contract. Exercises the same contract against /plan-design-review.
|
||||
@@ -15,10 +15,33 @@ import {
|
||||
assertReportAtBottomIfPlanWritten,
|
||||
} from './helpers/claude-pty-runner';
|
||||
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
|
||||
const describeE2E = shouldRun ? describe : describe.skip;
|
||||
|
||||
describeE2E('plan-design-review plan-mode smoke (gate)', () => {
|
||||
// UI-heavy seed with guaranteed design gaps (center-aligned everything, no
|
||||
// empty states, no responsive intent) so the review has real findings to
|
||||
// surface. Inline twin of the eng smoke's SEED_PLAN_FORCING_FINDINGS —
|
||||
// FORCING_FLOOR_DESIGN from forcing-finding-seeds.ts is NOT reusable here:
|
||||
// it embeds a write-to-/tmp instruction shaped for the floor check's
|
||||
// followUpPrompt, which would trip strictPlanWrites as a silent_write.
|
||||
const SEED_PLAN_UI_HEAVY = `
|
||||
# Plan: Marketing landing page
|
||||
|
||||
## Layout
|
||||
All headings, taglines, and body copy will be center-aligned for a
|
||||
"clean modern look." The hero h1 sits 8px above the subhead; the CTA
|
||||
button has the same visual weight as the "Learn more" link beside it.
|
||||
|
||||
## Pages
|
||||
- / (hero, 3-column features grid, testimonials carousel, footer)
|
||||
- /pricing (3 tier cards)
|
||||
|
||||
## States
|
||||
Only the happy path is designed. No empty states, no error states,
|
||||
no loading states. Mobile: "stacks on mobile."
|
||||
`;
|
||||
|
||||
describeE2E('plan-design-review plan-mode smoke (periodic)', () => {
|
||||
test('reaches a terminal outcome (asked or plan_ready) without silent writes', async () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName: 'plan-design-review',
|
||||
@@ -37,4 +60,40 @@ describeE2E('plan-design-review plan-mode smoke (gate)', () => {
|
||||
expect(['asked', 'plan_ready']).toContain(obs.outcome);
|
||||
assertReportAtBottomIfPlanWritten(obs);
|
||||
}, 360_000);
|
||||
|
||||
// Plan-mode scope-gate bypass: with a seeded UI-heavy plan in plan mode,
|
||||
// the gate must NOT render its "What should I review?" menu — it
|
||||
// auto-selects B and announces it, then proceeds to the pre-review audit
|
||||
// and mockups. Mirrors the eng smoke's seeded STOP-gate test, without
|
||||
// --disallowedTools (native AUQ available is the common path here).
|
||||
test('scope gate auto-selects B when a plan is seeded in plan mode', async () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName: 'plan-design-review',
|
||||
inPlanMode: true,
|
||||
initialPlanContent: SEED_PLAN_UI_HEAVY,
|
||||
timeoutMs: 300_000,
|
||||
});
|
||||
|
||||
if (
|
||||
obs.outcome === 'wrote_findings_before_asking' ||
|
||||
obs.outcome === 'auto_decided' ||
|
||||
obs.outcome === 'silent_write' ||
|
||||
obs.outcome === 'exited' ||
|
||||
obs.outcome === 'timeout'
|
||||
) {
|
||||
throw new Error(
|
||||
`plan-design plan-mode bypass FAILED: outcome=${obs.outcome}\n` +
|
||||
`summary: ${obs.summary}\nelapsed: ${obs.elapsedMs}ms\n` +
|
||||
`--- evidence (last 2KB) ---\n${obs.evidence}`,
|
||||
);
|
||||
}
|
||||
|
||||
expect(['asked', 'plan_ready']).toContain(obs.outcome);
|
||||
assertReportAtBottomIfPlanWritten(obs);
|
||||
|
||||
// The bypass contract (exception ordering makes this deterministic even
|
||||
// though the seed arrives as a pasted user message).
|
||||
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
|
||||
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(true);
|
||||
}, 360_000);
|
||||
});
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* /plan-eng-review AskUserQuestion floor regression (gate, paid, real-PTY).
|
||||
* /plan-eng-review AskUserQuestion floor regression (periodic, paid, real-PTY).
|
||||
*
|
||||
* Catches the May 2026 transcript bug where /plan-eng-review wrote a
|
||||
* multi-section review plan to ~/.claude/plans/ and called ExitPlanMode
|
||||
@@ -11,7 +11,7 @@
|
||||
* render. See claude-pty-runner.ts for why this is separate from the
|
||||
* runPlanSkillCounting harness used by periodic finding-count tests.
|
||||
*
|
||||
* Tier: gate. Budget: 10 min (early exit on success ~30-90s typical).
|
||||
* Tier: periodic. Budget: 10 min (early exit on success ~30-90s typical).
|
||||
* Cost: ~$0.50-$1.50 per run depending on early-exit timing.
|
||||
*/
|
||||
|
||||
@@ -19,10 +19,10 @@ import { describe, test } from 'bun:test';
|
||||
import { runPlanSkillFloorCheck } from './helpers/claude-pty-runner';
|
||||
import { FORCING_FLOOR_ENG } from './fixtures/forcing-finding-seeds';
|
||||
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
|
||||
const describeE2E = shouldRun ? describe : describe.skip;
|
||||
|
||||
describeE2E('/plan-eng-review AskUserQuestion floor (gate)', () => {
|
||||
describeE2E('/plan-eng-review AskUserQuestion floor (periodic)', () => {
|
||||
test(
|
||||
'seeded forcing finding causes the agent to fire at least one AskUserQuestion',
|
||||
async () => {
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* plan-eng-review plan-mode smoke (gate, paid, real-PTY).
|
||||
* plan-eng-review plan-mode smoke (periodic, paid, real-PTY).
|
||||
*
|
||||
* See test/skill-e2e-plan-ceo-plan-mode.test.ts for the shared assertion
|
||||
* contract. This file exercises the same contract against /plan-eng-review.
|
||||
@@ -12,7 +12,7 @@ import {
|
||||
assertReportAtBottomIfPlanWritten,
|
||||
} from './helpers/claude-pty-runner';
|
||||
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
|
||||
const describeE2E = shouldRun ? describe : describe.skip;
|
||||
|
||||
// SEED_PLAN_FORCING_FINDINGS: 8+ files + custom-vs-builtin smell forces the
|
||||
@@ -45,7 +45,7 @@ Ignore Bun's native --shard flag because we want full control.
|
||||
None planned — will add later.
|
||||
`;
|
||||
|
||||
describeE2E('plan-eng-review plan-mode smoke (gate)', () => {
|
||||
describeE2E('plan-eng-review plan-mode smoke (periodic)', () => {
|
||||
test('reaches a terminal outcome (asked or plan_ready) without silent writes', async () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName: 'plan-eng-review',
|
||||
@@ -108,5 +108,15 @@ describeE2E('plan-eng-review plan-mode smoke (gate)', () => {
|
||||
|
||||
expect(['asked', 'plan_ready']).toContain(obs.outcome);
|
||||
assertReportAtBottomIfPlanWritten(obs);
|
||||
|
||||
// Plan-mode scope-gate bypass: with a seeded plan in plan mode, the gate
|
||||
// must NOT render its "What should I review?" menu — it auto-selects B
|
||||
// and announces it. Exception ordering in the template (plan-mode branch
|
||||
// first) makes this deterministic even though the seed arrives as a
|
||||
// pasted user message. Unseeded test 1 keeps its lenient contract: with
|
||||
// no plan drafted, the "ask as normal" fallback legitimately renders the
|
||||
// question.
|
||||
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
|
||||
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(true);
|
||||
}, 360_000);
|
||||
});
|
||||
|
||||
@@ -1,17 +1,33 @@
|
||||
/**
|
||||
* Plan-mode-info no-op regression (gate tier, paid, real-PTY).
|
||||
*
|
||||
* Asserts: when /plan-ceo-review is invoked OUTSIDE plan mode (no
|
||||
* Asserts: when a plan-review skill is invoked OUTSIDE plan mode (no
|
||||
* --permission-mode plan flag, no plan-mode reminder injected), the skill
|
||||
* still reaches a terminal outcome ('asked' or 'plan_ready'). This is the
|
||||
* negative coverage to the per-skill plan-mode smokes — if the
|
||||
* plan-mode-info preamble section ever starts misfiring for non-plan-mode
|
||||
* sessions (e.g., gating questions on a phrase that isn't there), this
|
||||
* test catches it.
|
||||
* negative coverage to the per-skill plan-mode smokes — if plan-mode-keyed
|
||||
* behavior ever starts misfiring for non-plan-mode sessions (e.g., gating
|
||||
* questions on a phrase that isn't there, or the plan-eng/plan-design
|
||||
* scope-gate auto-select-B bypass firing without plan mode), this test
|
||||
* catches it.
|
||||
*
|
||||
* Why this matters: outside plan mode, claude doesn't render a native
|
||||
* confirmation UI. The skill must drive its own AskUserQuestion. Same
|
||||
* runner, same outcome contract — just `inPlanMode: false`.
|
||||
*
|
||||
* Coverage grew with the scope-gate bypass (plan-mode auto-select B):
|
||||
* - plan-ceo-review: original preamble-misfire regression.
|
||||
* - plan-eng-review / plan-design-review: the bypass must NOT fire outside
|
||||
* plan mode (scopeGateAutoSelectObserved stays false), and when the run
|
||||
* ends in 'asked', the question that fired must be the scope gate itself
|
||||
* (outside plan mode with no named target, the gate is the FIRST
|
||||
* question by contract).
|
||||
* - named-target case: a pasted draft (initialPlanContent) IS an
|
||||
* explicitly-named target, so the gate must NOT ask — and the review
|
||||
* must actually consume the pasted content.
|
||||
*
|
||||
* Cost note: 4 sequential PTY runs (~3-5 min each) in the gate lane, up
|
||||
* from 1 pre-bypass. Selected only when plan-ceo/eng/design or the runner
|
||||
* change (see 'plan-mode-no-op' in touchfiles.ts).
|
||||
*/
|
||||
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
@@ -20,29 +36,114 @@ import { runPlanSkillObservation } from './helpers/claude-pty-runner';
|
||||
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
|
||||
const describeE2E = shouldRun ? describe : describe.skip;
|
||||
|
||||
const PLAN_MODE_REMINDER =
|
||||
'Plan mode is active. The user indicated that they do not want you to execute yet';
|
||||
|
||||
// Distinctive token proves the pasted target was consumed by the review —
|
||||
// not just that no question fired. Nonsense-unique so it can't appear by
|
||||
// coincidence in skill output.
|
||||
const SEED_TOKEN = 'ZephyrLedgerWidget';
|
||||
const NAMED_TARGET_SEED = `
|
||||
# Plan: ${SEED_TOKEN} settings panel
|
||||
|
||||
## Scope
|
||||
Add a ${SEED_TOKEN} settings panel with a single toggle that enables
|
||||
weekly export emails. One new component, one route, one test file.
|
||||
|
||||
## Files
|
||||
- src/components/${SEED_TOKEN}.tsx (new)
|
||||
- src/routes/settings.tsx (add panel)
|
||||
- test/${SEED_TOKEN}.test.tsx (new)
|
||||
`;
|
||||
|
||||
describeE2E('plan-mode-info no-op outside plan mode (gate regression)', () => {
|
||||
test('skill reaches a terminal outcome outside plan mode', async () => {
|
||||
for (const skillName of ['plan-ceo-review', 'plan-eng-review', 'plan-design-review'] as const) {
|
||||
test(`${skillName} reaches a terminal outcome outside plan mode`, async () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName,
|
||||
inPlanMode: false,
|
||||
timeoutMs: 300_000,
|
||||
// eng/design: force the prose-fallback path. The unconditional
|
||||
// gate-must-ask assert below pins the render shape the detector
|
||||
// anchors on, and only the --disallowedTools prose fallback makes
|
||||
// that shape CONTRACTUAL ("use exactly this shape" in the template);
|
||||
// native AskUserQuestion could render terse option labels that a
|
||||
// correct run would fail on (red-team finding).
|
||||
...(skillName === 'plan-ceo-review'
|
||||
? {}
|
||||
: { extraArgs: ['--disallowedTools', 'AskUserQuestion'] }),
|
||||
});
|
||||
|
||||
if (obs.outcome === 'silent_write' || obs.outcome === 'exited' || obs.outcome === 'timeout') {
|
||||
throw new Error(
|
||||
`plan-mode no-op regression FAILED (${skillName}): outcome=${obs.outcome}\n` +
|
||||
`summary: ${obs.summary}\n` +
|
||||
`elapsed: ${obs.elapsedMs}ms\n` +
|
||||
`--- evidence (last 2KB visible) ---\n${obs.evidence}`,
|
||||
);
|
||||
}
|
||||
expect(['asked', 'plan_ready']).toContain(obs.outcome);
|
||||
|
||||
// Negative regression: the rendered output must NOT echo the plan-mode
|
||||
// distinctive reminder phrase. If it does, the plan-mode preamble
|
||||
// section is leaking outside plan mode.
|
||||
expect(obs.evidence).not.toContain(PLAN_MODE_REMINDER);
|
||||
|
||||
if (skillName !== 'plan-ceo-review') {
|
||||
// Scope-gate bypass must not misfire: no auto-select announcement
|
||||
// outside plan mode.
|
||||
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(false);
|
||||
// UNCONDITIONAL: outside plan mode with no named target, the gate is
|
||||
// a hard STOP before any tool call, so the gate question must have
|
||||
// rendered no matter which terminal outcome fired. Gating this on
|
||||
// outcome === 'asked' would let a silent-bypass run that reaches
|
||||
// plan_ready (isPlanReadyVisible also matches common prose) sail
|
||||
// through — the exact regression this test exists to catch.
|
||||
expect(obs.scopeGateQuestionObserved ?? false).toBe(true);
|
||||
}
|
||||
}, 360_000);
|
||||
}
|
||||
|
||||
// Named-target exception (outside plan mode): a pasted draft IS an
|
||||
// explicitly-named target, so the scope gate must NOT ask — and the
|
||||
// review must consume the pasted content (seed token visible in the
|
||||
// review output), proving the target was used rather than the question
|
||||
// merely skipped. Also the over-trigger guard for the tightened
|
||||
// "explicit-only" exception wording.
|
||||
test('plan-eng-review skips the scope gate for an explicitly-pasted target', async () => {
|
||||
const obs = await runPlanSkillObservation({
|
||||
skillName: 'plan-ceo-review',
|
||||
skillName: 'plan-eng-review',
|
||||
inPlanMode: false,
|
||||
initialPlanContent: NAMED_TARGET_SEED,
|
||||
trackTokens: [SEED_TOKEN],
|
||||
timeoutMs: 300_000,
|
||||
});
|
||||
|
||||
if (obs.outcome === 'silent_write' || obs.outcome === 'exited' || obs.outcome === 'timeout') {
|
||||
if (
|
||||
obs.outcome === 'wrote_findings_before_asking' ||
|
||||
obs.outcome === 'silent_write' ||
|
||||
obs.outcome === 'exited' ||
|
||||
obs.outcome === 'timeout'
|
||||
) {
|
||||
throw new Error(
|
||||
`plan-mode no-op regression FAILED: outcome=${obs.outcome}\n` +
|
||||
`named-target no-op FAILED: outcome=${obs.outcome}\n` +
|
||||
`summary: ${obs.summary}\n` +
|
||||
`elapsed: ${obs.elapsedMs}ms\n` +
|
||||
`--- evidence (last 2KB visible) ---\n${obs.evidence}`,
|
||||
);
|
||||
}
|
||||
expect(['asked', 'plan_ready']).toContain(obs.outcome);
|
||||
|
||||
// Negative regression: the rendered output must NOT echo the plan-mode
|
||||
// distinctive reminder phrase. If it does, the plan-mode preamble
|
||||
// section is leaking outside plan mode.
|
||||
const PLAN_MODE_REMINDER =
|
||||
'Plan mode is active. The user indicated that they do not want you to execute yet';
|
||||
expect(obs.evidence).not.toContain(PLAN_MODE_REMINDER);
|
||||
|
||||
// The pasted doc is the named target: gate question must not render,
|
||||
// no plan-mode announcement either (we are NOT in plan mode).
|
||||
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
|
||||
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(false);
|
||||
|
||||
// Target consumption via high-water token tracking over the CUMULATIVE
|
||||
// buffer — the 2KB evidence tail is lossy and the plan-file fallback is
|
||||
// unreachable outside plan mode (extractPlanFilePath only matches
|
||||
// plan-mode save renders).
|
||||
expect(obs.tokensObserved?.[SEED_TOKEN] ?? false).toBe(true);
|
||||
}, 360_000);
|
||||
});
|
||||
|
||||
@@ -134,6 +134,15 @@ describe('SKILL.md command validation', () => {
|
||||
const result = validateSkill(skill);
|
||||
expect(result.snapshotFlagErrors).toHaveLength(0);
|
||||
});
|
||||
|
||||
test('autoplan section skip list includes the scope gate', () => {
|
||||
// autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md
|
||||
// verbatim; without this skip-list entry it ingests their scope gate — a
|
||||
// hard-STOP AskUserQuestion that contradicts autoplan's auto-decide
|
||||
// contract. Nothing else pins the skip-list contents.
|
||||
const md = fs.readFileSync(path.join(ROOT, 'autoplan', 'SKILL.md'), 'utf-8');
|
||||
expect(md).toContain('- Scope gate (the plan under review is already the target)');
|
||||
});
|
||||
});
|
||||
|
||||
describe('Command registry consistency', () => {
|
||||
|
||||
Reference in New Issue
Block a user