v1.62.0.0 feat: plan-mode auto-select at the review scope gate (#2533)

* fix(evals): align plan-eng/design plan-mode + finding-floor smokes to their declared periodic tier

The #2077 demotion of these four stochastic tests to 'periodic' was inert:
E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate',
so they kept running in the blocking gate lane and never in the weekly lane.

Flip the four self-gates to 'periodic' (headers/describe labels updated), add
a free static tier-alignment invariant test (dep-list filename mapping;
unmapped self-gated files are reported, never silently skipped), and name the
two plan-mode test files in their own touchfiles dep lists so the invariant
binds for them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(pty-runner): scope-gate question/auto-select detectors + observation flags

Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5
collapsed-form handling): isScopeGateQuestionVisible requires the question text
PLUS option-body text (native AskUserQuestion renders numbered options, prose
fallback renders lettered — the option body appears in both; narration doesn't),
and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the
selected-B token.

runPlanSkillObservation gains scopeGateQuestionObserved /
scopeGateAutoSelectObserved high-water flags (attached at every return path) so
paid smokes can assert gate behavior across the whole run instead of the lossy
2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render
toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven
questions, and the gate could fire inside the 3s pre-target window.

Unit fixtures pin clean/native/collapsed positives, narration negatives, and
the verbatim template announcement string (template rewording fails here first,
before the paid smokes degrade to vacuous asserts).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(plan-eng/design-review): auto-select B in plan mode at the scope gate

In plan mode the scope gate's "What should I review? A/B/C" question is pure
friction: there is no branch diff and the target is the plan being drafted.
Both gates gain an ordered exceptions block, checked BEFORE asking:

1. Plan mode → auto-select B: review the active plan (in context or pasted),
   announce it in one line ("Scope gate: plan mode — auto-selected B
   (reviewing <target>)") so the user can interrupt; an explicitly different
   user-named target still wins; no plan drafted yet → ask as normal.
2. User-named target (outside plan mode): explicit-only — a path, a pasted
   doc, or the literal words "branch diff". A passing mention is not naming;
   when in doubt, ask.

Outside plan mode with no explicitly-named target, nothing changes. Plan-mode
is checked FIRST because the PTY harness seeds drafts as pasted user messages
(claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic.

Pinning: seeded plan-mode smokes assert no gate render + announcement rendered
(eng test 2; new design seeded test); plan-mode-no-op extends to eng/design
(bypass must not misfire outside plan mode; first question must be the gate)
plus a named-target case proving the pasted target is consumed; a drift-guard
asserts the two hand-duplicated exceptions blocks stay identical modulo the
two variant slots and carry the announcement string the detectors pin.

Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union
ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(autoplan): skip the scope gate when following loaded review skills

autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim,
and its section skip list omitted the scope gate — so autoplan ingested a
hard-STOP AskUserQuestion that contradicts its every-question-auto-decides
contract. One skip-list line fixes it; a static toContain pin in
skill-validation keeps the entry load-bearing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: file scope-gate resolver-extraction TODO (eng-review D5 follow-up)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pty-runner): positional floor exclusion, flag builder, outcome union, token tracking

Review-army + adversarial findings on the scope-gate observability work,
all verified before fixing:

- Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate
  exclusion scanned only the 1500-byte tail, so an early gate render satisfied
  the floor vacuously once ~1.5KB of output accumulated (found independently
  by 4 review passes; predicate reproduced). Acceptance now scans only content
  APPENDED after the first gate render (positional anchor), and the LLM-judge
  'waiting' shortcut no longer fires while the gate menu is the pending render.
- High-water flags are built once and spread at every return path — the
  hand-spread pattern had already drifted (judge-waiting return omitted two
  flags), which made must-stay-false asserts vacuous on those paths.
- isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects
  token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases)
  and quoted-occurrence rejection (a model verbatim-quoting the announcement
  while declining must not trip must-stay-FALSE asserts). Fixtures added for
  both directions.
- PlanSkillObservation outcome union gains 'wrote_findings_before_asking'
  (returned at runtime via classifyVisible but missing from the type).
- trackTokens/tokensObserved: cumulative-buffer token high-water for
  consumption asserts (the 2KB evidence tail is lossy and the plan-file
  fallback is unreachable outside plan mode).
- New scope-gate-floor unit pins (from the ship coverage audit): both gate
  render forms trip acceptance and exclusion; a genuine finding AUQ is not
  excluded; tail-scoping semantics pinned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(evals): harden no-op asserts, close tier-invariant fail-open holes, pin gate question strings

- no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the
  outcome==='asked' conditional let a silent-bypass plan_ready run sail
  through); eng/design cases force --disallowedTools so the pinned prose
  shape is contractual rather than hoping native AUQ renders match; the
  named-target case uses trackTokens for consumption and lists
  wrote_findings_before_asking in its diagnostic throw branch.
- tier-alignment invariant: both quote styles matched; zero-self-gate,
  mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED
  instead of silently skipped (the fail-open holes three reviewers found).
- drift-guard: the generated gate menus must carry the exact question/option
  strings the PTY question detector anchors on — free CI fails before the
  paid smokes can go vacuous on a menu reword.
- touchfiles: corrected the no-op cost note for CI concurrency + retry
  semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): register plan-eng/design-review skills in PTY eval containers

The extended plan-mode-no-op smoke invokes /plan-eng-review and
/plan-design-review, but the fresh CI containers registered only
office-hours and plan-ceo-review — both new runs would return
'Unknown command' and fail every PR's gate job (Codex structured
review P1, verified against evals.yml). Registration loops, the
dangling-target fail-fast list, and the frontmatter checks (now a
loop over the same skill list, so the lists can't drift) all cover
the two skills.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(plan-eng/design-review): harden scope-gate exceptions against injection and ambiguity

Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex
cross-confirmation), applied to both gate templates + regen:

- Host-anchored mode signal: only the host's own system messages (plan-mode
  reminder or active plan file path) arm the auto-select; plan-shaped text
  inside pasted documents, tool results, or fetched pages does NOT count —
  injected content can't disarm the consent gate or nominate the target.
- Multiple plan candidates: the host-referenced plan file wins; still
  ambiguous means ask.
- The DIFFERENT-target override carries the passing-mention guard.
- Plan mode + explicitly named target + no drafted plan resolves to the
  named target instead of a contradictory re-ask.
- The numbered ask-path rules are qualified ('When no exception above
  applied:') so they no longer restate an unconditional MUST-ask that
  contradicts the exceptions.
- 'Whenever this gate does ask — in any mode — it is a hard STOP.'
- Shared preamble: 'any AskUserQuestion the skill fires is the workflow
  operating within plan mode' (was 'the first AskUserQuestion is the
  workflow entering plan mode', which framed the opposite of the bypass);
  regenerates every skill.
- Ceilings ratcheted with attribution: plan-eng union ratio 1.10,
  investigate 1.10 (the ~250B shared-preamble reword lands the
  closest-to-ceiling skill at 1.092).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.62.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pty-runner): active-render gate veto in the floor check + honest periodic-wiring docs

Codex re-review P2s on the fix wave, both verified:

- A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting,
  no further output) was vetoed by the blanket tail exclusion until timeout.
  The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last
  cursor menu, so only a pending GATE menu vetoes; the judge fallback shares
  the same check. Residual (documented): prose gate + prose finding inside
  one tail — floors run the native-menu path in practice.
- The four demoted periodic tests are not in evals-periodic.yml's explicit
  matrix (a named instance of the pre-existing periodic-orphans TODO), so
  they run locally/manually until the PTY-capable periodic job lands.
  CHANGELOG claim softened accordingly; TODO filed with the wiring recipe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.62.0.0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: apply codex doc-review fixes for v1.62.0.0

- CLAUDE.md: scope the tier-alignment invariant claim (mapped files
  enforced, unmapped files reported)
- docs/skills.md: document the plan-mode auto-select scope gate for
  /plan-eng-review and /plan-design-review
- evals.yml: fix stale comment (PTY smokes register four skills, not two)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: refresh ship golden baselines for the plan-mode preamble reword

The generate-completion-status.ts wording change ('any AskUserQuestion the
skill fires…') intentionally regenerates every SKILL.md; the byte-compare
goldens carry the generator's output and refresh with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ship): custom-hooks-path detection false-negatives on git worktrees

The pre-push guard's HOOKS_IN_GIT_DIR check compared the hooks dir against
--absolute-git-dir, which in a linked worktree is .git/worktrees/<name>
while hooks resolve to the COMMON .git/hooks — so every Conductor worktree
read as a 'custom hooks path' and the consented guard install was skipped.
Match against the resolved --git-common-dir too (with a /nonexistent
fallback so a failed resolution can't collapse the case pattern into
match-everything). Verified live: this worktree now reports yes (was no),
and the main checkout still reports yes. Goldens refreshed (--host all).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: changelog bullet for the worktree hooks-detection fix

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): give the plan-ceo plan-mode smoke real budget headroom

Measured 2026-08-11: a clean isolated pass took 295.7s against the 300s
inner budget (4s of margin) and the same test timed out at ~308s three
times under concurrent eval load — a budget-edge flake in the gate lane,
not a behavior regression (it passed isolated on both this branch and
main). Inner budget 300s -> 420s, outer bun timeout 360s -> 480s, and the
test file is now named in its own touchfiles dep list so the tier-alignment
invariant binds for it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): 300s budget floor for the two 90s design-consultation SDK tests

Root cause of PR #2533's e2e-design CI failure: design-consultation-preview
failed 3 attempts at 0 turns/$0.00/93s — the session was up but the model's
first completion queued past the 90s inner budget under concurrent API load
(11 matrix jobs; the sibling research test booted its first tool at 4s, so
this is API-side queuing, not CPU boot contention). The test was selected
only because touchfiles.ts is a global touchfile; the tested behavior is
untouched by this branch.

90s budgets cannot absorb one slow first completion. Both 90s tests in the
file move to the repo's saturated-runner standard (300s inner / 360s outer,
matching review-dashboard-via and retro-base-branch). Deliberately NOT
re-arming the runner's inner timer on first stream event: an audit found
~100 outer bun-timeout literals sized inner+30-60s that a re-arm would
silently break — the structural options are written up in TODOS.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-12 11:12:28 -07:00
committed by GitHub
co-authored by Claude Fable 5
parent 94993f7401
commit d078622b73
81 changed files with 1097 additions and 137 deletions
+91
View File
@@ -0,0 +1,91 @@
/**
* Tier-alignment invariant (free, static).
*
* Kills the "inert demotion" defect class: E2E_TIERS declares a test's tier,
* but the paid test files also self-gate on `process.env.EVALS_TIER === '<tier>'`.
* When the two disagree, the touchfiles declaration is dead metadata — the
* #2077 demotion of the plan-mode/finding-floor smokes to 'periodic' was inert
* for months because the files still gated on 'gate' and ran in the blocking
* lane on every gate run.
*
* Mapping rule (test filenames do NOT map mechanically to tier keys): for each
* `test/skill-e2e-*.test.ts` with an EVALS_TIER self-gate, search the
* E2E_TOUCHFILES / LLM_JUDGE_TOUCHFILES dep lists for the exact file path. If
* found under key K, the file's self-gate tier must equal E2E_TIERS[K]. Files
* not named in any dep list are REPORTED as unmapped (a nudge to add them to
* their eval's dep list), never silently skipped.
*/
import { describe, test, expect } from 'bun:test';
import { readdirSync, readFileSync } from 'fs';
import * as path from 'path';
import { E2E_TOUCHFILES, E2E_TIERS, LLM_JUDGE_TOUCHFILES } from './helpers/touchfiles';
const TEST_DIR = import.meta.dir;
// Both quote styles — a mechanical refactor to double quotes must not
// silently drop a file from the invariant (fail-open is the defect class
// this test exists to kill).
const SELF_GATE_RE = /EVALS_TIER\s*===\s*['"](gate|periodic)['"]/g;
describe('E2E tier alignment (touchfiles declaration vs test self-gate)', () => {
const testFiles = readdirSync(TEST_DIR)
.filter((f) => f.startsWith('skill-e2e-') && f.endsWith('.test.ts'))
.sort();
const allDeps: Record<string, string[]> = { ...E2E_TOUCHFILES, ...LLM_JUDGE_TOUCHFILES };
test('every self-gated test file named in a dep list matches its declared tier', () => {
const misaligned: string[] = [];
const reported: string[] = [];
for (const file of testFiles) {
const content = readFileSync(path.join(TEST_DIR, file), 'utf-8');
const tiers = new Set<string>();
for (const m of content.matchAll(SELF_GATE_RE)) tiers.add(m[1]);
const repoPath = `test/${file}`;
if (tiers.size === 0) {
// Every skill-e2e file is expected to self-gate; zero matches means
// either a genuinely ungated file or a gate shape the regex can't
// see — both worth a visible report, never a silent skip.
reported.push(`${repoPath}: no detectable EVALS_TIER self-gate`);
continue;
}
if (tiers.size > 1) {
reported.push(`${repoPath}: mixed-tier self-gates (${[...tiers].join(', ')}) — not tier-checked`);
continue;
}
const selfTier = [...tiers][0];
const owningKeys = Object.keys(allDeps).filter((k) => allDeps[k].includes(repoPath));
if (owningKeys.length === 0) {
reported.push(`${repoPath} (self-gates '${selfTier}'): not named in any touchfiles dep list`);
continue;
}
for (const k of owningKeys) {
const declared = E2E_TIERS[k];
if (!declared) {
// A dep-list key with no E2E_TIERS entry (e.g. an LLM-judge key)
// can't tier-check this file — report instead of silently passing.
reported.push(`${repoPath}: matched key '${k}' which has no E2E_TIERS entry`);
continue;
}
if (declared !== selfTier) {
misaligned.push(
`${repoPath}: self-gates on '${selfTier}' but E2E_TIERS['${k}'] declares '${declared}' — the declaration is inert`,
);
}
}
}
// Reported, not asserted: coverage holes the invariant can see but not
// arbitrate. Add the test file to its eval's dep list (or a tier entry
// for the key) to bring it under the invariant.
if (reported.length > 0) {
console.warn(
`[tier-alignment] ${reported.length} file(s) outside the invariant:\n ` + reported.join('\n '),
);
}
expect(misaligned).toEqual([]);
});
});
+8 -2
View File
@@ -150,7 +150,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
## Skill Invocation During Plan Mode
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
@@ -1272,9 +1272,15 @@ _HOOK_INSTALLED="no"
# committed hook and write a machine-local wrapper into the working tree.
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
# resolve to the COMMON .git/hooks, so match against the common dir too or
# every Conductor worktree false-negatives as a "custom hooks path". The
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
# (match-everything) when resolution fails.
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
_HOOKS_IN_GIT_DIR="no"
case "$_HOOKS_DIR" in
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
esac
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
+8 -2
View File
@@ -136,7 +136,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
## Skill Invocation During Plan Mode
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
@@ -2455,9 +2455,15 @@ _HOOK_INSTALLED="no"
# committed hook and write a machine-local wrapper into the working tree.
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
# resolve to the COMMON .git/hooks, so match against the common dir too or
# every Conductor worktree false-negatives as a "custom hooks path". The
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
# (match-everything) when resolution fails.
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
_HOOKS_IN_GIT_DIR="no"
case "$_HOOKS_DIR" in
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
esac
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
+8 -2
View File
@@ -138,7 +138,7 @@ In plan mode, allowed because they inform the plan: `$B`, `$D`, `codex exec`/`co
## Skill Invocation During Plan Mode
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; the first AskUserQuestion is the workflow entering plan mode, not a violation of it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If the user invokes a skill in plan mode, the skill takes precedence over generic plan mode behavior. **Treat the skill file as executable instructions, not reference.** Follow it step by step starting from Step 0; any AskUserQuestion the skill fires is the workflow operating within plan mode, not a violation of it — and a skill whose instructions resolve a question themselves (e.g. a plan-mode auto-select) may legitimately not ask it. AskUserQuestion (any variant — `mcp__*__AskUserQuestion` or native; see "AskUserQuestion Format → Tool resolution") satisfies plan mode's end-of-turn requirement. If AskUserQuestion is unavailable or a call fails, follow the AskUserQuestion Format failure fallback: `headless` → BLOCKED; `interactive` → the prose fallback (also satisfies end-of-turn). At a STOP point, stop immediately. Do not continue the workflow or call ExitPlanMode there. Commands marked "PLAN MODE EXCEPTION — ALWAYS RUN" execute. Call ExitPlanMode only after the skill workflow completes, or if the user tells you to cancel the skill or leave plan mode.
If `PROACTIVE` is `"false"`, do not auto-invoke or proactively suggest skills. If a skill seems useful, ask: "I think /skillname might help here — want me to run it?"
@@ -2861,9 +2861,15 @@ _HOOK_INSTALLED="no"
# committed hook and write a machine-local wrapper into the working tree.
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
_GIT_DIR=$(git rev-parse --absolute-git-dir 2>/dev/null || echo "")
# Linked worktrees: --absolute-git-dir is .git/worktrees/<name> but hooks
# resolve to the COMMON .git/hooks, so match against the common dir too or
# every Conductor worktree false-negatives as a "custom hooks path". The
# /nonexistent fallback keeps the case pattern from collapsing to "/*"
# (match-everything) when resolution fails.
_GIT_COMMON=$(cd "$(git rev-parse --git-common-dir 2>/dev/null || echo /nonexistent)" 2>/dev/null && pwd || echo /nonexistent)
_HOOKS_IN_GIT_DIR="no"
case "$_HOOKS_DIR" in
"$_GIT_DIR"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
"$_GIT_DIR"/*|"$_GIT_COMMON"/*|hooks|.git/hooks) _HOOKS_IN_GIT_DIR="yes" ;;
esac
_PREPUSH_PROMPTED=$([ -f "${GSTACK_HOME:-$HOME/.gstack}/.redact-prepush-prompted" ] && echo "yes" || echo "no")
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
+60
View File
@@ -3244,6 +3244,66 @@ describe('EXIT PLAN MODE GATE placement', () => {
});
});
describe('scope-gate exceptions drift-guard', () => {
// The plan-mode auto-select-B exceptions block is hand-duplicated in the
// plan-eng-review and plan-design-review templates (matching the gate
// around it, which predates this block). The two copies must stay
// byte-identical modulo exactly two known variant slots:
// 1. the plan-mode bullet's action tail (Design Doc Check vs pre-review
// audit + mockups),
// 2. the named-target vocabulary ("a path, a doc" vs "a path, a page, a doc").
// A future edit to one copy that silently misses the other fails here
// instead of drifting. The real fix (shared {{SCOPE_GATE}} resolver) is a
// filed TODO — this guard is the stopgap that makes the duplication safe.
const START_MARKER = '**Exceptions — check in this order, BEFORE asking:**';
const END_MARKER = 'in any mode — it is a hard STOP.';
function extractExceptionsBlock(skill: string): string {
const md = fs.readFileSync(path.join(ROOT, skill, 'SKILL.md'), 'utf-8');
const start = md.indexOf(START_MARKER);
expect(start, `${skill}/SKILL.md: exceptions block start marker present`).toBeGreaterThan(-1);
const end = md.indexOf(END_MARKER, start);
expect(end, `${skill}/SKILL.md: exceptions block end marker present`).toBeGreaterThan(start);
return md.slice(start, end + END_MARKER.length);
}
const normalizeVariantSlots = (block: string) =>
block
.replace('Then run the Design Doc Check and Step 0 against that plan.', '<ACTION_TAIL>')
.replace('Then run the pre-review audit, mockups, and Step 0 against that plan.', '<ACTION_TAIL>')
.replace('a path, a page, a doc they pasted,', 'a path, a doc they pasted,');
test('eng and design exceptions blocks are identical modulo the two variant slots', () => {
const eng = normalizeVariantSlots(extractExceptionsBlock('plan-eng-review'));
const design = normalizeVariantSlots(extractExceptionsBlock('plan-design-review'));
expect(eng).toBe(design);
// The action tail must actually have been normalized in both (guards
// against a rewording that bypasses the normalizer and vacuously passes).
expect(eng).toContain('<ACTION_TAIL>');
});
test('exceptions block carries the announcement string the PTY detectors pin', () => {
for (const skill of ['plan-eng-review', 'plan-design-review']) {
const block = extractExceptionsBlock(skill);
expect(block, `${skill}: verbatim announcement`).toContain(
'Scope gate: plan mode — auto-selected B (reviewing <target>).',
);
}
});
test('gate menu carries the question strings the PTY question detector pins', () => {
// isScopeGateQuestionVisible (claude-pty-runner.ts) anchors on the
// question text + option A's body. If the menu is reworded without
// updating the detector, the paid smokes' must-stay-false assertions go
// vacuous — this free pin fails first.
for (const skill of ['plan-eng-review', 'plan-design-review']) {
const md = fs.readFileSync(path.join(ROOT, skill, 'SKILL.md'), 'utf-8');
expect(md, `${skill}: gate question text`).toContain('What should I review?');
expect(md, `${skill}: option A body text`).toContain('The current branch diff');
}
});
});
describe('GSTACK REVIEW REPORT mandatory unresolved-decisions status', () => {
// Report text rides in PLAN_FILE_REVIEW_REPORT → every report consumer gets it.
// devex-review is a report consumer but NOT a gate consumer, so the two target
+8 -3
View File
@@ -164,7 +164,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
},
behavioral: 'plan',
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
maxSkeletonBytes: 67_000,
// +~1 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
maxSkeletonBytes: 68_000,
minUnionBytes: 70_000,
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
@@ -172,7 +173,10 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
// default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
// prose, replacing the smaller opt-in question) land this at ~6.6% over the
// v1.53.0.0 baseline. Headroom for those intentional additions.
maxSizeRatio: 1.08,
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
// hardening: host-anchored mode signal, precedence, passing-mention
// guards) and the plan-mode preamble reword land the union at 1.092.
maxSizeRatio: 1.10,
},
'plan-design-review': {
skill: 'plan-design-review',
@@ -189,7 +193,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
maxSkeletonBytes: 88_000,
// +~1.3 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
maxSkeletonBytes: 89_000,
minUnionBytes: 70_000,
mustContain: ['design', 'visual'],
maxSizeRatio: 1.07,
@@ -0,0 +1,143 @@
/**
* Scope-gate floor-exclusion regression pins (free, static).
*
* runPlanSkillFloorCheck's acceptance condition changed with the plan-mode
* auto-select-B work: a render only satisfies the finding floor when
*
* (isNumberedOptionListVisible(visible) || isProseAUQVisible(visible))
* && !isPermissionDialogVisible(tail)
* && !isScopeGateQuestionVisible(tail) // <- new exclusion
*
* where tail = visible.slice(-TAIL_SCAN_BYTES). The composition lives inline
* in the paid PTY loop, so these tests pin the load-bearing behavior of each
* detector on the exact render shapes the floor passes them:
*
* 1. Both scope-gate render forms (native numbered UI, prose lettered
* fallback) trip the acceptance detectors WITHOUT the exclusion the
* gate would trivially satisfy the floor inside the 3s pre-target
* window. The exclusion must catch both forms.
* 2. A genuine finding-driven AskUserQuestion must NOT trip the exclusion,
* or the floor becomes unsatisfiable.
* 3. The exclusion is TAIL-scoped by design: an early gate render that has
* scrolled past TAIL_SCAN_BYTES must not suppress a later real finding
* AskUserQuestion.
*
* Also closes the untested OR-branch of isScopeGateAutoSelectVisible: the
* fully-collapsed hyphen-less 'autoselectedb' form.
*/
import { describe, test, expect } from 'bun:test';
import {
TAIL_SCAN_BYTES,
isNumberedOptionListVisible,
isProseAUQVisible,
isPermissionDialogVisible,
isScopeGateQuestionVisible,
isScopeGateAutoSelectVisible,
parseNumberedOptions,
} from './claude-pty-runner';
// The gate's native AskUserQuestion render (numbered options + cursor) —
// what fires inside the floor check's 3s window before the seed arrives.
const GATE_NATIVE_RENDER = `
What should I review?
1. The current branch diff the work in progress on this branch.
2. A plan or design doc I'll paste or point you to.
3. A specific file, directory, or path.
`;
// The gate's prose fallback render (lettered options under --disallowedTools).
const GATE_PROSE_RENDER = `
What should I review?
A) The current branch diff the work in progress on this branch.
B) A plan or design doc I'll paste or point you to.
C) A specific file, directory, or path.
Recommendation: A when a branch diff exists, otherwise B.
`;
// A genuine finding-driven AskUserQuestion — the render the floor MEASURES.
const FINDING_AUQ_RENDER = `
Finding 1: the plan reimplements test sharding that Bun provides natively.
1. Use Bun's native --shard flag (recommended)
2. Keep the custom scheduler as planned
3. Defer this decision to implementation
`;
describe('floor-check scope-gate exclusion (acceptance-condition regression)', () => {
test('native gate render trips the acceptance detector — the exclusion is load-bearing', () => {
// Pre-exclusion, this render satisfied the floor by itself.
expect(isNumberedOptionListVisible(GATE_NATIVE_RENDER)).toBe(true);
expect(isPermissionDialogVisible(GATE_NATIVE_RENDER)).toBe(false);
// The new exclusion catches it.
expect(isScopeGateQuestionVisible(GATE_NATIVE_RENDER)).toBe(true);
});
test('prose gate render trips the prose-AUQ arm — the exclusion catches that form too', () => {
expect(isProseAUQVisible(GATE_PROSE_RENDER)).toBe(true);
expect(isPermissionDialogVisible(GATE_PROSE_RENDER)).toBe(false);
expect(isScopeGateQuestionVisible(GATE_PROSE_RENDER)).toBe(true);
});
test('a genuine finding AskUserQuestion is NOT excluded — the floor stays satisfiable', () => {
expect(isNumberedOptionListVisible(FINDING_AUQ_RENDER)).toBe(true);
expect(isPermissionDialogVisible(FINDING_AUQ_RENDER)).toBe(false);
expect(isScopeGateQuestionVisible(FINDING_AUQ_RENDER)).toBe(false);
});
test('tail-scoping: an early gate render scrolled out of the tail does not suppress a later finding AUQ', () => {
// Gate render, then >TAIL_SCAN_BYTES of review output, then the real
// finding AskUserQuestion — the shape the TAIL-scoped exclusion exists for.
const filler = 'Reading the plan and auditing the design system.\n'.repeat(
Math.ceil(TAIL_SCAN_BYTES / 48) + 4,
);
const visible = GATE_NATIVE_RENDER + filler + FINDING_AUQ_RENDER;
const tail = visible.slice(-TAIL_SCAN_BYTES);
// Full buffer still remembers the gate (scrollback)…
expect(isScopeGateQuestionVisible(visible)).toBe(true);
// …but the floor's exclusion looks only at the tail, which is clean:
expect(isScopeGateQuestionVisible(tail)).toBe(false);
// and the acceptance arm (full-buffer scan) sees the finding AUQ.
expect(isNumberedOptionListVisible(visible)).toBe(true);
expect(isPermissionDialogVisible(tail)).toBe(false);
});
test('a gate render inside the tail IS suppressed (no false floor pass)', () => {
const tail = GATE_NATIVE_RENDER.slice(-TAIL_SCAN_BYTES);
expect(isScopeGateQuestionVisible(tail)).toBe(true);
});
test('active-render veto: a finding AUQ close after the gate is NOT vetoed (codex P2 re-review)', () => {
// The finding menu renders <TAIL_SCAN_BYTES after the gate, then the
// model waits (no further output). A blanket tail veto would suppress
// this until timeout; the active-render veto anchors on the LAST cursor
// menu, which is the finding AUQ, so the floor is satisfiable.
const visible = GATE_NATIVE_RENDER + '\nAuditing the plan…\n' + FINDING_AUQ_RENDER;
const activeMenu = parseNumberedOptions(visible);
expect(activeMenu.length).toBeGreaterThan(0);
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
expect(gateIsActiveRender).toBe(false);
});
test('active-render veto: the gate as the pending menu IS vetoed', () => {
const visible = 'booting…\n' + GATE_NATIVE_RENDER;
const activeMenu = parseNumberedOptions(visible);
expect(activeMenu.length).toBeGreaterThan(0);
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
expect(gateIsActiveRender).toBe(true);
});
});
describe('isScopeGateAutoSelectVisible collapsed hyphen-less branch', () => {
test("matches the fully-collapsed 'autoselectedb' form (hyphen lost in TTY reflow)", () => {
const sample = 'Scopegate:planmode—autoselectedB(reviewingPLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('hyphen-less token without the announcement prefix stays false', () => {
const sample = 'The agent autoselectedB from the menu without announcing a scope gate decision.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
});
+162 -10
View File
@@ -618,6 +618,55 @@ export function isProseAUQVisible(visible: string): boolean {
return false;
}
// ---------------------------------------------------------------------------
// Scope-gate render detectors (plan-eng-review / plan-design-review)
// ---------------------------------------------------------------------------
//
// Both anchor on the RENDER SHAPE, not bare keywords, so model narration
// about the gate ("normally I'd ask what should I review…") stays false.
// Matching is whitespace-squished + lowercased because stripAnsi collapses
// TTY cursor-positioning escapes unpredictably (the same failure mode the
// Pattern-4/5 collapsed-form handling above exists for).
/**
* True when the scope-gate QUESTION is actually rendered: the question text
* plus option A's body text. Option-body anchoring (not `A)`/`B)` markers)
* because native AskUserQuestion renders NUMBERED options in the TTY while
* the --disallowedTools prose fallback renders lettered ones the option
* body appears in both renders; narration rarely quotes both the question
* and an option body.
*/
export function isScopeGateQuestionVisible(visible: string): boolean {
const squished = visible.replace(/\s+/g, '').toLowerCase();
return squished.includes('whatshouldireview') && squished.includes('currentbranchdiff');
}
/**
* True when the plan-mode auto-select announcement is rendered:
* "Scope gate: plan mode — auto-selected B (reviewing <target>)."
* Requires BOTH the announcement prefix and an auto-select-B token so
* narration ("in plan mode I'd auto-select B") stays false. The token is
* tense-tolerant (selected/selecting/selects) because the smokes assert
* must-be-TRUE on it a semantically-perfect paraphrase must not fail a
* paid run while the prefix stays exact so paraphrase narration without
* the announcement frame stays false. A prefix immediately preceded by a
* quote character is a QUOTATION (e.g. the model explaining why it is NOT
* announcing), not a render the announcement line itself never renders
* quoted.
*/
export function isScopeGateAutoSelectVisible(visible: string): boolean {
const squished = visible.replace(/\s+/g, '').toLowerCase();
const QUOTES = ['"', "'", '`', '“', ''];
const re = /scopegate:planmode/g;
let m: RegExpExecArray | null;
while ((m = re.exec(squished)) !== null) {
const before = m.index > 0 ? squished[m.index - 1]! : '';
if (QUOTES.includes(before)) continue; // quoted occurrence — narration, keep scanning
if (/auto-?select(?:ed|ing|s)?b/.test(squished.slice(m.index))) return true;
}
return false;
}
/**
* Parse a rendered numbered-option list out of the visible TTY text.
*
@@ -1476,10 +1525,20 @@ export interface PlanSkillObservation {
* "Ready to execute" confirmation
* - 'silent_write' a Write/Edit landed BEFORE any prompt, to a path
* outside the sanctioned plan/project directories
* - 'wrote_findings_before_asking' strictPlanWrites only (seeded runs):
* the plan file was rewritten with findings before any
* AskUserQuestion render (the May-2026 transcript bug)
* - 'exited' claude process died before any of the above
* - 'timeout' none of the above within budget
*/
outcome: 'asked' | 'auto_decided' | 'plan_ready' | 'silent_write' | 'exited' | 'timeout';
outcome:
| 'asked'
| 'auto_decided'
| 'plan_ready'
| 'silent_write'
| 'wrote_findings_before_asking'
| 'exited'
| 'timeout';
/** Human-readable summary. */
summary: string;
/** Visible terminal text since the slash command was sent (last 2KB). */
@@ -1516,6 +1575,28 @@ export interface PlanSkillObservation {
* Haiku judge fallback rather than the regex detector.
*/
waitingEverObserved?: boolean;
/**
* High-water-mark flag: did the scope-gate QUESTION ("What should I
* review?" plus option-body text) ever render during the run? Same
* lossy-2KB-evidence rationale as proseAUQEverObserved. The plan-mode
* smokes assert this stays false (gate bypassed via auto-select B); the
* no-op regression asserts it fires outside plan mode.
*/
scopeGateQuestionObserved?: boolean;
/**
* High-water-mark flag: did the plan-mode auto-select announcement
* ("Scope gate: plan mode — auto-selected B …") ever render? The
* plan-mode smokes assert true; the no-op regression asserts false.
*/
scopeGateAutoSelectObserved?: boolean;
/**
* High-water map for opts.trackTokens: token did it EVER appear in the
* cumulative visible buffer? Consumption asserts (e.g. "the pasted target's
* distinctive token shows up in the review output") must not depend on the
* lossy 2KB evidence tail plan-file fallbacks are unreachable outside
* plan mode (extractPlanFilePath only matches plan-mode save renders).
*/
tokensObserved?: Record<string, boolean>;
}
/**
@@ -1576,6 +1657,10 @@ export async function runPlanSkillObservation(opts: {
/** Override the spawned model. Defaults via launchClaudePty's chain
* (opts.model ?? EVALS_MODEL ?? 'claude-sonnet-4-6'). */
model?: string;
/** Literal tokens to track as high-water marks over the CUMULATIVE visible
* buffer (case-sensitive). Results land in obs.tokensObserved. Use for
* consumption asserts that must survive the 2KB evidence tail. */
trackTokens?: string[];
}): Promise<PlanSkillObservation> {
const startedAt = Date.now();
const session = await launchClaudePty({
@@ -1619,6 +1704,21 @@ export async function runPlanSkillObservation(opts: {
// even if the current state is 'working'.
let proseAUQEverObserved = false;
let waitingEverObserved = false;
let scopeGateQuestionObserved = false;
let scopeGateAutoSelectObserved = false;
const tokensObserved: Record<string, boolean> = {};
for (const t of opts.trackTokens ?? []) tokensObserved[t] = false;
// Single source for the high-water flags at EVERY return site. Hand-
// spreading them per-site already drifted once (the judge-waiting return
// omitted the prose/waiting flags); a site that forgets a must-stay-false
// flag makes `obs.flag ?? false` negative assertions pass vacuously.
const highWaterFlags = () => ({
proseAUQEverObserved,
waitingEverObserved,
scopeGateQuestionObserved,
scopeGateAutoSelectObserved,
...(opts.trackTokens?.length ? { tokensObserved } : {}),
});
const JUDGE_AFTER_MS = 60_000;
const JUDGE_INTERVAL_MS = 30_000;
while (Date.now() - start < budgetMs) {
@@ -1631,6 +1731,7 @@ export async function runPlanSkillObservation(opts: {
summary: `claude exited (code=${session.exitCode()}) before reaching a terminal outcome`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
if (visible.includes('Unknown command:')) {
@@ -1639,6 +1740,7 @@ export async function runPlanSkillObservation(opts: {
summary: `claude rejected /${opts.skillName} as unknown command (skill not registered in this cwd)`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
@@ -1652,6 +1754,18 @@ export async function runPlanSkillObservation(opts: {
tag: 'prose-auq-surfaced',
});
}
// Scope-gate render tracking (same high-water shape). Full-run
// detection matters because the 2KB evidence tail usually scrolls
// past the gate render before the outcome fires.
if (!scopeGateQuestionObserved && isScopeGateQuestionVisible(visible)) {
scopeGateQuestionObserved = true;
}
if (!scopeGateAutoSelectObserved && isScopeGateAutoSelectVisible(visible)) {
scopeGateAutoSelectObserved = true;
}
for (const t of opts.trackTokens ?? []) {
if (!tokensObserved[t] && visible.includes(t)) tokensObserved[t] = true;
}
const classified = classifyVisible(visible, {
strictPlanWrites: !!opts.initialPlanContent,
@@ -1661,8 +1775,7 @@ export async function runPlanSkillObservation(opts: {
...classified,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
// Capture the plan file path on any outcome where one may have been
// written. Gating only on 'plan_ready' missed two cases: (1) the
@@ -1693,6 +1806,7 @@ export async function runPlanSkillObservation(opts: {
summary: `LLM judge: ${lastJudgeVerdict.reasoning} (state=waiting after ${Math.round(elapsed / 1000)}s)`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
}
@@ -1714,8 +1828,7 @@ export async function runPlanSkillObservation(opts: {
: ''),
evidence: finalVisible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
}
return {
@@ -1727,8 +1840,7 @@ export async function runPlanSkillObservation(opts: {
: ''),
evidence: finalVisible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
} finally {
await session.close();
@@ -2099,11 +2211,23 @@ export async function runPlanSkillFloorCheck(opts: {
const start = Date.now();
let lastJudgeAt = 0;
let lastJudgeVerdict: PtyStateVerdict | null = null;
// Positional anchor for the scope-gate exclusion. The visible buffer is
// append-only (old renders never leave scrollback), so a gate question
// rendered in the 3s pre-target window would keep satisfying the
// full-buffer acceptance checks forever while a tail-only exclusion
// stops seeing it after ~TAIL_SCAN_BYTES of output — a vacuous
// auq_observed (found independently by 4 review passes). Once the gate
// render is seen, acceptance only counts AUQ renders in content APPENDED
// after that point.
let gateSeenIdx = -1;
const JUDGE_AFTER_MS = 60_000;
const JUDGE_INTERVAL_MS = 30_000;
while (Date.now() - start < timeoutMs) {
await Bun.sleep(2000);
const visible = session.visibleSince(since);
if (gateSeenIdx === -1 && isScopeGateQuestionVisible(visible)) {
gateSeenIdx = visible.length;
}
if (session.exited()) {
return {
@@ -2129,10 +2253,34 @@ export async function runPlanSkillFloorCheck(opts: {
// OR via prose-rendered options under --disallowedTools when no MCP
// variant is callable (isProseAUQVisible). Both surface the question
// to the user; the bug we're catching is "fired zero AUQs."
//
// Scope-gate renders do NOT count: the gate's "What should I review?"
// can fire inside the 3s pre-target window and would trivially satisfy
// the floor, but the floor measures FINDING-driven questions. Once a
// gate render has been seen, acceptance scans only the content APPENDED
// after it (positional anchor above) — the buffer is append-only, so a
// whole-buffer acceptance would keep matching the stale gate render
// forever.
//
// The gate veto is ACTIVE-RENDER-aware, not blanket-tail: when a
// numbered menu is up, parseNumberedOptions anchors on the LAST cursor
// line, so we veto only when the pending menu IS the gate — a finding
// AUQ that renders within TAIL_SCAN_BYTES of the gate (model waiting,
// no further output) still satisfies the floor. Prose renders have no
// cursor anchor, so the prose path falls back to the tail check
// (accepted residual: prose gate + prose finding inside one tail can
// suppress until timeout; floors run the native-menu path in practice).
const tail = visible.slice(-TAIL_SCAN_BYTES);
const acceptWindow = gateSeenIdx === -1 ? visible : visible.slice(gateSeenIdx);
const activeMenu = parseNumberedOptions(visible);
const gateIsActiveRender =
activeMenu.length > 0
? activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label))
: isScopeGateQuestionVisible(tail);
if (
(isNumberedOptionListVisible(visible) || isProseAUQVisible(visible)) &&
!isPermissionDialogVisible(tail)
(isNumberedOptionListVisible(acceptWindow) || isProseAUQVisible(acceptWindow)) &&
!isPermissionDialogVisible(tail) &&
!gateIsActiveRender
) {
return {
auqObserved: true,
@@ -2154,7 +2302,11 @@ export async function runPlanSkillFloorCheck(opts: {
lastJudgeAt = Date.now();
logPtySnapshot(visible, { testName: opts.skillName, elapsedMs: elapsed, tag: 'floor-judge-tick' });
lastJudgeVerdict = judgePtyState(visible, { testName: opts.skillName });
if (lastJudgeVerdict.state === 'waiting') {
// The judge can't tell a scope-gate question from a finding question,
// so a 'waiting' verdict while the gate menu is the pending render
// must NOT satisfy the floor — same active-render exclusion as the
// regex path.
if (lastJudgeVerdict.state === 'waiting' && !gateIsActiveRender) {
return {
auqObserved: true,
outcome: 'auq_observed',
+109
View File
@@ -28,6 +28,8 @@ import {
isPermissionDialogVisible,
isNumberedOptionListVisible,
isProseAUQVisible,
isScopeGateQuestionVisible,
isScopeGateAutoSelectVisible,
isPlanReadyVisible,
parseNumberedOptions,
classifyVisible,
@@ -194,6 +196,113 @@ describe('isNumberedOptionListVisible', () => {
});
});
describe('scope-gate render detectors', () => {
// The verbatim announcement string from the plan-eng/plan-design SKILL.md
// templates. If the template rewording drifts, THIS fixture fails first —
// before the paid plan-mode smokes silently degrade to vacuous asserts.
const TEMPLATE_ANNOUNCEMENT =
'Scope gate: plan mode — auto-selected B (reviewing <target>).';
describe('isScopeGateQuestionVisible', () => {
test('matches the clean prose gate render (question + option bodies)', () => {
const sample = `
What should I review?
A) The current branch diff the work in progress on this branch.
B) A plan or design doc I'll paste or point you to.
C) A specific file, directory, or path.
Recommendation: A when a branch diff exists, otherwise B.
`;
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('matches the native numbered render (no lettered markers)', () => {
const sample = `
What should I review?
1. The current branch diff the work in progress on this branch.
2. A plan or design doc I'll paste or point you to.
3. A specific file, directory, or path.
`;
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('matches the PTY-collapsed render (stripAnsi squished spaces)', () => {
const sample = 'WhatshouldIreview?A)Thecurrentbranchdiff—theworkinprogress';
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('stays false on narration quoting only the question', () => {
const sample =
"Normally I'd ask 'What should I review?' but plan mode is active, so I'm proceeding.";
expect(isScopeGateQuestionVisible(sample)).toBe(false);
});
test('stays false on unrelated review prose', () => {
const sample = 'I will review the current branch diff and report findings.';
expect(isScopeGateQuestionVisible(sample)).toBe(false);
});
});
describe('isScopeGateAutoSelectVisible', () => {
test('matches the verbatim template announcement', () => {
expect(isScopeGateAutoSelectVisible(TEMPLATE_ANNOUNCEMENT)).toBe(true);
});
test('matches a real announcement with a concrete target', () => {
const sample =
'Scope gate: plan mode — auto-selected B (reviewing ~/.claude/plans/my-feature.md). Running the Design Doc Check next.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('matches the PTY-collapsed announcement', () => {
const sample = 'Scopegate:planmode—auto-selectedB(reviewingPLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('stays false on narration about the behavior', () => {
const sample = "In plan mode I'd auto-select B and review the active plan.";
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('stays false on a VERBATIM QUOTE of the announcement (negation narration)', () => {
// The exact announcement line sits quoted in the skill context, so a
// model explaining why it is NOT firing it can reproduce it byte-exact
// inside quotes — that must not trip a must-stay-false assert.
const sample =
'Not in plan mode, so I won\'t announce "Scope gate: plan mode — auto-selected B (reviewing <target>)." and will ask instead.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('a later real render still matches after an earlier quoted mention', () => {
const sample =
'Earlier I said I would render "Scope gate: plan mode — auto-selected B (…)" and now:\n' +
'Scope gate: plan mode — auto-selected B (reviewing PLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('matches tense paraphrases WITH the announcement prefix (auto-selecting / auto-selects)', () => {
expect(
isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selecting B (reviewing the drafted plan).'),
).toBe(true);
expect(isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selects B.')).toBe(true);
});
test('stays false on tense paraphrases WITHOUT the announcement prefix', () => {
expect(isScopeGateAutoSelectVisible('Auto-selecting B since we are in plan mode.')).toBe(false);
});
test('stays false on AUTO_DECIDE preamble output', () => {
const sample = 'Auto-decided scope question → B (your preference). Change with /plan-tune.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('stays false on a bare "selected B" without the announcement prefix', () => {
const sample = 'I selected B as the review target.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
});
});
describe('isProseAUQVisible', () => {
test('matches 4 lettered options A) B) C) D) at line starts (plan-eng prose AUQ shape)', () => {
const sample = `
+4 -1
View File
@@ -234,7 +234,10 @@ const MONOLITH_INVARIANTS: ParityInvariant[] = [
// cross-session decision-memory nudge) lands this skill just over the strict 1.05;
// headroom for the shared preamble additions (matches the carved-skill overrides).
// v1.2.0 activation lift adds the first-run-guidance section on top.
maxSizeRatio: 1.09,
// 1.09 → 1.10: the plan-mode preamble reword (scope-gate auto-select-B
// change) adds ~250 B to every skill's shared preamble; investigate was
// the closest to its ceiling (landed 1.092).
maxSizeRatio: 1.10,
minBytes: 30_000,
},
{
+10 -4
View File
@@ -98,11 +98,17 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// include question-tuning.ts and generate-ask-user-format.ts because the
// AUTO_DECIDE preamble injection lives there and changes can flip the
// regression test outcome between 'asked' and 'auto_decided'.
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-ceo-plan-mode.test.ts'],
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-eng-plan-mode.test.ts'],
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts'],
'plan-devex-review-plan-mode': ['plan-devex-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-mode-no-op': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts'],
// Covers ceo (preamble misfire) + eng/design (scope-gate bypass must not
// fire outside plan mode) + the named-target exception case. 4 PTY runs;
// in CI these run CONCURRENT with the rest of the pty-plan-smoke suite
// (--max-concurrency + --retry 2), so worst-case cost is ~3x a single
// pass of each, sharing the API budget with sibling tests — not the
// sequential ~+10min a local read suggests.
'plan-mode-no-op': ['plan-ceo-review/**', 'plan-eng-review/**', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-mode-no-op.test.ts'],
// v1.21+ AskUserQuestion-blocked regression tests — Conductor launches
// claude with `--disallowedTools AskUserQuestion --permission-mode default`
+12 -4
View File
@@ -176,7 +176,13 @@ Include: color trends, typography patterns, and layout conventions you observed.
Do NOT generate a full DESIGN.md just research notes.`,
workingDirectory: researchDir,
maxTurns: 8,
timeout: 90_000,
// 300s, not 90s: saturated-runner class (same as review-dashboard-via /
// retro-base-branch). PR #2533 CI observed the sibling preview test at
// 0 turns/$0.00 for 93s x3 attempts — session up, first completion
// queued past the budget under concurrent API load. 90s budgets cannot
// absorb one slow first completion; 300s is the repo's standard floor
// for CI SDK tests. Outer timeout below rises to 360s for headroom.
timeout: 300_000,
testName: 'design-consultation-research',
runId,
});
@@ -206,7 +212,7 @@ Do NOT generate a full DESIGN.md — just research notes.`,
}
try { fs.rmSync(researchDir, { recursive: true, force: true }); } catch {}
}, 120_000);
}, 360_000);
testConcurrentIfSelected('design-consultation-existing', async () => {
// Pre-create a minimal DESIGN.md (independent of core test)
@@ -274,7 +280,9 @@ Write a single HTML file to ${previewDir}/design-preview.html that shows:
Do NOT write DESIGN.md only the preview HTML.`,
workingDirectory: previewDir,
maxTurns: 8,
timeout: 90_000,
// 300s, not 90s: this is the test that failed 3x at 0 turns/$0.00/93s
// on PR #2533 CI — see the research test's comment for the class.
timeout: 300_000,
testName: 'design-consultation-preview',
runId,
});
@@ -303,7 +311,7 @@ Do NOT write DESIGN.md — only the preview HTML.`,
}
try { fs.rmSync(previewDir, { recursive: true, force: true }); } catch {}
}, 120_000);
}, 360_000);
});
// --- Plan Design Review E2E (plan-mode) ---
+7 -2
View File
@@ -47,7 +47,12 @@ describeE2E('plan-ceo-review plan-mode smoke (gate)', () => {
const obs = await runPlanSkillObservation({
skillName: 'plan-ceo-review',
inPlanMode: true,
timeoutMs: 300_000,
// 420s, not 300s: measured 2026-08-11, a clean isolated pass took
// 295.7s (80s on a quiet main run) — 4s under the old budget — and the
// same run timed out at ~308s three times under concurrent eval load.
// Same runner-contention class as review-dashboard-via/retro-base-
// branch; headroom instead of a budget-edge flake in the gate lane.
timeoutMs: 420_000,
env: { QUESTION_TUNING: 'false', EXPLAIN_LEVEL: 'default' },
});
@@ -72,5 +77,5 @@ describeE2E('plan-ceo-review plan-mode smoke (gate)', () => {
);
}
assertReportAtBottomIfPlanWritten(obs);
}, 360_000);
}, 480_000);
});
@@ -1,5 +1,5 @@
/**
* /plan-design-review AskUserQuestion floor regression (gate, paid, real-PTY).
* /plan-design-review AskUserQuestion floor regression (periodic, paid, real-PTY).
*
* See test/skill-e2e-plan-eng-finding-floor.test.ts for the contract.
*/
@@ -8,10 +8,10 @@ import { describe, test } from 'bun:test';
import { runPlanSkillFloorCheck } from './helpers/claude-pty-runner';
import { FORCING_FLOOR_DESIGN } from './fixtures/forcing-finding-seeds';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
const describeE2E = shouldRun ? describe : describe.skip;
describeE2E('/plan-design-review AskUserQuestion floor (gate)', () => {
describeE2E('/plan-design-review AskUserQuestion floor (periodic)', () => {
test(
'seeded forcing finding causes the agent to fire at least one AskUserQuestion',
async () => {
+62 -3
View File
@@ -1,5 +1,5 @@
/**
* plan-design-review plan-mode smoke (gate, paid, real-PTY).
* plan-design-review plan-mode smoke (periodic, paid, real-PTY).
*
* See test/skill-e2e-plan-ceo-plan-mode.test.ts for the shared assertion
* contract. Exercises the same contract against /plan-design-review.
@@ -15,10 +15,33 @@ import {
assertReportAtBottomIfPlanWritten,
} from './helpers/claude-pty-runner';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
const describeE2E = shouldRun ? describe : describe.skip;
describeE2E('plan-design-review plan-mode smoke (gate)', () => {
// UI-heavy seed with guaranteed design gaps (center-aligned everything, no
// empty states, no responsive intent) so the review has real findings to
// surface. Inline twin of the eng smoke's SEED_PLAN_FORCING_FINDINGS —
// FORCING_FLOOR_DESIGN from forcing-finding-seeds.ts is NOT reusable here:
// it embeds a write-to-/tmp instruction shaped for the floor check's
// followUpPrompt, which would trip strictPlanWrites as a silent_write.
const SEED_PLAN_UI_HEAVY = `
# Plan: Marketing landing page
## Layout
All headings, taglines, and body copy will be center-aligned for a
"clean modern look." The hero h1 sits 8px above the subhead; the CTA
button has the same visual weight as the "Learn more" link beside it.
## Pages
- / (hero, 3-column features grid, testimonials carousel, footer)
- /pricing (3 tier cards)
## States
Only the happy path is designed. No empty states, no error states,
no loading states. Mobile: "stacks on mobile."
`;
describeE2E('plan-design-review plan-mode smoke (periodic)', () => {
test('reaches a terminal outcome (asked or plan_ready) without silent writes', async () => {
const obs = await runPlanSkillObservation({
skillName: 'plan-design-review',
@@ -37,4 +60,40 @@ describeE2E('plan-design-review plan-mode smoke (gate)', () => {
expect(['asked', 'plan_ready']).toContain(obs.outcome);
assertReportAtBottomIfPlanWritten(obs);
}, 360_000);
// Plan-mode scope-gate bypass: with a seeded UI-heavy plan in plan mode,
// the gate must NOT render its "What should I review?" menu — it
// auto-selects B and announces it, then proceeds to the pre-review audit
// and mockups. Mirrors the eng smoke's seeded STOP-gate test, without
// --disallowedTools (native AUQ available is the common path here).
test('scope gate auto-selects B when a plan is seeded in plan mode', async () => {
const obs = await runPlanSkillObservation({
skillName: 'plan-design-review',
inPlanMode: true,
initialPlanContent: SEED_PLAN_UI_HEAVY,
timeoutMs: 300_000,
});
if (
obs.outcome === 'wrote_findings_before_asking' ||
obs.outcome === 'auto_decided' ||
obs.outcome === 'silent_write' ||
obs.outcome === 'exited' ||
obs.outcome === 'timeout'
) {
throw new Error(
`plan-design plan-mode bypass FAILED: outcome=${obs.outcome}\n` +
`summary: ${obs.summary}\nelapsed: ${obs.elapsedMs}ms\n` +
`--- evidence (last 2KB) ---\n${obs.evidence}`,
);
}
expect(['asked', 'plan_ready']).toContain(obs.outcome);
assertReportAtBottomIfPlanWritten(obs);
// The bypass contract (exception ordering makes this deterministic even
// though the seed arrives as a pasted user message).
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(true);
}, 360_000);
});
@@ -1,5 +1,5 @@
/**
* /plan-eng-review AskUserQuestion floor regression (gate, paid, real-PTY).
* /plan-eng-review AskUserQuestion floor regression (periodic, paid, real-PTY).
*
* Catches the May 2026 transcript bug where /plan-eng-review wrote a
* multi-section review plan to ~/.claude/plans/ and called ExitPlanMode
@@ -11,7 +11,7 @@
* render. See claude-pty-runner.ts for why this is separate from the
* runPlanSkillCounting harness used by periodic finding-count tests.
*
* Tier: gate. Budget: 10 min (early exit on success ~30-90s typical).
* Tier: periodic. Budget: 10 min (early exit on success ~30-90s typical).
* Cost: ~$0.50-$1.50 per run depending on early-exit timing.
*/
@@ -19,10 +19,10 @@ import { describe, test } from 'bun:test';
import { runPlanSkillFloorCheck } from './helpers/claude-pty-runner';
import { FORCING_FLOOR_ENG } from './fixtures/forcing-finding-seeds';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
const describeE2E = shouldRun ? describe : describe.skip;
describeE2E('/plan-eng-review AskUserQuestion floor (gate)', () => {
describeE2E('/plan-eng-review AskUserQuestion floor (periodic)', () => {
test(
'seeded forcing finding causes the agent to fire at least one AskUserQuestion',
async () => {
+13 -3
View File
@@ -1,5 +1,5 @@
/**
* plan-eng-review plan-mode smoke (gate, paid, real-PTY).
* plan-eng-review plan-mode smoke (periodic, paid, real-PTY).
*
* See test/skill-e2e-plan-ceo-plan-mode.test.ts for the shared assertion
* contract. This file exercises the same contract against /plan-eng-review.
@@ -12,7 +12,7 @@ import {
assertReportAtBottomIfPlanWritten,
} from './helpers/claude-pty-runner';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'periodic';
const describeE2E = shouldRun ? describe : describe.skip;
// SEED_PLAN_FORCING_FINDINGS: 8+ files + custom-vs-builtin smell forces the
@@ -45,7 +45,7 @@ Ignore Bun's native --shard flag because we want full control.
None planned will add later.
`;
describeE2E('plan-eng-review plan-mode smoke (gate)', () => {
describeE2E('plan-eng-review plan-mode smoke (periodic)', () => {
test('reaches a terminal outcome (asked or plan_ready) without silent writes', async () => {
const obs = await runPlanSkillObservation({
skillName: 'plan-eng-review',
@@ -108,5 +108,15 @@ describeE2E('plan-eng-review plan-mode smoke (gate)', () => {
expect(['asked', 'plan_ready']).toContain(obs.outcome);
assertReportAtBottomIfPlanWritten(obs);
// Plan-mode scope-gate bypass: with a seeded plan in plan mode, the gate
// must NOT render its "What should I review?" menu — it auto-selects B
// and announces it. Exception ordering in the template (plan-mode branch
// first) makes this deterministic even though the seed arrives as a
// pasted user message. Unseeded test 1 keeps its lenient contract: with
// no plan drafted, the "ask as normal" fallback legitimately renders the
// question.
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(true);
}, 360_000);
});
+116 -15
View File
@@ -1,17 +1,33 @@
/**
* Plan-mode-info no-op regression (gate tier, paid, real-PTY).
*
* Asserts: when /plan-ceo-review is invoked OUTSIDE plan mode (no
* Asserts: when a plan-review skill is invoked OUTSIDE plan mode (no
* --permission-mode plan flag, no plan-mode reminder injected), the skill
* still reaches a terminal outcome ('asked' or 'plan_ready'). This is the
* negative coverage to the per-skill plan-mode smokes if the
* plan-mode-info preamble section ever starts misfiring for non-plan-mode
* sessions (e.g., gating questions on a phrase that isn't there), this
* test catches it.
* negative coverage to the per-skill plan-mode smokes if plan-mode-keyed
* behavior ever starts misfiring for non-plan-mode sessions (e.g., gating
* questions on a phrase that isn't there, or the plan-eng/plan-design
* scope-gate auto-select-B bypass firing without plan mode), this test
* catches it.
*
* Why this matters: outside plan mode, claude doesn't render a native
* confirmation UI. The skill must drive its own AskUserQuestion. Same
* runner, same outcome contract just `inPlanMode: false`.
*
* Coverage grew with the scope-gate bypass (plan-mode auto-select B):
* - plan-ceo-review: original preamble-misfire regression.
* - plan-eng-review / plan-design-review: the bypass must NOT fire outside
* plan mode (scopeGateAutoSelectObserved stays false), and when the run
* ends in 'asked', the question that fired must be the scope gate itself
* (outside plan mode with no named target, the gate is the FIRST
* question by contract).
* - named-target case: a pasted draft (initialPlanContent) IS an
* explicitly-named target, so the gate must NOT ask and the review
* must actually consume the pasted content.
*
* Cost note: 4 sequential PTY runs (~3-5 min each) in the gate lane, up
* from 1 pre-bypass. Selected only when plan-ceo/eng/design or the runner
* change (see 'plan-mode-no-op' in touchfiles.ts).
*/
import { describe, test, expect } from 'bun:test';
@@ -20,29 +36,114 @@ import { runPlanSkillObservation } from './helpers/claude-pty-runner';
const shouldRun = !!process.env.EVALS && process.env.EVALS_TIER === 'gate';
const describeE2E = shouldRun ? describe : describe.skip;
const PLAN_MODE_REMINDER =
'Plan mode is active. The user indicated that they do not want you to execute yet';
// Distinctive token proves the pasted target was consumed by the review —
// not just that no question fired. Nonsense-unique so it can't appear by
// coincidence in skill output.
const SEED_TOKEN = 'ZephyrLedgerWidget';
const NAMED_TARGET_SEED = `
# Plan: ${SEED_TOKEN} settings panel
## Scope
Add a ${SEED_TOKEN} settings panel with a single toggle that enables
weekly export emails. One new component, one route, one test file.
## Files
- src/components/${SEED_TOKEN}.tsx (new)
- src/routes/settings.tsx (add panel)
- test/${SEED_TOKEN}.test.tsx (new)
`;
describeE2E('plan-mode-info no-op outside plan mode (gate regression)', () => {
test('skill reaches a terminal outcome outside plan mode', async () => {
for (const skillName of ['plan-ceo-review', 'plan-eng-review', 'plan-design-review'] as const) {
test(`${skillName} reaches a terminal outcome outside plan mode`, async () => {
const obs = await runPlanSkillObservation({
skillName,
inPlanMode: false,
timeoutMs: 300_000,
// eng/design: force the prose-fallback path. The unconditional
// gate-must-ask assert below pins the render shape the detector
// anchors on, and only the --disallowedTools prose fallback makes
// that shape CONTRACTUAL ("use exactly this shape" in the template);
// native AskUserQuestion could render terse option labels that a
// correct run would fail on (red-team finding).
...(skillName === 'plan-ceo-review'
? {}
: { extraArgs: ['--disallowedTools', 'AskUserQuestion'] }),
});
if (obs.outcome === 'silent_write' || obs.outcome === 'exited' || obs.outcome === 'timeout') {
throw new Error(
`plan-mode no-op regression FAILED (${skillName}): outcome=${obs.outcome}\n` +
`summary: ${obs.summary}\n` +
`elapsed: ${obs.elapsedMs}ms\n` +
`--- evidence (last 2KB visible) ---\n${obs.evidence}`,
);
}
expect(['asked', 'plan_ready']).toContain(obs.outcome);
// Negative regression: the rendered output must NOT echo the plan-mode
// distinctive reminder phrase. If it does, the plan-mode preamble
// section is leaking outside plan mode.
expect(obs.evidence).not.toContain(PLAN_MODE_REMINDER);
if (skillName !== 'plan-ceo-review') {
// Scope-gate bypass must not misfire: no auto-select announcement
// outside plan mode.
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(false);
// UNCONDITIONAL: outside plan mode with no named target, the gate is
// a hard STOP before any tool call, so the gate question must have
// rendered no matter which terminal outcome fired. Gating this on
// outcome === 'asked' would let a silent-bypass run that reaches
// plan_ready (isPlanReadyVisible also matches common prose) sail
// through — the exact regression this test exists to catch.
expect(obs.scopeGateQuestionObserved ?? false).toBe(true);
}
}, 360_000);
}
// Named-target exception (outside plan mode): a pasted draft IS an
// explicitly-named target, so the scope gate must NOT ask — and the
// review must consume the pasted content (seed token visible in the
// review output), proving the target was used rather than the question
// merely skipped. Also the over-trigger guard for the tightened
// "explicit-only" exception wording.
test('plan-eng-review skips the scope gate for an explicitly-pasted target', async () => {
const obs = await runPlanSkillObservation({
skillName: 'plan-ceo-review',
skillName: 'plan-eng-review',
inPlanMode: false,
initialPlanContent: NAMED_TARGET_SEED,
trackTokens: [SEED_TOKEN],
timeoutMs: 300_000,
});
if (obs.outcome === 'silent_write' || obs.outcome === 'exited' || obs.outcome === 'timeout') {
if (
obs.outcome === 'wrote_findings_before_asking' ||
obs.outcome === 'silent_write' ||
obs.outcome === 'exited' ||
obs.outcome === 'timeout'
) {
throw new Error(
`plan-mode no-op regression FAILED: outcome=${obs.outcome}\n` +
`named-target no-op FAILED: outcome=${obs.outcome}\n` +
`summary: ${obs.summary}\n` +
`elapsed: ${obs.elapsedMs}ms\n` +
`--- evidence (last 2KB visible) ---\n${obs.evidence}`,
);
}
expect(['asked', 'plan_ready']).toContain(obs.outcome);
// Negative regression: the rendered output must NOT echo the plan-mode
// distinctive reminder phrase. If it does, the plan-mode preamble
// section is leaking outside plan mode.
const PLAN_MODE_REMINDER =
'Plan mode is active. The user indicated that they do not want you to execute yet';
expect(obs.evidence).not.toContain(PLAN_MODE_REMINDER);
// The pasted doc is the named target: gate question must not render,
// no plan-mode announcement either (we are NOT in plan mode).
expect(obs.scopeGateQuestionObserved ?? false).toBe(false);
expect(obs.scopeGateAutoSelectObserved ?? false).toBe(false);
// Target consumption via high-water token tracking over the CUMULATIVE
// buffer — the 2KB evidence tail is lossy and the plan-file fallback is
// unreachable outside plan mode (extractPlanFilePath only matches
// plan-mode save renders).
expect(obs.tokensObserved?.[SEED_TOKEN] ?? false).toBe(true);
}, 360_000);
});
+9
View File
@@ -134,6 +134,15 @@ describe('SKILL.md command validation', () => {
const result = validateSkill(skill);
expect(result.snapshotFlagErrors).toHaveLength(0);
});
test('autoplan section skip list includes the scope gate', () => {
// autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md
// verbatim; without this skip-list entry it ingests their scope gate — a
// hard-STOP AskUserQuestion that contradicts autoplan's auto-decide
// contract. Nothing else pins the skip-list contents.
const md = fs.readFileSync(path.join(ROOT, 'autoplan', 'SKILL.md'), 'utf-8');
expect(md).toContain('- Scope gate (the plan under review is already the target)');
});
});
describe('Command registry consistency', () => {