v1.62.0.0 feat: plan-mode auto-select at the review scope gate (#2533)

* fix(evals): align plan-eng/design plan-mode + finding-floor smokes to their declared periodic tier

The #2077 demotion of these four stochastic tests to 'periodic' was inert:
E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate',
so they kept running in the blocking gate lane and never in the weekly lane.

Flip the four self-gates to 'periodic' (headers/describe labels updated), add
a free static tier-alignment invariant test (dep-list filename mapping;
unmapped self-gated files are reported, never silently skipped), and name the
two plan-mode test files in their own touchfiles dep lists so the invariant
binds for them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(pty-runner): scope-gate question/auto-select detectors + observation flags

Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5
collapsed-form handling): isScopeGateQuestionVisible requires the question text
PLUS option-body text (native AskUserQuestion renders numbered options, prose
fallback renders lettered — the option body appears in both; narration doesn't),
and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the
selected-B token.

runPlanSkillObservation gains scopeGateQuestionObserved /
scopeGateAutoSelectObserved high-water flags (attached at every return path) so
paid smokes can assert gate behavior across the whole run instead of the lossy
2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render
toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven
questions, and the gate could fire inside the 3s pre-target window.

Unit fixtures pin clean/native/collapsed positives, narration negatives, and
the verbatim template announcement string (template rewording fails here first,
before the paid smokes degrade to vacuous asserts).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(plan-eng/design-review): auto-select B in plan mode at the scope gate

In plan mode the scope gate's "What should I review? A/B/C" question is pure
friction: there is no branch diff and the target is the plan being drafted.
Both gates gain an ordered exceptions block, checked BEFORE asking:

1. Plan mode → auto-select B: review the active plan (in context or pasted),
   announce it in one line ("Scope gate: plan mode — auto-selected B
   (reviewing <target>)") so the user can interrupt; an explicitly different
   user-named target still wins; no plan drafted yet → ask as normal.
2. User-named target (outside plan mode): explicit-only — a path, a pasted
   doc, or the literal words "branch diff". A passing mention is not naming;
   when in doubt, ask.

Outside plan mode with no explicitly-named target, nothing changes. Plan-mode
is checked FIRST because the PTY harness seeds drafts as pasted user messages
(claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic.

Pinning: seeded plan-mode smokes assert no gate render + announcement rendered
(eng test 2; new design seeded test); plan-mode-no-op extends to eng/design
(bypass must not misfire outside plan mode; first question must be the gate)
plus a named-target case proving the pasted target is consumed; a drift-guard
asserts the two hand-duplicated exceptions blocks stay identical modulo the
two variant slots and carry the announcement string the detectors pin.

Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union
ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(autoplan): skip the scope gate when following loaded review skills

autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim,
and its section skip list omitted the scope gate — so autoplan ingested a
hard-STOP AskUserQuestion that contradicts its every-question-auto-decides
contract. One skip-list line fixes it; a static toContain pin in
skill-validation keeps the entry load-bearing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: file scope-gate resolver-extraction TODO (eng-review D5 follow-up)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pty-runner): positional floor exclusion, flag builder, outcome union, token tracking

Review-army + adversarial findings on the scope-gate observability work,
all verified before fixing:

- Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate
  exclusion scanned only the 1500-byte tail, so an early gate render satisfied
  the floor vacuously once ~1.5KB of output accumulated (found independently
  by 4 review passes; predicate reproduced). Acceptance now scans only content
  APPENDED after the first gate render (positional anchor), and the LLM-judge
  'waiting' shortcut no longer fires while the gate menu is the pending render.
- High-water flags are built once and spread at every return path — the
  hand-spread pattern had already drifted (judge-waiting return omitted two
  flags), which made must-stay-false asserts vacuous on those paths.
- isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects
  token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases)
  and quoted-occurrence rejection (a model verbatim-quoting the announcement
  while declining must not trip must-stay-FALSE asserts). Fixtures added for
  both directions.
- PlanSkillObservation outcome union gains 'wrote_findings_before_asking'
  (returned at runtime via classifyVisible but missing from the type).
- trackTokens/tokensObserved: cumulative-buffer token high-water for
  consumption asserts (the 2KB evidence tail is lossy and the plan-file
  fallback is unreachable outside plan mode).
- New scope-gate-floor unit pins (from the ship coverage audit): both gate
  render forms trip acceptance and exclusion; a genuine finding AUQ is not
  excluded; tail-scoping semantics pinned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(evals): harden no-op asserts, close tier-invariant fail-open holes, pin gate question strings

- no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the
  outcome==='asked' conditional let a silent-bypass plan_ready run sail
  through); eng/design cases force --disallowedTools so the pinned prose
  shape is contractual rather than hoping native AUQ renders match; the
  named-target case uses trackTokens for consumption and lists
  wrote_findings_before_asking in its diagnostic throw branch.
- tier-alignment invariant: both quote styles matched; zero-self-gate,
  mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED
  instead of silently skipped (the fail-open holes three reviewers found).
- drift-guard: the generated gate menus must carry the exact question/option
  strings the PTY question detector anchors on — free CI fails before the
  paid smokes can go vacuous on a menu reword.
- touchfiles: corrected the no-op cost note for CI concurrency + retry
  semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): register plan-eng/design-review skills in PTY eval containers

The extended plan-mode-no-op smoke invokes /plan-eng-review and
/plan-design-review, but the fresh CI containers registered only
office-hours and plan-ceo-review — both new runs would return
'Unknown command' and fail every PR's gate job (Codex structured
review P1, verified against evals.yml). Registration loops, the
dangling-target fail-fast list, and the frontmatter checks (now a
loop over the same skill list, so the lists can't drift) all cover
the two skills.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(plan-eng/design-review): harden scope-gate exceptions against injection and ambiguity

Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex
cross-confirmation), applied to both gate templates + regen:

- Host-anchored mode signal: only the host's own system messages (plan-mode
  reminder or active plan file path) arm the auto-select; plan-shaped text
  inside pasted documents, tool results, or fetched pages does NOT count —
  injected content can't disarm the consent gate or nominate the target.
- Multiple plan candidates: the host-referenced plan file wins; still
  ambiguous means ask.
- The DIFFERENT-target override carries the passing-mention guard.
- Plan mode + explicitly named target + no drafted plan resolves to the
  named target instead of a contradictory re-ask.
- The numbered ask-path rules are qualified ('When no exception above
  applied:') so they no longer restate an unconditional MUST-ask that
  contradicts the exceptions.
- 'Whenever this gate does ask — in any mode — it is a hard STOP.'
- Shared preamble: 'any AskUserQuestion the skill fires is the workflow
  operating within plan mode' (was 'the first AskUserQuestion is the
  workflow entering plan mode', which framed the opposite of the bypass);
  regenerates every skill.
- Ceilings ratcheted with attribution: plan-eng union ratio 1.10,
  investigate 1.10 (the ~250B shared-preamble reword lands the
  closest-to-ceiling skill at 1.092).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.62.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pty-runner): active-render gate veto in the floor check + honest periodic-wiring docs

Codex re-review P2s on the fix wave, both verified:

- A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting,
  no further output) was vetoed by the blanket tail exclusion until timeout.
  The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last
  cursor menu, so only a pending GATE menu vetoes; the judge fallback shares
  the same check. Residual (documented): prose gate + prose finding inside
  one tail — floors run the native-menu path in practice.
- The four demoted periodic tests are not in evals-periodic.yml's explicit
  matrix (a named instance of the pre-existing periodic-orphans TODO), so
  they run locally/manually until the PTY-capable periodic job lands.
  CHANGELOG claim softened accordingly; TODO filed with the wiring recipe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.62.0.0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: apply codex doc-review fixes for v1.62.0.0

- CLAUDE.md: scope the tier-alignment invariant claim (mapped files
  enforced, unmapped files reported)
- docs/skills.md: document the plan-mode auto-select scope gate for
  /plan-eng-review and /plan-design-review
- evals.yml: fix stale comment (PTY smokes register four skills, not two)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: refresh ship golden baselines for the plan-mode preamble reword

The generate-completion-status.ts wording change ('any AskUserQuestion the
skill fires…') intentionally regenerates every SKILL.md; the byte-compare
goldens carry the generator's output and refresh with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ship): custom-hooks-path detection false-negatives on git worktrees

The pre-push guard's HOOKS_IN_GIT_DIR check compared the hooks dir against
--absolute-git-dir, which in a linked worktree is .git/worktrees/<name>
while hooks resolve to the COMMON .git/hooks — so every Conductor worktree
read as a 'custom hooks path' and the consented guard install was skipped.
Match against the resolved --git-common-dir too (with a /nonexistent
fallback so a failed resolution can't collapse the case pattern into
match-everything). Verified live: this worktree now reports yes (was no),
and the main checkout still reports yes. Goldens refreshed (--host all).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: changelog bullet for the worktree hooks-detection fix

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): give the plan-ceo plan-mode smoke real budget headroom

Measured 2026-08-11: a clean isolated pass took 295.7s against the 300s
inner budget (4s of margin) and the same test timed out at ~308s three
times under concurrent eval load — a budget-edge flake in the gate lane,
not a behavior regression (it passed isolated on both this branch and
main). Inner budget 300s -> 420s, outer bun timeout 360s -> 480s, and the
test file is now named in its own touchfiles dep list so the tier-alignment
invariant binds for it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): 300s budget floor for the two 90s design-consultation SDK tests

Root cause of PR #2533's e2e-design CI failure: design-consultation-preview
failed 3 attempts at 0 turns/$0.00/93s — the session was up but the model's
first completion queued past the 90s inner budget under concurrent API load
(11 matrix jobs; the sibling research test booted its first tool at 4s, so
this is API-side queuing, not CPU boot contention). The test was selected
only because touchfiles.ts is a global touchfile; the tested behavior is
untouched by this branch.

90s budgets cannot absorb one slow first completion. Both 90s tests in the
file move to the repo's saturated-runner standard (300s inner / 360s outer,
matching review-dashboard-via and retro-base-branch). Deliberately NOT
re-arming the runner's inner timer on first stream event: an audit found
~100 outer bun-timeout literals sized inner+30-60s that a re-arm would
silently break — the structural options are written up in TODOS.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-12 11:12:28 -07:00
committed by GitHub
co-authored by Claude Fable 5
parent 94993f7401
commit d078622b73
81 changed files with 1097 additions and 137 deletions
+8 -3
View File
@@ -164,7 +164,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
},
behavioral: 'plan',
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
maxSkeletonBytes: 67_000,
// +~1 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
maxSkeletonBytes: 68_000,
minUnionBytes: 70_000,
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
@@ -172,7 +173,10 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
// default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
// prose, replacing the smaller opt-in question) land this at ~6.6% over the
// v1.53.0.0 baseline. Headroom for those intentional additions.
maxSizeRatio: 1.08,
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
// hardening: host-anchored mode signal, precedence, passing-mention
// guards) and the plan-mode preamble reword land the union at 1.092.
maxSizeRatio: 1.10,
},
'plan-design-review': {
skill: 'plan-design-review',
@@ -189,7 +193,8 @@ export const CARVE_GUARDS: Record<string, CarveGuard> = {
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
maxSkeletonBytes: 88_000,
// +~1.3 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
maxSkeletonBytes: 89_000,
minUnionBytes: 70_000,
mustContain: ['design', 'visual'],
maxSizeRatio: 1.07,
@@ -0,0 +1,143 @@
/**
* Scope-gate floor-exclusion regression pins (free, static).
*
* runPlanSkillFloorCheck's acceptance condition changed with the plan-mode
* auto-select-B work: a render only satisfies the finding floor when
*
* (isNumberedOptionListVisible(visible) || isProseAUQVisible(visible))
* && !isPermissionDialogVisible(tail)
* && !isScopeGateQuestionVisible(tail) // <- new exclusion
*
* where tail = visible.slice(-TAIL_SCAN_BYTES). The composition lives inline
* in the paid PTY loop, so these tests pin the load-bearing behavior of each
* detector on the exact render shapes the floor passes them:
*
* 1. Both scope-gate render forms (native numbered UI, prose lettered
* fallback) trip the acceptance detectors — WITHOUT the exclusion the
* gate would trivially satisfy the floor inside the 3s pre-target
* window. The exclusion must catch both forms.
* 2. A genuine finding-driven AskUserQuestion must NOT trip the exclusion,
* or the floor becomes unsatisfiable.
* 3. The exclusion is TAIL-scoped by design: an early gate render that has
* scrolled past TAIL_SCAN_BYTES must not suppress a later real finding
* AskUserQuestion.
*
* Also closes the untested OR-branch of isScopeGateAutoSelectVisible: the
* fully-collapsed hyphen-less 'autoselectedb' form.
*/
import { describe, test, expect } from 'bun:test';
import {
TAIL_SCAN_BYTES,
isNumberedOptionListVisible,
isProseAUQVisible,
isPermissionDialogVisible,
isScopeGateQuestionVisible,
isScopeGateAutoSelectVisible,
parseNumberedOptions,
} from './claude-pty-runner';
// The gate's native AskUserQuestion render (numbered options + cursor) —
// what fires inside the floor check's 3s window before the seed arrives.
const GATE_NATIVE_RENDER = `
What should I review?
1. The current branch diff — the work in progress on this branch.
2. A plan or design doc I'll paste or point you to.
3. A specific file, directory, or path.
`;
// The gate's prose fallback render (lettered options under --disallowedTools).
const GATE_PROSE_RENDER = `
What should I review?
A) The current branch diff — the work in progress on this branch.
B) A plan or design doc I'll paste or point you to.
C) A specific file, directory, or path.
Recommendation: A when a branch diff exists, otherwise B.
`;
// A genuine finding-driven AskUserQuestion — the render the floor MEASURES.
const FINDING_AUQ_RENDER = `
Finding 1: the plan reimplements test sharding that Bun provides natively.
1. Use Bun's native --shard flag (recommended)
2. Keep the custom scheduler as planned
3. Defer this decision to implementation
`;
describe('floor-check scope-gate exclusion (acceptance-condition regression)', () => {
test('native gate render trips the acceptance detector — the exclusion is load-bearing', () => {
// Pre-exclusion, this render satisfied the floor by itself.
expect(isNumberedOptionListVisible(GATE_NATIVE_RENDER)).toBe(true);
expect(isPermissionDialogVisible(GATE_NATIVE_RENDER)).toBe(false);
// The new exclusion catches it.
expect(isScopeGateQuestionVisible(GATE_NATIVE_RENDER)).toBe(true);
});
test('prose gate render trips the prose-AUQ arm — the exclusion catches that form too', () => {
expect(isProseAUQVisible(GATE_PROSE_RENDER)).toBe(true);
expect(isPermissionDialogVisible(GATE_PROSE_RENDER)).toBe(false);
expect(isScopeGateQuestionVisible(GATE_PROSE_RENDER)).toBe(true);
});
test('a genuine finding AskUserQuestion is NOT excluded — the floor stays satisfiable', () => {
expect(isNumberedOptionListVisible(FINDING_AUQ_RENDER)).toBe(true);
expect(isPermissionDialogVisible(FINDING_AUQ_RENDER)).toBe(false);
expect(isScopeGateQuestionVisible(FINDING_AUQ_RENDER)).toBe(false);
});
test('tail-scoping: an early gate render scrolled out of the tail does not suppress a later finding AUQ', () => {
// Gate render, then >TAIL_SCAN_BYTES of review output, then the real
// finding AskUserQuestion — the shape the TAIL-scoped exclusion exists for.
const filler = 'Reading the plan and auditing the design system.\n'.repeat(
Math.ceil(TAIL_SCAN_BYTES / 48) + 4,
);
const visible = GATE_NATIVE_RENDER + filler + FINDING_AUQ_RENDER;
const tail = visible.slice(-TAIL_SCAN_BYTES);
// Full buffer still remembers the gate (scrollback)…
expect(isScopeGateQuestionVisible(visible)).toBe(true);
// …but the floor's exclusion looks only at the tail, which is clean:
expect(isScopeGateQuestionVisible(tail)).toBe(false);
// and the acceptance arm (full-buffer scan) sees the finding AUQ.
expect(isNumberedOptionListVisible(visible)).toBe(true);
expect(isPermissionDialogVisible(tail)).toBe(false);
});
test('a gate render inside the tail IS suppressed (no false floor pass)', () => {
const tail = GATE_NATIVE_RENDER.slice(-TAIL_SCAN_BYTES);
expect(isScopeGateQuestionVisible(tail)).toBe(true);
});
test('active-render veto: a finding AUQ close after the gate is NOT vetoed (codex P2 re-review)', () => {
// The finding menu renders <TAIL_SCAN_BYTES after the gate, then the
// model waits (no further output). A blanket tail veto would suppress
// this until timeout; the active-render veto anchors on the LAST cursor
// menu, which is the finding AUQ, so the floor is satisfiable.
const visible = GATE_NATIVE_RENDER + '\nAuditing the plan…\n' + FINDING_AUQ_RENDER;
const activeMenu = parseNumberedOptions(visible);
expect(activeMenu.length).toBeGreaterThan(0);
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
expect(gateIsActiveRender).toBe(false);
});
test('active-render veto: the gate as the pending menu IS vetoed', () => {
const visible = 'booting…\n' + GATE_NATIVE_RENDER;
const activeMenu = parseNumberedOptions(visible);
expect(activeMenu.length).toBeGreaterThan(0);
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
expect(gateIsActiveRender).toBe(true);
});
});
describe('isScopeGateAutoSelectVisible collapsed hyphen-less branch', () => {
test("matches the fully-collapsed 'autoselectedb' form (hyphen lost in TTY reflow)", () => {
const sample = 'Scopegate:planmode—autoselectedB(reviewingPLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('hyphen-less token without the announcement prefix stays false', () => {
const sample = 'The agent autoselectedB from the menu without announcing a scope gate decision.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
});
+162 -10
View File
@@ -618,6 +618,55 @@ export function isProseAUQVisible(visible: string): boolean {
return false;
}
// ---------------------------------------------------------------------------
// Scope-gate render detectors (plan-eng-review / plan-design-review)
// ---------------------------------------------------------------------------
//
// Both anchor on the RENDER SHAPE, not bare keywords, so model narration
// about the gate ("normally I'd ask what should I review…") stays false.
// Matching is whitespace-squished + lowercased because stripAnsi collapses
// TTY cursor-positioning escapes unpredictably (the same failure mode the
// Pattern-4/5 collapsed-form handling above exists for).
/**
* True when the scope-gate QUESTION is actually rendered: the question text
* plus option A's body text. Option-body anchoring (not `A)`/`B)` markers)
* because native AskUserQuestion renders NUMBERED options in the TTY while
* the --disallowedTools prose fallback renders lettered ones — the option
* body appears in both renders; narration rarely quotes both the question
* and an option body.
*/
export function isScopeGateQuestionVisible(visible: string): boolean {
const squished = visible.replace(/\s+/g, '').toLowerCase();
return squished.includes('whatshouldireview') && squished.includes('currentbranchdiff');
}
/**
* True when the plan-mode auto-select announcement is rendered:
* "Scope gate: plan mode — auto-selected B (reviewing <target>)."
* Requires BOTH the announcement prefix and an auto-select-B token so
* narration ("in plan mode I'd auto-select B") stays false. The token is
* tense-tolerant (selected/selecting/selects) because the smokes assert
* must-be-TRUE on it — a semantically-perfect paraphrase must not fail a
* paid run — while the prefix stays exact so paraphrase narration without
* the announcement frame stays false. A prefix immediately preceded by a
* quote character is a QUOTATION (e.g. the model explaining why it is NOT
* announcing), not a render — the announcement line itself never renders
* quoted.
*/
export function isScopeGateAutoSelectVisible(visible: string): boolean {
const squished = visible.replace(/\s+/g, '').toLowerCase();
const QUOTES = ['"', "'", '`', '“', ''];
const re = /scopegate:planmode/g;
let m: RegExpExecArray | null;
while ((m = re.exec(squished)) !== null) {
const before = m.index > 0 ? squished[m.index - 1]! : '';
if (QUOTES.includes(before)) continue; // quoted occurrence — narration, keep scanning
if (/auto-?select(?:ed|ing|s)?b/.test(squished.slice(m.index))) return true;
}
return false;
}
/**
* Parse a rendered numbered-option list out of the visible TTY text.
*
@@ -1476,10 +1525,20 @@ export interface PlanSkillObservation {
* "Ready to execute" confirmation
* - 'silent_write' — a Write/Edit landed BEFORE any prompt, to a path
* outside the sanctioned plan/project directories
* - 'wrote_findings_before_asking' — strictPlanWrites only (seeded runs):
* the plan file was rewritten with findings before any
* AskUserQuestion render (the May-2026 transcript bug)
* - 'exited' — claude process died before any of the above
* - 'timeout' — none of the above within budget
*/
outcome: 'asked' | 'auto_decided' | 'plan_ready' | 'silent_write' | 'exited' | 'timeout';
outcome:
| 'asked'
| 'auto_decided'
| 'plan_ready'
| 'silent_write'
| 'wrote_findings_before_asking'
| 'exited'
| 'timeout';
/** Human-readable summary. */
summary: string;
/** Visible terminal text since the slash command was sent (last 2KB). */
@@ -1516,6 +1575,28 @@ export interface PlanSkillObservation {
* Haiku judge fallback rather than the regex detector.
*/
waitingEverObserved?: boolean;
/**
* High-water-mark flag: did the scope-gate QUESTION ("What should I
* review?" plus option-body text) ever render during the run? Same
* lossy-2KB-evidence rationale as proseAUQEverObserved. The plan-mode
* smokes assert this stays false (gate bypassed via auto-select B); the
* no-op regression asserts it fires outside plan mode.
*/
scopeGateQuestionObserved?: boolean;
/**
* High-water-mark flag: did the plan-mode auto-select announcement
* ("Scope gate: plan mode — auto-selected B …") ever render? The
* plan-mode smokes assert true; the no-op regression asserts false.
*/
scopeGateAutoSelectObserved?: boolean;
/**
* High-water map for opts.trackTokens: token → did it EVER appear in the
* cumulative visible buffer? Consumption asserts (e.g. "the pasted target's
* distinctive token shows up in the review output") must not depend on the
* lossy 2KB evidence tail — plan-file fallbacks are unreachable outside
* plan mode (extractPlanFilePath only matches plan-mode save renders).
*/
tokensObserved?: Record<string, boolean>;
}
/**
@@ -1576,6 +1657,10 @@ export async function runPlanSkillObservation(opts: {
/** Override the spawned model. Defaults via launchClaudePty's chain
* (opts.model ?? EVALS_MODEL ?? 'claude-sonnet-4-6'). */
model?: string;
/** Literal tokens to track as high-water marks over the CUMULATIVE visible
* buffer (case-sensitive). Results land in obs.tokensObserved. Use for
* consumption asserts that must survive the 2KB evidence tail. */
trackTokens?: string[];
}): Promise<PlanSkillObservation> {
const startedAt = Date.now();
const session = await launchClaudePty({
@@ -1619,6 +1704,21 @@ export async function runPlanSkillObservation(opts: {
// even if the current state is 'working'.
let proseAUQEverObserved = false;
let waitingEverObserved = false;
let scopeGateQuestionObserved = false;
let scopeGateAutoSelectObserved = false;
const tokensObserved: Record<string, boolean> = {};
for (const t of opts.trackTokens ?? []) tokensObserved[t] = false;
// Single source for the high-water flags at EVERY return site. Hand-
// spreading them per-site already drifted once (the judge-waiting return
// omitted the prose/waiting flags); a site that forgets a must-stay-false
// flag makes `obs.flag ?? false` negative assertions pass vacuously.
const highWaterFlags = () => ({
proseAUQEverObserved,
waitingEverObserved,
scopeGateQuestionObserved,
scopeGateAutoSelectObserved,
...(opts.trackTokens?.length ? { tokensObserved } : {}),
});
const JUDGE_AFTER_MS = 60_000;
const JUDGE_INTERVAL_MS = 30_000;
while (Date.now() - start < budgetMs) {
@@ -1631,6 +1731,7 @@ export async function runPlanSkillObservation(opts: {
summary: `claude exited (code=${session.exitCode()}) before reaching a terminal outcome`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
if (visible.includes('Unknown command:')) {
@@ -1639,6 +1740,7 @@ export async function runPlanSkillObservation(opts: {
summary: `claude rejected /${opts.skillName} as unknown command (skill not registered in this cwd)`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
@@ -1652,6 +1754,18 @@ export async function runPlanSkillObservation(opts: {
tag: 'prose-auq-surfaced',
});
}
// Scope-gate render tracking (same high-water shape). Full-run
// detection matters because the 2KB evidence tail usually scrolls
// past the gate render before the outcome fires.
if (!scopeGateQuestionObserved && isScopeGateQuestionVisible(visible)) {
scopeGateQuestionObserved = true;
}
if (!scopeGateAutoSelectObserved && isScopeGateAutoSelectVisible(visible)) {
scopeGateAutoSelectObserved = true;
}
for (const t of opts.trackTokens ?? []) {
if (!tokensObserved[t] && visible.includes(t)) tokensObserved[t] = true;
}
const classified = classifyVisible(visible, {
strictPlanWrites: !!opts.initialPlanContent,
@@ -1661,8 +1775,7 @@ export async function runPlanSkillObservation(opts: {
...classified,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
// Capture the plan file path on any outcome where one may have been
// written. Gating only on 'plan_ready' missed two cases: (1) the
@@ -1693,6 +1806,7 @@ export async function runPlanSkillObservation(opts: {
summary: `LLM judge: ${lastJudgeVerdict.reasoning} (state=waiting after ${Math.round(elapsed / 1000)}s)`,
evidence: visible.slice(-2000),
elapsedMs: Date.now() - startedAt,
...highWaterFlags(),
};
}
}
@@ -1714,8 +1828,7 @@ export async function runPlanSkillObservation(opts: {
: ''),
evidence: finalVisible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
}
return {
@@ -1727,8 +1840,7 @@ export async function runPlanSkillObservation(opts: {
: ''),
evidence: finalVisible.slice(-2000),
elapsedMs: Date.now() - startedAt,
proseAUQEverObserved,
waitingEverObserved,
...highWaterFlags(),
};
} finally {
await session.close();
@@ -2099,11 +2211,23 @@ export async function runPlanSkillFloorCheck(opts: {
const start = Date.now();
let lastJudgeAt = 0;
let lastJudgeVerdict: PtyStateVerdict | null = null;
// Positional anchor for the scope-gate exclusion. The visible buffer is
// append-only (old renders never leave scrollback), so a gate question
// rendered in the 3s pre-target window would keep satisfying the
// full-buffer acceptance checks forever while a tail-only exclusion
// stops seeing it after ~TAIL_SCAN_BYTES of output — a vacuous
// auq_observed (found independently by 4 review passes). Once the gate
// render is seen, acceptance only counts AUQ renders in content APPENDED
// after that point.
let gateSeenIdx = -1;
const JUDGE_AFTER_MS = 60_000;
const JUDGE_INTERVAL_MS = 30_000;
while (Date.now() - start < timeoutMs) {
await Bun.sleep(2000);
const visible = session.visibleSince(since);
if (gateSeenIdx === -1 && isScopeGateQuestionVisible(visible)) {
gateSeenIdx = visible.length;
}
if (session.exited()) {
return {
@@ -2129,10 +2253,34 @@ export async function runPlanSkillFloorCheck(opts: {
// OR via prose-rendered options under --disallowedTools when no MCP
// variant is callable (isProseAUQVisible). Both surface the question
// to the user; the bug we're catching is "fired zero AUQs."
//
// Scope-gate renders do NOT count: the gate's "What should I review?"
// can fire inside the 3s pre-target window and would trivially satisfy
// the floor, but the floor measures FINDING-driven questions. Once a
// gate render has been seen, acceptance scans only the content APPENDED
// after it (positional anchor above) — the buffer is append-only, so a
// whole-buffer acceptance would keep matching the stale gate render
// forever.
//
// The gate veto is ACTIVE-RENDER-aware, not blanket-tail: when a
// numbered menu is up, parseNumberedOptions anchors on the LAST cursor
// line, so we veto only when the pending menu IS the gate — a finding
// AUQ that renders within TAIL_SCAN_BYTES of the gate (model waiting,
// no further output) still satisfies the floor. Prose renders have no
// cursor anchor, so the prose path falls back to the tail check
// (accepted residual: prose gate + prose finding inside one tail can
// suppress until timeout; floors run the native-menu path in practice).
const tail = visible.slice(-TAIL_SCAN_BYTES);
const acceptWindow = gateSeenIdx === -1 ? visible : visible.slice(gateSeenIdx);
const activeMenu = parseNumberedOptions(visible);
const gateIsActiveRender =
activeMenu.length > 0
? activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label))
: isScopeGateQuestionVisible(tail);
if (
(isNumberedOptionListVisible(visible) || isProseAUQVisible(visible)) &&
!isPermissionDialogVisible(tail)
(isNumberedOptionListVisible(acceptWindow) || isProseAUQVisible(acceptWindow)) &&
!isPermissionDialogVisible(tail) &&
!gateIsActiveRender
) {
return {
auqObserved: true,
@@ -2154,7 +2302,11 @@ export async function runPlanSkillFloorCheck(opts: {
lastJudgeAt = Date.now();
logPtySnapshot(visible, { testName: opts.skillName, elapsedMs: elapsed, tag: 'floor-judge-tick' });
lastJudgeVerdict = judgePtyState(visible, { testName: opts.skillName });
if (lastJudgeVerdict.state === 'waiting') {
// The judge can't tell a scope-gate question from a finding question,
// so a 'waiting' verdict while the gate menu is the pending render
// must NOT satisfy the floor — same active-render exclusion as the
// regex path.
if (lastJudgeVerdict.state === 'waiting' && !gateIsActiveRender) {
return {
auqObserved: true,
outcome: 'auq_observed',
+109
View File
@@ -28,6 +28,8 @@ import {
isPermissionDialogVisible,
isNumberedOptionListVisible,
isProseAUQVisible,
isScopeGateQuestionVisible,
isScopeGateAutoSelectVisible,
isPlanReadyVisible,
parseNumberedOptions,
classifyVisible,
@@ -194,6 +196,113 @@ describe('isNumberedOptionListVisible', () => {
});
});
describe('scope-gate render detectors', () => {
// The verbatim announcement string from the plan-eng/plan-design SKILL.md
// templates. If the template rewording drifts, THIS fixture fails first —
// before the paid plan-mode smokes silently degrade to vacuous asserts.
const TEMPLATE_ANNOUNCEMENT =
'Scope gate: plan mode — auto-selected B (reviewing <target>).';
describe('isScopeGateQuestionVisible', () => {
test('matches the clean prose gate render (question + option bodies)', () => {
const sample = `
What should I review?
A) The current branch diff — the work in progress on this branch.
B) A plan or design doc I'll paste or point you to.
C) A specific file, directory, or path.
Recommendation: A when a branch diff exists, otherwise B.
`;
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('matches the native numbered render (no lettered markers)', () => {
const sample = `
What should I review?
1. The current branch diff — the work in progress on this branch.
2. A plan or design doc I'll paste or point you to.
3. A specific file, directory, or path.
`;
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('matches the PTY-collapsed render (stripAnsi squished spaces)', () => {
const sample = 'WhatshouldIreview?A)Thecurrentbranchdiff—theworkinprogress';
expect(isScopeGateQuestionVisible(sample)).toBe(true);
});
test('stays false on narration quoting only the question', () => {
const sample =
"Normally I'd ask 'What should I review?' but plan mode is active, so I'm proceeding.";
expect(isScopeGateQuestionVisible(sample)).toBe(false);
});
test('stays false on unrelated review prose', () => {
const sample = 'I will review the current branch diff and report findings.';
expect(isScopeGateQuestionVisible(sample)).toBe(false);
});
});
describe('isScopeGateAutoSelectVisible', () => {
test('matches the verbatim template announcement', () => {
expect(isScopeGateAutoSelectVisible(TEMPLATE_ANNOUNCEMENT)).toBe(true);
});
test('matches a real announcement with a concrete target', () => {
const sample =
'Scope gate: plan mode — auto-selected B (reviewing ~/.claude/plans/my-feature.md). Running the Design Doc Check next.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('matches the PTY-collapsed announcement', () => {
const sample = 'Scopegate:planmode—auto-selectedB(reviewingPLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('stays false on narration about the behavior', () => {
const sample = "In plan mode I'd auto-select B and review the active plan.";
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('stays false on a VERBATIM QUOTE of the announcement (negation narration)', () => {
// The exact announcement line sits quoted in the skill context, so a
// model explaining why it is NOT firing it can reproduce it byte-exact
// inside quotes — that must not trip a must-stay-false assert.
const sample =
'Not in plan mode, so I won\'t announce "Scope gate: plan mode — auto-selected B (reviewing <target>)." and will ask instead.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('a later real render still matches after an earlier quoted mention', () => {
const sample =
'Earlier I said I would render "Scope gate: plan mode — auto-selected B (…)" and now:\n' +
'Scope gate: plan mode — auto-selected B (reviewing PLAN.md).';
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
});
test('matches tense paraphrases WITH the announcement prefix (auto-selecting / auto-selects)', () => {
expect(
isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selecting B (reviewing the drafted plan).'),
).toBe(true);
expect(isScopeGateAutoSelectVisible('Scope gate: plan mode — auto-selects B.')).toBe(true);
});
test('stays false on tense paraphrases WITHOUT the announcement prefix', () => {
expect(isScopeGateAutoSelectVisible('Auto-selecting B since we are in plan mode.')).toBe(false);
});
test('stays false on AUTO_DECIDE preamble output', () => {
const sample = 'Auto-decided scope question → B (your preference). Change with /plan-tune.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
test('stays false on a bare "selected B" without the announcement prefix', () => {
const sample = 'I selected B as the review target.';
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
});
});
});
describe('isProseAUQVisible', () => {
test('matches 4 lettered options A) B) C) D) at line starts (plan-eng prose AUQ shape)', () => {
const sample = `
+4 -1
View File
@@ -234,7 +234,10 @@ const MONOLITH_INVARIANTS: ParityInvariant[] = [
// cross-session decision-memory nudge) lands this skill just over the strict 1.05;
// headroom for the shared preamble additions (matches the carved-skill overrides).
// v1.2.0 activation lift adds the first-run-guidance section on top.
maxSizeRatio: 1.09,
// 1.09 → 1.10: the plan-mode preamble reword (scope-gate auto-select-B
// change) adds ~250 B to every skill's shared preamble; investigate was
// the closest to its ceiling (landed 1.092).
maxSizeRatio: 1.10,
minBytes: 30_000,
},
{
+10 -4
View File
@@ -98,11 +98,17 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
// include question-tuning.ts and generate-ask-user-format.ts because the
// AUTO_DECIDE preamble injection lives there and changes can flip the
// regression test outcome between 'asked' and 'auto_decided'.
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-ceo-review-plan-mode': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-ceo-plan-mode.test.ts'],
'plan-eng-review-plan-mode': ['plan-eng-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-eng-plan-mode.test.ts'],
'plan-design-review-plan-mode': ['plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts'],
'plan-devex-review-plan-mode': ['plan-devex-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts'],
'plan-mode-no-op': ['plan-ceo-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts'],
// Covers ceo (preamble misfire) + eng/design (scope-gate bypass must not
// fire outside plan mode) + the named-target exception case. 4 PTY runs;
// in CI these run CONCURRENT with the rest of the pty-plan-smoke suite
// (--max-concurrency + --retry 2), so worst-case cost is ~3x a single
// pass of each, sharing the API budget with sibling tests — not the
// sequential ~+10min a local read suggests.
'plan-mode-no-op': ['plan-ceo-review/**', 'plan-eng-review/**', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/preamble.ts', 'test/helpers/claude-pty-runner.ts', 'test/skill-e2e-plan-mode-no-op.test.ts'],
// v1.21+ AskUserQuestion-blocked regression tests — Conductor launches
// claude with `--disallowedTools AskUserQuestion --permission-mode default`