mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-13 09:40:21 +02:00
* fix(evals): align plan-eng/design plan-mode + finding-floor smokes to their declared periodic tier The #2077 demotion of these four stochastic tests to 'periodic' was inert: E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate', so they kept running in the blocking gate lane and never in the weekly lane. Flip the four self-gates to 'periodic' (headers/describe labels updated), add a free static tier-alignment invariant test (dep-list filename mapping; unmapped self-gated files are reported, never silently skipped), and name the two plan-mode test files in their own touchfiles dep lists so the invariant binds for them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(pty-runner): scope-gate question/auto-select detectors + observation flags Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5 collapsed-form handling): isScopeGateQuestionVisible requires the question text PLUS option-body text (native AskUserQuestion renders numbered options, prose fallback renders lettered — the option body appears in both; narration doesn't), and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the selected-B token. runPlanSkillObservation gains scopeGateQuestionObserved / scopeGateAutoSelectObserved high-water flags (attached at every return path) so paid smokes can assert gate behavior across the whole run instead of the lossy 2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven questions, and the gate could fire inside the 3s pre-target window. Unit fixtures pin clean/native/collapsed positives, narration negatives, and the verbatim template announcement string (template rewording fails here first, before the paid smokes degrade to vacuous asserts). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(plan-eng/design-review): auto-select B in plan mode at the scope gate In plan mode the scope gate's "What should I review? A/B/C" question is pure friction: there is no branch diff and the target is the plan being drafted. Both gates gain an ordered exceptions block, checked BEFORE asking: 1. Plan mode → auto-select B: review the active plan (in context or pasted), announce it in one line ("Scope gate: plan mode — auto-selected B (reviewing <target>)") so the user can interrupt; an explicitly different user-named target still wins; no plan drafted yet → ask as normal. 2. User-named target (outside plan mode): explicit-only — a path, a pasted doc, or the literal words "branch diff". A passing mention is not naming; when in doubt, ask. Outside plan mode with no explicitly-named target, nothing changes. Plan-mode is checked FIRST because the PTY harness seeds drafts as pasted user messages (claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic. Pinning: seeded plan-mode smokes assert no gate render + announcement rendered (eng test 2; new design seeded test); plan-mode-no-op extends to eng/design (bypass must not misfire outside plan mode; first question must be the gate) plus a named-target case proving the pasted target is consumed; a drift-guard asserts the two hand-duplicated exceptions blocks stay identical modulo the two variant slots and carry the announcement string the detectors pin. Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): skip the scope gate when following loaded review skills autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim, and its section skip list omitted the scope gate — so autoplan ingested a hard-STOP AskUserQuestion that contradicts its every-question-auto-decides contract. One skip-list line fixes it; a static toContain pin in skill-validation keeps the entry load-bearing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: file scope-gate resolver-extraction TODO (eng-review D5 follow-up) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): positional floor exclusion, flag builder, outcome union, token tracking Review-army + adversarial findings on the scope-gate observability work, all verified before fixing: - Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate exclusion scanned only the 1500-byte tail, so an early gate render satisfied the floor vacuously once ~1.5KB of output accumulated (found independently by 4 review passes; predicate reproduced). Acceptance now scans only content APPENDED after the first gate render (positional anchor), and the LLM-judge 'waiting' shortcut no longer fires while the gate menu is the pending render. - High-water flags are built once and spread at every return path — the hand-spread pattern had already drifted (judge-waiting return omitted two flags), which made must-stay-false asserts vacuous on those paths. - isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases) and quoted-occurrence rejection (a model verbatim-quoting the announcement while declining must not trip must-stay-FALSE asserts). Fixtures added for both directions. - PlanSkillObservation outcome union gains 'wrote_findings_before_asking' (returned at runtime via classifyVisible but missing from the type). - trackTokens/tokensObserved: cumulative-buffer token high-water for consumption asserts (the 2KB evidence tail is lossy and the plan-file fallback is unreachable outside plan mode). - New scope-gate-floor unit pins (from the ship coverage audit): both gate render forms trip acceptance and exclusion; a genuine finding AUQ is not excluded; tail-scoping semantics pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(evals): harden no-op asserts, close tier-invariant fail-open holes, pin gate question strings - no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the outcome==='asked' conditional let a silent-bypass plan_ready run sail through); eng/design cases force --disallowedTools so the pinned prose shape is contractual rather than hoping native AUQ renders match; the named-target case uses trackTokens for consumption and lists wrote_findings_before_asking in its diagnostic throw branch. - tier-alignment invariant: both quote styles matched; zero-self-gate, mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED instead of silently skipped (the fail-open holes three reviewers found). - drift-guard: the generated gate menus must carry the exact question/option strings the PTY question detector anchors on — free CI fails before the paid smokes can go vacuous on a menu reword. - touchfiles: corrected the no-op cost note for CI concurrency + retry semantics. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): register plan-eng/design-review skills in PTY eval containers The extended plan-mode-no-op smoke invokes /plan-eng-review and /plan-design-review, but the fresh CI containers registered only office-hours and plan-ceo-review — both new runs would return 'Unknown command' and fail every PR's gate job (Codex structured review P1, verified against evals.yml). Registration loops, the dangling-target fail-fast list, and the frontmatter checks (now a loop over the same skill list, so the lists can't drift) all cover the two skills. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(plan-eng/design-review): harden scope-gate exceptions against injection and ambiguity Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex cross-confirmation), applied to both gate templates + regen: - Host-anchored mode signal: only the host's own system messages (plan-mode reminder or active plan file path) arm the auto-select; plan-shaped text inside pasted documents, tool results, or fetched pages does NOT count — injected content can't disarm the consent gate or nominate the target. - Multiple plan candidates: the host-referenced plan file wins; still ambiguous means ask. - The DIFFERENT-target override carries the passing-mention guard. - Plan mode + explicitly named target + no drafted plan resolves to the named target instead of a contradictory re-ask. - The numbered ask-path rules are qualified ('When no exception above applied:') so they no longer restate an unconditional MUST-ask that contradicts the exceptions. - 'Whenever this gate does ask — in any mode — it is a hard STOP.' - Shared preamble: 'any AskUserQuestion the skill fires is the workflow operating within plan mode' (was 'the first AskUserQuestion is the workflow entering plan mode', which framed the opposite of the bypass); regenerates every skill. - Ceilings ratcheted with attribution: plan-eng union ratio 1.10, investigate 1.10 (the ~250B shared-preamble reword lands the closest-to-ceiling skill at 1.092). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.62.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): active-render gate veto in the floor check + honest periodic-wiring docs Codex re-review P2s on the fix wave, both verified: - A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting, no further output) was vetoed by the blanket tail exclusion until timeout. The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last cursor menu, so only a pending GATE menu vetoes; the judge fallback shares the same check. Residual (documented): prose gate + prose finding inside one tail — floors run the native-menu path in practice. - The four demoted periodic tests are not in evals-periodic.yml's explicit matrix (a named instance of the pre-existing periodic-orphans TODO), so they run locally/manually until the PTY-capable periodic job lands. CHANGELOG claim softened accordingly; TODO filed with the wiring recipe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.62.0.0 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply codex doc-review fixes for v1.62.0.0 - CLAUDE.md: scope the tier-alignment invariant claim (mapped files enforced, unmapped files reported) - docs/skills.md: document the plan-mode auto-select scope gate for /plan-eng-review and /plan-design-review - evals.yml: fix stale comment (PTY smokes register four skills, not two) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh ship golden baselines for the plan-mode preamble reword The generate-completion-status.ts wording change ('any AskUserQuestion the skill fires…') intentionally regenerates every SKILL.md; the byte-compare goldens carry the generator's output and refresh with it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): custom-hooks-path detection false-negatives on git worktrees The pre-push guard's HOOKS_IN_GIT_DIR check compared the hooks dir against --absolute-git-dir, which in a linked worktree is .git/worktrees/<name> while hooks resolve to the COMMON .git/hooks — so every Conductor worktree read as a 'custom hooks path' and the consented guard install was skipped. Match against the resolved --git-common-dir too (with a /nonexistent fallback so a failed resolution can't collapse the case pattern into match-everything). Verified live: this worktree now reports yes (was no), and the main checkout still reports yes. Goldens refreshed (--host all). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: changelog bullet for the worktree hooks-detection fix Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): give the plan-ceo plan-mode smoke real budget headroom Measured 2026-08-11: a clean isolated pass took 295.7s against the 300s inner budget (4s of margin) and the same test timed out at ~308s three times under concurrent eval load — a budget-edge flake in the gate lane, not a behavior regression (it passed isolated on both this branch and main). Inner budget 300s -> 420s, outer bun timeout 360s -> 480s, and the test file is now named in its own touchfiles dep list so the tier-alignment invariant binds for it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): 300s budget floor for the two 90s design-consultation SDK tests Root cause of PR #2533's e2e-design CI failure: design-consultation-preview failed 3 attempts at 0 turns/$0.00/93s — the session was up but the model's first completion queued past the 90s inner budget under concurrent API load (11 matrix jobs; the sibling research test booted its first tool at 4s, so this is API-side queuing, not CPU boot contention). The test was selected only because touchfiles.ts is a global touchfile; the tested behavior is untouched by this branch. 90s budgets cannot absorb one slow first completion. Both 90s tests in the file move to the repo's saturated-runner standard (300s inner / 360s outer, matching review-dashboard-via and retro-base-branch). Deliberately NOT re-arming the runner's inner timer on first stream event: an audit found ~100 outer bun-timeout literals sized inner+30-60s that a re-arm would silently break — the structural options are written up in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
144 lines
6.5 KiB
TypeScript
144 lines
6.5 KiB
TypeScript
/**
|
||
* Scope-gate floor-exclusion regression pins (free, static).
|
||
*
|
||
* runPlanSkillFloorCheck's acceptance condition changed with the plan-mode
|
||
* auto-select-B work: a render only satisfies the finding floor when
|
||
*
|
||
* (isNumberedOptionListVisible(visible) || isProseAUQVisible(visible))
|
||
* && !isPermissionDialogVisible(tail)
|
||
* && !isScopeGateQuestionVisible(tail) // <- new exclusion
|
||
*
|
||
* where tail = visible.slice(-TAIL_SCAN_BYTES). The composition lives inline
|
||
* in the paid PTY loop, so these tests pin the load-bearing behavior of each
|
||
* detector on the exact render shapes the floor passes them:
|
||
*
|
||
* 1. Both scope-gate render forms (native numbered UI, prose lettered
|
||
* fallback) trip the acceptance detectors — WITHOUT the exclusion the
|
||
* gate would trivially satisfy the floor inside the 3s pre-target
|
||
* window. The exclusion must catch both forms.
|
||
* 2. A genuine finding-driven AskUserQuestion must NOT trip the exclusion,
|
||
* or the floor becomes unsatisfiable.
|
||
* 3. The exclusion is TAIL-scoped by design: an early gate render that has
|
||
* scrolled past TAIL_SCAN_BYTES must not suppress a later real finding
|
||
* AskUserQuestion.
|
||
*
|
||
* Also closes the untested OR-branch of isScopeGateAutoSelectVisible: the
|
||
* fully-collapsed hyphen-less 'autoselectedb' form.
|
||
*/
|
||
|
||
import { describe, test, expect } from 'bun:test';
|
||
import {
|
||
TAIL_SCAN_BYTES,
|
||
isNumberedOptionListVisible,
|
||
isProseAUQVisible,
|
||
isPermissionDialogVisible,
|
||
isScopeGateQuestionVisible,
|
||
isScopeGateAutoSelectVisible,
|
||
parseNumberedOptions,
|
||
} from './claude-pty-runner';
|
||
|
||
// The gate's native AskUserQuestion render (numbered options + cursor) —
|
||
// what fires inside the floor check's 3s window before the seed arrives.
|
||
const GATE_NATIVE_RENDER = `
|
||
What should I review?
|
||
|
||
❯ 1. The current branch diff — the work in progress on this branch.
|
||
2. A plan or design doc I'll paste or point you to.
|
||
3. A specific file, directory, or path.
|
||
`;
|
||
|
||
// The gate's prose fallback render (lettered options under --disallowedTools).
|
||
const GATE_PROSE_RENDER = `
|
||
What should I review?
|
||
A) The current branch diff — the work in progress on this branch.
|
||
B) A plan or design doc I'll paste or point you to.
|
||
C) A specific file, directory, or path.
|
||
Recommendation: A when a branch diff exists, otherwise B.
|
||
`;
|
||
|
||
// A genuine finding-driven AskUserQuestion — the render the floor MEASURES.
|
||
const FINDING_AUQ_RENDER = `
|
||
Finding 1: the plan reimplements test sharding that Bun provides natively.
|
||
|
||
❯ 1. Use Bun's native --shard flag (recommended)
|
||
2. Keep the custom scheduler as planned
|
||
3. Defer this decision to implementation
|
||
`;
|
||
|
||
describe('floor-check scope-gate exclusion (acceptance-condition regression)', () => {
|
||
test('native gate render trips the acceptance detector — the exclusion is load-bearing', () => {
|
||
// Pre-exclusion, this render satisfied the floor by itself.
|
||
expect(isNumberedOptionListVisible(GATE_NATIVE_RENDER)).toBe(true);
|
||
expect(isPermissionDialogVisible(GATE_NATIVE_RENDER)).toBe(false);
|
||
// The new exclusion catches it.
|
||
expect(isScopeGateQuestionVisible(GATE_NATIVE_RENDER)).toBe(true);
|
||
});
|
||
|
||
test('prose gate render trips the prose-AUQ arm — the exclusion catches that form too', () => {
|
||
expect(isProseAUQVisible(GATE_PROSE_RENDER)).toBe(true);
|
||
expect(isPermissionDialogVisible(GATE_PROSE_RENDER)).toBe(false);
|
||
expect(isScopeGateQuestionVisible(GATE_PROSE_RENDER)).toBe(true);
|
||
});
|
||
|
||
test('a genuine finding AskUserQuestion is NOT excluded — the floor stays satisfiable', () => {
|
||
expect(isNumberedOptionListVisible(FINDING_AUQ_RENDER)).toBe(true);
|
||
expect(isPermissionDialogVisible(FINDING_AUQ_RENDER)).toBe(false);
|
||
expect(isScopeGateQuestionVisible(FINDING_AUQ_RENDER)).toBe(false);
|
||
});
|
||
|
||
test('tail-scoping: an early gate render scrolled out of the tail does not suppress a later finding AUQ', () => {
|
||
// Gate render, then >TAIL_SCAN_BYTES of review output, then the real
|
||
// finding AskUserQuestion — the shape the TAIL-scoped exclusion exists for.
|
||
const filler = 'Reading the plan and auditing the design system.\n'.repeat(
|
||
Math.ceil(TAIL_SCAN_BYTES / 48) + 4,
|
||
);
|
||
const visible = GATE_NATIVE_RENDER + filler + FINDING_AUQ_RENDER;
|
||
const tail = visible.slice(-TAIL_SCAN_BYTES);
|
||
|
||
// Full buffer still remembers the gate (scrollback)…
|
||
expect(isScopeGateQuestionVisible(visible)).toBe(true);
|
||
// …but the floor's exclusion looks only at the tail, which is clean:
|
||
expect(isScopeGateQuestionVisible(tail)).toBe(false);
|
||
// and the acceptance arm (full-buffer scan) sees the finding AUQ.
|
||
expect(isNumberedOptionListVisible(visible)).toBe(true);
|
||
expect(isPermissionDialogVisible(tail)).toBe(false);
|
||
});
|
||
|
||
test('a gate render inside the tail IS suppressed (no false floor pass)', () => {
|
||
const tail = GATE_NATIVE_RENDER.slice(-TAIL_SCAN_BYTES);
|
||
expect(isScopeGateQuestionVisible(tail)).toBe(true);
|
||
});
|
||
|
||
test('active-render veto: a finding AUQ close after the gate is NOT vetoed (codex P2 re-review)', () => {
|
||
// The finding menu renders <TAIL_SCAN_BYTES after the gate, then the
|
||
// model waits (no further output). A blanket tail veto would suppress
|
||
// this until timeout; the active-render veto anchors on the LAST cursor
|
||
// menu, which is the finding AUQ, so the floor is satisfiable.
|
||
const visible = GATE_NATIVE_RENDER + '\nAuditing the plan…\n' + FINDING_AUQ_RENDER;
|
||
const activeMenu = parseNumberedOptions(visible);
|
||
expect(activeMenu.length).toBeGreaterThan(0);
|
||
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
|
||
expect(gateIsActiveRender).toBe(false);
|
||
});
|
||
|
||
test('active-render veto: the gate as the pending menu IS vetoed', () => {
|
||
const visible = 'booting…\n' + GATE_NATIVE_RENDER;
|
||
const activeMenu = parseNumberedOptions(visible);
|
||
expect(activeMenu.length).toBeGreaterThan(0);
|
||
const gateIsActiveRender = activeMenu.some((o) => /current\s*branch\s*diff/i.test(o.label));
|
||
expect(gateIsActiveRender).toBe(true);
|
||
});
|
||
});
|
||
|
||
describe('isScopeGateAutoSelectVisible collapsed hyphen-less branch', () => {
|
||
test("matches the fully-collapsed 'autoselectedb' form (hyphen lost in TTY reflow)", () => {
|
||
const sample = 'Scopegate:planmode—autoselectedB(reviewingPLAN.md).';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(true);
|
||
});
|
||
|
||
test('hyphen-less token without the announcement prefix stays false', () => {
|
||
const sample = 'The agent autoselectedB from the menu without announcing a scope gate decision.';
|
||
expect(isScopeGateAutoSelectVisible(sample)).toBe(false);
|
||
});
|
||
});
|