mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-13 09:40:21 +02:00
* fix(evals): align plan-eng/design plan-mode + finding-floor smokes to their declared periodic tier The #2077 demotion of these four stochastic tests to 'periodic' was inert: E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate', so they kept running in the blocking gate lane and never in the weekly lane. Flip the four self-gates to 'periodic' (headers/describe labels updated), add a free static tier-alignment invariant test (dep-list filename mapping; unmapped self-gated files are reported, never silently skipped), and name the two plan-mode test files in their own touchfiles dep lists so the invariant binds for them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(pty-runner): scope-gate question/auto-select detectors + observation flags Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5 collapsed-form handling): isScopeGateQuestionVisible requires the question text PLUS option-body text (native AskUserQuestion renders numbered options, prose fallback renders lettered — the option body appears in both; narration doesn't), and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the selected-B token. runPlanSkillObservation gains scopeGateQuestionObserved / scopeGateAutoSelectObserved high-water flags (attached at every return path) so paid smokes can assert gate behavior across the whole run instead of the lossy 2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven questions, and the gate could fire inside the 3s pre-target window. Unit fixtures pin clean/native/collapsed positives, narration negatives, and the verbatim template announcement string (template rewording fails here first, before the paid smokes degrade to vacuous asserts). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(plan-eng/design-review): auto-select B in plan mode at the scope gate In plan mode the scope gate's "What should I review? A/B/C" question is pure friction: there is no branch diff and the target is the plan being drafted. Both gates gain an ordered exceptions block, checked BEFORE asking: 1. Plan mode → auto-select B: review the active plan (in context or pasted), announce it in one line ("Scope gate: plan mode — auto-selected B (reviewing <target>)") so the user can interrupt; an explicitly different user-named target still wins; no plan drafted yet → ask as normal. 2. User-named target (outside plan mode): explicit-only — a path, a pasted doc, or the literal words "branch diff". A passing mention is not naming; when in doubt, ask. Outside plan mode with no explicitly-named target, nothing changes. Plan-mode is checked FIRST because the PTY harness seeds drafts as pasted user messages (claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic. Pinning: seeded plan-mode smokes assert no gate render + announcement rendered (eng test 2; new design seeded test); plan-mode-no-op extends to eng/design (bypass must not misfire outside plan mode; first question must be the gate) plus a named-target case proving the pasted target is consumed; a drift-guard asserts the two hand-duplicated exceptions blocks stay identical modulo the two variant slots and carry the announcement string the detectors pin. Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): skip the scope gate when following loaded review skills autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim, and its section skip list omitted the scope gate — so autoplan ingested a hard-STOP AskUserQuestion that contradicts its every-question-auto-decides contract. One skip-list line fixes it; a static toContain pin in skill-validation keeps the entry load-bearing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: file scope-gate resolver-extraction TODO (eng-review D5 follow-up) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): positional floor exclusion, flag builder, outcome union, token tracking Review-army + adversarial findings on the scope-gate observability work, all verified before fixing: - Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate exclusion scanned only the 1500-byte tail, so an early gate render satisfied the floor vacuously once ~1.5KB of output accumulated (found independently by 4 review passes; predicate reproduced). Acceptance now scans only content APPENDED after the first gate render (positional anchor), and the LLM-judge 'waiting' shortcut no longer fires while the gate menu is the pending render. - High-water flags are built once and spread at every return path — the hand-spread pattern had already drifted (judge-waiting return omitted two flags), which made must-stay-false asserts vacuous on those paths. - isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases) and quoted-occurrence rejection (a model verbatim-quoting the announcement while declining must not trip must-stay-FALSE asserts). Fixtures added for both directions. - PlanSkillObservation outcome union gains 'wrote_findings_before_asking' (returned at runtime via classifyVisible but missing from the type). - trackTokens/tokensObserved: cumulative-buffer token high-water for consumption asserts (the 2KB evidence tail is lossy and the plan-file fallback is unreachable outside plan mode). - New scope-gate-floor unit pins (from the ship coverage audit): both gate render forms trip acceptance and exclusion; a genuine finding AUQ is not excluded; tail-scoping semantics pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(evals): harden no-op asserts, close tier-invariant fail-open holes, pin gate question strings - no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the outcome==='asked' conditional let a silent-bypass plan_ready run sail through); eng/design cases force --disallowedTools so the pinned prose shape is contractual rather than hoping native AUQ renders match; the named-target case uses trackTokens for consumption and lists wrote_findings_before_asking in its diagnostic throw branch. - tier-alignment invariant: both quote styles matched; zero-self-gate, mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED instead of silently skipped (the fail-open holes three reviewers found). - drift-guard: the generated gate menus must carry the exact question/option strings the PTY question detector anchors on — free CI fails before the paid smokes can go vacuous on a menu reword. - touchfiles: corrected the no-op cost note for CI concurrency + retry semantics. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): register plan-eng/design-review skills in PTY eval containers The extended plan-mode-no-op smoke invokes /plan-eng-review and /plan-design-review, but the fresh CI containers registered only office-hours and plan-ceo-review — both new runs would return 'Unknown command' and fail every PR's gate job (Codex structured review P1, verified against evals.yml). Registration loops, the dangling-target fail-fast list, and the frontmatter checks (now a loop over the same skill list, so the lists can't drift) all cover the two skills. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(plan-eng/design-review): harden scope-gate exceptions against injection and ambiguity Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex cross-confirmation), applied to both gate templates + regen: - Host-anchored mode signal: only the host's own system messages (plan-mode reminder or active plan file path) arm the auto-select; plan-shaped text inside pasted documents, tool results, or fetched pages does NOT count — injected content can't disarm the consent gate or nominate the target. - Multiple plan candidates: the host-referenced plan file wins; still ambiguous means ask. - The DIFFERENT-target override carries the passing-mention guard. - Plan mode + explicitly named target + no drafted plan resolves to the named target instead of a contradictory re-ask. - The numbered ask-path rules are qualified ('When no exception above applied:') so they no longer restate an unconditional MUST-ask that contradicts the exceptions. - 'Whenever this gate does ask — in any mode — it is a hard STOP.' - Shared preamble: 'any AskUserQuestion the skill fires is the workflow operating within plan mode' (was 'the first AskUserQuestion is the workflow entering plan mode', which framed the opposite of the bypass); regenerates every skill. - Ceilings ratcheted with attribution: plan-eng union ratio 1.10, investigate 1.10 (the ~250B shared-preamble reword lands the closest-to-ceiling skill at 1.092). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.62.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pty-runner): active-render gate veto in the floor check + honest periodic-wiring docs Codex re-review P2s on the fix wave, both verified: - A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting, no further output) was vetoed by the blanket tail exclusion until timeout. The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last cursor menu, so only a pending GATE menu vetoes; the judge fallback shares the same check. Residual (documented): prose gate + prose finding inside one tail — floors run the native-menu path in practice. - The four demoted periodic tests are not in evals-periodic.yml's explicit matrix (a named instance of the pre-existing periodic-orphans TODO), so they run locally/manually until the PTY-capable periodic job lands. CHANGELOG claim softened accordingly; TODO filed with the wiring recipe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.62.0.0 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply codex doc-review fixes for v1.62.0.0 - CLAUDE.md: scope the tier-alignment invariant claim (mapped files enforced, unmapped files reported) - docs/skills.md: document the plan-mode auto-select scope gate for /plan-eng-review and /plan-design-review - evals.yml: fix stale comment (PTY smokes register four skills, not two) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh ship golden baselines for the plan-mode preamble reword The generate-completion-status.ts wording change ('any AskUserQuestion the skill fires…') intentionally regenerates every SKILL.md; the byte-compare goldens carry the generator's output and refresh with it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): custom-hooks-path detection false-negatives on git worktrees The pre-push guard's HOOKS_IN_GIT_DIR check compared the hooks dir against --absolute-git-dir, which in a linked worktree is .git/worktrees/<name> while hooks resolve to the COMMON .git/hooks — so every Conductor worktree read as a 'custom hooks path' and the consented guard install was skipped. Match against the resolved --git-common-dir too (with a /nonexistent fallback so a failed resolution can't collapse the case pattern into match-everything). Verified live: this worktree now reports yes (was no), and the main checkout still reports yes. Goldens refreshed (--host all). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: changelog bullet for the worktree hooks-detection fix Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): give the plan-ceo plan-mode smoke real budget headroom Measured 2026-08-11: a clean isolated pass took 295.7s against the 300s inner budget (4s of margin) and the same test timed out at ~308s three times under concurrent eval load — a budget-edge flake in the gate lane, not a behavior regression (it passed isolated on both this branch and main). Inner budget 300s -> 420s, outer bun timeout 360s -> 480s, and the test file is now named in its own touchfiles dep list so the tier-alignment invariant binds for it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): 300s budget floor for the two 90s design-consultation SDK tests Root cause of PR #2533's e2e-design CI failure: design-consultation-preview failed 3 attempts at 0 turns/$0.00/93s — the session was up but the model's first completion queued past the 90s inner budget under concurrent API load (11 matrix jobs; the sibling research test booted its first tool at 4s, so this is API-side queuing, not CPU boot contention). The test was selected only because touchfiles.ts is a global touchfile; the tested behavior is untouched by this branch. 90s budgets cannot absorb one slow first completion. Both 90s tests in the file move to the repo's saturated-runner standard (300s inner / 360s outer, matching review-dashboard-via and retro-base-branch). Deliberately NOT re-arming the runner's inner timer on first stream event: an audit found ~100 outer bun-timeout literals sized inner+30-60s that a re-arm would silently break — the structural options are written up in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
342 lines
17 KiB
TypeScript
342 lines
17 KiB
TypeScript
/**
|
|
* Canonical carved-skill guard registry — the single source of truth for which
|
|
* skills are carved (skeleton SKILL.md + on-demand sections/*.md) and what each
|
|
* carve must guarantee.
|
|
*
|
|
* PURE LEAF DATA MODULE (codex outside-voice #1, refined-plan pass): this file
|
|
* has NO runtime imports — `import type` only. parity-harness.ts and
|
|
* skill-size-budget.test.ts derive their carved-skill lists FROM here (no
|
|
* parallel hand-maintained lists), so a runtime import back into either of them
|
|
* would create a cycle. Keep it data.
|
|
*
|
|
* Consumers:
|
|
* - test/carve-section-ordering.test.ts (E2, gate) → staticInvariants
|
|
* - test/carve-section-loading.test.ts (T2, periodic) → requiredReads + scenario
|
|
* - test/carve-guard-completeness.test.ts (E1, gate) → the set must equal the
|
|
* filesystem carved set
|
|
* - test/carve-guards-negative.test.ts (ET1, gate) → injects a broken fixture
|
|
* - test/helpers/parity-harness.ts → sectioned/maxSkeletonBytes/minBytes/mustContain
|
|
* - test/skill-size-budget.test.ts → SECTIONS_EXTRACTED = CARVED_SKILLS
|
|
*
|
|
* Adding a carve = add one entry here (atomically, in the same commit as the
|
|
* skeleton + manifest + sections — codex #4 — so E1's bidirectional parity never
|
|
* false-positives mid-commit).
|
|
*/
|
|
|
|
/** Static (skeleton-shape) invariants the per-PR ordering guard (E2) asserts. */
|
|
export interface CarveStaticInvariants {
|
|
/**
|
|
* Substrings that MUST remain in the always-loaded skeleton. Empty = skip
|
|
* (the skill has no distinctive pre-STOP anchor worth pinning beyond the
|
|
* universal STOP/section-index checks E2 already runs).
|
|
*/
|
|
mustStayInSkeleton: string[];
|
|
/**
|
|
* Substrings that MUST appear in the skeleton BEFORE the first STOP-Read
|
|
* (earliest-use, codex #6). For cso: mode-dispatch directives (## Arguments,
|
|
* ## Mode Resolution) must be resolved before any section is read — a dispatch
|
|
* directive stranded after the STOP can't govern which sections to read.
|
|
* Empty/undefined = skip (most skills).
|
|
*/
|
|
mustPrecedeStop?: string[];
|
|
/**
|
|
* Substrings that MUST be in the union (skeleton + sections) but MUST NOT be in
|
|
* the skeleton — i.e. the heavy body that the carve relocated. Empty = skip.
|
|
*/
|
|
mustMoveToSection: string[];
|
|
/**
|
|
* If set, this marker must appear in the skeleton AFTER the last STOP-Read
|
|
* directive (e.g. the EXIT PLAN MODE GATE that fires once section work returns).
|
|
* Undefined = the skill has no post-STOP gate (operational/conversational carve).
|
|
*/
|
|
gateAfterStop?: string;
|
|
}
|
|
|
|
export interface CarveGuard {
|
|
skill: string;
|
|
/** Section .md filenames the manifest lists and the skeleton must STOP-Read. */
|
|
expectedSections: string[];
|
|
/**
|
|
* Sections the behavioral test (T2) asserts the agent actually Read when driven
|
|
* by `scenario`. A non-empty subset of expectedSections — the ones the scenario
|
|
* is built to require. The registry owns this so "registered ⇒ asserted" is
|
|
* structural (codex #2), not policed.
|
|
*/
|
|
requiredReads: string[];
|
|
/**
|
|
* Fixture prompt that drives a real `claude -p` run down the STOP-Read path for
|
|
* this skill (codex #7). The behavioral test asserts the run reached the STOP
|
|
* (read requiredReads), not merely that nothing was read.
|
|
*/
|
|
scenario: string;
|
|
staticInvariants: CarveStaticInvariants;
|
|
/**
|
|
* How the behavioral guard (T2) exercises this skill:
|
|
* - 'plan' → write a PLAN.md fixture, run the review against it
|
|
* - 'prompt' → no fixture file; the scenario prompt alone drives the run
|
|
* - 'external' → covered by a dedicated bespoke test (complex fixtures, e.g.
|
|
* ship's git/VERSION/CHANGELOG state). The data-driven loop
|
|
* skips it; E1 asserts `externalTest` exists instead.
|
|
*/
|
|
behavioral: 'plan' | 'prompt' | 'external';
|
|
/** Required when behavioral === 'external': path (repo-relative) to the dedicated test. */
|
|
externalTest?: string;
|
|
/** Parity: max bytes for the always-loaded skeleton (asserts the carve shrank it). */
|
|
maxSkeletonBytes: number;
|
|
/** Parity: min bytes for the skeleton+sections union (total behavior preserved). */
|
|
minUnionBytes: number;
|
|
/** Parity: content phrases the union must preserve. */
|
|
mustContain: string[];
|
|
/**
|
|
* Parity: optional per-skill override for the union size-growth ceiling vs the
|
|
* v1.53.0.0 baseline (default 1.05). Bumped only when a deliberate cross-cutting
|
|
* preamble feature legitimately grows a smaller carved skeleton past 5%.
|
|
*/
|
|
maxSizeRatio?: number;
|
|
}
|
|
|
|
export const CARVE_GUARDS: Record<string, CarveGuard> = {
|
|
ship: {
|
|
skill: 'ship',
|
|
expectedSections: [
|
|
'tests.md',
|
|
'test-coverage.md',
|
|
'plan-completion.md',
|
|
'review-army.md',
|
|
'greptile.md',
|
|
'adversarial.md',
|
|
'changelog.md',
|
|
'pr-body.md',
|
|
],
|
|
requiredReads: ['review-army.md', 'changelog.md'],
|
|
scenario:
|
|
'This is a FRESH version-changing ship: the branch has a real code change, VERSION still equals the base version (needs a bump), and CHANGELOG.md needs a new entry. Follow the skill flow for a version-changing ship: run the pre-landing review and prepare the CHANGELOG entry. Produce the ship plan / review report. Do NOT actually commit, push, or open a PR.',
|
|
staticInvariants: {
|
|
// The PR-title-version invariant MUST stay always-loaded: the v1.54.0.0
|
|
// carve stranded it in pr-body.md and PRs started landing with bare titles
|
|
// (CI backstop: test/pr-title-sync-workflow-safety.test.ts).
|
|
mustStayInSkeleton: ['v$NEW_VERSION', 'gstack-pr-title-rewrite'],
|
|
// ...while the full create/update procedure stays carved into pr-body.md
|
|
// (out of the skeleton, present in the union). Asserts BOTH PR paths
|
|
// survive: the create path and the idempotent update path.
|
|
mustMoveToSection: ['gh pr create --base', 'gh pr edit --title'],
|
|
// ship is operational (multi-STOP, not a plan review); no single post-STOP gate.
|
|
gateAfterStop: undefined,
|
|
},
|
|
behavioral: 'external',
|
|
externalTest: 'test/skill-e2e-ship-section-loading.test.ts',
|
|
maxSkeletonBytes: 90_000,
|
|
minUnionBytes: 120_000,
|
|
mustContain: ['VERSION', 'CHANGELOG', 'review', 'merge', 'PR'],
|
|
// v1.58.5.0: pre-push-guard install (#2077) stacks on the shared first-run-guidance preamble.
|
|
maxSizeRatio: 1.08,
|
|
},
|
|
'plan-ceo-review': {
|
|
skill: 'plan-ceo-review',
|
|
expectedSections: ['review-sections.md'],
|
|
requiredReads: ['review-sections.md'],
|
|
scenario:
|
|
'Review the plan in PLAN.md. Hold the current scope (HOLD SCOPE mode) — do not challenge or expand scope. Run the full CEO review and produce the review report.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: ['## Step 0: Nuclear Scope Challenge'],
|
|
mustMoveToSection: ['### Section 1: Architecture Review', '## Mode Quick Reference'],
|
|
gateAfterStop: 'EXIT PLAN MODE GATE',
|
|
},
|
|
behavioral: 'external',
|
|
externalTest: 'test/skill-e2e-plan-ceo-review-section-loading.test.ts',
|
|
maxSkeletonBytes: 90_000,
|
|
minUnionBytes: 80_000,
|
|
mustContain: ['SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'HOLD SCOPE', 'SCOPE REDUCTION'],
|
|
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
|
// prose replacing the smaller opt-in question) lands this ~5.2% over baseline.
|
|
maxSizeRatio: 1.08,
|
|
},
|
|
'plan-eng-review': {
|
|
skill: 'plan-eng-review',
|
|
expectedSections: ['review-sections.md'],
|
|
requiredReads: ['review-sections.md'],
|
|
scenario:
|
|
'Review the plan in PLAN.md. Accept the current scope. Run the full engineering review (architecture, code quality, tests, performance) and produce the review report.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: ['### Step 0: Scope Challenge'],
|
|
mustMoveToSection: ['### 1. Architecture review'],
|
|
gateAfterStop: 'EXIT PLAN MODE GATE',
|
|
},
|
|
behavioral: 'plan',
|
|
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
|
|
// +~1 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
|
|
maxSkeletonBytes: 68_000,
|
|
minUnionBytes: 70_000,
|
|
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
|
|
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
|
|
// decision-memory nudge + the v1.57.4.0 Boil-the-Ocean rename) plus the
|
|
// default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
|
// prose, replacing the smaller opt-in question) land this at ~6.6% over the
|
|
// v1.53.0.0 baseline. Headroom for those intentional additions.
|
|
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
|
|
// hardening: host-anchored mode signal, precedence, passing-mention
|
|
// guards) and the plan-mode preamble reword land the union at 1.092.
|
|
maxSizeRatio: 1.10,
|
|
},
|
|
'plan-design-review': {
|
|
skill: 'plan-design-review',
|
|
expectedSections: ['review-sections.md'],
|
|
requiredReads: ['review-sections.md'],
|
|
scenario:
|
|
'Review the plan in PLAN.md for design and UX. Accept the current scope. Run the full design review passes and produce the review report.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: [],
|
|
mustMoveToSection: ['### Pass 1: Information Architecture'],
|
|
gateAfterStop: 'EXIT PLAN MODE GATE',
|
|
},
|
|
behavioral: 'plan',
|
|
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
|
|
// always-loaded AskUserQuestion Format section.
|
|
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
|
|
// +~1.3 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
|
|
maxSkeletonBytes: 89_000,
|
|
minUnionBytes: 70_000,
|
|
mustContain: ['design', 'visual'],
|
|
maxSizeRatio: 1.07,
|
|
},
|
|
'plan-devex-review': {
|
|
skill: 'plan-devex-review',
|
|
expectedSections: ['review-sections.md'],
|
|
requiredReads: ['review-sections.md'],
|
|
scenario:
|
|
'Review the plan in PLAN.md for developer experience. Accept the current scope. Run the full DX review passes and produce the review report.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: [],
|
|
mustMoveToSection: ['### Pass 1: Getting Started Experience'],
|
|
gateAfterStop: 'EXIT PLAN MODE GATE',
|
|
},
|
|
behavioral: 'plan',
|
|
// +Conductor AUQ-default-prose rule + one-way/destructive prose safety +
|
|
// continuation protocol in the always-loaded AskUserQuestion Format section.
|
|
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
|
|
maxSkeletonBytes: 80_000,
|
|
minUnionBytes: 70_000,
|
|
mustContain: ['developer experience', 'Getting Started'],
|
|
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
|
|
// prose replacing the smaller opt-in question) lands this ~5.7% over baseline.
|
|
maxSizeRatio: 1.08,
|
|
},
|
|
'office-hours': {
|
|
skill: 'office-hours',
|
|
expectedSections: ['design-and-handoff.md'],
|
|
requiredReads: ['design-and-handoff.md'],
|
|
scenario:
|
|
'Run office hours for this product idea through to the end: have the diagnostic conversation, explore alternatives, then write the design doc and run the relationship handoff (Phases 5-6).',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: [],
|
|
mustMoveToSection: [],
|
|
// office-hours is conversational; the design-doc/handoff section has no
|
|
// post-STOP review gate in the skeleton.
|
|
gateAfterStop: undefined,
|
|
},
|
|
behavioral: 'prompt',
|
|
// v1.2.0 activation lift: first-run-guidance section in the shared preamble,
|
|
// plus the P1 office-hours closing handoff (AUQ that launches the next skill).
|
|
maxSkeletonBytes: 98_000,
|
|
minUnionBytes: 70_000,
|
|
mustContain: ['design doc', 'problem statement'],
|
|
maxSizeRatio: 1.07,
|
|
},
|
|
'document-release': {
|
|
skill: 'document-release',
|
|
expectedSections: ['release-body.md'],
|
|
requiredReads: ['release-body.md'],
|
|
scenario:
|
|
'A PR has shipped a new CLI flag and touched README.md and CHANGELOG.md. Skip the git pre-flight shell commands (assume the diff adds --new-flag and updates those two docs). Run the documentation workflow: build the coverage map, then audit the docs, apply updates, and polish the CHANGELOG voice. Produce the documentation health summary.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: ['## Step 1: Pre-flight', '## Step 1.5: Coverage Map'],
|
|
mustMoveToSection: ['## Step 2: Per-File Documentation Audit', '## Step 5: CHANGELOG Voice Polish'],
|
|
// Operational skill (no plan-mode review gate).
|
|
gateAfterStop: undefined,
|
|
},
|
|
behavioral: 'prompt',
|
|
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
|
|
// always-loaded AskUserQuestion Format section.
|
|
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
|
|
maxSkeletonBytes: 56_000,
|
|
minUnionBytes: 55_000,
|
|
mustContain: ['CHANGELOG', 'Diataxis', 'coverage'],
|
|
// Two intentional additions stack on this small skill: the AUQ-failure prose
|
|
// fallback (v1.57.2.0, ~2KB to every preamble) AND the new default-on Codex
|
|
// documentation-review section (codexPreflight + prompt + apply-gate, carved
|
|
// into release-body so the SKELETON stays under maxSkeletonBytes). On a ~55KB
|
|
// baseline that whole new capability is ~18.6% of union bytes. The doc review
|
|
// is a deliberate new feature, not preamble creep; the union ceiling is raised
|
|
// to match while the skeleton budget (50_000) still holds the always-loaded
|
|
// cost flat.
|
|
maxSizeRatio: 1.20,
|
|
},
|
|
'design-consultation': {
|
|
skill: 'design-consultation',
|
|
expectedSections: ['proposal-and-preview.md'],
|
|
requiredReads: ['proposal-and-preview.md'],
|
|
scenario:
|
|
'The user gave product context (a B2B analytics dashboard for ops teams) and declined the research phase. Skip browser/design tool setup. Proceed to build the complete design-system proposal, then write DESIGN.md. Produce the proposal and the DESIGN.md content.',
|
|
staticInvariants: {
|
|
mustStayInSkeleton: ['## Phase 0: Pre-checks', '## Phase 1: Product Context', '## Phase 2: Research'],
|
|
mustMoveToSection: ['## Phase 3: The Complete Proposal', '## Phase 6: Write DESIGN.md'],
|
|
gateAfterStop: undefined,
|
|
},
|
|
behavioral: 'prompt',
|
|
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
|
|
// always-loaded AskUserQuestion Format section.
|
|
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
|
|
maxSkeletonBytes: 69_000,
|
|
minUnionBytes: 72_000,
|
|
mustContain: ['Typography', 'Color', 'Aesthetic Direction'],
|
|
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback ~2KB +
|
|
// the cross-session decision-memory nudge) lands this carved skeleton just over
|
|
// the strict 1.05; headroom for the shared preamble additions.
|
|
maxSizeRatio: 1.07,
|
|
},
|
|
cso: {
|
|
skill: 'cso',
|
|
expectedSections: ['audit-phases.md'],
|
|
requiredReads: ['audit-phases.md'],
|
|
scenario:
|
|
'Run a security audit on this repository in --owasp mode (OWASP Top 10 only). Resolve the mode, do the Phase 0 stack detection and Phase 1 attack-surface census, then run the scoped audit phases and produce the findings report. Skip any step that needs network access.',
|
|
staticInvariants: {
|
|
// Dispatch + always-run + FP-filtering phases are ALWAYS loaded (security).
|
|
mustStayInSkeleton: [
|
|
'## Arguments',
|
|
'## Mode Resolution',
|
|
'### Phase 0',
|
|
'### Phase 1',
|
|
'### Phase 12',
|
|
'### Phase 13',
|
|
'### Phase 14',
|
|
],
|
|
// Earliest-use: mode must be resolvable before any section is read (codex #6).
|
|
mustPrecedeStop: ['## Arguments', '## Mode Resolution'],
|
|
// Scope-dependent audit detail moved to the section.
|
|
mustMoveToSection: [
|
|
'### Phase 2: Secrets Archaeology',
|
|
'### Phase 9: OWASP Top 10 Assessment',
|
|
'### Phase 10: STRIDE Threat Model',
|
|
],
|
|
gateAfterStop: undefined,
|
|
},
|
|
behavioral: 'prompt',
|
|
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
|
|
// always-loaded AskUserQuestion Format section.
|
|
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
|
|
maxSkeletonBytes: 75_000,
|
|
minUnionBytes: 72_000,
|
|
mustContain: ['OWASP', 'STRIDE', 'daily', 'comprehensive', 'verif'],
|
|
// cso keeps its mode-dispatch + FP-filtering phases always-loaded, so the
|
|
// cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback ~2KB + the
|
|
// decision-memory nudge) lands it just over 1.05; headroom for the shared additions.
|
|
maxSizeRatio: 1.07,
|
|
},
|
|
};
|
|
|
|
/** Sorted carved-skill names. Consumers derive their lists from this — no parallel lists. */
|
|
export const CARVED_SKILLS: readonly string[] = Object.freeze(
|
|
Object.keys(CARVE_GUARDS).sort(),
|
|
);
|