Files
gstack/test/helpers/carve-guards.ts
T
Garry TanandClaude Fable 5 a3749bfa4b v1.70.1.0 fix: ship names the /document-release subagent at every decision point (tripwire + gate E2E) (#2700)
* fix(ship): name the /document-release subagent at every Step 18 decision point

The v1.54.0.0 carve moved Step 18 (documentation sync) into
ship/sections/pr-body.md and the Claude-host skeleton stopped saying
"document-release" anywhere in the workflow body — the dispatch became
invisible at exactly the moments an agent decides whether to open the
section. Restore visibility at three touchpoints, all subagent-framed
(never bare-slash-framed, which would invite an inline Skill invocation
that bypasses the fresh-context subagent + JSON contract):

- manifest trigger (renders into the section-index row AND the STOP
  pointer): "dispatching the /document-release subagent to sync docs
  (Step 18) and then creating or updating the PR/MR (Step 19)"
- Step 17 handoff line names Step 18's dispatch explicitly
- new hoisted doc-sync invariant beside the PR-title invariant: the
  dispatch itself is never skipped; only a failed subagent is
  non-blocking

Pin it in carve-guards: 'the /document-release subagent' (all three
touchpoints) + 'dispatches the /document-release subagent' (invariant)
must stay in the skeleton; the carved imperative 'Dispatch
/document-release as a subagent' must stay carved. Skeleton cap
91,600 → 92,300 (measured 91,764; trigger renders twice). Goldens
regenerated for all three hosts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: pin the ship→document-release Step 18 wiring with a free tripwire

Five substring/structure asserts across the carved section, the Claude
skeleton's three touchpoints, the manifest trigger, and the codex/factory
goldens (inlined Step 18 ordered before Step 19). Claude-golden asserts
deliberately omitted: host-config.test.ts already enforces golden ==
generated byte-for-byte.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: gate-tier E2E proving /ship dispatches the document-release subagent

New skill-e2e-ship-docsync: a live agent gets the sliced Step 17→19 tail
of the generated ship skeleton in a bare-remote git fixture (Steps 0-16
"done"), under a fake HOME so the STOP pointer and the Step 18 subagent
prompt resolve to planted copies, with a stub document-release skill that
returns the empty-result JSON contract. Hard assert: an Agent/Task
tool-call matching /document-release/i exists in result.toolCalls and
precedes any `gh pr create`. Neutral prompt (no STOP-Read priming, no
document-release mention — the prompt echoes into the transcript, so
asserts read toolCalls only).

Hardening from review: throw-on-marker-drift fixture slice; per-test
GSTACK_HOME + .redact-prepush-prompted marker (routes Step 17's
credential guard to its silent branch — the hermetic GSTACK_HOME pin
defeats a HOME-only override); 480s/540s timeouts (nested subagent adds
wall clock the 300s sibling never carried); 'timeout' accepted in
exitReason only because the dispatch assert is independently hard;
whole-file describeE2ETier('gate') composed with diff selection (keeps
the file out of the periodic shard census, which sits at its ceiling,
and under the hard tier-alignment invariant).

Registered as 'ship-docsync' in E2E_TOUCHFILES + E2E_TIERS (gate) in the
same commit — touchfiles.test.ts rejects either half landing first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fix stale document-release TODOS entry + three review-deferred items

The SHIPPED entry still described the deleted Step 8.5 post-PR cat-delegation
design from v0.8.4; replace with the current Step 18 subagent design and its
test pins. Add the three P3 items deferred from the v1.69 plan review:
dispatch receipt enforcement, land-and-deploy→canary dispatch-pin pattern,
and the periodic shard-census boundary.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: pre-landing review fixes

Testing-specialist findings, all mechanical: (1) pin the E2E fixture's git
branch (-b main / init.defaultBranch=main) and assert every setup command's
exit status so operator git config can't silently corrupt a paid run;
(2) tighten the dispatch matcher to Step 18-prompt-specific markers
(document-release/SKILL.md | executing the /document-release workflow) so a
subagent merely quoting section text can't false-pass the regression assert
(verified against recorded burn-in transcripts); (3) replace the subsumed
carve-guards anchor with three non-overlapping per-touchpoint anchors
(gerund/imperative/3rd-person) so each touchpoint is independently enforced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: red-team review fixes

Five informational findings: TODOS shard-census arithmetic corrected (census
is 67 with one free ungated slot; the SECOND ungated file trips the floor)
and version pointer fixed (v0.18.2.0, not v0.18.1.0); the free tripwire now
pins the two dispatch-matcher marker strings so a pr-body prompt reword
fails the free suite instead of surfacing as a paid-tier mystery; the E2E
matcher gains a section-paste exclusion (scaffold strings disqualify) —
verified against all recorded runs; the E2E header documents the tierless
test:evals invisibility tradeoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: adversarial review fixes

Pin the E2E matcher's two EXCLUSION markers in the free tripwire (an
unpinned 'Parent processing:' reword would silently deaden the
section-paste guard while every test stayed green); add an ordering pin
(the hoisted doc-sync invariant must sit above the pr-body STOP pointer —
presence-only anchors can't catch drift below it); plant a third
cwd-relative pr-body copy inside the fixture repo, gitignored so the agent
never tries to commit test scaffolding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.70.1.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: CHANGELOG accuracy fixes from the doc-release review

Three factual corrections the Step 18 doc subagent caught in the fresh
v1.70.1.0 entry: 5 tripwire tests (not 6), cost floor $0.63 per the cited
eval store (not $0.59), and the visibility claim scoped to decision points
(the re-run checklist mention survived the carve). Plus the E2E header's
stale pending-burn-in note replaced with the observed numbers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: raise bun-polyfill subprocess budget to 60s for degraded Windows runners

The 50ms-sleep test blew the 20s budget on BOTH bun retry attempts on PR
#2700's windows-latest runner (run 32989821401) — sustained AV/runner
pressure, not just the documented cold-start. Same flake passed-on-rerun on
the prompt-token-load-reduction branch yesterday. Budget only; every
assertion still checks exact output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): run the ship-docsync gate E2E in the evals matrix + silent-skip tripwire

The evals.yml matrix is hand-enumerated and the Run step never exported
EVALS_TIER, so the new whole-file-gated ship-docsync E2E would have
self-skipped even with a row — a hollow green one layer deeper than the
documented rehomed-monolith incident. Add the e2e-ship-docsync row with a
row-level `tier: gate` property, exported as EVALS_TIER by the Run step
(empty = unset for every existing row: all readers are `=== '<tier>'` or
truthiness).

New free tripwire test/evals-workflow-matrix.test.ts ratchets the class:
matrix files must exist; gate-hosting files must have a row; whole-file-gated
matrix files must carry a matching row tier; and the burn-down lists enforce
their own cleanup. It enumerates the PRE-EXISTING holes found while wiring
this (8 gate-hosting files with no row; codex/gemini rows running zero tests;
the pty-plan-smoke row hollow since its files adopted describeE2ETier) —
tracked in TODOS as the CI gate-lane hollow-coverage burn-down.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 08:46:41 -07:00

404 lines
21 KiB
TypeScript

/**
* Canonical carved-skill guard registry — the single source of truth for which
* skills are carved (skeleton SKILL.md + on-demand sections/*.md) and what each
* carve must guarantee.
*
* PURE LEAF DATA MODULE (codex outside-voice #1, refined-plan pass): this file
* has NO runtime imports — `import type` only. parity-harness.ts and
* skill-size-budget.test.ts derive their carved-skill lists FROM here (no
* parallel hand-maintained lists), so a runtime import back into either of them
* would create a cycle. Keep it data.
*
* Consumers:
* - test/carve-section-ordering.test.ts (E2, gate) → staticInvariants
* - test/carve-section-loading.test.ts (T2, periodic) → requiredReads + scenario
* - test/carve-guard-completeness.test.ts (E1, gate) → the set must equal the
* filesystem carved set
* - test/carve-guards-negative.test.ts (ET1, gate) → injects a broken fixture
* - test/helpers/parity-harness.ts → sectioned/maxSkeletonBytes/minBytes/mustContain
* - test/skill-size-budget.test.ts → SECTIONS_EXTRACTED = CARVED_SKILLS
*
* Adding a carve = add one entry here (atomically, in the same commit as the
* skeleton + manifest + sections — codex #4 — so E1's bidirectional parity never
* false-positives mid-commit).
*/
/** Static (skeleton-shape) invariants the per-PR ordering guard (E2) asserts. */
export interface CarveStaticInvariants {
/**
* Substrings that MUST remain in the always-loaded skeleton. Empty = skip
* (the skill has no distinctive pre-STOP anchor worth pinning beyond the
* universal STOP/section-index checks E2 already runs).
*/
mustStayInSkeleton: string[];
/**
* Substrings that MUST appear in the skeleton BEFORE the first STOP-Read
* (earliest-use, codex #6). For cso: mode-dispatch directives (## Arguments,
* ## Mode Resolution) must be resolved before any section is read — a dispatch
* directive stranded after the STOP can't govern which sections to read.
* Empty/undefined = skip (most skills).
*/
mustPrecedeStop?: string[];
/**
* Substrings that MUST be in the union (skeleton + sections) but MUST NOT be in
* the skeleton — i.e. the heavy body that the carve relocated. Empty = skip.
*/
mustMoveToSection: string[];
/**
* If set, this marker must appear in the skeleton AFTER the last STOP-Read
* directive (e.g. the EXIT PLAN MODE GATE that fires once section work returns).
* Undefined = the skill has no post-STOP gate (operational/conversational carve).
*/
gateAfterStop?: string;
}
export interface CarveGuard {
skill: string;
/** Section .md filenames the manifest lists and the skeleton must STOP-Read. */
expectedSections: string[];
/**
* Sections the behavioral test (T2) asserts the agent actually Read when driven
* by `scenario`. A non-empty subset of expectedSections — the ones the scenario
* is built to require. The registry owns this so "registered ⇒ asserted" is
* structural (codex #2), not policed.
*/
requiredReads: string[];
/**
* Fixture prompt that drives a real `claude -p` run down the STOP-Read path for
* this skill (codex #7). The behavioral test asserts the run reached the STOP
* (read requiredReads), not merely that nothing was read.
*/
scenario: string;
staticInvariants: CarveStaticInvariants;
/**
* How the behavioral guard (T2) exercises this skill:
* - 'plan' → write a PLAN.md fixture, run the review against it
* - 'prompt' → no fixture file; the scenario prompt alone drives the run
* - 'external' → covered by a dedicated bespoke test (complex fixtures, e.g.
* ship's git/VERSION/CHANGELOG state). The data-driven loop
* skips it; E1 asserts `externalTest` exists instead.
*/
behavioral: 'plan' | 'prompt' | 'external';
/** Required when behavioral === 'external': path (repo-relative) to the dedicated test. */
externalTest?: string;
/** Parity: max bytes for the always-loaded skeleton (asserts the carve shrank it). */
maxSkeletonBytes: number;
/** Parity: min bytes for the skeleton+sections union (total behavior preserved). */
minUnionBytes: number;
/** Parity: content phrases the union must preserve. */
mustContain: string[];
/**
* Parity: optional per-skill override for the union size-growth ceiling vs the
* v1.53.0.0 baseline (default 1.05). Bumped only when a deliberate cross-cutting
* preamble feature legitimately grows a smaller carved skeleton past 5%.
*/
maxSizeRatio?: number;
}
export const CARVE_GUARDS: Record<string, CarveGuard> = {
ship: {
skill: 'ship',
expectedSections: [
'apple-release.md',
'tests.md',
'test-coverage.md',
'plan-completion.md',
'review-army.md',
'greptile.md',
'adversarial.md',
'changelog.md',
'pr-body.md',
],
requiredReads: ['review-army.md', 'changelog.md'],
scenario:
'This is a FRESH version-changing ship: the branch has a real code change, VERSION still equals the base version (needs a bump), and CHANGELOG.md needs a new entry. Follow the skill flow for a version-changing ship: run the pre-landing review and prepare the CHANGELOG entry. Produce the ship plan / review report. Do NOT actually commit, push, or open a PR.',
staticInvariants: {
// The PR-title-version invariant MUST stay always-loaded: the v1.54.0.0
// carve stranded it in pr-body.md and PRs started landing with bare titles
// (CI backstop: test/pr-title-sync-workflow-safety.test.ts).
// Same carve also stranded the Step 18 /document-release dispatch out of
// sight — the skeleton never named it and the handoff "got lost" (#2666
// follow-up). Three NON-OVERLAPPING anchors pin the restored visibility,
// one per touchpoint (no anchor is a substring of another, so each is
// independently enforced — a subsumed anchor adds zero enforcement):
// gerund form → manifest trigger (renders 2x: section index + STOP)
// imperative → Step 17 handoff line
// 3rd person → hoisted doc-sync invariant
// Matching is case-sensitive String.includes — "dispatching the" does NOT
// contain "dispatch the" — so update anchors in lockstep with any
// touchpoint rewording.
mustStayInSkeleton: [
'v$NEW_VERSION',
'gstack-pr-title-rewrite',
'dispatching the /document-release subagent to sync docs',
'dispatch the /document-release subagent to sync docs',
'dispatches the /document-release subagent',
],
// ...while the full create/update procedure stays carved into pr-body.md
// (out of the skeleton, present in the union). Asserts BOTH PR paths
// survive: the create path and the idempotent update path. The Step 18
// dispatch imperative stays carved too — pasting that literal into the
// skeleton (correctly) fails this guard; the skeleton speaks of "the
// /document-release subagent", never the carved imperative.
mustMoveToSection: [
'gh pr create --base',
'gh pr edit --title',
'Dispatch /document-release as a subagent',
],
// ship is operational (multi-STOP, not a plan review); no single post-STOP gate.
gateAfterStop: undefined,
},
behavioral: 'external',
externalTest: 'test/skill-e2e-ship-section-loading.test.ts',
maxSkeletonBytes: 92_300, // document-release visibility restore: named trigger (renders twice) + Step 17 handoff + hoisted doc-sync invariant; measured 91,764
minUnionBytes: 120_000,
mustContain: ['VERSION', 'CHANGELOG', 'review', 'merge', 'PR'],
// v1.58.5.0: pre-push-guard install (#2077) stacks on the shared first-run-guidance preamble.
// Fork port wave 2: multi-ecosystem test-detection evidence (Django/JVM
// markers, test-file census — e3259078 port) + the #1079 gh pr edit REST
// fallback grew the union to 1.090x; the third-party web-actions
// contract (consent-gated browser drive for API-key registration etc.)
// adds ~2.3KB inline judgment, measured 1.103x. The Apple release
// adapter (14.8KB carved section, 21 live releases of judgment — the
// wave's headline capability) grows the union to 1.195x. Deliberate:
// the section is on-demand (loads only for Apple store targets), so
// per-invocation cost for non-iOS ships is one manifest line.
maxSizeRatio: 1.22,
},
'plan-ceo-review': {
skill: 'plan-ceo-review',
expectedSections: ['review-sections.md'],
requiredReads: ['review-sections.md'],
scenario:
'Review the plan in PLAN.md. Hold the current scope (HOLD SCOPE mode) — do not challenge or expand scope. Run the full CEO review and produce the review report.',
staticInvariants: {
mustStayInSkeleton: ['## Step 0: Nuclear Scope Challenge'],
mustMoveToSection: ['### Section 1: Architecture Review', '## Mode Quick Reference'],
gateAfterStop: 'EXIT PLAN MODE GATE',
},
behavioral: 'external',
externalTest: 'test/skill-e2e-plan-ceo-review-section-loading.test.ts',
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
// Fork port wave 2 (#703): the repo-doc-preference block in the design
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
maxSkeletonBytes: 93_900, // v1.68 fix wave: #2402 learnings capture + spool queue-depth lines; measured 93,345
minUnionBytes: 80_000,
mustContain: ['SCOPE EXPANSION', 'SELECTIVE EXPANSION', 'HOLD SCOPE', 'SCOPE REDUCTION'],
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
// prose replacing the smaller opt-in question) lands this ~5.2% over baseline.
maxSizeRatio: 1.08,
},
'plan-eng-review': {
skill: 'plan-eng-review',
expectedSections: ['review-sections.md'],
requiredReads: ['review-sections.md'],
scenario:
'Review the plan in PLAN.md. Accept the current scope. Run the full engineering review (architecture, code quality, tests, performance) and produce the review report.',
staticInvariants: {
mustStayInSkeleton: ['### Step 0: Scope Challenge'],
mustMoveToSection: ['### 1. Architecture review'],
gateAfterStop: 'EXIT PLAN MODE GATE',
},
behavioral: 'plan',
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
// +~1 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
// Fork port wave 2 (#703): the repo-doc-preference block in the design
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
maxSkeletonBytes: 71_800, // v1.68 fix wave (#2402); measured 71,228
minUnionBytes: 70_000,
mustContain: ['Architecture', 'Code Quality', 'Test', 'Performance'],
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback + the
// decision-memory nudge + the v1.57.4.0 Boil-the-Ocean rename) plus the
// default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
// prose, replacing the smaller opt-in question) land this at ~6.6% over the
// v1.53.0.0 baseline. Headroom for those intentional additions.
// 1.08 → 1.10: the scope-gate exceptions block (+ its adversarial-review
// hardening: host-anchored mode signal, precedence, passing-mention
// guards) and the plan-mode preamble reword land the union at 1.092.
maxSizeRatio: 1.12, // measured 1.103
},
'plan-design-review': {
skill: 'plan-design-review',
expectedSections: ['review-sections.md'],
requiredReads: ['review-sections.md'],
scenario:
'Review the plan in PLAN.md for design and UX. Accept the current scope. Run the full design review passes and produce the review report.',
staticInvariants: {
mustStayInSkeleton: [],
mustMoveToSection: ['### Pass 1: Information Architecture'],
gateAfterStop: 'EXIT PLAN MODE GATE',
},
behavioral: 'plan',
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift (shared first-run-guidance preamble) + #2077 ask-first scope gate.
// +~1.3 KB: plan-mode auto-select-B scope-gate exceptions (2026-08).
// Fork port wave 2 (D1): evidence directive adds ~0.45KB to every
// tier-2+ skeleton (measured 89,184). Main's v1.64.0.0 adds ~340 B more
// (telemetry --error-message/--failed-step preamble prose, PR #769).
// Budget covers the sum of both waves.
maxSkeletonBytes: 91_700, // v1.68 fix wave (#2402); measured 91,176
minUnionBytes: 70_000,
mustContain: ['design', 'visual'],
maxSizeRatio: 1.12, // D1 1.104 + main's ~0.008
},
'plan-devex-review': {
skill: 'plan-devex-review',
expectedSections: ['review-sections.md'],
requiredReads: ['review-sections.md'],
scenario:
'Review the plan in PLAN.md for developer experience. Accept the current scope. Run the full DX review passes and produce the review report.',
staticInvariants: {
mustStayInSkeleton: [],
mustMoveToSection: ['### Pass 1: Getting Started Experience'],
gateAfterStop: 'EXIT PLAN MODE GATE',
},
behavioral: 'plan',
// +Conductor AUQ-default-prose rule + one-way/destructive prose safety +
// continuation protocol in the always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
// Fork port wave 2 (#703): the repo-doc-preference block in the design
// check grew every plan-review skeleton ~0.7KB. Measured values noted.
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
maxSkeletonBytes: 83_500, // v1.68 fix wave (#2402); measured 82,941
minUnionBytes: 70_000,
mustContain: ['developer experience', 'Getting Started'],
// Default-on Codex outside-voice (codexPreflight block + CODEX_MODE branch
// prose replacing the smaller opt-in question) lands this ~5.7% over baseline.
maxSizeRatio: 1.08,
},
'office-hours': {
skill: 'office-hours',
expectedSections: ['design-and-handoff.md'],
requiredReads: ['design-and-handoff.md'],
scenario:
'Run office hours for this product idea through to the end: have the diagnostic conversation, explore alternatives, then write the design doc and run the relationship handoff (Phases 5-6).',
staticInvariants: {
mustStayInSkeleton: [],
mustMoveToSection: [],
// office-hours is conversational; the design-doc/handoff section has no
// post-STOP review gate in the skeleton.
gateAfterStop: undefined,
},
behavioral: 'prompt',
// v1.2.0 activation lift: first-run-guidance section in the shared preamble,
// plus the P1 office-hours closing handoff (AUQ that launches the next skill).
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
// Fork port wave 2: the third-party web-actions contract sits inline
// (judgment must be visible before the workflow directs the user to a
// vendor site), plus the #703 dual-write + repo-doc-preference block and
// the #538 opt-out + D1 evidence directive — ratio 1.104 measured.
// #2499 project-scope MCP jq in the brain-sync block grew every tier-2+
// skeleton ~1.5KB (entry resolution emitted once per SKILL.md).
maxSkeletonBytes: 102_800, // v1.68 fix wave (#2402); measured 102,220
minUnionBytes: 70_000,
mustContain: ['design doc', 'problem statement'],
maxSizeRatio: 1.12,
},
'document-release': {
skill: 'document-release',
expectedSections: ['release-body.md'],
requiredReads: ['release-body.md'],
scenario:
'A PR has shipped a new CLI flag and touched README.md and CHANGELOG.md. Skip the git pre-flight shell commands (assume the diff adds --new-flag and updates those two docs). Run the documentation workflow: build the coverage map, then audit the docs, apply updates, and polish the CHANGELOG voice. Produce the documentation health summary.',
staticInvariants: {
mustStayInSkeleton: ['## Step 1: Pre-flight', '## Step 1.5: Coverage Map'],
mustMoveToSection: ['## Step 2: Per-File Documentation Audit', '## Step 5: CHANGELOG Voice Polish'],
// Operational skill (no plan-mode review gate).
gateAfterStop: undefined,
},
behavioral: 'prompt',
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
maxSkeletonBytes: 57_900, // v1.68 fix wave (#2402); measured 57,385
minUnionBytes: 55_000,
mustContain: ['CHANGELOG', 'Diataxis', 'coverage'],
// Two intentional additions stack on this small skill: the AUQ-failure prose
// fallback (v1.57.2.0, ~2KB to every preamble) AND the new default-on Codex
// documentation-review section (codexPreflight + prompt + apply-gate, carved
// into release-body so the SKELETON stays under maxSkeletonBytes). On a ~55KB
// baseline that whole new capability is ~18.6% of union bytes. The doc review
// is a deliberate new feature, not preamble creep; the union ceiling is raised
// to match while the skeleton budget (50_000) still holds the always-loaded
// cost flat.
maxSizeRatio: 1.20,
},
'design-consultation': {
skill: 'design-consultation',
expectedSections: ['proposal-and-preview.md'],
requiredReads: ['proposal-and-preview.md'],
scenario:
'The user gave product context (a B2B analytics dashboard for ops teams) and declined the research phase. Skip browser/design tool setup. Proceed to build the complete design-system proposal, then write DESIGN.md. Produce the proposal and the DESIGN.md content.',
staticInvariants: {
mustStayInSkeleton: ['## Phase 0: Pre-checks', '## Phase 1: Product Context', '## Phase 2: Research'],
mustMoveToSection: ['## Phase 3: The Complete Proposal', '## Phase 6: Write DESIGN.md'],
gateAfterStop: undefined,
},
behavioral: 'prompt',
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
// v1.65 merge: provisional larger-of-both-waves budget; re-measured below.
// v1.64.1.0: shared-preamble prose from the two parallel v1.64 waves lands
// the skeleton at 69,022 B; +~1 KB headroom.
maxSkeletonBytes: 71_400, // v1.68 fix wave (#2402); measured 70,815
minUnionBytes: 72_000,
mustContain: ['Typography', 'Color', 'Aesthetic Direction'],
// Cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback ~2KB +
// the cross-session decision-memory nudge) lands this carved skeleton just over
// the strict 1.05; headroom for the shared preamble additions.
// v1.64+v1.65 merge sums both waves' preamble growth; measured 1.073.
maxSizeRatio: 1.08,
},
cso: {
skill: 'cso',
expectedSections: ['audit-phases.md'],
requiredReads: ['audit-phases.md'],
scenario:
'Run a security audit on this repository in --owasp mode (OWASP Top 10 only). Resolve the mode, do the Phase 0 stack detection and Phase 1 attack-surface census, then run the scoped audit phases and produce the findings report. Skip any step that needs network access.',
staticInvariants: {
// Dispatch + always-run + FP-filtering phases are ALWAYS loaded (security).
mustStayInSkeleton: [
'## Arguments',
'## Mode Resolution',
'### Phase 0',
'### Phase 1',
'### Phase 12',
'### Phase 13',
'### Phase 14',
],
// Earliest-use: mode must be resolvable before any section is read (codex #6).
mustPrecedeStop: ['## Arguments', '## Mode Resolution'],
// Scope-dependent audit detail moved to the section.
mustMoveToSection: [
'### Phase 2: Secrets Archaeology',
'### Phase 9: OWASP Top 10 Assessment',
'### Phase 10: STRIDE Threat Model',
],
gateAfterStop: undefined,
},
behavioral: 'prompt',
// +Conductor AUQ-default-prose rule + one-way/continuation safety in the
// always-loaded AskUserQuestion Format section.
// v1.2.0 activation lift: first-run-guidance section in the shared preamble.
maxSkeletonBytes: 77_300, // v1.68 fix wave (#2402); measured 76,705
minUnionBytes: 72_000,
mustContain: ['OWASP', 'STRIDE', 'daily', 'comprehensive', 'verif'],
// cso keeps its mode-dispatch + FP-filtering phases always-loaded, so the
// cross-cutting preamble growth (v1.57.2.0 AUQ-failure prose fallback ~2KB + the
// decision-memory nudge) lands it just over 1.05; headroom for the shared additions.
// v1.64+v1.65 merge sums both waves' preamble growth; measured 1.073.
maxSizeRatio: 1.08,
},
};
/** Sorted carved-skill names. Consumers derive their lists from this — no parallel lists. */
export const CARVED_SKILLS: readonly string[] = Object.freeze(
Object.keys(CARVE_GUARDS).sort(),
);