Files
gstack/scripts/gen-skill-docs.ts
T
Garry Tan df89475b17 v1.91.11.0 refactor: one state-root rule, browse route table, shared shard engine, PTY harness split, MECE review resolvers (#3002)
* refactor(resolvers): split review.ts into MECE resolver modules (pure move)

Move every function from scripts/resolvers/review.ts, unchanged, into:
- review-dashboard.ts: review dashboard, plan-file review report
- plan-gates.ts: approval check, exit-plan-mode gate, plan-file discovery,
  plan-completion audit/gate (ship + review), plan verification exec
- spec-review.ts: both spec review loops, benefits-from, anti-shortcut clause
- outside-voice-steps.ts: Codex second opinion, adversarial step, Codex plan
  review, Codex doc review, disabled-outside record
- review-scope.ts: scope drift, cross-review dedup, shared-code reuse

review.ts is deleted; index.ts imports the new modules. gen-skill-docs
output is byte-identical for every host (--host all). Test imports and
source-path references are re-pointed; the two source-text report/gate
tests in gen-skill-docs.test.ts become behavioral renders across every
consuming skill and host. All 46 touchfile entries that named review.ts
now name all five modules, guarded by a recorded selection golden.

* test(browse): black-box auth matrix for every server route and both surfaces

Drives buildFetchHandler fetchLocal/fetchTunnel with no token, wrong token,
root token, scoped token and the SSE cookie for all 33 routes, plus unmatched
paths and wrong methods. Denials assert today's exact status, body and content
type; allowed credentials assert the handler was reached. Written against the
unchanged if-chain server so the W3 route-table refactor must keep it green.

* refactor(shard-engine): move scripts/test-strict-output.ts to scripts/lib/shard-engine.ts

The shared shard engine grows from the existing strict-output module
(runShardChild, killProcessGroup, signal forwarding, strict classifier).
scripts/test-strict-output.ts stays as a re-export so existing importers,
mock.module paths and the strict-output/run-shard-child tests are unchanged.
The engine inherits the global touchfile entry; the free runner's CLI-routing
fixture copies the new module.

* refactor(resolvers): decompose the three >150-line review resolvers (output-neutral)

Split generateAdversarialStep, generateCodexPlanReview and
generatePlanCompletionAuditInner into per-section helpers whose template
literals are copied verbatim, so every function in the new modules is at
or under 150 lines. gen-skill-docs output is byte-identical for every host
(--host all, compared against 96764e80 with a fixed --link-root).

* refactor(resolvers): one outside-voice failure policy (deliberate prose unification)

outsideVoiceFailurePolicy(ctx, opts) in outside-voice.ts now renders the
auth / timeout / empty-response bullets for all four call sites that
hand-typed them (Codex second opinion, adversarial step, Codex plan
review, design outside voices). Options are explicit per site
(timeoutMinutes, onTimeout, stderrOnEmpty, fallback, escape) with no
defaults.

Deliberate generated-prose changes (every host):
- office-hours: 'Fall back to <native> subagent.' becomes
  'Fall back to the <native> subagent below.'
- plan-devex-review: the plain 'Auth failure (stderr contains ...)'
  bullets become the canonical bold bullets; auth also triggers on
  'API key'; 'auth failed' becomes 'authentication failed'.
- review/ship adversarial: 'exceeded 9 minutes and was terminated'
  becomes 'timed out after 9 minutes and was terminated'; the timeout
  is still MISSING COVERAGE.
- design outside voices: unchanged.

Adds ratchet (d) (test/outside-voice-failure-policy.test.ts) with a
reasoned allowlist for /codex's own CLI errors, the MISSING COVERAGE
retention test, refreshed codex/factory ship goldens, and outside-voice.ts
in every touchfile entry of review.ts and design.ts (selection golden
extended).

* test(pty): fake PTY session driver with an injectable clock through the runner launch seam

The three plan-skill runners take an optional PtyDriver (launch, now,
monotonic, sleep); omitted, they use the real launcher and clocks exactly as
before. test/helpers/pty/fake-session.ts feeds scripted frames through that
seam, and claude-pty-runner.runners.unit.test.ts runs observation, counting
and floor for success, deadline timeout, permission prompt and plan-ready
outcomes with no CLI or real timers. These cases must stay green unchanged
through the W4 split and the runPtySession extraction.

Touchfiles: every entry that lists claude-pty-runner.ts or pty-screen.ts now
also lists test/helpers/pty/**.

* refactor(shard-engine): run both lanes on the shared engine; lane policy injected

Engine (scripts/lib/shard-engine.ts) gains the W2 primitives: per-shard
tmp/Chromium sandbox + async cleanup backstop, log-path allocation and
full-stream log capture, one duration-seed reader/writer with a lane
predicate, LanePolicy (seed predicate + zero-execution verdict),
strictShardStatus, and the shared CLI flag loop. runShardChild takes an
optional companion (signal/settle) and waits a bounded 250ms to reap a
wall-killed child.

Free lane stops spawning shards itself: runFreeShard uses runShardChild
with trackShardBrowser as the companion (win32 path unchanged: no process
group, no negative-pid kill). Its sync state-dir removal stays lane policy.
Paid lane uses the sandbox, log, seed, verdict and flag primitives; the
hollow-shard guard applies PAID_LANE_POLICY. Lane outcomes are unchanged
(free keeps >= 0 seeds and file-count zero-exec rule; paid keeps > 0 seeds,
warning under selection and passed-empty under EVALS_ALL).

paid-free-boundary's closure assertion now names the engine module, where
the strict classifier lives.

* test(shard-engine): engine unit tests, fixture-corpus equivalence, per-lane CLI parity

- test/shard-engine.test.ts: failing/unhandled/module-load output fails both
  lanes, per-lane zero-execution and seed rules, whole-group kill on a wall
  timeout (both lanes), mocked-win32 path with no negative-pid kill,
  companion settle order, log capture, sandbox isolation, flag loop.
- test/shard-engine-equivalence.test.ts + test/fixtures/shard-equivalence:
  seven outcome fixtures plus one real shard, run through both lanes and
  compared with classifications recorded from the base runners (96764e80).
- test/shard-cli-parity.test.ts + test/fixtures/shard-cli-parity: flag set,
  defaults, validation errors and the Unknown argument error per lane match
  the base runners.

* refactor(shard-engine): decompose runFreeShard and runPaidShard to <= 150 lines

Output-neutral extraction under the fixture-corpus equivalence and runner
tests: captureFreeStream, explainFreeVerdict and logFreeRecovery (free);
paidShardCommand, settleShardSpool, settleBootstrapRetention and
printLogTail (paid). The bootstrap scope-creation block that
bootstrap-retention.test.ts evaluates stays verbatim.

* refactor(pty): split claude-pty-runner.ts into test/helpers/pty/* behind a barrel

Pure move: every line of the former 5,047-line runner lands verbatim in one
module (four private helpers gain `export` for cross-module use):
binary, screen (absorbs test/helpers/pty-screen.ts, which now re-exports it),
launch, session (PtyDriver), judge, classify, auq, plan-native, boundaries,
runners/{observation,counting,floor}. claude-pty-runner.ts re-exports the
original public surface by name; pty/ modules import siblings directly.

Tests that read the runner's source text:
- rewritten as behavioral: the unit test's model-pin tripwire (fake CLI argv:
  fallback chain, --model before extraArgs, hermetic --strict-mcp-config),
  pty-skill-seeding-wiring (runners through the fake driver; launcher through
  a fake CLI reporting CLAUDE_CONFIG_DIR). The "three wrappers forward model"
  grep is replaced by the runners' fake-driver launch assertions.
- pty-screen-session / pty-screen-supervision: stop copying runner source;
  they mock.module the real pty/screen.ts (and the fixture cleanup) instead.
- re-pointed to the owning module (they execute a sliced runner body with
  injected boundaries; no seam exists for those boundaries yet):
  eng-seeded-completion-ai, plan-floor-permission, plan-create-prepublication,
  plan-count-completion; hermetic-wiring's source guard now reads pty/launch.ts
  and scans every pty/ module for raw process.env spreads.
- plan-count-timeout and pty-output-wake mock the viewport at pty/screen.ts.

* test(ratchet-c): enforcing module/function size ratchet and moved-code touchfile coverage

Ratchet (c) ships enforcing: test/helpers/module-size.ts counts file and
top-level function lengths by brace matching over masked source (strings,
comments, regex literals and template text masked; ${} expressions kept),
covering function declarations, arrow functions assigned to consts and
route-table handler properties, with no parser dependency. Its self-test
uses template literals and code-fence braces copied from
scripts/resolvers/review.ts and design.ts. test/fixtures/module-size-ratchet.json
binds scripts/lib/shard-engine.ts (<= 800 lines, <= 150 per function) and
records the residual runner sizes (free 2352, paid 1921) as non-growth caps;
allowlist entries are keyed on file plus matched text and need a reason.
Failure output lists file:line, the rule, Fix: and the allowlist path.

touchfiles.test.ts gains the moved-code superset check over
test/fixtures/touchfile-move-goldens/ (W2 golden recorded at 96764e80:
test-strict-output.ts and test-paid-shards.ts global, test-free-shards.ts none).

* refactor(browse): declared route table replaces the buildFetchHandler if-chain

The ~1,300-line if-chain in buildFetchHandler becomes a route table:
each entry declares method, path, auth kind and surfaces, and one auth
gate in browse/src/routes/table.ts returns the per-kind denial (root-bearer,
scoped, root-or-sse-cookie: 401 Unauthorized; root-token: 403 Root token
required; extension-origin: 403 Forbidden). Unmatched requests take the
declared fallthrough (root-bearer check, then plain-text 404). Handlers move
to browse/src/routes/{core,pairing,pty,tokens,tunnel,activity,commands,files,
inspector}.ts and receive a RouteContext with auth checks as functions
instead of closing over factory locals. Dispatch order is unchanged:
tunnel filter, beforeRoute overlay, gate, handler. TUNNEL_PATHS stays a
literal in server.ts.

Behavior-preserving: the black-box auth matrix from the previous commit
passes unchanged. /memory and /inspector/events are declared root-bearer
because the blanket check always ran before their SSE-cookie branch.

Source-text route tests are rewritten as behavioral tests through
buildFetchHandler or a route's real handler with a stub RouteContext
(browse/test/route-test-harness.ts). Checks with no runtime seam are
re-pointed to the route modules: Surface type, /inspector/events SSE
helper, sanitizeReplacer imports, /pty-inject-scan sidecar-client import,
and the ngrok config lookup and startTunnel wiring that stay in server.ts.

* test(browse): stubbed-handler auth matrix and route inventory for the route table

Every ROUTES entry runs through the real dispatcher and gate with stub
handlers on each declared surface and six credentials; denials assert the
exact status and body each auth kind returned at 96764e8, admitted
credentials assert the handler ran (with the gate's TokenInfo for scoped
routes). Also pins the reviewed route inventory (method, path, auth kind,
surfaces), that every entry declares auth and surfaces, that the table's
tunnel paths equal the TUNNEL_PATHS literal with GET /connect admitted, the
unmatched fallthrough, and that the root token is rejected on every tunnel
route through buildFetchHandler.

* test(browse): ratchet (b) keeps route dispatch inside the route table

Scans browse/src/server.ts and browse/src/routes/*.ts for pathname
comparisons; only the table matcher and the tunnel-surface filter are
allowed, listed with reasons in browse/test/fixtures/route-dispatch-allowlist.json
(keyed on file plus line text). Also checks every entry declares auth and
surfaces and that gstack registers no beforeRoute overlay itself. Self-tests
plant a violation and assert the file:line, Fix: and allowlist path in the
message, that a shifted line stays allowlisted, and that a reasonless entry
is rejected.

* test: touchfile superset check for modules moved out of browse/src/server.ts

Records the paid evals selected by touching browse/src/server.ts at 96764e80
(17 E2E, 1 LLM judge) and asserts every browse/src/routes/*.ts module selects
a superset. The test reads every golden in test/fixtures/moved-module-selection/
so other moved-code goldens can sit beside it.

* test(shard-engine): give non-timeout corpus fixtures CI headroom; keep the POSIX golden off the Windows lane

Only the wall-timeout fixture keeps a 3s wall; the rest get 60s so a loaded
host cannot turn a pass into a timeout. Base and branch runners still agree
on every classification under the new walls. The Windows exclusion entry
moves the free runner's ratchet (c) residual cap to 2356 lines.

* refactor(pty): one runPtySession loop drives observation, counting and floor

test/helpers/pty/session.ts owns launch -> start -> (poll -> tick)* ->
timeout and the failure contract the three runners each hand-rolled: the
run's own error wins over capture and close errors, close always runs, owned
fixture cleanup runs last (also when launch fails). Each runner now supplies a
PtySessionPlan: its boot/command step, poll cadence (2s observation/floor
sleep; counting's output wake + 250ms coalesce), tick policy (permission
handling, native identity, terminal rules stay per runner because they differ)
and capture hooks. The runner bodies are decomposed into top-level steps so no
function exceeds 150 lines; behavior is unchanged and the fake-driver cases
from the first W4 commit pass unmodified.

The counting capture step and the native completion-summary predicate are now
named functions (countingCapture, isNativeCompletionSummary), so
plan-create-prepublication and plan-count-completion call them directly
instead of executing sliced source. The two harnesses that still execute a
sliced runner body with injected boundaries (eng-seeded-completion-ai,
plan-floor-permission) pass the PtyDriver seam instead of overriding
Date/Bun.sleep.

* test(ratchet-c): register route modules, review resolver modules and server.ts residual cap

* refactor(pty): decompose launchClaudePty and engNumberedFindingAUQ under 150 lines

launchClaudePty (349 lines) becomes launch preparation (args, hermetic
child env, owned state roots), recorder creation, spawn, the trust-dialog
watcher, close, and the session handle over one PtyProcess state object. The
failure order is unchanged: abort the viewport, dispose any recorders created
so far, dispose the viewport, rethrow. The --model / --strict-mcp-config
ordering and seedSkills wiring stay pinned by the behavioral fake-CLI tests.

engNumberedFindingAUQ (345 lines) keeps its guards and dispatch; each
self-contained issue family (declared cache, library retry hooks, cache
owner, injected singleton, shared writers, injected export) moves verbatim
into its own function. Every pty/ module is now <= 800 lines and every
top-level function <= 150 lines.

* test(pty): split claude-pty-runner.unit.test.ts along the pty/ module seams

The 188 unit tests move verbatim into claude-pty-runner.{screen,classify,
auq,launch,plan-native,boundaries}.unit.test.ts (test names unchanged; each
file imports only what it uses from the barrel). The five files that no longer
read a SKILL.md template join the test-of-test ratchet baseline with a reason.

* test(touchfiles): moved PTY modules keep their paid-eval selection

test/fixtures/touchfile-selection/w4-pty.json records, at 96764e8, the paid
evals selected by touching test/helpers/claude-pty-runner.ts (20) and
test/helpers/pty-screen.ts (20). touchfiles.test.ts now asserts every .ts file
under test/helpers/pty/ (and pty/screen.ts for both sources) selects a
superset, reading every golden in that directory so later moves can add one;
a planted-violation case pins the report and its Fix line.

* fix(browse): unexchanged pair setup keys no longer authenticate bearer requests

validateToken accepted a gsk_setup_ key as a bearer on /command, /batch and
/file (found while building the W3 auth matrix). A setup key now only
authenticates the /connect exchange.

* W1: one state-root owner (lib/state-root.ts + bin/gstack-state-root.sh), gstack-paths --explain and fail-stop, parity tests

* W1: guarded migration of every executable state-root site; uninstall deletes only ~/.gstack

Bins, careful/freeze hooks, setup, upgrade migrations, browse/src, design,
ios-qa daemon, lib and scripts resolve the state root through
bin/gstack-state-root.sh (bash) or lib/state-root.ts (TS). Bins source the
twin and stop with a reinstall message when it is missing; hooks source it
and never spawn gstack-paths. browse/src/config.ts and lib/cso/state.ts
delegate to resolveStateRoot. Analytics writers and readers move together
so the usage log stays one file. gstack-uninstall deletes state only at
~/.gstack, refuses (exit 2) when it resolves to /, $HOME or an ancestor,
the checkout or the git root, and leaves any other resolved root in place
with the removal command. Fixtures that copy single bins now copy the twin.

* W1: privacy keys and trust-policy deny tiers merge across state roots; gstack-config reporting; test hermeticity

readConfigKey / gstack_read_config_key return the most restrictive
telemetry, memorable_recall, codex_reviews and update_check across the
resolved root and ~/.gstack; other keys read the resolved root only.
gstack-config set reports an overriding root with the exact override
command, list shows the winning root and a root-variable disagreement line.
gstack-gbrain-repo-policy get merges deny/read-only tiers. gstack-egress
reads through readConfigKey. test-setup.ts strips inherited
GSTACK_STATE_ROOT/GSTACK_STATE_DIR and redirects the legacy root.

* W1: shared hook logging helper (hosts/claude/hooks/hook-log.ts)

One hook-errors.log writer: root from resolveStateRoot, 0600 on every
append, opt-in rate limit used only by memorable-user-prompt. The five
hooks route through it.

* W1: docs/state-root.md and README troubleshooting pointer

Precedence table, a real --explain example, the move-your-state recipe,
merged privacy keys, the uninstall rule, the resolver-failure fix, and the
plugin-mode note (evidence gate: no official plugin distribution).

* W1b: template and resolver prose resolve state through guarded gstack-paths; ratchet (a)

Every gstack-paths eval in templates and resolvers carries the fail-stop
guard; executable ~/.gstack paths in bash blocks (context recovery preamble,
eureka log, analytics, project artifacts, upgrade snooze, setup-gbrain lock,
retro snapshots, ship consent marker) use $GSTACK_STATE_ROOT, and the writer
prose that pairs with them points at the printed PROJECT_DIR / RETRO_FILE.
ship drops export GSTACK_STATE_ROOT. SKILL.md regenerated (claude + codex),
ship goldens re-pinned, parity and context-budget caps raised to the measured
sizes with notes. test/state-root-ratchet.test.ts enforces the rule with a
reasoned allowlist; W1 touchfile entries plus a superset golden.

* refactor: apply W1 state-root edits in W2/W3/W5-owned files; one moved-code touchfile golden for all workstreams

* test: fold the moved-code touchfile golden into touchfiles.test.ts; fix integration fixture closure and caps

* v1.91.11.0: CHANGELOG, TODOS, docs and conventions for the refactor wave

* test: re-measure plan-ceo/design-consultation caps and ship goldens after the guarded plan-discovery and spec-review blocks; add the state-root twin to the workflow-boundaries fixture

* fix(windows): migrations resolve their directory with either path separator; state-root parity compares under the HOME Git Bash actually sees

* fix(review,ship): state plan-check timing after smoke expiry and test_stub Skip semantics (review workflow judge clarity)

* test(qa-eval): webhook fix eval asks for the fix loop's post-repair probes; eight-scenario coverage stays in the report-only case and the harness recheck

* test(qa-eval): re-pin the webhook prompt contract to the fix-loop stage; R29 coverage omissions stay bound by the report-only case

* fix(review,ship): plan checks publish a checkpoint before each probe; only the smoke expiry stop is skipped

* fix(qa): carry #2999's checkpoint receipt link, report-template line and full-revision placeholder (identical hunks)

* test(qa-callers): disable git auto maintenance in the caller fixture

Git 2.47+ runs auto maintenance detached after commit; on the CI runner's git
2.55 it rewrote .git/objects fan-out directories while the write observer was
running, which surfaced as unauthorized mutations. Same gc.auto=0 /
maintenance.auto=false guard the shared-libs fixture already uses.

* test(plan-mode-no-op): require prose evidence for the prose-fallback members so a spinner-frame judge verdict cannot end the run as asked

* test(ship-docsync): carry #2999's seeded-attempt docsync harness (identical files)

The doc-sync fault cases replayed attempt 1 before reaching their gate and ran
out of their 285s budget. The fixture now seeds attempt 1 and the parent starts
at the gate under test. Taken byte-identical from origin/capy/audit-fix-wave
(fb526898, e6ac813d, 6ce10ff7, d0c53577, 77cce3be). Local: stale-before,
recovery and late-result 6/6 PASS (97-164s); the whole file 12/12 PASS.
2026-10-01 11:57:48 -07:00

1107 lines
50 KiB
TypeScript

#!/usr/bin/env bun
/**
* Generate SKILL.md files from .tmpl templates.
*
* Pipeline:
* read .tmpl → find {{PLACEHOLDERS}} → resolve from source → format → write .md
*
* Supports --dry-run: generate to memory, exit 1 if different from committed file.
* Used by skill:check and CI freshness checks.
*/
import { discoverTemplates, discoverSectionTemplates, includesSkill } from './discover-skills';
import { externalSkillName, extractNameAndDescription } from './external-skill-names';
export { extractNameAndDescription } from './external-skill-names';
import { generateLlmsTxt } from './gen-llms-txt';
import { generateAgentsDigest, DIGEST_RELPATH, DIGEST_BYTE_BUDGET } from './gen-agents-digest';
import { generateDesignChecklistMd } from './resolvers/design-checklist';
import { DOM_DUMP_SCRIPT, DOM_DUMP_FILE } from '../lib/dom-dump-script';
import * as fs from 'fs';
import * as path from 'path';
import type { Host, TemplateContext } from './resolvers/types';
import { HOST_PATHS } from './resolvers/types';
import { RESOLVERS } from './resolvers/index';
import { usesLazySections } from './resolvers/sections';
import { ALL_HOST_NAMES, resolveHostArg, getHostConfig } from '../hosts/index';
import type { HostConfig } from './host-config';
const ROOT = path.resolve(import.meta.dir, '..');
import { ALL_MODEL_NAMES, resolveModel, type Model } from './models';
import { resolveStateRoot } from '../lib/state-root';
type HostArg = Host | 'all';
/** Internal render settings. Inputs always come from ROOT; output routing and
* content links are separate so checks can render canonical bytes into scratch. */
export interface GenerationOptions {
host?: HostArg;
dryRun?: boolean;
outputRoot?: string;
contentLinkRoot?: string | null;
model?: Model | null;
catalogMode?: 'trim' | 'full';
explainLevel?: 'default' | 'terse';
respectDetection?: boolean;
log?: (message: string) => void;
}
interface RenderOptions {
outputRoot: string;
contentLinkRoot: string | null;
model: Model | null;
catalogMode: 'trim' | 'full';
explainLevel: 'default' | 'terse';
gbrainDetected: boolean;
}
export interface GeneratedArtifact {
relativePath: string;
kind: 'skill' | 'section' | 'metadata' | 'openclaw' | 'index' | 'digest' | 'asset';
host?: Host;
}
export interface GenerationDiagnostic {
kind: 'stale' | 'error' | 'warning' | 'skipped';
message: string;
host?: Host;
relativePath?: string;
}
export interface GenerationResult {
exitCode: 0 | 1;
artifacts: GeneratedArtifact[];
diagnostics: GenerationDiagnostic[];
}
/** Canonical generation never reads local detection state unless opted in. */
function loadGbrainOverride(respectDetection: boolean): boolean {
if (!respectDetection) return false;
const stateDir = resolveStateRoot();
try {
const json = JSON.parse(fs.readFileSync(path.join(stateDir, 'gbrain-detection.json'), 'utf-8'));
// Slow, remote, and locked engines are still usable (#1964/#2051/#2456).
return ['ok', 'timeout', 'thin-client', 'engine-locked'].includes(json.gbrain_local_status ?? '');
} catch {
return false;
}
}
function effectiveSuppressedResolvers(hostConfig: HostConfig, options: RenderOptions): Set<string> {
let list = hostConfig.suppressedResolvers || [];
if (options.gbrainDetected) {
list = list.filter(r => r !== 'GBRAIN_CONTEXT_LOAD' && r !== 'GBRAIN_SAVE_RESULTS');
}
return new Set(list);
}
/** Parse CLI settings only when executing, never when imported by tests/checks. */
function parseGenerationArgs(args: string[]): GenerationOptions {
const value = (flag: string): string | undefined => {
const index = args.findIndex(arg => arg === flag || arg.startsWith(`${flag}=`));
if (index < 0) return undefined;
const arg = args[index];
const result = arg.startsWith(`${flag}=`) ? arg.slice(flag.length + 1) : args[index + 1];
if (!result || result.startsWith('--')) throw new Error(`${flag} requires a value`);
return result;
};
const hostValue = value('--host') ?? 'claude';
const host = hostValue === 'all' ? 'all' : resolveHostArg(hostValue) as Host;
const modelValue = value('--model');
const model = modelValue === undefined ? null : resolveModel(modelValue);
if (modelValue !== undefined && !model) {
throw new Error(`Unknown model: ${modelValue}. Use ${ALL_MODEL_NAMES.join(', ')}, or a family variant (e.g., claude-opus-4-7, gpt-5.4-mini, o3).`);
}
const catalogMode = value('--catalog-mode') ?? 'trim';
if (catalogMode !== 'trim' && catalogMode !== 'full') {
throw new Error(`Unknown catalog mode: ${catalogMode}. Use 'trim' (default) or 'full'.`);
}
const explainLevel = value('--explain-level') ?? 'default';
if (explainLevel !== 'default' && explainLevel !== 'terse') {
throw new Error(`Unknown explain level: ${explainLevel}. Use 'default' or 'terse'.`);
}
const outDir = value('--out-dir');
const linkRoot = value('--link-root');
// Swap-in callers use --link-root for the FINAL serving path (#2692).
// Direct --out-dir callers retain their existing links into the render.
return {
host, model, catalogMode, explainLevel,
dryRun: args.includes('--dry-run'),
respectDetection: args.includes('--respect-detection'),
outputRoot: outDir === undefined ? ROOT : path.resolve(outDir),
contentLinkRoot: linkRoot !== undefined ? path.resolve(linkRoot)
: outDir !== undefined ? path.resolve(outDir) : null,
};
}
/** Repoint only Claude section links, retaining global bin/browse/doc paths. */
function rewriteSectionBase(content: string, linkRoot: string | null): string {
if (!linkRoot) return content;
// Callback replacement preserves literal $ sequences in configured paths.
return content.replace(
/~\/\.claude\/skills\/gstack\/([^\s)`"'*]+\/sections\/)/g,
(_m, p1: string) => `${linkRoot}/${p1}`,
);
}
// HostPaths, HOST_PATHS, and TemplateContext imported from ./resolvers/types (line 7-8)
// Design constants (AI_SLOP_BLACKLIST, OPENAI_HARD_REJECTIONS, OPENAI_LITMUS_CHECKS)
// live in ./resolvers/constants and are consumed by resolvers directly.
// ─── External Host Helpers ───────────────────────────────────
// ─── Voice Trigger Processing ────────────────────────────────
/**
* Extract voice-triggers YAML list from frontmatter.
* Returns an array of trigger strings, or [] if no voice-triggers field.
*/
function extractVoiceTriggers(content: string): string[] {
const fmStart = content.indexOf('---\n');
if (fmStart !== 0) return [];
const fmEnd = content.indexOf('\n---', fmStart + 4);
if (fmEnd === -1) return [];
const frontmatter = content.slice(fmStart + 4, fmEnd);
const triggers: string[] = [];
let inVoice = false;
for (const line of frontmatter.split('\n')) {
if (/^voice-triggers:/.test(line)) { inVoice = true; continue; }
if (inVoice) {
const m = line.match(/^\s+-\s+"(.+)"$/);
if (m) triggers.push(m[1]);
else if (!/^\s/.test(line)) break;
}
}
return triggers;
}
/**
* Preprocess voice triggers: fold voice-triggers YAML field into description,
* then strip the field from frontmatter. Must run BEFORE transformFrontmatter
* and extractNameAndDescription so all hosts see the updated description.
*/
function processVoiceTriggers(content: string): string {
const triggers = extractVoiceTriggers(content);
if (triggers.length === 0) return content;
// Strip voice-triggers block from frontmatter
content = content.replace(/^voice-triggers:\n(?:\s+-\s+"[^"]*"\n?)*/m, '');
// Get current description (after stripping voice-triggers, so it's clean)
const { description } = extractNameAndDescription(content);
if (!description) return content;
// Build new description with voice triggers appended
const voiceLine = `Voice triggers (speech-to-text aliases): ${triggers.map(t => `"${t}"`).join(', ')}.`;
const newDescription = description + '\n' + voiceLine;
// Replace old indented description with new in frontmatter
const oldIndented = description.split('\n').map(l => ` ${l}`).join('\n');
const newIndented = newDescription.split('\n').map(l => ` ${l}`).join('\n');
content = content.replace(oldIndented, newIndented);
return content;
}
// Export for testing
export { extractVoiceTriggers, processVoiceTriggers };
// ─── Catalog Trim (v1.45.0.0 T4) ─────────────────────────────
//
// Frontmatter `description:` blocks today pack: a one-line outcome, "Use when
// asked to..." voice triggers, "Proactively..." routing guidance, and a
// "(gstack)" tag. This pile is the always-loaded catalog surface — every
// session pays for the full text. The catalog trim splits the description
// into a one-line catalog entry (lead sentence + "(gstack)") that stays in
// the frontmatter, and a "## When to invoke" body section that holds the
// routing/voice triggers prose for in-skill discovery.
//
// Opt-out: `--catalog-mode=full` keeps v1.44 behavior (no trim, full
// description in frontmatter). Use when debugging routing regressions or
// when shipping skills to hosts that depend on the legacy fat catalog.
export interface CatalogParts {
lead: string; // First sentence — kept in catalog
routingProse: string; // "Use when asked to...", "Proactively..." paragraphs
voiceLine: string | null; // "Voice triggers (speech-to-text aliases): ..." line if present
hasGstackTag: boolean;
}
export function splitCatalogDescription(description: string): CatalogParts {
// Voice triggers line (folded in by processVoiceTriggers earlier)
const voiceMatch = description.match(/Voice triggers \(speech-to-text aliases\):[^\n]+/);
const voiceLine = voiceMatch ? voiceMatch[0] : null;
let working = voiceLine ? description.replace(voiceLine, '').trim() : description.trim();
const hasGstackTag = /\(gstack\)/.test(working);
if (hasGstackTag) working = working.replace(/\(gstack\)/, '').trim();
// Lead = first sentence, ending at the first `.`/`!`/`?` that is followed by
// whitespace or end-of-text. Terminator chars NOT followed by whitespace/end
// (embedded periods in "TODOS.md", URLs, "v1.45.0.0") are consumed by the
// second alternative `[.!?](?!\s|$)` and do NOT end the sentence. The two
// alternatives are disjoint character classes, so there is no ambiguity and
// no catastrophic-backtracking risk. If no terminator-followed-by-boundary
// exists at all, we fall back to a 20-word cut below.
// First normalize to single-line for sentence detection, then back out.
const collapsed = working.replace(/\s+/g, ' ').trim();
const sentenceMatch = collapsed.match(/^((?:[^.!?]|[.!?](?!\s|$))*[.!?])(?:\s|$)/);
// sentenceLead is the FULL first sentence (no truncation). We compute routing
// from this position, then optionally truncate the displayed lead afterwards.
// Truncating first then computing routing was the v1.45.0.0 bug — when the
// first sentence exceeded 200 chars, the routing extraction would lose the
// entire tail of the description (design-consultation's "Use when..."
// routing prose silently dropped).
const sentenceLead = sentenceMatch ? sentenceMatch[1].trim() : collapsed.split(/\s/).slice(0, 20).join(' ');
// Routing prose: everything AFTER the first sentence boundary in the collapsed view.
const leadInCollapsed = collapsed.indexOf(sentenceLead);
const routingCollapsed = leadInCollapsed >= 0
? collapsed.slice(leadInCollapsed + sentenceLead.length).trim()
: '';
// Now produce the displayed lead — truncated if too long. The original
// sentenceLead is preserved for routing extraction below.
let lead = sentenceLead;
if (lead.length > 200) {
const trunc = lead.slice(0, 197);
const lastSpace = trunc.lastIndexOf(' ');
lead = (lastSpace > 60 ? trunc.slice(0, lastSpace) : trunc) + '...';
}
// Restore line breaks for routing prose by mapping back to original layout.
// Use original whitespace structure where possible; fall back to collapsed.
// Anchor recovery on sentenceLead (the untruncated first sentence) — not
// `lead` (which may have a "..." suffix and won't substring-match `working`).
let routingProse = routingCollapsed;
const collapsedLeadIdx = working.replace(/\s+/g, ' ').indexOf(sentenceLead);
if (collapsedLeadIdx >= 0) {
let consumed = 0;
let cut = 0;
for (let i = 0; i < working.length && consumed < collapsedLeadIdx + sentenceLead.length; i++) {
if (/\s/.test(working[i])) {
if (i === 0 || /\s/.test(working[i - 1])) continue;
consumed += 1;
} else {
consumed += 1;
}
cut = i + 1;
}
const tail = working.slice(cut).trim();
if (tail.length > 0) routingProse = tail;
}
return { lead, routingProse, voiceLine, hasGstackTag };
}
/** Build the catalog-trimmed `description:` block. */
export function buildTrimmedDescription(parts: CatalogParts): string {
const lead = parts.lead.trim();
const suffix = parts.hasGstackTag ? ' (gstack)' : '';
return `${lead}${suffix}`;
}
/** Build the body section that holds the routing/voice prose. */
export function buildWhenToInvokeSection(parts: CatalogParts): string {
const lines: string[] = ['## When to invoke this skill', ''];
if (parts.routingProse) {
lines.push(parts.routingProse);
lines.push('');
}
if (parts.voiceLine) {
lines.push(parts.voiceLine);
lines.push('');
}
return lines.join('\n');
}
/**
* Render a string as a YAML inline scalar value (the text after `key: `),
* quoting only when a plain scalar would be invalid or ambiguous.
*
* The bug this guards (#1778): a description like "Ship workflow: detect..."
* emitted as a plain scalar has an interior ": " that a strict YAML parser
* (Codex/OpenAI skill loading) reads as a nested mapping and rejects with
* "mapping values are not allowed in this context". When quoting is needed we
* fall back to JSON.stringify, which produces a double-quoted scalar that YAML
* accepts verbatim (YAML is a superset of JSON for flow scalars). Strings that
* are already valid plain scalars pass through unchanged to keep regen diffs small.
*/
export function toYamlInlineScalar(s: string): string {
const needsQuote =
s.length === 0 ||
s !== s.trim() || // leading/trailing whitespace
/:(\s|$)/.test(s) || // "foo: bar" / trailing colon → mapping ambiguity
/\s#/.test(s) || // " #" → inline comment
/\.\.\./.test(s) || // "..." → document-end marker; strict parsers reject mid-scalar (catalog-trim truncation appends it)
/^[\s>|&*!%@`"'#,\[\]{}?-]/.test(s); // leading YAML indicator char
return needsQuote ? JSON.stringify(s) : s;
}
/**
* Apply catalog trim to a SKILL.md body:
* - shorten frontmatter `description:` to lead + (gstack)
* - insert "## When to invoke" body section AFTER the generated header
* (so it lands near the top of body content, where routing guidance
* belongs)
*
* Returns the rewritten content plus the extracted parts.
*/
export function applyCatalogTrim(content: string, skillName: string): { content: string; parts: CatalogParts } | null {
// Locate description block in frontmatter
if (!content.startsWith('---\n')) return null;
const fmEnd = content.indexOf('\n---', 4);
if (fmEnd === -1) return null;
const frontmatter = content.slice(4, fmEnd);
// Match `description: |` block + indented body lines
const descMatch = frontmatter.match(/^description:\s*\|?\s*\n((?:\s{2,}.*(?:\n|$))+)/m)
|| frontmatter.match(/^description:\s+(.+)$/m);
if (!descMatch) return null;
// Extract full description text
let descText: string;
if (descMatch[0].startsWith('description: |') || /^description:\s*\|/.test(descMatch[0])) {
descText = descMatch[1].split('\n').map(l => l.replace(/^\s{2}/, '')).join('\n').trim();
} else {
descText = descMatch[1].trim();
}
// Skip skills with very short descriptions (already trimmed or no routing prose).
// Below ~120 chars, splitting adds no value.
if (descText.length < 120) return null;
const parts = splitCatalogDescription(descText);
// If lead + (gstack) is already most of the text, no trim needed.
const trimmedLen = buildTrimmedDescription(parts).length;
if (trimmedLen >= descText.length - 20) return null;
// Replace description in frontmatter — keep trailing newline so the next
// YAML field doesn't collide on the same line as the description value.
// Quote the value when it would be an invalid YAML plain scalar (the common
// case: an interior ": " like "Ship workflow: detect..." which a strict YAML
// parser reads as a nested mapping and rejects — #1778). toYamlInlineScalar
// only quotes when needed, so descriptions without special chars stay plain.
const newDesc = buildTrimmedDescription(parts);
// Function replacer (not a string) so a `$` in the description — e.g. a future
// skill referencing `$B`/`$D` — can't be interpreted as a `$&`/`$1` replacement
// pattern and silently corrupt the frontmatter.
const newDescLine = `description: ${toYamlInlineScalar(newDesc)}\n`;
const newFrontmatter = frontmatter.replace(descMatch[0], () => newDescLine);
let newContent = '---\n' + newFrontmatter + content.slice(fmEnd);
// Insert body section after frontmatter (after the closing ---\n and any
// existing GENERATED header). We insert before the first non-comment line.
const bodyStart = newContent.indexOf('\n---\n') + 5;
const whenToInvoke = '\n' + buildWhenToInvokeSection(parts).trim() + '\n';
// Skip past the generated header if present (it lives after frontmatter close)
const headerMatch = newContent.slice(bodyStart).match(/^(<!--[^>]*-->\s*\n)+/);
const insertAt = bodyStart + (headerMatch ? headerMatch[0].length : 0);
newContent = newContent.slice(0, insertAt) + whenToInvoke + '\n' + newContent.slice(insertAt);
return { content: newContent, parts };
}
const OPENAI_SHORT_DESCRIPTION_LIMIT = 120;
function condenseOpenAIShortDescription(description: string): string {
const firstParagraph = description.split(/\n\s*\n/)[0] || description;
const collapsed = firstParagraph.replace(/\s+/g, ' ').trim();
if (collapsed.length <= OPENAI_SHORT_DESCRIPTION_LIMIT) return collapsed;
const truncated = collapsed.slice(0, OPENAI_SHORT_DESCRIPTION_LIMIT - 3);
const lastSpace = truncated.lastIndexOf(' ');
const safe = lastSpace > 40 ? truncated.slice(0, lastSpace) : truncated;
return `${safe}...`;
}
function generateOpenAIYaml(displayName: string, shortDescription: string): string {
return `interface:
display_name: ${JSON.stringify(displayName)}
short_description: ${JSON.stringify(shortDescription)}
default_prompt: ${JSON.stringify(`Use ${displayName} for this task.`)}
policy:
allow_implicit_invocation: true
`;
}
/**
* Transform frontmatter for external hosts.
* Claude: strips `sensitive:` field (only Factory uses it).
* Codex: keeps name + description only, enforces 1024-char limit.
* Factory: keeps name + description + user-invocable, conditionally adds disable-model-invocation.
*/
function transformFrontmatter(content: string, host: Host): string {
const hostConfig = getHostConfig(host);
const fm = hostConfig.frontmatter;
if (fm.mode === 'denylist') {
// Denylist mode: strip listed fields, keep everything else
for (const field of fm.stripFields || []) {
if (field === 'voice-triggers') {
content = content.replace(/^voice-triggers:\n(?:\s+-\s+"[^"]*"\n?)*/m, '');
} else {
content = content.replace(new RegExp(`^${field}:\\s*.*\\n`, 'm'), '');
}
}
return content;
}
// Allowlist mode: reconstruct frontmatter with only allowed fields
const fmStart = content.indexOf('---\n');
if (fmStart !== 0) return content;
const fmEnd = content.indexOf('\n---', fmStart + 4);
if (fmEnd === -1) return content;
const frontmatter = content.slice(fmStart + 4, fmEnd);
const body = content.slice(fmEnd + 4);
const { name, description } = extractNameAndDescription(content);
// Description limit enforcement
if (fm.descriptionLimit) {
const behavior = fm.descriptionLimitBehavior || 'error';
if (description.length > fm.descriptionLimit) {
if (behavior === 'error') {
throw new Error(
`${hostConfig.displayName} description for "${name}" is ${description.length} chars (max ${fm.descriptionLimit}). ` +
`Compress the description in the .tmpl file.`
);
} else if (behavior === 'warn') {
console.warn(`WARNING: ${hostConfig.displayName} description for "${name}" exceeds ${fm.descriptionLimit} chars`);
}
// 'truncate' — silently proceed
}
}
// Build frontmatter with allowed fields
const indentedDesc = description.split('\n').map(l => ` ${l}`).join('\n');
let newFm = `---\nname: ${name}\ndescription: |\n${indentedDesc}\n`;
// Add extra fields (host-wide)
if (fm.extraFields) {
for (const [key, value] of Object.entries(fm.extraFields)) {
if (key !== 'name' && key !== 'description') {
newFm += `${key}: ${value}\n`;
}
}
}
// Add conditional fields
if (fm.conditionalFields) {
for (const rule of fm.conditionalFields) {
const match = Object.entries(rule.if).every(([k, v]) =>
new RegExp(`^${k}:\\s*${v}`, 'm').test(frontmatter)
);
if (match) {
for (const [key, value] of Object.entries(rule.add)) {
newFm += `${key}: ${value}\n`;
}
}
}
}
// Preserve additional keepFields beyond name and description
if (fm.keepFields) {
for (const field of fm.keepFields) {
if (field === 'name' || field === 'description') continue;
// Match YAML field with possible multi-line/array value (indented lines after colon)
const fieldMatch = frontmatter.match(new RegExp(`^${field}:(.*(?:\\n(?:[ \\t]+.+))*)`, 'm'));
if (fieldMatch) {
newFm += `${field}:${fieldMatch[1]}\n`;
}
}
}
// Rename fields (copy values from template frontmatter with new keys)
if (fm.renameFields) {
for (const [oldName, newName] of Object.entries(fm.renameFields)) {
const fieldMatch = frontmatter.match(new RegExp(`^${oldName}:(.+(?:\\n(?:\\s+.+)*)?)`, 'm'));
if (fieldMatch) {
newFm += `${newName}:${fieldMatch[1]}\n`;
}
}
}
newFm += '---';
return newFm + body;
}
/**
* Extract hook descriptions from frontmatter for inline safety prose.
* Returns a description of what the hooks do, or null if no hooks.
*/
function extractHookSafetyProse(tmplContent: string): string | null {
if (!tmplContent.match(/^hooks:/m)) return null;
// Parse the hook matchers to build a human-readable safety description
const matchers: string[] = [];
const matcherRegex = /matcher:\s*"(\w+)"/g;
let m;
while ((m = matcherRegex.exec(tmplContent)) !== null) {
if (!matchers.includes(m[1])) matchers.push(m[1]);
}
if (matchers.length === 0) return null;
// Build safety prose based on what tools are hooked
const toolDescriptions: Record<string, string> = {
Bash: 'check bash commands for destructive operations (rm -rf, DROP TABLE, force-push, git reset --hard, etc.) before execution',
Edit: 'verify file edits are within the allowed scope boundary before applying',
Write: 'verify file writes are within the allowed scope boundary before applying',
};
const safetyChecks = matchers
.map(t => toolDescriptions[t] || `check ${t} operations for safety`)
.join(', and ');
return `> **Safety Advisory:** This skill includes safety checks that ${safetyChecks}. When using this skill, always pause and verify before executing potentially destructive operations. If uncertain about a command's safety, ask the user for confirmation before proceeding.`;
}
// ─── External Host Config (now derived from hosts/*.ts) ──────
// EXTERNAL_HOST_CONFIG replaced by getHostConfig() from hosts/index.ts
// ─── Template Processing ────────────────────────────────────
const GENERATED_HEADER = `<!-- AUTO-GENERATED from {{SOURCE}} — do not edit directly -->\n<!-- Regenerate: bun run gen:skill-docs -->\n`;
/**
* Apply a host's configured path + tool rewrites. Extracted so both SKILL.md
* (via processExternalHost) and section files (via processSectionTemplate) get
* identical per-host treatment — a section's cross-references must rewrite the
* same way the parent skill's do, or external hosts get wrong paths.
*/
function applyHostRewrites(content: string, hostConfig: HostConfig): string {
let result = content;
for (const rewrite of hostConfig.pathRewrites) {
result = result.replaceAll(rewrite.from, rewrite.to);
}
if (hostConfig.toolRewrites) {
for (const [from, to] of Object.entries(hostConfig.toolRewrites)) {
result = result.replaceAll(from, to);
}
}
return result;
}
/**
* Resolve {{PLACEHOLDER}} / {{NAME:arg}} tokens against the RESOLVERS registry,
* honoring host suppression and appliesTo gating, then assert nothing is left
* unresolved. Extracted so SKILL.md and section templates resolve through the
* exact same path — a security/sanitization fix to one can't miss the other.
*/
/**
* A second {{PREAMBLE}} in one template re-expands the entire ~12K-token
* preamble mid-document (#2508/#2362 — a PROSE mention of the macro in
* spec/SKILL.md.tmpl expanded it a second time, +43KB per /spec load).
* Resolution is context-blind, so any second occurrence — code fence, prose,
* anywhere — is a generation error, never intentional. Throw at render time
* so the mistake cannot reach a generated SKILL.md again.
*/
export function assertSinglePreamble(tmplContent: string, relTmplPath: string): void {
const count = (tmplContent.match(/\{\{PREAMBLE\}\}/g) || []).length;
if (count > 1) {
throw new Error(
`${relTmplPath} contains {{PREAMBLE}} ${count} times — a template may reference it `
+ `at most once (each occurrence expands the full preamble; see #2508/#2362). `
+ `Refer to "the preamble" in prose instead of the macro.`,
);
}
}
function resolvePlaceholders(
tmplContent: string,
ctx: TemplateContext,
hostConfig: HostConfig,
relTmplPath: string,
options: RenderOptions,
): string {
assertSinglePreamble(tmplContent, relTmplPath);
// effectiveSuppressedResolvers() honors --respect-detection: when gbrain is
// detected locally, GBRAIN_* resolvers un-suppress. Shared by SKILL.md and
// section generation so both paths get the same gbrain-aware behavior.
const suppressed = effectiveSuppressedResolvers(hostConfig, options);
const onePass = (input: string): string =>
input.replace(/\{\{(\w+(?::[^}]+)?)\}\}/g, (_match, fullKey) => {
const parts = fullKey.split(':');
const resolverName = parts[0];
const args = parts.slice(1);
if (suppressed.has(resolverName)) return '';
const resolve = RESOLVERS[resolverName];
if (!resolve) throw new Error(`Unknown placeholder {{${resolverName}}} in ${relTmplPath}`);
return args.length > 0 ? resolve(ctx, args) : resolve(ctx);
});
// Multi-pass: a resolver may emit content that itself contains {{TOKENS}} — the
// {{SECTION:id}} resolver inlines a section template (with its own resolvers)
// for non-Claude hosts. .replace() doesn't re-scan inserted text, so loop until
// the output stabilizes. Bounded to avoid an infinite loop if a resolver ever
// emits its own placeholder; 6 passes is far more nesting than any skill needs.
let content = tmplContent;
for (let pass = 0; pass < 6; pass++) {
const next = onePass(content);
if (next === content) break;
content = next;
}
const remaining = content.match(/\{\{(\w+(?::[^}]+)?)\}\}/g);
if (remaining) {
throw new Error(`Unresolved placeholders in ${relTmplPath}: ${remaining.join(', ')}`);
}
return content;
}
/**
* Build the TemplateContext from a template's frontmatter. Shared by SKILL.md
* and section generation so sections inherit the SAME context the parent skill
* resolves with (skillName, tier, benefitsFrom, interactive) — enforced by
* test/template-context-parity.test.ts. skillNameOverride lets section
* generation pin the parent skill's name instead of deriving "sections".
*/
function buildContext(
tmplContent: string,
tmplPath: string,
host: Host,
options: RenderOptions,
skillNameOverride?: string,
): TemplateContext {
const { name: extractedName } = extractNameAndDescription(tmplContent);
const skillName = skillNameOverride || extractedName || path.basename(path.dirname(tmplPath));
const benefitsMatch = tmplContent.match(/^benefits-from:\s*\[([^\]]*)\]/m);
const benefitsFrom = benefitsMatch
? benefitsMatch[1].split(',').map(s => s.trim()).filter(Boolean)
: undefined;
const tierMatch = tmplContent.match(/^preamble-tier:\s*(\d+)$/m);
const preambleTier = tierMatch ? parseInt(tierMatch[1], 10) : undefined;
const interactiveMatch = tmplContent.match(/^interactive:\s*(true|false)\s*$/m);
const interactive = interactiveMatch ? interactiveMatch[1] === 'true' : undefined;
return {
skillName, tmplPath, benefitsFrom, host, paths: HOST_PATHS[host],
preambleTier, model: options.model ?? getHostConfig(host).defaultModel, interactive, explainLevel: options.explainLevel,
};
}
/**
* Process external host output: routing, frontmatter, path rewrites, metadata.
* Shared between Codex and Factory (and future external hosts).
*/
function processExternalHost(
content: string,
tmplContent: string,
host: Host,
skillDir: string,
extractedDescription: string,
ctx: TemplateContext,
options: RenderOptions,
frontmatterName?: string,
): { content: string; outputPath: string; symlinkLoop: boolean; metadata?: { outputPath: string; content: string } } {
const hostConfig = getHostConfig(host);
const name = externalSkillName(skillDir === '.' ? '' : skillDir, frontmatterName);
// --out-dir mirrors the host tree (outputs only; inputs read from ROOT).
const outputDir = path.join(options.outputRoot, hostConfig.hostSubdir, 'skills', name);
const outputPath = path.join(outputDir, 'SKILL.md');
// Guard against symlink loops
let symlinkLoop = false;
const claudePath = ctx.tmplPath.replace(/\.tmpl$/, '');
try {
const resolvedClaude = fs.realpathSync(claudePath);
const resolvedExternal = path.join(fs.realpathSync(path.dirname(outputPath)), path.basename(outputPath));
if (resolvedClaude === resolvedExternal) {
symlinkLoop = true;
}
} catch {
// realpathSync fails if file doesn't exist yet — no symlink loop
}
// Extract hook safety prose BEFORE transforming frontmatter (which strips hooks)
const safetyProse = extractHookSafetyProse(tmplContent);
// Transform frontmatter (host-aware)
let result = transformFrontmatter(content, host);
// Insert safety advisory at the top of the body (after frontmatter)
if (safetyProse) {
const bodyStart = result.indexOf('\n---') + 4;
result = result.slice(0, bodyStart) + '\n' + safetyProse + '\n' + result.slice(bodyStart);
}
// Config-driven path + tool rewrites (shared with processSectionTemplate so
// section cross-references get the same per-host treatment as SKILL.md).
result = applyHostRewrites(result, hostConfig);
// Config-driven: generate metadata (e.g., openai.yaml for Codex)
const metadata = hostConfig.generation.generateMetadata && !symlinkLoop ? {
outputPath: path.join(outputDir, 'agents', 'openai.yaml'),
content: generateOpenAIYaml(name, condenseOpenAIShortDescription(extractedDescription)),
} : undefined;
return { content: result, outputPath, symlinkLoop, metadata };
}
function processTemplate(tmplPath: string, host: Host, options: RenderOptions): { outputPath: string; content: string; symlinkLoop?: boolean; metadata?: { outputPath: string; content: string } } {
// Normalize to LF at the entry point. Templates may have CRLF on disk when
// checked out on Windows with core.autocrlf=true. Downstream regexes
// (processVoiceTriggers, transformFrontmatter) hardcode \n, so without
// normalization they silently no-op on CRLF — producing different output
// than CI (Linux, LF) and breaking the Skill Docs Freshness check.
// (catalogParts left the return type with the proactive-suggestions
// retirement — merge of the two v1.64 waves.)
const tmplContent = fs.readFileSync(tmplPath, 'utf-8').replace(/\r\n/g, '\n');
const relTmplPath = path.relative(ROOT, tmplPath);
let outputPath = tmplPath.replace(/\.tmpl$/, '');
// Determine skill directory relative to ROOT
const skillDir = path.relative(ROOT, path.dirname(tmplPath));
// --out-dir: mirror the skill tree into the out-dir instead of writing in
// place (external hosts compute their own output paths below).
if (host === 'claude') {
outputPath = path.join(options.outputRoot, skillDir, path.basename(tmplPath).replace(/\.tmpl$/, ''));
}
// Extract name/description: name drives external skill naming + setup symlinks
// (and TemplateContext.skillName via buildContext); description feeds external
// host metadata. When frontmatter name: differs from directory name (e.g.
// run-tests/ with name: test), the frontmatter name wins.
const { name: extractedName, description: extractedDescription } = extractNameAndDescription(tmplContent);
const currentHostConfig = getHostConfig(host);
const ctx = buildContext(tmplContent, tmplPath, host, options);
const skillName = ctx.skillName;
// Replace placeholders + assert none remain (shared path with section generation).
let content = resolvePlaceholders(tmplContent, ctx, currentHostConfig, relTmplPath, options);
// Preprocess voice triggers: fold into description, strip field from frontmatter.
// Must run BEFORE transformFrontmatter so all hosts see the updated description,
// and BEFORE extractedDescription is used by external host metadata.
content = processVoiceTriggers(content);
// Re-extract description AFTER voice trigger preprocessing so Codex openai.yaml
// metadata gets the updated description with voice triggers included.
const postProcessDescription = extractNameAndDescription(content).description;
// For Claude: strip sensitive: field (only Factory uses it)
// For external hosts: route output, transform frontmatter, rewrite paths
let symlinkLoop = false;
let metadata: { outputPath: string; content: string } | undefined;
if (host === 'claude') {
content = transformFrontmatter(content, host);
} else {
const result = processExternalHost(content, tmplContent, host, skillDir, postProcessDescription, ctx, options, extractedName || undefined);
content = result.content;
outputPath = result.outputPath;
symlinkLoop = result.symlinkLoop;
metadata = result.metadata;
}
// Prepend generated header (after frontmatter)
const header = GENERATED_HEADER.replace('{{SOURCE}}', path.basename(tmplPath));
const fmEnd = content.indexOf('---', content.indexOf('---') + 3);
if (fmEnd !== -1) {
const insertAt = content.indexOf('\n', fmEnd) + 1;
content = content.slice(0, insertAt) + header + content.slice(insertAt);
} else {
content = header + content;
}
// Catalog trim (Claude only — external hosts have their own frontmatter shapes)
if (host === 'claude' && options.catalogMode === 'trim') {
const trimmed = applyCatalogTrim(content, skillName);
if (trimmed) content = trimmed.content;
}
// --out-dir: repoint section-base paths to the out-dir (no-op otherwise).
if (host === 'claude') content = rewriteSectionBase(content, options.contentLinkRoot);
return { outputPath, content, symlinkLoop, metadata };
}
/**
* Generate one on-demand section file (`<skill>/sections/<name>.md.tmpl` →
* `<name>.md`). Sections are BODY FRAGMENTS — no frontmatter, no catalog trim,
* no voice triggers. They resolve placeholders through the SAME path as
* SKILL.md (resolvePlaceholders) using the PARENT skill's TemplateContext
* (so appliesTo gating + tier behave identically — a section's {{PREAMBLE}}-
* style resolver renders the same content it would in the parent, not empty).
*
* Output routing mirrors SKILL.md: Claude writes in-tree at
* `<skill>/sections/<name>.md`; external hosts write to
* `<hostSubdir>/skills/<externalName>/sections/<name>.md`. External hosts get
* applyHostRewrites so cross-references resolve per host.
*/
function processSectionTemplate(
sectionTmplPath: string,
skillDir: string,
host: Host,
options: RenderOptions,
): { outputPath: string; content: string } {
const tmplContent = fs.readFileSync(sectionTmplPath, 'utf-8');
const relTmplPath = path.relative(ROOT, sectionTmplPath);
const hostConfig = getHostConfig(host);
// Read the owning SKILL.md.tmpl so the section inherits the parent's name +
// tier + benefits-from (TemplateContext parity). Fall back to the dir name.
const parentTmplPath = path.join(ROOT, skillDir, 'SKILL.md.tmpl');
const parentContent = fs.existsSync(parentTmplPath) ? fs.readFileSync(parentTmplPath, 'utf-8') : '';
const parentName = (parentContent && extractNameAndDescription(parentContent).name) || skillDir;
const ctx = buildContext(parentContent || tmplContent, parentTmplPath, host, options, parentName);
// Resolve placeholders against the section body (shared guard catches stragglers).
let content = resolvePlaceholders(tmplContent, ctx, hostConfig, relTmplPath, options);
// External hosts: rewrite cross-reference paths/tools (no frontmatter to transform).
if (host !== 'claude') {
content = applyHostRewrites(content, hostConfig);
} else {
// --out-dir: a section may cross-reference another section by absolute path;
// repoint those to the out-dir too (no-op when --out-dir is unset).
content = rewriteSectionBase(content, options.contentLinkRoot);
}
// Plain generated header (no frontmatter to insert after).
content = GENERATED_HEADER.replace('{{SOURCE}}', path.basename(sectionTmplPath)) + content;
const fileName = path.basename(sectionTmplPath).replace(/\.tmpl$/, '');
let outputPath: string;
if (host === 'claude') {
outputPath = path.join(options.outputRoot, skillDir, 'sections', fileName);
} else {
const externalName = externalSkillName(skillDir, parentName);
outputPath = path.join(options.outputRoot, hostConfig.hostSubdir, 'skills', externalName, 'sections', fileName);
}
return { outputPath, content };
}
// ─── Main ───────────────────────────────────────────────────
/** Render each artifact once. Artifact writes go through emit(); stale external
* caches are pruned only after a successful host render. Dry runs use the same
* inventory as normal generation, including metadata and
* shared outputs. Options are per invocation so imports/concurrent runs cannot
* inherit another caller's model, detection, or output paths.
*
* templates + host settings -> render -> emit -> dry-run: compare only
* | -> normal: mkdir + write
* shared index/digest ---------------+ -> artifact inventory + diagnostics
* successful external host -----------------> normal only: prune retired caches
*/
export async function runGeneration(settings: GenerationOptions = {}): Promise<GenerationResult> {
const options: RenderOptions = {
outputRoot: path.resolve(settings.outputRoot ?? ROOT),
contentLinkRoot: settings.contentLinkRoot ?? null,
model: settings.model ?? null,
catalogMode: settings.catalogMode ?? 'trim',
explainLevel: settings.explainLevel ?? 'default',
gbrainDetected: loadGbrainOverride(settings.respectDetection ?? false),
};
const hosts = settings.host === 'all' ? ALL_HOST_NAMES as Host[] : [settings.host ?? 'claude'];
const log = settings.log ?? (() => {});
const artifacts: GeneratedArtifact[] = [];
const diagnostics: GenerationDiagnostic[] = [];
const templates = discoverTemplates(ROOT);
const sections = discoverSectionTemplates(ROOT);
const rel = (outputPath: string) => path.relative(options.outputRoot, outputPath).split(path.sep).join('/');
function emit(outputPath: string, content: string, kind: GeneratedArtifact['kind'], host?: Host): void {
const relativePath = rel(outputPath);
artifacts.push({ relativePath, kind, ...(host ? { host } : {}) });
try {
if (settings.dryRun) {
let existing: string | undefined;
try {
existing = fs.readFileSync(outputPath, 'utf-8');
} catch (error) {
if ((error as NodeJS.ErrnoException).code !== 'ENOENT') throw error;
// Windows reports ENOENT for a child of a regular file. Distinguish
// that filesystem error from a missing artifact without writing.
let parent = path.dirname(outputPath);
while (true) {
try {
if (!fs.statSync(parent).isDirectory()) {
throw Object.assign(new Error(`ENOTDIR: output ancestor is not a directory: ${parent}`, { cause: error }), { code: 'ENOTDIR' });
}
break;
} catch (ancestorError) {
if ((ancestorError as NodeJS.ErrnoException).code !== 'ENOENT') throw ancestorError;
const next = path.dirname(parent);
if (next === parent) throw error;
parent = next;
}
}
}
if (existing !== content) {
diagnostics.push({ kind: 'stale', relativePath, host, message: `STALE: ${relativePath}` });
log(`STALE: ${relativePath}`);
} else {
log(`FRESH: ${relativePath}`);
}
} else {
fs.mkdirSync(path.dirname(outputPath), { recursive: true });
fs.writeFileSync(outputPath, content);
log(`GENERATED: ${relativePath}`);
}
} catch (error) {
throw Object.assign(new Error(`${relativePath}: ${(error as Error).message}`), { relativePath });
}
}
function failed(error: unknown, host?: Host): void {
const message = error instanceof Error ? error.message : String(error);
diagnostics.push({ kind: 'error', host, message, relativePath: (error as { relativePath?: string })?.relativePath });
}
for (const host of hosts) {
try {
const hostConfig = getHostConfig(host);
const tokenBudget: Array<{ skill: string; lines: number; tokens: number }> = [];
const renderedNames = new Set<string>();
for (const template of templates) {
const skillDir = path.dirname(template.tmpl);
if (!includesSkill(hostConfig, skillDir)) continue;
const result = processTemplate(path.join(ROOT, template.tmpl), host, options);
const relativePath = rel(result.outputPath);
if (host !== 'claude') renderedNames.add(path.basename(path.dirname(result.outputPath)));
if (result.symlinkLoop) {
diagnostics.push({ kind: 'skipped', relativePath, host, message: `SKIPPED (symlink loop): ${relativePath}` });
log(`SKIPPED (symlink loop): ${relativePath}`);
continue;
}
emit(result.outputPath, result.content, 'skill', host);
if (result.metadata) emit(result.metadata.outputPath, result.metadata.content, 'metadata', host);
if (skillDir === 'qa') {
const report = fs.readFileSync(path.join(ROOT, 'qa', 'templates', 'functional-report-template.md'), 'utf-8');
emit(path.join(path.dirname(result.outputPath), 'templates', 'functional-report-template.md'),
(host === 'claude' ? '' : GENERATED_HEADER.replace('{{SOURCE}}', 'qa/templates/functional-report-template.md')) + report, 'asset', host);
}
tokenBudget.push({ skill: relativePath, lines: result.content.split('\n').length, tokens: Math.round(result.content.length / 4) });
const TOKEN_CEILING_BYTES = 160_000;
if (result.content.length > TOKEN_CEILING_BYTES) {
const message = `⚠️ TOKEN CEILING: ${relativePath} is ${result.content.length} bytes (~${Math.round(result.content.length / 4)} tokens), exceeds ${TOKEN_CEILING_BYTES} byte ceiling (~40K tokens)`;
diagnostics.push({ kind: 'warning', host, relativePath, message });
}
}
for (const section of sections) {
if (!includesSkill(hostConfig, section.skillDir) || !usesLazySections(host, section.skillDir)) continue;
const result = processSectionTemplate(path.join(ROOT, section.tmpl), section.skillDir, host, options);
emit(result.outputPath, result.content, 'section', host);
tokenBudget.push({ skill: rel(result.outputPath), lines: result.content.split('\n').length, tokens: Math.round(result.content.length / 4) });
}
// Claude owns these catalog-derived runtime assets. Use the same host
// inclusion rule and compare-or-write path as every other artifact.
if (host === 'claude' && includesSkill(hostConfig, 'review')) {
emit(path.join(options.outputRoot, 'review', 'design-checklist.md'),
generateDesignChecklistMd(), 'asset', host);
emit(path.join(options.outputRoot, DOM_DUMP_FILE), DOM_DUMP_SCRIPT + '\n', 'asset', host);
}
if (host === 'openclaw') {
for (const variant of ['lite', 'full', 'plan'] as const) {
const fileName = `gstack-${variant}-CLAUDE.md`;
emit(path.join(options.outputRoot, 'openclaw', fileName),
fs.readFileSync(path.join(ROOT, 'openclaw', 'templates', fileName), 'utf-8'), 'openclaw', host);
}
}
// A failed render exits this try before pruning: its inventory is partial.
// Only remove generated directories owned by this host; sidecars and user
// skills survive. Dry runs never create, rewrite, or remove any directory.
if (!settings.dryRun && host !== 'claude') {
const skillsRoot = path.join(options.outputRoot, hostConfig.hostSubdir, 'skills');
let entries: fs.Dirent[] = [];
try {
entries = fs.readdirSync(skillsRoot, { withFileTypes: true });
} catch (error) {
if ((error as NodeJS.ErrnoException).code !== 'ENOENT') throw error;
}
for (const entry of entries) {
if (entry.isSymbolicLink() || !entry.isDirectory() || !entry.name.startsWith('gstack-') || renderedNames.has(entry.name)) continue;
// Keep the old render usable until setup has migrated installed copies/links.
if (entry.name === 'gstack-claude' && process.env.GSTACK_DEFER_CLAUDE_RENAME_PRUNE === '1') continue;
let generated = false;
try {
generated = fs.readFileSync(path.join(skillsRoot, entry.name, 'SKILL.md'), 'utf-8').includes('<!-- AUTO-GENERATED from');
} catch (error) {
if ((error as NodeJS.ErrnoException).code !== 'ENOENT') throw error;
}
if (!generated) {
log(` kept ${host} skills/${entry.name}: not a gstack render (no generated banner)`);
continue;
}
fs.rmSync(path.join(skillsRoot, entry.name), { recursive: true, force: true });
log(` pruned stale ${host} render: ${entry.name}`);
if (entry.name === 'gstack-claude') {
log(' /claude is now /claude-code. Run ./setup to migrate installed skill links; generation only updates render files.');
}
}
}
if (!settings.dryRun && tokenBudget.length > 0) {
tokenBudget.sort((a, b) => b.lines - a.lines);
log(`\nToken Budget (${host} host)`);
log('═'.repeat(60));
for (const item of tokenBudget) {
const name = item.skill.replace(/\/SKILL\.md$/, '').replace(`${hostConfig.hostSubdir}/skills/`, '');
log(` ${name.padEnd(30)} ${String(item.lines).padStart(5)} lines ~${String(item.tokens).padStart(6)} tokens`);
}
log('─'.repeat(60));
log(` ${'TOTAL'.padEnd(30)} ${String(tokenBudget.reduce((sum, t) => sum + t.lines, 0)).padStart(5)} lines ~${String(tokenBudget.reduce((sum, t) => sum + t.tokens, 0)).padStart(6)} tokens\n`);
}
} catch (error) {
failed(error, host);
}
}
// Shared artifacts are awaited in both modes. A failure must reach the CLI
// exit code, never disappear in a fire-and-forget auxiliary writer.
try {
const index = await generateLlmsTxt();
emit(path.join(options.outputRoot, 'gstack', 'llms.txt'), index.content, 'index');
for (const warning of index.warnings) diagnostics.push({ kind: 'warning', message: `[gen-llms-txt] ${warning}` });
} catch (error) {
failed(error);
}
try {
const digest = generateAgentsDigest();
emit(path.join(options.outputRoot, DIGEST_RELPATH), digest.content, 'digest');
if (!settings.dryRun) log(`[gen-agents-digest] ${DIGEST_RELPATH}: ${digest.bytes} bytes (budget ${DIGEST_BYTE_BUDGET})`);
} catch (error) {
failed(error);
}
return { exitCode: diagnostics.some(d => d.kind === 'error' || d.kind === 'stale') ? 1 : 0, artifacts, diagnostics };
}
/** Importing this module never executes generation or reads CLI/user settings.
* Async main is require()-compatible: only top-level await would break callers. */
export async function main(args = process.argv.slice(2)): Promise<number> {
try {
const settings = parseGenerationArgs(args);
const result = await runGeneration({ ...settings, log: console.log });
for (const diagnostic of result.diagnostics) {
if (diagnostic.kind === 'error') console.error(`ERROR${diagnostic.host ? ` (${diagnostic.host})` : ''}: ${diagnostic.message}`);
if (diagnostic.kind === 'warning') console.error(diagnostic.message);
}
if (result.diagnostics.some(d => d.kind === 'stale')) {
console.error(`\nGenerated files are stale. Run: bun run gen:skill-docs --host ${settings.host ?? 'claude'}`);
}
if (!settings.dryRun) {
try {
const config = fs.readFileSync(path.join(resolveStateRoot(), 'config.yaml'), 'utf-8');
if (/^skill_prefix:\s*true/m.test(config)) {
console.log('\nNote: skill_prefix is true. Run gstack-relink to re-apply name: patches (it patches both the install and any active gbrain render).');
}
} catch { /* optional local install note */ }
}
return result.exitCode;
} catch (error) {
console.error(`ERROR: ${(error as Error).message}`);
return 1;
}
}
if (import.meta.main) {
void main().then(code => { process.exitCode = code; });
}