mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-29 09:20:39 +02:00
v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders
interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate SKILL.md — dead frontmatter keys removed
Mechanical regen after hosts/claude.ts stripFields change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers
New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight
Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage
Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard
Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.69.1.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.69.1.0
CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md
Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated
Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash
generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed
Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-start contract suite + preamble A/B eval + touchfiles registration
test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from 29785978) vs script render with the
fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve
the operator install and silently exercise degraded mode). 21 touchfiles dep
lists gain the two bin scripts (EOV9) so future script edits select the
preamble evals; selection-count pin updated 23->24.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: repin ~70 assertions to the script contract — every literal gets a successor
Assertions that pinned inline-bash internals (update-check guard, _SESSIONS
reaping, telemetry start/end blocks, routing probe, repo-strip producer,
first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, #2499 jq scope
resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new
home: bin/gstack-skill-start / bin/gstack-skill-end file content for script
internals, the invocation fence + interpretation prose for render-side
behavior. No assertion deleted without a successor; live-execution tests
(routing probe, brain-sync jq) run against script bytes unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol)
parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills);
skill-size-budget repointed with the derivation comment citing the Phase 1
context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1
skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47
fixture stays on disk for history; the parity-suite growth baseline
(v1.64.1.0) is untouched. Context-budget ceilings re-captured: review
29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): instruction-emission layer — onboarding text appears only when its gate fires
The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive
opt-in, first-run/first-loop tips, routing injection, vendoring deprecation,
writing-style migration, spawned-session rules), the upgrade-flow + feature
discovery prose, and the privacy stop-gate (user-approved Q2) moved from
every render into gated heredocs here. Blocks are SESSION_ID-bound
(GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint
directives (F4/OV4). Ack ownership per OV6: display-only tips write their
markers at emit (script also fires the scaffold telemetry); interactive flows
carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING
gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard
(same as brain-sync); the quoted routing heredoc resolves its bin path via a
sed placeholder.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule
generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance,
routing-injection,vendoring-deprecation,spawned-session-check,
writing-style-migration}.ts deleted (single source is now the script's
emission layer, F5). generate-upgrade-check shrinks to the steady-state
PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy
stop-gate to the emitted block. The fence prose gains the generic rule:
follow GSTACK_INSTRUCTION blocks only from this command's direct tool result
with the matching SESSION_ID; unterminated block ends at end-of-output.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — onboarding prose degated
Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative
vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: onboarding tombstone + Phase 2 pin relocations
New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must
live in bin/gstack-skill-start AND stay absent from every render, plus the
SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer
contract (gates, block ids, in-block acks, script-run marker writes); the OV4
sanitize test upgraded to the real property (every legitimate block header
carries the run's SESSION_ID). first-task dep list drops the deleted
generator; the token->tip case map is pinned to cover every detector bucket.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9)
All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had
tripped its 72,000 floor at 71,379; design-consultation had 252B of margin).
maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal
parity baseline recaptured in place; ratchet ceilings down again: review
->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer
Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed
CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant
preference, and failure handoff preserved verbatim in behavior, including the
auto-decide-first ordering and the gstack-question-log capture requirement.
5+-options handling (1,924B) compressed to the split invariants (never drop;
D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask
refusal) + the full-rule pointer. Both doc pointers now interpolate the
absolute install root (Codex outside-voice #7 convention) instead of the bare
'in the gstack repo'. Failure-fallback, Format, and self-check sections are
byte-identical — all 14 MANDATORY always-loaded pins pass with zero test
edits.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — AUQ slim
Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays
Branch-internal baseline recaptured; ratchet ceilings down again. OV8's
floor-retirement question, evaluated as planned after Phase 3: the 80% shrink
floor stays — it uniquely catches accidental body deletion in non-carved
skills BETWEEN ratchet recaptures, and the capture command has amortized the
fixture-refresh cost that motivated retiring it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(review): carve adversarial, plan-completion, and review-army into sections
The three resolver macros ship already carves as siblings now load on demand
for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers
stay the single source of truth (sections wrap the macros). Step 0/1, scope
drift, critical pass, confidence calibration, and fix-first stay always-loaded.
Fixtures and pins follow the moved content (codex-hardening wrapped-sites,
review-army E2E fixture builds skeleton+sections with an empty-fixture guard).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(codex): carve the three mutually exclusive modes into sections
Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs)
load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode
dispatch, filesystem boundary, and a new always-loaded 'Synthesis
recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ
per-skill pins pass unchanged); the plan-file report + exit gate render after
the last section pointer per the gateAfterStop pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections
The once-per-repo dry-run validation, the pre-merge readiness gate, and the
merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB ->
55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run
section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained.
Zero content lost (line-coverage checked against HEAD).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5)
They never consume the tier-3 sections (repo-mode ownership, search-before-
building) but do fire AskUserQuestion, which tier >=2 provides — verified by
grep before the plan review. -2.2KB per skill. Render assertions pin the
demotion (tier-3 sections absent, AUQ format present).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow
CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total);
their MONOLITH_INVARIANTS entries retire (invariants now generate from the
registry, cso precedent). Touchfiles: carve-section-loading covers the three
new carves; the codex + land-and-deploy LLM-judge dep lists widen to their
sections. Regen + goldens + branch-internal baseline + ratchet ceilings
recaptured (review 24,052 -> skeleton-based ceiling; union floors hold).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(gen-skill-docs): review render pins read the carved union
The review carve's readSkillUnion conversions (same pattern its neighbor
carved-skill pins already use).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(autoplan): carve the four review phases + tasks aggregator into sections
Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks
aggregator load on demand; Design and DX stay separate sections because each
is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30%
always-loaded); the 6 decision principles, classification, sequencing, and
explicit skip-condition dispatch stay always-loaded. The chain E2E's
phase-complete markers now live only in sections, so its assertions double as
section-read proof (behavioral: external).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(spec): carve the post-confirmation gate-and-file tail into one section
Phases 1-4 are the turn-1 conversational spine — carving them would force the
Read on the first user message for zero real savings. The mechanical tail
(4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only
after draft confirmation: a genuine lazy boundary, kept as ONE section so the
gh-issue-create bash can never load without the fail-closed redaction gate
that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure
invariants migrated location-aware plus a new carve-shape suite (56 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(setup-gbrain): carve the branch-exclusive install paths into sections
Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and
CLAUDE.md persist load on demand — at most one install route ever runs.
Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay
always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es
extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate
counts scan the tmpl union (the third init site lives in engine-remediation).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow
CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec,
and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles:
setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the
stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch
baseline + ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format
RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the
skill prose), degraded path documented in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines
CARVE_GUARDS gains the three entries; qa's monolith invariant retires.
auq-format carve-safety now keys on the skeleton+sections union shipping
the AUQ block (first tier-1 carve: browse never renders it by design).
Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move)
Sol scope discipline stays pinned via the model overlay + completeness
section; the lake intro is now a single script-emitted blurb.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton)
A session runs exactly one mode, so a builder session never loads the
13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule
stay in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections
design-html 57→49KB, design-shotgun 53→50KB. Sections wrap
{{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of
truth; the pretext-patterns STOP sits at the top of Step 3 so the read
provably precedes the Write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills
Both design entries carry requiredReads + loading-eval scenarios (D3A
condition). office-hours phase sections are mode-exclusive, so only the
always-reached design/handoff section is a deterministic requiredRead.
Baselines and ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline
Moved: browser/sidebar/server internals, CHANGELOG release-summary format
spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw
publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction
scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG
entry rules), the machine-managed GBrain block (byte-identical), and the
'## Deploying to the active skill' header with gbrain-refresh in range
(pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): seed onboarding markers into the hermetic child GSTACK_HOME
EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME
seeding in e2e-helpers.ts no longer reaches hermetic children — the
emission layer fired lake-intro/telemetry prompts that burned turns and
stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed
by the telemetry question). Onboarding-specific tests pin their own
GSTACK_HOME per-test, which merges over this seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: raise carve-section-loading wall clock to 480s SDK / 540s bun
The heavy full-workflow scenarios satisfy their required section reads
inside 60s but need 300-450s to finish the report on slower sandboxes;
the 300s default read as a loading failure when the carve invariant held
(traces: plan-eng-review read its section at 8s, office-hours all three
at 24s, design-html both at 50s — all timed out mid-report).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): harden the skill-start trust boundary — review-army findings
Session ID gains a urandom suffix (block binding unforgeable by reflected
content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch
names are charset-clamped before JSON embedding (skill-start + skill-end);
.brain-last-push reads first line only with a charset clamp; the artifacts
URL echo routes through _sanitize; the privacy consent gate fires in
interactive sessions only (spawned auto-choose could accept consent no
human gave — emission order is not a safety property); the daily pull gets
non-interactive + slow-network git guards and stamps only when the
receipted path ran; ~/.claude.json gets a grep pre-filter before the jq
parse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep
The question-log block bound $_SESSION_ID, a shell variable the
consolidated fence never sets — hook-less hosts logged empty session_id,
breaking /plan-tune per-session grouping. It now uses the same
substitute-from-the-skill-start-echoes contract as the telemetry block.
Also: retired the pre-Phase-2 stop-gate docstring, repointed the
gbrain-local-status cross-reference at the script's inline jq, dropped an
orphaned section comment, documented retro-metrics' suffix-only census.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: regenerate renders for the question-log placeholder; goldens + baselines follow
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: hermetic update-check, onboarding gate sequencing, seeding parity
The contract test's child did a live git ls-remote + curl to github.com on
every bun run test (update_check config now gates it off); the headless
test gets a fresh GSTACK_HOME so the suppression is actually exercised; a
new OV6 test drives the script three times to pin ack-at-emit and gate
sequencing; hermetic seeding covers the config-keyed privacy gate; the
EVALS_HERMETIC=0 debug seeding reaches marker parity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix
Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git
shows a pre-Phase-1 sha), which only the periodic workflow provides — and
a static matrix entry so it can't silently never run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.70.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.70.0.0
ARCHITECTURE.md: the preamble section now describes the v1.70 runtime —
the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads
STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding
text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in
every render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim
docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when
the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md.
ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session
count and an "ELI16 mode" that no shipped code implements (the count
computation was deleted with the inline preamble); describe the real
touch-and-prune behavior instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): correct numeric claims against measured counts
50 of 62 installed skills dropped (fixture/alias entries have no preamble);
11 new carves + a deeper office-hours carve = 9→20; test counts match the
files (13 / 11 / 3 / 7).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression
Both attempts in run 33090283032 produced judge-praised DESIGN.md files
phrased as 'design principles'/'design language' without any of the four
original literals; inputs were identical to the prior passing run
32899975845 (design-consultation untouched by the intervening merge).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): stage design-consultation's sections/ into the E2E fixture
The skill has been carved since v1.57.0.0 — the DESIGN.md structure
prescription (the AESTHETIC proposal template) lives in
sections/proposal-and-preview.md behind a STOP-read. The fixture only
copied SKILL.md, so the agent improvised structure from the skeleton and
the section-synonym check has been a coin flip since the carve (CI run
33090283032 trace shows 'no sections dir'; the local eval store has the
same failure on 2026-08-25 while that day's CI run passed on lucky
vocabulary).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
a3749bfa4b
commit
394db326f2
+227
-106
@@ -28,6 +28,15 @@ function readShipUnion(): string {
|
||||
return readSkillUnion('ship');
|
||||
}
|
||||
|
||||
// Token-reduction Phase 1: the preamble's inline bash (session bookkeeping,
|
||||
// config echoes, telemetry producers, artifacts sync) moved into
|
||||
// bin/gstack-skill-start / bin/gstack-skill-end. The render carries a one-line
|
||||
// invocation fence + interpretation prose. Assertions that pinned inline-bash
|
||||
// internals now pin the scripts (the new home); render-side assertions pin the
|
||||
// fence + prose. Script behavior is pinned by test/gstack-skill-start.test.ts.
|
||||
const SKILL_START_SCRIPT = fs.readFileSync(path.join(ROOT, 'bin', 'gstack-skill-start'), 'utf-8');
|
||||
const SKILL_END_SCRIPT = fs.readFileSync(path.join(ROOT, 'bin', 'gstack-skill-end'), 'utf-8');
|
||||
|
||||
function extractDescription(content: string): string {
|
||||
const fmEnd = content.indexOf('\n---', 4);
|
||||
expect(fmEnd).toBeGreaterThan(0);
|
||||
@@ -117,8 +126,11 @@ const CLAUDE_SKIPPED = new Set(__getHostConfig('claude').generation.skipSkills ?
|
||||
const CLAUDE_GENERATED_SKILLS = ALL_SKILLS.filter(s => !CLAUDE_SKIPPED.has(s.dir));
|
||||
|
||||
describe('gen-skill-docs', () => {
|
||||
// Browse carve (token-reduction Phase 4): the command reference + snapshot
|
||||
// flags render into browse/sections/command-list.md now — read the
|
||||
// skeleton+sections union so these pins hold across the carve.
|
||||
test('generated SKILL.md contains all command categories', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md'), 'utf-8');
|
||||
const content = readSkillUnion('browse');
|
||||
const categories = new Set(Object.values(COMMAND_DESCRIPTIONS).map(d => d.category));
|
||||
for (const cat of categories) {
|
||||
expect(content).toContain(`### ${cat}`);
|
||||
@@ -126,7 +138,7 @@ describe('gen-skill-docs', () => {
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains all commands', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md'), 'utf-8');
|
||||
const content = readSkillUnion('browse');
|
||||
for (const [cmd, meta] of Object.entries(COMMAND_DESCRIPTIONS)) {
|
||||
const display = meta.usage || cmd;
|
||||
expect(content).toContain(display);
|
||||
@@ -134,7 +146,7 @@ describe('gen-skill-docs', () => {
|
||||
});
|
||||
|
||||
test('command table is sorted alphabetically within categories', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md'), 'utf-8');
|
||||
const content = readSkillUnion('browse');
|
||||
// Extract command names from the Navigation section as a test
|
||||
const navSection = content.match(/### Navigation\n\|.*\n\|.*\n([\s\S]*?)(?=\n###|\n## )/);
|
||||
expect(navSection).not.toBeNull();
|
||||
@@ -159,7 +171,7 @@ describe('gen-skill-docs', () => {
|
||||
});
|
||||
|
||||
test('snapshot flags section contains all flags', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md'), 'utf-8');
|
||||
const content = readSkillUnion('browse');
|
||||
for (const flag of SNAPSHOT_FLAGS) {
|
||||
expect(content).toContain(flag.short);
|
||||
expect(content).toContain(flag.description);
|
||||
@@ -302,10 +314,19 @@ describe('gen-skill-docs', () => {
|
||||
expect(rootTmpl).not.toContain('{{COMMAND_REFERENCE}}');
|
||||
expect(rootTmpl).not.toContain('{{SNAPSHOT_FLAGS}}');
|
||||
|
||||
// Browse carve: the reference resolvers moved into the on-demand section
|
||||
// template (so gen-skill-docs keeps them fresh from browse/src); the
|
||||
// skeleton points at the section instead of inlining the reference.
|
||||
const browseTmpl = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md.tmpl'), 'utf-8');
|
||||
expect(browseTmpl).toContain('{{COMMAND_REFERENCE}}');
|
||||
expect(browseTmpl).toContain('{{SNAPSHOT_FLAGS}}');
|
||||
expect(browseTmpl).not.toContain('{{COMMAND_REFERENCE}}');
|
||||
expect(browseTmpl).not.toContain('{{SNAPSHOT_FLAGS}}');
|
||||
expect(browseTmpl).toContain('{{SECTION:command-list}}');
|
||||
expect(browseTmpl).toContain('{{PREAMBLE}}');
|
||||
|
||||
const browseSectionTmpl = fs.readFileSync(
|
||||
path.join(ROOT, 'browse', 'sections', 'command-list.md.tmpl'), 'utf-8');
|
||||
expect(browseSectionTmpl).toContain('{{COMMAND_REFERENCE}}');
|
||||
expect(browseSectionTmpl).toContain('{{SNAPSHOT_FLAGS}}');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains operational self-improvement (replaced contributor mode)', () => {
|
||||
@@ -315,7 +336,9 @@ describe('gen-skill-docs', () => {
|
||||
expect(content).not.toContain('contributor-logs');
|
||||
expect(content).toContain('Operational Self-Improvement');
|
||||
expect(content).toContain('gstack-learnings-log');
|
||||
expect(content).toContain('gstack-learnings-search --limit 3');
|
||||
// The learnings-resurface call moved from the inline preamble bash into
|
||||
// the skill-start script (Phase 1) — same command, new home.
|
||||
expect(SKILL_START_SCRIPT).toContain('gstack-learnings-search" --limit 3');
|
||||
});
|
||||
|
||||
test('generated SKILL.md with LEARNINGS_LOG contains operational type', () => {
|
||||
@@ -324,41 +347,43 @@ describe('gen-skill-docs', () => {
|
||||
expect(content).toContain('operational');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains session awareness', () => {
|
||||
test('session awareness lives in gstack-skill-start (registry touch + stale cleanup)', () => {
|
||||
// The sessions registry moved from inline preamble bash into the script:
|
||||
// it records the harness pid (--parent-pid identity) and expires entries
|
||||
// older than 120 minutes.
|
||||
expect(SKILL_START_SCRIPT).toContain('sessions/$PARENT_PID');
|
||||
expect(SKILL_START_SCRIPT).toContain('-mmin +120');
|
||||
// The render keeps the completion-status protocol the sessions feed into.
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('_SESSIONS');
|
||||
expect(content).toContain('RECOMMENDATION');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains branch detection', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('_BRANCH');
|
||||
expect(content).toContain('git branch --show-current');
|
||||
test('branch detection lives in gstack-skill-start and is echoed as BRANCH', () => {
|
||||
expect(SKILL_START_SCRIPT).toContain('_BRANCH=$(git branch --show-current');
|
||||
expect(SKILL_START_SCRIPT).toContain('echo "BRANCH: $_BRANCH"');
|
||||
});
|
||||
|
||||
// #2001: update_check: false silences the binary but the upgrade-handling
|
||||
// instruction prose used to ship unconditionally. Every skill that carries
|
||||
// the runtime config-echo cluster must (a) echo UPDATE_CHECK so the
|
||||
// instruction layer can read it, and (b) gate the UPGRADE_AVAILABLE /
|
||||
// JUST_UPGRADED prose on it — the same echo-then-gate convention every other
|
||||
// flag (PROACTIVE, SKILL_PREFIX, EXPLAIN_LEVEL, QUESTION_TUNING) follows.
|
||||
test('update_check opt-out gates preamble echo and upgrade-handling prose (issue #2001)', () => {
|
||||
let checked = 0;
|
||||
for (const skill of CLAUDE_GENERATED_SKILLS) {
|
||||
const content = fs.readFileSync(path.join(ROOT, skill.dir, 'SKILL.md'), 'utf-8');
|
||||
// Scope: only skills that render the runtime config-echo cluster.
|
||||
if (!content.includes('echo "QUESTION_TUNING: $_QUESTION_TUNING"')) continue;
|
||||
checked++;
|
||||
expect(content, `${skill.dir} must echo UPDATE_CHECK`).toContain('echo "UPDATE_CHECK: $_UPDATE_CHECK"');
|
||||
expect(content, `${skill.dir} must read update_check config`).toContain('_UPDATE_CHECK=$(');
|
||||
// Whenever the upgrade-handling prose ships, it must gate on the flag.
|
||||
if (content.includes('UPGRADE_AVAILABLE <old> <new>')) {
|
||||
expect(content, `${skill.dir} upgrade prose must gate on UPDATE_CHECK`)
|
||||
.toContain('If `UPDATE_CHECK` is `"false"`');
|
||||
}
|
||||
}
|
||||
// Guard against the scope filter silently matching nothing.
|
||||
expect(checked).toBeGreaterThan(0);
|
||||
// instruction prose used to ship unconditionally. Token-reduction Phase 2
|
||||
// made the gate STRUCTURAL: the prose left the renders entirely (absence is
|
||||
// pinned by test/onboarding-moved-literals.test.ts) and now emits from
|
||||
// gstack-skill-start's instruction layer ONLY when the update-check binary
|
||||
// produced output — and that binary silences itself on update_check=false.
|
||||
// Opted-out installs can never see the prose, by construction.
|
||||
test('update_check opt-out gates the update binary and upgrade-flow emission (issue #2001)', () => {
|
||||
// The config-echo cluster lives in gstack-skill-start: the flag is still
|
||||
// read and echoed as a STATUS line for the model.
|
||||
expect(SKILL_START_SCRIPT, 'script must read update_check config').toContain('_UPDATE_CHECK=$(');
|
||||
expect(SKILL_START_SCRIPT, 'script must echo UPDATE_CHECK').toContain('echo "UPDATE_CHECK: $_UPDATE_CHECK"');
|
||||
// Gate half 1: the update-check binary exits silently when opted out.
|
||||
const updateCheck = fs.readFileSync(path.join(ROOT, 'bin', 'gstack-update-check'), 'utf-8');
|
||||
expect(updateCheck, 'binary must read update_check config').toContain('get update_check');
|
||||
expect(updateCheck, 'binary must exit silently on update_check=false')
|
||||
.toMatch(/if \[ "\$_UC" = "false" \]; then\n\s*exit 0/);
|
||||
// Gate half 2: the upgrade-flow instruction block emits only when the
|
||||
// binary emitted something (empty when opted out, cached, or up to date).
|
||||
expect(SKILL_START_SCRIPT, 'upgrade-flow must be gated on update-check output')
|
||||
.toMatch(/if \[ -n "\$_UPD" \]; then\n\s*_emit_block upgrade-flow/);
|
||||
});
|
||||
|
||||
test('tier 2+ skills contain ELI10 simplification rules (AskUserQuestion format)', () => {
|
||||
@@ -377,9 +402,12 @@ describe('gen-skill-docs', () => {
|
||||
expect(content).not.toContain('## Completeness Principle');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains telemetry line', () => {
|
||||
test('telemetry producer lives in the scripts; render documents the analytics sink', () => {
|
||||
// The skill-usage.jsonl producers moved into the scripts (Phase 1).
|
||||
expect(SKILL_START_SCRIPT).toContain('analytics/skill-usage.jsonl');
|
||||
expect(SKILL_END_SCRIPT).toContain('analytics/skill-usage.jsonl');
|
||||
// The render still tells the model where telemetry lands.
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('skill-usage.jsonl');
|
||||
expect(content).toContain('~/.gstack/analytics');
|
||||
});
|
||||
|
||||
@@ -499,20 +527,31 @@ describe('gen-skill-docs', () => {
|
||||
];
|
||||
for (const skill of PREAMBLE_SKILLS) {
|
||||
const content = fs.readFileSync(path.join(ROOT, skill.dir, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain(`"skill":"${skill.name}"`);
|
||||
// The skill name now travels as --skill into gstack-skill-start (the
|
||||
// preamble fence) and gstack-skill-end (the telemetry epilogue) — the
|
||||
// scripts write it into the JSONL events.
|
||||
expect(content, `${skill.dir} preamble fence must pass its own name`)
|
||||
.toMatch(new RegExp(`--skill "${skill.name}" --model`));
|
||||
expect(content, `${skill.dir} epilogue must pass its own name`)
|
||||
.toContain(`gstack-skill-end --skill "${skill.name}"`);
|
||||
}
|
||||
});
|
||||
|
||||
test('qa and qa-only templates use QA_METHODOLOGY placeholder', () => {
|
||||
const qaTmpl = fs.readFileSync(path.join(ROOT, 'qa', 'SKILL.md.tmpl'), 'utf-8');
|
||||
expect(qaTmpl).toContain('{{QA_METHODOLOGY}}');
|
||||
// qa carve: the macro moved into the section template (the skeleton
|
||||
// carries the STOP-Read pointer); qa-only remains an inline monolith.
|
||||
const qaSkeletonTmpl = fs.readFileSync(path.join(ROOT, 'qa', 'SKILL.md.tmpl'), 'utf-8');
|
||||
expect(qaSkeletonTmpl).toContain('{{SECTION:qa-patterns}}');
|
||||
expect(qaSkeletonTmpl).not.toContain('{{QA_METHODOLOGY}}');
|
||||
const qaSectionTmpl = fs.readFileSync(path.join(ROOT, 'qa', 'sections', 'qa-patterns.md.tmpl'), 'utf-8');
|
||||
expect(qaSectionTmpl).toContain('{{QA_METHODOLOGY}}');
|
||||
|
||||
const qaOnlyTmpl = fs.readFileSync(path.join(ROOT, 'qa-only', 'SKILL.md.tmpl'), 'utf-8');
|
||||
expect(qaOnlyTmpl).toContain('{{QA_METHODOLOGY}}');
|
||||
});
|
||||
|
||||
test('QA_METHODOLOGY appears expanded in both qa and qa-only generated files', () => {
|
||||
const qaContent = fs.readFileSync(path.join(ROOT, 'qa', 'SKILL.md'), 'utf-8');
|
||||
const qaContent = readSkillUnion('qa'); // carved: methodology lives in qa/sections/qa-patterns.md
|
||||
const qaOnlyContent = fs.readFileSync(path.join(ROOT, 'qa-only', 'SKILL.md'), 'utf-8');
|
||||
|
||||
// Both should contain the health score rubric
|
||||
@@ -629,8 +668,9 @@ describe('GitLab support in generated skills', () => {
|
||||
*/
|
||||
describe('description quality evals', () => {
|
||||
// Regression: snapshot flags lost value hints (-d <N>, -s <sel>, -o <path>)
|
||||
// Browse carve: the flag reference renders into browse/sections/command-list.md.
|
||||
test('snapshot flags with values include value hints in output', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'browse', 'SKILL.md'), 'utf-8');
|
||||
const content = readSkillUnion('browse');
|
||||
for (const flag of SNAPSHOT_FLAGS) {
|
||||
if (flag.takesValue) {
|
||||
expect(flag.valueHint).toBeDefined();
|
||||
@@ -773,7 +813,11 @@ describe('REVIEW_DASHBOARD resolver', () => {
|
||||
}
|
||||
|
||||
test('plan-ceo-review chaining mentions eng and design reviews', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'plan-ceo-review', 'SKILL.md'), 'utf-8');
|
||||
// Carved skill: the chaining prose lives in sections/*.md. (It used to
|
||||
// pass against the skeleton only because the preamble's routing-injection
|
||||
// rules incidentally named these skills — that prose moved into
|
||||
// bin/gstack-skill-start in token-reduction Phase 2.)
|
||||
const content = readSkillUnion('plan-ceo-review');
|
||||
expect(content).toContain('/plan-eng-review');
|
||||
expect(content).toContain('/plan-design-review');
|
||||
});
|
||||
@@ -803,7 +847,7 @@ describe('REVIEW_DASHBOARD resolver', () => {
|
||||
describe('TEST_COVERAGE_AUDIT placeholders', () => {
|
||||
const planSkill = readSkillUnion('plan-eng-review'); // carved
|
||||
const shipSkill = readShipUnion();
|
||||
const reviewSkill = fs.readFileSync(path.join(ROOT, 'review', 'SKILL.md'), 'utf-8');
|
||||
const reviewSkill = readSkillUnion('review'); // carved: Review Army moved to sections/review-army.md
|
||||
|
||||
test('plan and ship modes share codepath tracing methodology', () => {
|
||||
// Review mode delegates test coverage to the Testing specialist subagent (Review Army)
|
||||
@@ -1024,7 +1068,7 @@ describe('PLAN_FILE_REVIEW_REPORT resolver', () => {
|
||||
|
||||
describe('PLAN_COMPLETION_AUDIT placeholders', () => {
|
||||
const shipSkill = readShipUnion();
|
||||
const reviewSkill = fs.readFileSync(path.join(ROOT, 'review', 'SKILL.md'), 'utf-8');
|
||||
const reviewSkill = readSkillUnion('review'); // carved: plan-completion audit moved to sections/plan-completion.md
|
||||
|
||||
test('ship SKILL.md contains plan completion audit step', () => {
|
||||
expect(shipSkill).toContain('Plan Completion Audit');
|
||||
@@ -1107,7 +1151,7 @@ describe('PLAN_VERIFICATION_EXEC placeholder', () => {
|
||||
|
||||
describe('Coverage gate in ship', () => {
|
||||
const shipSkill = readShipUnion();
|
||||
const reviewSkill = fs.readFileSync(path.join(ROOT, 'review', 'SKILL.md'), 'utf-8');
|
||||
const reviewSkill = readSkillUnion('review'); // carved: testing.md specialist ref lives in sections/review-army.md
|
||||
|
||||
test('ship SKILL.md contains coverage gate with thresholds', () => {
|
||||
expect(shipSkill).toContain('Coverage gate');
|
||||
@@ -1152,7 +1196,7 @@ describe('Plan file discovery shared helper', () => {
|
||||
// The shared helper should appear in ship (via PLAN_COMPLETION_AUDIT_SHIP)
|
||||
// and in review (via PLAN_COMPLETION_AUDIT_REVIEW)
|
||||
const shipSkill = readShipUnion();
|
||||
const reviewSkill = fs.readFileSync(path.join(ROOT, 'review', 'SKILL.md'), 'utf-8');
|
||||
const reviewSkill = readSkillUnion('review'); // carved: plan-completion audit moved to sections/plan-completion.md
|
||||
|
||||
test('plan file discovery appears in both ship and review', () => {
|
||||
expect(shipSkill).toContain('Plan File Discovery');
|
||||
@@ -1173,7 +1217,9 @@ describe('Plan file discovery shared helper', () => {
|
||||
// --- Retro plan completion ---
|
||||
|
||||
describe('Retro plan completion section', () => {
|
||||
const retroSkill = fs.readFileSync(path.join(ROOT, 'retro', 'SKILL.md'), 'utf-8');
|
||||
// Carved: the narrative report format (incl. Plan Completion) lives in
|
||||
// retro/sections/report-format.md — read the skeleton+sections union.
|
||||
const retroSkill = readSkillUnion('retro');
|
||||
|
||||
test('retro SKILL.md contains plan completion section', () => {
|
||||
expect(retroSkill).toContain('### Plan Completion');
|
||||
@@ -1384,8 +1430,9 @@ describe('Codex filesystem boundary', () => {
|
||||
});
|
||||
|
||||
test('review.ts CODEX_BOUNDARY constant is interpolated into resolver output', () => {
|
||||
// The adversarial step resolver should include boundary text in codex exec prompts
|
||||
const reviewContent = fs.readFileSync(path.join(ROOT, 'review', 'SKILL.md'), 'utf-8');
|
||||
// The adversarial step resolver should include boundary text in codex exec
|
||||
// prompts. Carved: the adversarial step lives in sections/adversarial.md.
|
||||
const reviewContent = readSkillUnion('review');
|
||||
// Boundary should appear near codex exec invocations
|
||||
const boundaryIdx = reviewContent.indexOf(BOUNDARY_MARKER);
|
||||
const codexExecIdx = reviewContent.indexOf('codex exec');
|
||||
@@ -1560,55 +1607,66 @@ describe('parameterized resolver support', () => {
|
||||
|
||||
// --- Preamble routing injection tests ---
|
||||
|
||||
describe('preamble routing injection', () => {
|
||||
const shipContent = readShipUnion();
|
||||
describe('preamble routing injection (bin/gstack-skill-start emission layer)', () => {
|
||||
// Token-reduction Phase 2: the routing-injection prose left the rendered
|
||||
// preamble entirely — bin/gstack-skill-start probes, gates, and emits the
|
||||
// whole flow as a GSTACK_INSTRUCTION block (with the AUQ, the routing rules
|
||||
// to append, and the decline ack all INSIDE the block). Absence from the
|
||||
// renders is pinned by test/onboarding-moved-literals.test.ts (tombstone);
|
||||
// this suite pins the gate structure and the emitted block's content.
|
||||
const routingBlock = (() => {
|
||||
const start = SKILL_START_SCRIPT.indexOf('_emit_block routing-injection');
|
||||
expect(start).toBeGreaterThan(0);
|
||||
return SKILL_START_SCRIPT.slice(start, SKILL_START_SCRIPT.indexOf('\nEOI', start));
|
||||
})();
|
||||
|
||||
test('preamble bash checks for routing section in CLAUDE.md and AGENTS.md', () => {
|
||||
test('routing probe checks CLAUDE.md and AGENTS.md (now in gstack-skill-start)', () => {
|
||||
// #2500: the probe iterates CLAUDE.md AND AGENTS.md — non-Claude hosts
|
||||
// route skills via AGENTS.md, the cross-harness convention file.
|
||||
expect(shipContent).toContain('for _RF in CLAUDE.md AGENTS.md');
|
||||
expect(shipContent).toContain('grep -q "## Skill routing" "$_RF"');
|
||||
expect(shipContent).toContain('HAS_ROUTING');
|
||||
expect(SKILL_START_SCRIPT).toContain('for _RF in CLAUDE.md AGENTS.md');
|
||||
expect(SKILL_START_SCRIPT).toContain('grep -q "## Skill routing" "$_RF"');
|
||||
expect(SKILL_START_SCRIPT).toContain('echo "HAS_ROUTING: $_HAS_ROUTING"');
|
||||
});
|
||||
|
||||
test('preamble bash reads routing_declined config', () => {
|
||||
expect(shipContent).toContain('routing_declined');
|
||||
expect(shipContent).toContain('ROUTING_DECLINED');
|
||||
test('script reads and echoes routing_declined config', () => {
|
||||
expect(SKILL_START_SCRIPT).toMatch(/_ROUTING_DECLINED=\$\("\$_BIN\/gstack-config" get routing_declined/);
|
||||
expect(SKILL_START_SCRIPT).toContain('echo "ROUTING_DECLINED: $_ROUTING_DECLINED"');
|
||||
});
|
||||
|
||||
test('preamble includes routing injection AskUserQuestion', () => {
|
||||
expect(shipContent).toContain('Add routing rules to CLAUDE.md');
|
||||
expect(shipContent).toContain("I'll invoke skills manually");
|
||||
test('emitted block carries the routing injection AskUserQuestion', () => {
|
||||
expect(routingBlock).toContain('Add routing rules to CLAUDE.md');
|
||||
expect(routingBlock).toContain("I'll invoke skills manually");
|
||||
});
|
||||
|
||||
test('routing injection respects prior decline', () => {
|
||||
expect(shipContent).toContain('ROUTING_DECLINED');
|
||||
expect(shipContent).toMatch(/routing_declined.*true/);
|
||||
test('routing injection respects prior decline (gate + in-block ack)', () => {
|
||||
expect(SKILL_START_SCRIPT).toContain('[ "$_ROUTING_DECLINED" = "false" ]');
|
||||
expect(routingBlock).toMatch(/routing_declined.*true/);
|
||||
expect(routingBlock).toContain('re-enable with `__BIN__/gstack-config set routing_declined false`');
|
||||
});
|
||||
|
||||
test('routing injection only fires when all conditions met', () => {
|
||||
// Must be: HAS_ROUTING=no AND ROUTING_DECLINED=false AND PROACTIVE_PROMPTED=yes
|
||||
expect(shipContent).toContain('HAS_ROUTING');
|
||||
expect(shipContent).toContain('ROUTING_DECLINED');
|
||||
expect(shipContent).toContain('PROACTIVE_PROMPTED');
|
||||
expect(SKILL_START_SCRIPT).toContain(
|
||||
'if [ "$_HAS_ROUTING" = "no" ] && [ "$_ROUTING_DECLINED" = "false" ] && [ "$_PROACTIVE_PROMPTED" = "yes" ]; then',
|
||||
);
|
||||
});
|
||||
|
||||
test('routing section content includes key routing rules', () => {
|
||||
expect(shipContent).toContain('invoke /office-hours');
|
||||
expect(shipContent).toContain('invoke /investigate');
|
||||
expect(shipContent).toContain('invoke /ship');
|
||||
expect(shipContent).toContain('invoke /qa');
|
||||
expect(routingBlock).toContain('invoke /office-hours');
|
||||
expect(routingBlock).toContain('invoke /investigate');
|
||||
expect(routingBlock).toContain('invoke /ship');
|
||||
expect(routingBlock).toContain('invoke /qa');
|
||||
});
|
||||
|
||||
test('routing section uses renamed checkpoint skills (not stale /checkpoint)', () => {
|
||||
expect(shipContent).toContain('invoke /context-save');
|
||||
expect(shipContent).toContain('invoke /context-restore');
|
||||
expect(shipContent).not.toContain('invoke checkpoint');
|
||||
expect(routingBlock).toContain('invoke /context-save');
|
||||
expect(routingBlock).toContain('invoke /context-restore');
|
||||
expect(routingBlock).not.toContain('invoke checkpoint');
|
||||
});
|
||||
|
||||
test('routing section uses soft "when in doubt" policy, not hard "ALWAYS invoke"', () => {
|
||||
expect(shipContent).toContain('When in doubt, invoke the skill');
|
||||
expect(shipContent).not.toContain('Do NOT answer directly');
|
||||
expect(routingBlock).toContain('When in doubt, invoke the skill');
|
||||
expect(routingBlock).not.toContain('Do NOT answer directly');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1951,8 +2009,16 @@ describe('Codex generation (--host codex)', () => {
|
||||
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gstack-review', 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('GSTACK_ROOT');
|
||||
expect(content).toContain('$_ROOT/.agents/skills/gstack');
|
||||
expect(content).toContain('$GSTACK_BIN/gstack-config');
|
||||
expect(content).toContain('$GSTACK_ROOT/gstack-upgrade/SKILL.md');
|
||||
// Phase 1/2: config reads moved into gstack-skill-start — the fence itself
|
||||
// is the bin asset the preamble must resolve through $GSTACK_BIN, and the
|
||||
// question-preference runtime call still resolves the same way.
|
||||
expect(content).toContain('$GSTACK_BIN/gstack-skill-start');
|
||||
expect(content).toContain('$GSTACK_BIN/gstack-question-preference');
|
||||
// The upgrade-skill doc reference moved into the script's upgrade-flow
|
||||
// block, resolved $0-relative ($_ROOT_DIR) — host-neutral by construction,
|
||||
// so the Codex render no longer needs its own copy.
|
||||
expect(SKILL_START_SCRIPT).toContain('$_ROOT_DIR/gstack-upgrade/SKILL.md');
|
||||
expect(SKILL_START_SCRIPT).toContain('_ROOT_DIR=$(dirname "$_BIN")');
|
||||
expect(content).not.toContain('~/.codex/skills/gstack/bin/gstack-config get telemetry');
|
||||
});
|
||||
|
||||
@@ -2114,7 +2180,9 @@ describe('Codex generation (--host codex)', () => {
|
||||
expect(override.exitCode).toBe(0);
|
||||
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gstack-ship', 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('Model-Specific Behavioral Patch (claude)');
|
||||
expect(content).toContain('MODEL_OVERLAY: claude');
|
||||
// The overlay now travels as --model into gstack-skill-start, which
|
||||
// echoes MODEL_OVERLAY at runtime.
|
||||
expect(content).toContain('--model "claude"');
|
||||
} finally {
|
||||
// Restore the host-default render — later tests and the host-config
|
||||
// golden read this tree.
|
||||
@@ -2127,6 +2195,7 @@ describe('Codex generation (--host codex)', () => {
|
||||
}
|
||||
const restored = fs.readFileSync(path.join(AGENTS_DIR, 'gstack-ship', 'SKILL.md'), 'utf-8');
|
||||
expect(restored).toContain('Model-Specific Behavioral Patch (gpt)');
|
||||
expect(restored).toContain('--model "gpt"');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -2813,39 +2882,55 @@ describe('discover-skills hidden directory filtering', () => {
|
||||
});
|
||||
|
||||
describe('telemetry', () => {
|
||||
test('generated SKILL.md contains telemetry start block', () => {
|
||||
test('telemetry start block lives in gstack-skill-start; render notes the handoff keys', () => {
|
||||
// The start-block bash moved into the script (Phase 1): it reads the
|
||||
// config, mints the session identity, and echoes the STATUS keys.
|
||||
expect(SKILL_START_SCRIPT).toContain('_TEL_START=$(date +%s)');
|
||||
expect(SKILL_START_SCRIPT).toContain('_SESSION_ID=');
|
||||
expect(SKILL_START_SCRIPT).toContain('echo "TELEMETRY:');
|
||||
expect(SKILL_START_SCRIPT).toContain('echo "TEL_PROMPTED:');
|
||||
expect(SKILL_START_SCRIPT).toMatch(/gstack-config" get telemetry/);
|
||||
// The render must tell the model to carry SESSION_ID/TEL_START to skill end.
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('_TEL_START');
|
||||
expect(content).toContain('_SESSION_ID');
|
||||
expect(content).toContain('TELEMETRY:');
|
||||
expect(content).toContain('TEL_PROMPTED:');
|
||||
expect(content).toContain('gstack-config get telemetry');
|
||||
expect(content).toContain('SESSION_ID');
|
||||
expect(content).toContain('TEL_START');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains telemetry opt-in prompt', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('.telemetry-prompted');
|
||||
expect(content).toContain('Help gstack get better');
|
||||
expect(content).toContain('gstack-config set telemetry community');
|
||||
expect(content).toContain('gstack-config set telemetry anonymous');
|
||||
expect(content).toContain('gstack-config set telemetry off');
|
||||
test('telemetry opt-in prompt lives in gstack-skill-start (marker-gated emit)', () => {
|
||||
// Token-reduction Phase 2: the one-time consent prompt left the renders
|
||||
// (absence pinned by test/onboarding-moved-literals.test.ts); the script
|
||||
// gates it on the marker files and emits it as a GSTACK_INSTRUCTION block
|
||||
// with all three config-set outcomes and the ack INSIDE the block.
|
||||
expect(SKILL_START_SCRIPT).toContain(
|
||||
'if [ "$_TEL_PROMPTED" = "no" ] && [ "$_LAKE_SEEN" = "yes" ]; then',
|
||||
);
|
||||
expect(SKILL_START_SCRIPT).toContain('_emit_block telemetry-prompt');
|
||||
expect(SKILL_START_SCRIPT).toContain('gstack-config set telemetry community');
|
||||
expect(SKILL_START_SCRIPT).toContain('gstack-config set telemetry anonymous');
|
||||
expect(SKILL_START_SCRIPT).toContain('gstack-config set telemetry off');
|
||||
expect(SKILL_START_SCRIPT).toContain('touch "$_GH/.telemetry-prompted"');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains telemetry epilogue', () => {
|
||||
test('generated SKILL.md contains telemetry epilogue (one gstack-skill-end call)', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('Telemetry (run last)');
|
||||
expect(content).toContain('gstack-telemetry-log');
|
||||
expect(content).toContain('_TEL_END');
|
||||
expect(content).toContain('_TEL_DUR');
|
||||
expect(content).toContain('SKILL_NAME');
|
||||
expect(content).toContain('OUTCOME');
|
||||
expect(content).toContain('gstack-skill-end --skill "gstack" --outcome OUTCOME');
|
||||
expect(content).toContain('--tel-start "TEL_START"');
|
||||
expect(content).toContain('PLAN MODE EXCEPTION');
|
||||
// The duration math + remote-log dispatch moved into gstack-skill-end.
|
||||
expect(SKILL_END_SCRIPT).toContain('_TEL_END');
|
||||
expect(SKILL_END_SCRIPT).toContain('_TEL_DUR');
|
||||
expect(SKILL_END_SCRIPT).toContain('SKILL_NAME');
|
||||
expect(SKILL_END_SCRIPT).toContain('OUTCOME');
|
||||
expect(SKILL_END_SCRIPT).toContain('gstack-telemetry-log');
|
||||
});
|
||||
|
||||
test('generated SKILL.md contains pending marker handling', () => {
|
||||
const content = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
expect(content).toContain('.pending');
|
||||
expect(content).toContain('_pending_finalize');
|
||||
test('pending marker handling lives in the scripts', () => {
|
||||
// gstack-skill-start finalizes stale markers; gstack-skill-end clears the
|
||||
// session's own marker.
|
||||
expect(SKILL_START_SCRIPT).toContain("-name '.pending-*'");
|
||||
expect(SKILL_START_SCRIPT).toContain('_pending_finalize');
|
||||
expect(SKILL_END_SCRIPT).toContain('.pending-$SESSION_ID');
|
||||
});
|
||||
|
||||
test('telemetry blocks appear in all skill files that use PREAMBLE', () => {
|
||||
@@ -2854,8 +2939,9 @@ describe('telemetry', () => {
|
||||
const skillPath = path.join(ROOT, skill, 'SKILL.md');
|
||||
if (fs.existsSync(skillPath)) {
|
||||
const content = fs.readFileSync(skillPath, 'utf-8');
|
||||
expect(content).toContain('_TEL_START');
|
||||
expect(content).toContain('Telemetry (run last)');
|
||||
expect(content).toContain(`gstack-skill-end --skill "${skill}"`);
|
||||
expect(content).toContain('--tel-start "TEL_START"');
|
||||
}
|
||||
}
|
||||
});
|
||||
@@ -3061,6 +3147,10 @@ describe('codex commands must not use inline $(git rev-parse --show-toplevel) fo
|
||||
'ship/SKILL.md',
|
||||
'codex/SKILL.md.tmpl',
|
||||
'codex/SKILL.md',
|
||||
// codex's scoped invocations moved into the carved review-mode section
|
||||
// (T9) — keep sweeping both the .tmpl source and the generated section.
|
||||
'codex/sections/review-mode.md.tmpl',
|
||||
'codex/sections/review-mode.md',
|
||||
];
|
||||
|
||||
const violations: string[] = [];
|
||||
@@ -3295,6 +3385,29 @@ describe('voice-triggers processing', () => {
|
||||
const frontmatter = content.slice(0, fmEnd);
|
||||
expect(frontmatter).not.toContain('voice-triggers:');
|
||||
});
|
||||
|
||||
// Gen-time-only keys: interactive + benefits-from are read from the .tmpl by
|
||||
// buildContext; the generated copy has no reader (the host reads name/
|
||||
// description/allowed-tools/hooks; gbrain: is runtime-read and NOT stripped).
|
||||
// Pin the strip so a stripFields refactor can't silently re-add the always-on
|
||||
// frontmatter weight — mirrors the voice-triggers pins above.
|
||||
test('generated SKILL.md strips gen-time-only keys the .tmpl still declares', () => {
|
||||
const tmpl = fs.readFileSync(path.join(ROOT, 'plan-ceo-review', 'SKILL.md.tmpl'), 'utf-8');
|
||||
const tmplFm = tmpl.slice(0, tmpl.indexOf('\n---', 4));
|
||||
expect(tmplFm).toContain('interactive:');
|
||||
expect(tmplFm).toContain('benefits-from:');
|
||||
|
||||
const generated = fs.readFileSync(path.join(ROOT, 'plan-ceo-review', 'SKILL.md'), 'utf-8');
|
||||
const genFm = generated.slice(0, generated.indexOf('\n---', 4));
|
||||
expect(genFm).not.toContain('interactive:');
|
||||
expect(genFm).not.toContain('benefits-from:');
|
||||
|
||||
// The runtime-read and host-read keys survive the strip.
|
||||
const investigate = fs.readFileSync(path.join(ROOT, 'investigate', 'SKILL.md'), 'utf-8');
|
||||
const invFm = investigate.slice(0, investigate.indexOf('\n---', 4));
|
||||
expect(invFm).toContain('hooks:');
|
||||
expect(invFm).toContain('gbrain:');
|
||||
});
|
||||
});
|
||||
|
||||
describe('plan-mode-info resolver (handshake-replacement)', () => {
|
||||
@@ -3370,12 +3483,16 @@ describe('plan-mode-info resolver (handshake-replacement)', () => {
|
||||
);
|
||||
|
||||
test('plan-mode-info is wired BEFORE generateUpgradeCheck in preamble', () => {
|
||||
// Token-reduction Phase 2: generateUpgradeCheck's render output is now
|
||||
// ONLY the steady-state PROACTIVE-false + SKILL_PREFIX rules (the
|
||||
// UPGRADE_AVAILABLE prose emits from bin/gstack-skill-start at runtime),
|
||||
// so those rules are the resolver's order marker.
|
||||
const content = fs.readFileSync(
|
||||
path.join(ROOT, 'plan-ceo-review', 'SKILL.md'),
|
||||
'utf-8',
|
||||
);
|
||||
const planModeIdx = content.indexOf(PLAN_MODE_INFO_MARKER);
|
||||
const upgradeIdx = content.indexOf('UPGRADE_AVAILABLE');
|
||||
const upgradeIdx = content.indexOf('If `PROACTIVE` is `"false"`');
|
||||
expect(planModeIdx).toBeGreaterThan(0);
|
||||
expect(upgradeIdx).toBeGreaterThan(0);
|
||||
expect(planModeIdx).toBeLessThan(upgradeIdx);
|
||||
@@ -3678,7 +3795,11 @@ describe('PREAMBLE resolution requires declared preamble-tier', () => {
|
||||
// user scope, so a correctly configured project-scoped brain was invisible.
|
||||
// ---------------------------------------------------------------------------
|
||||
describe('brain-sync block reads project-scoped MCP registrations (#2499)', () => {
|
||||
const rendered = fs.readFileSync(path.join(ROOT, 'SKILL.md'), 'utf-8');
|
||||
// Phase 1: the artifacts-sync bash (including the MCP-scope jq probe) moved
|
||||
// from the rendered SKILL.md into bin/gstack-skill-start. Pin the LIVE
|
||||
// script bytes — same assertions, new home. The render carries only the
|
||||
// ARTIFACTS_SYNC interpretation prose.
|
||||
const rendered = fs.readFileSync(path.join(ROOT, 'bin', 'gstack-skill-start'), 'utf-8');
|
||||
|
||||
test('rendered _GBRAIN_MCP_ENTRY jq resolves project scope with nearest-ancestor cwd match', () => {
|
||||
const line = rendered.split('\n').find((l) => l.includes('_GBRAIN_MCP_ENTRY=$('));
|
||||
|
||||
Reference in New Issue
Block a user