mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-28 17:10:26 +02:00
* feat(gen): strip gen-time-only frontmatter keys from Claude renders
interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate SKILL.md — dead frontmatter keys removed
Mechanical regen after hosts/claude.ts stripFields change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers
New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight
Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage
Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard
Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.69.1.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.69.1.0
CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md
Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated
Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash
generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed
Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-start contract suite + preamble A/B eval + touchfiles registration
test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from 29785978) vs script render with the
fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve
the operator install and silently exercise degraded mode). 21 touchfiles dep
lists gain the two bin scripts (EOV9) so future script edits select the
preamble evals; selection-count pin updated 23->24.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: repin ~70 assertions to the script contract — every literal gets a successor
Assertions that pinned inline-bash internals (update-check guard, _SESSIONS
reaping, telemetry start/end blocks, routing probe, repo-strip producer,
first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, #2499 jq scope
resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new
home: bin/gstack-skill-start / bin/gstack-skill-end file content for script
internals, the invocation fence + interpretation prose for render-side
behavior. No assertion deleted without a successor; live-execution tests
(routing probe, brain-sync jq) run against script bytes unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol)
parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills);
skill-size-budget repointed with the derivation comment citing the Phase 1
context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1
skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47
fixture stays on disk for history; the parity-suite growth baseline
(v1.64.1.0) is untouched. Context-budget ceilings re-captured: review
29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): instruction-emission layer — onboarding text appears only when its gate fires
The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive
opt-in, first-run/first-loop tips, routing injection, vendoring deprecation,
writing-style migration, spawned-session rules), the upgrade-flow + feature
discovery prose, and the privacy stop-gate (user-approved Q2) moved from
every render into gated heredocs here. Blocks are SESSION_ID-bound
(GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint
directives (F4/OV4). Ack ownership per OV6: display-only tips write their
markers at emit (script also fires the scaffold telemetry); interactive flows
carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING
gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard
(same as brain-sync); the quoted routing heredoc resolves its bin path via a
sed placeholder.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule
generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance,
routing-injection,vendoring-deprecation,spawned-session-check,
writing-style-migration}.ts deleted (single source is now the script's
emission layer, F5). generate-upgrade-check shrinks to the steady-state
PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy
stop-gate to the emitted block. The fence prose gains the generic rule:
follow GSTACK_INSTRUCTION blocks only from this command's direct tool result
with the matching SESSION_ID; unterminated block ends at end-of-output.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — onboarding prose degated
Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative
vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: onboarding tombstone + Phase 2 pin relocations
New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must
live in bin/gstack-skill-start AND stay absent from every render, plus the
SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer
contract (gates, block ids, in-block acks, script-run marker writes); the OV4
sanitize test upgraded to the real property (every legitimate block header
carries the run's SESSION_ID). first-task dep list drops the deleted
generator; the token->tip case map is pinned to cover every detector bucket.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9)
All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had
tripped its 72,000 floor at 71,379; design-consultation had 252B of margin).
maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal
parity baseline recaptured in place; ratchet ceilings down again: review
->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer
Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed
CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant
preference, and failure handoff preserved verbatim in behavior, including the
auto-decide-first ordering and the gstack-question-log capture requirement.
5+-options handling (1,924B) compressed to the split invariants (never drop;
D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask
refusal) + the full-rule pointer. Both doc pointers now interpolate the
absolute install root (Codex outside-voice #7 convention) instead of the bare
'in the gstack repo'. Failure-fallback, Format, and self-check sections are
byte-identical — all 14 MANDATORY always-loaded pins pass with zero test
edits.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — AUQ slim
Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays
Branch-internal baseline recaptured; ratchet ceilings down again. OV8's
floor-retirement question, evaluated as planned after Phase 3: the 80% shrink
floor stays — it uniquely catches accidental body deletion in non-carved
skills BETWEEN ratchet recaptures, and the capture command has amortized the
fixture-refresh cost that motivated retiring it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(review): carve adversarial, plan-completion, and review-army into sections
The three resolver macros ship already carves as siblings now load on demand
for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers
stay the single source of truth (sections wrap the macros). Step 0/1, scope
drift, critical pass, confidence calibration, and fix-first stay always-loaded.
Fixtures and pins follow the moved content (codex-hardening wrapped-sites,
review-army E2E fixture builds skeleton+sections with an empty-fixture guard).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(codex): carve the three mutually exclusive modes into sections
Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs)
load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode
dispatch, filesystem boundary, and a new always-loaded 'Synthesis
recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ
per-skill pins pass unchanged); the plan-file report + exit gate render after
the last section pointer per the gateAfterStop pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections
The once-per-repo dry-run validation, the pre-merge readiness gate, and the
merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB ->
55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run
section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained.
Zero content lost (line-coverage checked against HEAD).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5)
They never consume the tier-3 sections (repo-mode ownership, search-before-
building) but do fire AskUserQuestion, which tier >=2 provides — verified by
grep before the plan review. -2.2KB per skill. Render assertions pin the
demotion (tier-3 sections absent, AUQ format present).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow
CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total);
their MONOLITH_INVARIANTS entries retire (invariants now generate from the
registry, cso precedent). Touchfiles: carve-section-loading covers the three
new carves; the codex + land-and-deploy LLM-judge dep lists widen to their
sections. Regen + goldens + branch-internal baseline + ratchet ceilings
recaptured (review 24,052 -> skeleton-based ceiling; union floors hold).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(gen-skill-docs): review render pins read the carved union
The review carve's readSkillUnion conversions (same pattern its neighbor
carved-skill pins already use).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(autoplan): carve the four review phases + tasks aggregator into sections
Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks
aggregator load on demand; Design and DX stay separate sections because each
is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30%
always-loaded); the 6 decision principles, classification, sequencing, and
explicit skip-condition dispatch stay always-loaded. The chain E2E's
phase-complete markers now live only in sections, so its assertions double as
section-read proof (behavioral: external).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(spec): carve the post-confirmation gate-and-file tail into one section
Phases 1-4 are the turn-1 conversational spine — carving them would force the
Read on the first user message for zero real savings. The mechanical tail
(4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only
after draft confirmation: a genuine lazy boundary, kept as ONE section so the
gh-issue-create bash can never load without the fail-closed redaction gate
that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure
invariants migrated location-aware plus a new carve-shape suite (56 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(setup-gbrain): carve the branch-exclusive install paths into sections
Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and
CLAUDE.md persist load on demand — at most one install route ever runs.
Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay
always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es
extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate
counts scan the tmpl union (the third init site lives in engine-remediation).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow
CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec,
and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles:
setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the
stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch
baseline + ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format
RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the
skill prose), degraded path documented in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines
CARVE_GUARDS gains the three entries; qa's monolith invariant retires.
auq-format carve-safety now keys on the skeleton+sections union shipping
the AUQ block (first tier-1 carve: browse never renders it by design).
Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move)
Sol scope discipline stays pinned via the model overlay + completeness
section; the lake intro is now a single script-emitted blurb.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton)
A session runs exactly one mode, so a builder session never loads the
13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule
stay in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections
design-html 57→49KB, design-shotgun 53→50KB. Sections wrap
{{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of
truth; the pretext-patterns STOP sits at the top of Step 3 so the read
provably precedes the Write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills
Both design entries carry requiredReads + loading-eval scenarios (D3A
condition). office-hours phase sections are mode-exclusive, so only the
always-reached design/handoff section is a deterministic requiredRead.
Baselines and ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline
Moved: browser/sidebar/server internals, CHANGELOG release-summary format
spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw
publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction
scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG
entry rules), the machine-managed GBrain block (byte-identical), and the
'## Deploying to the active skill' header with gbrain-refresh in range
(pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): seed onboarding markers into the hermetic child GSTACK_HOME
EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME
seeding in e2e-helpers.ts no longer reaches hermetic children — the
emission layer fired lake-intro/telemetry prompts that burned turns and
stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed
by the telemetry question). Onboarding-specific tests pin their own
GSTACK_HOME per-test, which merges over this seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: raise carve-section-loading wall clock to 480s SDK / 540s bun
The heavy full-workflow scenarios satisfy their required section reads
inside 60s but need 300-450s to finish the report on slower sandboxes;
the 300s default read as a loading failure when the carve invariant held
(traces: plan-eng-review read its section at 8s, office-hours all three
at 24s, design-html both at 50s — all timed out mid-report).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): harden the skill-start trust boundary — review-army findings
Session ID gains a urandom suffix (block binding unforgeable by reflected
content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch
names are charset-clamped before JSON embedding (skill-start + skill-end);
.brain-last-push reads first line only with a charset clamp; the artifacts
URL echo routes through _sanitize; the privacy consent gate fires in
interactive sessions only (spawned auto-choose could accept consent no
human gave — emission order is not a safety property); the daily pull gets
non-interactive + slow-network git guards and stamps only when the
receipted path ran; ~/.claude.json gets a grep pre-filter before the jq
parse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep
The question-log block bound $_SESSION_ID, a shell variable the
consolidated fence never sets — hook-less hosts logged empty session_id,
breaking /plan-tune per-session grouping. It now uses the same
substitute-from-the-skill-start-echoes contract as the telemetry block.
Also: retired the pre-Phase-2 stop-gate docstring, repointed the
gbrain-local-status cross-reference at the script's inline jq, dropped an
orphaned section comment, documented retro-metrics' suffix-only census.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: regenerate renders for the question-log placeholder; goldens + baselines follow
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: hermetic update-check, onboarding gate sequencing, seeding parity
The contract test's child did a live git ls-remote + curl to github.com on
every bun run test (update_check config now gates it off); the headless
test gets a fresh GSTACK_HOME so the suppression is actually exercised; a
new OV6 test drives the script three times to pin ack-at-emit and gate
sequencing; hermetic seeding covers the config-keyed privacy gate; the
EVALS_HERMETIC=0 debug seeding reaches marker parity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix
Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git
shows a pre-Phase-1 sha), which only the periodic workflow provides — and
a static matrix entry so it can't silently never run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.70.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.70.0.0
ARCHITECTURE.md: the preamble section now describes the v1.70 runtime —
the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads
STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding
text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in
every render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim
docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when
the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md.
ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session
count and an "ELI16 mode" that no shipped code implements (the count
computation was deleted with the inline preamble); describe the real
touch-and-prune behavior instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): correct numeric claims against measured counts
50 of 62 installed skills dropped (fixture/alias entries have no preamble);
11 new carves + a deeper office-hours carve = 9→20; test counts match the
files (13 / 11 / 3 / 7).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression
Both attempts in run 33090283032 produced judge-praised DESIGN.md files
phrased as 'design principles'/'design language' without any of the four
original literals; inputs were identical to the prior passing run
32899975845 (design-consultation untouched by the intervening merge).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): stage design-consultation's sections/ into the E2E fixture
The skill has been carved since v1.57.0.0 — the DESIGN.md structure
prescription (the AESTHETIC proposal template) lives in
sections/proposal-and-preview.md behind a STOP-read. The fixture only
copied SKILL.md, so the agent improvised structure from the skeleton and
the section-synonym check has been a coin flip since the carve (CI run
33090283032 trace shows 'no sections dir'; the local eval store has the
same failure on 2026-08-25 while that day's CI run passed on lucky
vocabulary).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
596 lines
33 KiB
Markdown
596 lines
33 KiB
Markdown
# Contributing to gstack
|
|
|
|
Thanks for wanting to make gstack better. Whether you're fixing a typo in a skill prompt or building an entirely new workflow, this guide will get you up and running fast.
|
|
|
|
## Quick start
|
|
|
|
gstack skills are Markdown files that Claude Code discovers from a `skills/` directory. Normally they live at `~/.claude/skills/gstack/` (your global install). But when you're developing gstack itself, you want Claude Code to use the skills *in your working tree* — so edits take effect instantly without copying or deploying anything.
|
|
|
|
That's what dev mode does. It symlinks your repo into the local `.claude/skills/` directory so Claude Code reads skills straight from your checkout.
|
|
|
|
```bash
|
|
git clone https://github.com/garrytan/gstack.git && cd gstack
|
|
bun install # install dependencies
|
|
bin/dev-setup # activate dev mode
|
|
```
|
|
|
|
> **Full clone vs shallow.** The README's user-facing install uses `--depth 1` for speed. As a contributor, use a full clone (no `--depth` flag) — you'll need history for `git log`, `git blame`, `git bisect`, and reviewing PRs against earlier versions. If you already have a `--depth 1` clone from following the README, promote it to a full clone with `git fetch --unshallow`.
|
|
|
|
Now edit any `SKILL.md`, invoke it in Claude Code (e.g. `/review`), and see your changes live. When you're done developing:
|
|
|
|
```bash
|
|
bin/dev-teardown # deactivate — back to your global install
|
|
```
|
|
|
|
## Operational self-improvement
|
|
|
|
gstack automatically learns from failures. At the end of every skill session, the agent
|
|
reflects on what went wrong (CLI errors, wrong approaches, project quirks) and logs
|
|
operational learnings to `~/.gstack/projects/{slug}/learnings.jsonl`. Future sessions
|
|
surface these learnings automatically, so gstack gets smarter on your codebase over time.
|
|
|
|
No setup needed. Learnings are logged automatically. View them with `/learn`.
|
|
|
|
### The contributor workflow
|
|
|
|
1. **Use gstack normally** — operational learnings are captured automatically
|
|
2. **Check your learnings:** `/learn` or `ls ~/.gstack/projects/*/learnings.jsonl`
|
|
3. **Fork and clone gstack** (if you haven't already)
|
|
4. **Symlink your fork into the project where you hit the bug:**
|
|
```bash
|
|
# In your core project (the one where gstack annoyed you)
|
|
ln -sfn /path/to/your/gstack-fork .claude/skills/gstack
|
|
cd .claude/skills/gstack && bun install && bun run build && ./setup
|
|
```
|
|
Setup creates per-skill directories with SKILL.md symlinks inside (`qa/SKILL.md -> gstack/qa/SKILL.md`),
|
|
links each skill's runtime assets alongside (sections/, templates, checklists — everything except
|
|
SKILL.md, tests, build output, and `.tmpl` sources), and asks your prefix preference.
|
|
Pass `--no-prefix` to skip the prompt and use short names.
|
|
5. **Fix the issue** — your changes are live immediately in this project
|
|
6. **Test by actually using gstack** — do the thing that annoyed you, verify it's fixed
|
|
7. **Open a PR from your fork**
|
|
|
|
This is the best way to contribute: fix gstack while doing your real work, in the
|
|
project where you actually felt the pain.
|
|
|
|
### Session awareness
|
|
|
|
When you have 3+ gstack sessions open simultaneously, every question tells you which project, which branch, and what's happening. No more staring at a question thinking "wait, which window is this?" The format is consistent across all skills.
|
|
|
|
## Working on gstack inside the gstack repo
|
|
|
|
When you're editing gstack skills and want to test them by actually using gstack
|
|
in the same repo, `bin/dev-setup` wires this up. It creates `.claude/skills/`
|
|
symlinks (gitignored) pointing back to your working tree, so Claude Code uses
|
|
your local edits instead of the global install.
|
|
|
|
```
|
|
gstack/ <- your working tree
|
|
├── .claude/skills/ <- created by dev-setup (gitignored)
|
|
│ ├── gstack -> ../../ <- symlink back to repo root
|
|
│ ├── review/ <- real directory (short name, default)
|
|
│ │ └── SKILL.md -> gstack/review/SKILL.md
|
|
│ ├── ship/ <- or gstack-review/, gstack-ship/ if --prefix
|
|
│ │ └── SKILL.md -> gstack/ship/SKILL.md
|
|
│ └── ... <- one directory per skill
|
|
├── review/
|
|
│ └── SKILL.md <- edit this, test with /review
|
|
├── ship/
|
|
│ └── SKILL.md
|
|
├── browse/
|
|
│ ├── src/ <- TypeScript source
|
|
│ └── dist/ <- compiled binary (gitignored)
|
|
└── ...
|
|
```
|
|
|
|
Setup creates real directories (not symlinks) at the top level with a SKILL.md
|
|
symlink inside, plus links to each skill's runtime assets (sections/, templates,
|
|
checklists). Alias skills (`_gstack-command`, `connect-chrome`) install as
|
|
rewritten copies, never symlinks — editing a symlinked alias would corrupt the
|
|
generated source. This ensures Claude discovers them as top-level skills, not nested
|
|
under `gstack/`. Names depend on your prefix setting (`~/.gstack/config.yaml`).
|
|
Short names (`/review`, `/ship`) are the default. Run `./setup --prefix` if you
|
|
prefer namespaced names (`/gstack-review`, `/gstack-ship`).
|
|
|
|
## Day-to-day workflow
|
|
|
|
```bash
|
|
# 1. Enter dev mode
|
|
bin/dev-setup
|
|
|
|
# 2. Edit a skill
|
|
vim review/SKILL.md
|
|
|
|
# 3. Test it in Claude Code — changes are live
|
|
# > /review
|
|
|
|
# 4. Editing browse source? Rebuild the binary
|
|
bun run build
|
|
|
|
# 5. Done for the day? Tear down
|
|
bin/dev-teardown
|
|
```
|
|
|
|
### Brain-aware blocks in a dev workspace (gbrain installed)
|
|
|
|
If gbrain is installed and usable (`bin/gstack-gbrain-detect --is-ok` exits 0),
|
|
`bin/dev-setup` keeps your tracked `SKILL.md` files canonical and renders the
|
|
brain-aware variant (the `GBRAIN_CONTEXT_LOAD` / `GBRAIN_SAVE_RESULTS` blocks)
|
|
into `.claude/gstack-rendered/` (gitignored, per-workspace). It then repoints the
|
|
workspace's `SKILL.md` symlinks at that render, so your Claude sessions get the
|
|
full gbrain experience while `git status` stays clean. Under the hood, dev-setup
|
|
passes `GSTACK_SKIP_GBRAIN_REGEN=1` inline to the nested `./setup` (so it never
|
|
dirties tracked source) and runs `gen:skill-docs:user --out-dir .claude/gstack-rendered`,
|
|
which rewrites only the section-base paths to point at the render. `bin/dev-teardown`
|
|
removes the render. To make the blocks live across your *other* projects' Claude
|
|
sessions, run `gstack-config gbrain-refresh`, which renders them to a user render
|
|
dir (`${GSTACK_USER_RENDER_DIR:-~/.gstack/render/claude}`, swapped in only on a
|
|
successful render) and repoints the installed skills at it via `gstack-relink` —
|
|
the global install checkout stays git-clean, and the refresh is guarded so it
|
|
never touches a symlinked or non-gstack directory.
|
|
|
|
## Testing & evals
|
|
|
|
### Setup
|
|
|
|
```bash
|
|
# 1. Copy .env.example and add your API key
|
|
cp .env.example .env
|
|
# Edit .env → set ANTHROPIC_API_KEY=sk-ant-...
|
|
|
|
# 2. Install deps (if you haven't already)
|
|
bun install
|
|
```
|
|
|
|
Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` from the main worktree automatically (see "Conductor workspaces" below).
|
|
|
|
### Test tiers
|
|
|
|
| Tier | Command | Cost | What it tests |
|
|
|------|---------|------|---------------|
|
|
| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
|
|
| 2 — E2E | `bun run test:e2e` | ~$3.85 | Full skill execution via `claude -p` subprocess |
|
|
| 3 — LLM eval | `EVALS=1 bun test test/skill-llm-eval.test.ts` | ~$0.15 standalone | LLM-as-judge scoring of generated SKILL.md docs |
|
|
| 2+3 | `bun run test:evals` | ~$4 combined | E2E + LLM-as-judge (runs both) |
|
|
|
|
```bash
|
|
bun run test # Tier 1 only (run before every commit, ~90-100s for the full ~7,000-test suite)
|
|
bun run test:e2e # Tier 2: E2E only (needs EVALS=1, can't run inside Claude Code)
|
|
bun run test:evals # Tier 2 + 3 combined (~$4/run)
|
|
```
|
|
|
|
### Tier 1: Static validation (free)
|
|
|
|
Runs with `bun run test`, which routes through `scripts/test-free-shards.ts`: N
|
|
concurrent shard processes under a strict output contract — a shard that exits
|
|
without bun's own terminal summary line, or a crashed worker, fails the run, so
|
|
silent truncation can never report green. Pass `--verbose` to forward the full
|
|
child stream; `--wall-timeout <secs>` overrides the per-shard kill deadline.
|
|
Don't type bare `bun test` for the suite: it walks the whole repo, loads paid
|
|
eval files, and misses the strict classifier. No API keys needed.
|
|
|
|
- **Skill parser tests** (`test/skill-parser.test.ts`) — Extracts every `$B` command from SKILL.md bash code blocks and validates against the command registry in `browse/src/commands.ts`. Catches typos, removed commands, and invalid snapshot flags.
|
|
- **Skill validation tests** (`test/skill-validation.test.ts`) — Validates that SKILL.md files reference only real commands and flags, and that command descriptions meet quality thresholds.
|
|
- **Generator tests** (`test/gen-skill-docs.test.ts`) — Tests the template system: verifies placeholders resolve correctly, output includes value hints for flags (e.g. `-d <N>` not just `-d`), enriched descriptions for key commands (e.g. `is` lists valid states, `press` lists key examples).
|
|
- **Tier-alignment invariant** (`test/e2e-tier-alignment.test.ts`) — For every self-gated `test/skill-e2e-*.test.ts` named in a touchfiles dep list, the file's `EVALS_TIER` self-gate must match its declared tier in `E2E_TIERS`. Kills the "inert demotion" class where a test is re-tiered in `touchfiles.ts` but the file still gates on the old tier and keeps running in the wrong lane. Unmapped or mixed-tier files are reported, never silently skipped.
|
|
- **Catalog budget** (`test/catalog-budget.test.ts`) — Caps the aggregate discovery surface: the sum of every skill's frontmatter `name` + `description` (what every host loads at discovery, every session) must stay under 1,150 token-equivalents, with a 260-byte per-skill cap. Counting goes through the shared census in `test/helpers/skill-census.ts` (physical files vs authored skills vs registry entries — three deliberately different counts). Adding a skill? The failure message carries the re-measure + ratchet protocol.
|
|
- **Context-budget ratchet** (`test/context-budget-ratchet.test.ts`) — CI ceilings on the two token ledgers the catalog budget doesn't cover: the always-on full-frontmatter aggregate and each skill's per-invocation eager tokens (SKILL.md + forced-read references), graded against `test/fixtures/context-budget.json` via `lib/context-bill.ts`. New skills fail until they have a ceiling; ceilings for removed skills must be pruned. Legitimate growth or a landed reduction: re-run `bun test/helpers/capture-context-budget.ts` and commit the refreshed fixture in the same commit, so the change is a visible decision in the diff.
|
|
|
|
### Tier 2: E2E via `claude -p` (~$3.85/run)
|
|
|
|
Spawns `claude -p` as a subprocess with `--output-format stream-json --verbose`, streams NDJSON for real-time progress, and scans for browse errors. This is the closest thing to "does this skill actually work end-to-end?"
|
|
|
|
```bash
|
|
# Must run from a plain terminal — can't nest inside Claude Code or Conductor
|
|
EVALS=1 bun test test/skill-e2e-*.test.ts
|
|
```
|
|
|
|
- Gated by `EVALS=1` env var (prevents accidental expensive runs)
|
|
- Auto-skips if running inside Claude Code (`claude -p` can't nest)
|
|
- API connectivity pre-check — fails fast on ConnectionRefused before burning budget
|
|
- Real-time progress to stderr: `[Ns] turn T tool #C: Name(...)`
|
|
- Saves full NDJSON transcripts and failure JSON for debugging
|
|
- Tests live in `test/skill-e2e-*.test.ts` (split by category), runner logic in `test/helpers/session-runner.ts`
|
|
|
|
**Hermetic by default.** Every E2E runner (claude -p, the real-PTY plan-mode
|
|
runner, the Agent SDK runner, plus the codex and gemini runners) spawns its child
|
|
through `test/helpers/hermetic-env.ts`: an allowlist-scrubbed environment, a fresh
|
|
seeded `CLAUDE_CONFIG_DIR`, a temp `GSTACK_HOME`, and `--strict-mcp-config`. Your
|
|
operator `~/.claude` config, MCP servers (gbrain, Conductor), skills, `~/.gstack`
|
|
decision logs, and `CONDUCTOR_*` env never leak into the child, so local eval
|
|
signal matches CI instead of disagreeing for reasons unrelated to the code under
|
|
test. The hermetic `CLAUDE_CONFIG_DIR` seeds no skills by default; a PTY test
|
|
that types a `/skill` slash command passes `seedSkills: true` to the PTY runner,
|
|
which swaps in `hermeticSkillsConfigDir()` — a seeded skill registry that
|
|
symlinks the LIVE working tree's SKILL.md files (by design: the skills are the
|
|
subject under test, so a snapshot would measure stale copies). Set
|
|
`EVALS_HERMETIC=0` to debug against your real operator state (this also
|
|
drops `--strict-mcp-config`). The wiring is pinned by `test/hermetic-wiring.test.ts`
|
|
(a free static tripwire), two gate-tier isolation canaries in
|
|
`test/skill-e2e-hermetic-canary.test.ts`, and the skill-seeding tripwires in
|
|
`test/hermetic-skills-seeding.test.ts` / `test/pty-skill-seeding-wiring.test.ts`.
|
|
|
|
### E2E observability
|
|
|
|
When E2E tests run, they produce machine-readable artifacts in `~/.gstack-dev/`:
|
|
|
|
| Artifact | Path | Purpose |
|
|
|----------|------|---------|
|
|
| Heartbeat | `e2e-live.json` | Current test status (updated per tool call) |
|
|
| Partial results | `evals/_partial-e2e.json` | Completed tests (survives kills) |
|
|
| Progress log | `e2e-runs/{runId}/progress.log` | Append-only text log |
|
|
| NDJSON transcripts | `e2e-runs/{runId}/{test}.ndjson` | Raw `claude -p` output per test |
|
|
| Failure JSON | `e2e-runs/{runId}/{test}-failure.json` | Diagnostic data on failure |
|
|
|
|
**Live dashboard:** Run `bun run eval:watch` in a second terminal to see a live dashboard showing completed tests, the currently running test, and cost. Use `--tail` to also show the last 10 lines of progress.log.
|
|
|
|
**Eval history tools:**
|
|
|
|
```bash
|
|
bun run eval:list # list all eval runs (turns, duration, cost per run)
|
|
bun run eval:compare # compare two runs — shows per-test deltas + Takeaway commentary
|
|
bun run eval:summary # aggregate stats + per-test efficiency averages across runs
|
|
```
|
|
|
|
**Detached runs for agents and long suites.** When an agent (or you, for a run
|
|
you don't want to babysit) launches a long eval, use the `eval:bg*` scripts. They
|
|
wrap the eval command in `bin/gstack-detach`: a fresh session that escapes a
|
|
turn-boundary SIGTERM, a `caffeinate` wrapper that blocks idle-sleep, a machine-wide
|
|
`gstack-evals` lock so concurrent worktrees serialize instead of saturating the
|
|
model API, a run-scoped log under `~/.gstack-dev/eval-runs/`, a per-tier watchdog,
|
|
and a guaranteed `### gstack-detach EXIT=<code> ###` sentinel so a poller never
|
|
mistakes silence for success.
|
|
|
|
```bash
|
|
bun run eval:bg # detached test:evals (diff-based)
|
|
bun run eval:bg:all # detached test:evals:all
|
|
bun run eval:bg:gate # detached gate-tier suite
|
|
bun run eval:bg:periodic # detached periodic-tier suite
|
|
```
|
|
|
|
Each prints its log path. The gate and periodic variants run their tier through
|
|
the sharded paid runner (`scripts/test-paid-shards.ts`, also available directly
|
|
as `bun run test:gate:sharded` / `bun run test:periodic:sharded`): one Bun
|
|
process per test file, an external wall-clock timeout that kills the shard's
|
|
whole process group (stray `claude`/`codex` grandchildren included), a per-shard
|
|
eval dir (`GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/`), and an aggregate that
|
|
distinguishes failed vs timed-out vs never-started shards. The runner also
|
|
selects by diff: shards untouched by your branch are reported as
|
|
skipped-by-diff, with a selection banner naming the reason (`EVALS_ALL=1`
|
|
forces everything). `EVALS_JOBS` sets how many shard processes run at once
|
|
(default 4); `EVALS_CONCURRENCY` is bun's concurrency WITHIN a shard — they
|
|
are deliberately separate knobs. `eval:list`,
|
|
`eval:compare`, and `eval:summary` are shard-aware. Humans running
|
|
`bun run test:evals` foreground in their own terminal don't need this — Ctrl-C
|
|
is intended there.
|
|
|
|
**Eval comparison commentary:** `eval:compare` generates natural-language Takeaway sections interpreting what changed between runs — flagging regressions, noting improvements, calling out efficiency gains (fewer turns, faster, cheaper), and producing an overall summary. This is driven by `generateCommentary()` in `eval-store.ts`.
|
|
|
|
Artifacts are never cleaned up — they accumulate in `~/.gstack-dev/` for post-mortem debugging and trend analysis.
|
|
|
|
### Tier 3: LLM-as-judge (~$0.15/run)
|
|
|
|
Uses Claude Sonnet to score generated SKILL.md docs on three dimensions.
|
|
Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`:
|
|
|
|
- **Clarity** — Can an AI agent understand the instructions without ambiguity?
|
|
- **Completeness** — Are all commands, flags, and usage patterns documented?
|
|
- **Actionability** — Can the agent execute tasks using only the information in the doc?
|
|
|
|
Each dimension is scored 1-5. Threshold: every dimension must score **≥ 4**. There's also a regression test that compares generated docs against the hand-maintained baseline from `origin/main` — generated must score equal or higher.
|
|
|
|
```bash
|
|
# Needs ANTHROPIC_API_KEY in .env — included in bun run test:evals
|
|
```
|
|
|
|
- Uses `claude-sonnet-4-6` for scoring stability
|
|
- Tests live in `test/skill-llm-eval.test.ts`
|
|
- Calls the Anthropic API directly (not `claude -p`), so it works from anywhere including inside Claude Code
|
|
|
|
### CI
|
|
|
|
A GitHub Action (`.github/workflows/skill-docs.yml`) runs `bun run gen:skill-docs --dry-run` on every push and PR. If the generated SKILL.md files differ from what's committed, CI fails. This catches stale docs before they merge.
|
|
|
|
Supply-chain gates run alongside it:
|
|
|
|
- **Quality gate** (`.github/workflows/quality-gate.yml`, every PR and push) — scans the diff's added lines for credentials using gstack's own redact engine (`.github/scripts/gate-secret-scan.mjs`). HIGH findings fail the job; MEDIUM findings surface as an advisory count. Fails closed if the scan can't produce a report. Also gates critical dependency advisories and runs ShellCheck on the setup/build boundaries.
|
|
- **Dependency review** (`.github/workflows/dependency-review.yml`) — reviews dependency changes on PRs that touch lockfiles or workflow files.
|
|
- **OSV scanner** (`.github/workflows/osv-scanner.yml`) — weekly vulnerability scan against the OSV database (config in `.osv-scanner.toml`).
|
|
- **Dependabot** (`.github/dependabot.yml`) — grouped dependency update PRs.
|
|
|
|
The supply-chain workflows pin their third-party actions to commit SHAs. The PR template (`.github/PULL_REQUEST_TEMPLATE.md`) asks for evidence — tests run, eval output — not promises.
|
|
|
|
Tests run against the browse binary directly — they don't require dev mode.
|
|
|
|
## Editing SKILL.md files
|
|
|
|
SKILL.md files are **generated** from `.tmpl` templates. Don't edit the `.md` directly — your changes will be overwritten on the next build.
|
|
|
|
```bash
|
|
# 1. Edit the template
|
|
vim SKILL.md.tmpl # or browse/SKILL.md.tmpl
|
|
|
|
# 2. Regenerate for all hosts
|
|
bun run gen:skill-docs --host all
|
|
|
|
# 3. Check health (reports all hosts)
|
|
bun run skill:check
|
|
|
|
# Or use watch mode — auto-regenerates on save
|
|
bun run dev:skill
|
|
```
|
|
|
|
For template authoring best practices (natural language over bash-isms, dynamic branch detection, `{{BASE_BRANCH_DETECT}}` usage), see CLAUDE.md's "Writing SKILL templates" section.
|
|
|
|
To add a browse command, add it to `browse/src/commands.ts`. To add a snapshot flag, add it to `SNAPSHOT_FLAGS` in `browse/src/snapshot.ts`. Then rebuild.
|
|
|
|
**Don't bundle puppeteer/Chromium in a skill.** `browse` is the one shared
|
|
Chromium per box, including offline local-render workloads. A skill that needs to
|
|
rasterize its own HTML/JSON (diagrams, cards, og-images) should route through
|
|
`browse` — `screenshot --selector` for visual output, `load-html` + `js --out` for
|
|
bytes a render function returns — instead of `npm i puppeteer` and downloading a
|
|
second Chromium that drifts out of version sync. One install to pin, one daemon to
|
|
manage.
|
|
|
|
## Jargon list (V1 writing style)
|
|
|
|
gstack's Writing Style section (injected into every tier-≥2 skill's preamble)
|
|
glosses technical terms on first use per skill invocation. The list of terms
|
|
that qualify for glossing lives at `scripts/jargon-list.json` — ~50 curated
|
|
high-frequency terms (idempotent, race condition, N+1, backpressure, etc.).
|
|
Terms not on the list are assumed plain-English enough.
|
|
|
|
**Adding or removing a term:** open a PR editing `scripts/jargon-list.json`.
|
|
Run `bun run gen:skill-docs` after the edit — terms are baked into every
|
|
generated SKILL.md at gen time, so changes take effect only after regeneration.
|
|
No runtime loading; no user-side override. The repo list is the source of truth.
|
|
|
|
Good candidates for addition: high-frequency terms that non-technical users
|
|
encounter in review output without context (common database/concurrency
|
|
terminology, security jargon, frontend framework concepts). Don't add terms
|
|
that only appear in one or two niche skills — the cost-to-value trade isn't
|
|
worth the review overhead.
|
|
|
|
## Multi-host development
|
|
|
|
gstack generates SKILL.md files for 8 hosts from one set of `.tmpl` templates.
|
|
Each host is a typed config in `hosts/*.ts`. The generator reads these configs
|
|
to produce host-appropriate output (different frontmatter, paths, tool names).
|
|
|
|
**Supported hosts:** Claude (primary), Codex, Factory, Kiro, OpenCode, Slate, Cursor, OpenClaw.
|
|
|
|
### Generating for all hosts
|
|
|
|
```bash
|
|
# Generate for a specific host
|
|
bun run gen:skill-docs # Claude (default)
|
|
bun run gen:skill-docs --host codex # Codex
|
|
bun run gen:skill-docs --host opencode # OpenCode
|
|
bun run gen:skill-docs --host all # All 8 hosts
|
|
|
|
# Or use build, which does all hosts + compiles binaries
|
|
bun run build
|
|
```
|
|
|
|
### What changes between hosts
|
|
|
|
Each host config (`hosts/*.ts`) controls:
|
|
|
|
| Aspect | Example (Claude vs Codex) |
|
|
|--------|---------------------------|
|
|
| Output directory | `{skill}/SKILL.md` vs `.agents/skills/gstack-{skill}/SKILL.md` |
|
|
| Frontmatter | Full (name, description, hooks, version) vs minimal (name + description) |
|
|
| Paths | `~/.claude/skills/gstack` vs `$GSTACK_ROOT` |
|
|
| Tool names | "use the Bash tool" vs same (Factory rewrites to "run this command") |
|
|
| Hook skills | `hooks:` frontmatter vs inline safety advisory prose |
|
|
| Suppressed sections | None vs Codex self-invocation sections stripped |
|
|
| Model overlay | `claude` vs `gpt` (per-host `defaultModel`; `--model` or, at setup time, the Codex `config.toml` model overrides) |
|
|
|
|
See `scripts/host-config.ts` for the full `HostConfig` interface.
|
|
|
|
### Testing host output
|
|
|
|
```bash
|
|
# Run all static tests (includes parameterized smoke tests for all hosts)
|
|
bun run test
|
|
|
|
# Check freshness for all hosts
|
|
bun run gen:skill-docs --host all --dry-run
|
|
|
|
# Health dashboard covers all hosts
|
|
bun run skill:check
|
|
```
|
|
|
|
### Adding a new host
|
|
|
|
See [docs/ADDING_A_HOST.md](docs/ADDING_A_HOST.md) for the full guide. Short version:
|
|
|
|
1. Create `hosts/myhost.ts` (copy from `hosts/opencode.ts`)
|
|
2. Add to `hosts/index.ts`
|
|
3. Add `.myhost/` to `.gitignore`
|
|
4. Run `bun run gen:skill-docs --host myhost`
|
|
5. Run `bun run test` (parameterized tests auto-cover it)
|
|
|
|
Zero generator, setup, or tooling code changes needed.
|
|
|
|
### Adding a new skill
|
|
|
|
When you add a new skill template, all hosts get it automatically:
|
|
1. Create `{skill}/SKILL.md.tmpl`
|
|
2. Run `bun run gen:skill-docs --host all`
|
|
3. The dynamic template discovery picks it up, no static list to update
|
|
4. Budget it: run `bun test/helpers/capture-context-budget.ts` and commit the refreshed `test/fixtures/context-budget.json` — the context-budget ratchet fails any skill without a ceiling
|
|
5. Commit `{skill}/SKILL.md`, external host output is generated at setup time and gitignored
|
|
|
|
## Conductor workspaces
|
|
|
|
If you're using [Conductor](https://conductor.build) to run multiple Claude Code sessions in parallel, `conductor.json` wires up workspace lifecycle automatically:
|
|
|
|
| Hook | Script | What it does |
|
|
|------|--------|-------------|
|
|
| `setup` | `bin/dev-setup` | Copies `.env` from main worktree, installs deps, symlinks skills, runs `./setup` non-interactively, and (if gbrain is installed) renders brain-aware blocks into `.claude/gstack-rendered/` without dirtying tracked source |
|
|
| `archive` | `bin/dev-teardown` | Removes skill symlinks, the `.claude/gstack-rendered/` render, and cleans up `.claude/` directory |
|
|
|
|
When Conductor creates a new workspace, `bin/dev-setup` runs automatically. It detects the main worktree (via `git worktree list`), copies your `.env` so API keys carry over, and sets up dev mode — no manual steps needed.
|
|
|
|
`bin/dev-setup` runs `./setup` fully non-interactively (it passes `--plan-tune-hooks=prompt` and closes stdin), so a forwarded Conductor TTY can never hang on a hidden setup prompt. It also never installs the plan-tune Claude Code hooks, which means a throwaway workspace can't rewrite your global `~/.claude/settings.json` to point at an ephemeral worktree path. To install the plan-tune hooks deliberately, run `./setup --plan-tune-hooks` outside dev-setup (or `gstack-config set plan_tune_hooks yes`). The explicit flag counts as an explicit decision: setup's Conductor auto-opt-in for AskUserQuestion hooks fires only on the true silent fall-through (no flag, no `GSTACK_PLAN_TUNE_HOOKS` env var, no `plan_tune_hooks` key literally present in config, checked via `gstack-config has`), so it can never override dev-setup into installing hooks. One stated repair exception: setup's heal-first pass (`gstack-settings-hook prune-stale --repoint`) may prune dead gstack hook entries and re-point existing ones at the stable `~/.claude/skills/gstack` install. That is strictly convergent repair, never a new registration, and registration itself is canonical-only, so an ephemeral tree path can never be baked into settings.json.
|
|
|
|
**First-time setup:** Put your `ANTHROPIC_API_KEY` in `.env` in the main repo (see `.env.example`). Every Conductor workspace inherits it automatically.
|
|
|
|
**`GSTACK_*` env prefix (Conductor-injected keys).** Conductor explicitly strips `ANTHROPIC_API_KEY` and `OPENAI_API_KEY` from every workspace's process env. The `.env` copy path doesn't restore them either — the strip happens after env inheritance. Users who want paid evals, `/sync-gbrain` embeddings, or `claude-agent-sdk` calls to work in a Conductor workspace must set `GSTACK_ANTHROPIC_API_KEY` and `GSTACK_OPENAI_API_KEY` in Conductor's workspace env config; Conductor passes those through untouched. On the gstack side, TS entry points import `lib/conductor-env-shim.ts` as a side effect, which promotes `GSTACK_FOO_API_KEY` to `FOO_API_KEY` when the canonical name is empty. If you add a new TS entry point that hits a paid API, add `import "../lib/conductor-env-shim";` to the top of the file. Today the shim is imported from `bin/gstack-gbrain-sync.ts`, `bin/gstack-model-benchmark`, `scripts/preflight-agent-sdk.ts`, and `test/helpers/e2e-helpers.ts`.
|
|
|
|
## Things to know
|
|
|
|
- **SKILL.md files are generated.** Edit the `.tmpl` template, not the `.md`. Run `bun run gen:skill-docs` to regenerate.
|
|
- **TODOS.md is the unified backlog.** Organized by skill/component with P0-P4 priorities. `/ship` auto-detects completed items. All planning/review/retro skills read it for context.
|
|
- **Browse source changes need a rebuild.** If you touch `browse/src/*.ts`, run `bun run build`.
|
|
- **Dev mode shadows your global install.** Project-local skills take priority over `~/.claude/skills/gstack`. `bin/dev-teardown` restores the global one.
|
|
- **Conductor workspaces are independent.** Each workspace is its own git worktree. `bin/dev-setup` runs automatically via `conductor.json`.
|
|
- **`.env` propagates across worktrees.** Set it once in the main repo, all Conductor workspaces get it.
|
|
- **`.claude/skills/` is gitignored.** The symlinks never get committed.
|
|
- **Never write raw `ln -snf` in `setup`.** Every link site in `setup` MUST route through the `_link_or_copy SRC DST` helper near the `IS_WINDOWS` detection. The helper preserves `ln -snf` on Unix and switches to `cp -R` / `cp -f` on Windows without Developer Mode, where plain `ln -snf` produces frozen file copies that don't refresh on `git pull`. `test/setup-windows-fallback.test.ts` enforces this with a static invariant — a single raw `ln` call outside the helper body fails CI.
|
|
|
|
## Testing your changes in a real project
|
|
|
|
**This is the recommended way to develop gstack.** Symlink your gstack checkout
|
|
into the project where you actually use it, so your changes are live while you
|
|
do real work.
|
|
|
|
### Step 1: Symlink your checkout
|
|
|
|
```bash
|
|
# In your core project (not the gstack repo)
|
|
ln -sfn /path/to/your/gstack-checkout .claude/skills/gstack
|
|
```
|
|
|
|
### Step 2: Run setup to create per-skill symlinks
|
|
|
|
The `gstack` symlink alone isn't enough. Claude Code discovers skills through
|
|
individual top-level directories (`qa/SKILL.md`, `ship/SKILL.md`, etc.), not through
|
|
the `gstack/` directory itself. Run `./setup` to create them:
|
|
|
|
```bash
|
|
cd .claude/skills/gstack && bun install && bun run build && ./setup
|
|
```
|
|
|
|
Setup will ask whether you want short names (`/qa`) or namespaced (`/gstack-qa`).
|
|
Your choice is saved to `~/.gstack/config.yaml` and remembered for future runs.
|
|
To skip the prompt, pass `--no-prefix` (short names) or `--prefix` (namespaced).
|
|
|
|
### Step 3: Develop
|
|
|
|
Edit a template, run `bun run gen:skill-docs`, and the next `/review` or `/qa`
|
|
call picks it up immediately. No restart needed.
|
|
|
|
### Going back to the stable global install
|
|
|
|
Remove the project-local symlink. Claude Code falls back to `~/.claude/skills/gstack/`:
|
|
|
|
```bash
|
|
rm .claude/skills/gstack
|
|
```
|
|
|
|
The per-skill directories (`qa/`, `ship/`, etc.) contain SKILL.md symlinks that point
|
|
to `gstack/...`, so they'll resolve to the global install automatically.
|
|
|
|
### Switching prefix mode
|
|
|
|
If you installed gstack with one prefix setting and want to switch:
|
|
|
|
```bash
|
|
cd .claude/skills/gstack && ./setup --no-prefix # switch to /qa, /ship
|
|
cd .claude/skills/gstack && ./setup --prefix # switch to /gstack-qa, /gstack-ship
|
|
```
|
|
|
|
Setup cleans up the old symlinks automatically. No manual cleanup needed.
|
|
|
|
### Alternative: point your global install at a branch
|
|
|
|
If you don't want per-project symlinks, you can switch the global install:
|
|
|
|
```bash
|
|
cd ~/.claude/skills/gstack
|
|
git fetch origin
|
|
git checkout origin/<branch>
|
|
bun install && bun run build && ./setup
|
|
```
|
|
|
|
This affects all projects. To revert: `git checkout main && git pull && bun run build && ./setup`.
|
|
|
|
## Community PR triage (wave process)
|
|
|
|
When community PRs accumulate, batch them into themed waves:
|
|
|
|
1. **Categorize** — group by theme (security, features, infra, docs)
|
|
2. **Deduplicate** — if two PRs fix the same thing, pick the one that
|
|
changes fewer lines. Close the other with a note pointing to the winner.
|
|
3. **Collector branch** — create `pr-wave-N`, merge clean PRs, resolve
|
|
conflicts for dirty ones, verify with `bun run test && bun run build`
|
|
4. **Close with context** — every closed PR gets a comment explaining
|
|
why and what (if anything) supersedes it. Contributors did real work;
|
|
respect that with clear communication.
|
|
5. **Ship as one PR** — single PR to main with all attributions preserved
|
|
in merge commits. Include a summary table of what merged and what closed.
|
|
|
|
See [PR #205](../../pull/205) (v0.8.3) for the first wave as an example.
|
|
|
|
## Upgrade migrations
|
|
|
|
When a release changes on-disk state (directory structure, config format, stale
|
|
files) in ways that `./setup` alone can't fix, add a migration script so existing
|
|
users get a clean upgrade.
|
|
|
|
### When to add a migration
|
|
|
|
- Changed how skill directories are created (symlinks vs real dirs)
|
|
- Renamed or moved config keys in `~/.gstack/config.yaml`
|
|
- Need to delete orphaned files from a previous version
|
|
- Changed the format of `~/.gstack/` state files
|
|
|
|
Don't add a migration for: new features (users get them automatically), new
|
|
skills (setup discovers them), or code-only changes (no on-disk state).
|
|
|
|
### How to add one
|
|
|
|
1. Create `gstack-upgrade/migrations/v{VERSION}.sh` where `{VERSION}` matches
|
|
the VERSION file for the release that needs the fix.
|
|
2. Make it executable: `chmod +x gstack-upgrade/migrations/v{VERSION}.sh`
|
|
3. The script must be **idempotent** (safe to run multiple times) and
|
|
**non-fatal** (failures are logged but don't block the upgrade).
|
|
4. Include a comment block at the top explaining what changed, why the
|
|
migration is needed, and which users are affected.
|
|
|
|
Example:
|
|
|
|
```bash
|
|
#!/usr/bin/env bash
|
|
# Migration: v0.15.2.0 — Fix skill directory structure
|
|
# Affected: users who installed with --no-prefix before v0.15.2.0
|
|
set -euo pipefail
|
|
SCRIPT_DIR="$(cd "$(dirname "$0")/../.." && pwd)"
|
|
"$SCRIPT_DIR/bin/gstack-relink" 2>/dev/null || true
|
|
```
|
|
|
|
### How it runs
|
|
|
|
During `/gstack-upgrade`, after `./setup` completes (Step 4.75), the upgrade
|
|
skill scans `gstack-upgrade/migrations/` and runs every `v*.sh` script whose
|
|
version is newer than the user's old version. Scripts run in version order.
|
|
Failures are logged but never block the upgrade.
|
|
|
|
### Testing migrations
|
|
|
|
Migrations are tested as part of `bun run test` (tier 1, free). The test suite
|
|
verifies that all migration scripts in `gstack-upgrade/migrations/` are
|
|
executable and parse without syntax errors.
|
|
|
|
## Shipping your changes
|
|
|
|
When you're happy with your skill edits:
|
|
|
|
```bash
|
|
/ship
|
|
```
|
|
|
|
This runs tests, reviews the diff, triages Greptile comments (with 2-tier escalation), manages TODOS.md, bumps the version, and opens a PR. See `ship/SKILL.md` for the full workflow.
|