mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-29 09:20:39 +02:00
v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders
interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate SKILL.md — dead frontmatter keys removed
Mechanical regen after hosts/claude.ts stripFields change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers
New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight
Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage
Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard
Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.69.1.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.69.1.0
CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md
Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated
Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash
generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed
Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-start contract suite + preamble A/B eval + touchfiles registration
test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from 29785978) vs script render with the
fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve
the operator install and silently exercise degraded mode). 21 touchfiles dep
lists gain the two bin scripts (EOV9) so future script edits select the
preamble evals; selection-count pin updated 23->24.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: repin ~70 assertions to the script contract — every literal gets a successor
Assertions that pinned inline-bash internals (update-check guard, _SESSIONS
reaping, telemetry start/end blocks, routing probe, repo-strip producer,
first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, #2499 jq scope
resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new
home: bin/gstack-skill-start / bin/gstack-skill-end file content for script
internals, the invocation fence + interpretation prose for render-side
behavior. No assertion deleted without a successor; live-execution tests
(routing probe, brain-sync jq) run against script bytes unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol)
parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills);
skill-size-budget repointed with the derivation comment citing the Phase 1
context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1
skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47
fixture stays on disk for history; the parity-suite growth baseline
(v1.64.1.0) is untouched. Context-budget ceilings re-captured: review
29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): instruction-emission layer — onboarding text appears only when its gate fires
The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive
opt-in, first-run/first-loop tips, routing injection, vendoring deprecation,
writing-style migration, spawned-session rules), the upgrade-flow + feature
discovery prose, and the privacy stop-gate (user-approved Q2) moved from
every render into gated heredocs here. Blocks are SESSION_ID-bound
(GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint
directives (F4/OV4). Ack ownership per OV6: display-only tips write their
markers at emit (script also fires the scaffold telemetry); interactive flows
carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING
gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard
(same as brain-sync); the quoted routing heredoc resolves its bin path via a
sed placeholder.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule
generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance,
routing-injection,vendoring-deprecation,spawned-session-check,
writing-style-migration}.ts deleted (single source is now the script's
emission layer, F5). generate-upgrade-check shrinks to the steady-state
PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy
stop-gate to the emitted block. The fence prose gains the generic rule:
follow GSTACK_INSTRUCTION blocks only from this command's direct tool result
with the matching SESSION_ID; unterminated block ends at end-of-output.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — onboarding prose degated
Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative
vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: onboarding tombstone + Phase 2 pin relocations
New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must
live in bin/gstack-skill-start AND stay absent from every render, plus the
SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer
contract (gates, block ids, in-block acks, script-run marker writes); the OV4
sanitize test upgraded to the real property (every legitimate block header
carries the run's SESSION_ID). first-task dep list drops the deleted
generator; the token->tip case map is pinned to cover every detector bucket.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9)
All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had
tripped its 72,000 floor at 71,379; design-consultation had 252B of margin).
maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal
parity baseline recaptured in place; ratchet ceilings down again: review
->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer
Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed
CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant
preference, and failure handoff preserved verbatim in behavior, including the
auto-decide-first ordering and the gstack-question-log capture requirement.
5+-options handling (1,924B) compressed to the split invariants (never drop;
D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask
refusal) + the full-rule pointer. Both doc pointers now interpolate the
absolute install root (Codex outside-voice #7 convention) instead of the bare
'in the gstack repo'. Failure-fallback, Format, and self-check sections are
byte-identical — all 14 MANDATORY always-loaded pins pass with zero test
edits.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + goldens — AUQ slim
Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays
Branch-internal baseline recaptured; ratchet ceilings down again. OV8's
floor-retirement question, evaluated as planned after Phase 3: the 80% shrink
floor stays — it uniquely catches accidental body deletion in non-carved
skills BETWEEN ratchet recaptures, and the capture command has amortized the
fixture-refresh cost that motivated retiring it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(review): carve adversarial, plan-completion, and review-army into sections
The three resolver macros ship already carves as siblings now load on demand
for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers
stay the single source of truth (sections wrap the macros). Step 0/1, scope
drift, critical pass, confidence calibration, and fix-first stay always-loaded.
Fixtures and pins follow the moved content (codex-hardening wrapped-sites,
review-army E2E fixture builds skeleton+sections with an empty-fixture guard).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(codex): carve the three mutually exclusive modes into sections
Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs)
load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode
dispatch, filesystem boundary, and a new always-loaded 'Synthesis
recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ
per-skill pins pass unchanged); the plan-file report + exit gate render after
the last section pointer per the gateAfterStop pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections
The once-per-repo dry-run validation, the pre-merge readiness gate, and the
merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB ->
55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run
section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained.
Zero content lost (line-coverage checked against HEAD).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5)
They never consume the tier-3 sections (repo-mode ownership, search-before-
building) but do fire AskUserQuestion, which tier >=2 provides — verified by
grep before the plan review. -2.2KB per skill. Render assertions pin the
demotion (tier-3 sections absent, AUQ format present).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow
CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total);
their MONOLITH_INVARIANTS entries retire (invariants now generate from the
registry, cso precedent). Touchfiles: carve-section-loading covers the three
new carves; the codex + land-and-deploy LLM-judge dep lists widen to their
sections. Regen + goldens + branch-internal baseline + ratchet ceilings
recaptured (review 24,052 -> skeleton-based ceiling; union floors hold).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(gen-skill-docs): review render pins read the carved union
The review carve's readSkillUnion conversions (same pattern its neighbor
carved-skill pins already use).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(autoplan): carve the four review phases + tasks aggregator into sections
Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks
aggregator load on demand; Design and DX stay separate sections because each
is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30%
always-loaded); the 6 decision principles, classification, sequencing, and
explicit skip-condition dispatch stay always-loaded. The chain E2E's
phase-complete markers now live only in sections, so its assertions double as
section-read proof (behavioral: external).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(spec): carve the post-confirmation gate-and-file tail into one section
Phases 1-4 are the turn-1 conversational spine — carving them would force the
Read on the first user message for zero real savings. The mechanical tail
(4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only
after draft confirmation: a genuine lazy boundary, kept as ONE section so the
gh-issue-create bash can never load without the fail-closed redaction gate
that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure
invariants migrated location-aware plus a new carve-shape suite (56 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(setup-gbrain): carve the branch-exclusive install paths into sections
Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and
CLAUDE.md persist load on demand — at most one install route ever runs.
Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay
always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es
extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate
counts scan the tmpl union (the third init site lives in engine-remediation).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow
CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec,
and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles:
setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the
stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch
baseline + ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format
RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the
skill prose), degraded path documented in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines
CARVE_GUARDS gains the three entries; qa's monolith invariant retires.
auq-format carve-safety now keys on the skeleton+sections union shipping
the AUQ block (first tier-1 carve: browse never renders it by design).
Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move)
Sol scope discipline stays pinned via the model overlay + completeness
section; the lake intro is now a single script-emitted blurb.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton)
A session runs exactly one mode, so a builder session never loads the
13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule
stay in the skeleton.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections
design-html 57→49KB, design-shotgun 53→50KB. Sections wrap
{{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of
truth; the pretext-patterns STOP sits at the top of Step 3 so the read
provably precedes the Write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills
Both design entries carry requiredReads + loading-eval scenarios (D3A
condition). office-hours phase sections are mode-exclusive, so only the
always-reached design/handoff section is a deterministic requiredRead.
Baselines and ratchet recaptured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline
Moved: browser/sidebar/server internals, CHANGELOG release-summary format
spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw
publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction
scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG
entry rules), the machine-managed GBrain block (byte-identical), and the
'## Deploying to the active skill' header with gbrain-refresh in range
(pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): seed onboarding markers into the hermetic child GSTACK_HOME
EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME
seeding in e2e-helpers.ts no longer reaches hermetic children — the
emission layer fired lake-intro/telemetry prompts that burned turns and
stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed
by the telemetry question). Onboarding-specific tests pin their own
GSTACK_HOME per-test, which merges over this seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: raise carve-section-loading wall clock to 480s SDK / 540s bun
The heavy full-workflow scenarios satisfy their required section reads
inside 60s but need 300-450s to finish the report on slower sandboxes;
the 300s default read as a loading failure when the carve invariant held
(traces: plan-eng-review read its section at 8s, office-hours all three
at 24s, design-html both at 50s — all timed out mid-report).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): harden the skill-start trust boundary — review-army findings
Session ID gains a urandom suffix (block binding unforgeable by reflected
content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch
names are charset-clamped before JSON embedding (skill-start + skill-end);
.brain-last-push reads first line only with a charset clamp; the artifacts
URL echo routes through _sanitize; the privacy consent gate fires in
interactive sessions only (spawned auto-choose could accept consent no
human gave — emission order is not a safety property); the daily pull gets
non-interactive + slow-network git guards and stamps only when the
receipted path ran; ~/.claude.json gets a grep pre-filter before the jq
parse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep
The question-log block bound $_SESSION_ID, a shell variable the
consolidated fence never sets — hook-less hosts logged empty session_id,
breaking /plan-tune per-session grouping. It now uses the same
substitute-from-the-skill-start-echoes contract as the telemetry block.
Also: retired the pre-Phase-2 stop-gate docstring, repointed the
gbrain-local-status cross-reference at the script's inline jq, dropped an
orphaned section comment, documented retro-metrics' suffix-only census.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: regenerate renders for the question-log placeholder; goldens + baselines follow
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: hermetic update-check, onboarding gate sequencing, seeding parity
The contract test's child did a live git ls-remote + curl to github.com on
every bun run test (update_check config now gates it off); the headless
test gets a fresh GSTACK_HOME so the suppression is actually exercised; a
new OV6 test drives the script three times to pin ack-at-emit and gate
sequencing; hermetic seeding covers the config-keyed privacy gate; the
EVALS_HERMETIC=0 debug seeding reaches marker parity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix
Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git
shows a pre-Phase-1 sha), which only the periodic workflow provides — and
a static matrix entry so it can't silently never run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.70.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.70.0.0
ARCHITECTURE.md: the preamble section now describes the v1.70 runtime —
the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads
STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding
text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in
every render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim
docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when
the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md.
ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session
count and an "ELI16 mode" that no shipped code implements (the count
computation was deleted with the inline preamble); describe the real
touch-and-prune behavior instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): correct numeric claims against measured counts
50 of 62 installed skills dropped (fixture/alias entries have no preamble);
11 new carves + a deeper office-hours carve = 9→20; test counts match the
files (13 / 11 / 3 / 7).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression
Both attempts in run 33090283032 produced judge-praised DESIGN.md files
phrased as 'design principles'/'design language' without any of the four
original literals; inputs were identical to the prior passing run
32899975845 (design-consultation untouched by the intervening merge).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): stage design-consultation's sections/ into the E2E fixture
The skill has been carved since v1.57.0.0 — the DESIGN.md structure
prescription (the AESTHETIC proposal template) lives in
sections/proposal-and-preview.md behind a STOP-read. The fixture only
copied SKILL.md, so the agent improvised structure from the skeleton and
the section-synonym check has been a coin flip since the carve (CI run
33090283032 trace shows 'no sections dir'; the local eval store has the
same failure on 2026-08-25 while that day's CI run passed on lucky
vocabulary).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
a3749bfa4b
commit
394db326f2
@@ -0,0 +1,163 @@
|
||||
# Browser / sidebar / server internals
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction). These are the
|
||||
load-bearing invariants for `browse/src/server.ts`, the Chrome extension,
|
||||
the sidebar PTY, SSE endpoints, CDP sessions, and the sidebar security
|
||||
stack. Every rule here is additionally pinned by a CI tripwire test named
|
||||
in its paragraph.
|
||||
|
||||
**Sidebar architecture:** Before modifying `sidepanel.js`, `background.js`,
|
||||
`content.js`, `terminal-agent.ts`, or sidebar-related server endpoints,
|
||||
read `docs/designs/SIDEBAR_MESSAGE_FLOW.md`. The sidebar has one primary
|
||||
surface — the **Terminal** pane (interactive `claude` PTY) — with
|
||||
Activity / Refs / Inspector as debug overlays behind the footer's
|
||||
`debug` toggle. The chat queue path was ripped once the PTY proved out;
|
||||
`sidebar-agent.ts` and the `/sidebar-command` / `/sidebar-chat` /
|
||||
`/sidebar-agent/event` endpoints are gone. The doc covers the WS auth
|
||||
flow, dual-token model, and threat-model boundary — silent failures
|
||||
here usually trace to not understanding the cross-component flow.
|
||||
|
||||
**Embedder terminal-agent ownership** (v1.42.1.0+, identity-based kill v1.44.0.0+).
|
||||
`buildFetchHandler` in `browse/src/server.ts` accepts `ServerConfig.ownsTerminalAgent?:
|
||||
boolean` (default `true`). When `true`, factory shutdown runs the full teardown:
|
||||
identity-based kill via `killAgentByRecord(readAgentRecord(stateDir))` from
|
||||
`browse/src/terminal-agent-control.ts` plus `safeUnlinkQuiet` on
|
||||
`<stateDir>/terminal-port`, `<stateDir>/terminal-internal-token`, and
|
||||
`<stateDir>/terminal-agent-pid` (the per-boot agent record introduced in v1.44).
|
||||
Embedders (e.g. the gbrowser phoenix overlay) that pre-launch their own PTY
|
||||
server must pass `false` so their discovery files survive gstack teardown cycles.
|
||||
The flag is the third caller-owned teardown gate in `ServerConfig` (alongside
|
||||
`xvfb?` and `proxyBridge?`); polarity is inverted (explicit bool vs presence) and
|
||||
documented in the field's JSDoc. CLI `start()` always passes `true` explicitly —
|
||||
the static-grep test in `browse/test/server-embedder-terminal-port.test.ts` fails
|
||||
CI if a refactor drops it. Pre-v1.44 used `pkill -f terminal-agent\.ts` (regex
|
||||
match) which would kill sibling gstack sessions on the same host; the new
|
||||
`browse/test/terminal-agent-pid-identity.test.ts` static-grep tripwire fails CI
|
||||
if any source file re-introduces `pkill ... terminal-agent` or `spawnSync('pkill', ...)`.
|
||||
|
||||
**WebSocket auth uses Sec-WebSocket-Protocol, not cookies.** Browsers
|
||||
can't set `Authorization` on a WebSocket upgrade, but they CAN set
|
||||
`Sec-WebSocket-Protocol` via `new WebSocket(url, [token])`. The agent
|
||||
reads it, validates against `validTokens`, and MUST echo the protocol
|
||||
back in the upgrade response — without the echo, Chromium closes the
|
||||
connection immediately. `Set-Cookie: gstack_pty=...` is kept as a
|
||||
fallback for non-browser callers (the cross-port `SameSite=Strict`
|
||||
cookie path doesn't survive from a chrome-extension origin).
|
||||
|
||||
**Cross-pane PTY injection.** The toolbar's Cleanup button and the
|
||||
Inspector's "Send to Code" action both pipe text into the live claude
|
||||
PTY via `window.gstackInjectToTerminal(text)`, exposed by
|
||||
`sidepanel-terminal.js`. No `/sidebar-command` POST — the live REPL is
|
||||
the only execution surface in the sidebar now.
|
||||
|
||||
**`/health` MUST NOT surface any token — and it no longer does** (v1.63+).
|
||||
The historical headed-mode leak of `AUTH_TOKEN` is fixed: `GET /health` is
|
||||
liveness/status only in every mode. Token bootstrap is `POST /extension-token`,
|
||||
which validates the caller's Origin against the pinned extension identity
|
||||
(the `key` field in `extension/manifest.json` pins the extension ID —
|
||||
`GSTACK_EXTENSION_ID` in `browse/src/server.ts`, derivation reproducible via
|
||||
`bun browse/scripts/extension-id.ts`) plus a loopback Host. PTY auth still
|
||||
flows through `POST /pty-session` only. Don't add any token to `/health`.
|
||||
|
||||
**Transport-layer security** (v1.6.0.0+). When `pair-agent` starts an ngrok tunnel,
|
||||
the daemon binds two HTTP listeners: a local listener (127.0.0.1, full command
|
||||
surface, never forwarded) and a tunnel listener (locked allowlist: `/connect`,
|
||||
`/command` with a scoped token + 26-command browser-driving allowlist,
|
||||
`/sidebar-chat`). ngrok forwards only the tunnel port. Root tokens over the tunnel
|
||||
return 403. SSE endpoints use a 30-minute HttpOnly `gstack_sse` cookie minted via
|
||||
`POST /sse-session` (never valid against `/command`). Tunnel-surface rejections go
|
||||
to `~/.gstack/security/attempts.jsonl` via `tunnel-denial-log.ts`. Before editing
|
||||
`server.ts`, `sse-session-cookie.ts`, or `tunnel-denial-log.ts`, read
|
||||
[ARCHITECTURE.md](../ARCHITECTURE.md#dual-listener-tunnel-architecture-v1600) —
|
||||
the module boundary (no imports from `token-registry.ts` into `sse-session-cookie.ts`)
|
||||
is load-bearing for scope isolation.
|
||||
|
||||
**Unicode sanitization at server egress** (v1.38.0.0+). Every server egress that
|
||||
ships page-content-derived strings MUST go through `JSON.stringify(payload,
|
||||
sanitizeReplacer)` for object payloads or `sanitizeLoneSurrogates(body)` for text
|
||||
bodies. Lone UTF-16 surrogate halves from CDP page content otherwise reach the
|
||||
Anthropic API as `\uD800`-style escapes and trigger a 400. Wired at four egress
|
||||
points today: `handleCommandInternal` (HTTP + batch via a sanitizing wrapper around
|
||||
`handleCommandInternalImpl`) and both SSE producers (`/activity/stream`,
|
||||
`/inspector/events`). Post-stringify regex is a no-op — `JSON.stringify` has
|
||||
already escaped the surrogate before regex could match, so the replacer must run
|
||||
inside the encoding pipeline. Before adding a new SSE/WebSocket writer or HTTP
|
||||
response in `server.ts`, read
|
||||
[ARCHITECTURE.md](../ARCHITECTURE.md#unicode-sanitization-at-server-egress-v13800).
|
||||
`browse/test/server-sanitize-surrogates.test.ts` pins the wiring with invariant
|
||||
tests, so bypasses fail CI.
|
||||
|
||||
**SSE endpoint helper** (v1.51.0.0+). New SSE endpoints in `server.ts` MUST route
|
||||
through `createSseEndpoint(req, config)` from `browse/src/sse-helpers.ts`. The
|
||||
helper owns the cleanup contract (abort + enqueue-throw + heartbeat-throw, all
|
||||
idempotent) and bakes in `sanitizeLoneSurrogates` on every JSON.stringify, so
|
||||
new subscribers can't accidentally regress either invariant. Inline
|
||||
`ReadableStream` wiring leaked subscribers when the TCP connection died without
|
||||
firing `req.signal.abort` (Chromium MV3 service-worker suspend, intermediate
|
||||
proxy half-close). `/activity/stream`, `/inspector/events`, and `/memory`
|
||||
(SSE-eligible) all route through it. `browse/test/sse-helpers.test.ts` pins the
|
||||
cleanup contract.
|
||||
|
||||
**CDP session lifecycle** (v1.51.0.0+). Direct `page.context().newCDPSession(page)`
|
||||
calls outside `browse/src/cdp-bridge.ts` fail CI via the static-grep tripwire in
|
||||
`browse/test/cdp-session-cleanup.test.ts`. Use `withCdpSession(page, async (s) => {...})`
|
||||
for one-shot CDP work (try/finally detach) or `getOrCreateCdpSession(page, cache)`
|
||||
for cached sessions tied to a page's lifetime (close-detach via `Map<page, session>`).
|
||||
Three sites migrated: cdp-bridge frame events, write-commands archive capture,
|
||||
cdp-inspector. The helpers prevent the per-session leak class where successful-path
|
||||
detach happened but error-path detach was missed.
|
||||
|
||||
**Setup symlink hardening** (v1.38.0.0+). Every link site in `setup` MUST route
|
||||
through the `_link_or_copy SRC DST` helper near the `IS_WINDOWS` detection. On
|
||||
Windows without Developer Mode, plain `ln -snf` produces frozen file copies that
|
||||
don't refresh on `git pull` — silent staleness across every host adapter. The
|
||||
helper preserves `ln -snf` on Unix and switches to `cp -R` / `cp -f` on Windows.
|
||||
`test/setup-windows-fallback.test.ts` enforces a static invariant: a single raw
|
||||
`ln` call outside the helper body fails CI. Windows users get a one-line note
|
||||
from `_print_windows_copy_note_once` reminding them to re-run `./setup` after
|
||||
every `git pull`.
|
||||
|
||||
**Sidebar security stack** (layered defense against prompt injection):
|
||||
|
||||
| Layer | Module | Lives in |
|
||||
|-------|--------|----------|
|
||||
| L1-L3 | `content-security.ts` | server + read path — datamarking, hidden element strip, ARIA regex, URL blocklist, envelope wrapping |
|
||||
| L4 | `security-classifier.ts` (TestSavantAI ONNX) | **security sidecar subprocess only** (`security-sidecar-entry.ts`, driven by `security-sidecar-client.ts` from server.ts) |
|
||||
| Canary | `security.ts` (generate/inject/detect) | pure utilities — no production injector today (the chat prompt-builder that injected them was ripped) |
|
||||
| Combiner | `security.ts` (combineVerdict + THRESHOLDS) | pure, tested; retains transcript/deberta vote handling for LayerSignal inputs no live layer produces anymore |
|
||||
|
||||
History note: an L4b Haiku transcript classifier and an opt-in DeBERTa ensemble
|
||||
(`GSTACK_SECURITY_ENSEMBLE=deberta`) existed until the chat-path agent that
|
||||
invoked them was ripped; both were deleted as dead code (zero production
|
||||
callers). Do not re-document them as live.
|
||||
|
||||
**Critical constraint:** `security-classifier.ts` CANNOT be imported from the
|
||||
compiled browse binary. `@huggingface/transformers` v4 requires `onnxruntime-node`
|
||||
which fails to `dlopen` from Bun compile's temp extract dir — hence the sidecar
|
||||
subprocess. Only `security.ts` (pure-string operations — canary utilities,
|
||||
verdict combiner, status) is safe for `server.ts`. See
|
||||
`~/.gstack/projects/garrytan-gstack/ceo-plans/2026-04-19-prompt-injection-guard.md`
|
||||
§"Pre-Impl Gate 1 Outcome" for the original architectural decision.
|
||||
|
||||
**Thresholds** (in `security.ts`): `BLOCK: 0.85`, `WARN: 0.75`, `LOG_ONLY: 0.40`,
|
||||
`SOLO_CONTENT_BLOCK: 0.92` (label-less content classifiers can't distinguish
|
||||
"injection" from "phishing aimed at the user", so their solo bar is higher).
|
||||
The live L4 path applies these in server.ts's sidecar-scan handling; canary
|
||||
leak always BLOCKs (deterministic).
|
||||
|
||||
**Env knobs:**
|
||||
- `GSTACK_SECURITY_OFF=1` — emergency kill switch. Classifier stays off even if
|
||||
warmed; the L1-L3 filters keep running.
|
||||
- Classifier model cache: `~/.gstack/models/testsavant-small/` (112MB, first run only)
|
||||
- Attack log: `~/.gstack/security/attempts.jsonl` — written by
|
||||
`tunnel-denial-log.ts` (tunnel-surface rejections; rotates at 10MB, 5 generations)
|
||||
|
||||
History note (#2557): the cross-process session state
|
||||
(`~/.gstack/security/session-state.json`), `getStatus()`, the `/health`
|
||||
`security` field, and the sidepanel SEC shield were all removed — the state
|
||||
file lost its only writer when sidebar-agent.ts was ripped, so the shield
|
||||
reported a permanent 'inactive' or a stale false-green 'protected' from
|
||||
leftover disk state. The live defenses (L1-L3 filters, L4 sidecar on the
|
||||
inject-scan path) report through their own call sites, never through
|
||||
/health. `browse/test/server-security-surface.test.ts` pins both the
|
||||
removal and the live L4 wiring. Do not re-document these as live.
|
||||
@@ -0,0 +1,58 @@
|
||||
# CHANGELOG entry format
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction). Read this BEFORE
|
||||
writing any `## [X.Y.Z]` CHANGELOG entry.
|
||||
|
||||
### Release-summary format (every `## [X.Y.Z]` entry)
|
||||
|
||||
Every version entry in `CHANGELOG.md` MUST start with a release-summary section in
|
||||
the GStack/Garry voice, one viewport's worth of prose + tables that lands like a
|
||||
verdict, not marketing. The itemized changelog (subsections, bullets, files) goes
|
||||
BELOW that summary, separated by a `### Itemized changes` header.
|
||||
|
||||
The release-summary section gets read by humans, by the auto-update agent, and by
|
||||
anyone deciding whether to upgrade. The itemized list is for agents that need to
|
||||
know exactly what changed.
|
||||
|
||||
Structure for the top of every `## [X.Y.Z]` entry:
|
||||
|
||||
1. **Two-line bold headline** (10-14 words total). Should land like a verdict, not
|
||||
marketing. Sound like someone who shipped today and cares whether it works.
|
||||
2. **Lead paragraph** (3-5 sentences). What shipped, what changed for the user.
|
||||
Specific, concrete, no AI vocabulary, no em dashes, no hype.
|
||||
3. **A "The X numbers that matter" section** with:
|
||||
- One short setup paragraph naming the source of the numbers (real production
|
||||
deployment OR a reproducible benchmark, name the file/command to run).
|
||||
- A table of 3-6 key metrics with BEFORE / AFTER / Δ columns.
|
||||
- A second optional table for per-category breakdown if relevant.
|
||||
- 1-2 sentences interpreting the most striking number in concrete user terms.
|
||||
4. **A "What this means for [audience]" closing paragraph** (2-4 sentences) tying
|
||||
the metrics to a real workflow shift. End with what to do.
|
||||
|
||||
Voice rules for the release summary:
|
||||
- No em dashes (use commas, periods, "...").
|
||||
- No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or
|
||||
banned phrases ("here's the kicker", "the bottom line", etc.).
|
||||
- Real numbers, real file names, real commands. Not "fast" but "~30s on 30K pages."
|
||||
- Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
|
||||
- Connect to user outcomes: "the agent does ~3x less reading" beats "improved precision."
|
||||
- Be direct about quality. "Well-designed" or "this is a mess." No dancing.
|
||||
|
||||
Source material:
|
||||
- CHANGELOG previous entry for prior context.
|
||||
- Benchmark files or `/retro` output for headline numbers.
|
||||
- Recent commits (`git log <prev-version>..HEAD --oneline`) for what shipped.
|
||||
- Don't make up numbers. If a metric isn't in a benchmark or production data,
|
||||
don't include it. Say "no measurement yet" if asked.
|
||||
|
||||
Target length: ~250-350 words for the summary. Should render as one viewport.
|
||||
|
||||
### Itemized changes (below the release summary)
|
||||
|
||||
Write `### Itemized changes` and continue with the detailed subsections (Added,
|
||||
Changed, Fixed, For contributors). Same rules as the user-facing voice guidance
|
||||
above, plus:
|
||||
|
||||
- **Always credit community contributions.** When an entry includes work from a
|
||||
community PR, name the contributor with `Contributed by @username`. Contributors
|
||||
did real work. Thank them publicly every time, no exceptions.
|
||||
@@ -0,0 +1,26 @@
|
||||
# Publishing native OpenClaw skills to ClawHub
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction).
|
||||
|
||||
## Workflow
|
||||
|
||||
Native OpenClaw skills live in `openclaw/skills/gstack-openclaw-*/SKILL.md`. These are
|
||||
hand-crafted methodology skills (not generated by the pipeline) published to ClawHub
|
||||
so any OpenClaw user can install them.
|
||||
|
||||
**Publishing:** The command is `clawhub publish` (NOT `clawhub skill publish`):
|
||||
|
||||
```bash
|
||||
clawhub publish openclaw/skills/gstack-openclaw-office-hours \
|
||||
--slug gstack-openclaw-office-hours --name "gstack Office Hours" \
|
||||
--version 1.0.0 --changelog "description of changes"
|
||||
```
|
||||
|
||||
Repeat for each skill: `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`,
|
||||
`gstack-openclaw-retro`. Bump `--version` on each update.
|
||||
|
||||
**Auth:** `clawhub login` (opens browser for GitHub auth). `clawhub whoami` to verify.
|
||||
|
||||
**Updating:** Same `clawhub publish` command with a higher `--version` and `--changelog`.
|
||||
|
||||
**Verification:** `clawhub search gstack` to confirm they're live.
|
||||
@@ -0,0 +1,79 @@
|
||||
# Project structure
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction).
|
||||
|
||||
## Directory tree
|
||||
|
||||
```
|
||||
gstack/
|
||||
├── browse/ # Headless browser CLI (Playwright)
|
||||
│ ├── src/ # CLI + server + commands
|
||||
│ │ ├── commands.ts # Command registry (single source of truth)
|
||||
│ │ └── snapshot.ts # SNAPSHOT_FLAGS metadata array
|
||||
│ ├── test/ # Integration tests + fixtures
|
||||
│ └── dist/ # Compiled binary
|
||||
├── hosts/ # Typed host configs (one per AI agent)
|
||||
│ ├── claude.ts # Primary host config
|
||||
│ ├── codex.ts, factory.ts, kiro.ts # Existing hosts
|
||||
│ ├── opencode.ts, slate.ts, cursor.ts, openclaw.ts # IDE hosts
|
||||
│ ├── hermes.ts, gbrain.ts # Agent runtime hosts
|
||||
│ └── index.ts # Registry: exports all, derives Host type
|
||||
├── scripts/ # Build + DX tooling
|
||||
│ ├── gen-skill-docs.ts # Template → SKILL.md generator (config-driven)
|
||||
│ ├── host-config.ts # HostConfig interface + validator
|
||||
│ ├── host-config-export.ts # Shell bridge for setup script
|
||||
│ ├── resolvers/ # Template resolver modules (preamble, design, review, gbrain, etc.)
|
||||
│ ├── skill-check.ts # Health dashboard
|
||||
│ ├── test-paid-shards.ts # Sharded paid-tier runner (one Bun process per shard)
|
||||
│ └── dev-skill.ts # Watch mode
|
||||
├── test/ # Skill validation + eval tests
|
||||
│ ├── helpers/ # skill-parser.ts, session-runner.ts, llm-judge.ts, eval-store.ts
|
||||
│ ├── fixtures/ # Ground truth JSON, planted-bug fixtures, eval baselines
|
||||
│ ├── skill-validation.test.ts # Tier 1: static validation (free, <1s)
|
||||
│ ├── gen-skill-docs.test.ts # Tier 1: generator quality (free, <1s)
|
||||
│ ├── skill-llm-eval.test.ts # Tier 3: LLM-as-judge (~$0.15/run)
|
||||
│ └── skill-e2e-*.test.ts # Tier 2: E2E via claude -p (~$3.85/run, split by category)
|
||||
├── qa-only/ # /qa-only skill (report-only QA, no fixes)
|
||||
├── plan-design-review/ # /plan-design-review skill (report-only design audit)
|
||||
├── design-review/ # /design-review skill (design audit + fix loop)
|
||||
├── ship/ # Ship workflow skill
|
||||
├── review/ # PR review skill
|
||||
├── plan-ceo-review/ # /plan-ceo-review skill
|
||||
├── plan-eng-review/ # /plan-eng-review skill
|
||||
├── autoplan/ # /autoplan skill (auto-review pipeline: CEO → design → eng)
|
||||
├── benchmark/ # /benchmark skill (performance regression detection)
|
||||
├── canary/ # /canary skill (post-deploy monitoring loop)
|
||||
├── codex/ # /codex skill (multi-AI second opinion via OpenAI Codex CLI)
|
||||
├── land-and-deploy/ # /land-and-deploy skill (merge → deploy → canary verify)
|
||||
├── office-hours/ # /office-hours skill (YC Office Hours — startup diagnostic + builder brainstorm)
|
||||
├── investigate/ # /investigate skill (systematic root-cause debugging)
|
||||
├── spec/ # /spec skill (five-phase spec → GitHub issue, optional agent spawn, /ship auto-closes)
|
||||
├── retro/ # Retrospective skill (includes /retro global cross-project mode)
|
||||
├── bin/ # CLI utilities (gstack-repo-mode, gstack-slug, gstack-config, gstack-wtree, gstack-evidence, gstack-issue-guard, etc.)
|
||||
├── document-release/ # /document-release skill (post-ship doc updates + Diataxis coverage map)
|
||||
├── document-generate/ # /document-generate skill (Diataxis doc generator: tutorial/how-to/reference/explanation)
|
||||
├── cso/ # /cso skill (OWASP Top 10 + STRIDE security audit)
|
||||
├── design-consultation/ # /design-consultation skill (design system from scratch)
|
||||
├── design-shotgun/ # /design-shotgun skill (visual design exploration)
|
||||
├── open-gstack-browser/ # /open-gstack-browser skill (launch GStack Browser)
|
||||
├── connect-chrome/ # symlink → open-gstack-browser (backwards compat)
|
||||
├── design/ # Design binary CLI (GPT Image API)
|
||||
│ ├── src/ # CLI + commands (generate, variants, compare, serve, etc.)
|
||||
│ ├── test/ # Integration tests
|
||||
│ └── dist/ # Compiled binary
|
||||
├── extension/ # Chrome extension (side panel + activity feed + CSS inspector)
|
||||
├── lib/ # Shared libraries (worktree.ts, egress-receipt.ts, context-bill.ts, redact-engine.ts, tracker-guard.ts, version-source.ts, code-intelligence/)
|
||||
├── patches/ # bun `patchedDependencies` patches (playwright-core windowsHide)
|
||||
├── docs/designs/ # Design documents
|
||||
├── setup-deploy/ # /setup-deploy skill (one-time deploy config)
|
||||
├── .github/ # CI workflows + Docker image
|
||||
│ ├── workflows/ # evals.yml (E2E on Ubicloud), quality-gate.yml (secret scan), dependency-review.yml, osv-scanner.yml, skill-docs.yml, actionlint.yml, and 8 more (windows, periodic evals, release gates, ci-image)
|
||||
│ └── docker/ # Dockerfile.ci (pre-baked toolchain + Playwright/Chromium)
|
||||
├── contrib/ # Contributor-only tools (never installed for users)
|
||||
│ └── add-host/ # /gstack-contrib-add-host skill
|
||||
├── setup # One-time setup: build binary + symlink skills
|
||||
├── SKILL.md # Generated from SKILL.md.tmpl (don't edit directly)
|
||||
├── SKILL.md.tmpl # Template: edit this, run gen:skill-docs
|
||||
├── ETHOS.md # Builder philosophy (Boil the Ocean, Search Before Building)
|
||||
└── package.json # Build scripts for browse
|
||||
```
|
||||
@@ -0,0 +1,47 @@
|
||||
# Slop-scan: what to fix, what to leave
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction). Read before acting
|
||||
on any slop-scan finding.
|
||||
|
||||
### What to fix (genuine quality improvements)
|
||||
|
||||
- **Empty catches around file ops** — use `safeUnlink()` (ignores ENOENT, rethrows
|
||||
EPERM/EIO). A swallowed EPERM in cleanup means silent data loss.
|
||||
- **Empty catches around process kills** — use `safeKill()` (ignores ESRCH, rethrows
|
||||
EPERM). A swallowed EPERM means you think you killed something you didn't.
|
||||
- **Redundant `return await`** — remove when there's no enclosing try block. Saves a
|
||||
microtask, signals intent.
|
||||
- **Typed exception catches** — `catch (err) { if (!(err instanceof TypeError)) throw err }`
|
||||
is genuinely better than `catch {}` when the try block does URL parsing or DOM work.
|
||||
You know what error you expect, so say so.
|
||||
|
||||
### What NOT to fix (linter gaming, not quality)
|
||||
|
||||
- **String-matching on error messages** — `err.message.includes('closed')` is brittle.
|
||||
Playwright/Chrome can change wording anytime. If a fire-and-forget operation can fail
|
||||
for ANY reason and you don't care, `catch {}` is the correct pattern.
|
||||
- **Adding comments to exempt pass-through wrappers** — "alias for active session" above
|
||||
a method just to trip slop-scan's exemption rule is noise, not documentation.
|
||||
- **Converting extension catch-and-log to selective rethrow** — Chrome extensions crash
|
||||
entirely on uncaught errors. If the catch logs and continues, that IS the right pattern
|
||||
for extension code. Don't make it throw.
|
||||
- **Tightening best-effort cleanup paths** — shutdown, emergency cleanup, and disconnect
|
||||
code should use `safeUnlinkQuiet()` (swallows ALL errors). A cleanup path that throws
|
||||
on EPERM means the rest of cleanup doesn't run. That's worse.
|
||||
|
||||
### Utilities in `browse/src/error-handling.ts`
|
||||
|
||||
| Function | Use when | Behavior |
|
||||
|----------|----------|----------|
|
||||
| `safeUnlink(path)` | Normal file deletion | Ignores ENOENT, rethrows others |
|
||||
| `safeUnlinkQuiet(path)` | Shutdown/emergency cleanup | Swallows all errors |
|
||||
| `safeKill(pid, signal)` | Sending signals | Ignores ESRCH, rethrows others |
|
||||
| `isProcessAlive(pid)` | Boolean process checks | Returns true/false, never throws |
|
||||
|
||||
### Score tracking
|
||||
|
||||
Baseline (2026-04-09, before cleanup): 100 findings, 432.8 score, 2.38 score/file.
|
||||
After cleanup: 90 findings, 358.1 score, 1.96 score/file.
|
||||
|
||||
Don't chase the number. Fix patterns that represent actual code quality problems.
|
||||
Accept findings where the "sloppy" pattern is the correct engineering choice.
|
||||
@@ -0,0 +1,43 @@
|
||||
# Testing internals: env keys, hermetic E2E
|
||||
|
||||
Moved verbatim from CLAUDE.md (token-load reduction). Read this before
|
||||
writing or debugging E2E tests, passing `env:` to a runner, or touching
|
||||
`test/helpers/hermetic-env.ts`.
|
||||
|
||||
**Env keys in Conductor workspaces.** The `GSTACK_*` env-shim (v1.39.2.0+,
|
||||
`lib/conductor-env-shim.ts`) promotes `GSTACK_ANTHROPIC_API_KEY` /
|
||||
`GSTACK_OPENAI_API_KEY` to their canonical names inside gstack's TS binaries.
|
||||
Tests run through gstack entrypoints inherit this promotion automatically.
|
||||
Don't echo the key value to stdout, logs, or shell history. The historical
|
||||
"never pass `env:` to `runAgentSdkTest`" rule is retired: the failure was
|
||||
partial-env replacement (the SDK's `Options.env` REPLACES the child's entire
|
||||
environment, so an object without the key broke auth). The runner now always
|
||||
passes a COMPLETE hermetic env with per-test `env:` merged last, so per-test
|
||||
overrides are safe; ambient `process.env.ANTHROPIC_API_KEY` mutation also
|
||||
still works (the env builder reads process.env at call time).
|
||||
|
||||
**Hermetic local E2E (default).** Every E2E runner (claude -p, PTY, Agent
|
||||
SDK, codex, gemini) spawns children through `test/helpers/hermetic-env.ts`:
|
||||
allowlist-scrubbed env (operator `CONDUCTOR_*`, `CLAUDE_*`, `GSTACK_*`,
|
||||
`MCP_*`, `GBRAIN_*`, and credentials like `GH_TOKEN` never reach children),
|
||||
a fresh seeded `CLAUDE_CONFIG_DIR` (no operator `~/.claude` CLAUDE.md /
|
||||
MCP servers / skills), a temp `GSTACK_HOME`, and `--strict-mcp-config`.
|
||||
Local eval signal matches CI. Debug against real operator state with
|
||||
`EVALS_HERMETIC=0` (restores the legacy env AND drops the strict-MCP flag).
|
||||
Per-test `env:` overrides merge last, so deliberate contamination
|
||||
(`CONDUCTOR_WORKSPACE_PATH`, per-test `GSTACK_HOME`) keeps working. The
|
||||
hermetic config dir seeds NO skills by default; a PTY test that types a
|
||||
`/skill` slash command must pass `seedSkills: true` to the PTY runner, which
|
||||
points the child's `CLAUDE_CONFIG_DIR` at `hermeticSkillsConfigDir()` — a
|
||||
seeded registry that symlinks the LIVE working tree's SKILL.md files (by
|
||||
design: the skills ARE the subject under test; a snapshot would measure stale
|
||||
copies). Wiring is pinned by `test/hermetic-wiring.test.ts` (static tripwire),
|
||||
two gate-tier canaries in `test/skill-e2e-hermetic-canary.test.ts`, and the
|
||||
seeding tripwires in `test/hermetic-skills-seeding.test.ts` /
|
||||
`test/pty-skill-seeding-wiring.test.ts`.
|
||||
|
||||
E2E tests stream progress in real-time (tool-by-tool via `--output-format stream-json
|
||||
--verbose`). Results are persisted to `~/.gstack/projects/<slug>/evals/` (legacy
|
||||
fallback `~/.gstack-dev/evals/`) with auto-comparison
|
||||
against the previous finalized run (in-flight `_partial` files are never used as
|
||||
a baseline, so a run can't compare against itself).
|
||||
Reference in New Issue
Block a user