b1485d8897 v1.74.0.0 test/CI overhaul: green means green, suites restructured for speed (#2721)
* fix(ci): free-tests lane actually runs the make-pdf e2e gates

The 9 make-pdf/test/e2e gate tests probe make-pdf/dist/pdf,
browse/dist/browse, and the diagram-render bundle, then self-skip when
absent. The required free-tests lane never built any of them, so the
gates silently skipped on Linux for their entire life (verified: 9 of
14 skip, exit 0). make-pdf-gate.yml's justification for deleting its
Linux leg claimed the free lane covered this — it didn't.

- new build:gates script: exactly the three artifacts the gates probe
  (full bun run build compiles five binaries; ~60-90s tax on the only
  required check is not warranted)
- free-tests.yml: build:gates step + poppler-utils +
  fonts-noto-color-emoji (fonts must precede the first browse daemon
  launch — Chromium snapshots fontconfig at startup; verified live:
  a warm daemon renders tofu, a fresh one embeds NotoColorEmoji)
- make-pdf/test/e2e/ci-prereqs.test.ts: GSTACK_EXPECT_BINARIES=1 (set
  by the workflow) inverts the skip polarity in CI — dropping the
  build step or poppler fails the lane instead of re-opening the
  silent-skip hole

Pre-flight: all 9 gates green on Linux locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): kill the three zero-test eval jobs (hollow green)

- delete the vestigial e2e-codex / e2e-gemini matrix rows: both files
  are whole-file periodic-tier, so with no row tier: they ran ZERO
  tests and reported green on every PR (~2 min of runner each, pure
  false confidence; the periodic lane owns those suites)
- e2e-pty-plan-smoke gains tier: gate — its two files are whole-file
  describeE2ETier('gate'), so the job burned ~7 min of container setup
  then skipped every describe
- KNOWN_TIER_UNSET burned down to empty; the ratchet stays armed so a
  future row/file tier mismatch fails the suite instead of shipping
  hollow green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): least-privilege permissions + fork-safe concurrency keys

- evals.yml / evals-periodic.yml evals jobs: explicit contents:read +
  packages:read (container-image pull) and persist-credentials:false —
  the jobs that execute PR-authored code with three provider API keys
  ran on the repo-default token grant with the token written into
  .git/config
- permissions blocks for the 4 workflows that had none (skill-docs,
  make-pdf-gate, windows-free-tests, windows-setup-e2e)
- fork-safe concurrency keys: actionlint, skill-docs, make-pdf-gate,
  windows-setup-e2e switch from head_ref to PR-number keying — a bare
  branch name carries no fork prefix, so same-name branches from two
  forks shared one group and cancelled each other's runs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): one bun version everywhere + drift tripwire

Lanes disagreed four ways: 1.3.13 (free-tests, windows, Dockerfile.ci),
latest (quality-gate, make-pdf-gate), unpinned (skill-docs,
version-gate — setup-bun installs latest), 1.3.10 (.gitlab-ci.yml).
Different Bun versions change the runner output shapes the strict
classifiers regex-match, spawn semantics, and shell parsing — a lane on
a different Bun tests a different product; Dockerfile.ci's own comment
records this class biting once already (silent 1.3.13/1.3.14 drift).

All surfaces pinned to 1.3.13; test/bun-version-drift.test.ts scans
every workflow setup-bun stanza + Dockerfile.ci + .gitlab-ci.yml and
fails on any mismatch or unpinned stanza. skill-docs also gains
--frozen-lockfile (was bare bun install).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): bind the three-way image-tag hashFiles() expressions

evals.yml, evals-periodic.yml, and ci-image.yml each compute the CI
image tag from hashFiles('.github/docker/Dockerfile.ci', 'bun.lock',
'patches/**') — synced by comment only (TODOS.md 'CI three-way
image-tag drift'). If one input list drifts, that workflow computes a
different tag for the same content: eval lanes silently rebuild the
image every run, or ci-image prebuilds a tag nobody looks up. The test
extracts each tag-computation site and fails on any mismatch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): ci-image stops rebuilding the identical image every ship

- package.json out of the trigger paths: the tag hash deliberately
  excludes it (version bumps every ship), so every merge rebuilt and
  re-pushed the IDENTICAL tag (~2m26s for zero content change);
  patches/** added (it IS a tag input)
- manifest existence check (mirrors evals.yml): tag already exists →
  skip the build
- concurrency group: two rapid main pushes raced pushing the same
  :latest/:buildcache tags
- cron staggered 06:00→04:00 Monday: it shared the exact minute with
  evals-periodic, which could race a half-pushed tag or duplicate the
  build
- timeout-minutes: 30 (was unbounded → 360-min default for a hung
  docker build)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): quality-gate drops the 74s full-history checkout

fetch-depth:0 cost 74 of the job's 92 seconds; the three gates it feeds
take ~12s combined. Shallow checkout + exact-SHA fetches for the diff's
base/head (an exact-SHA fetch, not a guessed depth — long-lived
branches and merge queues still resolve), with a --deepen fallback for
push events whose 'before' is unusable. timeout right-sized 20→10 min.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): small-lane batch — timeouts, right-sizing, windows cache warm-start

- timeout-minutes on the 6 remaining unbounded jobs (actionlint 5,
  skill-docs 10, version-gate 10, make-pdf-gate 15, pr-title-sync 5,
  evals build-image 15) — a hung step sat on GitHub's 360-min default
- right-size measured-over-long timeouts: dependency-review 10→5,
  windows-setup-e2e 15→10
- dependency-review: 2-core runner (28s API call on an 8-core box) and
  drop .github/workflows/** from its trigger paths (workflow edits have
  no dependencies to review)
- windows caches gain restore-keys: a lockfile bump paid the 26s/43s
  restore for a guaranteed cold miss

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): scope GSTACK_HOME to each file's execution window

Five files assigned process.env.GSTACK_HOME at module scope. Shard
processes evaluate sibling modules before running their tests, so the
assignment leaked into every other file in the shard — the damage was
already visible in defensive workarounds (relink.test.ts:28 'fresh
install test saw a neighbor's skill_prefix'; cdp-e2e's own comment
documents a sibling's temp dir baked into artifacts).

Pattern: save original, assign in beforeAll, restore in afterAll
(cdp-e2e already restored but still assigned at load — its window now
matches the others). GSTACK_TELEMETRY_OFF and GSTACK_PROJECT_SLUG get
the same treatment where they rode along. Victim files' defenses stay
in place (cheap insurance).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: tripwire against module-scope GSTACK_HOME assignments

Column-0 assignment of GSTACK_HOME / GSTACK_STATE_ROOT in any tracked
*.test.ts fails with the file:line and the fix (beforeAll + afterAll
restore). Kills the cross-file env-leak class the previous commit
swept.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): e2e-harness-audit derives its skill census from disk

The hand-maintained 39-name SKILL_GLOBS list had drifted to 39 of 54
SKILL.md.tmpl on disk. No live gap today (none of the 15 unlisted
skills is interactive), but the next interactive skill would have
landed unguarded with zero signal. The audit now walks top-level dirs
for SKILL.md.tmpl (statSync so symlinked dirs like connect-chrome
count), so new skills are in scope the commit they appear.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): judges honor the eval-model resolution chain + real 429 backoff

callJudge inlined GSTACK_EVAL_MODEL_JUDGE || sonnet, silently ignoring
the global GSTACK_EVAL_MODEL override every other eval call site honors
via lib/eval-model.ts. New 'judge' kind in DEFAULTS (sonnet — the D1a
pin-on-regressors calibration stands; model CHOICE unchanged) and
callJudge resolves through it: explicit arg > GSTACK_EVAL_MODEL_JUDGE >
GSTACK_EVAL_MODEL > default.

429 handling upgraded from one fixed 1s retry (reliably lost races at
CI concurrency) to three jittered exponential retries (~1s/4s/16s),
honoring the server's retry-after when present.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): the two expect(true) paid stubs become test.todo

skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s)
reported PASS on every periodic run while asserting nothing. Deleting
them would remove the periodic-tier selector surface they exist to
register (diff-based selection for spec/ changes), so they become
test.todo — reported as todo/skip, never pass — with the v1.1
implementation specs kept in-file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): reactivate 5 quarantined browse tests (2 security)

extension-sender-auth's two privileged-message denial tests (content
script + missing sender.url — the extension's security boundary) and
snapshot's three skips were quarantined 'pre-existing' failures. Root
cause: machine-local state on the quarantining dev machines — the test
and gate code are byte-identical between the quarantining commit
(410b4928) and HEAD, and all five pass deterministically on a clean
checkout (68/68 across both files, multiple runs). No assertions
weakened, no product changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): activate the 4 paid test files that could never run anywhere

carve-section-loading, codex-e2e-plan-format,
codex-e2e-recommendation-substance, and llm-judge-recommendation gated
on EVALS/tier (free suite loads them as describe.skip) but their names
fell outside PAID_TEST_GLOBS, so no paid lane ever selected them — net
execution zero, forever. The existing matrix tripwire filtered on
isPaidTestFile() first, so it was blind to exactly this class (the same
bug that hid the pre-split monolith's gate tests for ~8 releases).

- PAID_TEST_GLOBS: codex-e2e* + skill-llm-eval* wildcards (replacing
  exact names) + llm-judge-recommendation + carve-section-loading;
  package.json's six test-script glob lists mirrored
- codex-e2e-plan-format gains the explicit periodic tier gate its
  siblings carry (external-service rule) — without it the sharded
  runner's no-guard default would spawn Codex in the gate tier per PR
- eval:bg:periodic --timeout 32400→37800: the census growth pushed the
  periodic worst case to 35910s; the old value had 270s of headroom
  BEFORE this change and would now kill healthy runs mid-flight
- new test/paid-orphan-tripwire.test.ts: any EVALS/tier-gated test file
  outside the globs fails the free suite (reasoned SCANNER_EXEMPT for
  the gate helpers + meta-tests) — the class-killer
- paid-shards pins updated: the four orphans now assert INSIDE the
  census

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): restrictDirectoryPermissions warns and skips symlinked dirs

Closes the Windows Free Tests red: recent lane failures showed a
platform-unguarded POSIX mode-bit assertion ('Expected: 493' — a
symlink-skip test) from PR-branch variants; the KNOWN_WINDOWS_SAFE
force-include reason ('mode-bitmask hits are POSIX-branch only') did
not hold for that shape, and main had neither the guard nor the
behavior.

- product: lstat first; a symlinked dir gets a warning and a skip on
  both platforms — chmod AND icacls dereference the link, so
  restricting through a symlink hardens an unvetted target (and
  /inheritance:r could lock out its real owner). All callers already
  treat hardening as best-effort (try/catch).
- test: the symlink regression test, platform-aware — symlinkSync in
  the house try/catch skip pattern (Windows runners without Developer
  Mode can't create symlinks), mode-bit assertion guarded off win32,
  behavior assertions (no throw, warning text, target readable)
  everywhere; POSIX still proves the skip (0o755 unchanged, not 0o700)
- KNOWN_WINDOWS_SAFE reason updated to the now-true premise

20/20 pass on Linux.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): unique tmp dirs for plan artifacts + audited live-repo cwd sites

Six paid PTY tests wrote their expected plan artifact to a FIXED shared
/tmp path ('/tmp/gstack-test-plan-<mode>.md') and rmSync'd it in
finally — under --retry 1, EVALS_JOBS>1, or two concurrent worktrees, a
sibling's cleanup deletes this run's artifact and the D19 'agent did
not produce expected plan file' assertion fires spuriously. Each test
now mkdtemps its own dir, interpolates the unique path into the agent
prompt (fixture-sourced prompts get a replaceAll + drift guard that
throws if the fixture's literal ever moves), and cleans up its own dir.

The 18 cwd:-into-the-live-repo sites were audited: all deliberate
(skill registry + hermetic pre-trusted dir, in-repo gen renders, git
history reads, slug resolution) — each now carries a
'// LIVE-REPO CWD: <reason>' comment so the next audit can tell
deliberate from accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): trim the seven over-wall 1700s timeouts to the 1500s physical ceiling

1,700,000ms (28.3 min) exceeded every wall these tests run inside: the
25-min CI job timeout and the 1800s sharded-runner wall (which also
leaves --retry 1 zero room for a second attempt). Budget above the wall
is fiction, not headroom — a test that actually used it produced a
job-level kill (no bun summary, no artifact) instead of a clean
per-test timeout. No recorded p95 exists for this family (they are
being retiered to periodic in the re-platform wave); the trim stops at
the physical ceiling rather than guessing lower. Final policy lands in
the Wave-2 eval-budgets constants module.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(gen): main() guard — importing gen-skill-docs no longer regenerates the tree

The generator's whole body executed at module load, so any import of it
(test/gen-skill-docs.test.ts pulls assertSinglePreamble via require();
test/catalog-trim.test.ts imports helpers) regenerated all 71 SKILL.md
in place — the root cause of half the TREE_MUTATING serial-shard
entries (hazard class #2532). The body now lives in an exported
main(): number behind if (import.meta.main).

Semantics preserved exactly: failure exits are immediate (matching the
old top-level process.exit), success leaves the event loop to drain so
the llms.txt fire-and-forget IIFE finishes its write, and the module
stays synchronous/require()-able. Proofs: byte-identical --host all
output (git status clean), --dry-run stale-tree still exits 1 (the
skill-docs freshness lane depends on it), and the new
test/gen-skill-docs-import-purity.test.ts pins load-time purity via a
subprocess probe (mtime-based, so a dirty worktree can't false-fail).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): --out-dir renders every host, outputs-only

--out-dir was Claude-host-only (gen-skill-docs.ts:842), which forced
the codex/factory-regenerating tests (gen-skill-docs, skill-validation,
host-config) to mutate the live tree — the reason they sit in the
TREE_MUTATING serial shard. The flag now mirrors ALL outputs into the
out-dir: external-host trees (.agents/.factory/... via
processExternalHost), external section files, openclaw docs, and
gstack/llms.txt (a catalog-mode render must never rewrite the tracked
index). OUTPUTS ONLY — inputs (templates, sections/, host configs) are
always read from ROOT, so an empty out-dir can never feed the render.
rewriteSectionBase stays Claude-only (external hosts have their own
path grammar).

Proofs: in-place --host all is byte-identical (tree clean);
--host all --out-dir <mkdtemp> renders the full multi-host tree with
ROOT untouched; gen-skill-docs-out-dir tests + 415/415
gen-skill-docs.test.ts green (bin/dev-setup's claude rendering
byte-compat).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evals): every E2E key's dep list names its own declaring test file

129-of-177 keys omitted their own test file, so editing only a test's
prompt or assertions selected NOTHING — the changed test never ran on
the change that changed it. 135 keys self-registered (110 E2E + 25
LLM-judge), resolved by strict declaration evidence (testName:/
testIfSelected/judge call sites), with skill-name false positives
excluded.

e2e-tier-alignment's warn-only branch for unregistered files is now a
hard failure with a 4-entry KNOWN_UNREGISTERED ratchet (template-
literal testNames, fail-open-safe) + a burn-down test so the set only
shrinks. Selection sanity: a one-file diff on skill-e2e-qa-workflow now
selects its 4 tests (was 0); skill-llm-eval 0 → 25.

Known follow-ups (filed): 15 E2E + 2 judge PHANTOM keys select tests
that exist nowhere; codex-e2e-plan-format's testIfSelected names have
no map keys (run-all only).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(evals): ratchet the 8 newly-visible gate-matrix gaps

The self-registration sweep made these eight files' gate-tier keys
visible to the census for the first time — their gate tests run in NO
CI lane today (pre-existing hole, newly measurable). Ratcheted into
KNOWN_MATRIX_GAPS with the burn-down note: the paid-lane re-platform
runs every gate file by construction and retires this ratchet class.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): duration-aware LPT shard packing for the free suite

Hash sharding balances file COUNTS (1.15x spread) but not cost — the
Playwright-launching files landed 4/3/4/1/2/1 across 6 shards, giving a
measured 28s–97s shard spread and ~40s of idle tail on every run.
Full-suite mode now packs by recorded per-file durations
(longest-processing-time-first) when the committed seed
scripts/free-test-durations.json exists.

- ONE store, no overlay: the seed is refreshed occasionally via the new
  --record-durations mode (each file timed in its own child — exact,
  and immune to bun's stream buffering, where silent passers print no
  header to timestamp); GSTACK_FREE_TEST_DURATIONS overrides the path
  for experiments; CI never records
- seed is a hint: missing → silent hash-shard fallback; corrupt (bad
  merge) → one warning + fallback; unknown files → 75th-percentile
  pessimism so a surprise long-runner can't recreate the tail
- packed shards get duration-aware walls (max(base, predicted x 3)) —
  LPT decouples count from cost BY DESIGN, so the 5s/file heuristic
  would undersize a shard holding few expensive files
- one log line per shard (files + predicted seconds) so packing
  regressions are diagnosable from any run log
- the --shard CI-matrix path is untouched: stable hash indices are its
  contract
- successor note in-code: bun >=1.3.14 ships native --timings/--shard
  LPT — swap this packer when the repo unpins 1.3.13

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): decouple slop:diff from bun run test; quality-gate runs it per PR

'bun run test' silently appended up to two 120s npx slop-scan runs plus
a git worktree add/remove after the suite (2>/dev/null || true) —
invisible in the documented '~90-100s' timing and pure friction in the
pre-commit loop. Decoupling is not coverage removal: quality-gate.yml
now runs slop:diff on every PR (advisory, matching its in-repo 'never
blocking' contract), and /review already invokes it explicitly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): eval-budgets timeout tiers + fit/ceiling policy test

Five named tiers (JUDGE 120s / CAPTURE 300s / CAPTURE_LONG 600s /
PTY 900s / PTY_LONG 1200s) replace hand-ratcheted sprawl (46x300s,
46x120s, 44x360s, 44x180s, 27x240s, 19x150s, 13x420s, 12x600s...),
much of it inflated to paper over the old 40-way in-shard concurrency
that the sharded runner's 1-file-per-shard model kills. Policy test
pins: every tier fits the shard wall minus 120s overhead (the
structural fix for budgets-above-the-wall fiction), tiers stay ordered,
and no paid literal exceeds PTY_LONG x1.25 — oversized tests get split,
not budgeted past the wall.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): shared runBin helper for bin-script unit tests

~36 free test files each carry a near-identical local run() (spawnSync
+ utf-8 + {status, stdout, stderr}) differing only in env composition,
cwd, and timeout. runBin absorbs the invariant core; options carry the
variance (gstackHome sets BOTH GSTACK_HOME and GSTACK_STATE_DIR — the
config-precedence trap several locals rediscovered independently; home
for $HOME-anchored bins; input/trim/timeout/maxBuffer). Free-test-only
by design so it never becomes a de facto global touchfile. Migration of
the 36 call sites lands separately (mechanical batches).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): runBin trim assertion — trim shapes stream ends, not interior

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(test): mechanical sweep — 298 paid-test timeouts onto eval-budget tiers

69 files, both shapes (trailing bun-test budgets and runner
timeout/timeoutMs options), ROUND-UP ONLY so nothing that passed can
start failing: 75 → JUDGE_MS, 137 → CAPTURE_MS, 74 → CAPTURE_LONG_MS,
9 → PTY_MS, 3 → PTY_LONG_MS. Raw >=60s literal count in the paid scope:
395 → 97, of which 51 are non-timeout noise (fixture dates, run IDs)
and 46 are enumerated justified holds (comment-carrying calibrated
budgets, poll-loop constants, utility spawn waits, and the seven
physical-ceiling 1_500_000 sites). The eval-budgets policy ratchet
keeps the residue from regrowing.

Known collapse: where an inner runner budget and its enclosing test
budget now share a tier, the old stagger is gone — an overrun surfaces
as a bun test timeout instead of a graceful runner timeout
(diagnosability trade, not a correctness one).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: coverage fill — 95 tests for six zero-coverage surfaces

- eval CLI family (eval-list/compare/summary + eval-select smoke): the
  primary interface to eval results had no tests; isolation via a fake
  gstack-slug under a mkdtemp HOME (the scripts' real resolution path —
  they do NOT honor GSTACK_EVAL_DIR; only EvalCollector does). Pinned
  current behavior: eval-list does NOT exclude _partial runs (documented
  improvement candidate)
- slop-diff (runs on every /review + quality-gate): fixture git repo +
  first-on-PATH npx stub (never downloads real slop-scan); no-diff
  early exit, missing-scanner fallback, fingerprint line-insensitivity,
  merge-base worktree scan
- bin/gstack-code-intelligence CLI arg surface (lib was covered, the
  284-line CLI wasn't): select/consent/suggest/index/search gating;
  pinned: --help routes to usage failure exit 1 (no handler)
- browse media-extract: the page.evaluate callback exercised in-process
  against a mock DOM (no exports added) — lazy-src fallback chain,
  HLS/DASH detection, bg-image url() parsing, 500-element cap
- browse session-cookie-store: factory contract (cookieName/ttlMs/
  maxSessions eviction, cross-store isolation, mint→validate
  round-trip); store is in-memory — no fs cases exist
- lib/version-source direct unit tests (gstack-version-bump.test.ts
  spawns the bin, never imports the lib): parse/format/cmp/bump
  coercion, npm 4→3 translation, #2501 mangled-JSON regression class

All hermetic (mkdtemp homes, runBin child isolation); windows curation
correctly partitions the six.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(test): first runBin migration batch (3 of ~36 run() duplicates)

explain-level-config, benchmark-cli, evidence move onto the shared
helper; each file's remaining special-case spawnSync sites (raw-buffer
probes, env-scrub probes) stay put deliberately. 55/55 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(evals): paid shards spool to disk + shared runShardChild lifecycle

- runPaidShard no longer buffers whole 30-min stream-json streams in
  RAM (x concurrent jobs): every byte tees to a per-shard log file
  (slug-named, path printed at START for mid-run inspection and on the
  FAILED terminal line); failures print a 64KiB tail read back from
  disk; passing shards stay quiet (the file is the record) — the free
  runner's proven contract. Classification unchanged: the strict
  classifier still sees every byte first.
- the ~35 duplicated spawn/group-kill/wall-timer/finally-reap lines
  move into runShardChild in test-strict-output.ts (detached-per-
  platform spawn, signal forwarding, SIGKILL group kill at the wall,
  drain-before-verdict); designed so the free runner can migrate later
- expectedFiles drift fixed toward ENFORCEMENT: the injected-command
  exemption is gone — a fake command exiting 0 without bun's terminal
  summary now reads FAILED (pinned: silent-pass → failed)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(evals): parent-computed selection propagates to shard children

The sharded runner computed diff selection once, then each of its 48-73
children recomputed it at module load — including, on touchfiles-diff
branches, a per-child bun subprocess evaluating the old data file (20s
timeout each). The parent now serializes {version, selected, reason} as
EVALS_SELECTION_JSON into the shard env; e2e-helpers adopts it at load.
Fail-open preserved: any parse/shape violation → ONE stderr warning +
local recompute; absent env → silent local compute (non-sharded
entrypoints unchanged). Drift test pins parent→child round-trip to
identical selection decisions plus the malformed/absent cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): kill the four worst fixed sleeps (300s/30s/30s/20s)

- watchdog.test: the 20s blind wait for one production parent-watchdog
  tick becomes BROWSE_PARENT_WATCHDOG_INTERVAL_MS=250 (new env knob in
  server.ts, NaN-safe, production default unchanged) + polls for the
  boot line and the tick's stay-alive log — strictly stronger (the old
  form never proved a tick observed the parent death). 24s → 3.6s.
- stop-dead-daemon / terminal-agent-owner-watchdog: the 300s/30s
  stand-in child lifetimes become stdin-EOF-bound — the child can never
  self-exit mid-test on a slow runner (spurious-failure class) and
  self-reaps instantly if the test dies (no 300s orphans). Node-compat
  stdin APIs (owner-watchdog runs on the Windows lane).
- browser-skill-commands: the sleeper fixture's 30s self-time becomes
  8s (no stdin pipe exists in runToFiles) — far above the 1s product
  timeout it must outlive, below the test ceiling, so a timeout-kill
  regression fails on clean assertions instead of an opaque bun
  timeout; added: stdout must NOT contain 'done'.

45/45 green across the four files + server tripwires.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): gen-skill-docs + catalog-trim leave the serial mutator shard

gen-skill-docs.test.ts's 15 in-place generator spawns now render into
mkdtemp out-dirs (gitignored-artifact reads repointed; the handshake
scan's silent console.warn degrade became a hard assertion); its
tracked-tree reads (freshness dry-run, SKILL.md content pins) stay
reads. catalog-trim needed no change beyond the earlier main() guard —
its import is now side-effect-free (pinned by the import-purity test).
Both TREE_MUTATING entries deleted in this commit, per the transition
rule: an entry leaves in the same commit as the file's last in-place
write.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): skill-validation renders codex host into an out-dir

Its 3 in-place --host codex regeneration sites collapse into one
module-level --out-dir render; assertions untouched. TREE_MUTATING
entry deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): host-config self-provisions goldens (ordering dependency severed)

Its goldens were 'produced by gen-skill-docs.test.ts' with a
when-missing beforeAll fallback that wrote the live tree — an
inter-test ordering dependency the serial shard hid. It now renders
codex+factory UNCONDITIONALLY into its own out-dir and reads goldens
only from there (the Claude golden deliberately keeps reading tracked
ship/SKILL.md — a read; out-dir claude renders repoint section-base
paths by design). TREE_MUTATING entry deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): gbrain-detection-override drops mutate-then-git-restore

regenAndSnapshot renders --host claude --out-dir <mkdtemp> (+
--respect-detection) and snapshots probes from the out-dir. The
git-restore machinery is deleted outright — it restored only
PROBE_FILES of the 71 files each call wrote, so a stale tree kept the
other 68 dirty (the partial-restore bug), and its 'no output-path arg'
comment had been false since --out-dir landed. TREE_MUTATING entry
deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): catalog-mode-full renders to out-dir; restore machinery deleted

The full-catalog smoke no longer rewrites all 71 SKILL.md then
regenerates to restore (with its 'CRITICAL: failed to restore' prayer
path) — it renders into a mkdtemp and additionally asserts tracked
ship/SKILL.md is byte-unchanged. TREE_MUTATING entry deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): idempotency proof strengthens to two-out-dir recursive diff

Two renders into two separate out-dirs, EVERY file diffed byte-for-byte
(claude-only and --host all; normalization only for each dir's own
sanctioned section-base repoint; presence-sanity lists guard against a
vacuous empty-dir pass) — strictly stronger than the old in-place
double-regen that sampled 5 files. TREE_MUTATING entry deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): spec-template-sync compares an out-dir render, not an in-place one

TREE_MUTATING entry deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): the serial tree-mutating shard dissolves — TREE_MUTATING is empty

Zero mutators remain (all eight render into out-dirs now), so the four
ratchet READERS (parity caps, size budgets, carve parity/ordering) get
a quiet tree by construction in any shard and rejoin the parallel
phase. The ~35-40s serial tail on every full-suite run is gone. The
mechanism stays: a future test that genuinely must write shared
artifacts in place earns an entry with a reason and is serialized
again; the census pin still fails on renamed keys.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(gen): out-dir byte-identity + tree-clean pins for external hosts

codex render: porcelain unchanged AND out-dir gstack-ship/SKILL.md
byte-identical to a fresh in-place render (+openai.yaml presence);
--host all render: exit 0, porcelain unchanged, claude + .agents +
.factory + llms.txt + openclaw docs all present in the out-dir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): commit the initial free-test durations seed (496 files)

Recorded via --record-durations on a quiescent tree: 479s serial
total, p50 92ms / p90 1.8s / max 31.4s — the top-heavy cost shape LPT
packing exists for. A hint, not a contract: refresh opportunistically
with bun run test:free --record-durations.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(evals): planner/executor/report modes — the CI re-platform surface

One PLANNER computes diff selection + the slice plan ONCE and writes a
manifest (--emit-plan <path> --slices K); K executors consume it
(--plan <path> --slice i), never self-selecting, and write slice-result
artifacts; a REPORT reconciles results against the manifest (--report
<dir>) fail-closed: a slice whose artifact never landed is a FAILURE,
a planned shard nobody reported fails, wrong-slice/duplicate/cross-tier
results fail. Kills per-slice selector divergence and hollow-lane
aggregation at the root.

- hollow-shard guard: under EVALS_ALL, exit 0 with ZERO executed tests
  (bun's 'Ran N tests' now captured by the classifier — additive) is
  'passed-empty' and fails the run; selective runs keep it 'passed'
  with one warning (in-file diff/tier self-skips are legitimate there);
  unknown counts are never guessed hollow
- retry parity: --retry 1 default + RETRY_OVERRIDES literals for the
  three files whose old matrix rows earned retries: 2 (stale entries
  pinned against disk)
- live smoke: gate plan = 48 shards across 6 slices; report mode exits
  1 on a fabricated missing slice, 0 when complete

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): sliced paid lane (planner -> 6 executors -> fail-closed report)

The parity-phase re-platform: evals.yml gains a second, sliced lane
driven by scripts/test-paid-shards.ts — the SAME engine local
eval:bg:gate uses, so CI and local share one selection engine.

- plan-slices: ONE planner (fetch-depth 0 — the only job needing
  history) emits the manifest; selection fails open to run-all, never
  per-slice (the divergence class is structurally dead)
- eval-slices: 6-way matrix consuming the manifest; PTY seed +
  skill-registration steps run unconditionally (idempotent — a sliced
  lane cannot key them on suite names); aggregate spawn budget
  6 x EVALS_JOBS=2 x EVALS_CONCURRENCY=2 = 24 lane-wide (the matrix's
  40-way per row queued session startup behind 39 siblings — the
  timeout-flake family root); slice results + spooled shard logs
  uploaded as artifacts
- slices-report: reconciles slice artifacts against the manifest
  FAIL-CLOSED via --report — a slice whose artifact never landed, or a
  planned shard nobody reported, is a failure, not an absence
- sequenced needs: evals so provider concurrency never doubles while
  both lanes coexist; the matrix + its ratchets are deleted after
  demonstrated parity (intersection + expected-additions comparison)
- workflow_dispatch gains evals_all (default true) for parity runs and
  post-merge smokes — a dispatch can never silently select zero

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): weekly periodic lane runs EVERY periodic test + gate census backstop

evals-periodic.yml re-platforms onto the sharded runner: planner
manifest → 6 executor slices → FAIL-CLOSED report. This IS the coverage
contract: all ~70 periodic-tier files weekly (EVALS_ALL=1), killing the
silent-rot class where a hard-coded 9-file matrix left ~57 files
running NOWHERE (the autoplan E2E rotted invisibly for months).

- test/helpers/periodic-exclude-data.ts: reasoned exclusions in their
  OWN literals file (deliberately not touchfiles-data — map-diff
  evaluates old versions of that file standalone). Every entry carries
  reason + tracking with a re-entry condition; the runner surfaces each
  exclusion per run; policy test pins real-file + non-empty fields.
  Initial: ship-idempotency + brain-privacy-gate (documented-red,
  never green) and skill-e2e-ios (manual hardware). The TODOS 'sidebar
  E2E trio' turned out already deleted — only tombstone tests remain.
- gate-census job: weekly EVALS_ALL gate-tier run — PR lanes are
  diff-billed, so without this the full gate census might never execute
  anywhere; with the hollow-shard guard it is a census-health check
  (exit 0 + zero executed tests fails), not just a test run.
- failure notification is a concrete gh issue UPSERT (one tracking
  issue, commented per red week — never issue-per-week spam), with
  issues:write scoped to the report job.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: TESTING_INTERNALS covers the 2026-08 runner overhaul

LPT-packed free suite + --record-durations, the emptied TREE_MUTATING
mechanism, the sharded paid runner as the single selection engine,
CI planner/executor/report with the fail-closed report and hollow-shard
guard, the weekly coverage contract + exclusions policy, and the
eval-budgets timeout tiers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(CLAUDE.md): testing prose matches the overhauled runners

- bun run test: duration-packed shards + --record-durations; the
  trailing serial tree-mutating shard no longer exists
- two-tier system: the sliced CI lanes (one engine local+CI), the
  weekly all-periodic coverage contract + exclusions, the gate census
- periodic detach timeout 32400 → 37800

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(TODOS): close the absorbed test-infra items, file the overhaul follow-ups

Closed with receipts: the periodic coverage contract (implemented as
full weekly coverage + exclusions), the eval-harness observability P1
(verified already landed: heartbeat, incremental _partial persistence,
live stderr + eval-watch), and the sidebar trio (already deleted —
tombstones remain). Filed: matrix deletion after parity, the
required-check maintainer decision, browse /tmp-namespace hardening,
PTY boot-readiness waits, the single typed test registry, bun-native
LPT swap, runBin/free-runner migrations, eval-list partial exclusion,
phantom key cleanup, duration-weighted slicing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v1.73.0.0: test/CI overhaul — green means green, suites restructured for speed

Version + release notes for the audit-and-overhaul branch: every
silently-skipping or never-running test class fixed and tripwired, the
free suite duration-packed with the serial mutator shard dissolved, the
paid lane re-platformed onto the sharded runner (planner/slices/
fail-closed report, parity phase), the weekly all-periodic coverage
contract, eval-budget timeout tiers, and 95 new coverage tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): first-live-run fixes — executor history + two environment-blind assertions

The sliced lane's first run (PR #2721) did its job: the planner and
report worked, the manifest governed, and every failure had a name.
Three were fixable on the spot:

- executor + gate-census checkouts get fetch-depth: 0 — files with
  SELF-derived selection (the LLM-judge map, routing) walk git at
  module load, and selection is deliberately fail-closed on git errors,
  so the shallow checkout crashed those shards ('ambiguous argument
  main...HEAD'). The manifest still governs WHICH shards run.
- landscape --toc gate: the exact toBe(3) landscape-page count was
  font-metric-dependent (3 on Amazon Linux, 2 on ubuntu CI — the same
  disease the file's own page-index comment warns about). Now a
  comparative invariant: --toc must not CHANGE the landscape count vs
  a baseline render.
- paid-run-manifest parse test builds its manifest under EVALS_ALL so
  it never walks git (proven with GIT_DIR=/nonexistent).

Remaining first-run failures are newly-exposed rot in gate files that
had never executed in CI (skillify D1 refusal, session-intelligence
context-restore, one tpa-apple-ban retry flake) — being probed
separately; they are the lane WORKING, not the lane failing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(TODOS): file the three first-execution findings from the sliced lane's live run

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v1.74.0.0: queue-advance — #2722 claims the v1.73.0.0 slot

The version gate caught a live queue collision (its whole job); same
MINOR bump level, next free slot per bin/gstack-next-version.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): per-shard CHROMIUM_PROFILE — the collision class duration packing exposed

Nine test files launch in-process persistent contexts or daemons that
default to the SHARED ~/.gstack/chromium-profile. Two concurrent shard
processes on one profile dir kill each other's browser — observed live
on CI once duration packing recomposed shards: handoff's
launchPersistentContext died 'Target page, context or browser has been
closed' (--user-data-dir=~/.gstack/chromium-profile in the call log)
while a sibling shard's daemon logged 'Chromium process crashed'. Hash
sharding had masked the collision by chance placement; handoff passes
standalone everywhere.

Fix at the runner, not per file: each shard child gets
CHROMIUM_PROFILE=<shard-state>/chromium-profile (the documented env
knob, same isolation idea as the existing per-shard TMPDIR). Files
within a shard run serially, so sharing the per-shard profile is safe;
config.test's resolution-order tests save/restore the env around their
assertions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): landscape --toc gate asserts promotion PRESENCE, not counts

Two rounds of CI receipts: the exact toBe(3) was font-metric-coupled
(3 on Amazon Linux, 2 on ubuntu), and the baseline-comparison repair
then failed 2-vs-3 across renders SECONDS apart in one CI job while the
sibling no-toc test saw 3 — per-render image-promotion timing makes any
count assertion here a coin flip. The sibling test owns exact promotion
counts; this test's actual invariant is that --toc does not break the
promotion machinery: >=1 landscape page + the TOC rendered. Also drops
the second render (halves the test's runtime).

Flaky per-render image promotion itself is worth its own look — noted
in TODOS with these receipts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(TODOS): file the per-render image-promotion nondeterminism (receipts from PR #2721)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): per-FILE Chromium profiles for the nine in-process launcher files

Completes the profile-isolation work: the per-shard CHROMIUM_PROFILE
stopped cross-shard kills; these nine files launch in-process
persistent contexts and could still collide with a lingering daemon a
sibling file spawned on the SAME shard profile. Each now scopes a
mkdtemp profile via beforeAll/afterAll (the module-scope-tripwire-safe
pattern), cleaned up per file. All nine green solo and in combined
runs, except the pre-existing commands+snapshot pairing — proven
identical WITH and WITHOUT these edits (baseline receipts) — which is
the daemon-lifecycle follow-up now extended in TODOS with this
session's receipts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): Chromium-crash exit is daemon-only — embedded launches never kill their host

handleChromiumDisconnect unconditionally process.exit()ed. Correct for
the standalone daemon (its supervisor/user must notice); suicidal when
a TEST launches BrowserManager in-process: a mid-suite Chromium death
exited the whole bun shard with no terminal summary — the exact
truncation class the strict runner flags (observed live: CI shard 1 on
eb233299 died at cache-concurrent-refresh right after a daemon-spawning
gate test; with this fix the same pairing runs to completion and
REPORTS instead of dying).

The standalone entrypoint opts in via markDaemonProcess() under
server.ts's import.meta.main gate — the same embedder contract its
signal handlers already use (gbrowser phoenix keeps its own handlers).
Embedded contexts now get the disconnect log line and continue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): context-restore assertion is evidence-based, not prose-matching

The test failed twice per run in TWO CI cycles while passing locally
4/4: the prompt said 'present the content' and the check grepped the
FINAL message for exact phrases — local runs quoted the file, CI runs
paraphrased ('the most recent context is from branch-b...') and the
substring check lost the coin flip.

- prompt now demands machine-checkable output: the newest file's
  '## Working on:' heading VERBATIM + a literal 'RESTORED: <filename>'
  marker (the mtime-scramble and cross-branch subject matter untouched)
- assertion ordered strongest-first: RESTORED marker → legacy content
  phrases → tool-call corroboration (Read/Bash input naming the newer
  file, credited ONLY when the older file was never read — a
  both-files run must still present the right one)
- the older-file negative got STRONGER: an explicit RESTORED marker
  naming the older file fails even if wintermute words appear elsewhere
- sibling scan: context-recovery-artifacts got the additive prompt-side
  treatment only (quote the matched literals verbatim); its lenient
  1-of-6 assertion deliberately unchanged

3/3 consecutive local green with all evidence classes firing
(marker=true, content=true, toolNewer=true, toolOlder=false).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): skillify family — HOME==cwd broke project-skill registration

Root cause (forensically pinned from stream-json init events + a
kill-after-init probe): with HOME set EQUAL to the child's cwd, claude
resolves <cwd>/.claude/skills as the PERSONAL skills directory and the
seeded project-tier skills never register — the Skill tool returned
'Unknown skill'. The provenance-refusal test then improvised a refusal
whose wording missed the regex (the deterministic CI+local red); the
happy-path and approval-reject siblings passed only because their
agents self-recovered by Reading SKILL.md manually — silently not
exercising the Skill-tool path at all.

All three tests now use HOME=<workDir>/home (a fresh subdir keeps the
override's intent: child ~/.gstack writes land in the assertable
sandbox, without the cwd collision). Refusal test additionally: a
'not registered/unknown skill' tripwire (a not-loaded skill can never
pass as a refusal) and the refusal regex now matches assistant text
only — the skill BODY echoed into the transcript contains the exact
refusal message, so the old full-surface match could pass vacuously
once the skill loaded. Sibling disk assertions sweep both $HOME/.gstack
and cwd .gstack roots (positives and negatives).

Verified paid: refusal 2x consecutive green with the skill's EXACT
message rendered ('Launching skill: skillify' in-transcript), then the
full file 5/5 green (~$1.35) with both siblings driving real Skill
calls (25-27 turns each).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(TODOS): two of three first-execution findings fixed (skillify family, context-restore)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): context-restore gets a private home — the REAL root cause was fixture sharing

The evidence-based assertion fix was treating a symptom. The slice
artifact's embedded transcript showed the CI agent restoring
20260829-context-save-skill-test.md — the checkpoint the SIBLING
context-save test wrote into the SHARED gstackHome checkpoints dir,
which by filename-prefix ordering genuinely IS the newest. The agent
behaved CORRECTLY; the test's fixture set was open to concurrent
sibling writes, and bun --concurrent ordering differs between CI (save
finished first) and local (restore listed first) — the entire
local-green/CI-red split explained.

The restore test now uses its own .gstack-restore-home (the whole home
moves, not just the handed path — an agent deriving the dir from
GSTACK_HOME/projects/<slug> must land in the closed set too). Full file
4/4 paid green with all evidence flags firing.

Also: the on-failure shard-log artifact glob uploaded nothing — the
Fix-bun-temp step points TMPDIR at /home/runner/.cache, so the spool
lands there, not /tmp. Both eval workflows now glob both locations
(this gap is why diagnosing THIS failure required digging transcripts
out of the slice-results artifact).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(evidence): carry the real index mtime onto gstack-wtree's temp copy

The stat-cache seed (cp of the real index) stamped the temp index "now",
which defeats git's racy-git protection: an entry is only re-hashed when
its cached mtime is not older than the index file itself, so a same-size
rewrite landing in the same second as the last real index write looked
non-racy, kept its stale stat-cache entry, and vanished from the
fingerprint — evidence stayed FRESH after a source change. This is the
CI flake in test/evidence.test.ts "allow-paths carve-out" (sub-second
alignment on fast runners: expected STALE exit 1, got FRESH exit 0).

touch -r restores the original index timestamp, reinstating the exact
racy window git itself uses. Deterministic regression pin in
test/review-log.test.ts reproduces the miss with pinned zero-nsec
timestamps (fails on the old script, passes now); receipts: manual
probe shows the fresh-stamped copy returning the clean tree for a
same-size 'hello'→'howdy' rewrite while the mtime-carried copy detects
it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): landscape gate bounds the promotion count instead of pinning 3

The alt-hinted image promotion rides the per-render measurement race
already filed in TODOS (2-vs-3 landscape pages on renders seconds
apart — CI receipts from PR #2721, now reproduced locally). Pin the
two deterministic promotions as the floor and the three promotable
blocks as the ceiling (anything above 3 means the veto leaked); the
veto/portrait assertions remain exact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Test <test@test.com>
2026-08-29 09:06:54 -07:00
2026-03-12 01:32:16 -07:00

gstack

"I don't think I've typed like a line of code probably since December, basically, which is an extremely large change." — Andrej Karpathy, No Priors podcast, March 2026

When I heard Karpathy say this, I wanted to find out how. How does one person ship like a team of twenty? Peter Steinberger built OpenClaw — 247K GitHub stars — essentially solo with AI agents. The revolution is here. A single builder with the right tooling can move faster than a traditional team.

I'm Garry Tan, President & CEO of Y Combinator. I've worked with thousands of startups — Coinbase, Instacart, Rippling — when they were one or two people in a garage. Before YC, I was one of the first eng/PM/designers at Palantir, cofounded Posterous (sold to Twitter), and built Bookface, YC's internal social network.

gstack is my answer. I've been building products for twenty years, and right now I'm shipping more products than I ever have. In the last 60 days: 3 production services, 40+ shipped features, part-time, while running YC full-time. On logical code change — not raw LOC, which AI inflates — my 2026 run rate is ~810× my 2013 pace (11,417 vs 14 logical lines/day). Year-to-date (through April 18), 2026 has already produced 240× the entire 2013 year. Measured across 40 public + private garrytan/* repos including Bookface, after excluding one demo repo. AI wrote most of it. The point isn't who typed it, it's what shipped.

The LOC critics aren't wrong that raw line counts inflate with AI. They are wrong that normalized-for-inflation, I'm less productive. I'm more productive, by a lot. Full methodology, caveats, and reproduction script: On the LOC Controversy.

2026 — 1,237 contributions and counting:

GitHub contributions 2026 — 1,237 contributions, massive acceleration in Jan-Mar

2013 — when I built Bookface at YC (772 contributions):

GitHub contributions 2013 — 772 contributions building Bookface at YC

Same person. Different era. The difference is the tooling.

gstack is how I do it. It turns Claude Code into a virtual engineering team — a CEO who rethinks the product, an eng manager who locks architecture, a designer who catches AI slop, a reviewer who finds production bugs, a QA lead who opens a real browser, a security officer who runs OWASP + STRIDE audits, and a release engineer who ships the PR. Twenty-three specialists and eight power tools, all slash commands, all Markdown, all free, MIT license.

This is my open source software factory. I use it every day. I'm sharing it because these tools should be available to everyone.

Fork it. Improve it. Make it yours. And if you want to hate on free open source software — you're welcome to, but I'd rather you just try it first.

Who this is for:

  • Founders and CEOs — especially technical ones who still want to ship
  • First-time Claude Code users — structured roles instead of a blank prompt
  • Tech leads and staff engineers — rigorous review, QA, and release automation on every PR

Quick start

  1. Install gstack (30 seconds — see below)
  2. Run /office-hours — describe what you're building
  3. Run /plan-ceo-review on any feature idea
  4. Run /review on any branch with changes
  5. Run /qa on your staging URL
  6. Stop there. You'll know if this is for you.

Install — 30 seconds

Requirements: Claude Code, Git, Bun v1.0+, Node.js (Windows only)

Step 1: Install on your machine

Open Claude Code and paste this. Claude does the rest.

Install gstack: run git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack && cd ~/.claude/skills/gstack && ./setup then add a "gstack" section to CLAUDE.md that says to use the /browse skill from gstack for all web browsing, never use mcp__claude-in-chrome__* tools, and lists the available skills: /office-hours, /plan-ceo-review, /plan-eng-review, /plan-design-review, /design-consultation, /design-shotgun, /design-html, /review, /ship, /land-and-deploy, /canary, /benchmark, /browse, /connect-chrome, /qa, /qa-only, /design-review, /setup-browser-cookies, /setup-deploy, /setup-gbrain, /retro, /investigate, /document-release, /document-generate, /codex, /cso, /autoplan, /plan-devex-review, /devex-review, /careful, /freeze, /guard, /unfreeze, /gstack-upgrade, /learn. Then ask the user if they also want to add gstack to the current project so teammates get it.

From inside your repo, paste this. Switches you to team mode, bootstraps the repo so teammates get gstack automatically, and commits the change:

(cd ~/.claude/skills/gstack && ./setup --team) && ~/.claude/skills/gstack/bin/gstack-team-init required && git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"

No vendored files in your repo, no version drift, no manual upgrades. Every Claude Code session starts with a fast auto-update check (throttled to once/hour, network-failure-safe, completely silent).

Swap required for optional if you'd rather nudge teammates than block them.

OpenClaw

OpenClaw spawns Claude Code sessions via ACP, so every gstack skill just works when Claude Code has gstack installed. Paste this to your OpenClaw agent:

Install gstack: run git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack && cd ~/.claude/skills/gstack && ./setup to install gstack for Claude Code. Then add a "Coding Tasks" section to AGENTS.md that says: when spawning Claude Code sessions for coding work, tell the session to use gstack skills. Include these examples — security audit: "Load gstack. Run /cso", code review: "Load gstack. Run /review", QA test a URL: "Load gstack. Run /qa https://...", build a feature end-to-end: "Load gstack. Run /autoplan, implement the plan, then run /ship", plan before building: "Load gstack. Run /office-hours then /autoplan. Save the plan, don't implement."

After setup, just talk to your OpenClaw agent naturally:

You say What happens
"Fix the typo in README" Simple — Claude Code session, no gstack needed
"Run a security audit on this repo" Spawns Claude Code with Run /cso
"Build me a notifications feature" Spawns Claude Code with /autoplan → implement → /ship
"Help me plan the v2 API redesign" Spawns Claude Code with /office-hours → /autoplan, saves plan

See docs/OPENCLAW.md for advanced dispatch routing and the gstack-lite/gstack-full prompt templates.

Native OpenClaw Skills (via ClawHub)

Four methodology skills that work directly in your OpenClaw agent, no Claude Code session needed. Install from ClawHub:

clawhub install gstack-openclaw-office-hours gstack-openclaw-ceo-review gstack-openclaw-investigate gstack-openclaw-retro
Skill What it does
gstack-openclaw-office-hours Product interrogation with 6 forcing questions
gstack-openclaw-ceo-review Strategic challenge with 4 scope modes
gstack-openclaw-investigate Root cause debugging methodology
gstack-openclaw-retro Weekly engineering retrospective

These are conversational skills. Your OpenClaw agent runs them directly via chat.

Other AI Agents

gstack works on 10 AI coding agents, not just Claude. Setup auto-detects which agents you have installed:

git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/gstack
cd ~/gstack && ./setup

Or target a specific agent with ./setup --host <name>:

Agent Flag Skills install to
OpenAI Codex CLI --host codex ${CODEX_HOME:-~/.codex}/skills/gstack-*/
OpenCode --host opencode ~/.config/opencode/skills/gstack-*/
Cursor --host cursor ~/.cursor/skills/gstack-*/
Factory Droid --host factory ~/.factory/skills/gstack-*/
Slate --host slate ~/.slate/skills/gstack-*/
Kiro --host kiro ~/.kiro/skills/gstack-*/
Hermes --host hermes ~/.hermes/skills/gstack-*/
GBrain (mod) --host gbrain ~/.gbrain/skills/gstack-*/

For Codex, setup reads the top-level model from ${CODEX_HOME:-~/.codex}/config.toml and generates the matching behavioral profile. gpt-5.6-sol automatically receives bounded-scope instructions that finish the requested lake without expanding into adjacent cleanup or speculative hardening. The Sol profile is exact-match only: dated snapshots and other 5.6 variants get the generic GPT profile, and setup warns on near-misses like gpt-5.6-sol-2026-08-01. Override detection with ./setup --host codex --model <id> — the override applies to that run only; set model in your Codex config.toml to make it stick across upgrades. After changing your Codex model, rerun ./setup --host codex to regenerate the skills.

Want to add support for another agent? See docs/ADDING_A_HOST.md. It's one TypeScript config file, zero code changes.

See it work

You:    I want to build a daily briefing app for my calendar.
You:    /office-hours
Claude: [asks about the pain — specific examples, not hypotheticals]

You:    Multiple Google calendars, events with stale info, wrong locations.
        Prep takes forever and the results aren't good enough...

Claude: I'm going to push back on the framing. You said "daily briefing
        app." But what you actually described is a personal chief of
        staff AI.
        [extracts 5 capabilities you didn't realize you were describing]
        [challenges 4 premises — you agree, disagree, or adjust]
        [generates 3 implementation approaches with effort estimates]
        RECOMMENDATION: Ship the narrowest wedge tomorrow, learn from
        real usage. The full vision is a 3-month project — start with
        the daily briefing that actually works.
        [writes design doc → feeds into downstream skills automatically]

You:    /plan-ceo-review
        [reads the design doc, challenges scope, runs 10-section review]

You:    /plan-eng-review
        [ASCII diagrams for data flow, state machines, error paths]
        [test matrix, failure modes, security concerns]

You:    Approve plan. Exit plan mode.
        [writes 2,400 lines across 11 files. ~8 minutes.]

You:    /review
        [AUTO-FIXED] 2 issues. [ASK] Race condition → you approve fix.

You:    /qa https://staging.myapp.com
        [opens real browser, clicks through flows, finds and fixes a bug]

You:    /ship
        Tests: 42 → 51 (+9 new). PR: github.com/you/app/pull/42

You said "daily briefing app." The agent said "you're building a chief of staff AI" — because it listened to your pain, not your feature request. Eight commands, end to end. That is not a copilot. That is a team.

The sprint

gstack is a process, not a collection of tools. The skills run in the order a sprint runs:

Think → Plan → Build → Review → Test → Ship → Reflect

Each skill feeds into the next. /office-hours writes a design doc that /plan-ceo-review reads. /plan-eng-review writes a test plan that /qa picks up. /review catches bugs that /ship verifies are fixed. Nothing falls through the cracks because every step knows what came before it.

Skill Your specialist What they do
/office-hours YC Office Hours Start here. Six forcing questions that reframe your product before you write code. Pushes back on your framing, challenges premises, generates implementation alternatives. Design doc feeds into every downstream skill.
/plan-ceo-review CEO / Founder Rethink the problem. Find the 10-star product hiding inside the request. Four modes: Expansion, Selective Expansion, Hold Scope, Reduction.
/plan-eng-review Eng Manager Lock in architecture, data flow, diagrams, edge cases, and tests. Forces hidden assumptions into the open.
/plan-design-review Senior Designer Rates each design dimension 0-10, explains what a 10 looks like, then edits the plan to get there. AI Slop detection. Interactive — one AskUserQuestion per design choice.
/plan-devex-review Developer Experience Lead Interactive DX review: explores developer personas, benchmarks against competitors' TTHW, designs your magical moment, traces friction points step by step. Three modes: DX EXPANSION, DX POLISH, DX TRIAGE. 20-45 forcing questions.
/design-consultation Design Partner Build a complete design system from scratch. Researches the landscape, proposes creative risks, generates realistic product mockups.
/review Staff Engineer Find the bugs that pass CI but blow up in production. Auto-fixes the obvious ones. Flags completeness gaps.
/investigate Debugger Systematic root-cause debugging. Iron Law: no fixes without investigation. Traces data flow, tests hypotheses, stops after 3 failed fixes.
/design-review Designer Who Codes Same audit as /plan-design-review, then fixes what it finds. Atomic commits, before/after screenshots.
/devex-review DX Tester Live developer experience audit. Actually tests your onboarding: navigates docs, tries the getting started flow, times TTHW, screenshots errors. Compares against /plan-devex-review scores — the boomerang that shows if your plan matched reality.
/design-shotgun Design Explorer "Show me options." Generates 4-6 AI mockup variants, opens a comparison board in your browser, collects your feedback, and iterates. Taste memory learns what you like. Repeat until you love something, then hand it to /design-html.
/design-html Design Engineer Turn a mockup into production HTML that actually works. Pretext computed layout: text reflows, heights adjust, layouts are dynamic. 30KB, zero deps. Detects React/Svelte/Vue. Smart API routing per design type (landing page vs dashboard vs form). The output is shippable, not a demo.
/qa QA Lead Test your app, find bugs, fix them with atomic commits, re-verify. Auto-generates regression tests for every fix.
/qa-only QA Reporter Same methodology as /qa but report only. Pure bug report without code changes.
/pair-agent Multi-Agent Coordinator Share your browser with any AI agent. One command, one paste, connected. Works with OpenClaw, Hermes, Codex, Cursor, or anything that can curl. Each agent gets its own tab. Auto-launches headed mode so you watch everything. Auto-starts ngrok tunnel for remote agents. Scoped tokens, tab isolation, rate limiting, activity attribution.
/cso Chief Security Officer OWASP Top 10 + STRIDE threat model. Zero-noise: 17 false positive exclusions, 8/10+ confidence gate, independent finding verification. Each finding includes a concrete exploit scenario.
/ship Release Engineer Sync main, run tests, audit coverage, push, open PR. Bootstraps test frameworks if you don't have one.
/land-and-deploy Release Engineer Merge the PR, wait for CI and deploy, verify production health. One command from "approved" to "verified in production."
/canary SRE Post-deploy monitoring loop. Watches for console errors, performance regressions, and page failures.
/benchmark Performance Engineer Baseline page load times, Core Web Vitals, and resource sizes. Compare before/after on every PR.
/document-release Technical Writer Update all project docs to match what you just shipped. Catches stale READMEs automatically. Builds a Diataxis coverage map (reference / how-to / tutorial / explanation) so gaps are visible in the PR body.
/document-generate Documentation Author Generate missing docs from scratch using the Diataxis framework. Researches the codebase first, then writes reference / how-to / tutorial / explanation docs that actually match the code. Invokable standalone or chained from /document-release when the coverage map finds gaps. Learn more: tutorialhow-towhy Diataxis.
/retro Eng Manager Team-aware weekly retro. Per-person breakdowns, shipping streaks, test health trends, growth opportunities. /retro global runs across all your projects and AI tools (Claude Code, Codex, Gemini).
/browse QA Engineer Give the agent eyes. Real Chromium browser, real clicks, real screenshots. ~100ms per command. /open-gstack-browser launches GStack Browser with sidebar, anti-bot stealth, and auto model routing.
/setup-browser-cookies Session Manager Import cookies from your real browser (Chrome, Arc, Brave, Edge) into the headless session. Test authenticated pages.
/autoplan Review Pipeline One command, fully reviewed plan. Runs CEO → design → eng review automatically with encoded decision principles. Surfaces only taste decisions for your approval.
/spec Spec Author Turn vague intent into a precise, executable spec in five phases (why, scope, technical with mandatory code-reading, draft, file). Codex quality gate before file (blocks below 7/10), fail-closed secret redaction, dedupe against existing issues, archive to $GSTACK_STATE_ROOT/projects/$SLUG/specs/ for team-corpus recall. --execute spawns claude -p in a fresh worktree; /ship auto-closes the source issue on merge. Plan-mode aware.
/learn Memory Manage what gstack learned across sessions. Review, search, prune, and export project-specific patterns, pitfalls, and preferences. Learnings compound across sessions so gstack gets smarter on your codebase over time.
/make-pdf Publisher Markdown in, publication-quality document out. Mermaid and excalidraw fences render as vector diagrams, fully offline. Images scale to the page and never truncate; wide diagrams get their own landscape page. --to html emits one self-contained file, --to docx a Word doc.
/diagram Diagram Maker English in, editable diagram out. Emits a triplet: mermaid source, .excalidraw you can open and edit on excalidraw.com (hand-drawn style), and rendered SVG/PNG. Zero network. Embed the source in markdown and /make-pdf renders it.

Which review should I use?

Building for... Plan stage (before code) Live audit (after shipping)
End users (UI, web app, mobile) /plan-design-review /design-review
Developers (API, CLI, SDK, docs) /plan-devex-review /devex-review
Architecture (data flow, perf, tests) /plan-eng-review /review
All of the above /autoplan (runs CEO → design → eng → DX, auto-detects which apply)

Power tools

Skill What it does
/codex Second Opinion — independent code review from OpenAI Codex CLI. Three modes: review (pass/fail gate), adversarial challenge, and open consultation. Cross-model analysis when both /review and /codex have run.
/careful Safety Guardrails — warns before destructive commands (rm -rf, DROP TABLE, force-push). Say "be careful" to activate. Override any MEDIUM warning; root/home recursive deletes and default-branch force-pushes are hard-denied.
/freeze Edit Lock — restrict file edits to one directory. Prevents accidental changes outside scope while debugging.
/guard Full Safety/careful + /freeze in one command. Maximum safety for prod work.
/unfreeze Unlock — remove the /freeze boundary.
/open-gstack-browser GStack Browser — launch GStack Browser with sidebar, anti-bot stealth, auto model routing (Sonnet for actions, Opus for analysis), one-click cookie import, and Claude Code integration. Clean up pages, take smart screenshots, edit CSS, and pass info back to your terminal.
/setup-deploy Deploy Configurator — one-time setup for /land-and-deploy. Detects your platform, production URL, and deploy commands.
/setup-gbrain GBrain Onboarding — from zero to running gbrain in under 5 minutes. PGLite local, Supabase existing URL, or auto-provision a new Supabase project via Management API. MCP registration for Claude Code + per-repo trust triad (read-write/read-only/deny). Full guide.
/sync-gbrain Keep Brain Current — re-index this repo's code into gbrain via gbrain sources add + gbrain sync --strategy code, refresh the ## GBrain Search Guidance block in CLAUDE.md, and auto-remove guidance when the capability check fails. --incremental (default), --full, --dry-run. Idempotent; safe to re-run.
/gstack-upgrade Self-Updater — upgrade gstack to latest. Detects global vs vendored install, syncs both, shows what changed.
/ios-qa iOS Live-Device QA (v1.43.0.0+) — drive a real iPhone over USB CoreDevice via an embedded StateServer in the app. Read Swift source, codegen typed @Observable accessors, run the agent loop. Optional --tailnet flag exposes the device to OpenClaw or any HTTP-capable agent on your Tailscale tailnet so remote agents can run iOS QA without ever touching the hardware. Capability-tier allowlist (observe/interact/mutate/restore), per-device session lock, audit log.
/ios-fix, /ios-design-review, /ios-clean, /ios-sync iOS bug-fix loop, designer's-eye HIG audit, debug-bridge cleanup, and accessor resync. See docs/skills.md. End-to-end walkthrough: docs/howto-ios-testing-with-gstack.md.

Standalone binaries

Beyond the slash-command skills, gstack ships standalone CLIs for workflows that don't belong inside a session:

Command What it does
gstack-model-benchmark Cross-model benchmark — run the same prompt through Claude, GPT (via Codex CLI), and Gemini; compare latency, tokens, cost, and (optionally) LLM-judge quality score. Auth detected per provider, unavailable providers skip cleanly. Output as table, JSON, or markdown. --dry-run validates flags + auth without spending API calls.
gstack-taste-update Design taste learning — writes approvals and rejections from /design-shotgun into a persistent per-project taste profile. Decays 5%/week. Feeds back into future variant generation so the system learns what you actually pick.
gstack-egress Egress receipt auditor — every gstack-initiated off-machine send writes a tamper-evident, hash-chained receipt to ~/.gstack/security/egress.jsonl before the send. list shows what gstack attempted to send and to which host, grants shows the standing consent settings plus the exact command that revokes each, verify recomputes the hash chain and exits 3 on tamper (catches edits, reordering, and mid-chain deletion; truncating or deleting the ledger itself is out of scope — it's a forensic log, not tamper-proof storage).
gstack-context-bill Token bill-of-materials — read-only, offline audit of what an installed skills tree costs in tokens: always-on frontmatter every session pays vs per-invocation SKILL.md + forced references. --diff compares two trees, --budget enforces a ceiling, --exact opts into Anthropic count_tokens (sends file text off-machine; writes an egress receipt first, degrades to the offline estimate if the receipt can't be written).
gstack-code-intelligence Code-intelligence provider picker — wraps GBrain, Sourcebot, and Graphify behind one interface: options/status to see what's available, select to pick one, index/search to use it, suggest to check whether the one-time indexing offer should fire here. The offer triggers on large repos (1,000+ tracked files; a decline is persisted). Non-local providers refuse to index or search until you record per-repo consent (consent <repo> yes|no — the query text is repo-derived content), the per-repo trust policy's deny and read-only tiers veto write-class operations regardless of consent, and every off-machine send writes an egress receipt. Fully optional — with nothing selected, gstack falls back to grep.
gstack-verify-gate Verification stop hook (opt-in) — blocks a Claude Code turn from ending until the project's declared verify command passes (after 3 blocked re-entries it yields with a loud still-RED warning instead of looping forever). Declare it on one line in CLAUDE.md: <!-- gstack:verify: bun test -->. Hooks bypass the permission system, so a declared command never runs until you trust it once per repo (gstack-verify-gate --trust); editing the command invalidates trust until re-granted, and every grant is audit-logged. ./setup never registers it for you — opt in with gstack-settings-hook add-event --event Stop --command ~/.claude/skills/gstack/bin/gstack-verify-gate --source verify-gate, remove with gstack-settings-hook remove-source --source verify-gate.
gstack-wtree Working-tree fingerprint — prints a content hash of what's actually on disk (temp index seeded from the stat cache, ~40x cheaper than a full re-hash; untracked source counts, gitignored scratch doesn't). Identical content fingerprints identically through commits, rebases, amends, and squashes — it's what binds reviews and test evidence to content instead of commit SHAs.
gstack-evidence Verification-evidence ledgerrun --label <lane> -- <cmd> transparently wraps any test command (the child's exit code always passes through) and records what ran against which working-tree fingerprint; check grades each label FRESH/STALE/MISSING with --expect-cmd, --max-age, and --allow-paths binding. /ship and /land-and-deploy cite fresh evidence instead of re-running suites. Per-run logs are 0600, capped at 2MB, pruned after 30 days; the ledger and logs stay machine-local by design.
gstack-issue-guard Tracker-text trust envelope — fetches GitHub issue/PR text (issue <n>, pr-body, pr-comments, or --stdin) and wraps it in a labeled envelope so agents treat it as data: injection-shaped lines get labeled even through fullwidth and invisible-character evasion, and forged envelope banners are defused. Every tracker-text ingress in gstack routes through it, enforced by a CI scanner.
gstack-ios-qa-daemon iOS QA daemon — Mac-side broker between an agent and a connected iPhone over USB CoreDevice. Loopback by default; --tailnet opens a Tailscale-facing listener with identity-gated capability tiers. Single-instance via flock on ~/.gstack/ios-qa-daemon.pid. See docs/howto-ios-testing-with-gstack.md.
gstack-ios-qa-mint iOS allowlist manager — owner-grant CLI for the tailnet allowlist. grant/revoke/list against ~/.gstack/ios-qa-allowlist.json (mode 0600). Remote agents never auto-allowlist; this is the explicit-intent path.
gstack-ios-qa-regen iOS bridge regenerator — deterministically installs the canonical DebugBridge package, generates typed state accessors, and records the installed gstack version. Safe to rerun after source changes or upgrades.

./setup also registers one default-on Stop hook in ~/.claude/settings.json: gstack-timeline-stop (closes dangling session-timeline entries when a session is interrupted; fail-open — 2s internal budget, always exits 0, can never block a session). Skip it with ./setup --no-team, remove it with gstack-settings-hook remove-source --source gstack-timeline-stop; gstack-uninstall removes it too.

Hook registration is canonical-only: every hook command points at the stable ~/.claude/skills/gstack install, never the tree setup ran from, so deleting a worktree or Conductor workspace can't leave dead hooks erroring in your sessions. Every ./setup run also heals first: gstack-settings-hook prune-stale --repoint removes dead gstack hook entries, re-points stale ones at the stable install, and collapses duplicates, printing one line (and writing a backup beside the file) only when it changed something.

Continuous checkpoint mode (opt-in, local by default)

Set gstack-config set checkpoint_mode continuous and skills auto-commit your work as you go with a WIP: prefix plus a structured [gstack-context] body (decisions, remaining work, failed approaches). Survives crashes and context switches. /context-restore reads those commits to reconstruct session state. /ship filter-squashes WIP commits before the PR (preserving non-WIP commits) so bisect stays clean. Push is opt-in via checkpoint_push=true — default is local-only so you don't trigger CI on every WIP commit.

Domain skills + raw CDP escape hatch

Two new browser primitives compound the gstack agent over time:

  • $B domain-skill save — agent saves a per-site note (e.g., "LinkedIn's Apply button lives in an iframe") that fires automatically next time it visits that hostname. Quarantined → active after 3 successful uses → optional cross-project promotion via $B domain-skill promote-to-global. Storage lives alongside /learn's per-project learnings file. Full reference: docs/domain-skills.md.
  • $B cdp <Domain.method> — raw Chrome DevTools Protocol escape hatch for the rare case curated commands miss. Deny-default: methods must be explicitly added to browse/src/cdp-allowlist.ts with a one-line justification. Two-tier mutex serializes browser-scoped CDP calls against per-tab work. Output for data-exfil methods is wrapped in the UNTRUSTED envelope.

Want raw CDP with no rails, no allowlist, no daemon — just thin transport from agent to Chrome? browser-use/browser-harness-js is a different philosophy (agent-authored helpers vs gstack's curated commands) and a good fit if you don't want gstack's security stack. The two can coexist: gstack's $B cdp and harness can both attach to the same Chrome via Playwright's newCDPSession.

Deep dives with examples and philosophy for every skill →

Karpathy's four failure modes? Already covered.

Andrej Karpathy's AI coding rules (17K stars) nail four failure modes: wrong assumptions, overcomplexity, orthogonal edits, imperative over declarative. gstack's workflow skills enforce all four. /office-hours forces assumptions into the open before code is written. The Confusion Protocol stops Claude from guessing on architectural decisions. /review catches unnecessary complexity and drive-by edits. /ship transforms tasks into verifiable goals with test-first execution. If you already use Karpathy-style CLAUDE.md rules, gstack is the workflow enforcement layer that makes them stick across entire sprints, not just single prompts.

Parallel sprints

gstack works well with one sprint. It gets interesting with ten running at once.

Design is at the heart. /design-consultation builds your design system from scratch, researches what's out there, proposes creative risks, and writes DESIGN.md. But the real magic is the shotgun-to-HTML pipeline.

/design-shotgun is how you explore. You describe what you want. It generates 4-6 AI mockup variants using GPT Image. Then it opens a comparison board in your browser with all variants side by side. You pick favorites, leave feedback ("more whitespace", "bolder headline", "lose the gradient"), and it generates a new round. Repeat until you love something. Taste memory kicks in after a few rounds so it starts biasing toward what you actually like. No more describing your vision in words and hoping the AI gets it. You see options, pick the good ones, and iterate visually.

/design-html makes it real. Take that approved mockup (from /design-shotgun, a CEO plan, a design review, or just a description) and turn it into production-quality HTML/CSS. Not the kind of AI HTML that looks fine at one viewport width and breaks everywhere else. This uses Pretext for computed text layout: text actually reflows on resize, heights adjust to content, layouts are dynamic. 30KB overhead, zero dependencies. It detects your framework (React, Svelte, Vue) and outputs the right format. Smart API routing picks different Pretext patterns depending on whether it's a landing page, dashboard, form, or card layout. The output is something you'd actually ship, not a demo.

/qa was a massive unlock. It let me go from 6 to 12 parallel workers. Claude Code saying "I SEE THE ISSUE" and then actually fixing it, generating a regression test, and verifying the fix — that changed how I work. The agent has eyes now.

Smart review routing. Just like at a well-run startup: CEO doesn't have to look at infra bug fixes, design review isn't needed for backend changes. gstack tracks what reviews are run, figures out what's appropriate, and just does the smart thing. The Review Readiness Dashboard tells you where you stand before you ship.

Test everything. /ship bootstraps test frameworks from scratch if your project doesn't have one. Every /ship run produces a coverage audit. Every /qa bug fix generates a regression test. 100% test coverage is the goal — tests make vibe coding safe instead of yolo coding.

/document-release is the engineer you never had. It reads every doc file in your project, cross-references the diff, and updates everything that drifted. README, ARCHITECTURE, CONTRIBUTING, CLAUDE.md, TODOS — all kept current automatically. And now /ship auto-invokes it — docs stay current without an extra command.

Real browser mode. /open-gstack-browser launches GStack Browser, an AI-controlled Chromium with anti-bot stealth, custom branding, and the sidebar extension baked in. Sites like Google and NYTimes work without captchas. The menu bar says "GStack Browser" instead of "Chrome for Testing." Your regular Chrome stays untouched. All existing browse commands work unchanged. $B disconnect returns to headless. The browser stays alive as long as the window is open... no idle timeout killing it while you're working.

Sidebar agent — your AI browser assistant. Type natural language in the Chrome side panel and a child Claude instance executes it. "Navigate to the settings page and screenshot it." "Fill out this form with test data." "Go through every item in this list and extract the prices." The sidebar auto-routes to the right model: Sonnet for fast actions (click, navigate, screenshot) and Opus for reading and analysis. Each task gets up to 5 minutes. The sidebar agent runs in an isolated session, so it won't interfere with your main Claude Code window. One-click cookie import right from the sidebar footer.

Personal automation. The sidebar agent isn't just for dev workflows. Example: "Browse my kid's school parent portal and add all the other parents' names, phone numbers, and photos to my Google Contacts." Two ways to get authenticated: (1) log in once in the headed browser, your session persists, or (2) click the "cookies" button in the sidebar footer to import cookies from your real Chrome. Once authenticated, Claude navigates the directory, extracts the data, and creates the contacts.

Prompt injection defense. Hostile web pages try to hijack your sidebar agent. gstack ships a layered defense: content filters (datamarking, hidden-element stripping, ARIA scrubbing, URL blocklist) on every page read, plus a 22MB ML classifier running locally in a sidecar subprocess that scans page-derived content before the agent sees it, with a verdict combiner that requires classifier agreement before blocking (prevents single-model false positives on Stack Overflow-style instruction pages). Everything runs on your machine, no network calls. Emergency kill switch: GSTACK_SECURITY_OFF=1. See ARCHITECTURE.md for the full stack.

Browser handoff when the AI gets stuck. Hit a CAPTCHA, auth wall, or MFA prompt? $B handoff opens a visible Chrome at the exact same page with all your cookies and tabs intact. Solve the problem, tell Claude you're done, $B resume picks up right where it left off. The agent even suggests it automatically after 3 consecutive failures.

/pair-agent is cross-agent coordination. You're in Claude Code. You also have OpenClaw running. Or Hermes. Or Codex. You want them both looking at the same website. Type /pair-agent, pick your agent, and a GStack Browser window opens so you can watch. The skill prints a block of instructions. Paste that block into the other agent's chat. It exchanges a one-time setup key for a session token, creates its own tab, and starts browsing. You see both agents working in the same browser, each in their own tab, neither able to interfere with the other. If ngrok is installed, the tunnel starts automatically so the other agent can be on a completely different machine. Same-machine agents get a zero-friction shortcut that writes credentials directly. This is the first time AI agents from different vendors can coordinate through a shared browser with real security: scoped tokens, tab isolation, rate limiting, domain restrictions, and activity attribution.

Multi-AI second opinion. /codex gets an independent review from OpenAI's Codex CLI — a completely different AI looking at the same diff. Three modes: code review with a pass/fail gate, adversarial challenge that actively tries to break your code, and open consultation with session continuity. When both /review (Claude) and /codex (OpenAI) have reviewed the same branch, you get a cross-model analysis showing which findings overlap and which are unique to each.

Safety guardrails on demand. Say "be careful" and /careful warns before any destructive command — rm -rf, DROP TABLE, force-push, git reset --hard. /freeze locks edits to one directory while debugging so Claude can't accidentally "fix" unrelated code. /guard activates both. /investigate auto-freezes to the module being investigated.

Proactive skill suggestions. gstack notices what stage you're in — brainstorming, reviewing, debugging, testing — and suggests the right skill. Don't like it? Say "stop suggesting" and it remembers across sessions.

10-15 parallel sprints

gstack is powerful with one sprint. It is transformative with ten running at once.

Conductor runs multiple Claude Code sessions in parallel — each in its own isolated workspace. One session running /office-hours on a new idea, another doing /review on a PR, a third implementing a feature, a fourth running /qa on staging, and six more on other branches. All at the same time. I regularly run 10-15 parallel sprints — that's the practical max right now.

The sprint structure is what makes parallelism work. Without a process, ten agents is ten sources of chaos. With a process — think, plan, build, review, test, ship — each agent knows exactly what to do and when to stop. You manage them the way a CEO manages a team: check in on the decisions that matter, let the rest run.

Voice input (AquaVoice, Whisper, etc.)

gstack skills have voice-friendly trigger phrases. Say what you want naturally — "run a security check", "test the website", "do an engineering review" — and the right skill activates. You don't need to remember slash command names or acronyms.

Uninstall

Option 1: Run the uninstall script

If gstack is installed on your machine:

~/.claude/skills/gstack/bin/gstack-uninstall

This handles skills, symlinks, global state (~/.gstack/), project-local state, browse daemons, and temp files. Use --keep-state to preserve config and analytics. Use --force to skip confirmation.

Option 2: Manual removal (no local repo)

If you don't have the repo cloned (e.g. you installed via a Claude Code paste and later deleted the clone):

# 1. Stop browse daemons
pkill -f "gstack.*browse" 2>/dev/null || true

# 2. Remove per-skill directories whose SKILL.md points into gstack/
#    (rm -rf, not rmdir — installed dirs also contain runtime-asset links)
find ~/.claude/skills -mindepth 1 -maxdepth 1 -type d ! -name gstack 2>/dev/null |
while IFS= read -r dir; do
  link="$dir/SKILL.md"
  [ -L "$link" ] || continue
  target=$(readlink "$link" 2>/dev/null) || continue
  case "$target" in
    gstack/*|*/gstack/*)
      rm -rf "$dir"
      ;;
  esac
done
# Alias skills install as copies (no symlink to detect) — remove by name
rm -rf ~/.claude/skills/_gstack-command ~/.claude/skills/connect-chrome 2>/dev/null

# 3. Remove gstack
rm -rf ~/.claude/skills/gstack

# 4. Remove global state
rm -rf ~/.gstack

# 5. Remove integrations (skip any you never installed)
rm -rf "${CODEX_HOME:-$HOME/.codex}/skills/gstack"* 2>/dev/null
rm -rf ~/.factory/skills/gstack* 2>/dev/null
rm -rf ~/.kiro/skills/gstack* 2>/dev/null
rm -rf ~/.openclaw/skills/gstack* 2>/dev/null
rm -rf ~/.cursor/skills/gstack* 2>/dev/null
rm -rf ~/.config/opencode/skills/gstack* 2>/dev/null

# 6. Remove temp files
rm -f /tmp/gstack-* 2>/dev/null

# 7. Per-project cleanup (run from each project root)
rm -rf .gstack .gstack-worktrees .claude/skills/gstack 2>/dev/null
rm -rf .agents/skills/gstack* .factory/skills/gstack* 2>/dev/null

Manual removal leaves gstack's hook entries behind in ~/.claude/settings.json (the uninstall script removes all of them for you, including entries whose _gstack_source tag was stripped). Edit that file and delete every hook whose command path points into .claude/skills/gstack/: the SessionStart auto-update hook, the AskUserQuestion PreToolUse/PostToolUse hooks, and the Stop hooks (session timeline, plus verify-gate if you opted in). Left in place, they error on every matching event once the install directory is gone.

Clean up CLAUDE.md

The uninstall script does not edit CLAUDE.md. In each project where gstack was added, remove the ## gstack and ## Skill routing sections.

Playwright

~/Library/Caches/ms-playwright/ (macOS) is left in place because other tools may share it. Remove it if nothing else needs it.


Free, MIT licensed, open source. No premium tier, no waitlist.

I open sourced how I build software. You can fork it and make it your own.

We're hiring. Want to ship real products at AI-coding speed and help harden gstack? Come work at YC — ycombinator.com/software Extremely competitive salary and equity. San Francisco, Dogpatch District.

GBrain — persistent knowledge for your coding agent

GBrain is a persistent knowledge base for AI agents — think of it as the memory your agent actually keeps between sessions. GStack gives you a one-command path from zero to "it's running, my agent can call it."

/setup-gbrain

Four paths, pick one:

  • Supabase, existing URL — your cloud agent already provisioned a brain; paste the Session Pooler URL, now this laptop uses the same data.
  • Supabase, auto-provision — paste a Supabase Personal Access Token; the skill creates a new project, polls to healthy, fetches the pooler URL, hands it to gbrain init. ~90 seconds end-to-end.
  • PGLite local — zero accounts, zero network, ~30 seconds. Isolated brain on this Mac only. Great for try-first; migrate to Supabase later with /setup-gbrain --switch.
  • Remote gbrain MCP — your brain runs on another machine (Tailscale, ngrok, internal LAN) or a teammate's server; paste an MCP URL and bearer token. Optionally pair with a local PGLite for symbol-aware code search in split-engine mode. Best for cross-machine memory without standing up a local DB.

After init, the skill offers to register gbrain as an MCP server for Claude Code (claude mcp add gbrain -- gbrain serve) so gbrain search, gbrain put, etc. show up as first-class typed tools — not bash shell-outs.

Keeping the brain current. Run /sync-gbrain from any repo to re-index its code into gbrain (incremental by default, --full for a full reindex, --dry-run to preview). The skill registers the cwd as a federated source via gbrain sources add, runs gbrain sync --strategy code, and writes a ## GBrain Search Guidance block to your project's CLAUDE.md so the agent prefers gbrain search/code-def/code-refs over Grep. The block is removed automatically if the capability check fails — no stale guidance pointing at tools that aren't installed.

Per-remote trust policy. Each repo on your machine gets one of three tiers:

  • read-write — agent can search the brain AND write new pages back from this repo
  • read-only — agent can search but never writes (best for multi-client consultants: search the shared brain, don't contaminate it with Client A's work while in Client B's repo)
  • deny — no gbrain interaction at all

The skill asks once per repo. The decision is sticky across worktrees and branches of the same remote.

GStack memory sync (different feature, same private-repo infra). Optionally pushes your gstack state (learnings, CEO plans, design docs, retros, developer profile) to a private git repo so your memory follows you across machines, with a one-time privacy prompt (everything allowlisted / artifacts only / off) and a defense-in-depth secret scanner that blocks AWS keys, tokens, PEM blocks, and JWTs before they leave your machine.

gstack-artifacts-init

Running gstack in Conductor? Conductor explicitly strips ANTHROPIC_API_KEY and OPENAI_API_KEY from every workspace's process env, so paid evals and gbrain embeddings won't work out of the box. Set GSTACK_ANTHROPIC_API_KEY and GSTACK_OPENAI_API_KEY in Conductor's workspace env config instead — gstack's TS entry points promote them to canonical names at runtime. Full details and the contributor checklist for adding the import to new entry points: Conductor + GSTACK_* env vars.

Full monty — every scenario, every flag, every bin helper, every troubleshooting step: USING_GBRAIN_WITH_GSTACK.md

Other references: docs/gbrain-sync.md (sync-specific guide) • docs/gbrain-sync-errors.md (error index)

Docs

Doc What it covers
Skill Deep Dives Philosophy, examples, and workflow for every skill (includes Greptile integration)
Diagrams & Document Formats Mermaid/excalidraw fences in PDFs, image sizing and safety defaults, --to html|docx, /diagram triplets
Builder Ethos Builder philosophy: Boil the Ocean, Search Before Building, three layers of knowledge
Using GBrain with GStack Every path, flag, bin helper, and troubleshooting step for /setup-gbrain
GBrain Sync Cross-machine memory setup, privacy modes, troubleshooting
Architecture Design decisions and system internals
Browser Reference Full command reference for /browse
Contributing Dev setup, testing, contributor mode, and dev mode
Changelog What's new in every version

Privacy & Telemetry

gstack includes opt-in usage telemetry to help improve the project. Here's exactly what happens:

  • Default is off. Nothing is sent anywhere unless you explicitly say yes.
  • On first run, gstack asks if you want to share anonymous usage data. You can say no.
  • What's sent (if you opt in): skill name, duration, success/fail, gstack version, OS. That's it.
  • What's never sent: code, file paths, repo names, branch names, prompts, or any user-generated content.
  • Change anytime: gstack-config set telemetry off disables everything instantly.
  • Every off-machine send is receipted. Any gstack-initiated network send — telemetry included — writes a hash-chained, tamper-evident receipt to ~/.gstack/security/egress.jsonl before the send; sensitive sinks refuse to send at all if the receipt can't be written. Audit with gstack-egress list, verify the chain with gstack-egress verify (exit 3 on tamper), see the standing consent settings with gstack-egress grants. The ledger records attempted sends so accidents are auditable — it's an audit trail, not a network firewall.

Data is stored in Supabase (open source Firebase alternative). The schema is in supabase/migrations/ — you can verify exactly what's collected. The Supabase publishable key in the repo is a public key (like a Firebase API key) — row-level security policies deny all direct access. Telemetry flows through validated edge functions that enforce schema checks, event type allowlists, and field length limits.

Local analytics are always available. Run gstack-analytics to see your personal usage dashboard from the local JSONL file — no remote data needed.

Troubleshooting

Skill not showing up? cd ~/.claude/skills/gstack && ./setup

/browse fails? cd ~/.claude/skills/gstack && bun install && bun run build

Stale install? Run /gstack-upgrade — or set auto_upgrade: true in ~/.gstack/config.yaml

Want shorter commands? cd ~/.claude/skills/gstack && ./setup --no-prefix — switches from /gstack-qa to /qa. Your choice is remembered for future upgrades.

Want namespaced commands? cd ~/.claude/skills/gstack && ./setup --prefix — switches from /qa to /gstack-qa. Useful if you run other skill packs alongside gstack.

Codex says "Skipped loading skill(s) due to invalid SKILL.md"? Your Codex skill descriptions are stale. Fix: cd "${CODEX_HOME:-$HOME/.codex}/skills/gstack" && git pull && ./setup --host codex — or for repo-local installs: cd "$(readlink -f .agents/skills/gstack)" && git pull && ./setup --host codex

Windows users: gstack works on Windows 11 via Git Bash or WSL. Node.js is required in addition to Bun — Bun has a known bug with Playwright's pipe transport on Windows (bun#4253). The browse server automatically falls back to Node.js. Make sure both bun and node are on your PATH.

On Windows without Developer Mode (MSYS2 / Git Bash), setup falls back to file copies instead of symlinks because ln -snf produces frozen copies that don't refresh on git pull. Re-run cd ~/.claude/skills/gstack && ./setup after every git pull so your skill files match the repo. setup prints a one-line note reminding you. Unix and WSL keep symlinks and don't need the re-run.

Claude says it can't see the skills? Make sure your project's CLAUDE.md has a gstack section. Add this:

## gstack
Use /browse from gstack for all web browsing. Never use mcp__claude-in-chrome__* tools.
Available skills: /office-hours, /plan-ceo-review, /plan-eng-review, /plan-design-review,
/design-consultation, /design-shotgun, /design-html, /review, /ship, /land-and-deploy,
/canary, /benchmark, /browse, /open-gstack-browser, /qa, /qa-only, /design-review,
/setup-browser-cookies, /setup-deploy, /setup-gbrain, /sync-gbrain, /retro, /investigate,
/document-release, /document-generate, /codex, /cso, /autoplan, /pair-agent, /careful, /freeze,
/guard, /unfreeze, /gstack-upgrade, /learn.

License

MIT. Free forever. Go build something.

S
Description
No description provided
Readme MIT
384 MiB
Languages
TypeScript 82%
Go Template 9.1%
Shell 6.6%
JavaScript 1.4%
CSS 0.3%
Other 0.5%