mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-31 10:20:42 +02:00
* feat(autoplan): eng review always runs last — the gate reviews the final amended plan Reorder the pipeline to CEO -> Design (if UI scope) -> DX (if developer-facing scope) -> Eng. The old order (CEO -> Design -> Eng -> DX) let DX findings land AFTER the required gate signed off, so eng validated a stale plan. Accept-all semantics made explicit: every AskUserQuestion resolves to the recommended option; premises no longer pause the pipeline mid-run (clearly-wrong ones queue as User-Challenge items at the single Final Approval Gate). Eng's Codex voice now sees the DX consensus summary. New free static test pins the order; the chain E2E gains DX-between and Eng-terminal assertions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(review): simplification specialist — advisory over-engineering lens with ponytail's tag vocabulary New 8th Review Army specialist (DIFF_LINES > 100, --simplification force flag) hunting unrequested STRUCTURE only: delete/stdlib/native/speculative/shrink closed tags, one-line findings, lines_removable field. speculative: replaces ponytail's yagni: tag — we import the lens, not the posture; coverage stays sacred (Completeness Gaps owns it, suppressions inlined, shrink needs >=5 lines). Advisory carve-out in the merge step: advisory findings are excluded from quality_score and the findings-count header, render with an [ADVISORY] label, and are ASK-only in Fix-First. Zero-findings case prints the lens-scoped 'Simplification: lean already — nothing to cut.' from the PARENT (the specialist keeps the exact NO FINDINGS contract); with findings, the parent prints 'net: -N lines possible' summed from lines_removable. Tests: static pins for the carve-out + early-out contract (gen-skill-docs), two periodic e2e cases with planted fixtures — activation (over-build traps: hand-rolled Intl, one-impl abstract, dead config) and false-flag precision (a lean ETHOS 'choose A' diff must yield NO FINDINGS). Inspired by dietrichgebert/ponytail's /ponytail-review. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): reuse ladder in Search Before Building — rungs 2-5 of ponytail's ladder, completeness kept Tier-3+ skills gain a per-edit reflex the section only stated as research discipline: before writing new code, stop at the first rung that holds — repo helper, stdlib, native platform feature, installed dependency — then build the COMPLETE version of what remains. The closing clause is the explicit reconciliation with Boil the Ocean: the ladder governs structure, never coverage. Rungs 1/6/7 (YAGNI / one line / minimum that works) are deliberately NOT imported. Also ports ponytail's root-cause rule: one guard in the shared function beats a guard in every caller. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): bounded-closer output rule for tier-2+ skills After completing work, skills report in a few short lines — what changed, what was skipped, what to watch — and cut any explanation that outgrows the change. Explicit exemptions protect every mandated output: decision briefs, completion-status blocks, user-requested explanations, and report-shaped skills' report formats (the report IS the work in /qa-only, /plan-*-review, /retro, /document-generate). Rationale is signal-to-noise, not tokens: ponytail's own benchmark shows terse prose alone doesn't cut cost (caveman arm: -20% LOC, +7% tokens), and independent replications found its 'skipped on purpose' essays ate the code savings. Includes a good/bad closer example pair per the model-overlay guidance that a positive example beats a 'don't be verbose' instruction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(resolvers): terse-mode savings claim matches measurement — 2.6KB, not 3-5KB Measured on the v1.71 render: --explain-level=terse saves exactly 2,611 bytes per tier-2+ skill. The old ~3-5KB claim predated the preamble restructuring. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(retro,preamble): gstack-shortcut debt ledger — accepted shortcuts leave a joined trail When the user accepts an option that is BOTH Completeness <= 7 AND a durable-scope call, the decision ledger entry (gstack-decision-log, ceiling + upgrade trigger in the rationale) is the source of truth, and the agent marks each cut corner in code with gstack-shortcut(dec-<id>): <ceiling>, upgrade when <trigger> — same edit, no follow-up question, never agent-initiated. /retro Step 11.5 harvests markers into a debt ledger (grep || true — zero matches is the healthy case; skill installs and docs excluded), joins on the decision id so nothing double-counts, tags unlinked and no-trigger rot risks, and closes with 'N markers, M with no trigger.' /review suppressions: a marker with ceiling+trigger downgrades a would-be Completeness Gaps finding to acknowledged debt. Redaction test pins that the marker ships untouched (the ledger is the point) — it does not match the TODO(owner) hygiene shape. Format from dietrichgebert/ponytail's ponytail-debt; store inverted to gstack's existing decision ledger. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh golden ship baselines after preamble additions (reuse ladder + bounded closer) The golden-file regression test pins the rendered ship skill byte-for-byte; the WS3/WS7 preamble sections are deliberate changes, so the baselines re-capture per the goldens' own update protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(hosts): instruction-only tier — a 2KB committed rules digest any agent host can read New agents-digest/gstack-AGENTS.md (1,765 bytes, hard 2,048-byte budget): gstack's ethos one-liners, the reuse ladder, and voice rules for hosts with no install arm — Zed, Amp, Jules, or any AGENTS.md-reading agent. Generated by scripts/gen-agents-digest.ts, auto-refreshed by gen:skill-docs, committed like llms.txt so setup's explainer arms can point at it before any toolchain exists. First line carries the gstack version as its own staleness nudge. Delivery is print-path + user-performed copy ONLY: setup never writes or overwrites a user's AGENTS.md (a test pins this — no cp/ln/mv/redirect into AGENTS.md anywhere in setup). openclaw and hermes explainer arms print the path; slate keeps routing to the full Claude install and gbrain ships from its own repo. HostConfig gains the optional install.instructionTier slot, declared by both instruction-tier hosts. README host table now matches what setup actually does. Inspired by dietrichgebert/ponytail's instruction-tier AGENTS.md fallback — one generated source, never per-host hand copies. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): AskUserQuestion repetition cut — gated, passed NOT-WORSE A/B Removes the duplicate statements v1.71's compaction left in the AskUserQuestion Format section: the completeness rule restated in the prose triad, the auto-decide marker syntax stated twice, the Conductor-flakiness explanation stated twice, and the self-check's full triad restatement. Every verbosity floor and all 14 format pins stay (Layer 0 green). The gate this decision rested on ran before landing (new periodic skill-e2e-auq-repetition-cut-ab.test.ts, pre-cut ref3263fffevs this render, same harness as auq-verbose-vs-carved-ab): POST 7/7 format elements, substance 5 — identical to PRE. No degradation; the load-bearing-repetition hypothesis did not hold for these duplicates. Net: -236 bytes per tier-2+ skill (~9.7KB corpus). Golden ship baselines re-captured for the deliberate change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): with-skill vs without-skill arm benchmark — measures whether gstack's behavioral layer earns its tokens Ponytail's honest-benchmark method pointed at gstack itself: 3 build-shaped tasks (native-platform over-build trap, CRUD endpoint, bug fix with planted decoys) x 2 arms, real claude -p sessions, scored on the git diff left behind. A research instrument, not a release gate — no assertion compares arm scores. Arms use the PROVEN project-scope pattern: the with-arm installs a build-discipline skill (extracted reuse-ladder + bounded-closer content, not whole-file copies) into the fixture's .claude/skills/ with a CLAUDE.md routing line and an explicit invocation; a live spike confirmed claude -p discovers and invokes project-scope skills via the Skill tool (3 turns, exact-output probe). Fixtures are git init + local bare origin; diff capture is three lines of git, no worktree machinery. Failure taxonomy: zero-diff arms are VALID scored cells (deterministic 0/none, no API call), harvest failures record harvest:null, judge_error cells are excluded from aggregates but named in the report — nothing drops silently. armJudge: fixed sonnet judge, 0-3 unrequested-structure rubric, must name the construct or say none, bounded retry-on-malformed; callJudge gains optional temperature/max_tokens (defaults unchanged). recordE2E now populates tokens_used for every E2E. Eval schema v2: harvest gains {insertions, deletions, net}, tolerant reads keep v1 runs comparable. Registered periodic in E2E_TIERS + touchfiles (with the auq-repetition-cut A/B); periodic detach timeout raised to the new shard-census floor. Free selftest (8 tests, zero API) pins fixtures, extraction, arm asymmetry, diff capture, judge plumbing, and the retry bound. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: absorb the ponytail-import wave into the guard fixtures — ceilings, schema pin, triad phrasing Skeleton ceilings re-captured for the 17 carved skills the wave deliberately grew (reuse ladder + bounded closer + shortcut trail, net of the gated -236B AUQ cut), each with its measured size in the comment per the carve-guards protocol. eval-store schema pin updated to v2 (harvest gains insertions/deletions/net). The AUQ prose-triad keeps its pinned per-choice phrasing ('explicit on EACH choice') while still deferring the score scale to the canonical Format rule — the shipped cut is strictly closer to the pre-cut text than the render that already passed the NOT-WORSE gate. Autoplan carve anchors follow the Phase 2.5 renumbering. Golden ship baselines re-captured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: observability partial-file pin follows eval-store schema v2 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test-runner): GSTACK_FREE_JOBS + opt-in flaky-retry pass for syscall-supervised sandboxes GSTACK_FREE_JOBS overrides the computed shard count (the free runner's analogue of the paid runner's EVALS_JOBS). On Vercel sandboxes, PID 1 installs a seccomp filter whose supervisor spuriously fails access(2) for busy processes — measured: 200/200 git-init probes fail 'Cannot access work tree: Permission denied' while the suite runs at 6 shards, 0/200 idle; statx succeeds while access fails on the same path in the same process. One serial mega-shard maximizes per-process pressure and fails too; 2 shards is the measured sweet spot. GSTACK_FREE_RETRY_FLAKY=1 (default OFF — dev boxes should see flakes) re-runs attributed failures once, serially, capped at 5 files; a clean retry downgrades to a loud FLAKY-PASS naming the offenders, a repeat failure stays red, timeouts and unattributed failures never retry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): portable temp paths — TEMP_DIRS allowlist, tmpdir()-based test files Local path validation now accepts os.tmpdir() alongside the classic /tmp (new TEMP_DIRS in platform.ts): on macOS os.tmpdir() is /var/folders/..., and TMPDIR-honoring CI/sandbox environments point it elsewhere entirely — both are legitimate scratch space. Remote file serving (TEMP_ONLY) stays pinned to TEMP_DIR alone; no change to the exfil boundary. commands.test.ts drops 41 hardcoded /tmp literals for a tmpp() helper on os.tmpdir() (two message assertions now reference the same variable), and path-validation's symlink-escape test targets /etc/hosts instead of /etc/crontab — the target must EXIST for realpath to resolve the link (a dangling target falls back to the link's own path and passes vacuously), and /etc/crontab is absent on Amazon Linux. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): portable sha256 — Linux ships sha256sum, not shasum resolve-user-slug and endpoint hashing exited 127 on Amazon Linux (shasum is a macOS/perl tool). New _sha256_hex helper prefers sha256sum and falls back to shasum, matching gstack-verify-gate's existing pattern; both call sites converted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): only trust ls-remote when origin is actually configured Without the guard, git DWIMs the literal 'origin' as an ssh host/path; on hosts whose transport launders exit codes the probe 'succeeds' with zero branches and the allocator silently sees an empty queue — the exact duplicate-allocation failure (#2545) fetchGitClaimed exists to prevent. git remote get-url origin gates the probe; absence falls through to the existing local-refs path with its staleness warning. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(testing): sandbox-doctor — one command makes a cloud sandbox run the suite green Measured failure taxonomy for Vercel/Conductor sandboxes (missing /dev/fd, 64M /dev/shm, seccomp-supervisor access(2) EACCES under load, uid-1000 processes with FULL capabilities defeating chmod-denial tests, no X server, no git identity, Conductor git-shim exit-code laundering) plus the idempotent script that treats all of it and seeds the run recipe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): converge on main's self-contained sha8_of — its tests extract the function standalone The merge kept a branch-local _sha256_hex helper; main's v1.72 landed the same portability fix inline WITH tests that extract sha8_of()'s text and run it under a shim-only PATH — a helper call can't satisfy that shape. Adopt the landed implementation at both hash sites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for GSTACK_FREE_JOBS override and failingFiles attribution Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for TEMP_DIRS widening and remote-serving TEMP_ONLY asymmetry Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for gstack-shortcut marker grammar and retro harvest joint Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for sandbox-doctor shell syntax and idempotency guards Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): empty-shard outcome carries failingFiles; harden flaky-retry list The empty-shard early return omitted the (required) failingFiles field — tsc TS2741 — feeding undefined into the flaky-retry flatMap. Also drop the dead 'else if (worst !== 0)' guard (the enclosing if already pins it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(release): version-bump write regenerates the version-stamped agents digest agents-digest/gstack-AGENTS.md embeds VERSION in its first line and is byte-freshness-gated (test/agents-digest.test.ts + Skill Docs Freshness CI), but nothing in the release path regenerated it — every version-bumping ship of this repo would land red. write now spawns the repo's own generator when present (agentsDigest true/false/null in the output JSON), and ship's evidence gate allow-lists the digest alongside VERSION/package.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): instruction-tier explainer prints the script-anchored digest path $(pwd) printed a nonexistent path when setup ran from any other directory; both arms now share one print_instruction_tier() using SOURCE_GSTACK_DIR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(digest): broaden AGENTS.md writer tripwire; pin digest-resolver ladder lockstep The print-path-only guard now catches tee/install/rsync/dd/truncate, >> appends, and laundered variable-destination writes. New test ties the digest's hand-rendered reuse-ladder text to the preamble resolver so an edit to either fails CI instead of shipping drift. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(retro): shortcut harvest drops placeholder markers and convention docs The Step 11.5 grep matched documentation mentions (dec-<id>, dec-*) in checklists, resolver sources, and convention tests, reporting phantom debt rows on gstack itself. A trailing filter kills placeholder forms; prose tells the agent to discard convention-quoting hits. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): advisory findings count in per-specialist stats Without this, simplification (all-advisory by construction) would log findings:0 every run and auto-gate itself into permanent silence after 10 dispatches. The advisory carve-out governs score and header only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): arm-benchmark harvest and judge hardening - Harvest diffs against the recorded seed SHA (origin/main is movable by an agent that commits AND pushes; a recorded SHA is not). - Fixtures get a node_modules .gitignore and the git wrapper a 64MB maxBuffer, so a vendored-dependency arm is scored instead of killing the cell. - The judge diff cap is a named constant with loud truncation (log + judge_reasoning suffix). - Judge prompt block markers carry a per-call random sentinel, so a diff containing a faked closing marker cannot escape the data block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): AUQ A/B vendored pre-cut arm + judge-error inconclusive taxonomy - The PRE arm read a branch-local SHA (3263fffe) that becomes unreachable on fresh clones after the squash-merge; the pre-cut render is now a vendored fixture. - A judge failure on one side no longer coerces substance to 0 (which fabricated DEGRADATION on POST-side failures and masked regressions on PRE-side failures): null substance = inconclusive, format still gates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: regression pin for the originConfigured guard vs laundering git shims On healthy hosts the guarded and unguarded paths behave identically, so a revert passes the suite; only a shim that makes 'git ls-remote' exit 0 with empty output (the Conductor wrapper's observed behavior) exposes it. Pins that the empty 'successful' probe is never trusted as an empty queue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): missing /dev/shm no longer aborts the doctor under set -eu Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(touchfiles): close dep-list gaps for the new evals - arm-benchmark entries gain ship/SKILL.md (buildBehavioralSkill extracts sections from the rendered ship skill) - review-army-simplification entries gain their planted fixtures + test file - auq-repetition-cut-ab gains llm-judge.ts and the vendored PRE fixture Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: re-capture context-budget fixture — lock the WS6-3 reduction and Step 9 deltas Per the ratchet protocol: the AUQ repetition cut shrank per-skill eager tokens but the fixture was never re-captured, leaving the win unlocked. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(release): digest regen is an explicit --regen-digest opt-in, not presence-sniffed code exec Review (security) caught the cycle-1 fix executing any repo's scripts/gen-agents-digest.ts on plain 'write' — arbitrary code exec from a hostile clone on a routine bump, contradicting the binary's own containment posture. The regen still runs the TARGET repo's generator (a 'trusted' copy beside the binary would false-red the freshness gate on version drift), but only under the flag: /ship passes it deliberately, in a repo whose code the operator already executes (its test suite). Plain write is side-effect-free again. Also: uniform output shape (agentsDigest: null on the JSON-manifest branch), a REAL generator round-trip test replacing the misnamed lockstep check, and land-and-deploy's evidence gate gets the same digest allow-path as ship so the two grading surfaces agree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): flaky-retry vetoes on ANY unattributable failure evidence The gate equated 'some failure attributed' with 'all failures attributed': a shard with one attributed failure plus a headerless failure, an unhandled error between tests, or a truncated run (no terminal summary) qualified for retry — re-running only failingFiles and masking the rest as FLAKY-PASS, re-opening the silent-truncation hole the strict classifier closes. FreeShardOutcome now carries unattributedFailures; nonzero vetoes the retry. Pins: mixed shard, truncated-with-attributed shard, empty-shard field values. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): a configured origin advertising zero heads is never trusted The originConfigured guard covered only the no-origin laundering case. With origin configured (the normal Conductor worktree state), the laundering shim makes a failed ls-remote exit 0 with empty stdout — read as 'the queue is empty', the exact duplicate-allocation bug (#2545) one layer up. A reachable remote always advertises at least its default branch, so an exit-0 zero-head probe now falls back to local refs/remotes/origin with a laundering-specific warning. Regression test shims git for both configurations. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): loud on git-shim patch drift; document the retry-contract override - The /conductor/bin/git patch was a silent no-op if the shim's bytes drift from the exact pattern — now warns that laundering is NOT fixed. - The bashrc block documents why GSTACK_FREE_RETRY_FLAKY=1 deliberately overrides the runner's default-OFF contract on this sandbox, and how to undo it. - Test pins the guarded shm form (missing /dev/shm must not abort set -eu). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(digest): pin the script-anchored explainer path; catch declaration-prefixed writers - Asserts $SOURCE_GSTACK_DIR/agents-digest path and forbids $(pwd)/agents-digest (the cycle-1 fix was revertible without failing anything). - The laundered-assignment arm now matches local/export/declare/readonly/typeset prefixed assignments — the likeliest in-function writer shape in setup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(evals): arm-benchmark selftest runs FREE on every PR The selftest lived inside the paid skill-e2e-* file, so fixture-integrity and plumbing pins executed weekly at best — a broken fixture would ship past every gating check and be discovered when the periodic run burned money on a dead instrument. Harness extracted to test/helpers/arm-benchmark-harness.ts, selftest to test/arm-benchmark-selftest.test.ts (free suite). Touchfiles: harness added to the three benchmark dep lists; the auq-repetition-cut-ab tier comment now states the MANUAL re-run obligation honestly (periodic runs force EVALS_ALL, so dep lists cannot auto-trigger it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: re-capture context-budget fixture after cycle-2 template deltas Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): keep both heredoc bodies under the 512B pipe-deadlock window The cycle-2 additions pushed the python-patch and bashrc heredocs into the 512-65536B window test/heredoc-pipe-deadlock.test.ts guards (sh scripts get no BASH_COMPAT escape hatch). Same content, tighter prose; the drift warning now reuses the patch pattern variable instead of a second literal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): a gstack-shortcut marker only suppresses findings when its decision id resolves in the ledger Cross-model catch (Claude adversarial + Codex agreed): any diff author could fabricate a marker and silence Completeness review of that gap. Reviewers now resolve the dec-id via gstack-decision-search; an orphan marker is reported as a forged suppression, not honored as debt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): define the B2 gate path — accepted premise challenges amend the plan and re-run Eng The final gate offered B2 (respond to User Challenges) but the option handler table omitted it, leaving accepted challenges with no amendment or Eng re-review path. B2 now walks challenges one at a time; an accepted one amends the plan and re-runs Eng (the gate always reviews the final plan), sharing D's 3-cycle cap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): arm benchmark runs each fixture's functional oracle — correctness before LOC The plan's metric order is diff-quality FIRST, but cells never ran the fixtures' own run-tests.js, so a refusal, a broken implementation, and working code were indistinguishable in aggregates (Codex adversarial catch). Tasks with an oracle declare checkCmd; every cell records checks=pass|fail|none in the report line and eval store. Selftest pins the oracle declarations and that the planted bug fails its own check pre-fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): check the bump's agentsDigest result; state the --regen-digest trust envelope honestly A failed digest regen warned and moved on — ship now instructs re-running the generator and staging the digest with the bump (the freshness check stays red otherwise). The 'no-op everywhere else' phrasing oversold safety: the step now names what executes and why that is inside the envelope Step 5 already opened (the repo's own test suite). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): GSTACK_FREE_JOBS accepts digits only — parseInt truncation defeated the loud-failure contract '2abc' silently became 2 and '3.7' became 3 despite the error text claiming a positive-integer requirement. Strict /^\d+$/ pre-check; both shapes pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): atomic git-shim patch, :99-socket Xvfb check, dnf gate, non-interactive sudo - The /conductor/bin/git patch writes tmp-then-rename with a .orig backup — a concurrently spawned git can never exec a truncated shim. - Xvfb running-check looks for the :99 socket, not any-display pgrep. - Xvfb install is dnf-gated so non-dnf distros degrade to a warning instead of aborting the remaining fixes under set -eu. - The bashrc /dev/fd restore uses sudo -n || true — no password prompt at every shell start on non-passwordless machines. - BASH_COMPAT=50 keeps heredoc bodies off the bash pipe window. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(build): a failed agents-digest regen fails gen-skill-docs instead of deferring the red to CI Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): an untrustable TMPDIR (/, $HOME, a cwd ancestor) never widens the local allowlist TEMP_DIRS honors os.tmpdir() at daemon start; a daemon launched with TMPDIR=/ would have trusted the whole filesystem for local path validation for its lifetime. Subprocess pins cover /, $HOME, cwd-ancestor rejection and that a benign distinct TMPDIR (the sandbox recipe's $HOME/tmp) stays honored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: zero-heads warning names the benign cause too; digest path declaration made load-bearing; ratchet re-capture - The ls-remote zero-heads warning no longer accuses an empty remote of running a laundering shim. - instructionTier.rulesFile now must equal the generator's DIGEST_RELPATH (and setup must print it) — the declaration fails with the real path instead of lying silently. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: file ship-time follow-ups in TODOS skillify HOME-override gate red (pre-existing, proven on main), the auq-verbose-vs-carved-ab branch-local ref, eval-store harvest union, evidence digest allow-path scoping, and the WS6-2 dead-frontmatter live-host verification deferral. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.73.0.0 chore: version bump + CHANGELOG — ponytail import wave Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: raise ship skeleton parity ceiling — measured 75,592 after the v1.73 release-step prose The --regen-digest trust-envelope paragraph (Step 12) and the evidence-gate digest note (Step 16) grew the ship skeleton past the previous 75,420 ceiling. Re-measured per the deliberate-change protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.73.0.0 - README.md, docs/skills.md, AGENTS.md: /autoplan phase order corrected to CEO → design → DX → eng (eng always last); /review rows note the advisory simplification lens - docs/PROJECT_STRUCTURE.md: add agents-digest/, gen-agents-digest.ts, sandbox-doctor.sh, test-free-shards.ts to the annotated tree - CONTRIBUTING.md: document GSTACK_FREE_JOBS, GSTACK_FREE_RETRY_FLAKY, and the sandbox-doctor one-command fixer in the Tier 1 test section Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply cross-model doc-review fixes for v1.73.0.0 - README.md: host table gains the OpenClaw explainer arm row (setup has the arm; the table claimed to match setup) - docs/skills.md: /review completeness-gaps section documents the gstack-shortcut(dec-<id>) acknowledged-debt suppression and orphan-marker flagging; /autoplan deep-dive states the recommended-option default with the 6 principles as tie-breakers - CONTRIBUTING.md: host count 8 -> 10 (Hermes, GBrain), supported-hosts list completed - docs/TESTING_INTERNALS.md: sandbox recipe says to source ~/.bashrc after the doctor seeds it; GSTACK_FREE_JOBS wording fixed from "caps" to "overrides in either direction" (matches the un-clamped runner) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): temp-dirs asymmetry pins are topology-aware; TMPDIR probes are POSIX-only CI exposed two wrong assumptions in the new temp-dirs tests, neither a product bug: - The remote-serving asymmetry test assumed a distinct os.tmpdir() lies OUTSIDE TEMP_DIR, but the free-shard runner nests each child's TMPDIR inside /tmp on CI — a file there is under TEMP_DIR, so serving it remotely is legitimate. The test now pins the actual exfil boundary on every topology (a cwd project file is locally readable, never remotely servable) and branches the os.tmpdir() case on nested-vs-outside. Reproduced locally with TMPDIR=/tmp/nested-tmp before fixing. - The untrustable-TMPDIR subprocess probes set TMPDIR, which Windows os.tmpdir() ignores (reads TEMP/TMP) — and on Windows TEMP_DIR is DEFINED as os.tmpdir(), so the fixed+movable two-dir topology the guard filters does not exist there. Probes now skip on Windows with that rationale; the benign-TMPDIR assertion compares realpaths. Verified under all three POSIX topologies: TMPDIR=$HOME/tmp (outside), TMPDIR=/tmp/nested-tmp (CI shard shape), TMPDIR unset (identical). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(build): DIGEST_RELPATH is a forward-slash literal on every platform path.join built it with backslashes on Windows, so the wiring test's string comparisons against setup and hosts/*.ts (which carry the forward-slash literal) could never match there — windows-free-tests red. path.join(root, DIGEST_RELPATH) at the write site normalizes fine. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): bashrc block re-heals the /dev/shm remount on sandbox restart The 4G remount does not survive restarts; a reverted 64M shm made the multi-tab browse handoff test fail consistently under suite concurrency (observed live: two consecutive full-run failures, green in isolation, green again after remounting). Same guarded arithmetic as the doctor body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): close the cross-shard porcelain race that failed Windows CI Two-part fix for the gen-skill-docs-out-dir isolation-pin failure: - cookie-import-browser built its scratch cookie DBs inside the TRACKED browse/test/fixtures/ dir (created in beforeAll, deleted in afterAll), so they flash as untracked files mid-run — a concurrent shard's porcelain snapshot caught the window on Windows. The DBs now live in a per-run tmpdir; zero source-tree writes. - gen-skill-docs-out-dir is the free suite's only LIVE porcelain-snapshot test, so it joins TREE_MUTATING (the serial quiet window): any concurrent transient tree-write can race it, and its own spawned render rewrites llms.txt/agents-digest in place (idempotent on a fresh tree). The race is pre-existing; this branch's +5 test files reshuffled shard composition and exposed it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.75.0.0 chore: queue-advance rebump — perth-v2 landed v1.74.0.0 on main The v1.73.0.0 slot this branch claimed was superseded when #2721 merged; same MINOR level relative to main per the versioning invariant. CHANGELOG entry renumbered (1.73.0.0 was branch-internal and never landed on main), digest restamped via --regen-digest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): duration-packed walls keep the per-file floor — predictions don't transfer across machines The committed duration seed is recorded on fast CI; a syscall-supervised sandbox replays the same files 2-4x slower. Observed post-merge: a 253-file shard predicted ~242s was wall-killed at its predicted-x3 725s wall while genuinely progressing (the old count heuristic guaranteed 1265s). Packed walls may be looser than the count floor, never tighter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
545 lines
52 KiB
Markdown
545 lines
52 KiB
Markdown
# gstack
|
||
|
||
> "I don't think I've typed like a line of code probably since December, basically, which is an extremely large change." — [Andrej Karpathy](https://fortune.com/2026/03/21/andrej-karpathy-openai-cofounder-ai-agents-coding-state-of-psychosis-openclaw/), No Priors podcast, March 2026
|
||
|
||
When I heard Karpathy say this, I wanted to find out how. How does one person ship like a team of twenty? Peter Steinberger built [OpenClaw](https://github.com/openclaw/openclaw) — 247K GitHub stars — essentially solo with AI agents. The revolution is here. A single builder with the right tooling can move faster than a traditional team.
|
||
|
||
I'm [Garry Tan](https://x.com/garrytan), President & CEO of [Y Combinator](https://www.ycombinator.com/). I've worked with thousands of startups — Coinbase, Instacart, Rippling — when they were one or two people in a garage. Before YC, I was one of the first eng/PM/designers at Palantir, cofounded Posterous (sold to Twitter), and built Bookface, YC's internal social network.
|
||
|
||
**gstack is my answer.** I've been building products for twenty years, and right now I'm shipping more products than I ever have. In the last 60 days: 3 production services, 40+ shipped features, part-time, while running YC full-time. On logical code change — not raw LOC, which AI inflates — my 2026 run rate is **~810× my 2013 pace** (11,417 vs 14 logical lines/day). Year-to-date (through April 18), 2026 has already produced **240× the entire 2013 year**. Measured across 40 public + private `garrytan/*` repos including Bookface, after excluding one demo repo. AI wrote most of it. The point isn't who typed it, it's what shipped.
|
||
|
||
> The LOC critics aren't wrong that raw line counts inflate with AI. They are wrong that normalized-for-inflation, I'm less productive. I'm more productive, by a lot. Full methodology, caveats, and reproduction script: **[On the LOC Controversy](docs/ON_THE_LOC_CONTROVERSY.md)**.
|
||
|
||
**2026 — 1,237 contributions and counting:**
|
||
|
||

|
||
|
||
**2013 — when I built Bookface at YC (772 contributions):**
|
||
|
||

|
||
|
||
Same person. Different era. The difference is the tooling.
|
||
|
||
**gstack is how I do it.** It turns Claude Code into a virtual engineering team — a CEO who rethinks the product, an eng manager who locks architecture, a designer who catches AI slop, a reviewer who finds production bugs, a QA lead who opens a real browser, a security officer who runs OWASP + STRIDE audits, and a release engineer who ships the PR. Twenty-three specialists and eight power tools, all slash commands, all Markdown, all free, MIT license.
|
||
|
||
This is my open source software factory. I use it every day. I'm sharing it because these tools should be available to everyone.
|
||
|
||
Fork it. Improve it. Make it yours. And if you want to hate on free open source software — you're welcome to, but I'd rather you just try it first.
|
||
|
||
**Who this is for:**
|
||
- **Founders and CEOs** — especially technical ones who still want to ship
|
||
- **First-time Claude Code users** — structured roles instead of a blank prompt
|
||
- **Tech leads and staff engineers** — rigorous review, QA, and release automation on every PR
|
||
|
||
## Quick start
|
||
|
||
1. Install gstack (30 seconds — see below)
|
||
2. Run `/office-hours` — describe what you're building
|
||
3. Run `/plan-ceo-review` on any feature idea
|
||
4. Run `/review` on any branch with changes
|
||
5. Run `/qa` on your staging URL
|
||
6. Stop there. You'll know if this is for you.
|
||
|
||
## Install — 30 seconds
|
||
|
||
**Requirements:** [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [Git](https://git-scm.com/), [Bun](https://bun.sh/) v1.0+, [Node.js](https://nodejs.org/) (Windows only)
|
||
|
||
### Step 1: Install on your machine
|
||
|
||
Open Claude Code and paste this. Claude does the rest.
|
||
|
||
> Install gstack: run **`git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack && cd ~/.claude/skills/gstack && ./setup`** then add a "gstack" section to CLAUDE.md that says to use the /browse skill from gstack for all web browsing, never use mcp\_\_claude-in-chrome\_\_\* tools, and lists the available skills: /office-hours, /plan-ceo-review, /plan-eng-review, /plan-design-review, /design-consultation, /design-shotgun, /design-html, /review, /ship, /land-and-deploy, /canary, /benchmark, /browse, /connect-chrome, /qa, /qa-only, /design-review, /setup-browser-cookies, /setup-deploy, /setup-gbrain, /retro, /investigate, /document-release, /document-generate, /codex, /cso, /autoplan, /plan-devex-review, /devex-review, /careful, /freeze, /guard, /unfreeze, /gstack-upgrade, /learn. Then ask the user if they also want to add gstack to the current project so teammates get it.
|
||
|
||
### Step 2: Team mode — auto-update for shared repos (recommended)
|
||
|
||
From inside your repo, paste this. Switches you to team mode, bootstraps the repo so teammates get gstack automatically, and commits the change:
|
||
|
||
```bash
|
||
(cd ~/.claude/skills/gstack && ./setup --team) && ~/.claude/skills/gstack/bin/gstack-team-init required && git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
|
||
```
|
||
|
||
No vendored files in your repo, no version drift, no manual upgrades. Every Claude Code session starts with a fast auto-update check (throttled to once/hour, network-failure-safe, completely silent).
|
||
|
||
Swap `required` for `optional` if you'd rather nudge teammates than block them.
|
||
|
||
### OpenClaw
|
||
|
||
OpenClaw spawns Claude Code sessions via ACP, so every gstack skill just works
|
||
when Claude Code has gstack installed. Paste this to your OpenClaw agent:
|
||
|
||
> Install gstack: run `git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack && cd ~/.claude/skills/gstack && ./setup` to install gstack for Claude Code. Then add a "Coding Tasks" section to AGENTS.md that says: when spawning Claude Code sessions for coding work, tell the session to use gstack skills. Include these examples — security audit: "Load gstack. Run /cso", code review: "Load gstack. Run /review", QA test a URL: "Load gstack. Run /qa https://...", build a feature end-to-end: "Load gstack. Run /autoplan, implement the plan, then run /ship", plan before building: "Load gstack. Run /office-hours then /autoplan. Save the plan, don't implement."
|
||
|
||
**After setup, just talk to your OpenClaw agent naturally:**
|
||
|
||
| You say | What happens |
|
||
|---------|-------------|
|
||
| "Fix the typo in README" | Simple — Claude Code session, no gstack needed |
|
||
| "Run a security audit on this repo" | Spawns Claude Code with `Run /cso` |
|
||
| "Build me a notifications feature" | Spawns Claude Code with /autoplan → implement → /ship |
|
||
| "Help me plan the v2 API redesign" | Spawns Claude Code with /office-hours → /autoplan, saves plan |
|
||
|
||
See [docs/OPENCLAW.md](docs/OPENCLAW.md) for advanced dispatch routing and
|
||
the gstack-lite/gstack-full prompt templates.
|
||
|
||
### Native OpenClaw Skills (via ClawHub)
|
||
|
||
Four methodology skills that work directly in your OpenClaw agent, no Claude Code
|
||
session needed. Install from ClawHub:
|
||
|
||
```
|
||
clawhub install gstack-openclaw-office-hours gstack-openclaw-ceo-review gstack-openclaw-investigate gstack-openclaw-retro
|
||
```
|
||
|
||
| Skill | What it does |
|
||
|-------|-------------|
|
||
| `gstack-openclaw-office-hours` | Product interrogation with 6 forcing questions |
|
||
| `gstack-openclaw-ceo-review` | Strategic challenge with 4 scope modes |
|
||
| `gstack-openclaw-investigate` | Root cause debugging methodology |
|
||
| `gstack-openclaw-retro` | Weekly engineering retrospective |
|
||
|
||
These are conversational skills. Your OpenClaw agent runs them directly via chat.
|
||
|
||
### Other AI Agents
|
||
|
||
gstack works on 10 AI coding agents, not just Claude. Setup auto-detects which
|
||
agents you have installed:
|
||
|
||
```bash
|
||
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/gstack
|
||
cd ~/gstack && ./setup
|
||
```
|
||
|
||
Or target a specific agent with `./setup --host <name>`:
|
||
|
||
| Agent | Flag | What you get |
|
||
|-------|------|--------------|
|
||
| OpenAI Codex CLI | `--host codex` | Full install → `${CODEX_HOME:-~/.codex}/skills/gstack-*/` |
|
||
| OpenCode | `--host opencode` | Full install → `~/.config/opencode/skills/gstack-*/` |
|
||
| Cursor | `--host cursor` | Full install → `~/.cursor/skills/gstack-*/` |
|
||
| Factory Droid | `--host factory` | Full install → `~/.factory/skills/gstack-*/` |
|
||
| Kiro | `--host kiro` | Full install → `~/.kiro/skills/gstack-*/` |
|
||
| Slate | `--host slate` | Pointer to the Claude install (Slate reads `.claude/skills` as a fallback) |
|
||
| OpenClaw | `--host openclaw` | ACP spawn pointers + methodology artifacts via `gen:skill-docs --host openclaw` + the instruction-only digest below (full guide: [docs/OPENCLAW.md](docs/OPENCLAW.md)) |
|
||
| Hermes | `--host hermes` | Methodology artifacts via `gen:skill-docs --host hermes` + the instruction-only digest below |
|
||
| GBrain (mod) | `--host gbrain` | Brain-aware skill variants, shipped from the GBrain repo |
|
||
|
||
**Instruction-only tier (any rules-reading agent — Zed, Amp, Jules, side projects):**
|
||
copy the 2KB digest at [`agents-digest/gstack-AGENTS.md`](agents-digest/gstack-AGENTS.md)
|
||
into a location your agent reads (for example, append it to your project's `AGENTS.md`).
|
||
It carries gstack's ethos, reuse ladder, and voice rules — no install required. The
|
||
digest's first line shows its gstack version; re-copy it after upgrading.
|
||
|
||
For Codex, setup reads the top-level `model` from
|
||
`${CODEX_HOME:-~/.codex}/config.toml` and generates the matching behavioral
|
||
profile. `gpt-5.6-sol` automatically receives bounded-scope instructions that
|
||
finish the requested lake without expanding into adjacent cleanup or speculative
|
||
hardening. The Sol profile is exact-match only: dated snapshots and other 5.6
|
||
variants get the generic GPT profile, and setup warns on near-misses like
|
||
`gpt-5.6-sol-2026-08-01`. Override detection with `./setup --host codex --model <id>` — the
|
||
override applies to that run only; set `model` in your Codex `config.toml` to
|
||
make it stick across upgrades. After changing your Codex model, rerun
|
||
`./setup --host codex` to regenerate the skills.
|
||
|
||
**Want to add support for another agent?** See [docs/ADDING_A_HOST.md](docs/ADDING_A_HOST.md).
|
||
It's one TypeScript config file, zero code changes.
|
||
|
||
## See it work
|
||
|
||
```
|
||
You: I want to build a daily briefing app for my calendar.
|
||
You: /office-hours
|
||
Claude: [asks about the pain — specific examples, not hypotheticals]
|
||
|
||
You: Multiple Google calendars, events with stale info, wrong locations.
|
||
Prep takes forever and the results aren't good enough...
|
||
|
||
Claude: I'm going to push back on the framing. You said "daily briefing
|
||
app." But what you actually described is a personal chief of
|
||
staff AI.
|
||
[extracts 5 capabilities you didn't realize you were describing]
|
||
[challenges 4 premises — you agree, disagree, or adjust]
|
||
[generates 3 implementation approaches with effort estimates]
|
||
RECOMMENDATION: Ship the narrowest wedge tomorrow, learn from
|
||
real usage. The full vision is a 3-month project — start with
|
||
the daily briefing that actually works.
|
||
[writes design doc → feeds into downstream skills automatically]
|
||
|
||
You: /plan-ceo-review
|
||
[reads the design doc, challenges scope, runs 10-section review]
|
||
|
||
You: /plan-eng-review
|
||
[ASCII diagrams for data flow, state machines, error paths]
|
||
[test matrix, failure modes, security concerns]
|
||
|
||
You: Approve plan. Exit plan mode.
|
||
[writes 2,400 lines across 11 files. ~8 minutes.]
|
||
|
||
You: /review
|
||
[AUTO-FIXED] 2 issues. [ASK] Race condition → you approve fix.
|
||
|
||
You: /qa https://staging.myapp.com
|
||
[opens real browser, clicks through flows, finds and fixes a bug]
|
||
|
||
You: /ship
|
||
Tests: 42 → 51 (+9 new). PR: github.com/you/app/pull/42
|
||
```
|
||
|
||
You said "daily briefing app." The agent said "you're building a chief of staff AI" — because it listened to your pain, not your feature request. Eight commands, end to end. That is not a copilot. That is a team.
|
||
|
||
## The sprint
|
||
|
||
gstack is a process, not a collection of tools. The skills run in the order a sprint runs:
|
||
|
||
**Think → Plan → Build → Review → Test → Ship → Reflect**
|
||
|
||
Each skill feeds into the next. `/office-hours` writes a design doc that `/plan-ceo-review` reads. `/plan-eng-review` writes a test plan that `/qa` picks up. `/review` catches bugs that `/ship` verifies are fixed. Nothing falls through the cracks because every step knows what came before it.
|
||
|
||
| Skill | Your specialist | What they do |
|
||
|-------|----------------|--------------|
|
||
| `/office-hours` | **YC Office Hours** | Start here. Six forcing questions that reframe your product before you write code. Pushes back on your framing, challenges premises, generates implementation alternatives. Design doc feeds into every downstream skill. |
|
||
| `/plan-ceo-review` | **CEO / Founder** | Rethink the problem. Find the 10-star product hiding inside the request. Four modes: Expansion, Selective Expansion, Hold Scope, Reduction. |
|
||
| `/plan-eng-review` | **Eng Manager** | Lock in architecture, data flow, diagrams, edge cases, and tests. Forces hidden assumptions into the open. |
|
||
| `/plan-design-review` | **Senior Designer** | Rates each design dimension 0-10, explains what a 10 looks like, then edits the plan to get there. AI Slop detection. Interactive — one AskUserQuestion per design choice. |
|
||
| `/plan-devex-review` | **Developer Experience Lead** | Interactive DX review: explores developer personas, benchmarks against competitors' TTHW, designs your magical moment, traces friction points step by step. Three modes: DX EXPANSION, DX POLISH, DX TRIAGE. 20-45 forcing questions. |
|
||
| `/design-consultation` | **Design Partner** | Build a complete design system from scratch. Researches the landscape, proposes creative risks, generates realistic product mockups. |
|
||
| `/review` | **Staff Engineer** | Find the bugs that pass CI but blow up in production. Auto-fixes the obvious ones. Flags completeness gaps. Advisory simplification lens flags over-built code — never blocks, never auto-applies. |
|
||
| `/investigate` | **Debugger** | Systematic root-cause debugging. Iron Law: no fixes without investigation. Traces data flow, tests hypotheses, stops after 3 failed fixes. |
|
||
| `/design-review` | **Designer Who Codes** | Same audit as /plan-design-review, then fixes what it finds. Atomic commits, before/after screenshots. |
|
||
| `/devex-review` | **DX Tester** | Live developer experience audit. Actually tests your onboarding: navigates docs, tries the getting started flow, times TTHW, screenshots errors. Compares against `/plan-devex-review` scores — the boomerang that shows if your plan matched reality. |
|
||
| `/design-shotgun` | **Design Explorer** | "Show me options." Generates 4-6 AI mockup variants, opens a comparison board in your browser, collects your feedback, and iterates. Taste memory learns what you like. Repeat until you love something, then hand it to `/design-html`. |
|
||
| `/design-html` | **Design Engineer** | Turn a mockup into production HTML that actually works. Pretext computed layout: text reflows, heights adjust, layouts are dynamic. 30KB, zero deps. Detects React/Svelte/Vue. Smart API routing per design type (landing page vs dashboard vs form). The output is shippable, not a demo. |
|
||
| `/qa` | **QA Lead** | Test your app, find bugs, fix them with atomic commits, re-verify. Auto-generates regression tests for every fix. |
|
||
| `/qa-only` | **QA Reporter** | Same methodology as /qa but report only. Pure bug report without code changes. |
|
||
| `/pair-agent` | **Multi-Agent Coordinator** | Share your browser with any AI agent. One command, one paste, connected. Works with OpenClaw, Hermes, Codex, Cursor, or anything that can curl. Each agent gets its own tab. Auto-launches headed mode so you watch everything. Auto-starts ngrok tunnel for remote agents. Scoped tokens, tab isolation, rate limiting, activity attribution. |
|
||
| `/cso` | **Chief Security Officer** | OWASP Top 10 + STRIDE threat model. Zero-noise: 17 false positive exclusions, 8/10+ confidence gate, independent finding verification. Each finding includes a concrete exploit scenario. |
|
||
| `/ship` | **Release Engineer** | Sync main, run tests, audit coverage, push, open PR. Bootstraps test frameworks if you don't have one. |
|
||
| `/land-and-deploy` | **Release Engineer** | Merge the PR, wait for CI and deploy, verify production health. One command from "approved" to "verified in production." |
|
||
| `/canary` | **SRE** | Post-deploy monitoring loop. Watches for console errors, performance regressions, and page failures. |
|
||
| `/benchmark` | **Performance Engineer** | Baseline page load times, Core Web Vitals, and resource sizes. Compare before/after on every PR. |
|
||
| `/document-release` | **Technical Writer** | Update all project docs to match what you just shipped. Catches stale READMEs automatically. Builds a Diataxis coverage map (reference / how-to / tutorial / explanation) so gaps are visible in the PR body. |
|
||
| `/document-generate` | **Documentation Author** | Generate missing docs from scratch using the Diataxis framework. Researches the codebase first, then writes reference / how-to / tutorial / explanation docs that actually match the code. Invokable standalone or chained from `/document-release` when the coverage map finds gaps. Learn more: [tutorial](docs/tutorial-document-generate.md) • [how-to](docs/howto-document-a-shipped-feature.md) • [why Diataxis](docs/explanation-diataxis-in-gstack.md). |
|
||
| `/retro` | **Eng Manager** | Team-aware weekly retro. Per-person breakdowns, shipping streaks, test health trends, growth opportunities. `/retro global` runs across all your projects and AI tools (Claude Code, Codex, Gemini). |
|
||
| `/browse` | **QA Engineer** | Give the agent eyes. Real Chromium browser, real clicks, real screenshots. ~100ms per command. `/open-gstack-browser` launches GStack Browser with sidebar, anti-bot stealth, and auto model routing. |
|
||
| `/setup-browser-cookies` | **Session Manager** | Import cookies from your real browser (Chrome, Arc, Brave, Edge) into the headless session. Test authenticated pages. |
|
||
| `/autoplan` | **Review Pipeline** | One command, fully reviewed plan. Runs CEO → design → DX → eng review automatically (eng always last, so the shipping gate reviews the final amended plan) with encoded decision principles. Surfaces only taste decisions for your approval. |
|
||
| `/spec` | **Spec Author** | Turn vague intent into a precise, executable spec in five phases (why, scope, technical with mandatory code-reading, draft, file). Codex quality gate before file (blocks below 7/10), fail-closed secret redaction, dedupe against existing issues, archive to `$GSTACK_STATE_ROOT/projects/$SLUG/specs/` for team-corpus recall. `--execute` spawns `claude -p` in a fresh worktree; `/ship` auto-closes the source issue on merge. Plan-mode aware. |
|
||
| `/learn` | **Memory** | Manage what gstack learned across sessions. Review, search, prune, and export project-specific patterns, pitfalls, and preferences. Learnings compound across sessions so gstack gets smarter on your codebase over time. |
|
||
| `/make-pdf` | **Publisher** | Markdown in, publication-quality document out. Mermaid and excalidraw fences render as vector diagrams, fully offline. Images scale to the page and never truncate; wide diagrams get their own landscape page. `--to html` emits one self-contained file, `--to docx` a Word doc. |
|
||
| `/diagram` | **Diagram Maker** | English in, editable diagram out. Emits a triplet: mermaid source, `.excalidraw` you can open and edit on excalidraw.com (hand-drawn style), and rendered SVG/PNG. Zero network. Embed the source in markdown and `/make-pdf` renders it. |
|
||
|
||
### Which review should I use?
|
||
|
||
| Building for... | Plan stage (before code) | Live audit (after shipping) |
|
||
|-----------------|--------------------------|----------------------------|
|
||
| **End users** (UI, web app, mobile) | `/plan-design-review` | `/design-review` |
|
||
| **Developers** (API, CLI, SDK, docs) | `/plan-devex-review` | `/devex-review` |
|
||
| **Architecture** (data flow, perf, tests) | `/plan-eng-review` | `/review` |
|
||
| **All of the above** | `/autoplan` (runs CEO → design → DX → eng, auto-detects which apply; eng always last) | — |
|
||
|
||
### Power tools
|
||
|
||
| Skill | What it does |
|
||
|-------|-------------|
|
||
| `/codex` | **Second Opinion** — independent code review from OpenAI Codex CLI. Three modes: review (pass/fail gate), adversarial challenge, and open consultation. Cross-model analysis when both `/review` and `/codex` have run. |
|
||
| `/careful` | **Safety Guardrails** — warns before destructive commands (rm -rf, DROP TABLE, force-push). Say "be careful" to activate. Override any MEDIUM warning; root/home recursive deletes and default-branch force-pushes are hard-denied. |
|
||
| `/freeze` | **Edit Lock** — restrict file edits to one directory. Prevents accidental changes outside scope while debugging. |
|
||
| `/guard` | **Full Safety** — `/careful` + `/freeze` in one command. Maximum safety for prod work. |
|
||
| `/unfreeze` | **Unlock** — remove the `/freeze` boundary. |
|
||
| `/open-gstack-browser` | **GStack Browser** — launch GStack Browser with sidebar, anti-bot stealth, auto model routing (Sonnet for actions, Opus for analysis), one-click cookie import, and Claude Code integration. Clean up pages, take smart screenshots, edit CSS, and pass info back to your terminal. |
|
||
| `/setup-deploy` | **Deploy Configurator** — one-time setup for `/land-and-deploy`. Detects your platform, production URL, and deploy commands. |
|
||
| `/setup-gbrain` | **GBrain Onboarding** — from zero to running gbrain in under 5 minutes. PGLite local, Supabase existing URL, or auto-provision a new Supabase project via Management API. MCP registration for Claude Code + per-repo trust triad (read-write/read-only/deny). [Full guide](USING_GBRAIN_WITH_GSTACK.md). |
|
||
| `/sync-gbrain` | **Keep Brain Current** — re-index this repo's code into gbrain via `gbrain sources add` + `gbrain sync --strategy code`, refresh the `## GBrain Search Guidance` block in CLAUDE.md, and auto-remove guidance when the capability check fails. `--incremental` (default), `--full`, `--dry-run`. Idempotent; safe to re-run. |
|
||
| `/gstack-upgrade` | **Self-Updater** — upgrade gstack to latest. Detects global vs vendored install, syncs both, shows what changed. |
|
||
| `/ios-qa` | **iOS Live-Device QA (v1.43.0.0+)** — drive a real iPhone over USB CoreDevice via an embedded `StateServer` in the app. Read Swift source, codegen typed `@Observable` accessors, run the agent loop. Optional `--tailnet` flag exposes the device to OpenClaw or any HTTP-capable agent on your Tailscale tailnet so remote agents can run iOS QA without ever touching the hardware. Capability-tier allowlist (observe/interact/mutate/restore), per-device session lock, audit log. |
|
||
| `/ios-fix`, `/ios-design-review`, `/ios-clean`, `/ios-sync` | iOS bug-fix loop, designer's-eye HIG audit, debug-bridge cleanup, and accessor resync. See `docs/skills.md`. End-to-end walkthrough: [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-testing-with-gstack.md). |
|
||
|
||
### Standalone binaries
|
||
|
||
Beyond the slash-command skills, gstack ships standalone CLIs for workflows that don't belong inside a session:
|
||
|
||
| Command | What it does |
|
||
|---------|-------------|
|
||
| `gstack-model-benchmark` | **Cross-model benchmark** — run the same prompt through Claude, GPT (via Codex CLI), and Gemini; compare latency, tokens, cost, and (optionally) LLM-judge quality score. Auth detected per provider, unavailable providers skip cleanly. Output as table, JSON, or markdown. `--dry-run` validates flags + auth without spending API calls. |
|
||
| `gstack-taste-update` | **Design taste learning** — writes approvals and rejections from `/design-shotgun` into a persistent per-project taste profile. Decays 5%/week. Feeds back into future variant generation so the system learns what you actually pick. |
|
||
| `gstack-egress` | **Egress receipt auditor** — every gstack-initiated off-machine send writes a tamper-evident, hash-chained receipt to `~/.gstack/security/egress.jsonl` before the send. `list` shows what gstack attempted to send and to which host, `grants` shows the standing consent settings plus the exact command that revokes each, `verify` recomputes the hash chain and exits 3 on tamper (catches edits, reordering, and mid-chain deletion; truncating or deleting the ledger itself is out of scope — it's a forensic log, not tamper-proof storage). |
|
||
| `gstack-context-bill` | **Token bill-of-materials** — read-only, offline audit of what an installed skills tree costs in tokens: always-on frontmatter every session pays vs per-invocation SKILL.md + forced references. `--diff` compares two trees, `--budget` enforces a ceiling, `--exact` opts into Anthropic `count_tokens` (sends file text off-machine; writes an egress receipt first, degrades to the offline estimate if the receipt can't be written). |
|
||
| `gstack-code-intelligence` | **Code-intelligence provider picker** — wraps GBrain, Sourcebot, and Graphify behind one interface: `options`/`status` to see what's available, `select` to pick one, `index`/`search` to use it, `suggest` to check whether the one-time indexing offer should fire here. The offer triggers on large repos (1,000+ tracked files; a decline is persisted). Non-local providers refuse to index *or search* until you record per-repo consent (`consent <repo> yes\|no` — the query text is repo-derived content), the per-repo trust policy's deny and read-only tiers veto write-class operations regardless of consent, and every off-machine send writes an egress receipt. Fully optional — with nothing selected, gstack falls back to grep. |
|
||
| `gstack-verify-gate` | **Verification stop hook (opt-in)** — blocks a Claude Code turn from ending until the project's declared verify command passes (after 3 blocked re-entries it yields with a loud still-RED warning instead of looping forever). Declare it on one line in CLAUDE.md: `<!-- gstack:verify: bun test -->`. Hooks bypass the permission system, so a declared command never runs until you trust it once per repo (`gstack-verify-gate --trust`); editing the command invalidates trust until re-granted, and every grant is audit-logged. `./setup` never registers it for you — opt in with `gstack-settings-hook add-event --event Stop --command ~/.claude/skills/gstack/bin/gstack-verify-gate --source verify-gate`, remove with `gstack-settings-hook remove-source --source verify-gate`. |
|
||
| `gstack-wtree` | **Working-tree fingerprint** — prints a content hash of what's actually on disk (temp index seeded from the stat cache, ~40x cheaper than a full re-hash; untracked source counts, gitignored scratch doesn't). Identical content fingerprints identically through commits, rebases, amends, and squashes — it's what binds reviews and test evidence to content instead of commit SHAs. |
|
||
| `gstack-evidence` | **Verification-evidence ledger** — `run --label <lane> -- <cmd>` transparently wraps any test command (the child's exit code always passes through) and records what ran against which working-tree fingerprint; `check` grades each label FRESH/STALE/MISSING with `--expect-cmd`, `--max-age`, and `--allow-paths` binding. /ship and /land-and-deploy cite fresh evidence instead of re-running suites. Per-run logs are 0600, capped at 2MB, pruned after 30 days; the ledger and logs stay machine-local by design. |
|
||
| `gstack-issue-guard` | **Tracker-text trust envelope** — fetches GitHub issue/PR text (`issue <n>`, `pr-body`, `pr-comments`, or `--stdin`) and wraps it in a labeled envelope so agents treat it as data: injection-shaped lines get labeled even through fullwidth and invisible-character evasion, and forged envelope banners are defused. Every tracker-text ingress in gstack routes through it, enforced by a CI scanner. |
|
||
| `gstack-ios-qa-daemon` | **iOS QA daemon** — Mac-side broker between an agent and a connected iPhone over USB CoreDevice. Loopback by default; `--tailnet` opens a Tailscale-facing listener with identity-gated capability tiers. Single-instance via flock on `~/.gstack/ios-qa-daemon.pid`. See [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-testing-with-gstack.md). |
|
||
| `gstack-ios-qa-mint` | **iOS allowlist manager** — owner-grant CLI for the tailnet allowlist. `grant`/`revoke`/`list` against `~/.gstack/ios-qa-allowlist.json` (mode 0600). Remote agents never auto-allowlist; this is the explicit-intent path. |
|
||
| `gstack-ios-qa-regen` | **iOS bridge regenerator** — deterministically installs the canonical DebugBridge package, generates typed state accessors, and records the installed gstack version. Safe to rerun after source changes or upgrades. |
|
||
|
||
`./setup` also registers one default-on Stop hook in `~/.claude/settings.json`:
|
||
`gstack-timeline-stop` (closes dangling session-timeline entries when a session
|
||
is interrupted; fail-open — 2s internal budget, always exits 0, can never block
|
||
a session). Skip it with `./setup --no-team`, remove it with
|
||
`gstack-settings-hook remove-source --source gstack-timeline-stop`;
|
||
`gstack-uninstall` removes it too.
|
||
|
||
Hook registration is canonical-only: every hook command points at the stable
|
||
`~/.claude/skills/gstack` install, never the tree setup ran from, so deleting
|
||
a worktree or Conductor workspace can't leave dead hooks erroring in your
|
||
sessions. Every `./setup` run also heals first: `gstack-settings-hook
|
||
prune-stale --repoint` removes dead gstack hook entries, re-points stale ones
|
||
at the stable install, and collapses duplicates, printing one line (and
|
||
writing a backup beside the file) only when it changed something.
|
||
|
||
### Continuous checkpoint mode (opt-in, local by default)
|
||
|
||
Set `gstack-config set checkpoint_mode continuous` and skills auto-commit your work as you go with a `WIP:` prefix plus a structured `[gstack-context]` body (decisions, remaining work, failed approaches). Survives crashes and context switches. `/context-restore` reads those commits to reconstruct session state. `/ship` filter-squashes WIP commits before the PR (preserving non-WIP commits) so bisect stays clean. Push is opt-in via `checkpoint_push=true` — default is local-only so you don't trigger CI on every WIP commit.
|
||
|
||
### Domain skills + raw CDP escape hatch
|
||
|
||
Two new browser primitives compound the gstack agent over time:
|
||
|
||
- **`$B domain-skill save`** — agent saves a per-site note (e.g., "LinkedIn's Apply button lives in an iframe") that fires automatically next time it visits that hostname. Quarantined → active after 3 successful uses → optional cross-project promotion via `$B domain-skill promote-to-global`. Storage lives alongside `/learn`'s per-project learnings file. Full reference: **[docs/domain-skills.md](docs/domain-skills.md)**.
|
||
- **`$B cdp <Domain.method>`** — raw Chrome DevTools Protocol escape hatch for the rare case curated commands miss. Deny-default: methods must be explicitly added to `browse/src/cdp-allowlist.ts` with a one-line justification. Two-tier mutex serializes browser-scoped CDP calls against per-tab work. Output for data-exfil methods is wrapped in the UNTRUSTED envelope.
|
||
|
||
> Want raw CDP with no rails, no allowlist, no daemon — just thin transport from agent to Chrome? [browser-use/browser-harness-js](https://github.com/browser-use/browser-harness-js) is a different philosophy (agent-authored helpers vs gstack's curated commands) and a good fit if you don't want gstack's security stack. The two can coexist: gstack's `$B cdp` and harness can both attach to the same Chrome via Playwright's `newCDPSession`.
|
||
|
||
**[Deep dives with examples and philosophy for every skill →](docs/skills.md)**
|
||
|
||
### Karpathy's four failure modes? Already covered.
|
||
|
||
Andrej Karpathy's [AI coding rules](https://github.com/forrestchang/andrej-karpathy-skills) (17K stars) nail four failure modes: wrong assumptions, overcomplexity, orthogonal edits, imperative over declarative. gstack's workflow skills enforce all four. `/office-hours` forces assumptions into the open before code is written. The Confusion Protocol stops Claude from guessing on architectural decisions. `/review` catches unnecessary complexity and drive-by edits. `/ship` transforms tasks into verifiable goals with test-first execution. If you already use Karpathy-style CLAUDE.md rules, gstack is the workflow enforcement layer that makes them stick across entire sprints, not just single prompts.
|
||
|
||
## Parallel sprints
|
||
|
||
gstack works well with one sprint. It gets interesting with ten running at once.
|
||
|
||
**Design is at the heart.** `/design-consultation` builds your design system from scratch, researches what's out there, proposes creative risks, and writes `DESIGN.md`. But the real magic is the shotgun-to-HTML pipeline.
|
||
|
||
**`/design-shotgun` is how you explore.** You describe what you want. It generates 4-6 AI mockup variants using GPT Image. Then it opens a comparison board in your browser with all variants side by side. You pick favorites, leave feedback ("more whitespace", "bolder headline", "lose the gradient"), and it generates a new round. Repeat until you love something. Taste memory kicks in after a few rounds so it starts biasing toward what you actually like. No more describing your vision in words and hoping the AI gets it. You see options, pick the good ones, and iterate visually.
|
||
|
||
**`/design-html` makes it real.** Take that approved mockup (from `/design-shotgun`, a CEO plan, a design review, or just a description) and turn it into production-quality HTML/CSS. Not the kind of AI HTML that looks fine at one viewport width and breaks everywhere else. This uses Pretext for computed text layout: text actually reflows on resize, heights adjust to content, layouts are dynamic. 30KB overhead, zero dependencies. It detects your framework (React, Svelte, Vue) and outputs the right format. Smart API routing picks different Pretext patterns depending on whether it's a landing page, dashboard, form, or card layout. The output is something you'd actually ship, not a demo.
|
||
|
||
**`/qa` was a massive unlock.** It let me go from 6 to 12 parallel workers. Claude Code saying *"I SEE THE ISSUE"* and then actually fixing it, generating a regression test, and verifying the fix — that changed how I work. The agent has eyes now.
|
||
|
||
**Smart review routing.** Just like at a well-run startup: CEO doesn't have to look at infra bug fixes, design review isn't needed for backend changes. gstack tracks what reviews are run, figures out what's appropriate, and just does the smart thing. The Review Readiness Dashboard tells you where you stand before you ship.
|
||
|
||
**Test everything.** `/ship` bootstraps test frameworks from scratch if your project doesn't have one. Every `/ship` run produces a coverage audit. Every `/qa` bug fix generates a regression test. 100% test coverage is the goal — tests make vibe coding safe instead of yolo coding.
|
||
|
||
**`/document-release` is the engineer you never had.** It reads every doc file in your project, cross-references the diff, and updates everything that drifted. README, ARCHITECTURE, CONTRIBUTING, CLAUDE.md, TODOS — all kept current automatically. And now `/ship` auto-invokes it — docs stay current without an extra command.
|
||
|
||
**Real browser mode.** `/open-gstack-browser` launches GStack Browser, an AI-controlled Chromium with anti-bot stealth, custom branding, and the sidebar extension baked in. Sites like Google and NYTimes work without captchas. The menu bar says "GStack Browser" instead of "Chrome for Testing." Your regular Chrome stays untouched. All existing browse commands work unchanged. `$B disconnect` returns to headless. The browser stays alive as long as the window is open... no idle timeout killing it while you're working.
|
||
|
||
**Sidebar agent — your AI browser assistant.** Type natural language in the Chrome side panel and a child Claude instance executes it. "Navigate to the settings page and screenshot it." "Fill out this form with test data." "Go through every item in this list and extract the prices." The sidebar auto-routes to the right model: Sonnet for fast actions (click, navigate, screenshot) and Opus for reading and analysis. Each task gets up to 5 minutes. The sidebar agent runs in an isolated session, so it won't interfere with your main Claude Code window. One-click cookie import right from the sidebar footer.
|
||
|
||
**Personal automation.** The sidebar agent isn't just for dev workflows. Example: "Browse my kid's school parent portal and add all the other parents' names, phone numbers, and photos to my Google Contacts." Two ways to get authenticated: (1) log in once in the headed browser, your session persists, or (2) click the "cookies" button in the sidebar footer to import cookies from your real Chrome. Once authenticated, Claude navigates the directory, extracts the data, and creates the contacts.
|
||
|
||
**Prompt injection defense.** Hostile web pages try to hijack your sidebar agent. gstack ships a layered defense: content filters (datamarking, hidden-element stripping, ARIA scrubbing, URL blocklist) on every page read, plus a 22MB ML classifier running locally in a sidecar subprocess that scans page-derived content before the agent sees it, with a verdict combiner that requires classifier agreement before blocking (prevents single-model false positives on Stack Overflow-style instruction pages). Everything runs on your machine, no network calls. Emergency kill switch: `GSTACK_SECURITY_OFF=1`. See [ARCHITECTURE.md](ARCHITECTURE.md#prompt-injection-defense-sidebar-agent) for the full stack.
|
||
|
||
**Browser handoff when the AI gets stuck.** Hit a CAPTCHA, auth wall, or MFA prompt? `$B handoff` opens a visible Chrome at the exact same page with all your cookies and tabs intact. Solve the problem, tell Claude you're done, `$B resume` picks up right where it left off. The agent even suggests it automatically after 3 consecutive failures.
|
||
|
||
**`/pair-agent` is cross-agent coordination.** You're in Claude Code. You also have OpenClaw running. Or Hermes. Or Codex. You want them both looking at the same website. Type `/pair-agent`, pick your agent, and a GStack Browser window opens so you can watch. The skill prints a block of instructions. Paste that block into the other agent's chat. It exchanges a one-time setup key for a session token, creates its own tab, and starts browsing. You see both agents working in the same browser, each in their own tab, neither able to interfere with the other. If ngrok is installed, the tunnel starts automatically so the other agent can be on a completely different machine. Same-machine agents get a zero-friction shortcut that writes credentials directly. This is the first time AI agents from different vendors can coordinate through a shared browser with real security: scoped tokens, tab isolation, rate limiting, domain restrictions, and activity attribution.
|
||
|
||
**Multi-AI second opinion.** `/codex` gets an independent review from OpenAI's Codex CLI — a completely different AI looking at the same diff. Three modes: code review with a pass/fail gate, adversarial challenge that actively tries to break your code, and open consultation with session continuity. When both `/review` (Claude) and `/codex` (OpenAI) have reviewed the same branch, you get a cross-model analysis showing which findings overlap and which are unique to each.
|
||
|
||
**Safety guardrails on demand.** Say "be careful" and `/careful` warns before any destructive command — rm -rf, DROP TABLE, force-push, git reset --hard. `/freeze` locks edits to one directory while debugging so Claude can't accidentally "fix" unrelated code. `/guard` activates both. `/investigate` auto-freezes to the module being investigated.
|
||
|
||
**Proactive skill suggestions.** gstack notices what stage you're in — brainstorming, reviewing, debugging, testing — and suggests the right skill. Don't like it? Say "stop suggesting" and it remembers across sessions.
|
||
|
||
## 10-15 parallel sprints
|
||
|
||
gstack is powerful with one sprint. It is transformative with ten running at once.
|
||
|
||
[Conductor](https://conductor.build) runs multiple Claude Code sessions in parallel — each in its own isolated workspace. One session running `/office-hours` on a new idea, another doing `/review` on a PR, a third implementing a feature, a fourth running `/qa` on staging, and six more on other branches. All at the same time. I regularly run 10-15 parallel sprints — that's the practical max right now.
|
||
|
||
The sprint structure is what makes parallelism work. Without a process, ten agents is ten sources of chaos. With a process — think, plan, build, review, test, ship — each agent knows exactly what to do and when to stop. You manage them the way a CEO manages a team: check in on the decisions that matter, let the rest run.
|
||
|
||
### Voice input (AquaVoice, Whisper, etc.)
|
||
|
||
gstack skills have voice-friendly trigger phrases. Say what you want naturally —
|
||
"run a security check", "test the website", "do an engineering review" — and the
|
||
right skill activates. You don't need to remember slash command names or acronyms.
|
||
|
||
## Uninstall
|
||
|
||
### Option 1: Run the uninstall script
|
||
|
||
If gstack is installed on your machine:
|
||
|
||
```bash
|
||
~/.claude/skills/gstack/bin/gstack-uninstall
|
||
```
|
||
|
||
This handles skills, symlinks, global state (`~/.gstack/`), project-local state, browse daemons, and temp files. Use `--keep-state` to preserve config and analytics. Use `--force` to skip confirmation.
|
||
|
||
### Option 2: Manual removal (no local repo)
|
||
|
||
If you don't have the repo cloned (e.g. you installed via a Claude Code paste and later deleted the clone):
|
||
|
||
```bash
|
||
# 1. Stop browse daemons
|
||
pkill -f "gstack.*browse" 2>/dev/null || true
|
||
|
||
# 2. Remove per-skill directories whose SKILL.md points into gstack/
|
||
# (rm -rf, not rmdir — installed dirs also contain runtime-asset links)
|
||
find ~/.claude/skills -mindepth 1 -maxdepth 1 -type d ! -name gstack 2>/dev/null |
|
||
while IFS= read -r dir; do
|
||
link="$dir/SKILL.md"
|
||
[ -L "$link" ] || continue
|
||
target=$(readlink "$link" 2>/dev/null) || continue
|
||
case "$target" in
|
||
gstack/*|*/gstack/*)
|
||
rm -rf "$dir"
|
||
;;
|
||
esac
|
||
done
|
||
# Alias skills install as copies (no symlink to detect) — remove by name
|
||
rm -rf ~/.claude/skills/_gstack-command ~/.claude/skills/connect-chrome 2>/dev/null
|
||
|
||
# 3. Remove gstack
|
||
rm -rf ~/.claude/skills/gstack
|
||
|
||
# 4. Remove global state
|
||
rm -rf ~/.gstack
|
||
|
||
# 5. Remove integrations (skip any you never installed)
|
||
rm -rf "${CODEX_HOME:-$HOME/.codex}/skills/gstack"* 2>/dev/null
|
||
rm -rf ~/.factory/skills/gstack* 2>/dev/null
|
||
rm -rf ~/.kiro/skills/gstack* 2>/dev/null
|
||
rm -rf ~/.openclaw/skills/gstack* 2>/dev/null
|
||
rm -rf ~/.cursor/skills/gstack* 2>/dev/null
|
||
rm -rf ~/.config/opencode/skills/gstack* 2>/dev/null
|
||
|
||
# 6. Remove temp files
|
||
rm -f /tmp/gstack-* 2>/dev/null
|
||
|
||
# 7. Per-project cleanup (run from each project root)
|
||
rm -rf .gstack .gstack-worktrees .claude/skills/gstack 2>/dev/null
|
||
rm -rf .agents/skills/gstack* .factory/skills/gstack* 2>/dev/null
|
||
```
|
||
|
||
Manual removal leaves gstack's hook entries behind in `~/.claude/settings.json`
|
||
(the uninstall script removes all of them for you, including entries whose
|
||
`_gstack_source` tag was stripped). Edit that file and delete every hook whose
|
||
command path points into `.claude/skills/gstack/`: the SessionStart auto-update
|
||
hook, the AskUserQuestion PreToolUse/PostToolUse hooks, and the Stop hooks
|
||
(session timeline, plus verify-gate if you opted in). Left in place, they error
|
||
on every matching event once the install directory is gone.
|
||
|
||
### Clean up CLAUDE.md
|
||
|
||
The uninstall script does not edit CLAUDE.md. In each project where gstack was added, remove the `## gstack` and `## Skill routing` sections.
|
||
|
||
### Playwright
|
||
|
||
`~/Library/Caches/ms-playwright/` (macOS) is left in place because other tools may share it. Remove it if nothing else needs it.
|
||
|
||
---
|
||
|
||
Free, MIT licensed, open source. No premium tier, no waitlist.
|
||
|
||
I open sourced how I build software. You can fork it and make it your own.
|
||
|
||
> **We're hiring.** Want to ship real products at AI-coding speed and help harden gstack?
|
||
> Come work at YC — [ycombinator.com/software](https://ycombinator.com/software)
|
||
> Extremely competitive salary and equity. San Francisco, Dogpatch District.
|
||
|
||
## GBrain — persistent knowledge for your coding agent
|
||
|
||
[GBrain](https://github.com/garrytan/gbrain) is a persistent knowledge base for AI agents — think of it as the memory your agent actually keeps between sessions. GStack gives you a one-command path from zero to "it's running, my agent can call it."
|
||
|
||
```bash
|
||
/setup-gbrain
|
||
```
|
||
|
||
Four paths, pick one:
|
||
|
||
- **Supabase, existing URL** — your cloud agent already provisioned a brain; paste the Session Pooler URL, now this laptop uses the same data.
|
||
- **Supabase, auto-provision** — paste a Supabase Personal Access Token; the skill creates a new project, polls to healthy, fetches the pooler URL, hands it to `gbrain init`. ~90 seconds end-to-end.
|
||
- **PGLite local** — zero accounts, zero network, ~30 seconds. Isolated brain on this Mac only. Great for try-first; migrate to Supabase later with `/setup-gbrain --switch`.
|
||
- **Remote gbrain MCP** — your brain runs on another machine (Tailscale, ngrok, internal LAN) or a teammate's server; paste an MCP URL and bearer token. Optionally pair with a local PGLite for symbol-aware code search in split-engine mode. Best for cross-machine memory without standing up a local DB.
|
||
|
||
After init, the skill offers to register gbrain as an MCP server for Claude Code (`claude mcp add gbrain -- gbrain serve`) so `gbrain search`, `gbrain put`, etc. show up as first-class typed tools — not bash shell-outs.
|
||
|
||
**Keeping the brain current.** Run `/sync-gbrain` from any repo to re-index its code into gbrain (incremental by default, `--full` for a full reindex, `--dry-run` to preview). The skill registers the cwd as a federated source via `gbrain sources add`, runs `gbrain sync --strategy code`, and writes a `## GBrain Search Guidance` block to your project's CLAUDE.md so the agent prefers `gbrain search`/`code-def`/`code-refs` over Grep. The block is removed automatically if the capability check fails — no stale guidance pointing at tools that aren't installed.
|
||
|
||
**Per-remote trust policy.** Each repo on your machine gets one of three tiers:
|
||
|
||
- `read-write` — agent can search the brain AND write new pages back from this repo
|
||
- `read-only` — agent can search but never writes (best for multi-client consultants: search the shared brain, don't contaminate it with Client A's work while in Client B's repo)
|
||
- `deny` — no gbrain interaction at all
|
||
|
||
The skill asks once per repo. The decision is sticky across worktrees and branches of the same remote.
|
||
|
||
**GStack memory sync (different feature, same private-repo infra).** Optionally pushes your gstack state (learnings, CEO plans, design docs, retros, developer profile) to a private git repo so your memory follows you across machines, with a one-time privacy prompt (everything allowlisted / artifacts only / off) and a defense-in-depth secret scanner that blocks AWS keys, tokens, PEM blocks, and JWTs before they leave your machine.
|
||
|
||
```bash
|
||
gstack-artifacts-init
|
||
```
|
||
|
||
**Running gstack in Conductor?** Conductor explicitly strips `ANTHROPIC_API_KEY` and `OPENAI_API_KEY` from every workspace's process env, so paid evals and gbrain embeddings won't work out of the box. Set `GSTACK_ANTHROPIC_API_KEY` and `GSTACK_OPENAI_API_KEY` in Conductor's workspace env config instead — gstack's TS entry points promote them to canonical names at runtime. Full details and the contributor checklist for adding the import to new entry points: [Conductor + GSTACK_* env vars](USING_GBRAIN_WITH_GSTACK.md#conductor--gstack_-env-vars).
|
||
|
||
**Full monty — every scenario, every flag, every bin helper, every troubleshooting step:** [USING_GBRAIN_WITH_GSTACK.md](USING_GBRAIN_WITH_GSTACK.md)
|
||
|
||
Other references: [docs/gbrain-sync.md](docs/gbrain-sync.md) (sync-specific guide) • [docs/gbrain-sync-errors.md](docs/gbrain-sync-errors.md) (error index)
|
||
|
||
## Docs
|
||
|
||
| Doc | What it covers |
|
||
|-----|---------------|
|
||
| [Skill Deep Dives](docs/skills.md) | Philosophy, examples, and workflow for every skill (includes Greptile integration) |
|
||
| [Diagrams & Document Formats](docs/howto-diagrams-and-formats.md) | Mermaid/excalidraw fences in PDFs, image sizing and safety defaults, `--to html\|docx`, `/diagram` triplets |
|
||
| [Builder Ethos](ETHOS.md) | Builder philosophy: Boil the Ocean, Search Before Building, three layers of knowledge |
|
||
| [Using GBrain with GStack](USING_GBRAIN_WITH_GSTACK.md) | Every path, flag, bin helper, and troubleshooting step for `/setup-gbrain` |
|
||
| [GBrain Sync](docs/gbrain-sync.md) | Cross-machine memory setup, privacy modes, troubleshooting |
|
||
| [Architecture](ARCHITECTURE.md) | Design decisions and system internals |
|
||
| [Browser Reference](BROWSER.md) | Full command reference for `/browse` |
|
||
| [Contributing](CONTRIBUTING.md) | Dev setup, testing, contributor mode, and dev mode |
|
||
| [Changelog](CHANGELOG.md) | What's new in every version |
|
||
|
||
## Privacy & Telemetry
|
||
|
||
gstack includes **opt-in** usage telemetry to help improve the project. Here's exactly what happens:
|
||
|
||
- **Default is off.** Nothing is sent anywhere unless you explicitly say yes.
|
||
- **On first run,** gstack asks if you want to share anonymous usage data. You can say no.
|
||
- **What's sent (if you opt in):** skill name, duration, success/fail, gstack version, OS. That's it.
|
||
- **What's never sent:** code, file paths, repo names, branch names, prompts, or any user-generated content.
|
||
- **Change anytime:** `gstack-config set telemetry off` disables everything instantly.
|
||
- **Every off-machine send is receipted.** Any gstack-initiated network send — telemetry included — writes a hash-chained, tamper-evident receipt to `~/.gstack/security/egress.jsonl` before the send; sensitive sinks refuse to send at all if the receipt can't be written. Audit with `gstack-egress list`, verify the chain with `gstack-egress verify` (exit 3 on tamper), see the standing consent settings with `gstack-egress grants`. The ledger records attempted sends so accidents are auditable — it's an audit trail, not a network firewall.
|
||
|
||
Data is stored in [Supabase](https://supabase.com) (open source Firebase alternative). The schema is in [`supabase/migrations/`](supabase/migrations/) — you can verify exactly what's collected. The Supabase publishable key in the repo is a public key (like a Firebase API key) — row-level security policies deny all direct access. Telemetry flows through validated edge functions that enforce schema checks, event type allowlists, and field length limits.
|
||
|
||
**Local analytics are always available.** Run `gstack-analytics` to see your personal usage dashboard from the local JSONL file — no remote data needed.
|
||
|
||
## Troubleshooting
|
||
|
||
**Skill not showing up?** `cd ~/.claude/skills/gstack && ./setup`
|
||
|
||
**`/browse` fails?** `cd ~/.claude/skills/gstack && bun install && bun run build`
|
||
|
||
**Stale install?** Run `/gstack-upgrade` — or set `auto_upgrade: true` in `~/.gstack/config.yaml`
|
||
|
||
**Want shorter commands?** `cd ~/.claude/skills/gstack && ./setup --no-prefix` — switches from `/gstack-qa` to `/qa`. Your choice is remembered for future upgrades.
|
||
|
||
**Want namespaced commands?** `cd ~/.claude/skills/gstack && ./setup --prefix` — switches from `/qa` to `/gstack-qa`. Useful if you run other skill packs alongside gstack.
|
||
|
||
**Codex says "Skipped loading skill(s) due to invalid SKILL.md"?** Your Codex skill descriptions are stale. Fix: `cd "${CODEX_HOME:-$HOME/.codex}/skills/gstack" && git pull && ./setup --host codex` — or for repo-local installs: `cd "$(readlink -f .agents/skills/gstack)" && git pull && ./setup --host codex`
|
||
|
||
**Windows users:** gstack works on Windows 11 via Git Bash or WSL. Node.js is required in addition to Bun — Bun has a known bug with Playwright's pipe transport on Windows ([bun#4253](https://github.com/oven-sh/bun/issues/4253)). The browse server automatically falls back to Node.js. Make sure both `bun` and `node` are on your PATH.
|
||
|
||
On Windows without Developer Mode (MSYS2 / Git Bash), `setup` falls back to file copies instead of symlinks because `ln -snf` produces frozen copies that don't refresh on `git pull`. **Re-run `cd ~/.claude/skills/gstack && ./setup` after every `git pull`** so your skill files match the repo. `setup` prints a one-line note reminding you. Unix and WSL keep symlinks and don't need the re-run.
|
||
|
||
**Claude says it can't see the skills?** Make sure your project's `CLAUDE.md` has a gstack section. Add this:
|
||
|
||
```
|
||
## gstack
|
||
Use /browse from gstack for all web browsing. Never use mcp__claude-in-chrome__* tools.
|
||
Available skills: /office-hours, /plan-ceo-review, /plan-eng-review, /plan-design-review,
|
||
/design-consultation, /design-shotgun, /design-html, /review, /ship, /land-and-deploy,
|
||
/canary, /benchmark, /browse, /open-gstack-browser, /qa, /qa-only, /design-review,
|
||
/setup-browser-cookies, /setup-deploy, /setup-gbrain, /sync-gbrain, /retro, /investigate,
|
||
/document-release, /document-generate, /codex, /cso, /autoplan, /pair-agent, /careful, /freeze,
|
||
/guard, /unfreeze, /gstack-upgrade, /learn.
|
||
```
|
||
|
||
## License
|
||
|
||
MIT. Free forever. Go build something.
|