Files
gstack/BROWSER.md
T
Garry TanandClaude Fable 5 c118e2402e v1.64.1.0 v1.64.1.0: the code-smell fix wave — every pipeline guard now provably fires (net −24,943 lines) (#2572)
* fix(ci): skill-docs freshness gate covers all 10 hosts and can actually fail

The Codex/Factory gates ran 'git diff --exit-code -- .agents/' / '-- .factory/',
but both paths are gitignored (.gitignore:16-17) — git diff on ignored untracked
paths is always empty, so those two gates were structurally incapable of failing
and 7 of 10 hosts had no gate at all.

New shape: one 'gen:skill-docs --host all' pass (the generator hard-fails on any
per-host error, gating all 10 hosts on generates-cleanly), byte-freshness via
git diff for tracked output, plus a porcelain check that fails on untracked
generated strays (git diff can't see brand-new files). The gitignored-hosts
byte-freshness limitation is documented in the workflow comment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): exorcise the sidebar-agent ghost from the test suite

browse/src/sidebar-agent.ts was deleted in the v1.14 sidebar refactor, but the
test suite kept testing it for 48 versions. Nothing noticed because the free
suite runs in no CI job and Bun-era module-load errors were suppressed in the
Windows shard runner via an exclusion pattern whose own comment documented the
breakage ('broken on every platform since v1.14 ... exit 0').

- Delete sidebar-security.test.ts + security-source-contracts.test.ts: crashed
  at module load (unguarded readFileSync of the deleted file); per-assertion
  triage confirmed every SERVER_SRC pin targeted the deleted chat prompt
  builder (zero hits in today's server.ts) — nothing to port.
- Delete sidebar-integration.test.ts: 11 of 13 tests exercised deleted
  endpoints (/sidebar-command queue, /sidebar-agent/event, chat buffer); the 2
  passing tests pinned only the blanket auth gate, covered by
  server-auth.test.ts + dual-listener.test.ts.
- Delete test/skill-e2e-sidebar.test.ts: E2E for the deleted queue flow.
- sidebar-ux.test.ts 1,669 -> 830 lines: 20 dead-chat describes + 15 dead
  tests removed (incl. 10 vacuous passes asserting on empty indexOf slices);
  2 stale pins on LIVE features fixed (content.js typed-catch CSSOM fallback,
  arrow-hint window widened). 95 pass / 0 fail.
- sidebar-tabs.test.ts: both failures were stale pins, not regressions —
  forceRestart's deliberate ws.close(4001) and the terminal-agent spawn that
  moved into spawnTerminalAgent() (identity-based kill refactor). 28 pass.
- touchfiles.ts: drop the three sidebar E2E entries from BOTH maps
  (E2E_TOUCHFILES + E2E_TIERS) — they pointed diff-selection at the deleted
  file, so those tests were unreachable by any diff.
- test-free-shards.ts: remove the now-dead sidebar-agent exclusion pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ci): run the free test suite in CI (it ran nowhere)

The full free suite (bun test: browse/test/ + test/ + make-pdf/test/) had no CI
job on any Linux/macOS runner — only Windows curated shards, paid evals, and
doc-freshness gates existed. That's how two module-load-crashing test files
survived 48 versions.

Same cached Dockerfile.ci image and container wiring as evals.yml (deps
restore, build, Chromium verify). Includes a module-load-error guard: older
Bun reported test-file import crashes with exit 0 on macOS/Linux, so the job
also fails on any nonzero 'N errors' count in the summary — future crash-class
regressions can't hide from the exact job built to catch them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): validate touchfile dependency paths exist on disk

New guard in touchfiles.test.ts: every non-glob dep path must exist, and every
glob's anchor directory must exist. This is the axis the 181-key two-map sync
discipline never covered — an entry can point at a long-deleted file and
diff-based selection then silently never triggers those tests (the sidebar
trio sat rotted for 48 versions).

First run immediately caught a fourth rotted entry: 'spec authored quality'
referenced test/fixtures/spec/** (directory does not exist) and selected for a
judge test that exists nowhere in the repo. Removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): remove deleted /sidebar-chat endpoint from tunnel allowlist

TUNNEL_PATHS is the audited tunnel attack surface — its own comment says every
addition widens it. '/sidebar-chat' stayed in the set after the endpoint was
deleted with the chat-queue path, meaning any future route matching that path
would have been silently tunnel-exposed. The set is now exactly the pair
ceremony (/connect) and the scoped command endpoint (/command), and the
dual-listener closed-set pin enforces that.

Also repairs a pre-existing red pin in dual-listener.test.ts: v1.63.0.0 made
the tunnel allowlist args-aware (canDispatchOverTunnel gained a second param)
without updating the test — red on main since then, invisible because the free
suite had no CI job.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): delete chain's shadow dispatcher that skipped every security gate

meta-commands.ts carried a 'CLI mode' fallback that re-implemented command
routing without the server pipeline's gates: no scope check, no domain check,
no tab ownership, no rate limit, no hidden-element stripping, no scoped-token
enveloping — and it called handleReadCommand without a BrowserManager, which
also skipped the JS-origin cookie-exfiltration assertion. It was unreachable
in production (server.ts always passes executeCommand) and one boolean away
from being live.

chain now hard-errors without a server context. handleReadCommand's bm param
is required and assertJsOriginAllowed runs unconditionally. The chain tests
that exercised the deleted fallback now route through a server-shaped
executeCommand adapter (real handlers + trust wrapping + {status,result}
envelope), so their behavioral coverage — sequencing, trust markers, pipe
format, aliases, error reporting — survives on the production-shaped path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extension): delete the dead chat-queue client surface

The sidebar-command handler in background.js POSTed to a server endpoint that
no longer exists (deleted with the chat queue) — ~35 lines of fully-wired dead
code including error handling for the permanent 404, plus its allowlist entry.
No sender in the extension ever emitted the message type.

chatEnabled leaves the /health contract (server hardcoded false, background.js
re-derived it, nothing consumed it — the chat input element it guarded is gone
from sidepanel.html). BROWSE_SIDEBAR_CHAT env flag had zero readers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): delete dead exports the ripped chat path left behind

Three-way split by importer class:

(a) Zero importers, deleted: the whole attack-attempt logging cluster in
security.ts (logAttempt, AttemptRecord, salted hashPayload + device-salt,
attempts.jsonl rotation, telemetry spawn plumbing incl.
buildTelemetrySpawnCommand/resolveBashBinary — the LIVE attempts.jsonl writer
is tunnel-denial-log.ts with its own rotation); the decision-file handshake
(writeDecision/readDecision/clearDecision/excerptForReview — written for
sidebar-agent's poll loop, which no longer exists); sidebar-utils.ts (whole
module — its sanitizeExtensionUrl 'sanitized before embedding in a prompt'
for the deleted prompt builder); 8 dead server.ts imports (sanitizeExtensionUrl,
generateCanary, injectCanary, writeDecision, rotateRoot, serializeRegistry,
restoreRegistry, clearAgentRecord); buildPtyClearCookie + buildSseClearCookie;
WEBDRIVER_MASK_SCRIPT (orphaned by the D7 stealth narrowing — applyStealth
never used it).

(b) Dead-pin tests edited with their exports: the 'still exported' pin in
stealth-layer-c, the string-content describe in stealth-webdriver (its live
applyStealth behavioral coverage untouched), the clear-cookie assertions,
security-review-flow.test.ts deleted whole (all 4 describes exercised the
dead decision mechanism, incl. a 'simulated sidebar-agent poll loop').

(c) KEPT deliberately: leaseCount (live behavioral coverage),
extractPtyCookie + validatePtySessionToken (extractPtyCookie is adopted by
the terminal-agent cookie-parse unification later in this wave),
resetSessionMarker + clearContentFilters (test-support API for the live
content-security layer).

Also fixes two pre-existing red pins found while here, invisible until the
free suite got a CI job: the v1.44 spawnClaude->maybeSpawnPty rename in
terminal-agent.test.ts, and a cross-file test-isolation bug where
content-security.test.ts's clearContentFilters() wiped the auto-registered
url-blocklist filter for every later file in the same bun process
(security-integration.test.ts failed on co-run; afterAll now restores it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): delete the dead ML layers — transcript classifier and DeBERTa ensemble

The L4b Haiku transcript classifier and the opt-in DeBERTa ensemble
(GSTACK_SECURITY_ENSEMBLE=deberta, a documented 721MB download) had ZERO
production callers since the chat-path agent that invoked them was ripped.
The only live ML path is scanPageContent (testsavant) inside the security
sidecar subprocess. Deleted by import graph:

- security-classifier.ts 614 -> 265 lines: HAIKU_MODEL, checkTranscript,
  shouldRunTranscriptCheck, loadDeberta, scanPageContentDeberta, ToolCallInput,
  all DEBERTA_* consts + load state. Header now states the live truth
  (imported only by security-sidecar-entry.ts). downloadFile kept, name
  intact — it is an enumerated egress sink (HF model download).
- security-bunnative.ts + test: a research skeleton self-described as 'NOT a
  production replacement', shipped into src/ with zero importers.
- security-bench-ensemble{,-live}.test.ts + the Haiku response fixture: a
  paid live-model benchmark for a layer that could not fire. The
  security-classifier-tdz test's only case exercised checkTranscript — gone.
- security.ts: layer-model header rewritten to the live architecture;
  StatusDetail.layers -> {testsavant, canary}; getStatus() no longer requires
  the impossible transcript==='ok' for 'protected' (old on-disk session state
  with a transcript key is tolerated on read, never re-emitted).
- security-sidecar-entry.ts needed zero changes: it serializes
  getClassifierStatus() verbatim and no consumer read .transcript (verified
  in sidecar-client + server.ts).
- BROWSER.md security section matches reality (ensemble knob gone, 112MB not
  22MB, sidecar hosting documented). combineVerdict/THRESHOLDS retained as
  the pure, tested combiner of record — comments now flag transcript/deberta
  votes as producer-less.

Net: 26 pass in security.test.ts incl. a NEW regression test for stale-
transcript disk tolerance; egress-receipt tripwire green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: scrub the sidebar-agent ghost from comments and CLAUDE.md

20+ comments across 10 files still described the deleted sidebar-agent.ts as a
live process — including load-bearing architecture claims ('IMPORTED ONLY BY
sidebar-agent.ts', 'sidebar-agent fills this in on first prompt-injection
load', 'kill sidebar-agent' in shutdown docs) and ~60 lines of tombstone
blocks in server.ts enumerating deleted identifiers by name (a false grep
surface: searching processAgentEvent hit server.ts and looked live).

CLAUDE.md's security-stack section now documents the LIVE architecture: L1-L3
content filters + testsavant via the security sidecar subprocess; the
L4b/ensemble rows, the GSTACK_SECURITY_ENSEMBLE knob, and the 721MB DeBERTa
download are gone (deleted as dead code this wave) with an explicit
do-not-re-document note; attempts.jsonl is correctly attributed to
tunnel-denial-log.ts; the no-live-writer status of classifierStatus is stated.

Comments that survive now describe what IS, not what WAS: the promotion gate
in domain-skills.ts explains why classifier_score>0 is load-bearing given no
L4 load-time scan exists; file-permissions.ts names real sensitive files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): delete the codex-helpers shadow module

gen-skill-docs.ts imported externalSkillName (unaliased) from
resolvers/codex-helpers.ts at line 21 and then re-declared the same function
locally — the import was silently shadowed, and the imported copy was the
STALE one (it lacked the frontmatterName param the local copy grew). Three
more functions were byte-identical duplicates, imported only under _-prefixed
aliases to keep the module 'referenced', and transformFrontmatter was a
superseded hardcoded-Codex variant. Nothing else imported the module.

Also drops three dead top-of-file imports (COMMAND_DESCRIPTIONS,
SNAPSHOT_FLAGS — which pulled the whole browse/src module graph into every
generator run for nothing — and an unused review-resolver trio).

Proof: bun run gen:skill-docs exits 0 with a byte-identical tree (zero-diff
regen); gen-skill-docs.test.ts 405/405 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): delete ServerConfig.idleTimeoutMs + chromiumProfile — documented, never read

Both fields carried JSDoc asserting embedder behavior that did not exist:
the idle check reads the module-level IDLE_TIMEOUT_MS env constant, and both
resolveChromiumProfile() call sites pass no argument. Worse than absent — an
embedder passing idleTimeoutMs: 5000 silently got 30 minutes.

Wiring them honestly is impossible today: the idle timer, activity state, and
shutdown target are module-global, so a per-factory value would lie for any
process running more than one handler. Deleted instead, with a ServerConfig
note pointing at the deferred singleton/route-table refactor where real
support belongs. BROWSE_IDLE_TIMEOUT and CHROMIUM_PROFILE env remain the
honest knobs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): wire appendSecureFile at the four real log-append sites

file-permissions.ts carries a 24-line rationale for why POSIX mode bits are
insufficient on Windows and implements appendSecureFile (0600 at create,
Windows ACL on first write only) — but its single caller was the dead
logAttempt, while the four REAL page-content log writers (console/network/
dialog logs in server.ts, the command audit log) used raw fs.appendFileSync
with no mode. Page-content-derived logs now get owner-only permissions from
birth on every platform.

Verified before wiring: mode applies atomically at create via appendFileSync
{mode}, and the ACL pass runs only on first write — no per-append subprocess
cost on the hot console-log path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(stealth): handoff() uses the shared profile resolution + lock cleanup

The headless-to-headed handoff path hardcoded ~/.gstack/chromium-profile,
silently ignoring $CHROMIUM_PROFILE and $GSTACK_HOME (gbrowser's gbd sets
per-workspace profiles), and skipped cleanSingletonLocks() — so a handoff
into a profile with a stale SingletonLock could hang where launchHeaded()
would have recovered.

This was the third live drift between the three Chromium launch paths; the
first two are documented in comments as shipped stealth regressions. Minimal
targeted fix — the full buildLaunchConfig() extraction stays in the deferred
queue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): resolver registry describes the template language again

Seven registered {{PLACEHOLDER}}s had zero uses in any .tmpl (checked in both
bare and :arg forms): REDACT_TAXONOMY_TABLE, TEST_COVERAGE_AUDIT_REVIEW,
MODEL_OVERLAY, QUESTION_PREFERENCE_CHECK, QUESTION_LOG, INLINE_TUNE_FEEDBACK,
MAKE_PDF_SETUP. The last two of those families are invoked programmatically by
preamble.ts (functions kept, registry entries dropped); the question-tuning
trio and the review coverage-audit wrapper were documented by their own module
as existing 'for unit testing' that no test performed — deleted, along with
generateRedactTaxonomyTable + its EXAMPLE/TIER_BLURB constants (its '/cso
renders the full table' comment was itself stale) and its test describe.

Also deletes the gated-resolver mechanism (ResolverEntry/appliesTo/
unwrapResolver + test/resolver-entry.test.ts): fully built, fully tested,
used by zero of the 65 registry entries — the generator loop simplifies to a
direct function call. CLAUDE.md's redact-doc line stops advertising the dead
token.

Proof: zero-diff regen (0 SKILL.md changed); gen-skill-docs + skill-validation
737 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): wire boundaryInstruction from host config; drop three no-op binDir ternaries

hosts/codex.ts declared boundaryInstruction and nothing read it — review.ts
kept its own byte-identical CODEX_BOUNDARY literal (verified equal + trailing
escaped newlines). The resolver now reads the config, so the boundary has one
owner. (autoplan's template carries deliberately generic variants, enforced by
gen-skill-docs.test.ts:1358 — untouched by design.)

The 'ctx.host === codex ? $GSTACK_BIN : ctx.paths.binDir' ternary appeared in
three resolvers and could never change the result: resolvers/types.ts already
sets binDir to $GSTACK_BIN for every usesEnvVars host including codex.

Proof: zero-diff regen for claude AND codex hosts; gen-skill-docs +
host-config suites green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test-infra): judge uses resolveClaudeBinary; eval:watch reads the real partials dir

judgePtyState spawned the bare string 'claude' three definitions below the
resolveClaudeBinary() helper this same file exports — broken under hermetic
PATHs where every other launch in the file resolves correctly.

eval:watch read _partial-e2e.json from the legacy global ~/.gstack-dev/evals/
while EvalCollector writes it into the per-project eval dir (or
GSTACK_EVAL_DIR) — so the dashboard's completed-tests panel was empty
whenever slug detection succeeded, i.e. the normal case. The heartbeat and
per-run progress logs stay global by design (session-runner.ts: 'heartbeat
stays global'). The three eval-CLI docstrings stop claiming the legacy dir
is the primary location.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): delete the superseded SDK ship-idempotency suite and three orphaned fixtures

test/skill-e2e-ship-idempotency.test.ts's own header documented that the
monolith's SDK-harness version tests a synthetic prompt while it exercises
the real /ship skill — the author knew the old suite was superseded and left
both running, two paid LLM runs for one behavior. The weaker copy is gone;
its 'ship-idempotency' diff-selection key goes with it (the dedicated file is
periodic-tier, which always runs under EVALS_ALL — the key had no remaining
consumer).

Fixture rot: test/fixtures/golden-ship-claude.md was a 128KB zero-reader
orphan that had drifted 46KB from its live successor
(test/fixtures/golden/claude-ship-SKILL.md) while looking authoritative;
parity-baseline-v1.46.0.0.json and v1.53.0.0.json had zero readers (three
tests pin three OTHER baseline versions — consolidation is queued, deletion
of the unreferenced two is free).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bin): delete zero-caller scripts; make host-config-export's docstring honest

- bin/gstack-open-url (14 lines): announced in a CHANGELOG entry, wired into
  nothing, ever. bin/gstack-platform-detect (27 lines): zero callers, and its
  hand-rolled host list was already stale (SLATE_HOST.md cites it as a
  problem). Note: the deprecated gstack-brain-consumer/reader pair the audit
  flagged was already deleted upstream in v1.63 with a stay-deleted tripwire.
- scripts/task-emission-schema.ts (61 lines): a typed schema module nothing
  imported; the tasks-section comment now documents the JSONL fields inline.
- scripts/host-config-export.ts claimed to be the 'shell bridge for the bash
  setup script' — setup never calls it (its hand-rolled host lists drifting
  is a known follow-up). Docstring now states what it IS: a standalone,
  test-pinned query CLI not yet wired into setup. Its validateValue +
  CLI_REGEX/PATH_REGEX internals were dead (defined for a guarantee the
  header claimed but nothing enforced).
- KEPT deliberately: scripts/preflight-agent-sdk.ts — a documented manual
  diagnostic (CONTRIBUTING.md + USING_GBRAIN_WITH_GSTACK.md reference it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): one lone-surrogate sanitizer, one sanitizeReplacer, one startTunnel

Three copies of the surrogate sanitizer existed with two algorithms
(sanitize.ts regex vs a hand-rolled charCodeAt walk in server.ts — verified
byte-identical across 11 edge cases before converging) plus two identical
sanitizeReplacer definitions each wrapping a different copy. sanitize.ts is
now the single source of truth; the runs-INSIDE-JSON.stringify egress
invariant is unchanged at every call site and its pin tests were adapted to
the new import shape without losing intent.

The ngrok tunnel-start sequence existed three times in server.ts — the
/tunnel/start route and the BROWSE_TUNNEL=1 autostart were line-for-line
equivalent (a comment admitted 'Same cleanup as /tunnel/start's error path').
One startTunnel() now owns the ephemeral loopback bind, the pre-send egress
receipt, the state-file RMW via tmpStatePath(), and the ordered error-path
cleanup; callers keep their distinct response surfaces. The
BROWSE_TUNNEL_LOCAL_ONLY test path shares nothing (no ngrok, different state
field) and deliberately stays separate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): one session-cookie registry implementation, two instances

pty-session-cookie.ts and sse-session-cookie.ts were byte-identical modulo
the cookie name — mint/validate/parse/prune/TTL, the exact code a security
fix would have to land in twice (and a third hand-rolled cookie parse in
terminal-agent.ts had already diverged; unified next commit).
createSessionCookieStore() owns the implementation; both modules become thin
instantiations keeping every exported name, their distinct threat-model
docstrings, and separate token spaces (an SSE-read cookie must never grant
PTY access). pty-session-lease.ts deliberately stays out — different contract
(sessionId/secret split, refresh, env TTL).

The factory imports nothing from token-registry (cookie-picker-auth-isolation
invariant, still pinned by sse-session-cookie.test.ts).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): terminal-agent uses the shared PTY cookie parser

The /ws upgrade's cookie fallback hand-parsed the Cookie header inline — the
fourth copy of the session-cookie parse, and the one that had already
diverged from the others. Parsing now goes through extractPtyCookie;
validation deliberately stays against the agent's own in-process validTokens
map (the server's registry lives in a different process). The ws-handler pin
test now pins the shared-parser call instead of the raw cookie-name literal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(hosts): defineHost() factory — 10 copy-paste host files become declarations

hosts/*.ts were ten copies of one file: runtimeRoot byte-identical in 9/10,
pathRewrites mechanically derivable from the host name for 7/10, the 11-entry
toolRewrites map byte-identical between openclaw and gbrain, and every asset
change a 10-file edit (cursor and slate had already fallen out of three other
hand-maintained lists). defineHost() owns the defaults; each host file now
declares only what makes it different (slate/cursor: 8 lines each). Shared
constants: CROSS_MODEL_RESOLVERS, GBRAIN_RESOLVERS, EXEC_STYLE_TOOL_REWRITES.
Genuinely-different things stayed explicit: codex/factory $GSTACK_ROOT
rewrites, hermes's tool vocabulary, claude's denylist+prefixable install,
opencode's wider runtimeRoot.

Proof: JSON.stringify(ALL_HOST_CONFIGS) dump-diff before/after EMPTY (and a
runtime walk confirmed no function-valued or undefined-keyed fields, so the
JSON diff is complete); gen:skill-docs --host all zero-diff; host-config +
gen-skill-docs + idempotency suites 485/485. Host files 595 -> 285 lines.
docs/ADDING_A_HOST.md teaches the factory pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(lib): fs-atomic — one atomic-write implementation, with the race actually fixed

Atomic tmp-write-then-rename was reimplemented ~20 times across lib/, bin/,
and browse/src with three tmp-suffix conventions. One of them was a latent
bug this commit closes: lib/worktree.ts used a bare '.tmp' suffix — the
deterministic-tmp collision race browse/src/server.ts documents having hit
in production (its fix, pid+random, was trapped in a comment at one site).

lib/fs-atomic.ts: atomicWriteSync (always throws, best-effort tmp cleanup,
pid+random suffix, optional mode applied at tmp creation so the file never
exists with looser permissions) + atomicWriteQuiet (shutdown paths only).
Unit tests pin the throw/quiet contracts, 0600 mode, tmp-name uniqueness
(captured via the read-only-dir failure path — Bun's fs exports are
readonly, no monkeypatching), and no-stray-tmp cleanup.

Migrated: lib/worktree.ts (the bare-.tmp bug), lib/gstack-decision.ts
(snapshot + compact log), lib/gbrain-local-status.ts (probe cache). browse
sites follow separately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): jsonl-store's docstring stops lying; mode option added; lib bypasses adopted

The header claimed 'single source of truth... the ONLY copy' with write-time
injection REJECTION — while appendJsonl never screened anything, only 1 of
~10 JSONL stores imported it, and a bypass appender lived in the same
directory. Now: the contract is explicit (screening is the CALLER's job via
hasInjection/firstInjectionMatch; the enforcing callers are named), a
option applies 0600 at create for sensitive stores, and the lib bypasses are
adopted (gstack-memory-helpers ×2, redact-audit-log — which keeps its chmod
backstop for files created looser by pre-mode versions). browse/src keeps
its own appenders by design (compiled-binary surface, own secure-append
helper) and the header now says so. gstack-decision's batched archive append
stays deliberate (single-write crash-window semantics appendJsonl's
one-record contract can't express).

New pins: 0600-at-create, and a test that documents appendJsonl does NOT
self-screen — so nobody can re-document it as self-screening without making
it true.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): migrate hand-rolled atomic writes to lib/fs-atomic

Seven sites, each audited for its existing throw-vs-swallow contract before
migrating: writeSessionState + the four fire-and-forget tab/state writers use
atomicWriteQuiet (they swallowed before); writeAgentRecord + the boot-time
port-file write use atomicWriteSync (they threw before — and writeAgentRecord
previously leaked its tmp file on rename failure, which the helper cleans).
All carry {mode: 0o600} plus restrictFilePermissions after successful writes,
preserving the Windows ACL hardening that writeSecureFile provided (mode bits
are POSIX-only). server.ts untouched: its three state writes route through
tmpStatePath(), pinned by server-tmp-state-path.test.ts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(hosts): delete five dead HostConfig fields

metadataFormat (generator hardcodes openai.yaml), sidecar (behavior lives in
setup's create_agents_sidecar — knowledge preserved as a comment in codex.ts),
install.prefixable (skill_prefix is implemented entirely in bin/gstack-config),
staticFiles (docstring cited a SOUL.md that never existed anywhere), and
adapter (its only would-be consumer, openclaw-adapter.ts, was fully dead —
with a test asserting the field was undefined). Kept: learningsMode (wired
next), linkingStrategy (validation reads it), coAuthorTrailer (consumed by
resolvers/utility.ts).

Proof: JSON dump diff shows ONLY the deleted keys vanishing; zero-diff regen
across all 10 hosts; host-config + gen-skill-docs suites green. Note: this
commit also carries chunk-23 edits to the shared hosts/claude.ts +
define-host.ts + host-config.test.ts files (skipSkills collapse, stale
line-number comment drops) — pathspec commits, concurrent prep.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): preamble tiers are explicit; silent ?? 4 default becomes an error; spec stops rendering its preamble twice

Eight skills (scrape, diagram, spec, skillify, pair-agent, landing-report,
open-gstack-browser + its connect-chrome symlink) silently received the
HEAVIEST tier-4 preamble because a missing frontmatter field defaulted to 4.
Tiers are now declared in every {{PREAMBLE}} template's frontmatter and a
missing declaration throws at generation time with the template path (the 5
templates without {{PREAMBLE}} never invoke the resolver). The stale
hand-written tier-map comment (wrong in 3 of 4 rows) is gone.

Bonus bug fixed: spec/SKILL.md.tmpl mentioned {{PREAMBLE}} in prose, so the
generator inlined the ENTIRE preamble a second time — spec/SKILL.md shrinks
127,462 -> 80,924 bytes (-46,538) from de-duplication alone. skill-size-budget
gains a reasoned INTENTIONAL_SHRINKS entry (its frozen baseline had measured
the doubled-preamble bug). New tests: missing-tier throw carries the path;
every {{PREAMBLE}} template declares a tier. (Carries chunk-23 edits in the
shared test/gen-skill-docs.test.ts.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): learningsMode is read from host config, not a hardcoded host name

resolvers/learnings.ts branched on ctx.host === 'codex' while every host
declared learningsMode — the field was decorative, and the 7 hosts configured
'basic' (cursor, slate, kiro, opencode, openclaw, hermes, gbrain) silently
received the 'full' cross-project flow their runtimes can't execute (it
depends on AskUserQuestion + gstack-config plumbing). Output now matches
declaration: basic hosts get the project-scoped search block.

Blast radius proof: all committed Claude SKILL.md files and the three golden
fixtures are byte-identical; the behavior diff lands only in the gitignored
external-host trees (hand-verified: .cursor review's learnings section swaps
the cross-project AskUserQuestion block for the project-scoped search).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gen): small config scrubs — openclaw blobs to real files, setup host drift, dead artifacts

- The three openclaw markdown blobs hardcoded inside gen-skill-docs.ts (which
  silently reverted any hand edit to their tracked outputs on regen) move to
  openclaw/templates/*.md source files; output shasums byte-identical.
- setup's --host allowlists gain cursor + slate — both fully registered hosts
  with generated output, but './setup --host cursor' exited 1 because two
  hand-rolled lists in setup had drifted from hosts/index.ts.
- scripts/proactive-suggestions.json deleted: 31KB regenerated on every run,
  read by nobody (the catalog-trim design's reader was never built); its
  emitter and three determinism tests (which guaranteed a file nothing reads
  didn't churn) retired with stays-retired pins.
- claude/SKILL.md.tmpl deleted: a complete 8.9KB skill that never generated
  output (directory name collides with the host id 'claude'), in no registry.
  Recoverable from git if ever wanted under a non-colliding name.
- openclaw's frozen extraFields.version '0.15.2.0' stamp dropped;
  includeSkills: [] no-ops omitted (the generator treats [] as absent);
  llms.txt 55 -> 54 skills.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): correct preamble tiers for the 8 silently-heaviest skills

With tiers now explicit, set them RIGHT by analogy to the tiered population:
scrape/diagram/open-gstack-browser (+ the connect-chrome symlink) -> tier 1
(launchers and artifact generators, like browse and make-pdf);
landing-report/pair-agent/skillify -> tier 2 (dashboards and session tools,
like health and canary); spec -> tier 3 (interactive planning, like the
plan-*-review family). Each tier-1 skill sheds 271 lines of onboarding
prose it never needed; tier-2 shed 20 each.

Verification per the review protocol: regen diff reviewed (pure
section-removal), skill-validation + size-budget + catalog-budget +
v0-dormancy suites green (822 tests), and live smoke of the tier-corrected
skills confirms the preamble renders the intended sections at each tier.
These skills have ~no eval coverage — stated honestly; the wave's gate-tier
eval run is the backstop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): e2e-gate — one tier-gate implementation, side-effect-free, with the trap pinned

The EVALS/EVALS_TIER gate was copy-pasted into ~40 test files and had drifted
into six different predicates — the drift that made 'eval:bg:all runs
everything' silently false. test/helpers/e2e-gate.ts owns the semantics now:
describeE2ETier(tier) + e2eTierEnabled(tier), env read at call time, zero
side effects (the existing e2e-helpers module runs a ~30s claude ping at
import under EVALS=1, so the gate lives in its own module; purity is pinned
by tests that scan imports and comment-stripped source).

The unit matrix pins all four env combos — including EVALS=1 with EVALS_TIER
unset -> SKIP, the exact trap that made eval:bg:all a non-run. The
tier-alignment tripwire gains a second regex for the helper shape (old shape
still detected — stragglers can't hide), and the sharded paid runner's
PRE-SPAWN tier classifier learns the helper shape too: without that, every
gate-sharded run would have spawned all 28 periodic shards just to skip them,
each paying the e2e-helpers import ping (~15 min of dead wall clock in the
CI-blocking lane). Verified: gate runs exclude the 29 periodic files,
periodic excludes the 8 gate files — identical to pre-migration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(test): migrate the 36 tier-gated eval files to describeE2ETier

Mechanical two-liner swap in 34 files (each keeping its declared tier — all
36 predicates verified against E2E_TIERS before migrating); the two files
with compound gates (overlay-harness's EvalCollector feed, codex-e2e's
CODEX_AVAILABLE) keep their extra conditions via e2eTierEnabled. Tier
rationale comments preserved. codex-e2e/gemini-e2e/benchmark-providers keep
their distinct stderr-message gate shapes by design.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(test): skill-e2e + skill-llm-eval adopt the shared selection machinery

Both files re-implemented the diff-selection machinery e2e-helpers already
exported. The helper gained computeDiffSelection() (extracted, identical
behavior) and a trailing optional selection param on the *IfSelected helpers
(defaults preserve all 30+ existing importers). skill-e2e.test.ts drops ~120
duplicated lines; skill-llm-eval keeps its LLM_JUDGE_TOUCHFILES selection and
test.concurrent semantics via testConcurrentIfSelected.

Deliberate deltas, stated: skill-e2e.test.ts now honors the EVALS_TIER
intersection its local copy lacked (affects only direct bun test invocations
of that file — it matches no eval-script glob); its recordE2E gains the
helper's three diagnostic fields; skill-llm-eval sharded solo now runs
e2e-helpers' module-scope preflight it already ran in combined processes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): kill the silent-truncation race; exempt the tier-corrected shrinks

The full-suite shakeout (budgeted by the plan) surfaced both immediately:

1. server-embedder-terminal-port.test.ts stubbed process.exit and restored
   the REAL exit in its finally — but shutdown() schedules async work that
   can call process.exit AFTER restoration, killing the entire bun process
   mid-suite with exit 0 and NO summary. This is the silent-truncation class
   the new free-suite CI job guards against, reproduced locally on the first
   full run. Exit now stays a logging no-op between tests (late async exits
   become visible stderr lines, not process death); the true exit returns in
   afterAll.

2. The 80%-of-baseline shrink guard correctly flagged the six tier-corrected
   skills — their baseline was measured at the silent tier-4 default. Added
   to INTENTIONAL_SHRINKS with the reason, joining spec's double-preamble
   entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* release: v1.64.0.0 — the code-smell fix wave

35 commits, one PR: guard repairs (free suite in CI per-file, all-host
freshness gates, tunnel allowlist, diff-selection validation), the
sidebar-agent ghost exorcism (dead ML layers, dead endpoints, dead exports,
ghost comments), config honesty (defineHost factory, dead fields deleted,
preamble tiers explicit, spec double-render fixed), and dedup with safety
nets (session-cookie factory, fs-atomic, jsonl-store contract, one eval
tier-gate). Net -24,943 lines across 183 files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): free-tests step runs under bash (container sh rejects pipefail)

Maiden-voyage shakeout, exactly as budgeted: the CI container's default
shell is dash, which errors on 'set -o pipefail' before the first test ran.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): free-tests curates 8 container-incompatible files with reasons

Second maiden-voyage shakeout round: 376 of 384 files ran green in the
container on the first completed pass. The 8 that can't run there yet are
excluded the same way the Windows shards curate POSIX-bound files — each
with its reason inline (headed-Chrome handoff, real-PTY round-trip, X server
management, extension-origin identity, the job's own TMPDIR override, and
three pre-existing env failures that fail on dev machines too). Anything
outside the list that fails still fails the job; trimming the list is
tracked follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): gstack-config-key-locale — suppress the skill_prefix auto-relink side effect

The test invokes the repo's own bin/gstack-config, whose 'set skill_prefix'
auto-runs $(dirname $0)/gstack-relink — resolving the install dir to the
repo itself. In any environment where the loop shares a working tree (the
free-tests CI container, a fresh-HOME run), gstack-patch-names rewrote all
52 tracked SKILL.md names to gstack- prefixed, poisoning five unrelated
suites downstream (hermetic-skills-seeding, host-config golden, skill-census,
skill-validation, spec-template-sync). GSTACK_SETUP_RUNNING=1 is the
documented suppression; relink behavior stays covered by relink.test.ts's
mock install.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bin): gstack-codex-session-import — empty sessions dir exits 0 on Linux

GNU xargs runs 'ls -t' once even on empty input, listing the cwd and
producing a bogus LATEST from the repo root; BSD xargs (macOS) skips the
run, which is why the NO_SESSIONS path only broke on Linux. xargs -r pins
the BSD behavior on both platforms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(parity): rebaseline v1.57.7.0 → v1.64.1.0 + skeleton-cap headroom

The two parallel v1.64 waves (code-smell fix wave + main's #2571) each
added shared-preamble prose, pushing document-release / design-consultation
/ cso past their size ratios on the v1.57.7.0 anchor and four carved
skeletons (plan-ceo-review, plan-eng-review, office-hours,
design-consultation) 22-280 B over their absolute caps. New baseline is
union-normalized (skeleton + sections/*.md, matching what the harness
measures); caps get +~1 KB headroom each with per-cap rationale. The
v1.57.7.0 fixture stays in test/fixtures/ for the audit trail, and
capture-parity-baseline.ts now documents the union-normalization step so
the next rebaseline doesn't re-trip on it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): free-tests container parity — tools, pinned bun, git identity, mutation tripwire

- Dockerfile.ci: add python3 (gstack-jsonl-merge/brain-sync/detach shell out
  to it), file (skill-validation's binary check), poppler-utils (make-pdf
  e2e gates hard-require pdftotext/pdffonts/pdfinfo), fonts-noto-color-emoji
  (emoji render gate, mirrors make-pdf-gate.yml). Fix the bun pin: the
  bun.sh installer ignores a BUN_VERSION env var, so the old form silently
  installed latest on every rebuild (observed 1.3.13/1.3.14 drift vs the
  1.3.10 devs run locally); pass the version as the positional arg.
- free-tests.yml: git identity + safe.directory for the git-exercising
  tests (container checkout is owned by a different uid than runner);
  post-loop tree-mutation tripwire that names a tracked-file-mutating test
  instead of letting downstream collateral confuse the report; skip the
  documented variants-retry-after timing flake.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bin): gstack-session-update — detached updater owns its stdio (SIGPIPE)

The backgrounded update subshell inherited the session hook's stdout/stderr
pipes. Once the hook exits and the caller closes them, any child that writes
— git pull's autostash notice, setup output — dies of SIGPIPE, logged as
PULL_FAILED exit=141 with an empty stderr capture (observed in the free-tests
container, and reachable by any production hook runner that closes stdio
promptly). Redirect the fork to /dev/null; all observability already flows
through the session-update log file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): gstack-decision-bins — explicit branch context for the scope filter

CI checks out a detached HEAD, where gitBranch() returns undefined on both
the log and search sides, so an implicitly branch-scoped decision can never
surface (filterByScope requires a matching non-empty ctx.branch). Pass the
branch explicitly on both sides — the filter logic is what's under test, not
git branch detection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): ring-buffer lease interplay — same TTL window, not same millisecond

Two back-to-back mintLease() calls each stamp Date.now() + TTL; when they
straddle a millisecond boundary the exact-equality assertion flakes
(observed in CI: expiries of ...525 vs ...526). Assert the expiries are
within a 50 ms window instead — the invariant under test is that leases
share a TTL policy, not that they mint in the same clock tick.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 09:37:04 -07:00

64 KiB
Raw Blame History

Browser — Complete Reference

gstack's browser surface in one document. Headless Chromium daemon, ~70+ commands, ref-based element selection, codifiable browser-skills, real-browser mode with a Chrome side panel, an in-sidebar Claude PTY, an ngrok pair-agent flow, and a layered prompt-injection defense — all behind a compiled CLI that prints plain text to stdout. ~100-200ms per call. Zero context-token overhead.

If you've used gstack in the last release or two, the productivity loop is the new headline: /scrape <intent> drives a page once, /skillify codifies the flow into a deterministic Playwright script, and the next /scrape on the same intent runs in ~200ms instead of ~30 seconds of agent re-exploration.


Quick start

# One-time: build the binary (browse/dist/browse, ~58MB)
bun install && bun run build

# Set $B once and forget about it
B=./browse/dist/browse           # or ~/.claude/skills/gstack/browse/dist/browse

# Drive a page
$B goto https://news.ycombinator.com
$B snapshot -i                   # @e refs you can click/fill/inspect later
$B click @e30                    # click ref 30 from the snapshot
$B text                          # get clean page text
$B screenshot /tmp/hn.png

# Codify a repeated flow
/scrape latest hacker news stories
/skillify                        # writes ~/.gstack/browser-skills/hn-front/...
/scrape hacker news front page   # second call: 200ms via the codified skill

# Watch Claude work in real time
$B connect                       # headed Chromium + Side Panel extension

Table of contents

  1. What it is
  2. The productivity loop — /scrape + /skillify
  3. Architecture
  4. Command reference
  5. Snapshot system + ref-based selection
  6. Browser-skills runtime
  7. Domain-skills (per-site agent notes)
  8. Real-browser mode ($B connect) — including --headed + --proxy + --navigate (v1.28.0.0)
  9. Side Panel + sidebar agent
  10. Pair-agent — remote agents over an ngrok tunnel
  11. Authentication + tokens
  12. Prompt-injection security stack (L1L6)
  13. Screenshots, PDFs, visual inspection
  14. Local HTML — goto file:// vs load-html
  15. Batch endpoint
  16. Console, network, dialog capture
  17. JS execution — js + eval
  18. Tabs, frames, state, watch, inbox
  19. CDP escape hatch + CSS inspector
  20. Performance + scale
  21. Multi-workspace isolation
  22. Environment variables
  23. Source map
  24. Development + testing
  25. Cross-references
  26. Acknowledgments

What it is

A compiled CLI binary that talks to a persistent local Chromium daemon over HTTP. The CLI is a thin client — it reads a state file, sends a command, prints the response to stdout. The daemon does the real work via Playwright.

Everything that was a Chrome MCP server in the early days now happens through plain stdout. No JSON-schema framing, no protocol negotiation, no persistent WebSocket — Claude's Bash tool already exists, so we use it.

Three escalating modes:

  • Headless (default). Daemon runs Chromium with no visible window. Fastest, cheapest, what skills like /qa, /design-review, /benchmark use by default.
  • Headed via $B connect. Same daemon, but Chromium is visible (rebranded as "GStack Browser") with the Side Panel extension auto-loaded. You watch every command tick through in real time.
  • Pair-agent over a tunnel. Daemon binds a second listener that ngrok forwards. A remote agent (Codex, OpenClaw, Hermes, anything that can speak HTTP) drives your local browser through a 26-command allowlist with a scoped, single-use token.

The productivity loop

The shipped headline of v1.19.0.0. Two gstack skills wrap the browser-skills runtime so the second time you ask Claude to scrape a page, it runs in ~200ms.

/scrape <intent>

One entry point for pulling page data. Three paths under the hood:

  1. Match path (~200ms) — agent runs $B skill list, semantically matches the intent against each skill's triggers: array + description + host, and runs $B skill run <name> if a confident match exists.
  2. Prototype path (~30s) — no match, agent drives the page with $B goto, $B text, $B html, $B links, etc., returns the JSON, and appends a one-line "say /skillify" suggestion.
  3. Mutating-intent refusal — verbs like submit, click, fill route to /automate (Phase 2b, P0 in TODOS.md). /scrape is read-only by contract.

/skillify

Codifies the most recent successful /scrape prototype into a permanent browser-skill on disk. Eleven steps, three locked contracts:

  • D1 — Provenance guard. Walks back ≤10 agent turns for a clearly-bounded /scrape result. Refuses with one specific message if cold. No silent synthesis from chat fragments.
  • D2 — Synthesis input slice. Extracts ONLY the final-attempt $B calls that produced the JSON the user accepted, plus the user's intent string. Drops failed selectors, drops chat, drops earlier-session content.
  • D3 — Atomic write. Stages everything to ~/.gstack/.tmp/skillify-<spawnId>/, runs $B skill test against the temp dir, and only renames into the final tier path on test pass + user approval. Test fail or rejection: rm -rf the temp dir entirely. No half-written skill ever appears in $B skill list.

Mutating-flow sibling /automate is split out as P0 in TODOS.md and ships on the next branch — same skillify machinery, per-mutating-step confirmation gate when running non-codified.

See docs/designs/BROWSER_SKILLS_V1.md for the full design + decision trail.


Architecture

┌─────────────────────────────────────────────────────────────────┐
│  Claude Code                                                    │
│                                                                 │
│  $B goto https://staging.myapp.com                              │
│       │                                                         │
│       ▼                                                         │
│  ┌──────────┐    HTTP POST     ┌──────────────┐                 │
│  │ browse   │ ──────────────── │ Bun HTTP     │                 │
│  │ CLI      │  127.0.0.1:rand  │ daemon       │                 │
│  │          │  Bearer token    │              │                 │
│  │ compiled │ ◄──────────────  │  Playwright  │──── Chromium    │
│  │ binary   │  plain text      │  API calls   │    (headless    │
│  └──────────┘                  └──────────────┘     or headed)  │
│   ~1ms startup                  persistent daemon               │
│                                 auto-starts on first call       │
│                                 auto-stops after 30 min idle    │
└─────────────────────────────────────────────────────────────────┘

Daemon lifecycle

  1. First call. CLI checks <project>/.gstack/browse.json for a running server. None found — it spawns bun run browse/src/server.ts in the background. Daemon launches headless Chromium via Playwright, picks a random port (1000060000), generates a bearer token, writes the state file (chmod 600), starts accepting requests. ~3 seconds.
  2. Subsequent calls. CLI reads the state file, sends an HTTP POST with the bearer token, prints the response. ~100-200ms round trip.
  3. Idle shutdown. After 30 minutes of no commands, daemon shuts down and cleans up the state file. Next call restarts it.
  4. Crash recovery. If Chromium crashes, the daemon exits immediately — no self-healing, don't hide failure. CLI detects the dead daemon on the next call and starts a fresh one.

Multi-workspace isolation

Each project root (detected via git rev-parse --show-toplevel) gets its own daemon, port, state file, cookies, and logs. No cross-workspace collisions. State at <project>/.gstack/browse.json.

Workspace State file Port
/code/project-a /code/project-a/.gstack/browse.json random (1000060000)
/code/project-b /code/project-b/.gstack/browse.json random (1000060000)

Command reference

~70 commands across read, write, and meta. Selectors accept CSS, @e refs from snapshot, or @c refs from snapshot -C. Full table:

Reading

Command Description
text [sel] Clean page text (or scoped to a selector)
html [sel] innerHTML, or full page HTML if no selector
links All links as text → href
forms Form fields as JSON
accessibility Full ARIA tree
media [--images|--videos|--audio] [sel] Media elements with URLs, dimensions, types
data [--jsonld|--og|--meta|--twitter] Structured data: JSON-LD, OG, Twitter Cards, meta tags

Inspection

Command Description
js <expr> [--out <file>] [--raw] Run inline JavaScript expression in page context, return as string. With --out <file> the result is written to disk instead of returned (a data:*;base64,... result is decoded to raw bytes unless --raw). --out makes the invocation a WRITE (needs write scope, never allowed over the tunnel).
eval <file> [--out <file>] [--raw] Run JS from a file (path under /tmp or cwd; same sandbox as js). --out/--raw behave as for js.
css <sel> <prop> Computed CSS value
attrs <sel|@ref> Element attributes as JSON
is <prop> <sel|@ref> State check: visible, hidden, enabled, disabled, checked, editable, focused
console [--clear|--errors] Captured console messages
network [--clear] Captured network requests
dialog [--clear] Captured dialog messages
cookies All cookies as JSON
storage / storage set <key> <val> Read both localStorage + sessionStorage; set localStorage
perf Page load timings
inspect [sel] [--all] [--history] Deep CSS via CDP — full rule cascade, box model, computed styles
ux-audit Page structure for behavioral analysis: site ID, nav, headings, text blocks, interactive elements
cdp <Domain.method> [json-params] Raw CDP method dispatch (deny-default; allowlist in cdp-allowlist.ts)

Navigation

Command Description
goto <url> Navigate to URL (http://, https://, file://)
load-html <file> Load local HTML in memory (no file:// URL; survives viewport scale changes)
back, forward, reload Standard nav
url Current page URL
wait <sel|--networkidle|--load> Wait for element, network idle, or page load (15s timeout)

Interaction

Command Description
click <sel|@ref> Click element
fill <sel> <val> Fill input
select <sel> <val> Select dropdown option (value, label, or visible text)
hover <sel> Hover element
type <text> Type into focused element
press <key> Playwright keyboard key (case-sensitive: Enter, Tab, ArrowUp, Shift+Enter, Control+A, ...)
scroll [sel|@ref] Scroll element into view, or jump to page bottom if no selector
viewport [<WxH>] [--scale <n>] Set viewport size + optional deviceScaleFactor 1-3 (retina screenshots)
upload <sel> <file> [...] Upload file(s)
dialog-accept [text] Auto-accept next alert/confirm/prompt; text is sent for prompts
dialog-dismiss Auto-dismiss next dialog

Style + cleanup

Command Description
style <sel> <prop> <val> Modify CSS property (with undo support)
style --undo [N] Undo last N style changes
cleanup [--ads|--cookies|--sticky|--social|--all] Remove page clutter
prettyscreenshot [--scroll-to <sel|text>] [--cleanup] [--hide <sel>...] [path] Clean screenshot with optional cleanup, scroll, hide

Visual

Command Description
screenshot [--selector <css>] [--viewport] [--clip x,y,w,h] [--base64] [sel|@ref] [path] Five modes: full page, viewport, element crop, region clip, base64
pdf [path] [--format letter|a4|legal] [...] PDF with full layout: format, width/height, margins, header/footer templates, page numbers, --tagged for accessibility, --toc waits for Paged.js
responsive [prefix] Three screenshots: mobile (375x812), tablet (768x1024), desktop (1280x720)
diff <url1> <url2> Text diff between two URLs

Cookies + headers

Command Description
cookie <name>=<value> Set cookie on current page domain
cookie-import <json> Import cookies from JSON file
cookie-import-browser [browser] [--domain d] Import from installed Chromium browsers (interactive picker, or --domain for direct import)
header <name>:<value> Set custom request header (sensitive values auto-redacted)
useragent <string> Set user agent (triggers context recreation, invalidates refs)

Tabs + frames

Command Description
tabs List open tabs
tab <id> Switch to tab
newtab [url] [--json] Open new tab; --json returns {tabId, url} for programmatic use
closetab [id] Close tab
tab-each <command> [args...] Fan out a command across every open tab; returns JSON
frame <sel|@ref|--name n|--url pattern|main> Switch to iframe context (or back to main); clears refs

Extraction

Command Description
download <url|@ref> [path] [--base64] Download URL or media element using browser cookies
scrape <images|videos|media> [--selector] [--dir] [--limit] Bulk download all media from page; writes manifest.json
archive [path] Save complete page as MHTML via CDP

Snapshot

Command Description
snapshot [-i] [-c] [-d N] [-s sel] [-D] [-a] [-o path] [-C] Accessibility tree with @e refs; -i interactive only, -c compact, -d N depth, -s scope, -D diff vs previous, -a annotated screenshot, -C cursor-interactive @c refs

Server lifecycle

Command Description
status Daemon health + mode (headless / headed / cdp)
stop Shut down daemon
restart Restart daemon
connect Launch headed GStack Browser with Side Panel extension
disconnect Close headed Chrome, return to headless
focus [@ref] Bring headed Chrome to foreground (macOS); @ref also scrolls into view
state save|load <name> Save or load browser state (cookies + URLs)
memory [--json] Snapshot Bun heap + per-tab JS heap + Chromium process tree + bounded buffer sizes. Use --json for programmatic consumers; text mode renders sorted top-10 tabs with "and N more" tail.

Handoff

Command Description
handoff [reason] Open visible Chrome at current page for user takeover (CAPTCHA, MFA, complex auth)
resume Re-snapshot after user takeover, return control to AI

Meta + chains

Command Description
chain (JSON via stdin) Run a sequence of commands. Pipe [["cmd","arg1",...],...] to $B chain. Stops at first error.
inbox [--clear] List messages from sidebar scout inbox
watch [stop] Passive observation — periodic snapshots while user browses; stop returns summary

Browser-skills runtime

Command Description
skill list List all browser-skills with resolved tier (project > global > bundled)
skill show <name> Print SKILL.md
skill run <name> [--arg k=v...] [--timeout=Ns] Spawn the skill script with a per-spawn scoped token
skill test <name> Run the skill's script.test.ts against bundled fixtures
skill rm <name> [--global] Tombstone a user-tier skill

Domain-skills

Command Description
domain-skill save|list|show|edit|promote-to-global|rollback|rm <host?> Per-site agent notes (host derived from active tab). Lifecycle: quarantined → active (after N=3 successful uses without classifier flag) → global (explicit promote)

Aliases: setcontent, set-content, setContentload-html (canonicalized before scope checks, so a read-scoped token can't use the alias to run a write command).


Snapshot system

The browser's key innovation is ref-based element selection built on Playwright's accessibility tree API. No DOM mutation. No injected scripts. Just Playwright's native AX API.

How @ref works

  1. page.locator(scope).ariaSnapshot() returns a YAML-like accessibility tree.
  2. The snapshot parser assigns refs (@e1, @e2, ...) to each element.
  3. For each ref, it builds a Playwright Locator (using getByRole + nth-child).
  4. The ref→Locator map is stored on BrowserManager.
  5. Later commands like click @e3 look up the Locator and call locator.click().

Ref staleness detection

SPAs can mutate the DOM without navigation (React router, tab switches, modals). When this happens, refs collected from a previous snapshot may point to elements that no longer exist. resolveRef() runs an async count() check before using any ref — if the element count is 0, it throws immediately with a message telling the agent to re-run snapshot. Fails fast (~5ms) instead of waiting for Playwright's 30-second action timeout.

Extended snapshot features

  • --diff (-D). Stores each snapshot as a baseline. On the next -D call, returns a unified diff showing what changed. Use this to verify that an action (click, fill, etc.) actually worked.
  • --annotate (-a). Injects temporary overlay divs at each ref's bounding box, takes a screenshot with ref labels visible, then removes the overlays. Use -o <path> to control the output.
  • --cursor-interactive (-C). Scans for non-ARIA interactive elements (divs with cursor:pointer, onclick, tabindex>=0) using page.evaluate. Assigns @c1, @c2... refs with deterministic nth-child CSS selectors. These are elements the ARIA tree misses but users can still click.

Browser-skills runtime

Per-task directories that codify a repeated browser flow into a deterministic Playwright script. The compounding layer.

Anatomy of a browser-skill

browser-skills/<name>/
├── SKILL.md                        # frontmatter + prose contract
├── script.ts                       # deterministic Playwright-via-browse-client logic
├── _lib/browse-client.ts           # vendored copy of the SDK (~3KB, byte-identical to canonical)
├── fixtures/<host>-<date>.html     # captured page for fixture-replay tests
└── script.test.ts                  # parser tests against the fixture (no daemon required)

The bundled reference is browser-skills/hackernews-frontpage/: scrapes the HN front page, returns 30 stories as JSON. Try it:

$B skill list                            # shows hackernews-frontpage (bundled)
$B skill show hackernews-frontpage
$B skill run hackernews-frontpage        # JSON of 30 stories in ~200ms
$B skill test hackernews-frontpage       # runs script.test.ts against fixture

Three-tier storage

$B skill list walks all three in priority order; first hit wins. Resolved tier is printed inline next to each skill name:

Tier Path When
Project <project>/.gstack/browser-skills/<name>/ Project-specific skills (committed or gitignored)
Global ~/.gstack/browser-skills/<name>/ Per-user skills, all projects
Bundled <gstack-install>/browser-skills/<name>/ Ships with gstack, read-only

Trust model

Two orthogonal axes — daemon-side capability and process-side env — independently configured.

Axis Mechanism Default
Daemon-side capability Per-spawn scoped token bound to read+write scope (browser-driving commands minus admin: eval, js, cookies, storage). Single-use clientId encodes skill name + spawn id. Revoked when spawn exits. Always scoped — never the daemon root token
Process-side env trusted: true frontmatter passes process.env minus GSTACK_TOKEN. trusted: false (default) drops everything except a minimal allowlist (LANG, LC_ALL, TERM, TZ) and pattern-strips secrets (TOKEN/KEY/SECRET/PASSWORD, AWS_, ANTHROPIC_, OPENAI_, GITHUB_, etc.) Untrusted (must opt in)

GSTACK_PORT and GSTACK_SKILL_TOKEN are injected last, so a parent process can't override them.

Output protocol

stdout = JSON. stderr = streaming logs. Exit 0 / non-zero. Default 60s timeout, override via --timeout=Ns. Max stdout 1MB (truncate + non-zero exit if exceeded). Matches gh / kubectl / docker conventions.

How the SDK distribution works

Each skill ships its own copy of browse-client.ts at _lib/browse-client.ts, byte-identical to the canonical browse/src/browse-client.ts. /skillify copies the canonical SDK alongside every generated script. Each skill is fully self-contained: copy the directory anywhere, it runs. Version drift impossible — the SDK is frozen at the version the skill was authored against.

Atomic write discipline (/skillify D3)

browse/src/browser-skill-write.ts provides three primitives:

  • stageSkill(opts) — writes files to ~/.gstack/.tmp/skillify-<spawnId>/<name>/ with restrictive perms.
  • commitSkill(opts) — atomic fs.renameSync into the final tier path. Refuses to follow symlinked staging dirs (lstat check), refuses to clobber existing skills, runs realpath discipline on the tier root.
  • discardStaged(stagedDir)rm -rf the staged dir + per-spawn wrapper. Idempotent. Called on test failure or approval rejection.

There is no "almost shipped" state. Tests pass + user approves = atomic rename. Tests fail or user rejects = staging vanishes.

See docs/designs/BROWSER_SKILLS_V1.md for the full design rationale.


Domain-skills

Different mental model from browser-skills: agent-authored notes about a site (not deterministic scripts). One per hostname. Lifecycle:

  1. domain-skill save <host> — agent writes a note about the site (e.g., "GitHub: PR creation needs --draft flag for non-staff", "X.com: timeline uses cursor pagination, not page numbers"). Default state: quarantined.
  2. After N=3 successful uses without the L4 prompt-injection classifier flagging the note, it auto-promotes to active.
  3. domain-skill promote-to-global <host> lifts it to the global tier (machine-wide, all projects).
  4. domain-skill rollback <host> demotes; domain-skill rm <host> tombstones.

The classifier flag is set automatically by the L4 prompt-injection scan; agents do not set it manually.

Storage:

  • Per-project: <project>/.gstack/domain-skills/<host>.md
  • Global: ~/.gstack/domain-skills/<host>.md

Source: browse/src/domain-skills.ts, domain-skill-commands.ts.


Real-browser mode

$B connect launches GStack Browser — a rebranded Chromium controlled by Playwright with the Side Panel extension auto-loaded and anti-bot stealth patches applied. You watch every command tick through a visible window in real time.

$B connect              # launches GStack Browser, headed
$B goto https://app.com # navigates in the visible window
$B snapshot -i          # refs from the real page
$B click @e3            # clicks in the real window
$B focus                # bring window to foreground (macOS)
$B status               # shows Mode: cdp
$B disconnect           # back to headless mode

The window has a subtle golden shimmer line at the top and a floating "gstack" pill in the bottom-right corner so you always know which Chrome window is being controlled.

What "GStack Browser" means

Not your daily Chrome — a Playwright-managed Chromium with custom branding in the Dock and menu bar (the .app name, Dock icon, and tray, NOT the UA string), always-on Layer C anti-bot stealth (most JS-observable automation tells are masked, so many anti-bot-protected sites load cleanly), a stock-Chrome user agent that reports the underlying Chromium version, and the gstack extension pre-loaded via launchPersistentContext. The UA no longer carries a GStackBrowser suffix — that branding string was itself a high-entropy tell, so the browser now reports a plain Chrome/<version> UA. Deepest-layer CDP-protocol detection still gets through (Google can still trigger captchas; see the CDP-patch item in TODOS.md). Your regular Chrome with your tabs and bookmarks stays untouched.

When to use headed mode

  • QA testing where you want to watch Claude click through your app
  • Design review where you need to see exactly what Claude sees
  • Debugging where headless behavior differs from real Chrome
  • Demos where you're sharing your screen
  • Pair-agent sessions (the remote agent drives your local browser)

CDP-aware skills

When in real-browser mode, /qa and /design-review automatically skip cookie import prompts and headless workarounds — the headed browser already has whatever session you logged into.

Headed mode + proxy + browser-native downloads (v1.28.0.0)

Three coordinated flags for sites that block headless browsers, fingerprint Playwright defaults, or sit behind authenticated upstream proxies:

# Visible Chromium. Auto-spawns Xvfb on Linux containers without DISPLAY.
$B --headed goto https://example.com

# SOCKS5 with auth — Chromium can't prompt for SOCKS5 creds, so $B runs a
# local 127.0.0.1 bridge that handles the auth handshake.
$B --proxy socks5://user:pass@residential.proxy.host:1080 goto https://example.com

# HTTP/HTTPS proxy passes through to Chromium directly.
$B --proxy http://corp-proxy:3128 goto https://example.com

# Browser-native download for Content-Disposition, redirect chains, anti-bot
# CDNs where page.request.fetch() falls over.
$B download "https://protected.example.com/file" /tmp/file.bin --navigate

# Combined.
$B --headed --proxy socks5://user:pass@host:1080 \
   download "https://protected.example.com/file" /tmp/file.bin --navigate

Credential policy. Pass creds via the URL (socks5://user:pass@host) OR the env vars BROWSE_PROXY_USER / BROWSE_PROXY_PASS — never both. $B refuses with a clear hint when both are set; silent override created "works on my machine" debugging traps.

Daemon discipline. --proxy and --headed are daemon-startup config. A running daemon with config A meeting a new invocation with config B exits 1 with a browse disconnect hint instead of silently restarting and dropping tab state, cookies, or sessions.

Stealth scope (Layer C, always on). Every context — headless launch, --headed/--proxy, handoff, and the useragent/viewport --scale rebuild (recreateContext) — gets the full Layer C mask, no opt-in flag. Layer C masks navigator.webdriver, restores the window.chrome.* shape (runtime, app, csi, loadTimes), aligns Notification.permission with the Permissions API, reports a per-install hardwareConcurrency/deviceMemory from the host profile, sweeps the known Selenium/Phantom/Nightmare/Playwright globals, and installs a Function.prototype.toString proxy so every patched getter reports [native code] even under the depth-3 recursion check. It still does NOT fake navigator.plugins or navigator.languages — modern fingerprinters cross-check those for consistency, and synthesizing fixed values flags MORE bot-like, not less. ChromeDriver's cdc_/__webdriver runtime artifacts and the Permissions notifications tell are also cleaned up on every path.

GSTACK_STEALTH=extended (also accepts 1 or true; off by default) layers six more aggressive patches on top — WebGL renderer spoof, a faked navigator.plugins PluginArray, navigator.mediaDevices. That mode actively lies and can break sites that reflect on those properties; use it only when the default triggers detection. For gbrowser builds with the C++ patches, the GSTACK_* host-profile env (GPU vendor/renderer, UA-CH platform/model, hardware) emits the Pack 1 --gstack-gpu-vendor / --gstack-gpu-renderer / --gstack-ua-platform / --gstack-ua-model / --gstack-hw-concurrency / --gstack-device-memory switches that push the GPU/UA-CH/hardware spoof down to native code, and GSTACK_CDP_STEALTH=on (or 1/true) emits the Pack 2 --gstack-suppress-prepare-stack-trace switch (closes the Cloudflare Error.prepareStackTrace canary). On stock Playwright Chromium every one of these switches is a safe no-op.

launchHeaded / handoff also strip Playwright's automation-tell launch defaults via ignoreDefaultArgs (STEALTH_IGNORE_DEFAULT_ARGS): --enable-automation (the "Chrome is being controlled by automated test software" infobar), --disable-extensions, --disable-component-extensions-with-background-pages, --disable-popup-blocking, --disable-component-update, and --disable-default-apps.

Container support. --headed on Linux without DISPLAY walks the display range (:99, :100, ...) until xdpyinfo reports a free slot, then spawns Xvfb. Cleanup-on-disconnect validates the recorded PID's /proc/<pid>/cmdline matches Xvfb AND start-time matches before sending any signal — no PID-reuse footguns. Skips spawn entirely when WAYLAND_DISPLAY is set (Chromium uses Wayland natively). Standard Debian/Ubuntu containers work out of the box; minimal images (alpine, distroless) may need fonts/dbus/gtk libs for headed Chromium to render.

Failure modes. SOCKS5 upstream rejected or unreachable — fail-fast at startup with a redacted error after 3 retries (5s budget). Mid-stream upstream drop — bridge kills the affected client connection only; no transport retries that could corrupt browser traffic.


Side Panel + sidebar agent

The Chrome extension that ships baked into GStack Browser shows a live activity feed of every browse command in a Side Panel, plus @ref overlays on the page, plus an interactive Claude PTY inside the sidebar.

The Terminal pane (the headline)

The Side Panel's primary surface is the Terminal pane — a live claude -p PTY you can type into directly from the sidebar. Activity / Refs / Inspector are debug overlays behind the footer's debug toggle. WebSocket auth uses Sec-WebSocket-Protocol (browsers can't set Authorization on a WebSocket upgrade), and the PTY session token is a 30-minute HttpOnly cookie minted via POST /pty-session.

The toolbar's Cleanup button and the Inspector's "Send to Code" action both pipe text into the live Claude PTY via window.gstackInjectToTerminal(text), exposed by sidepanel-terminal.js. There's no separate /sidebar-command POST — the live REPL is the only execution surface.

Activity feed

A scrolling feed of every browse command — name, args, duration, status, errors. Shows up in real time as Claude works. Backed by SSE (/activity/stream) that accepts the Bearer token OR the HttpOnly gstack_sse session cookie (30-minute stream-scope cookie minted via POST /sse-session).

Refs tab

After $B snapshot, shows the current @ref list (role + name) so you can see what Claude is targeting.

CSS Inspector

Powered by $B inspect (CDP-based). Click any element on the page to see the full CSS rule cascade, computed styles, box model, and modification history. The "Send to Code" button injects a description into the Claude PTY.

Sidebar architecture

Component Where it lives Notes
Side Panel UI extension/sidepanel.js, sidepanel-terminal.js Chrome extension surface
Background SW extension/background.js Manages tab events, port management
Content script extension/content.js Page overlays, gstack pill
Terminal agent browse/src/terminal-agent.ts PTY spawn, lifecycle, auth
Sidebar utilities browse/src/sidebar-utils.ts URL sanitization, helpers

Before modifying any of these, read the comment block in CLAUDE.md under "Sidebar architecture" — silent failures here usually trace to not understanding the cross-component flow.

Manual install (for your regular Chrome)

If you want the extension in your everyday Chrome (not the Playwright-controlled one):

bin/gstack-extension    # opens chrome://extensions, copies path to clipboard

Or do it manually: chrome://extensions → toggle Developer mode → Load unpacked → navigate to ~/.claude/skills/gstack/extension → pin the extension → enter the port from $B status.

v1.63 pinned the extension identity via the manifest key field, so existing unpacked installs get a new extension ID and panel-local state (saved port) resets once — a one-time in-product notice explains this.


Pair-agent

Remote AI agents (Codex, OpenClaw, Hermes, anything that speaks HTTP) can drive your local browser through an ngrok tunnel. The whole flow is gated by a 26-command allowlist, scoped tokens, and a denial log.

How it works

/pair-agent                     # generates a setup key, prints connection instructions
# Copy the instructions to the remote agent
# Remote agent runs:
#   POST <tunnel-url>/connect with setup key → gets a scoped token (24h, single client)
#   POST <tunnel-url>/command with token → runs allowed commands

Dual-listener architecture (v1.6.0.0+)

When pair-agent activates, the daemon binds two HTTP listeners:

  • Local listener (127.0.0.1:LOCAL_PORT). Full command surface. Never forwarded by ngrok. Used by your Claude Code, the Side Panel, anything on your machine.
  • Tunnel listener (127.0.0.1:TUNNEL_PORT). Locked allowlist — /connect, /command (scoped tokens + 26-command browser-driving allowlist), /sidebar-chat. ngrok forwards only this port.

Root tokens sent over the tunnel return 403. SSE endpoints use a 30-minute HttpOnly gstack_sse cookie (never valid against /command).

The 26-command tunnel allowlist

Defined in browse/src/server.ts as TUNNEL_COMMANDS. Pure gate function canDispatchOverTunnel(command) is exported for unit testing. Set:

goto, click, text, screenshot, html, links, forms, accessibility,
attrs, media, data, scroll, press, type, select, wait, eval,
newtab, tabs, back, forward, reload, snapshot, fill, url, closetab

Notably absent: pair, unpair, cookies, setup, launch, restart, stop, tunnel-start, token-mint, state, connect, disconnect. A remote agent that tries them gets a 403 plus a fresh entry in the denial log.

Tunnel denial log

~/.gstack/security/attempts.jsonl — append-only, salted SHA-256 of source

  • domain only (no raw IP, no full request body), rotates at 10MB with 5 generations. Per-device salt at ~/.gstack/security/device-salt (mode 0600).

Tunnel egress receipts (v1.63+)

Every tunnel session open writes a hash-chained egress receipt (sink browse-tunnel) to ~/.gstack/security/egress.jsonl BEFORE ngrok forwards anything. Fail-closed: if the receipt can't be written, the tunnel listener is torn down and the start is refused. Inspect the ledger with bin/gstack-egress list and verify chain integrity with bin/gstack-egress verify (exit 3 on tamper).

See docs/REMOTE_BROWSER_ACCESS.md for the full operator guide.

Tab ownership

Scoped tokens default to tabPolicy: 'own-only'. A paired agent can newtab to create its own tab and drive that tab freely, but it can't goto, fill, or click on tabs another caller owns. tabs lists ALL tab metadata (an accepted tradeoff — see ARCHITECTURE.md), but text/html/snapshot content of unowned tabs is blocked by ownership checks.


Authentication

Three token types, three lifetimes, three scopes.

Token Generated by Lifetime Scope
Root token Daemon startup (random UUID) Daemon process lifetime Full command surface, local listener only — 403 over tunnel
Setup key POST /pair 5 minutes, one-time use Single redemption: present at /connect, get a scoped token
Scoped token POST /connect (with setup key) 24 hours Per-client, allowlist-bound, optionally tab-scoped

The root token is written to <project>/.gstack/browse.json with chmod 600. Every command that mutates browser state must include Authorization: Bearer <token>.

SSE endpoints (/activity/stream, /inspector/events) accept the Bearer token OR a 30-minute HttpOnly gstack_sse cookie minted via POST /sse-session. The ?token=<ROOT> query-param auth is no longer supported. This is what lets the Chrome extension subscribe to the activity feed without putting the root token in extension storage.

The Terminal pane uses a separate session cookie, gstack_pty, minted via POST /pty-session. Different scope — can spawn / drive the live claude PTY, can't dispatch arbitrary /command calls. /health endpoint MUST NOT surface this token.

Extension token bootstrap (v1.63+)

GET /health is liveness/status only — it never carries a token, in any mode. The Side Panel extension bootstraps the root token via POST /extension-token on the local listener. The server releases the token only when the caller's Origin is exactly chrome-extension://<GSTACK_EXTENSION_ID> — the key field in extension/manifest.json pins the extension ID (GSTACK_EXTENSION_ID in browse/src/server.ts; derivation reproducible via bun browse/scripts/extension-id.ts) — AND the parsed Host hostname is loopback. Anything else gets a detail-free 403. The endpoint is never added to TUNNEL_PATHS, so the tunnel surface 404s it by default-deny.

Token registry

browse/src/token-registry.ts handles mint/validate/revoke for all three types, plus per-token rate limiting. Setup keys are single-use; scoped tokens have a sliding 24h window; the root token is rotated on each daemon startup.


Security stack

Layered defense against prompt injection on untrusted page content.

Layer Module Lives in
L1 Datamarking content-security.ts server + page-content read path
L2 Hidden-element strip content-security.ts server + page-content read path
L3 ARIA + URL blocklist + envelope wrapping content-security.ts server + page-content read path
L4 TestSavantAI ML classifier (112MB ONNX) security-classifier.ts security sidecar subprocess*
Canary token utilities security.ts pure functions — no live injector today
combineVerdict ensemble security.ts server (inline L4 verdict path)

* security-classifier.ts cannot be imported from the compiled browse binary — @huggingface/transformers v4 requires onnxruntime-node which fails to dlopen from Bun compile's temp extract dir. The compiled binary runs L1L3 plus the pure parts of security.ts; L4 runs in a plain-Node sidecar (security-sidecar-entry.ts, spawned lazily by security-sidecar-client.ts on the first /pty-inject-scan).

Thresholds

  • BLOCK: 0.85 — single-layer score that would cause BLOCK if cross-confirmed
  • WARN: 0.75 — cross-confirm threshold in combineVerdict
  • LOG_ONLY: 0.40 — log-only floor
  • SOLO_CONTENT_BLOCK: 0.92 — single-layer threshold for label-less content classifiers

Ensemble rule

combineVerdict retains multi-layer ensemble semantics (2-of-N block votes; single-layer high confidence degrades to WARN — the Stack Overflow instruction-writing FP mitigation), but only L4 (testsavant) is live today: the Haiku transcript and DeBERTa ensemble layers were removed along with the sidebar chat pipeline that hosted them. Canary leak always BLOCKs (deterministic).

Env knobs

  • GSTACK_SECURITY_OFF=1 — emergency kill switch. Classifier stays off even if warmed. Just the ML scan is skipped.
  • Classifier model cache: ~/.gstack/models/testsavant-small/ (112MB, first run only).
  • Attack log: ~/.gstack/security/attempts.jsonl (salted SHA-256 + domain only, rotates at 10MB, 5 generations).
  • Per-device salt: ~/.gstack/security/device-salt (0600).
  • Session state: ~/.gstack/security/session-state.json (cross-process, atomic).

A shield icon in the sidebar header shows the live status. See ARCHITECTURE.md § "Prompt injection defense" for the full threat model.


Screenshots, PDFs, visual

Screenshot modes

Mode Syntax Playwright API
Full page (default) screenshot [path] page.screenshot({ fullPage: true })
Viewport only screenshot --viewport [path] page.screenshot({ fullPage: false })
Element crop (flag) screenshot --selector <css> [path] locator.screenshot()
Element crop (positional) screenshot "#sel" [path] or screenshot @e3 [path] locator.screenshot()
Region clip screenshot --clip x,y,w,h [path] page.screenshot({ clip })

Element crop accepts CSS selectors (.class, #id, [attr]) or @e/@c refs. Tag selectors like button aren't caught by the positional heuristic — use the --selector flag form.

--base64 returns data:image/png;base64,... instead of writing to disk — composes with --selector, --clip, --viewport.

Mutual exclusion: --clip + selector, --viewport + --clip, and --selector + positional selector all throw.

Retina screenshots — viewport --scale

viewport --scale <n> sets Playwright's deviceScaleFactor (context-level, 13 cap):

$B viewport 480x600 --scale 2
$B load-html /tmp/card.html
$B screenshot /tmp/card.png --selector .card
# .card at 400x200 CSS pixels → card.png is 800x400 pixels

--scale N alone (no WxH) keeps the current viewport size. Scale changes trigger a context recreation, which invalidates @e/@c refs — rerun snapshot after. HTML loaded via load-html survives the recreation via in-memory replay. Rejected in headed mode (real browser controls scale).

PDF generation

pdf accepts the full Playwright surface plus a few additions:

  • Layout: --format letter|a4|legal, --width <dim>, --height <dim>, --margins <dim>, --margin-top/right/bottom/left <dim>
  • Structure: --toc (waits for Paged.js if loaded), --outline, --tagged (PDF/A accessibility), --print-background, --prefer-css-page-size
  • Branding: --header-template <html>, --footer-template <html>, --page-numbers
  • Tabs: --tab-id <N> to render a specific tab
  • Large payloads: --from-file <payload.json> (avoids shell argv limits)

Responsive screenshots

responsive [prefix] — three screenshots in one call: mobile (375x812), tablet (768x1024), desktop (1280x720). Saves as {prefix}-mobile.png etc.

prettyscreenshot

Combines cleanup + scroll + element hide in one call:

$B prettyscreenshot --cleanup --scroll-to "hero section" --hide ".cookie-banner" /tmp/clean.png

Local HTML

Two ways to render HTML that isn't on a web server:

Approach When URL after Relative assets
goto file://<abs-path> File already on disk file:///... Resolve against file's directory
goto file://./<rel>, goto file://~/<rel> Smart-parsed to absolute file:///... Same
load-html <file> HTML generated in memory, no parent-dir context needed about:blank Broken (self-contained HTML only)

Both are scoped to files under cwd or $TMPDIR via the same safe-dirs policy as eval. file:// URLs preserve query strings and fragments (SPA routes work).

load-html has an extension allowlist (.html, .htm, .xhtml, .svg) and a magic-byte sniff to reject binary files mis-renamed as HTML. 50MB size cap (override via GSTACK_BROWSE_MAX_HTML_BYTES).

load-html content survives later viewport --scale calls via in-memory replay (TabSession tracks the loaded HTML + waitUntil). The replay is purely in-memory — HTML is never persisted to disk via state save to avoid leaking secrets or customer data.


Batch endpoint

POST /batch sends multiple commands in a single HTTP request. Eliminates per-command round-trip latency — critical for remote agents over ngrok where each HTTP call costs 2-5s.

POST /batch
Authorization: Bearer <token>

{
  "commands": [
    {"command": "text", "tabId": 1},
    {"command": "text", "tabId": 2},
    {"command": "snapshot", "args": ["-i"], "tabId": 3},
    {"command": "click", "args": ["@e5"], "tabId": 4}
  ]
}

Each command routes through handleCommandInternal — full security pipeline (scope checks, domain validation, tab ownership, content wrapping) enforced per command. Per-command error isolation: one failure doesn't abort the batch. Max 50 commands per batch. Nested batches rejected. Rate limiting: 1 batch = 1 request against the per-agent limit.

Pattern: agent crawling 20 pages opens 20 tabs (individual newtab or batch), then POST /batch with 20 text commands → 20 page contents in ~2-3 seconds total vs ~40-100 seconds serial.


Capture

Console, network, and dialog events flow into O(1) circular buffers (50,000 capacity each), flushed to disk asynchronously via Bun.write():

  • Console: .gstack/browse-console.log
  • Network: .gstack/browse-network.log
  • Dialog: .gstack/browse-dialog.log

The console, network, and dialog commands read from the in-memory buffers (not disk) so capture is real-time even when disk is slow.

Dialogs (alert, confirm, prompt) are auto-accepted by default to prevent browser lockup. dialog-accept <text> controls prompt response text.


JS execution

js runs an inline expression. eval runs a JS file. Both run in the same JS sandbox — the only difference is inline-vs-file. Both support await — expressions containing await are auto-wrapped in an async context:

$B js "await fetch('/api/data').then(r => r.json())"   # auto-wrapped
$B js "document.title"                                  # no wrap needed
$B eval my-script.js                                    # file with await

For eval files, single-line files return the expression value directly. Multi-line files need explicit return when using await. Comments containing the literal token "await" don't trigger wrapping.

Path safety: eval rejects paths outside cwd or /tmp. js doesn't read files at all.


Tabs, frames, state

Tabs

$B tabs                          # list all open tabs
$B tab 3                         # switch to tab 3
$B newtab https://example.com    # open new tab, switch to it
$B newtab --json                 # programmatic: returns {"tabId":N,"url":...}
$B closetab                      # close current
$B closetab 2                    # close tab 2
$B tab-each "text"               # run "text" on every tab, return JSON

tab-each <command> fans out a command across every open tab and returns a JSON array — handy for "give me the text of every tab I have open."

Frames

$B frame "#stripe-iframe"        # switch to iframe by selector
$B frame @e7                     # by ref
$B frame --name "checkout"       # by name attribute
$B frame --url "stripe.com"      # by URL pattern match
$B frame main                    # back to top frame

Refs are cleared on switch (the iframe has its own AX tree).

State save/load

$B state save my-session         # save cookies + URLs to .gstack/browse-state-my-session.json
$B state load my-session         # restore

In-memory load-html content is intentionally NOT persisted (avoid leaking secrets to disk).

Watch

$B watch                         # passive observation: snapshot every 5s while user browses
$B watch stop                    # return summary of what changed

Useful when you're driving the browser manually and want Claude to see what you did at the end without spamming snapshot calls.

Inbox

$B inbox                         # list messages from sidebar scout
$B inbox --clear                 # clear after reading

The sidebar scout (a background process the Chrome extension can spawn) drops notes for Claude when the user surfaces something they want noticed. Stored in .gstack/browser-scout.jsonl.


CDP

$B cdp — raw Chrome DevTools Protocol dispatch

Deny-default. Only methods enumerated in browse/src/cdp-allowlist.ts (CDP_ALLOWLIST const) are reachable; any other method returns 403. Each allowlist entry declares scope (tab vs browser) and output (trusted vs untrusted). Untrusted methods (data-exfil-shaped, e.g. Network.getResponseBody) get UNTRUSTED-envelope wrapped output.

$B cdp Page.getLayoutMetrics
$B cdp Network.enable
$B cdp Accessibility.getFullAXTree --json '{"max_depth":5}'

To discover allowed methods: read browse/src/cdp-allowlist.ts.

$B inspect — CDP-based CSS inspector

$B inspect ".header"                # full rule cascade for the header
$B inspect ".header" --all          # include user-agent rules
$B inspect ".header" --history      # show modification history

Returns the matched rule cascade with specificity, computed styles, the box model, and (with --history) every CSS modification made via $B style since the page loaded. Powered by a persistent CDP session per page in browse/src/cdp-inspector.ts.

$B ux-audit

$B ux-audit

Returns JSON with site identity, navigation, headings (capped 50), text blocks, interactive elements (capped 200) — page structure for behavioral analysis without dumping the full HTML. Used by /qa and /design-review for cheap coverage maps.


Performance

Tool First call Subsequent calls Context overhead per call
Chrome MCP ~5s ~2-5s ~2000 tokens (schema + protocol)
Playwright MCP ~3s ~1-3s ~1500 tokens (schema + protocol)
gstack browse ~3s ~100-200ms 0 tokens (plain text stdout)
gstack browse + codified skill ~3s ~200ms 0 tokens (single skill invocation)

In a 20-command browser session, MCP tools burn 30,00040,000 tokens on protocol framing alone. gstack burns zero. The codified-skill path takes a 20-command session down to a single $B skill run call.

Why CLI over MCP

MCP works well for remote services. For local browser automation it adds pure overhead:

  • Context bloat — every MCP call includes full JSON schemas. A simple "get the page text" costs 10x more context tokens than it should.
  • Connection fragility — persistent WebSocket/stdio connections drop and fail to reconnect.
  • Unnecessary abstraction — Claude already has a Bash tool. A CLI that prints to stdout is the simplest possible interface.

gstack skips all of this. Compiled binary. Plain text in, plain text out. No protocol. No schema. No connection management.


Multi-workspace

Each project root (detected via git rev-parse --show-toplevel) gets its own daemon, port, state file, cookies, and logs. No cross-workspace collisions.

Workspace State file Port
/code/project-a /code/project-a/.gstack/browse.json random (1000060000)
/code/project-b /code/project-b/.gstack/browse.json random (1000060000)

Browser-skills three-tier lookup walks project → global → bundled, so a project-tier skill at /code/project-a/.gstack/browser-skills/foo/ shadows the global ~/.gstack/browser-skills/foo/ only inside project-a.


Environment variables

Variable Default Description
BROWSE_PORT 0 (random 1000060000) Fixed port for the HTTP server (debug override)
BROWSE_IDLE_TIMEOUT 1800000 (30 min) Idle shutdown timeout in ms
BROWSE_STATE_FILE .gstack/browse.json Path to state file
BROWSE_SERVER_SCRIPT auto-detected Path to server.ts
BROWSE_CDP_URL (none) Set to channel:chrome for real-browser mode
BROWSE_CDP_PORT 0 CDP port (used internally)
BROWSE_HEADLESS_SKIP 0 Skip Chromium launch entirely (test harness only)
BROWSE_TUNNEL 0 Activate the dual-listener tunnel architecture (requires NGROK_AUTHTOKEN)
BROWSE_TUNNEL_LOCAL_ONLY 0 Test-only — bind both listeners locally without ngrok
GSTACK_BROWSE_MAX_HTML_BYTES 52428800 (50MB) load-html size cap
GSTACK_SECURITY_OFF unset Emergency kill switch — disable ML classifier
GSTACK_STEALTH unset Set to extended (also accepts 1/true) to layer six aggressive patches (WebGL spoof, faked plugins, mediaDevices) on top of Layer C. Actively lies; can break sites.
GSTACK_CDP_STEALTH unset Set to on/1/true to emit --gstack-suppress-prepare-stack-trace (gbrowser Pack 2 / B11 C++ patch only; no-op on stock Chromium)
GSTACK_GPU_VENDOR, GSTACK_GPU_RENDERER, GSTACK_GPU_CHIPSET unset Per-install GPU spoof fed to the Pack 1 WebGL/UA-CH C++ patches. Set by gbd from the host profile; emitted as --gstack-gpu-vendor / --gstack-gpu-renderer / --gstack-ua-model cmdline switches only when present.
GSTACK_PLATFORM unset Host platform classification (MacARM/MacIntelmacOS, Win32Windows, Linux*Linux) emitted as --gstack-ua-platform
GSTACK_HW_CONCURRENCY, GSTACK_DEVICE_MEMORY host profile (fallback 8) Per-install hardwareConcurrency/deviceMemory reported by Layer C and emitted as --gstack-hw-concurrency / --gstack-device-memory for the worker-navigator C++ patch

Source map

browse/
├── src/
│   ├── cli.ts                   # Thin client — reads state, sends HTTP, prints
│   ├── server.ts                # Bun HTTP daemon — routes commands, dual-listener
│   ├── browser-manager.ts       # Chromium lifecycle, tabs, ref map, crash detection
│   ├── socks-bridge.ts          # Local 127.0.0.1 SOCKS5 bridge that handles auth handshakes Chromium can't speak
│   ├── proxy-config.ts          # --proxy URL parsing + cred resolution (URL vs env, fail-fast on both)
│   ├── proxy-redact.ts          # Cred-redaction helper for any proxy URL surfaced to logs/errors
│   ├── xvfb.ts                  # Xvfb auto-spawn + orphan cleanup with PID + start-time validation
│   ├── stealth.ts               # Layer C: webdriver mask + window.chrome.* + Notification/Permissions + per-install hardware + toString proxy + automation-global sweep; buildGStackLaunchArgs (GSTACK_* cmdline switches); GSTACK_STEALTH=extended opt-in
│   ├── browse-client.ts         # Canonical SDK — what skills import as _lib/browse-client.ts
│   ├── snapshot.ts              # AX tree → @e/@c refs → Locator map; -D/-a/-C handling
│   ├── read-commands.ts         # Non-mutating: text, html, links, js, css, is, dialog, ...
│   ├── write-commands.ts        # Mutating: goto, click, fill, upload, dialog-accept, ...
│   ├── meta-commands.ts         # state, watch, inbox, frame, ux-audit, chain, diff, ...
│   ├── browser-skills.ts        # 3-tier walk + frontmatter parser + tombstones
│   ├── browser-skill-commands.ts # $B skill list/show/run/test/rm + spawnSkill
│   ├── browser-skill-write.ts   # D3 atomic stage/commit/discard helper for /skillify
│   ├── skill-token.ts           # mintSkillToken / revokeSkillToken (per-spawn, scoped)
│   ├── domain-skills.ts         # Per-site agent notes (state machine: quarantined→active→global)
│   ├── domain-skill-commands.ts # $B domain-skill save/list/show/edit/promote/rollback/rm
│   ├── cdp-allowlist.ts         # Deny-default CDP method allowlist
│   ├── cdp-bridge.ts            # CDP session lifecycle bridge
│   ├── cdp-commands.ts          # $B cdp dispatcher
│   ├── cdp-inspector.ts         # $B inspect — persistent CDP session per page
│   ├── activity.ts              # ActivityEntry, CircularBuffer, SSE subscribers, privacy filtering
│   ├── buffers.ts               # Console/network/dialog circular buffers (O(1) ring)
│   ├── tab-session.ts           # Per-tab session state (load-html replay, ref map scope)
│   ├── token-registry.ts        # Mint/validate/revoke for root + setup keys + scoped tokens
│   ├── sse-session-cookie.ts    # 30-min HttpOnly cookie for /activity/stream + /inspector/events
│   ├── pty-session-cookie.ts    # Separate scope: live Claude PTY auth
│   ├── tunnel-denial-log.ts     # ~/.gstack/security/attempts.jsonl writer (salted)
│   ├── path-security.ts         # validateOutputPath / validateReadPath / validateTempPath
│   ├── url-validation.ts        # URL safety checks for goto
│   ├── content-security.ts      # L1-L3: datamarking, hidden strip, ARIA, URL blocklist, envelopes
│   ├── security.ts              # L5 canary + L6 verdict combiner + thresholds
│   ├── security-classifier.ts   # L4 ML classifier (TestSavantAI, runs in the security sidecar)
│   ├── terminal-agent.ts        # Side Panel Claude PTY manager (auth + lifecycle)
│   ├── sidebar-utils.ts         # Sidebar URL sanitization + helpers
│   ├── cookie-import-browser.ts # Decrypt + import cookies from real Chromium browsers
│   ├── cookie-picker-routes.ts  # HTTP routes for /cookie-picker/*
│   ├── cookie-picker-ui.ts      # Self-contained HTML/CSS/JS for cookie picker
│   ├── network-capture.ts       # Network request capture for $B network
│   ├── media-extract.ts         # Media element extraction for $B media
│   ├── project-slug.ts          # Project slug derivation for state paths
│   ├── error-handling.ts        # safeUnlink / safeKill / isProcessAlive
│   ├── platform.ts              # OS detection (macOS, Linux, Windows)
│   ├── telemetry.ts             # Anonymous opt-in usage telemetry
│   ├── find-browse.ts           # Locate running daemon or bootstrap
│   └── config.ts                # Config resolution (env / files)
├── test/                        # Integration tests + HTML fixtures
└── dist/
    └── browse                   # Compiled binary (~58MB, Bun --compile)

browser-skills/
└── hackernews-frontpage/        # Bundled reference skill
    ├── SKILL.md
    ├── script.ts
    ├── _lib/browse-client.ts
    ├── fixtures/hn-2026-04-26.html
    └── script.test.ts

scrape/SKILL.md.tmpl             # /scrape gstack skill — match-or-prototype entry point
skillify/SKILL.md.tmpl           # /skillify gstack skill — codify last /scrape into permanent skill

Development

Prerequisites

  • Bun v1.0+
  • Playwright's Chromium (installed automatically by bun install)

Quick start

bun install                      # install deps + Playwright Chromium
bun test                         # all integration tests (~3s for browse-only)
bun run dev <cmd>                # run CLI from source (no compile)
bun run build                    # compile to browse/dist/browse

Dev mode vs compiled binary

During development, use bun run dev instead of the compiled binary. It runs browse/src/cli.ts directly with Bun, so you get instant feedback:

bun run dev goto https://example.com
bun run dev text
bun run dev snapshot -i
bun run dev click @e3

The compiled binary (bun run build) is only needed for distribution. It produces a single ~58MB executable at browse/dist/browse using Bun's --compile flag.

Running tests

bun test                                    # all tests
bun test browse/test/commands               # command integration tests
bun test browse/test/snapshot               # snapshot tests
bun test browse/test/cookie-import-browser  # cookie import unit tests
bun test browse/test/browser-skill-write    # D3 atomic-write helper tests
bun test browse/test/tunnel-gate-unit       # canDispatchOverTunnel pure tests

Tests spin up a local HTTP server (browse/test/test-server.ts) serving HTML fixtures from browse/test/fixtures/, then exercise the CLI against those pages.

Adding a new command

  1. Add the handler in read-commands.ts (non-mutating) or write-commands.ts (mutating), or meta-commands.ts (server / lifecycle).
  2. Register the route in server.ts.
  3. Add the entry to COMMAND_DESCRIPTIONS in browse/src/commands.ts (with a clear description and usage — the gen-skill-docs validation suite enforces no | characters in description).
  4. Add a test case in browse/test/commands.test.ts with an HTML fixture if needed.
  5. Run bun test to verify.
  6. Run bun run build to compile.
  7. Run bun run gen:skill-docs to regenerate SKILL.md (the command appears in the command-reference table downstream).

Adding a new browser-skill

For a hand-written skill: copy browser-skills/hackernews-frontpage/, update SKILL.md frontmatter, rewrite script.ts against your target site, re-capture the fixture, update the parser test. bun test validates the SKILL.md contract (sibling SDK byte-identity, frontmatter schema).

For an agent-written skill: drive the page once with /scrape <intent>, say /skillify, accept the proposed name in the approval gate. The skill lands at ~/.gstack/browser-skills/<name>/ after the test passes.

Deploying to the active skill

The active skill lives at ~/.claude/skills/gstack/. After making changes:

cd ~/.claude/skills/gstack
git fetch origin && git reset --hard origin/main
bun run build

Or copy the binary directly:

cp browse/dist/browse ~/.claude/skills/gstack/browse/dist/browse

Cross-references

  • ARCHITECTURE.md — system-level architecture, dual-listener tunnel design, prompt-injection defense threat model
  • CLAUDE.md — project-level instructions, sidebar architecture notes, security-stack constraints
  • docs/REMOTE_BROWSER_ACCESS.md — operator guide for /pair-agent (setup keys, scoped tokens, denial log)
  • docs/designs/BROWSER_SKILLS_V1.md — design doc for browser-skills runtime (Phase 1 + 2a + roadmap)
  • scrape/SKILL.md/scrape skill: match-or-prototype data extraction
  • skillify/SKILL.md/skillify skill: codify last /scrape into permanent skill
  • TODOS.md/automate (Phase 2b P0), Phase 3 resolver injection, Phase 4 eval + sandbox

Acknowledgments

The browser automation layer is built on Playwright by Microsoft. Playwright's accessibility tree API, locator system, and headless Chromium management are what make ref-based interaction possible. The snapshot system — assigning @ref labels to AX tree nodes and mapping them back to Playwright Locators — is built entirely on top of Playwright's primitives. Thank you to the Playwright team for building such a solid foundation.

The prompt-injection L4 layer uses TestSavantAI/distilbert-v1.1-32 (112MB ONNX), run locally via @huggingface/transformers.

The CDP escape hatch is gated by an allowlist directly inspired by Codex's T2 outside-voice review during the v1.4 design pass: deny-default with an explicit allowlist, not allow-default with a denylist.