mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-11 07:29:00 +02:00
* fix(plan-tune): reject never-ask on one-way ids at --write --check already ignored those prefs; --write still stored them and --stats counted them as a working NEVER_ASK. Refuse the write and count leftover on-disk prefs as INERT_ONE_WAY. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix: gstack-config get returns "" with exit 0 for keys that have no default Skill preambles read configuration with VAR=$(gstack-config get <key> 2>/dev/null || echo "<default>") and that fallback only fires on a non-zero exit. lookup_default ended in a catch-all that echoed "" and returned 0, so for any key missing from the table VAR came back empty and the default written right there in the preamble was unreachable. The skill then branched on a value it never specified: "skip entirely if QUESTION_TUNING is false", reached with QUESTION_TUNING="". Four keys that skills actually read had no entry and took that path: question_tuning -> callers assume "false" repo_mode -> callers assume "unknown" team_mode -> callers assume "false" transcript_ingest_mode -> callers assume "off" Each default above is the value the call sites already substitute in their own `|| echo` fallback, so this only makes reachable what was already intended. The catch-all now returns non-zero. That is deliberately scoped to the unknown-key arm alone: keys whose default is intentionally empty still exit 0, because "" is their real answer and their callers depend on it -- cross_project_learnings ("unset triggers the first-time prompt"), redact_repo_visibility ("empty falls through to gh/glab detection"), salience_allowlist, user_slug_at_*. Making every empty answer an error would have broken those. test/gstack-config-defaults.test.ts pins the class rather than the four instances: it parses the case arms and asserts every `gstack-config get <key>` site in the tree is covered, so adding a read without a default fails CI. It also pins the exit-code contract in both directions. Verified failing against the pre-fix script, where it names exactly those four keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(redact): a typo'd subcommand no longer exits 0 having done nothing main() recognised exactly two subcommands and let everything else fall through to the stdin scan. On empty stdin that prints "(no findings)" and exits 0, so: $ gstack-redact install-prepush-hooks # plural typo gstack-redact scan — repo UNKNOWN (no findings) $ echo $? 0 No hook was installed, and the operator has every reason to believe the credential guard is armed. A guard that silently no-ops must never exit 0. Two smaller faults in the same dispatch, both of which lead people here: - There was no --help handler, so `gstack-redact --help` fell through to the scanner. Piping a credential to it scanned the secret and exited 3. - With no piped input and no --from-file, readInput() blocks on readSync(fd 0) until an EOF that an interactive terminal never sends. That prints nothing at all, so it reads as a hang rather than as "this is a filter, feed it". Now: --help/-h/help prints usage and exits 0; an unrecognised positional prints the offender and exits 1; a TTY with nothing piped in prints usage instead of blocking. "scan" stays accepted, because the human output header reads "gstack-redact scan — repo …" and that is what people type. Usage errors exit 1, deliberately not 2 or 3. Those mean MEDIUM and HIGH findings and callers gate dispatch on them, so a usage error exiting 2 would be read as "medium findings — prompt the user". A test pins that. Tests: 4 written failing first, then fixed. Full suite 7,722 pass / 0 fail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(browse): one ambiguous ref no longer kills the whole annotated screenshot `snapshot -a` exits 1 with "Selector matched multiple elements" on most real pages, so /qa, /canary and /land-and-deploy silently produce reports whose screenshots do not exist. Plain `screenshot <path>` is unaffected. Refs are built as getByRole(role, {name}) and disambiguated with .nth() when role+name repeats. That disambiguation cannot fire for a node with NO accessible name: the locator degrades to getByRole(role) with no name filter, and the count driving .nth() is taken from the FILTERED aria snapshot while getByRole matches the unfiltered DOM. Measured on a live page: the tree surfaced 2 unnamed paragraphs, the DOM had 9. Landmarks (banner/main/contentinfo) and paragraphs are correctly unnamed per ARIA, so this is the common case rather than an edge case. boundingBox() then hits Playwright strict mode, and the catch allowlisted only timeout/closed/Target/Execution-context messages — so the strict-mode error was re-thrown and aborted every remaining annotation. Two changes: - `.first()` before boundingBox(), so an ambiguous ref draws a box on its first match instead of aborting. The heatmap path below has always tolerated this via a bare `catch {}`; annotate was the only path that could be killed outright. - the catch no longer re-throws on unrecognised messages. A box we cannot measure is a box we do not draw, never a reason to lose the rest of the page. Set BROWSE_DEBUG to see what was skipped. Also: `-o` passed without `-a`/`-H` was silently ignored (exit 0, no file), which reads as "screenshots are broken" rather than "you forgot a flag". It now warns and points at `browse screenshot <path>`. Verified by rebuilding both ways against the same page with 51 refs present: before — "Selector matched multiple elements", no file written after — exit 0, 229KB PNG Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(version-bump): missing or empty VERSION no longer repairs a fabricated 0.0.0.0 into package.json repair now fails with exit 2 when the VERSION file is absent or empty instead of folding to DEFAULT ("0.0.0.0") — which passed VERSION_RE and regressed package.json below where it started. classify gains an additive versionFileExists field so /ship can tell a real 0.0.0.0 from a fabricated one. Re-derived from PR #2612 under the generated-file screening rule. Fixes #2600 (repair half; the path-configurability half landed in v1.67 via #2531). Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(memory-ingest): --probe counts post-attribution, through the same gate --bulk uses probeMode previously stat'd every walked file, so setup-gbrain gated its silent bulk ingest on pre-filter counts that the write path would never ingest (#2394). The attribution decision now lives in ONE shared gate (sessionIsAttributable — cheap-parse: cwd extraction + memoized resolveGitRemote, never a full page build) used by BOTH probeMode and preparePages, so the two stages' post-attribution counts are structurally identical. ProbeReport gains skipped_unattributed; the probe prints what it excluded and --include-unattributed restores raw counts. The parity is pinned at the prepare stage (probe post-attribution == transcripts reaching import), deliberately NOT == final written. Re-derived from PR #2612 under the generated-file screening rule; the shared-gate design and the remote memo are additions from the plan review. Fixes #2394. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(browse): allow CPU and network throttling for performance measurement Adds Emulation.setCPUThrottlingRate and Network.emulateNetworkConditions to CDP_ALLOWLIST. Motivation: diagnosing a real "uploads take 1-2 minutes" report, the only machine available was a fast developer workstation. Client-side processing measured 1.4s where the user experienced minutes, so the conclusion had to be reached arithmetically rather than observed. Throttling would have let the measurement reproduce the reporter's conditions directly. Both fit the existing posture rather than widening it: - Emulation already allows setDeviceMetricsOverride, clearDeviceMetricsOverride and setUserAgentOverride, which are equally mutating and scoped to the tab. - Neither method reads page content. setCPUThrottlingRate affects only timing; emulateNetworkConditions constrains traffic rather than inspecting it, so no request bodies, headers or cookies are exposed. Both are output: 'trusted' because they return no page-derived data. scope 'tab' for both, matching the surrounding Emulation entries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(session-update): lock pidfile records the live holder; hard TTL bounds every wedge (#2613) echo $$ inside the backgrounded subshell recorded the PARENT hook's PID — which exits immediately — so every subsequent session judged the lock stale and rm -rf'd a LIVE holder's lock, letting concurrent updaters run over each other. The pidfile now records ${BASHPID:-$(sh -c 'echo $PPID')} (macOS bash 3.2 has no BASHPID; the sh child's PPID is exactly this subshell). Staleness is now two independent detectors: PID liveness (as before, but against the real holder), and a 30-minute hard TTL on the heartbeat mtime — reclaimed regardless of kill -0, so a recycled PID or hung holder can't wedge the lock forever. The holder touches the pidfile after the pull and after setup, so a legitimately-slow run keeps itself alive. Empty and missing pidfiles are respected inside the TTL window (the mkdir→echo race) and reclaimed past it. Fixes #2613. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(browse): explicit windowsHide on every Bun.spawn site + census tripwire (#2575 residual) Bun.spawn sites were structurally outside the windowsHide census (it swept child_process bindings only). The runtime was already safe — native Bun hides consoles by default and bun-polyfill.cjs defaults windowsHide !== false since #2523/#2539 — but implicit defaults are exactly what regress silently. Every Bun.spawn/spawnSync in browse/src now carries the explicit flag (harmless on unix-only sites like Xvfb/xattr/open), and a second SWEEP in windows-spawn-hide.test.ts fails CI on any new flagless Bun.spawn site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain): brain worktree advances on the daily sync — no more silently stale brains (#2516) The daily pull refreshed only ~/.gstack itself, never the detached worktree at ~/.gstack-brain-worktree that gbrain actually indexes — so after setup the brain served stale pages forever unless setup-gbrain/sync-gbrain happened to run. brain-sync --once now advances the worktree once per 24h behind an ATTEMPT stamp (.brain-worktree-last-advance — a persistently-failing advance warns once a day, not at every skill boundary), inside the existing run lock and before any ingest step touches the worktree. The new gstack-gbrain-source-wireup --advance-only is built for the unattended cadence: git-only (no gbrain prereqs), pins every operation to the managed worktree (refuses paths that are not worktrees of the artifacts repo), refuses dirty worktrees, and never runs the force-remove recovery — a cron path must not be able to delete local changes. A static pin keeps the force-remove out. docs/gbrain-sync.md stops overclaiming the old cadence. Fixes #2516. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(memory-ingest): honor the per-remote deny/read-only trust policy (#2392) Transcript ingest now respects the same trust store as code import — the gate existed only in gstack-gbrain-sync's runCodeImport, so memory-ingest happily ingested transcripts from deny-listed repos. preparePages filters prepared transcript pages through ONE batch policy lookup (new 'get --batch' verb on bin/gstack-gbrain-repo-policy — the script owns URL normalization; the client adds repoPolicyTierBatch, one spawn for all distinct remotes, so large corpora never pay a 10s-timeout subprocess per remote). Outcomes match code-import semantics: read-only → clean skip (skipped_policy_readonly), deny → counted refusal (skipped_policy_deny), corrupted/unreadable store → HARD ERROR before any write (state, staging, egress receipt, and import all untouched) with the recovery command named — policy corruption must never read as successful ingestion. Artifacts are never policy-filtered (their git_remote is a project slug, not a remote). Fixes #2392. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): repo_mode keeps its empty no-default semantics (#2611 follow-up) The ported defaults table synthesized repo_mode → "unknown", but EMPTY is load-bearing for that key: gstack-repo-mode treats any non-empty answer as a user override and skips its own repo classification — the synthesized default turned the classifier into dead code (REPO_MODE=unknown everywhere; caught by test/gstack-repo-mode.test.ts via the wave's cross-agent blame protocol). repo_mode joins the empty-is-real carve-outs (empty output, exit 0). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): consent before killing a healthy headless daemon The pair-agent headed switch spawned 'connect --force-restart' unconditionally — auto-killing a live headless daemon (open tabs, cookies, logins) in direct contradiction of the iron rule it sits beside ('only an explicit --force-restart may kill a live daemon'). The CLI now captures daemon liveness BEFORE ensureServer (which can itself boot a fresh daemon) and relaunches only when the user passed --force-restart to pair-agent; otherwise it prints the tab count and continues against the existing daemon. The /pair-agent skill gains a matching one-way-door consent question (template half rides the wave's template block). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain-status): MCP scoping is per-project, and project-local beats user scope hasRemoteOnlyGbrainMcp scanned EVERY project's mcpServers in ~/.claude.json, so one project's remote gbrain registration reclassified broken local engines as thin-client machine-wide. It now reads user scope plus only the cwd's nearest-ancestor project key. The precedence itself was verified empirically and hermetically (fake HOME + CLAUDE_CONFIG_DIR fixtures, claude 2.1.233): with both scopes defining gbrain, 'claude mcp get gbrain' reports Scope: Local config — PROJECT-LOCAL WINS. Both in-repo consumers assumed the opposite; brain-cache's endpoint resolution flips to nearest-ancestor-project-first, and the stale user-first pin in brain-cache-roundtrip now pins the verified precedence. (The user-first jq in the brain-sync preamble resolver gets the same swap in the template block.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): gstack-slug matches remote-slug's owner-repo canonical form (live misfile bug) Found live during this wave's CEO review: bin/gstack-slug emitted SLUG=garrytan for this garrytan/gstack worktree while remote-slug correctly gave garrytan-gstack — decisions, timeline, ceo-plans, and learnings were filing into the wrong project store (observed polluting Context Recovery with another repo's decisions). Root cause: a stray empty ~/.git directory made the walk-up crown $HOME as the outermost project root; the remote lookup ran only against that root, failed silently, and the basename fallback cached 'garrytan' sticky. NOT worktree-specific — any strong marker on a non-repo ancestor triggered it. Fix: the walk now finds the outermost ancestor whose .git actually resolves an origin remote and derives owner-repo with remote-slug's byte-identical parse; marker-only ancestors keep anchoring the basename fallback but can no longer shadow a real remote. A new cache self-heal recomputes the poisoned shape (cached == basename of a marker root while a remote-bearing repo exists below), preserving legit #2212 stickiness. Nested-repo walk-up, no-remote and non-git fallbacks, and the SLUG=/BRANCH= eval contract are unchanged, pinned by a 10-case parity suite. Store migration for pre-fix data is tracked in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(brain-sync): per-record spool dir — the enqueue/drain race dies structurally Producers appended lines to .brain-queue.jsonl while the drain re-read and os.replace'd it; the in-code comment admitted a lockless append between the re-read and the replace was lost. Locks and rename-rotation designs were both reviewed and rejected (each retained a tail race); the shipped design is a maildir-style spool: one FILE per record in .brain-queue.d/ (tmp + atomic rename), the drain snapshots filenames, processes, and deletes exactly what it snapshotted. Writer and drainer never share an inode — nothing to race. Semantics: at-least-once (a crash between process and unlink re-drains; downstream content-hash dedup absorbs duplicates); retained (privacy-held) records keep their files; unparseable records are kept + warned, never destroyed. Legacy .brain-queue.jsonl migrates atomically on the next drain (crash-leftover .migrating files recovered too); status/drop-queue count both surfaces; discover-new writes spool records and advances its cursor per-record-written. The preamble's queue-depth line switches to spool count in this wave's template block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin-context): native slug fallback walks up like bash gstack-slug slugFromEnvironment derived the slug from the INNERMOST repo's origin while bash gstack-slug walks to the outermost project root — nested/vendored repos split their stores across the bash/native boundary (win32 hits the native path constantly). The native fallback now ports _outermost_project_root faithfully (strong/weak markers, outermost-strong-wins, 64-depth cap, fixed-point termination) plus the full resolution order: env override → walk-up → sticky cache with the #1125 self-heal → remote get-url → basename. Twelve mirrored scenarios drive BOTH implementations against the same fixtures and pin identical slugs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): git fallback queries the live remote, never mutates, and keeps 3-digit width The degraded path counted every remote-tracking ref on every remote — stale experiment branches and second remotes inflated version allocation, and a failed base read flipped 3-digit repos to 4-digit slots. Now: ls-remote --heads origin first (GIT_TERMINAL_PROMPT=0, 5s timeout, zero local ref mutation); on failure, local refs/remotes/origin ONLY with an explicit stale-refs warning; a failed base read zeroes at the LOCAL version file's width so a 3-digit repo allocates 0.0.1, not 0.0.1.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): hooks register the global-install path and re-point stale ones Registering hooks from a dev worktree baked that worktree's absolute path into settings.json — deleting the worktree left a dead hook erroring on every session stop, and the presence-only dedup (list-sources | grep) could never re-point it. setup's hook paths now route through _hook_install_path (global install preferred, source dir fallback), and the new ensure-event verb on gstack-settings-hook compares the registered command payload against canonical: identical → no write, different → single atomic replacement (never zero or two registrations). The plan-tune hooks had the same stale pattern and get the same fix without re-triggering their consent prompt. Also hardened: bun 1.3.13 turns an uncaught sync fs error in bun -e into a SILENT exit 0 — the registrar's write path now catches, prints, and exits 1, so a failed update can never report fake-green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): learnings capture is unconditional at completion (#2402) 43 of 44 learnings entries came from explicit /learn — the completion-status prose read 'if you discovered a durable project quirk... log it', which models treated as optional. The step now ALWAYS runs: review the session for durable learnings, log each one, and state 'No durable learnings this session' explicitly when the review comes up empty — an empty result, never a skipped step. Re-derived from PR #2612 under the generated-file screening rule. Fixes #2402. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(scrape): untrusted-content warning on the page-fetching skills (#2441) /scrape and /skillify consumed page content with zero injection guidance — the CHANGELOG claimed coverage the skills didn't have. The warning now lives in ONE exported const (UNTRUSTED_CONTENT_WARNING in resolvers/browse.ts), embedded in the browse COMMAND_REFERENCE as before AND injected standalone into both skills via the new {{UNTRUSTED_CONTENT_WARNING}} token — single source, wording can never drift between surfaces. Re-derived from PR #2612 under the generated-file screening rule. (Structural isolation for skillify-generated code is tracked as its own TODO.) Fixes #2441. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): checklist paths resolve from the installed skill root (#2518) /review Step 2 read .claude/skills/review/checklist.md — a path relative to the TARGET repo, which only resolves in gstack's own checkout. Every checklist/greptile-triage/TODOS-format reference (six across five templates — two more than the issue named, same class) now uses the installed-root form ~/.claude/skills/gstack/review/... that the templates' other references already use. The install-root class itself (non-default install dirs) is #1882, deliberately its own PR. Fixes #2518. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): one-way-door consent question before a daemon relaunch (template half) The skill flow now checks daemon liveness before Step 4 and asks an explicit one-way-door question (tabs/cookies/logins are lost) before passing --force-restart — never proceeding on a vague reply. Pairs with the CLI-half commit that stopped pair-agent auto-killing live daemons. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(codex): resume does not amortize the ~21K session prelude (#2387) Measured (#2387): every codex exec call pays Codex's session prelude, and a resumed call came in slightly ABOVE a fresh one — resume buys continuity, never token savings. The skill now says so where the resume flow lives: prefer one codex call per skill, batch questions into it. Fixes #2387. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(upgrade): fast-forward first; reset --hard only behind a proved-safe gate (#2517) /gstack-upgrade went straight to stash + reset --hard origin/main. Now it tries git pull --ff-only --autostash first (the same policy session-update's auto-upgrade uses). The destructive fallback runs unprompted ONLY when both git status --porcelain AND git rev-list origin/main..HEAD are empty — a clean tree with unpushed local commits is NOT safe, reset destroys them. Anything else requires an explicit one-way-door confirmation that lists every dirty file and unpushed commit being discarded. Fixes #2517. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): brain-sync block counts the spool queue and resolves MCP project-first Two resolver halves deferred from earlier wave commits: the queue-depth line counts .brain-queue.d/*.json spool records (plus legacy lines until the drain migrates them), and GBRAIN_MCP_ENTRY_JQ swaps its operands to nearest-ancestor-project-first — matching the empirically verified Claude Code precedence (project-local beats user scope) instead of the backwards user-first assumption. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + golden fixtures (single regen for the template block) Pure generator output for the six template/resolver commits above (learnings capture, untrusted-content warning, review paths, pair-agent consent, codex resume note, upgrade ff-only, brain-sync block) — bun run gen:skill-docs + --host codex + --host factory, with the three ship golden fixtures refreshed per the documented procedure. The three sidecar-path pins in gen-skill-docs.test.ts move to the new installed-root/$GSTACK_ROOT contract (#2518). Restores template freshness; full suite green from here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: TODOS.md — strike the six wave-fixed residuals, add two follow-ups The v1.67 adversarial-review residuals section shrinks to the one item the wave couldn't reach (iOS tap routing — needs real-device verification). New entries: skillify structural isolation (a prose warning is not a boundary for page-derived generated code) and the slug store migration (pre-fix sessions on stray-marker machines filed data under the degraded slug; post-fix reads go to the correct store, so history needs a merge/alias). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align cross-cutting pins with the wave's contracts Three suites pinned pre-wave behavior: browse's gstack-config test asserted the old unknown-key ''/exit-0 shape (#2611 made it exit 1); the Windows-paths suite pinned O_APPEND enqueue atomicity (the spool design satisfies the same invariant via tmp + os.replace, one file per record — pinned in its new form); and nine carve-guard skeleton ceilings absorbed the #2402 unconditional-learnings prose (~450B per skill), bumped with measured values per the guard's own protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: re-anchor the referenced-path scanner self-check to the gstack-rooted review refs The self-check pinned the review checklist as a class-1 alias-relative ref; #2518 moved those refs to the installed gstack root (class 2). The guard now proves the scanner sees them in their new class, so the class-2 assertion can't go vacuous. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: pin the wave's prose-tier behaviors (ship coverage-audit gap closure) The coverage audit found one regression-shaped gap: nothing pinned that the upgrade template's ff-only pull precedes the gated reset --hard (#2517) — a future template edit reverting to reset-first would fail nothing. Pinned: the ordering, the FF_OK gate, and the unpushed-commits check. Also pinned the two minor gaps: the {{UNTRUSTED_CONTENT_WARNING}} injection points in scrape/skillify (#2441) and brain-uninstall's spool-dir cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review round — 8 auto-fixes + 8 accepted findings hardened The ship review army (4 specialists + red-team + checklist, 29 findings) produced 8 mechanical auto-fixes and 11 decisions; the accepted set: - win32 slug parity completed: lib/bin-context.ts gains the remote-first outermost walk + degraded-cache self-heal the bash side got this wave — the two implementations now agree on the stray-marker live-bug shape, pinned by shared fixtures (multi-specialist 9/10 finding). - probe honors the plan's bounded-read decision: 256KB prefix, extraction semantics mirrored from parseTranscriptJsonl so probe/prepare can never diverge on the same file (>1MB transcript test). - policy normalize parity: bash normalize() now matches canonicalizeRemote on .git/-trailing and uppercase-.GIT shapes (7-shape corpus pinned two ways) — a deny for those shapes could previously slip the transcript gate. - session-update reclaim is TOCTOU-safe (atomic mv-aside on both branches). - settings-hook: unparseable settings.json errors instead of being replaced with {}; ensure-event keys on (event, source) so matcher changes update in place — never zero or two registrations. - dot-only slug guard at both parse sites (hostile 'url = ..' can't escape projects/); enqueue tmp-file janitor (1h TTL, inside the drain lock); brain-sync .migrating never clobbered; drop-queue/status count .migrating; snapshot -o warning correct + surfaced in diff mode; version-bump test order-dependence removed; uninstall clears the advance stamp. Deferred with record: slug heal-probe cost sentinel (P3 TODO), FF_OK conflation (noted, misdiagnosis-only). 270 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial round — the P0 finalize fail-safe and 12 hardened findings Three adversarial passes (Claude fresh-context, Codex chaos, Codex structured with P1 gate) on the full wave diff. Multi-source findings, all fixed: - P0: finalize_queue is now explicit-delete-only — a record is unlinked ONLY when classification proves it staged or dropped; a classifier crash, a missing class file, or a malformed pulled .brain-privacy-map.json (which previously nuked the whole snapshotted queue, remotely triggerable) now retains everything, warns, and re-drains next run. load_privacy_map treats corrupt maps as retain-all, never as empty. - next-version cannot silently drop a live claim: unreadable advertised refs get a targeted --depth=1 fetch + retry; still-unreadable claims surface as UNKNOWN warnings instead of duplicate-version silence. - session-update lock: ownership-checked EXIT trap (a TTL-reclaimed holder can no longer delete the new holder's lock) + a 5-min background heartbeat so a legitimately-slow pull/setup is never reclaimed while alive. - ensure-event collapses ALL same-(event,source) duplicates to one canonical entry; unique per-process tmp path; setup call sites surface (not swallow) the hardened refusals. - memory-ingest: --limit counts only policy-permitted pages (denied records no longer starve permitted ones); --probe applies the same policy filter as --bulk (skipped_policy_* fields on the report). - version-bump repair accepts a genuine literal 0.0.0.0 VERSION file. - slug heal restricted to the stray-.git shape — package.json-anchored wrapper roots keep their legit sticky identity (#2212 preserved). - brain-sync: idle fast path sees leftover .migrating records; unparseable spool records quarantine instead of warning forever; migration comment stops overclaiming the transition-window race. - CDP throttling justifications document override persistence (callers own restoration), pinned in the allowlist test. Deferred with record: deny retroactivity for already-ingested pages (P2 TODO, same semantics as the code-import gate); legacy-migration tail race (transition-window, requires pre-spool writers). 288 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + goldens (Windows-separator jq fix) Pure generator output for the brain-sync block's jq ancestor match now accepting backslash-formed Windows project keys — previously project-scoped brains were invisible on Windows while the TS scope resolvers saw them. Golden ship fixtures refreshed per the documented procedure. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: codex verify-pass residuals — chunked cwd read, post-filter partial count, migrating depth The verify re-review passed the P1 gate (0 P1s) and left three residuals, all applied: transcriptCwdFromPrefix reads in chunks until one complete record (4MB cap) so a giant first prompt can't truncate mid-JSON and break probe/bulk parity; partial_pages derives from the FINAL prepared set instead of the whole scanned corpus; the preamble queue-depth line counts leftover .brain-queue.jsonl.migrating records like the status path does (regen + goldens included). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.68.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.68.0.0 BROWSER.md: fix the $B cdp example (positional JSON params, not --json; depth is the real CDP param) and add the new perf-throttling examples (Emulation.setCPUThrottlingRate, Network.emulateNetworkConditions) with their clear-override counterparts. USING_GBRAIN_WITH_GSTACK.md: the state-files table row for the sync queue now names the maildir-style spool dir .brain-queue.d/ that replaced .brain-queue.jsonl this release. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align memory-pipeline probe pins with the #2394 stage-count contract The paid-tier E2E pinned the pre-fix contract (probe headline = raw discovered). Probe now counts post-attribution — the same gate --bulk uses — with an explicit unattributed-skip line. Adds the --include-unattributed companion pin so all 9 fixtures stay accounted for. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): batch missing-tip fetches — one bounded round trip, never a per-branch crawl The targeted-fetch retry for branches whose advertised tip has no local object ran ONE git fetch per branch (10s cap each). On a shallow clone against a busy remote that crawls the network for minutes — CI's shard deadline killed the free suite mid-file. Missing tips now collect into a single batched shallow fetch (15s cap); refs still missing after the batch (one unservable ref fails the whole transfer) get a capped per-branch retry, and anything past the cap warns as an UNKNOWN claim instead of fetching. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the batched fetch + make the offline-contract tests hermetic Two new G2 pins: N unfetched claim branches resolve with exactly ONE fetch spawn (PATH-shimmed git counts invocations), and one unservable ref no longer poisons the batch — live claims resolve via the bounded retry while only the ghost warns UNKNOWN. The #2545 offline-contract tests now run the CLI in a local fixture repo instead of the repo's own checkout: the checkout path did a live ls-remote against the real origin (operator-network-dependent, and the CI shard-deadline hang). The online-contract test gains a succeeding gh stub, so fallback:null is asserted deterministically instead of only when the operator happens to be authed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(redact-cli): derive the synthetic AWS-key fixture — no contiguous credential literal in source The CI quality gate scans every ADDED diff line with the redact engine, so the #2610 port's raw fixture literals failed the very gate they exist to test. The fixture is now assembled at runtime; the scanner still receives the identical bytes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the fixture's host via origin-URL sniff — kills the last environment dependence The hermetic offline-contract fixture had no origin remote, so detectHost() fell through to auth probes: a machine with glab authed passed via the gitlab path while a bare CI runner read host:unknown (offline stays false there) and failed. The fixture now pushes to a local bare origin at a path containing github.com — the URL sniff pins host:github identically everywhere, asserted explicitly in both tests, with every git call still local. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: benjamin beres <benjamin.beres@bienpreter.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Ricky <ricky@kinokostudio.com.hk> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: henbima <henbima@gmail.com>
1047 lines
44 KiB
TypeScript
1047 lines
44 KiB
TypeScript
/**
|
|
* Terminal Agent — PTY-backed Claude Code terminal for the gstack browser
|
|
* sidebar. Translates the phoenix gbrowser PTY (cmd/gbd/terminal.go) into
|
|
* Bun, with a few changes informed by codex's outside-voice review:
|
|
*
|
|
* - Lives in a separate non-compiled bun process from the browse daemon so
|
|
* a bug in WS framing or PTY cleanup can't take down the command surface.
|
|
* - Binds 127.0.0.1 only — never on the dual-listener tunnel surface.
|
|
* - Origin validation on the WS upgrade is REQUIRED (not defense-in-depth)
|
|
* because a localhost shell WS is a real cross-site WebSocket-hijacking
|
|
* target.
|
|
* - Cookie-based auth via /internal/grant from the parent server, not a
|
|
* token in /health.
|
|
* - Lazy spawn: claude PTY is not spawned until the WS receives its first
|
|
* data frame. Sidebar opens that never type don't burn a claude session.
|
|
* - PTY dies with WS close (one PTY per WS). v1.1 may add session
|
|
* survival; for v1 we match phoenix's lifecycle.
|
|
*
|
|
* The PTY uses Bun's `terminal:` spawn option (verified at impl time on
|
|
* Bun 1.3.10): pass cols/rows + a data callback; write input via
|
|
* `proc.terminal.write(buf)`; resize via `proc.terminal.resize(cols, rows)`.
|
|
*/
|
|
import * as fs from 'fs';
|
|
import * as path from 'path';
|
|
import * as crypto from 'crypto';
|
|
import { writeSecureFile, restrictFilePermissions, mkdirSecure } from './file-permissions';
|
|
import { atomicWriteSync, atomicWriteQuiet } from '../../lib/fs-atomic';
|
|
import { safeUnlink } from './error-handling';
|
|
import { writeAgentRecord, clearAgentRecord } from './terminal-agent-control';
|
|
import { findAvailablePort } from './port-allocator';
|
|
import { extractPtyCookie } from './pty-session-cookie';
|
|
|
|
const STATE_FILE = process.env.BROWSE_STATE_FILE || path.join(process.env.HOME || '/tmp', '.gstack', 'browse.json');
|
|
const PORT_FILE = path.join(path.dirname(STATE_FILE), 'terminal-port');
|
|
const BROWSE_SERVER_PORT = parseInt(process.env.BROWSE_SERVER_PORT || '0', 10);
|
|
const BROWSE_OWNER_PID = parseInt(process.env.BROWSE_OWNER_PID || '0', 10);
|
|
const OWNER_WATCHDOG_MS = parseInt(
|
|
process.env.GSTACK_TERMINAL_OWNER_WATCHDOG_MS || '15000',
|
|
10,
|
|
);
|
|
const EXTENSION_ID = process.env.BROWSE_EXTENSION_ID || ''; // optional: tighten Origin check
|
|
const INTERNAL_TOKEN = crypto.randomBytes(32).toString('base64url'); // shared with parent server via env at spawn
|
|
/**
|
|
* Per-boot generation identifier. Loopback /internal/* callers include
|
|
* `X-Browse-Gen: <CURRENT_GEN>` so a slow agent the watchdog respawned
|
|
* around can't service a stale grant from the prior generation. Absent
|
|
* header means "legacy caller" and is accepted (backward compat); a
|
|
* present-but-mismatched header returns 409 stale generation.
|
|
*/
|
|
const CURRENT_GEN = crypto.randomBytes(16).toString('base64url');
|
|
|
|
// In-memory attach-token registry. Parent posts /internal/grant after
|
|
// /pty-session; we validate WS upgrades against this map.
|
|
//
|
|
// v1.44+: each token is bound to a v1.44 sessionId (the stable, non-secret
|
|
// identifier from browse/src/pty-session-lease.ts). The token grants ONE
|
|
// attach for ONE session — re-attach within the lease window comes through
|
|
// /pty-session/reattach, which mints a fresh token for the same sessionId.
|
|
//
|
|
// Legacy callers can still pass `{token}` without sessionId (the value
|
|
// stays null and the WS upgrade still works); those callers don't get
|
|
// re-attach because there's no stable identifier to match against.
|
|
const validTokens = new Map<string, string | null>(); // token → sessionId
|
|
|
|
/**
|
|
* Reverse index for re-attach lookups: sessionId → live PtySession.
|
|
* Populated when a WS first attaches with a known sessionId; cleared when
|
|
* the session is disposed or the lease expires. Used by:
|
|
* - /ws upgrade: if the incoming attachToken maps to a sessionId that
|
|
* already has a live session, REPLACE its ws ref instead of spawning.
|
|
* - /internal/restart: enumerate by sessionId, dispose that one session.
|
|
*
|
|
* Kept separate from the WeakMap<ws,PtySession> so re-attach can find the
|
|
* session by id even after the original ws has gone.
|
|
*/
|
|
const sessionsById = new Map<string, PtySession>();
|
|
|
|
// Active PTY session per WS. One terminal per connection. Codex finding #4:
|
|
// uncaught handlers below catch bugs in framing/cleanup so they don't kill
|
|
// the listener loop.
|
|
process.on('uncaughtException', (err) => {
|
|
console.error('[terminal-agent] uncaughtException:', err);
|
|
});
|
|
process.on('unhandledRejection', (reason) => {
|
|
console.error('[terminal-agent] unhandledRejection:', reason);
|
|
});
|
|
|
|
export interface PtySession {
|
|
proc: any | null; // Bun.Subprocess once spawned
|
|
cols: number;
|
|
rows: number;
|
|
cookie: string;
|
|
/**
|
|
* Current attached websocket. Swapped on re-attach (Commit 3): when a new
|
|
* WS upgrade matches this session's sessionId, the old liveWs is gone
|
|
* and the new ws takes its place. The PTY on-data callback closes over
|
|
* `session`, not the original `ws`, so it always writes to the current
|
|
* liveWs (or skips the write when detached and liveWs is null).
|
|
*/
|
|
liveWs: any | null;
|
|
/**
|
|
* v1.44+ stable session identifier (from pty-session-lease). Null for
|
|
* legacy /internal/grant callers that didn't pass one. Used for
|
|
* targeted /internal/restart and Commit 3 re-attach lookups.
|
|
*/
|
|
sessionId: string | null;
|
|
spawned: boolean;
|
|
/**
|
|
* 25s server-side WS keepalive interval (v1.44+). Set in the WS `open`
|
|
* handler, cleared in `close`. We send `{type:"ping",ts}` text frames so
|
|
* NAT boxes, proxies, and Chrome's MV3 panel-suspend heuristics see the
|
|
* connection as active; the client either replies with `{type:"pong"}`
|
|
* or fires its own 25s `{type:"keepalive"}` cycle. Either path keeps
|
|
* the underlying TCP from being silently dropped.
|
|
*/
|
|
pingInterval: ReturnType<typeof setInterval> | null;
|
|
/**
|
|
* Commit 3 scrollback ring buffer. Each PTY write appends a frame; the
|
|
* total byte count is capped at RING_BUFFER_MAX_BYTES with oldest frames
|
|
* evicted first. On re-attach, the surviving frames are replayed as a
|
|
* single binary frame (prefixed with the v1.44 reset sequence) so the
|
|
* user sees their last screen of output. Frame boundaries preserve UTF-8
|
|
* + ANSI-CSI boundaries because each frame is the exact buffer that
|
|
* spawnClaude's on-data callback emitted.
|
|
*/
|
|
ringBuffer: Buffer[];
|
|
ringBufferBytes: number;
|
|
/**
|
|
* Tracks whether the PTY is currently in xterm alt-screen mode. claude's
|
|
* TUI enters alt-screen (CSI ?1049h) during tool calls and exits (CSI
|
|
* ?1049l) when returning to the main prompt. On re-attach, the replay
|
|
* prelude must re-enter alt-screen if the original PTY left it active,
|
|
* otherwise the replay renders against the main screen and the cursor
|
|
* + colors end up in the wrong place.
|
|
*/
|
|
altScreenActive: boolean;
|
|
/**
|
|
* Detach state machine (Commit 3). When the WS closes for a reason OTHER
|
|
* than the v1.44 intentional-restart code (4001), we keep the PtySession
|
|
* alive for the detach window (default 60s) so a re-attach within the
|
|
* window can resume the same PTY and replay the ring buffer. The timer
|
|
* disposes the session if no re-attach arrives in time.
|
|
*/
|
|
detached: boolean;
|
|
detachTimer: ReturnType<typeof setTimeout> | null;
|
|
}
|
|
|
|
/**
|
|
* WS keepalive interval. 25s is comfortably under the lowest common NAT
|
|
* idle timeout (typically 30-60s) and shorter than Chromium's WebSocket
|
|
* dead-peer threshold. Test-overridable via env so the v1.44 e2e tests
|
|
* can compress idle-window assertions to <1s without waiting half a
|
|
* minute per assertion.
|
|
*/
|
|
const KEEPALIVE_INTERVAL_MS = parseInt(
|
|
process.env.GSTACK_PTY_KEEPALIVE_INTERVAL_MS || '25000',
|
|
10,
|
|
);
|
|
|
|
/**
|
|
* Commit 3 scrollback ring buffer cap. 1 MB is enough for a full screen
|
|
* of dense claude output (including a recent tool result), small enough
|
|
* that a worst-case 10 detached sessions only cost ~10 MB of RSS.
|
|
* Env-overridable so e2e tests can verify eviction without writing 1 MB
|
|
* of fixture data per assertion.
|
|
*/
|
|
const RING_BUFFER_MAX_BYTES = parseInt(
|
|
process.env.GSTACK_PTY_RING_BUFFER_BYTES || `${1024 * 1024}`,
|
|
10,
|
|
);
|
|
|
|
/**
|
|
* Commit 3 detach window — how long to keep a session alive after WS
|
|
* close (with any code other than 4001 intentional-restart) so a
|
|
* re-attach can resume the same PTY. 60s is long enough to cover a
|
|
* Chrome MV3 service-worker suspend cycle, a wifi blip, or a brief
|
|
* laptop sleep; short enough that genuinely-closed sessions don't
|
|
* stack up unbounded.
|
|
*/
|
|
const DETACH_WINDOW_MS = parseInt(
|
|
process.env.GSTACK_PTY_DETACH_WINDOW_MS || '60000',
|
|
10,
|
|
);
|
|
|
|
/**
|
|
* Append a frame to a session's ring buffer, evicting oldest frames if
|
|
* the total byte count exceeds RING_BUFFER_MAX_BYTES. Eviction is at
|
|
* frame boundaries (one PTY write = one frame), so we never cut a
|
|
* multi-byte UTF-8 sequence or a partial ANSI CSI in half — claude's
|
|
* on-data callback emits coherent frames.
|
|
*
|
|
* Side effect: scans the appended chunk for alt-screen enter/exit
|
|
* sequences (CSI ?1049h / CSI ?1049l) and updates session.altScreenActive
|
|
* so the re-attach prelude knows whether to re-enter alt-screen.
|
|
*/
|
|
export function appendToRingBuffer(session: PtySession, frame: Buffer): void {
|
|
session.ringBuffer.push(frame);
|
|
session.ringBufferBytes += frame.length;
|
|
while (session.ringBufferBytes > RING_BUFFER_MAX_BYTES && session.ringBuffer.length > 1) {
|
|
const evicted = session.ringBuffer.shift()!;
|
|
session.ringBufferBytes -= evicted.length;
|
|
}
|
|
// Alt-screen tracking. Scan for the canonical xterm enter/exit pairs.
|
|
// We do this on every append (not just on attach) so the state is
|
|
// correct even if many frames have flowed since the last attach.
|
|
const ascii = frame.toString('latin1'); // single-byte view is enough — the codes are 7-bit ASCII
|
|
// Use lastIndexOf so trailing state wins when both appear in one frame
|
|
// (e.g., a quick tool-call open+close inside one render pass).
|
|
const enterIdx = ascii.lastIndexOf('\x1b[?1049h');
|
|
const exitIdx = ascii.lastIndexOf('\x1b[?1049l');
|
|
if (enterIdx >= 0 && enterIdx > exitIdx) session.altScreenActive = true;
|
|
else if (exitIdx >= 0 && exitIdx > enterIdx) session.altScreenActive = false;
|
|
}
|
|
|
|
/**
|
|
* Build the re-attach replay payload: server-side reset prelude + the
|
|
* accumulated ring buffer. The client side writes RIS (`\x1bc`) to xterm
|
|
* BEFORE feeding this payload in, so the layout is:
|
|
*
|
|
* 1. Client: `\x1bc` (RIS — full reset, clears pre-blip xterm content)
|
|
* 2. Server: `\x1b[!p` (DECSTR soft reset — re-defaults char attributes)
|
|
* 3. Server: optional `\x1b[?1049h` if we were in alt-screen at detach
|
|
* 4. Server: ring buffer contents, in append order
|
|
*
|
|
* The client coordinates the order by waiting for a `{type:"reattach-begin"}`
|
|
* text frame before treating the next binary frame as replay. That separation
|
|
* is what lets us prepend reset codes without clobbering the live stream
|
|
* that resumes immediately after.
|
|
*/
|
|
export function buildReplayPayload(session: PtySession): Buffer {
|
|
const parts: Buffer[] = [];
|
|
parts.push(Buffer.from('\x1b[!p'));
|
|
if (session.altScreenActive) parts.push(Buffer.from('\x1b[?1049h'));
|
|
for (const frame of session.ringBuffer) parts.push(frame);
|
|
return Buffer.concat(parts);
|
|
}
|
|
|
|
const sessions = new WeakMap<any, PtySession>(); // ws -> session
|
|
|
|
/** Find claude on PATH. */
|
|
function findClaude(): string | null {
|
|
// Test-only override. Lets the integration tests spawn /bin/bash instead
|
|
// of requiring claude to be installed on every CI runner. NEVER read in
|
|
// production (sidebar UI). Documented in browse/test/terminal-agent-integration.test.ts.
|
|
const override = process.env.BROWSE_TERMINAL_BINARY;
|
|
if (override && fs.existsSync(override)) return override;
|
|
// Bun.which is sync and respects PATH. Falls back to a small list of
|
|
// common install locations if PATH is stripped (e.g., launched from
|
|
// Conductor with a minimal env).
|
|
const which = (Bun as any).which?.('claude');
|
|
if (which) return which;
|
|
const candidates = [
|
|
'/opt/homebrew/bin/claude',
|
|
'/usr/local/bin/claude',
|
|
`${process.env.HOME}/.local/bin/claude`,
|
|
`${process.env.HOME}/.bun/bin/claude`,
|
|
`${process.env.HOME}/.npm-global/bin/claude`,
|
|
];
|
|
for (const c of candidates) {
|
|
try { fs.accessSync(c, fs.constants.X_OK); return c; } catch {}
|
|
}
|
|
return null;
|
|
}
|
|
|
|
/** Probe + persist claude availability for the bootstrap card. */
|
|
function writeClaudeAvailable(): void {
|
|
const stateDir = path.dirname(STATE_FILE);
|
|
try { mkdirSecure(stateDir); } catch {}
|
|
const found = findClaude();
|
|
const status = {
|
|
available: !!found,
|
|
path: found || undefined,
|
|
install_url: 'https://docs.anthropic.com/en/docs/claude-code',
|
|
checked_at: new Date().toISOString(),
|
|
};
|
|
const target = path.join(stateDir, 'claude-available.json');
|
|
// Fire-and-forget state file: a failed write must not break boot.
|
|
if (atomicWriteQuiet(target, JSON.stringify(status, null, 2), { mode: 0o600 })) {
|
|
restrictFilePermissions(target); // Windows ACL hardening
|
|
}
|
|
}
|
|
|
|
/**
|
|
* System-prompt hint passed to claude via --append-system-prompt. Tells
|
|
* claude what tab-awareness affordances exist in this session so it
|
|
* doesn't have to discover them by trial. The user can override anything
|
|
* here just by saying so — system prompt is a soft hint, not a contract.
|
|
*
|
|
* Two paths claude has:
|
|
* 1. Read live state from <stateDir>/tabs.json + active-tab.json
|
|
* (updated continuously by the gstack browser extension).
|
|
* 2. Run $B tab, $B tabs, $B tab-each <command> to act on tabs. The
|
|
* tab-each helper fans a single command across every open tab and
|
|
* returns per-tab results as JSON.
|
|
*/
|
|
function buildTabAwarenessHint(stateDir: string): string {
|
|
const tabsFile = path.join(stateDir, 'tabs.json');
|
|
const activeFile = path.join(stateDir, 'active-tab.json');
|
|
return [
|
|
'You are running inside the gstack browser sidebar with live access to the user\'s browser tabs.',
|
|
'',
|
|
'Tab state files (kept fresh automatically by the extension):',
|
|
` ${tabsFile} — all open tabs (id, url, title, active, pinned)`,
|
|
` ${activeFile} — the currently active tab`,
|
|
'Read these any time the user asks about "tabs", "the current page", or anything multi-tab. Do NOT shell out to $B tabs just to learn what\'s open — read the file.',
|
|
'',
|
|
'Tab manipulation commands (via $B):',
|
|
' $B tab <id> — switch to a tab',
|
|
' $B newtab [url] — open a new tab',
|
|
' $B closetab [id] — close a tab (current if no id)',
|
|
' $B tab-each <command> — fan out a command across every tab; returns JSON results',
|
|
'',
|
|
'When the user asks for multi-tab work, prefer $B tab-each. Examples:',
|
|
' $B tab-each snapshot -i — grab a snapshot from every tab',
|
|
' $B tab-each text — pull clean text from every tab',
|
|
' $B tab-each title — list every tab\'s title',
|
|
'',
|
|
'You\'re in a real terminal with a real PTY — slash commands, /resume, ANSI colors all work as in a normal claude session.',
|
|
].join('\n');
|
|
}
|
|
|
|
/** Spawn claude in a PTY. Returns null if claude not on PATH. */
|
|
function spawnClaude(cols: number, rows: number, onData: (chunk: Buffer) => void) {
|
|
const claudePath = findClaude();
|
|
if (!claudePath) return null;
|
|
|
|
// Match phoenix env so claude knows which browse server to talk to and
|
|
// doesn't try to autostart its own. BROWSE_HEADED=1 keeps the existing
|
|
// headed-mode browser; BROWSE_NO_AUTOSTART prevents claude's gstack
|
|
// tooling from racing to spawn another server.
|
|
const env: Record<string, string> = {
|
|
...process.env as any,
|
|
BROWSE_PORT: String(BROWSE_SERVER_PORT),
|
|
BROWSE_STATE_FILE: STATE_FILE,
|
|
BROWSE_NO_AUTOSTART: '1',
|
|
BROWSE_HEADED: '1',
|
|
TERM: 'xterm-256color',
|
|
COLORTERM: 'truecolor',
|
|
};
|
|
|
|
// --append-system-prompt is the right injection surface (per `claude --help`):
|
|
// it gets appended to the model's system prompt, so claude treats this as
|
|
// contextual guidance, not a user message. Don't use a leading PTY write
|
|
// for this — that would show up as if the user typed the hint, polluting
|
|
// the visible transcript.
|
|
const stateDir = path.dirname(STATE_FILE);
|
|
const tabHint = buildTabAwarenessHint(stateDir);
|
|
|
|
const proc = (Bun as any).spawn([claudePath, '--append-system-prompt', tabHint], {
|
|
windowsHide: true,
|
|
terminal: {
|
|
rows,
|
|
cols,
|
|
data(_terminal: any, chunk: Buffer) { onData(chunk); },
|
|
},
|
|
env,
|
|
});
|
|
return proc;
|
|
}
|
|
|
|
/** Cleanup a PTY session: SIGINT, then SIGKILL after 3s. */
|
|
function disposeSession(session: PtySession): void {
|
|
try { session.proc?.terminal?.close?.(); } catch {}
|
|
if (session.proc?.pid) {
|
|
try { session.proc.kill?.('SIGINT'); } catch {}
|
|
setTimeout(() => {
|
|
try {
|
|
if (session.proc && !session.proc.killed) session.proc.kill?.('SIGKILL');
|
|
} catch {}
|
|
}, 3000);
|
|
}
|
|
session.proc = null;
|
|
session.spawned = false;
|
|
}
|
|
|
|
/**
|
|
* Build the HTTP server. Two routes:
|
|
* POST /internal/grant — parent server pushes a fresh cookie token
|
|
* GET /ws — extension upgrades to WebSocket (PTY transport)
|
|
*
|
|
* Everything else returns 404. The listener binds 127.0.0.1 only.
|
|
*/
|
|
/**
|
|
* Validate a loopback /internal/* request. Returns null when the request
|
|
* is allowed; otherwise returns the Response to send back. Centralizes
|
|
* bearer auth + the v1.44 X-Browse-Gen generation check so adding a new
|
|
* /internal/* route is a one-liner.
|
|
*/
|
|
function checkInternalAuth(req: Request): Response | null {
|
|
const auth = req.headers.get('authorization');
|
|
if (auth !== `Bearer ${INTERNAL_TOKEN}`) {
|
|
return new Response('forbidden', { status: 403 });
|
|
}
|
|
const headerGen = req.headers.get('x-browse-gen');
|
|
if (headerGen && headerGen !== CURRENT_GEN) {
|
|
return new Response('stale generation', { status: 409 });
|
|
}
|
|
return null;
|
|
}
|
|
|
|
/**
|
|
* Wrap a JSON-bodied /internal/* handler with the standard bearer-auth +
|
|
* generation-check + json-parse + error-response boilerplate. The handler
|
|
* `fn` is called with the parsed body; whatever it returns is JSON-stringified
|
|
* into a 200 Response, or the handler can return a Response directly to
|
|
* customize status / headers. Throwing from `fn` collapses to a 400 "bad".
|
|
*
|
|
* Centralizing the dance kills the copy-paste pattern of bearer + gen check
|
|
* + req.json().then(...).catch(...) that every /internal/* route needs.
|
|
* New routes become a single call to internalHandler.
|
|
*/
|
|
async function internalHandler<T>(
|
|
req: Request,
|
|
fn: (body: any) => T | Promise<T> | Response | Promise<Response>,
|
|
): Promise<Response> {
|
|
const denied = checkInternalAuth(req);
|
|
if (denied) return denied;
|
|
let body: any;
|
|
try {
|
|
body = await req.json();
|
|
} catch {
|
|
return new Response('bad', { status: 400 });
|
|
}
|
|
try {
|
|
const result = await fn(body);
|
|
if (result instanceof Response) return result;
|
|
if (result === undefined || result === null) return new Response('ok');
|
|
return new Response(JSON.stringify(result), {
|
|
status: 200,
|
|
headers: { 'Content-Type': 'application/json' },
|
|
});
|
|
} catch {
|
|
return new Response('bad', { status: 400 });
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Spawn the claude PTY for a session if it hasn't been spawned yet.
|
|
* Used by both the legacy binary-frame spawn trigger and the v1.44 explicit
|
|
* `{type:"start"}` text-frame trigger. Idempotent on `session.spawned`.
|
|
*
|
|
* Returns true if claude is now running, false if spawn failed (e.g. claude
|
|
* binary not on PATH). On failure, the caller is expected to have already
|
|
* surfaced the error to the client (or will via the next frame).
|
|
*/
|
|
function maybeSpawnPty(ws: any, session: PtySession): boolean {
|
|
if (session.spawned) return true;
|
|
session.spawned = true;
|
|
let leftover = Buffer.alloc(0);
|
|
const proc = spawnClaude(session.cols, session.rows, (chunk) => {
|
|
const combined = Buffer.concat([leftover, Buffer.from(chunk)]);
|
|
// UTF-8 boundary detection (issue #1272). Look back at most 3 bytes
|
|
// for the start of an incomplete multibyte sequence and defer it.
|
|
let safeEnd = combined.length;
|
|
for (let i = combined.length - 1; i >= Math.max(0, combined.length - 3); i--) {
|
|
const b = combined[i];
|
|
if ((b & 0x80) === 0) { safeEnd = i + 1; break; }
|
|
if ((b & 0xC0) === 0x80) continue;
|
|
const expected = (b & 0xE0) === 0xC0 ? 2 : (b & 0xF0) === 0xE0 ? 3 : 4;
|
|
safeEnd = (combined.length - i >= expected) ? combined.length : i;
|
|
break;
|
|
}
|
|
const flush = combined.slice(0, safeEnd);
|
|
leftover = combined.slice(safeEnd);
|
|
if (flush.length) {
|
|
// Always record into the ring buffer (Commit 3) so re-attach can
|
|
// replay. session.liveWs is what changes across re-attaches — we
|
|
// close over `session`, not the original `ws`, so the write always
|
|
// goes to whichever ws is currently attached (or is skipped when
|
|
// detached and liveWs is null).
|
|
appendToRingBuffer(session, flush);
|
|
if (session.liveWs) {
|
|
try { session.liveWs.sendBinary(flush); } catch {}
|
|
}
|
|
}
|
|
});
|
|
if (!proc) {
|
|
try {
|
|
ws.send(JSON.stringify({
|
|
type: 'error',
|
|
code: 'CLAUDE_NOT_FOUND',
|
|
message: 'claude CLI not on PATH. Install: https://docs.anthropic.com/en/docs/claude-code',
|
|
}));
|
|
ws.close(4404, 'claude not found');
|
|
} catch {}
|
|
return false;
|
|
}
|
|
session.proc = proc;
|
|
proc.exited?.then?.(() => {
|
|
try { session.liveWs?.close(1000, 'pty exited'); } catch {}
|
|
});
|
|
return true;
|
|
}
|
|
|
|
function buildServer(port: number) {
|
|
return Bun.serve({
|
|
hostname: '127.0.0.1',
|
|
// #2314: allocated from the SAME fixed 10000-60000 scan range the main
|
|
// server uses (port-allocator.ts, decision 8) — never `port: 0`. Binding
|
|
// 0 drew from the OS EPHEMERAL range (49152-65535 on macOS), where this
|
|
// weeks-lived agent squatted ports that short-lived `app.listen(0)` test
|
|
// servers expected to receive, absorbing their traffic as phantom 404s.
|
|
port,
|
|
idleTimeout: 0, // PTY connections are long-lived; default idleTimeout would kill them
|
|
|
|
fetch(req, server) {
|
|
const url = new URL(req.url);
|
|
|
|
// /internal/grant — loopback-only handshake from parent server.
|
|
// v1.44+: accepts `{token, sessionId?}`. The sessionId binding lets
|
|
// the agent route re-attach attempts (same sessionId, fresh token)
|
|
// back to the same PtySession. Legacy callers passing just `{token}`
|
|
// still work — sessionId becomes null and re-attach is unavailable
|
|
// for that grant.
|
|
if (url.pathname === '/internal/grant' && req.method === 'POST') {
|
|
return internalHandler(req, (body) => {
|
|
if (typeof body?.token === 'string' && body.token.length > 16) {
|
|
const sid = typeof body?.sessionId === 'string' && body.sessionId.length > 0
|
|
? body.sessionId
|
|
: null;
|
|
validTokens.set(body.token, sid);
|
|
}
|
|
});
|
|
}
|
|
|
|
// /internal/revoke — drop a token (called on WS close or bootstrap reload)
|
|
if (url.pathname === '/internal/revoke' && req.method === 'POST') {
|
|
return internalHandler(req, (body) => {
|
|
if (typeof body?.token === 'string') validTokens.delete(body.token);
|
|
});
|
|
}
|
|
|
|
// /internal/restart — dispose the PtySession for a specific sessionId.
|
|
// Scoped to one caller (not enumerate-all). Server.ts /pty-restart
|
|
// posts here with the caller's sessionId; we kill ONLY that PTY,
|
|
// leaving any other live sidebar tabs untouched. Codex T2 of the
|
|
// eng review caught this gap — pre-spec the route would have
|
|
// disposed all sessions.
|
|
if (url.pathname === '/internal/restart' && req.method === 'POST') {
|
|
return internalHandler(req, (body) => {
|
|
const sid = typeof body?.sessionId === 'string' ? body.sessionId : null;
|
|
if (!sid) return { killed: 0 };
|
|
const session = sessionsById.get(sid);
|
|
if (!session) return { killed: 0 };
|
|
// Cancel any pending detach timer before disposal — otherwise it
|
|
// would fire later against an already-disposed session.
|
|
if (session.detachTimer) {
|
|
clearTimeout(session.detachTimer);
|
|
session.detachTimer = null;
|
|
}
|
|
disposeSession(session);
|
|
sessionsById.delete(sid);
|
|
return { killed: 1 };
|
|
});
|
|
}
|
|
|
|
// /internal/healthz — liveness probe used by the v1.44 watchdog.
|
|
// Returns this agent's pid + gen + active session count without
|
|
// touching claude binary lookup (which can fail for non-process
|
|
// reasons and isn't a useful liveness signal). GET — no body to parse,
|
|
// so it stays on the bare checkInternalAuth gate.
|
|
if (url.pathname === '/internal/healthz' && req.method === 'GET') {
|
|
const denied = checkInternalAuth(req);
|
|
if (denied) return denied;
|
|
return new Response(JSON.stringify({
|
|
pid: process.pid,
|
|
gen: CURRENT_GEN,
|
|
sessions: validTokens.size,
|
|
}), { status: 200, headers: { 'Content-Type': 'application/json' } });
|
|
}
|
|
|
|
// /claude-available — bootstrap card hits this when user clicks "I installed it".
|
|
if (url.pathname === '/claude-available' && req.method === 'GET') {
|
|
writeClaudeAvailable();
|
|
const found = findClaude();
|
|
return new Response(JSON.stringify({ available: !!found, path: found }), {
|
|
status: 200,
|
|
headers: { 'Content-Type': 'application/json' },
|
|
});
|
|
}
|
|
|
|
// /ws — WebSocket upgrade. CRITICAL gates:
|
|
// (1) Origin must be chrome-extension://<id>. Cross-site WS hijacking
|
|
// defense — required, not optional.
|
|
// (2) Token must be in validTokens. We accept the token via two
|
|
// transports for compatibility:
|
|
// - Sec-WebSocket-Protocol (preferred for browsers — the only
|
|
// auth header settable from the browser WebSocket API)
|
|
// - Cookie gstack_pty (works for non-browser callers and
|
|
// same-port browser callers; doesn't survive the cross-port
|
|
// jump from server.ts:34567 to the agent's random port
|
|
// when SameSite=Strict is set)
|
|
// Either path works; both verify against the same in-memory
|
|
// validTokens Set, populated by the parent server's
|
|
// authenticated /pty-session → /internal/grant chain.
|
|
if (url.pathname === '/ws') {
|
|
const origin = req.headers.get('origin') || '';
|
|
const isExtensionOrigin = origin.startsWith('chrome-extension://');
|
|
if (!isExtensionOrigin) {
|
|
return new Response('forbidden origin', { status: 403 });
|
|
}
|
|
if (EXTENSION_ID && origin !== `chrome-extension://${EXTENSION_ID}`) {
|
|
return new Response('forbidden origin', { status: 403 });
|
|
}
|
|
|
|
// Try Sec-WebSocket-Protocol first. Format: a single token, possibly
|
|
// with a `gstack-pty.` prefix (which we strip). Browsers send a
|
|
// comma-separated list when multiple were requested; we pick the
|
|
// first that matches a known token.
|
|
const protoHeader = req.headers.get('sec-websocket-protocol') || '';
|
|
let token: string | null = null;
|
|
for (const raw of protoHeader.split(',').map(s => s.trim()).filter(Boolean)) {
|
|
const candidate = raw.startsWith('gstack-pty.') ? raw.slice('gstack-pty.'.length) : raw;
|
|
if (validTokens.has(candidate)) {
|
|
token = candidate;
|
|
break;
|
|
}
|
|
}
|
|
|
|
// Fallback: Cookie gstack_pty (legacy / non-browser callers).
|
|
// Parsing is shared with the server via extractPtyCookie; VALIDATION
|
|
// deliberately stays against the agent's own validTokens map — the
|
|
// server's registry lives in a different process.
|
|
if (!token) {
|
|
const candidate = extractPtyCookie(req);
|
|
if (candidate && validTokens.has(candidate)) {
|
|
token = candidate;
|
|
}
|
|
}
|
|
|
|
if (!token) {
|
|
return new Response('unauthorized', { status: 401 });
|
|
}
|
|
|
|
// v1.44+: surface the token's sessionId binding to the upgraded ws.
|
|
// open() reads it via ws.data and registers the session in
|
|
// sessionsById so /internal/restart and (Commit 3) re-attach
|
|
// lookups can find it.
|
|
const sessionId = validTokens.get(token) ?? null;
|
|
// No explicit Sec-WebSocket-Protocol echo: Bun >= 1.3 auto-echoes the
|
|
// first offered protocol in the 101 response, so setting the header
|
|
// here produced a DUPLICATE header — strict clients (Chromium, python
|
|
// websockets) reject the handshake per RFC 6455 and the sidebar
|
|
// terminal could never connect. Verified on Bun 1.3.6.
|
|
const upgraded = server.upgrade(req, {
|
|
data: { cookie: token, sessionId },
|
|
});
|
|
return upgraded ? undefined : new Response('upgrade failed', { status: 500 });
|
|
}
|
|
|
|
return new Response('not found', { status: 404 });
|
|
},
|
|
|
|
websocket: {
|
|
/**
|
|
* Spawn the claude PTY for `session` if it hasn't been spawned yet.
|
|
* Called from both message paths: the legacy binary-frame trigger
|
|
* (any keystroke) AND the v1.44 explicit `{type:"start"}` trigger
|
|
* (forceRestart sends this on every fresh WS to get an eager prompt
|
|
* without requiring the user to type). Idempotent — a second call
|
|
* after `spawned: true` is a no-op.
|
|
*/
|
|
open(ws) {
|
|
const sessionId = (ws.data as any)?.sessionId ?? null;
|
|
const cookie = (ws.data as any)?.cookie || '';
|
|
|
|
// Commit 3 re-attach: if this sessionId already has a detached
|
|
// PtySession in sessionsById, REPLACE its liveWs ref and replay
|
|
// the ring buffer. The PTY process is unchanged — claude keeps
|
|
// running through the wifi blip / panel-suspend cycle.
|
|
if (sessionId) {
|
|
const existing = sessionsById.get(sessionId);
|
|
if (existing) {
|
|
if (existing.detachTimer) {
|
|
clearTimeout(existing.detachTimer);
|
|
existing.detachTimer = null;
|
|
}
|
|
existing.detached = false;
|
|
existing.liveWs = ws;
|
|
existing.cookie = cookie;
|
|
// Re-bind the WS-keyed map so resize/close/message handlers
|
|
// can still find this session via the new ws.
|
|
sessions.set(ws, existing);
|
|
// Restart keepalive on the new ws.
|
|
if (existing.pingInterval) clearInterval(existing.pingInterval);
|
|
existing.pingInterval = setInterval(() => {
|
|
try { ws.send(JSON.stringify({ type: 'ping', ts: Date.now() })); } catch {}
|
|
}, KEEPALIVE_INTERVAL_MS);
|
|
// Tell the client to prep its xterm (write RIS) before the
|
|
// replay binary arrives. Order matters — the binary frame
|
|
// immediately after this text frame IS the replay.
|
|
try { ws.send(JSON.stringify({ type: 'reattach-begin', sessionId })); } catch {}
|
|
try { ws.sendBinary(buildReplayPayload(existing)); } catch {}
|
|
return;
|
|
}
|
|
}
|
|
|
|
const session: PtySession = {
|
|
proc: null,
|
|
cols: 80,
|
|
rows: 24,
|
|
cookie,
|
|
liveWs: ws,
|
|
sessionId,
|
|
spawned: false,
|
|
pingInterval: null,
|
|
ringBuffer: [],
|
|
ringBufferBytes: 0,
|
|
altScreenActive: false,
|
|
detached: false,
|
|
detachTimer: null,
|
|
};
|
|
session.pingInterval = setInterval(() => {
|
|
try {
|
|
ws.send(JSON.stringify({ type: 'ping', ts: Date.now() }));
|
|
} catch {
|
|
// ws likely closed mid-tick; close handler clears the interval.
|
|
}
|
|
}, KEEPALIVE_INTERVAL_MS);
|
|
sessions.set(ws, session);
|
|
// Index by sessionId for /internal/restart + Commit 3 re-attach.
|
|
if (sessionId) sessionsById.set(sessionId, session);
|
|
},
|
|
|
|
message(ws, raw) {
|
|
let session = sessions.get(ws);
|
|
if (!session) {
|
|
// Fallback for any path where open() didn't fire (shouldn't happen
|
|
// in Bun.serve but keeps the spawn path safe). No keepalive on
|
|
// this branch — open() is the supported entry point.
|
|
session = {
|
|
proc: null,
|
|
cols: 80,
|
|
rows: 24,
|
|
cookie: (ws.data as any)?.cookie || '',
|
|
liveWs: ws,
|
|
sessionId: (ws.data as any)?.sessionId ?? null,
|
|
spawned: false,
|
|
pingInterval: null,
|
|
ringBuffer: [],
|
|
ringBufferBytes: 0,
|
|
altScreenActive: false,
|
|
detached: false,
|
|
detachTimer: null,
|
|
};
|
|
sessions.set(ws, session);
|
|
if (session.sessionId) sessionsById.set(session.sessionId, session);
|
|
}
|
|
|
|
// Text frames are control messages: {type: "resize", cols, rows},
|
|
// {type: "tabSwitch", tabId, url, title}, {type: "tabState", ...},
|
|
// or v1.44 keepalive frames: {type: "pong", ts}, {type: "keepalive"}.
|
|
// Binary frames are raw input bytes destined for the PTY stdin.
|
|
if (typeof raw === 'string') {
|
|
let msg: any;
|
|
try { msg = JSON.parse(raw); } catch { return; }
|
|
if (msg?.type === 'resize') {
|
|
const cols = Math.max(2, Math.floor(Number(msg.cols) || 80));
|
|
const rows = Math.max(2, Math.floor(Number(msg.rows) || 24));
|
|
session.cols = cols;
|
|
session.rows = rows;
|
|
try { session.proc?.terminal?.resize?.(cols, rows); } catch {}
|
|
return;
|
|
}
|
|
if (msg?.type === 'tabSwitch') {
|
|
handleTabSwitch(msg);
|
|
return;
|
|
}
|
|
if (msg?.type === 'tabState') {
|
|
handleTabState(msg);
|
|
return;
|
|
}
|
|
if (msg?.type === 'pong' || msg?.type === 'keepalive' || msg?.type === 'ping') {
|
|
// Keepalive frames — accepted and silently dropped. The mere
|
|
// fact that the WS carried this frame is the liveness signal;
|
|
// there's no application-level state to update at this layer.
|
|
// `ping` is acknowledged here too in case the client (or a
|
|
// future agent peer) mirrors our server-side ping shape.
|
|
return;
|
|
}
|
|
if (msg?.type === 'start') {
|
|
// v1.44 explicit spawn trigger. forceRestart sends this
|
|
// immediately on every fresh WS so claude boots without the
|
|
// user having to type a keystroke (pre-v1.44, the lazy-binary
|
|
// spawn made restart look stuck until the user typed). No-op
|
|
// if already spawned.
|
|
maybeSpawnPty(ws, session);
|
|
return;
|
|
}
|
|
// Unknown text frame — ignore.
|
|
return;
|
|
}
|
|
|
|
// Binary input. Lazy-spawn claude on the first byte if `start`
|
|
// wasn't sent first. Both paths land in the same maybeSpawnPty
|
|
// helper for behavior parity.
|
|
if (!session.spawned) {
|
|
if (!maybeSpawnPty(ws, session)) return;
|
|
}
|
|
try {
|
|
// raw is a Uint8Array; Bun.Terminal.write accepts string|Buffer.
|
|
// Convert to Buffer for safety.
|
|
session.proc?.terminal?.write?.(Buffer.from(raw as Uint8Array));
|
|
} catch (err) {
|
|
console.error('[terminal-agent] terminal.write failed:', err);
|
|
}
|
|
},
|
|
|
|
close(ws, code, _reason) {
|
|
const session = sessions.get(ws);
|
|
if (!session) return;
|
|
// Always drop the WS-keyed map entry and the per-attach
|
|
// attachToken — the attach grant was single-use.
|
|
sessions.delete(ws);
|
|
if (session.cookie) validTokens.delete(session.cookie);
|
|
// Keepalive lives with the WS — every attach starts a fresh one.
|
|
if (session.pingInterval) {
|
|
clearInterval(session.pingInterval);
|
|
session.pingInterval = null;
|
|
}
|
|
|
|
// Commit 3 detach state machine. If the close was intentional
|
|
// (code 4001 = restart, 4404 = no-claude error), dispose
|
|
// immediately — there's no value in keeping the PTY alive.
|
|
// Otherwise enter the detach window: claude keeps running, the
|
|
// ring buffer keeps accumulating, and a re-attach with the same
|
|
// sessionId within DETACH_WINDOW_MS picks back up. If the timer
|
|
// fires without a re-attach, the session is disposed normally.
|
|
//
|
|
// Sessions without a sessionId (legacy single-shot grants) can't
|
|
// re-attach by definition — fall through to immediate dispose.
|
|
const intentional = code === 4001 || code === 4404 || code === 1000;
|
|
if (intentional || !session.sessionId) {
|
|
disposeSession(session);
|
|
if (session.sessionId) sessionsById.delete(session.sessionId);
|
|
return;
|
|
}
|
|
|
|
// Mark detached and start the disposal timer. The session stays
|
|
// in sessionsById so the next /ws upgrade with the same
|
|
// sessionId can find and reattach to it.
|
|
session.detached = true;
|
|
session.liveWs = null;
|
|
session.detachTimer = setTimeout(() => {
|
|
if (!session.detached) return; // re-attached in the meantime
|
|
disposeSession(session);
|
|
if (session.sessionId) sessionsById.delete(session.sessionId);
|
|
}, DETACH_WINDOW_MS);
|
|
// setTimeout returns a Bun Timer; unref so the detach window
|
|
// doesn't keep the process alive past natural shutdown.
|
|
(session.detachTimer as any)?.unref?.();
|
|
},
|
|
},
|
|
});
|
|
}
|
|
|
|
/**
|
|
* Tab-switch helper: write the active tab to a state file (claude reads it)
|
|
* and notify the parent server so its activeTabId stays synced. Skips
|
|
* chrome:// and chrome-extension:// internal pages.
|
|
*/
|
|
/**
|
|
* Live tab snapshot. Writes <stateDir>/tabs.json (full list) and updates
|
|
* <stateDir>/active-tab.json (current active). claude can read these any
|
|
* time without invoking $B tabs — saves a round-trip when the model just
|
|
* needs to check the landscape before deciding what to do.
|
|
*/
|
|
function handleTabState(msg: {
|
|
active?: { tabId?: number; url?: string; title?: string } | null;
|
|
tabs?: Array<{ tabId?: number; url?: string; title?: string; active?: boolean; windowId?: number; pinned?: boolean; audible?: boolean }>;
|
|
reason?: string;
|
|
}): void {
|
|
const stateDir = path.dirname(STATE_FILE);
|
|
try { mkdirSecure(stateDir); } catch {}
|
|
|
|
// tabs.json — full list
|
|
if (Array.isArray(msg.tabs)) {
|
|
const payload = {
|
|
updatedAt: new Date().toISOString(),
|
|
reason: msg.reason || 'unknown',
|
|
tabs: msg.tabs.map(t => ({
|
|
tabId: t.tabId ?? null,
|
|
url: t.url || '',
|
|
title: t.title || '',
|
|
active: !!t.active,
|
|
windowId: t.windowId ?? null,
|
|
pinned: !!t.pinned,
|
|
audible: !!t.audible,
|
|
})),
|
|
};
|
|
const target = path.join(stateDir, 'tabs.json');
|
|
// Fire-and-forget state file: atomic write (via lib/fs-atomic) so
|
|
// claude never reads a half-written JSON document; failures swallowed.
|
|
if (atomicWriteQuiet(target, JSON.stringify(payload, null, 2), { mode: 0o600 })) {
|
|
restrictFilePermissions(target); // Windows ACL hardening
|
|
}
|
|
}
|
|
|
|
// active-tab.json — single active tab. Skip chrome-internal pages so
|
|
// claude doesn't see chrome:// or chrome-extension:// URLs as
|
|
// "current target."
|
|
const active = msg.active;
|
|
if (active && active.url && !active.url.startsWith('chrome://') && !active.url.startsWith('chrome-extension://')) {
|
|
const ctxFile = path.join(stateDir, 'active-tab.json');
|
|
const ok = atomicWriteQuiet(ctxFile, JSON.stringify({
|
|
tabId: active.tabId ?? null,
|
|
url: active.url,
|
|
title: active.title ?? '',
|
|
}), { mode: 0o600 });
|
|
if (ok) restrictFilePermissions(ctxFile); // Windows ACL hardening
|
|
}
|
|
}
|
|
|
|
function handleTabSwitch(msg: { tabId?: number; url?: string; title?: string }): void {
|
|
const url = msg.url || '';
|
|
if (!url || url.startsWith('chrome://') || url.startsWith('chrome-extension://')) return;
|
|
|
|
const stateDir = path.dirname(STATE_FILE);
|
|
const ctxFile = path.join(stateDir, 'active-tab.json');
|
|
// Fire-and-forget: atomic write via lib/fs-atomic, failures swallowed.
|
|
const ok = atomicWriteQuiet(ctxFile, JSON.stringify({
|
|
tabId: msg.tabId ?? null,
|
|
url,
|
|
title: msg.title ?? '',
|
|
}), { mode: 0o600 });
|
|
if (ok) restrictFilePermissions(ctxFile); // Windows ACL hardening
|
|
|
|
// Best-effort sync to parent server so its activeTabId tracking matches.
|
|
// No await; this is fire-and-forget.
|
|
if (BROWSE_SERVER_PORT > 0) {
|
|
fetch(`http://127.0.0.1:${BROWSE_SERVER_PORT}/command`, {
|
|
method: 'POST',
|
|
headers: {
|
|
'Content-Type': 'application/json',
|
|
'Authorization': `Bearer ${readBrowseToken()}`,
|
|
},
|
|
body: JSON.stringify({
|
|
command: 'tab',
|
|
args: [String(msg.tabId ?? ''), '--no-focus'],
|
|
}),
|
|
}).catch(() => {});
|
|
}
|
|
}
|
|
|
|
function readBrowseToken(): string {
|
|
try {
|
|
const raw = fs.readFileSync(STATE_FILE, 'utf-8');
|
|
const j = JSON.parse(raw);
|
|
return j.token || '';
|
|
} catch { return ''; }
|
|
}
|
|
|
|
// Boot.
|
|
async function main() {
|
|
writeClaudeAvailable();
|
|
// #2314: allocate from the shared fixed scan range, then bind. Probe-then-
|
|
// bind has a TOCTOU window — a concurrent process can take the port between
|
|
// the probe and Bun.serve, which throws and would kill the boot with no
|
|
// retry (main().catch → exit 1). Re-allocate and retry a few times; each
|
|
// iteration probes fresh, so only a genuine race lands here.
|
|
let server: ReturnType<typeof buildServer> | undefined;
|
|
let lastBindErr: unknown;
|
|
for (let attempt = 0; attempt < 5 && !server; attempt++) {
|
|
const allocatedPort = await findAvailablePort();
|
|
try {
|
|
server = buildServer(allocatedPort);
|
|
} catch (err) {
|
|
lastBindErr = err;
|
|
}
|
|
}
|
|
if (!server) {
|
|
console.error(`[terminal-agent] failed to bind after 5 attempts: ${lastBindErr}`);
|
|
process.exit(1);
|
|
}
|
|
const port = (server as any).port || (server as any).address?.port;
|
|
if (!port) {
|
|
console.error('[terminal-agent] failed to bind: no port');
|
|
process.exit(1);
|
|
}
|
|
|
|
// Write port file atomically so the parent server can pick it up.
|
|
// Throws on failure — a boot without a discoverable port file is broken.
|
|
const dir = path.dirname(PORT_FILE);
|
|
try { mkdirSecure(dir); } catch {}
|
|
atomicWriteSync(PORT_FILE, String(port), { mode: 0o600 });
|
|
restrictFilePermissions(PORT_FILE); // Windows ACL hardening
|
|
|
|
// Write identity-based agent record (pid + per-boot gen). Replaces the
|
|
// v1.43- `pkill -f terminal-agent\.ts` regex teardown that could kill
|
|
// sibling gstack sessions. Callers (cli.ts spawn site, server.ts
|
|
// shutdown, the v1.44 watchdog) now route through killAgentByRecord in
|
|
// terminal-agent-control.ts.
|
|
writeAgentRecord(dir, { pid: process.pid, gen: CURRENT_GEN, startedAt: Date.now() });
|
|
|
|
// Hand the parent the internal token so it can call /internal/grant.
|
|
// Parent learns INTERNAL_TOKEN via env (TERMINAL_AGENT_INTERNAL_TOKEN below).
|
|
// We just print it on stdout for the supervising process to pick up if it's
|
|
// not already in env. Defense against env races at spawn time.
|
|
console.log(`[terminal-agent] listening on 127.0.0.1:${port} pid=${process.pid} gen=${CURRENT_GEN}`);
|
|
|
|
// Cleanup port file + agent record on exit.
|
|
let cleaningUp = false;
|
|
const cleanup = () => {
|
|
if (cleaningUp) return;
|
|
cleaningUp = true;
|
|
safeUnlink(PORT_FILE);
|
|
safeUnlink(INTERNAL_TOKEN_FILE);
|
|
clearAgentRecord(dir);
|
|
process.exit(0);
|
|
};
|
|
process.on('SIGTERM', cleanup);
|
|
process.on('SIGINT', cleanup);
|
|
|
|
// The terminal agent is intentionally detached so it survives the short-lived
|
|
// CLI launcher, but its real owner is the persistent browse server. If that
|
|
// server crashes or is killed before running normal shutdown, the agent would
|
|
// otherwise be adopted by PID 1 and live forever. Poll the server PID and use
|
|
// the same cleanup path as an intentional shutdown when it disappears.
|
|
if (BROWSE_OWNER_PID > 0) {
|
|
const ownerWatchdog = setInterval(() => {
|
|
try {
|
|
process.kill(BROWSE_OWNER_PID, 0);
|
|
} catch {
|
|
cleanup();
|
|
}
|
|
}, OWNER_WATCHDOG_MS);
|
|
(ownerWatchdog as any)?.unref?.();
|
|
}
|
|
}
|
|
|
|
// Export the internal token so cli.ts can pass the SAME value to the parent
|
|
// server via env. Parent reads BROWSE_TERMINAL_INTERNAL_TOKEN and uses it
|
|
// for /internal/grant calls.
|
|
//
|
|
// In practice, the agent generates INTERNAL_TOKEN once at boot and writes it
|
|
// to a state file the parent reads. This avoids env-passing races. See main().
|
|
const INTERNAL_TOKEN_FILE = path.join(path.dirname(STATE_FILE), 'terminal-internal-token');
|
|
try {
|
|
mkdirSecure(path.dirname(INTERNAL_TOKEN_FILE));
|
|
writeSecureFile(INTERNAL_TOKEN_FILE, INTERNAL_TOKEN);
|
|
} catch {}
|
|
|
|
main().catch((err) => {
|
|
console.error(`[terminal-agent] boot failed: ${err instanceof Error ? err.message : String(err)}`);
|
|
process.exit(1);
|
|
});
|