mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 22:48:57 +02:00
* feat(aside): browser-driver contract, cookbook, research and fallback resolvers
{{ASIDE_SETUP}} (readiness probe + ten rules for driving the user's real browser), {{ASIDE_COOKBOOK}} (script shapes verified live against Aside CLI 1.26: one flow per aside repl script, CDP console hook before navigation, evidence lines, session-directory artifact handoff, GSTACK_STEP_OK sentinel), {{ASIDE_RESEARCH}} (research through aside exec, WebSearch when Aside is absent, knowledge otherwise) and {{BROWSE_FALLBACK}} (the fifteen-row Aside-step to $B-command table plus the rules that differ, so every browsing skill keeps working on gstack's own headless browser). test/aside-driver.test.ts pins the sentences and asserts every browsing skill carries the Aside block followed by the fallback; test/helpers/aside-available.ts is the shared live-Aside probe.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(render): Aside-first local-HTML renderer with the bundled browser as fallback
lib/aside-render.ts serves the HTML's directory on loopback (Aside refuses file:// URLs), opens it with waitUntil load, prints through CDP Page.printToPDF so tagged output, outlines, header/footer templates and page numbers survive, emulates device metrics for sized screenshots, and writes in-page evaluations to files; when Aside is absent it runs the same spec through the browse daemon (newtab, load, js, pdf, screenshot, closetab) and reports ENGINE=aside|browse. bin/gstack-render.ts is the CLI skill templates call. lib/claude-bin.ts and lib/error-handling.ts become the canonical copies (browse/src re-exports them).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(browse): /browse drives Aside first, with the $B reference behind the fallback
Contract, cookbook, mode choice (aside repl by default, aside exec for reading), report format, the fallback section, and the full command reference carved on demand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(qa): /qa and /qa-only drive Aside, fall back to $B
QA_METHODOLOGY runs every phase as Aside scripts (orient, explore, document, re-test, mobile viewport via CDP emulation, links via HEAD fetch); the authenticate phase is 'you are already signed in'; a 13th rule requires consent before mutating actions on non-local targets; the fallback section translates each step onto $B. The qa E2E tests run on whichever engine is present.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(design): design-review, design-consultation, design-shotgun, plan-design-review, design-html drive Aside
Design-system extraction is one script printing FONTS/COLORS/HEADINGS/TOUCH_TARGETS/NAV; competitor research confirms the exact URLs before opening them in the real browser and runs on the bundled browser when Aside is absent; design-html's viewport screenshots, sketches and comparison boards render through gstack-render.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(deploy): benchmark, canary, land-and-deploy Step 7, devex-review drive Aside
One aside repl script per page prints NAV/PAINT/LCP/RESOURCES/SCRIPTS/CSS/SUMMARY (benchmark), CONSOLE_ERRORS/NAV/TEXT + screenshot (canary, re-run every 60s), and the post-deploy check reads responseStatus from the navigation entry; each carries the $B fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(third-party-actions): Aside is the recommended driver; gstack's visible browser stays the fallback
The readiness probe is lifted from {{ASIDE_SETUP}} at gen time (byte-identity pinned) and rule 3 points at browse/SKILL.md for how to drive; the consent question offers Aside first and gstack's own visible browser (handoff/resume for sign-in) as the fallback, as v1.72 framed it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(scrape): /scrape reads pages through Aside; the browser-skills runtime rides the fallback
Look-then-extract scripts build the JSON inside the page and print it between JSON_START/JSON_END; aside exec for fuzzy intents; on the $B fallback the browser-skills match/prototype flow and /skillify apply as before.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(make-pdf): print through Aside first, the bundled browser otherwise
asideClient.ts replaces the direct $B client with one render() call per PDF (the exact option mapping the browse pdf command had: paper, margins, header/footer/page numbers, tagged, outline, printBackground, preferCSSPageSize, Paged.js wait); the diagram pre-pass, oversized-image downscale and DOCX rasters each run as one render script with per-fence try/catch; exit 4 now means no browser is available and names both remedies; $P setup reports which engine it found. The e2e gates run on whichever engine is present, so the Linux lane exercises the fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(diagram): the triplet is one gstack-render call
SVG, PNG and excalidraw from one invocation over the content-addressed bundle staged under /tmp/gstack-render; every diagram type gets an excalidraw export; gstack-render picks the engine and prints ENGINE=; the diagram E2E gates on either engine.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(research): web research runs in Aside first, WebSearch second
The planning, review, design, security and investigate skills research through {{ASIDE_RESEARCH}}; WebSearch stays in allowed-tools as the fallback; testing.ts's bootstrap step follows; skeleton ceilings ratcheted for the research block.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(setup,gen-skill-docs): prune renders of skills that no longer exist
setup gains _prune_stale_generated for every host tree and the doc generator removes gstack-* output dirs it did not write, so a skill removed from the source tree can never linger in an install.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: registries, budgets and suite reconciled for Aside-first with the $B fallback
Touchfiles + E2E tiers gain the Aside keys, coverage matrix and eval baselines updated, size budget re-baselined to parity-baseline-v1.80.0.0.json (the contract plus fallback ride in every browsing skill), parity ceilings ratcheted with measured values, LLM-judge prompts and the E2E fixtures speak Aside-first, browse-fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: Aside first, gstack browser fallback
README, BROWSER.md, docs/, CONTRIBUTING, CLAUDE.md, ARCHITECTURE, AGENTS.md, TODOS and the root router describe the one product story: Aside is the browser gstack drives first; the bundled headless browser is the automatic fallback (Linux, Windows, app closed) where cookie import, GStack Browser, pair-agent and browser-skills still apply.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore: regenerate SKILL.md docs, llms.txt, agents digest, ship goldens, context-budget fixture
bun run gen:skill-docs over the templates; goldens re-rendered; context-budget ceilings recaptured.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* v1.80.0.0: Aside is the browser gstack drives first; the bundled browser is the fallback
MINOR: new capability across ten skills, the renderer and research; nothing removed. CHANGELOG release summary + itemized changes; VERSION 1.80.0.0; package.json 1.80.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(todos): file non-Claude host ownership-gate and version-heading pin follow-ups
Two follow-ups from the /plan-ceo-review + /plan-eng-review pass on merging
PR #2804 with main's v1.80.0.0 ownership gate: bring the Codex/Factory/
OpenCode/Cursor/Kiro copy loops and the stale-render prune under the
.gstack-owned marker rule, and a free test pinning that the CHANGELOG top
heading equals VERSION (the collision that git cannot see).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: pre-landing review fixes for the Aside-first branch
Review army + adversarial passes (Claude and Codex) on the merged branch:
setup
- _prune_stale_generated scans the host dirs too (the generator already
removed the render before setup ran, so the host branch was dead), skips
symlinks in the render tree (rm -rf on a slash-terminated link empties its
target), removes a host symlink only when it resolves into gstack, cleans a
bannered real dir through _cleanup_weak_dir, recognizes frontmatter-renamed
skills, and logs through log. The always-run codex render passes every host
dir that may link to it.
- NEEDS_BUILD checks all three binaries (with $_EXE) and lib/ sources; the
browser hint and the bootstrap summary honor GSTACK_SKIP_ASIDE, treat a
requested skip as a request, and derive one skill list.
lib/aside-render.ts + bin/gstack-render.ts
- The loopback server carries a per-render secret path, checks containment on
the real path (symlink escapes are 403), and rejects malformed encoding.
- Inline eval results are one base64 line, so page text cannot forge
ASIDE_DIR= or the sentinel; the last ASIDE_DIR wins.
- runProc escalates SIGTERM to SIGKILL, bounds every wait, and clears every
timer (an uncleared one kept gstack-render alive after printing OK).
- renderTmpDir refuses a shared /tmp name owned by someone else; the work dir
and server are created inside try; goto's budget follows the render budget.
- probeAside classifies a present-but-failing CLI as ASIDE_NOT_RUNNING like
the skills' bash probe; render() retries on gstack's own browser when Aside
could not start or its private CDP bridge is gone (never on a page error
or a timeout of a running script); the CLI reports the engine that actually
rendered, exits 0 on --help, rejects non-numeric flags, documents
--wait-timeout, fences EVAL/PAGE_ERRORS as untrusted content, and names the
daemon's cookie-import JS lock remedy.
- The browse path passes --scale only when asked (a scale change rebuilds
the daemon context) and restores the viewport after a sized screenshot.
resolvers / templates
- The bash probe honors GSTACK_SKIP_ASIDE and has a perl deadline on stock
macOS; .local is no longer LOCAL (mDNS); same-origin filters compare parsed
origins; link status is HEAD-checked only on LOCAL targets; every
aside exec goes through the receipted _aside_exec prelude
({{ASIDE_EXEC_PRELUDE}}), including nine template blocks that called it
bare; the design sketch and diagram staging use private directories.
- The generator prunes only bannered renders and never a host whose
generation failed.
Docs, stale comments and dead code cleaned; goldens re-rendered; tests
updated and added for every behavior above.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: coverage for the render CLI, setup rebuild check, make-pdf exit codes, and prose $B spans
New free tests from the ship coverage audit: test/gstack-render-cli.test.ts
(argv guards, --help, output contract with a fake daemon, failure and
serve-root paths, no-browser case, prompt exit), test/setup-needs-build.test.ts
(every binary and source set flips NEEDS_BUILD, Windows suffixes),
make-pdf/test/cli-exit-codes.test.ts and setup-smoke.test.ts (error to exit
code mapping, runSetup stages, renderPdf's engine), and prose-span cases for
extractBrowseCommands in test/skill-parser.test.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG and TODOS cover the review fixes (v1.81.0.0)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: sync project docs with the v1.81.0.0 review fixes
BROWSER.md, ARCHITECTURE.md, CONTRIBUTING.md, README.md, CLAUDE.md,
docs/TESTING_INTERNALS.md and docs/PROJECT_STRUCTURE.md now describe the
shipped renderer and setup: the loopback render server's per-render secret
path and real-path containment, ENGINE= naming the engine that actually
rendered (mid-run retry on gstack's own browser), EVAL/PAGE_ERRORS fenced as
untrusted content, --wait-timeout and the CLI's argv guards, the receipted
_aside_exec prelude ({{ASIDE_EXEC_PRELUDE}} in the placeholder table), the
LOCAL host rule without .local, LOCAL-only HEAD checks in the links script,
GSTACK_SKIP_ASIDE across probe/renderer/setup, the ownership-gated
retired-skill prune, the widened NEEDS_BUILD check, and the new free tests
(gstack-render-cli, setup-prune-stale-generated, setup-browser-hint,
setup-needs-build, make-pdf cli-exit-codes and setup-smoke).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG states the precise mid-run retry rule
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): skill-e2e-bws slices the $B setup block from the Browser fallback section
browse/SKILL.md no longer has '## SETUP' / '## Core QA Patterns' (Aside is the
primary driver; the $B block moved under 'Browser fallback'), so the gate test
sliced an empty block and handed the agent nothing to run. Anchor on
'### Find the `$B` binary' up to the next heading. 7/7 pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): gate POSIX-only fixtures off Windows
windows-free-tests: the gstack-render CLI tests drive a shebang fake browse
that CreateProcess cannot exec, and two NEEDS_BUILD cases assert an execute
bit and a bare-name miss that MSYS bash does not have (test -x ignores mode
bits and resolves design -> design.exe). Those describes and cases now
self-skip on win32; argument guards, --help, the no-browser case, and every
other rebuild-check case still run there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(render): runProc waits for the exit code until the kill deadline; newtab retries once on a cold daemon
A process whose pipes have reached EOF is exiting, but runProc gave the exit
code only five seconds to arrive and then returned null, which run() reports
as a failed command. Under CI's six-shard load one such render failed with the
artifact already written. The SIGTERM/SIGKILL timers already bound the wait,
so the exit race now runs to the kill deadline.
The first CLI call auto-starts the browse daemon; on a cold start it can
answer 'Unable to connect' once while the server is still coming up. That
single case is retried after 1.5s; every other newtab failure is not.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(aside-render): warm the daemon before live fallback cases; failures name the render error
- Live fallback cases run 'goto about:blank' up to twice before asserting and
skip (never fail) when the daemon cannot come up.
- expectOk() puts r.error and the browse transcript into the assertion so a
failed render is diagnosable from the CI log.
- The argv-contract cases dump the fake's log on a miss.
- File default timeout is 30s: the subject is the CLI contract, not latency.
- Two cases pin the cold-daemon newtab retry and that other errors are not
retried.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG notes the cold-start tolerance of the bundled-browser renderer
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Sina <sdroid674+github@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
184 lines
8.6 KiB
Markdown
184 lines
8.6 KiB
Markdown
# gstack x OpenClaw Integration
|
|
|
|
gstack integrates with OpenClaw as a methodology source, not a ported codebase.
|
|
OpenClaw's ACP runtime spawns Claude Code sessions natively. gstack provides the
|
|
planning discipline and methodology that makes those sessions better.
|
|
|
|
This is a lightweight protocol encoded as prompt text. No daemon. No JSON-RPC.
|
|
No compatibility matrices. The prompt is the bridge.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
OpenClaw gstack repo
|
|
───────────────────── ──────────────
|
|
Orchestrator: messaging, Source of truth for
|
|
calendar, memory, EA methodology + planning
|
|
│ │
|
|
├── Native skills (conversational) ├── Generates native skills
|
|
│ office-hours, ceo-review, │ via gen-skill-docs pipeline
|
|
│ investigate, retro │
|
|
│ ├── Generates gstack-lite
|
|
├── sessions_spawn(runtime: "acp") │ (planning discipline)
|
|
│ │ │
|
|
│ └── Claude Code ├── Generates gstack-full
|
|
│ └── gstack installed at │ (complete pipeline)
|
|
│ ~/.claude/skills/gstack │
|
|
│ └── docs/OPENCLAW.md (this file)
|
|
└── Dispatch routing (AGENTS.md)
|
|
```
|
|
|
|
## Dispatch Routing
|
|
|
|
OpenClaw decides at spawn time which tier of gstack support to use:
|
|
|
|
| Tier | When | Prompt prefix |
|
|
|------|------|---------------|
|
|
| **Simple** | One-file edits, typos, config changes | No gstack context injected |
|
|
| **Medium** | Multi-file features, refactors | gstack-lite CLAUDE.md appended |
|
|
| **Heavy** | Specific gstack skill needed | "Load gstack. Run /X" |
|
|
| **Full** | Complete features, objectives, projects | gstack-full pipeline appended |
|
|
| **Plan** | "Help me plan a Claude Code project" | gstack-plan pipeline appended |
|
|
|
|
### Decision heuristic
|
|
|
|
- Can it be done in <10 lines of code? -> **Simple**
|
|
- Does it touch multiple files but the approach is obvious? -> **Medium**
|
|
- Does the user name a specific skill (/cso, /review, /qa)? -> **Heavy**
|
|
- Is it a feature, project, or objective (not a task)? -> **Full**
|
|
- Does the user want to PLAN something for Claude Code without implementing yet? -> **Plan**
|
|
|
|
### Dispatch routing guide (for AGENTS.md)
|
|
|
|
The complete ready-to-paste section lives in `openclaw/agents-gstack-section.md`.
|
|
Copy it into your OpenClaw AGENTS.md.
|
|
|
|
Key behavioral rules (these go ABOVE the dispatch tiers):
|
|
|
|
1. **Always spawn, never redirect.** When the user asks to use ANY gstack skill,
|
|
ALWAYS spawn a Claude Code session. Never tell the user to open Claude Code.
|
|
2. **Resolve the repo.** If the user names a repo, set the working directory. If
|
|
unknown, ask which repo.
|
|
3. **Autoplan runs end-to-end.** Spawn, let it run the full pipeline, report back
|
|
in chat. User should never have to leave Telegram.
|
|
|
|
### CLAUDE.md collision handling
|
|
|
|
When spawning Claude Code in a repo that already has a CLAUDE.md, APPEND
|
|
gstack-lite/full as a new section. Do not replace the repo's existing instructions.
|
|
|
|
## What gstack generates for OpenClaw
|
|
|
|
All artifacts live in the `openclaw/` directory and are generated by
|
|
`bun run gen:skill-docs --host openclaw`:
|
|
|
|
### gstack-lite (Medium tier)
|
|
`openclaw/gstack-lite-CLAUDE.md` — ~15 lines of planning discipline:
|
|
1. Read every file before modifying
|
|
2. Write a 5-line plan: what, why, which files, test case, risk
|
|
3. Resolve ambiguity using decision principles
|
|
4. Self-review before reporting done
|
|
5. Completion report: what shipped, decisions made, anything uncertain
|
|
|
|
A/B tested: 2x time, meaningfully better output.
|
|
|
|
### gstack-full (Full tier)
|
|
`openclaw/gstack-full-CLAUDE.md` — chains existing gstack skills:
|
|
1. Read CLAUDE.md and understand the project
|
|
2. Run /autoplan (CEO + eng + design review)
|
|
3. Implement the approved plan
|
|
4. Run /ship to create a PR
|
|
5. Report back with PR URL and decisions
|
|
|
|
### gstack-plan (Plan tier)
|
|
`openclaw/gstack-plan-CLAUDE.md` — full review gauntlet, no implementation:
|
|
1. Run /office-hours to produce a design doc
|
|
2. Run /autoplan (CEO + eng + design + DX reviews + codex adversarial)
|
|
3. Save the reviewed plan to `plans/<project-slug>-plan-<date>.md`
|
|
4. Report back: plan path, summary, key decisions, recommended next step
|
|
|
|
The orchestrator persists the plan link to its own memory store (brain repo,
|
|
knowledge base, or whatever is configured in AGENTS.md). When the user is
|
|
ready to build, spawn a FULL session that references the saved plan.
|
|
|
|
### Native methodology skills
|
|
Published to ClawHub. Install with `clawhub install`:
|
|
- `gstack-openclaw-office-hours` — Product interrogation (6 forcing questions)
|
|
- `gstack-openclaw-ceo-review` — Strategic challenge (10-section review, 4 modes)
|
|
- `gstack-openclaw-investigate` — Operational debugging (4-phase methodology)
|
|
- `gstack-openclaw-retro` — Operational retrospective (weekly review)
|
|
|
|
Source lives in `openclaw/skills/` in the gstack repo. These are hand-crafted
|
|
adaptations of the gstack methodology for OpenClaw's conversational context.
|
|
No gstack infrastructure (no browser, no telemetry, no preamble).
|
|
|
|
## Spawned session detection
|
|
|
|
When Claude Code runs inside a session spawned by OpenClaw, the `OPENCLAW_SESSION`
|
|
environment variable should be set. gstack detects this and adjusts:
|
|
- Skips interactive prompts (auto-chooses recommended options; destructive or
|
|
irreversible options are never auto-chosen — the conservative choice wins
|
|
and gets recorded in the completion report)
|
|
- Suppresses interactive-onboarding instruction blocks at emission (upgrade
|
|
checks, telemetry prompts, feature discovery, routing injection, tips), so
|
|
one-time prompts survive intact for the next human session
|
|
- Suppresses the Conductor prose signal (`CONDUCTOR_SESSION: true`) — a
|
|
spawned session inside a Conductor workspace auto-chooses instead of
|
|
rendering prose to nobody
|
|
- Focuses on task completion and prose reporting
|
|
|
|
Set the env var in sessions_spawn: `env: { OPENCLAW_SESSION: "1" }`
|
|
|
|
### Explicit override: GSTACK_SESSION_KIND
|
|
|
|
`GSTACK_SESSION_KIND=spawned` is the explicit per-command marker for the same
|
|
classification, outranking every ambient env marker (including
|
|
`OPENCLAW_SESSION` and `GSTACK_HEADLESS`). It exists because Claude Code
|
|
subagents inherit the parent session's env byte-for-byte (#2733) — a
|
|
dispatching skill marks its subagent by prefixing the preamble invocation on
|
|
the same command line:
|
|
|
|
```bash
|
|
GSTACK_SESSION_KIND=spawned "$_SS" --skill "document-release" ...
|
|
```
|
|
|
|
gstack itself uses this: `/ship` Step 18 dispatches the `/document-release`
|
|
subagent with this prefix so its interactive gates auto-choose instead of
|
|
prose-stopping. Deliberately narrow: only `spawned` is honored — `headless`
|
|
already has `GSTACK_HEADLESS`, and letting an env var force `interactive`
|
|
over CI markers would be a misclassification footgun. Empty or other values
|
|
are reserved and ignored (fall through to ambient detection). Note that hook
|
|
processes inherit the harness env, so a per-command prefix never reaches
|
|
PreToolUse/PostToolUse hooks — the hook texts carry a spawned escape sentence
|
|
for that topology (`hosts/claude/hooks/spawned-directive.ts`).
|
|
|
|
**Tamper visibility.** Any mechanism that injects session-wide env (a cloned
|
|
repo's `.claude/settings.json` env block, direnv, a CI wrapper) could set
|
|
`GSTACK_SESSION_KIND=spawned` for a real human's session and silently flip
|
|
its confirmation gates to auto-choose. When the env override drives the
|
|
classification, the preamble emits a loud `SPAWNED_OVERRIDE: env` status line
|
|
so the transcript shows WHY the session is spawned — audit
|
|
`.claude/settings.json` env blocks in untrusted repos (/cso covers this).
|
|
|
|
## Installation
|
|
|
|
For OpenClaw users: tell your OpenClaw agent "install gstack for openclaw."
|
|
|
|
The agent should:
|
|
1. Install gstack-lite CLAUDE.md into its coding session templates
|
|
2. Install the 4 native methodology skills
|
|
3. Add dispatch routing to AGENTS.md
|
|
4. Verify with a test spawn
|
|
|
|
For gstack developers: `./setup --host openclaw` outputs this documentation.
|
|
The actual artifacts are generated by `bun run gen:skill-docs --host openclaw`.
|
|
|
|
## What we don't do
|
|
|
|
- No dispatch daemon (ACP handles session spawning)
|
|
- No Clawvisor relay (no security layer needed)
|
|
- No bidirectional learnings bridge (brain repo is the knowledge store)
|
|
- No JSON schemas or protocol versioning
|
|
- No SOUL.md from gstack (OpenClaw has its own)
|
|
- No full skill porting (coding skills stay native to Claude Code)
|