Files
gstack/CLAUDE.md
T
+8 702a1a9b69 v1.78.0.0 fix: the two-red-lanes wave — AUQ collapse rooted, OSV green from 105, 18 community PRs absorbed, upgrade path can't eat installs (#2752)
* fix(auq): spawned trigger is objective — explicit declaration or STATUS echo, never inference (periodic-lane AUQ collapse)

The v1.76 spawned rule's parenthetical '(or your dispatch prompt marks this
session as spawned)' let the model INFER spawned status from a scripted-looking
prompt in a CI-looking session and silently auto-choose every review-phase
question: reviewCount=0 across the plan-review periodic E2Es (weekly run
33363624506, 9 of 14 failed shards; reproduced locally, zero AUQ fingerprints).
Env and hook paths were excluded by inspection: hermetic children echo
SESSION_KIND: interactive (CLAUDE_CODE_ENTRYPOINT=cli beats CI markers) and the
question-preference hook isn't installed there.

The trigger is now objective: the echoed SESSION_KIND: spawned STATUS line, or
an EXPLICIT dispatch-prompt declaration ("you are a SPAWNED subagent") —
declared, never inferred — with an absence-safe interactive fence: CI env vars,
scripted-looking or pasted prompts, and write-to-this-exact-file instructions
are NOT spawned markers. The prose channel stays because Task-tool subagents
inherit the parent env (no spawned prefix) — their dispatch prompt is the only
signal; #2733's env-prefix channel is untouched.

19 carve skeleton ceilings re-pinned with measured values (+~440 bytes/skill);
ship goldens refreshed for all three hosts; resolver pins extended with the
no-inference regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: mktemp failure aborts loudly at all three skill-content sites; failed upgrade swap restores the backup (#2679)

An empty $(mktemp) result silently disabled the redaction pass (redact-doc
resolver, ship pr-body) and made /gstack-upgrade's vendored path destructive:
clone lands at "/gstack", the swap mv fails, and rm -rf then deletes BOTH the
live install's backup and "". All three sites now guard the assignment with a
loud exit; the vendored block additionally restores the backup when the swap
fails (same failure class — backup deletion after a failed mv) and the GitLab
MR path sends the SCANNED file's bytes instead of re-rendering an unscanned
heredoc. bin/gstack-redact rejects an explicit empty --from-file path instead
of silently falling through to stdin.

Receipts: 6 of 8 new regression checks fail on a v1.77.0.0 scratch worktree.

Fixes #2679

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auq): the interactive fence classifies the session — it never nudges ask-count

Burn-in run 1 of the periodic repro overshot the review band (reviewCount=8 >
CEILING=7) with the fence's 'when unsure, ask' tail: that phrasing is a quota
nudge, not a classification default. The fence now states it only classifies
the session and never changes how many questions the skill asks. Pin added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): OSV suppression config actually loads — explicit global --config + expiring, reasoned ignores

The ignore file was inert from v1.65.0.0: OSV-Scanner only auto-discovers
configs named osv-scanner.toml (no leading dot) and applies them
per-directory, so the root config never covered lib/diagram-render/bun.lock
either way. The workflow now passes --config=.osv-scanner.toml globally.
Every IgnoredVulns entry carries a reason with an upgrade trigger and an
ignoreUntil expiry (~90 days) so suppressions must be re-justified. A wiring
test pins flag ↔ filename ↔ entry hygiene so the file can never silently go
inert again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(deps): dependency wave — 105 OSV advisories → 3 reasoned suppressions, all lanes verified on the pinned scanner

Root: overrides pin ip-address 10.3.1 (defeats BOTH nested nodes — socks'
range pull and express-rate-limit's exact 10.1.0 pin, which a top-level bump
provably cannot reach) and sharp 0.35.0 (GHSA-f88m, HIGH; transformers still
pins ^0.34 upstream — smoke-tested round-trip); marked ^18.0.11; full in-range
lockfile refresh clears hono, fast-uri, protobufjs, qs, body-parser, nanoid,
uuid, immutable and friends.

lib/diagram-render (via its own build-script contract: exact pins edited,
fresh lock, dist rebuilt): mermaid 11.16.1, @excalidraw/excalidraw 0.18.1,
@excalidraw/mermaid-to-excalidraw 1.1.2 → 2.2.2 — the 1.x line exact-pinned
mermaid 10.9.x and dragged the entire duplicate mermaid-10 advisory chain
(dompurify 3.1.6, nanoid 3.3.3, lodash-es); the bundle shrinks 9.96 → 7.59 MB
with the duplicate mermaid gone. Nested exact pins that survived get scoped
overrides (nanoid 5.1.16, lodash-es 4.18.1).

Verification: clean-worktree frozen-lockfile installs (root + nested) + the
SAME osv-scanner release the action pins (v2.3.8) with the workflow's exact
scan-args → exit 0, 'No issues found'. Smoke tests cover the override
surfaces (sharp round-trip, ip-address lockfile assertion, marked parse);
socks + diagram-drift suites already pin the rest.

Supersedes #2695 (its own lockfile kept socks/ip-address@10.2.0; @anupamme's
report credited for the parallel diagnosis).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(gbrain-sync): stub pgrep so the pin case is hermetic

The only non-dry-run --code-only child hits #1734's PATH-resolved
autopilot probe. A live host daemon is a correct refuse; the test
cannot inject processRunning. Neutralize pgrep in the fixture bindir
instead of adding a production env hatch.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(gbrain-sync): blank inherited GBRAIN_HOME in the pin child

Lock paths are checked before pgrep. Spreading process.env let a runner
GBRAIN_HOME with a live lock refuse the case before the stub ran.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: point ship design-checklist at installed gstack/review path

The /ship Design Review step skipped the checklist because the generated path omitted the gstack/ install segment. Sync the generated skill doc and pin a regression assertion.

Co-authored-by: Cursor <cursoragent@cursor.com>
Wave-amended: goldens regenerated against the wave tree (author's golden commit 8e7a03ca superseded)

* fix(codex): a CLI that cannot execute no longer reports CODEX_MODE: ready

Follow-up to #2477. The model probe it added does a real round trip, but its
final branch is the `else` of a "model 400" grep, so it swallowed spawn ENOENT,
non-executable binaries and missing vendor payloads alongside genuine network
timeouts. All three are deterministic — retrying never helps — yet they landed
in the fail-open bucket and resolved to `ready`, so every Codex pass was
skipped in silence and the review reported itself complete.

Observed live: @openai/codex was on PATH with an empty
vendor/aarch64-apple-darwin/codex/ directory. gstack said `ready` for two
months while no Codex pass ran.

Three changes:

- `_gstack_codex_model_probe` classifies deterministic install failures (exit
  126/127, or stderr matching ENOENT/ENOEXEC/EACCES/"cannot execute binary
  file") as MODEL_UNUSABLE_INSTALL, exit 2, never cached — a reinstall is
  picked up on the next probe. Exit 124 and genuine transients still fail open,
  which is what #2477 intended.

- The preflight chain captures the probe's code instead of testing it for
  truthiness, so exit 2 routes to a new `broken_install` mode whose remedy is
  `npm install -g @openai/codex` rather than "check your model pin". A missing
  binary and an unusable model are different problems with different fixes.

- `_gstack_codex_version_check` no longer reads a broken CLI as healthy. It ran
  `codex --version 2>/dev/null | head -1`, which captures head's status, not
  codex's — and 2>/dev/null discarded the one diagnostic available. It now
  captures the real exit code and warns on non-zero. Empty-but-successful
  output stays silent, per the existing "empty output → OK" case.

Tests: 6 added to test/codex-hardening.test.ts covering both broken-install
shapes, the exit-2 contract, no caching, the transient still failing open, the
model 400 still classifying as MODEL_UNUSABLE, and the version-check warning.
845 pass / 0 fail across all 8 suites touching the changed files.

Closes #2742

Wave-amended: autoplan hand-maintained preflight chain completed (tmpl+render); install-signature grep gated on failed spawn only; goldens regenerated against the wave tree (author's golden commit 5797d326 superseded); +2 tests

* feat(redact): add Groq, Tavily and Notion API key patterns

* fix(redact): stop reporting .env.local as an internal hostname

`internal.hostname` ends in `.local|.prod|.staging|…`, so `.env.local`
matches on `env.local` and a dotenv FILENAME is reported as a leaked
internal host.

The collision is not exotic. It fires on `--env-file=.env.local` in an npm
script, `.env.staging` in a README, `.env.prod` in a .gitignore — ordinary
lines on branches that leak nothing. Measured on one private repo, three of
four MEDIUM findings in a routine push were this, and the fourth was a
deleted localhost URL. That ratio is the real cost: a scanner that reports
package.json is one people learn to skim, and skimming is how the HIGH
finding it exists for gets missed.

The guard follows the `insideUuid` precedent and stays deliberately narrow —
it exempts only a span beginning `env.` immediately preceded by a dot, i.e.
the literal `.env.<suffix>` form. `api.corp.local`, `build-7.internal` and
`myenv.local` all still report.

The test pins both directions, and the negative controls are the point: an
exemption written as "any span ending .local" would pass the dotenv half
while quietly gutting the pattern for every real host. Verified red/green —
with the validate hook removed, exactly the 6 dotenv cases fail and all 9
real-host controls still pass.

* fix: don't flag git SSH remotes as pii.email

`pii.email` matches the `git@github.com` inside
`git@github.com:acme/widgets.git`. That is a transport user@host, not a
person's address, so any diff touching a clone URL -- a deploy config's
repo URL, a submodule entry, a README clone line -- draws a spurious
MEDIUM from the pre-push hook.

Suppressed by URL shape rather than by adding `git` to
EMAIL_ALLOW_LOCALPARTS. A bare `git@` allowlist entry would also
suppress a genuine address at a domain that merely begins with "git"
(git@gitmail.com), converting a false positive into a false negative --
the worse failure for a guardrail. Two shapes are accepted:

  - `<user>@<host>:<path>.git` for ANY host, covering self-hosted
    remotes, plus the equivalent ssh:// URL form.
  - `git@<known-host>` for github.com, gitlab.com, bitbucket.org and
    ssh.dev.azure.com, whose bare form appears in docs and in
    `ssh -T git@github.com` connectivity checks with no path at all.
    Matched exactly, so gitmail.com is unaffected.

emailAllowed now receives the normalized text and the span offset so it
can see that surrounding shape; it had only ever been passed the matched
span.

Tests pin both directions: the SSH remotes go quiet, and a real address
still fires -- including at a git host (alex@github.com) and at a
git-prefixed domain (git@gitmail.com).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(redact): install-prepush-hook refreshes a stale managed hook

The marker check returned before the only writer, so once a repo had the
hook, no later change to the wrapper could ever reach it. The `printf x`
fail-open fix (v1.64.0.0) has still not landed in any repo that received
the hook before it, and a wrapper naming a gstack that has since moved
stays pointed at a dead path for the same reason.

Compare the body against what this version generates: rewrite on drift,
stay a no-op when identical. The chained pre-push.local is untouched on
both paths.

The existing trailing-newline regression test cannot catch this — it
installs into a repo with no prior managed hook, the one case that was
never broken.

* fix(redact): install-prepush-hook refreshes a stale managed hook

The marker check returned before the only writer, so once a repo had the
hook, no later change to the wrapper could ever reach it. The `printf x`
fail-open fix (v1.64.0.0) has still not landed in any repo that received
the hook before it, and a wrapper naming a gstack that has since moved
stays pointed at a dead path for the same reason.

Compare the body against what this version generates: rewrite on drift,
stay a no-op when identical. The chained pre-push.local is untouched on
both paths.

The existing trailing-newline regression test cannot catch this — it
installs into a repo with no prior managed hook, the one case that was
never broken.

Wave-amended: spawnSync timeouts added to the new tests (v1.77 sync-spawn tripwire)

* refactor(redact): name the SSH-remote path lookahead constant

Wave polish on the #2734 absorption: the 512-char scp-path lookahead window
follows the UUID_CONTEXT_CHARS named-constant convention instead of a magic
number at the slice site.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(config): reject malformed cross_project_learnings at set

A typo was stored with exit 0, so the feature stayed off and the first-run prompt never returned. Reject like codex_reviews; do not coerce.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(gbrain-detect): classify gbrain >= 0.43 held-lock refusal as engine-locked

gbrain 0.43+ refuses a held PGLite lock with exit 1 and the message
"GBrain's local database is already open through `gbrain serve` (MCP,
PID N)" instead of the pre-0.43 exit 124 + "connect timed out" that
the #2194 branch matches. The message matches no known pattern, so the
classifier falls through to the defensive broken-config default — and
Step 1.5 of /setup-gbrain and /sync-gbrain then tell the user to move a
perfectly healthy config.json aside and re-init the engine.

Reproduced live on gbrain 0.43.0.0, 0.44.0.0 and 0.46.30.0: with a
serve holding the lock, gstack-gbrain-detect reports
gbrain_local_status=broken-config; after stopping the serve it reports
ok with the same untouched config.

Match on the stable substring "already open through", mirroring the
existing #2194 branch semantics: engine-locked for pglite, broken-db
otherwise. Adds a fake-gbrain behavior for the 0.43+ refusal plus two
cases (pglite -> engine-locked, postgres -> broken-db).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory-helpers): a slow gitleaks probe no longer disables secret scanning

`gitleaksAvailable()` cached every failure the same way, so a 2s timeout on
`gitleaks version` was recorded as "the binary is absent" for the rest of the
process. One busy moment and the whole ingest ran unscanned behind a single
stderr line — a fail-open outcome decided by machine load rather than by
anything about the machine's setup. The caller only acts on
`scanner === "gitleaks"`, so every later file was written with no scan and no
second warning.

The probe now classifies three outcomes. ENOENT (and a present-but-unusable
binary: bad exit, EACCES) stays cached — that is a fact about the box, and
re-probing it per file would be waste. A timeout gets one retry on a 10s
budget, and if that also expires nothing is cached: the file is reported
unscanned, the warning says so in those words, and the next file probes again.

Observed under the 7-way sharded free-test runner, where spawning a shell
script inside a temp bin dir took longer than the 2s budget.

Tests: the retry path, the no-cache-on-timeout path (the second call must
re-probe), and the cached-absent path. The fake gitleaks hangs for 30s rather
than racing a short sleep against a short budget, and the budgets are chosen so
load cannot flip an outcome: 30s where the retry MUST answer, 800ms where the
probe MUST expire. An earlier draft used 1s/5s and flaked under the same shard
runner this commit is about. The existing probe test pinned `detect` to
calls[1], which a retry breaks; it now asserts the order instead of the index.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0111Mq3JGwZDcstn5wYcbhSw

* fix(make-pdf): pdftotext version and flavor probe returns unknown on poppler

describeBinary reports version="unknown" flavor="unknown" for every poppler
install, so logDiagnostics prints nothing useful on the most common
implementation. Two independent causes:

1. poppler writes the -v banner to stderr and exits 0. execFileSync returns
   stdout (empty) and does not throw on a zero exit, so the stderr fallback in
   the catch block is unreachable. The in-code comment already notes poppler
   exits 0, but only the throwing path reads stderr.

2. flavor is matched against the version line alone. poppler prints
   "pdftotext version 26.06.0" on line 1 and names itself on line 2,
   "Copyright ... The Poppler Developers", so even a working stderr read
   yields "unknown".

Switch the probe to spawnSync, which returns both streams regardless of exit
status, match the version banner rather than assuming line 0, and derive the
flavor from the full output.

Measured on poppler 26.06.0 (Homebrew, macOS), same machine and binary:

  before: { version: "unknown",                  flavor: "unknown" }
  after:  { version: "pdftotext version 26.06.0", flavor: "poppler" }

xpdf is unaffected: it exits non-zero and names itself on line 1, so it
resolved correctly before and still does.

Tests use shell shims reproducing each vendor's banner, stream and exit status,
since a real pdftotext cannot be assumed present in CI. Two of the four fail on
this commit's parent; the xpdf and no-banner cases pass there and are included
as regression guards rather than red-proofs.

* fix(open-gstack-browser): pre-flight cleanup never killed the stale daemon

Step 0 read the old pid with `grep -o '"pid":[0-9]*'` and Step 2 read the port
the same way. Neither can match. Every writer of that file in
browse/src/server.ts serializes with `JSON.stringify(state, null, 2)`, so the
bytes on disk are `"pid": 12060` — colon, space, digits.

The failure was silent in the worst way. `_OLD_PID` came back empty, the kill
never ran, browse.json was deleted anyway, and the next `connect` died with
"existing daemon has different config (proxy/headed mismatch)" — an error
pointing at proxy/headed flags rather than at the cleanup that no-opped. Caught
against a daemon left over from a reboot: the operator was told to check flags
they had never passed.

Both patterns now accept optional whitespace. The new tripwire does not match
strings — it RUNS the snippets the skill hands the agent, against a state file
written exactly the way the server writes one, and asserts pid and port come
back out. A third case pins the coupling to `JSON.stringify(state, null, 2)`,
so a switch to compact JSON surfaces as a failing expectation rather than as
silence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0111Mq3JGwZDcstn5wYcbhSw

* test(make-pdf): clean up the pdftotext shim tmpdir after the suite

Wave polish on the #2690 absorption: the describe-scope mkdtemp left one
directory per run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): honour CHROMIUM_PROFILE in cli profile-lock cleanup

cli.ts resolved the Chromium profile dir with a hardcoded
$HOME/.gstack/chromium-profile, while browser-manager launches the profile
returned by config.resolveChromiumProfile(), which honours CHROMIUM_PROFILE
and GSTACK_HOME.

killOrphanChromium() and cleanChromiumProfileLocks() are called with no
argument, so whenever CHROMIUM_PROFILE was set they cleaned locks for, and
killed Chromium on, the DEFAULT profile rather than the one being launched.
Starting a browser with a custom profile therefore evicted an unrelated
browser running on the default profile.

Delegating to resolveChromiumProfile() also picks up GSTACK_HOME and
os.homedir(), so the cleanup path now matches the launch path on Windows
where HOME is frequently unset.

* fix(auq): the interactive fence is quota-silent — it defers to the skill's own decision points

Burn-in calibration: run 1 (fence tail 'when unsure, ask') overshot the
plan-ceo review band at reviewCount=8; run 2 (tail mentioning 'HOW MANY
questions') undershot at 1. Any ask-count language in the fence anchors the
model in one direction or the other. The tail now says only: classify as
interactive, then follow the skill's own decision-point instructions exactly
as written. Pins updated to forbid count language in either direction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(browse): pin cli.ts profile-dir wiring to the canonical resolver

Wave-added coverage for the #2732 absorption: a 6-line fix with zero tests is
how the hardcoded path shipped in the first place. resolveChromiumProfile's
env behavior is already pinned in config.test.ts; this pins cli.ts's
delegation and forbids the hardcoded path from returning.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): preserve return value for async IIFE expressions in js/eval (#2727)

Wave-amended: test moved to browse/test/ (browse unit-test convention); trailing-semicolon normalization kept — it is load-bearing for the expression wrapper

* fix: bin writers drop data on Windows paths with an apostrophe

Two independent Windows git-bash bugs in the bin writers, both silent
because callers invoke these scripts with 2>/dev/null and do not check
the exit status — a hard failure was indistinguishable from success.

Bug 1 — apostrophe in the checkout path breaks the bun -e program.
gstack-learnings-log, gstack-question-log and gstack-telemetry-log build
a bun -e program as a double-quoted shell string and interpolate
SCRIPT_DIR into a single-quoted JS import specifier. A path such as
C:/Users/Someone's PC/... closes the JS string literal early and Bun
fails to parse ("Expected ; but found s"). Every learning write and every
plan-tune question event no-oped; telemetry error redaction fell to its
fail-closed null path. The #1950 cygpath -m guard did not cover this —
cygpath normalises the drive form but does not remove the apostrophe.

Fixed by not interpolating the path at all: cd into the module root and
use a relative import specifier, which is immune to apostrophes, spaces,
backslashes and MSYS paths alike. The one remaining interpolated data
path in gstack-developer-profile (readFileSync of PROFILE_FILE) is passed
via the environment instead, matching do_log_session in the same file.

Bug 2 — gstack-developer-profile --derive fails on an MSYS-form
GSTACK_HOME. GSTACK_HOME defaults to $HOME/.gstack, which under git-bash
is /c/Users/..., and Bun on Windows cannot open that form (ENOENT). This
script carried no cygpath guard at all. Fixed by normalising GSTACK_HOME
once, before PROFILE_FILE / LEGACY_FILE / the events path are derived
from it, so all three pick up the normalised value.

Adds test/hostile-path-writers.test.ts, which runs the bins from a
directory whose name contains an apostrophe and asserts that rows are
ACTUALLY WRITTEN (not merely that the exit code is 0 — exit-code-only
checks are what masked bug 1). The apostrophe repro is OS-independent:
SCRIPT_DIR derives from the script's own location, so a copied checkout
under a hostile directory name reproduces bug 1 on Linux/macOS CI too.

Wave-amended: all four writers unified on the env-var import pattern the PR already used in gstack-developer-profile (no CWD-dependent module resolution)
Wave-amended: all four writers unified on the env-var import pattern the PR already used in gstack-developer-profile (apostrophe-safe without CWD-dependent module resolution); import-shape pin updated

* fix(memory-ingest): stop two silent transcript-ingest failures

Two independent bugs made transcript pages silently fail to reach the brain.

1. Frontmatter fence gluing. buildTranscriptPage() built the closing "---"
   with no trailing newline, and session bodies always start with "## ", so
   the rendered page ended "...---## User". gbrain's frontmatter matcher
   (/^---\r?\n([\s\S]*?)\r?\n---(\r?\n|$)/ in src/core/markdown.ts) requires
   the closing "---" to end its own line, so it skipped the glued fence,
   latched onto the next standalone "---" in the transcript body, parsed the
   prose between as YAML, and dropped the page with "Invalid YAML frontmatter".
   Transcripts with no later "---" fell back to body-only, silently losing
   their frontmatter. Fix: emit the fence on its own line with a blank
   separator, matching renderPageBody()'s artifact branch.

2. Slug collisions. Two source files can map to one path-derived slug (a
   session resumed under the same id on one day, or two ids sharing a 12-char
   prefix). writeStaged() names each file "${slug}.md", so the second
   overwrote the first; gbrain collected N-1 of N staged files and the
   reconciliation guard failed the whole batch every run. Fix:
   disambiguateSlugs() keeps the first occurrence and gives each later collider
   a stable "-<sha8(source_path)>" suffix (deterministic, and slug + page_slug
   move together so writeStaged, the failure mapping, and state recording agree).

Exports buildTranscriptPage, renderPageBody, and disambiguateSlugs for tests.
Adds regression tests for both failures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wave-amended: contributor's local-workaround docblock note removed; issue refs retargeted #2653 (closed by its author) -> #2724 (the live 887-staged-to-0-ingested report)

* fix: keep feature markers in GStack state

* fix: align feature marker seeding with GStack state

Wave-amended: seeding relocation re-applied to the composite action (v1.77 moved CI seeding out of the inline workflow steps the original commit edited); wiring tripwire re-pointed accordingly; stale marker comment updated

* chore(upgrade): migrate feature-discovery markers to GSTACK_HOME

Follow-through on the #2748 absorption: existing installs answered the
continuous-checkpoint and model-overlay prompts with markers beside the
install; v1.78 reads them from GSTACK_HOME. Copy them once so nobody gets
re-prompted. Idempotent, non-fatal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(land-and-deploy): check fork branch in head repo

Wave-amended: gh leaves .headRepository.nameWithOwner empty (verified live against gh 2.83) — owner/name now composed from headRepositoryOwner.login + headRepository.name so reconciliation is not a permanent no-op; fork branches get report-not-delete (maintainers lack fork push rights); pins updated

* test: spawn timeouts on the #2748 marker tests (v1.77 sync-spawn tripwire)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: spawn timeouts on absorbed-PR tests (v1.77 sync-spawn tripwire)

The absorbed community tests (#2748, #2676, #2714, #2720) were authored
before the v1.77 tripwire required a timeout on every sync spawn in the test
trees.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gbrain): a slow --version probe classifies as timeout, never no-cli (#2716)

resolveGbrainBin's bare catch collapsed 'gbrain missing' and 'gbrain present
but the 2s --version budget expired' into the same null — freshClassify then
said no-cli, which the --is-ok whitelist from #1964 does NOT forgive, so a
bun-shim install on a loaded POSIX box silently lost every brain-aware block.
The probe now returns a discriminated result (cached per-process, same
lifetime the old null had) using the same killed/SIGTERM/ETIMEDOUT
discrimination the sources-list probe below already uses; timeout routes to
the forgiven 'timeout' status. GSTACK_GBRAIN_VERSION_PROBE_TIMEOUT_MS test
override added (same precedent as the sources-probe override).

Receipt: the slow-but-present sibling test fails on a v1.77.0.0 scratch
worktree (classifies no-cli there).

Fixes #2716

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(codex): close the consult-mode fence, report turn.failed as a failure, capture exit codes portably (#2671, #2669)

Three defects in the codex skill sections:

- The resumed-session bash block never closed its fence; every fenced region
  after it inverted (prose rendered as code, the synthesis-recommendation tail
  rendered inert). A repo-wide fence-pairing test now scans every generated
  SKILL.md and sections/*.md with a CommonMark-faithful state machine (an
  info-string opener inside a fence is literal content — nested template
  examples in document-generate/make-pdf stay legal; a file ending inside a
  fence fails).

- The JSONL parsers had no turn.failed branch: a turn that STATED its failure
  was reported as 'possible mid-stream disconnect'. Challenge and consult now
  print the event's error and run a three-way completeness check (failed-with-
  reason / silent-disconnect / ok); consult previously had no completeness
  check at all.

- ${PIPESTATUS[0]} is empty under zsh, so hang detection never fired and
  every clean run printed a spurious '[codex exit ]'. All three capture sites
  use ${PIPESTATUS[0]:-${pipestatus[1]}}, pinned statically and EXECUTED
  under real bash and zsh in the new test. Expect a step-change in
  codex_timeout telemetry — the counter starts firing for zsh users.

Receipt: the portability pin fails on a v1.77.0.0 scratch worktree; the fence
fix is structural (17 → 18 fence lines, tail no longer inside a block).

Fixes #2671
Fixes #2669

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: outside-voice fallback is labeled honestly — same model family, not cross-model (#2735)

When Codex is unavailable, the plan-review outside voice falls back to a
Claude subagent and the copy sold it as 'cross-model coverage' with 'genuine
independence'. Fresh context is real; cross-model validation is not — a user
weighing 'both reviewers agree' deserves to know both reviewers share a model
family. Six canonical strings fixed at the resolver source (constants.ts
not_installed/not_authed, review.ts outside-voice bullet + three dispatch
paragraphs); ~10 generated docs and the ship goldens regenerated. Printing
the resolved fallback model at dispatch time is descoped as a functional
change (follow-up in the wave dispositions).

Fixes #2735

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(relink): skill_prefix patches the gbrain render too — the file the host actually serves (#2738)

gstack-relink linked SKILL.md from RENDER_DIR when a gbrain render was active
but ran gstack-patch-names only on INSTALL_DIR, so the served frontmatter kept
the unprefixed name and skill_prefix=true silently no-oped for every
brain-aware skill. The render tree (user-owned, untracked) is now patched too;
gstack-patch-names is idempotent so repeat relinks never double-prefix. The
gen-skill-docs note that pointed users at relink now describes what relink
actually covers.

Receipt: the new test fails on a v1.77.0.0 scratch worktree (served render
keeps 'name: qa').

Fixes #2738

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(render): section refs point at the FINAL render dir, never the tmp swap dir (#2692)

gen-skill-docs bakes its --out-dir into rendered CONTENT (rewriteSectionBase),
and both swap-in callers (setup, gstack-config gbrain-refresh) render into
claude.tmp.<pid> before the #2569 atomic rename — so every rendered skill
carried ~9 dead section Read paths that pointed at a directory the swap had
just deleted. New --link-root flag names the final serving dir (defaults to
--out-dir for direct-render callers: bin/dev-setup, dev-skill.ts, mkdtemp
tests — full caller audit in the wave notes); the rewrite now uses a
replacement callback so a $-bearing configured path can't expand as $& in a
replacement string. The swap logic itself stays byte-identical. Tests pin the
generator contract (tmp out-dir files reference the final dir, $-bearing
path included) and both callers' wiring.

Fixes #2692

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(setup): persistent timeline Stop hook opt-out — timeline_stop_hook config gate (#2677)

--no-team is a one-shot teardown, so every later bare ./setup (including the
ones /gstack-upgrade runs) re-registered the timeline Stop hook with no way
to say 'never'. New gate mirrors the plan_tune_hooks pattern: flag
(--timeline-stop-hook/--no-timeline-stop-hook) > env
(GSTACK_TIMELINE_STOP_HOOK) > saved config (timeline_stop_hook) > default
yes. An explicit flag persists to config so the decision survives upgrades;
an explicit 'no' also removes a live registration (reconciliation), so the
opt-out works against installs registered by an older setup. --no-team
semantics unchanged (NO_TEAM_MODE is never initialized from config). Full
gstack-config surface: DEFAULTS entry, header docs, list/defaults
enumeration, warn-and-default validation.

Fixes #2677

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): tame the macOS headless GPU spin + reap the lock-less headless Chromium on stop (#2709)

Two defects in one report. On macOS 26 / Apple Silicon the headless-shell GPU
process pegs ~800% CPU indefinitely after real page work and --disable-gpu
alone is not enough; the reporter validated that adding
--disable-software-rasterizer/--disable-gpu-compositing/--disable-gpu-watchdog
drops it to 0.0% with screenshots still working. The flag block is a pure
platform-parameterized function (unit-tested on any host), darwin-gated,
headless-only (buildGStackLaunchArgs feeds the headed/GBrowser paths where
GPU-off is wrong), with a GSTACK_DISABLE_GPU=off escape.

Separately: the headless launch has no userDataDir, so it never writes the
SingletonLock that killOrphanChromium walks — 'browse stop' reported success
while the orphan kept spinning. The daemon now records the launched child's
pid + wall-clock start time in the state file (the xvfbPid/xvfbStartTime
contract), and stop paths reap a survivor only after verifying BOTH the
recorded start time and a Chromium-looking cmdline — a recycled PID, even one
running a different legitimate Chromium, is never killed (identity tests
include the coreutils-shebang trap that defeats argv0 renames).

macOS efficacy is per the reporter's validation; live re-verification on
Apple silicon is tracked in TODOS.md.

Refs #2709

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(wtree): a failed touch falls through to the HEAD seed instead of reopening the racy window (#2687)

The v1.74 racy-git fix carries the real index's mtime onto the temp copy —
but its 'touch -r … || true' meant a FAILED touch silently kept the copy's
fresh stamp, marking every entry non-racy and reopening the exact same-size-
rewrite hole. A failed touch now discards the copy and seeds from read-tree
HEAD (slower; every entry re-hashed; fingerprint stays honest).

Verification for #2687 itself: the reporter's same-size-rewrite repro run 20
iterations against this tree — 0 misses (the underlying race was fixed by
v1.74's b1485d88 with its own regression test; this wave verifies and closes,
it does not claim that fix). Receipt: the stubbed-touch test fails on a
v1.77.0.0 scratch worktree.

Fixes #2687

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: rewrite gate pin follows the LINK_ROOT rename (#2692)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: v1.78 fix-wave deferrals filed in TODOS.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v1.78.0.0 release metadata: VERSION, package.json translation, CHANGELOG wave entry, agents digest

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(osv): ignore ledger names its filed tracking issues (#2753, #2754)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(redact): large reports survive the pipe — exitCode instead of process.exit; inert test payload

The wave's PR quality gate failed closed: gate-secret-scan.mjs pipes the
diff's added lines into gstack-redact and parses the JSON report, but
process.exit() discards stdout still buffered in the pipe — this wave's
646-finding report (202 KB) is the first big enough to arrive truncated
(~145 KB) at node's collector, so JSON.parse failed and the gate read
'no report' as HIGH. The report and auto-redact body paths now set
process.exitCode and let the runtime drain stdout; exit-code contract
unchanged (verified 0/2/3 end-to-end). Also: the C1 test's stdin payload no
longer uses a provider-prefix credential shape (the gate correctly flagged
it; the content was never read on the error path under test).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(memory-ingest): fence regression test survives Windows tmpdirs

Wave polish on the #2699 absorption: the hand-built JSONL interpolated the
raw tmpdir into a JSON string — on Windows (D:\a\...) that's an invalid
escape, the user line was silently dropped, and the body started at
'## Assistant' (Windows Free Tests red). JSON.stringify the path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auq): the interactive fence ends at classification — all behavioral tails removed

The pinned-container periodic lane proved the collapse dead (reviewCount 0 →
7/8/5 across the AUQ suite) but flagged the fence's remaining behavioral
clause: 'never adds, removes, or batches the skill's decision points' broke
the paired-finding control (5 > 4 — it suppressed the batching that fixture
expects), and the band overshot its ceiling (8 > 7). Every behavioral tail
tried so far skewed counts somewhere ('when unsure, ask' → 8; 'HOW MANY
questions' → 1; 'never batches' → paired control red). The fence now ends at
'When unsure, default to interactive.' — classification only, zero behavior
words. Pins forbid every tried-and-failed phrasing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auq): the spawned trigger is the STATUS echo, nothing else — prose channel removed from the eager path

Two pinned-container periodic rounds showed that ANY dispatch-prompt
declaration channel in rule 1 keeps question counts unstable (round 1, fence
with behavioral clause: paired control 5>4, band 8>7; round 2, bare fence:
intermittent 0s return, paired control breaks both directions). The stable
regime CI was calibrated against had no spawned prose in the eager path at
all. Rule 1 now keys on exactly one machine-verifiable thing: the preamble's
own SESSION_KIND: spawned STATUS echo. No text from a dispatch prompt, file,
or page can flip a session to auto-choose (the strongest anti-injection
form). Subagents that missed the env marker are caught at FAILURE time by
the AUQ hooks' spawned escape (explicit declaration, never inference) — a
channel that never enters an interactive session's eager reasoning.

This reverses the wave's earlier explicit-declaration middle ground (and
adopts the outside voice's twice-made echo-only argument) on the new
evidence. #2733 protected: skill-e2e-docsync-spawned (gate) passes 1/1 on
this prose — the ship Step-18 dispatch forces the env prefix, so the echo
fires there.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: periodic-lane stabilization residual filed (#2756)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): chromium reap works off-Linux and on every stale-state path

readPidCmdline fell back to '' on darwin (no /proc), so the identity gate
never matched and reapRecordedChromium was inert on the platform #2709's
GPU-spin reap actually targets — it now falls back to ps -o command=.
readPidStartTime no longer throws when ps is missing (Windows): a launch
must never die to a reap-bookkeeping probe. Three stale-state cleanup
paths (dead-daemon stop, startServer stale cleanup, headed-connect) now
reap the recorded chromium BEFORE unlinking the state file instead of
orphaning it, and the stop-path wait polls (100ms steps, 1s cap) instead
of sleeping a fixed 500ms. Wiring pinned: server-state pid/start-time
write, all five cli.ts reap call sites, headless-only GPU-flag push.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(browse): chained IIFE + second statement no longer misclassified as one expression

isSingleParenOrIifeExpression accepted any tail after the initial group's
close as long as trailing chars looked chain-ish, so
`(async()=>{await 1})().then(x=>x); console.log('done')` classified as a
single expression and the expression wrapper emitted a SyntaxError. The
tail is now consumed as a strict member/call/index/optional-chain walk to
END of input via a shared string/escape-aware findBalancedClose scanner;
anything else (';', operators) demotes to the block wrapper. Negative +
positive tests added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(browse): move headlessGpuArgs below the import block

The #2709 helper landed between two import statements; imports now stay
contiguous. No behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(codex-probe): timed-out probe (124) keeps its fail-open contract

Exit 124 reached the string-signature branch before the timeout fail-open,
so a slow probe whose partial output happened to quote 'permission denied'
classified as MODEL_UNUSABLE_INSTALL — a deterministic-broken verdict from
a transient condition. 124 is now excluded from the signature branch, and
the detect/display greps share one hoisted _BROKEN_SIG regex (they had
already drifted: display dropped 'not executable').

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(codex): JSONL parser initializes its state vars in both modes

challenge-mode initialized turn_completed_count but tested turn_failed via
'in dir()'; consult-mode initialized neither and rebuilt the counter with
a dir() conditional per event. Both parsers now init turn_completed_count
and turn_failed up front and use plain checks — same semantics, no
module-globals introspection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ship): PR/MR create aborts on a missing or empty scanned body file

Both the gh and glab send blocks now guard [ -s "$PR_BODY_FILE" ] and the
prose restates that the variable comes from the scan block — bash blocks
run in separate shells, and an unset/empty path would previously send an
empty body (gh) or cat's error output (glab) instead of the scanned bytes.
Codex/factory ship goldens regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(upgrade): abort when a stale .bak already exists at the install path

A leftover $INSTALL_DIR.bak from a crashed upgrade would make the mv nest
the live install inside it, and the failure-restore arm would 'restore'
the stale backup — possibly deleting the only good copy. The upgrade now
refuses to start and tells the user to inspect/salvage the backup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(redact-doc): mktemp-failure message names what it refuses to send

'refusing to send unscanned <noun>' read as if 'unscanned' modified a
missing word for sink nouns like 'the spec body'; now 'refusing to send
<noun> unscanned'. Generated spec section refreshed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(setup): typo'd timeline-stop-hook value warns instead of persisting

--timeline-stop-hook=noo silently normalized to yes AND wrote yes to
config — a persisted decision the user never made. Unrecognized values now
warn (naming the source), apply the default for this run only, and skip
the config write. The opt-out log line names the actual decision source
(flag/env/config) and no longer claims a removal that may not have
happened.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory-helpers): slow-probe warning no longer suppresses the absent warning

One shared _gitleaksWarned flag served two different messages: a 'machine
under load, retrying next file' warning early in a run permanently
silenced the later 'gitleaks not in PATH; secret scanning disabled'
warning — the user never learned scanning was off for good. Split into
per-message flags.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(lib): shared isExecTimeout helper; export GbrainBinProbe

The killed/SIGTERM/ETIMEDOUT discrimination was hand-rolled at three sites
(gbrain version probe, engine classifier, gitleaks probe) and free to
drift; it now lives once in lib/gbrain-exec.ts. GbrainBinProbe is exported
(it's the return type of exported probeGbrainBin) and the cache carries a
rationale comment: caching a timeout for process lifetime is deliberate —
the memo dedupes the ~3 probes of one short-lived preamble process.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(gen-skill-docs): extract parsePathFlag; fix rewriteSectionBase docstring

--out-dir and --link-root shared near-identical inline parsing; one helper
now owns it. The rewriteSectionBase docstring said 'no-op when --out-dir
is unset' but the gate is the link root (which --link-root can set
independently) — it now describes the real behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(memory-ingest): reunite preparePages with its docblock; pin disambiguateSlugs wiring

The #2724 disambiguateSlugs block was inserted between preparePages'
docblock and the function, orphaning the secret-scanning policy doc onto
the wrong symbol. Reordered. A call-site pin now asserts the prepare→stage
flow actually invokes disambiguateSlugs, so a refactor can't drop the call
while every unit test stays green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(hostile-path): per-run mkdtemp root; telemetry-log redaction coverage

The suite used a FIXED tmpdir name, so concurrent runs (sharded runner,
sibling worktrees) tore down each other's trees mid-flight — now a
per-run mkdtemp root with the apostrophe dir inside. gstack-telemetry-log
was the one bin named in the suite header with no test: it now must append
a real row under the hostile path with the credential span redacted
(<REDACTED-github.pat>) and the rest of the message preserved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(config): signal-killed spawns map to -1, not exit 0

Both cfg() helpers defaulted a null spawn status to 0 — a child killed by
signal would read as success and mask real failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(upgrade): v1.78.0.0 feature-marker migration suite

The only migration without a dedicated test. Covers copy-when-absent
(script must mkdir GSTACK_HOME itself), destination-wins (never
overwrites), clean no-op, and two-run idempotence — asserting file
existence and contents, not just exit codes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(auq): pin the explicit-declaration-only spawned escape sentence

SPAWNED_ESCAPE_SENTENCE's tightened wording had no pin: positive pins on
the explicit-declaration clause, negative pins on the retired v1.76 loose
parentheticals ('e.g. your dispatch prompt says', 'marks this session as
spawned'), and a drift guard that both hook directives embed the constant
verbatim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(gbrain): invalid version-probe timeout env falls back to the default

GSTACK_GBRAIN_VERSION_PROBE_TIMEOUT_MS set to 'abc', '-1', or '0' must use
the default budget — exercised behaviorally through probeGbrainBin with a
fresh PATH per case (the memo keys on PATH).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(codex): execute the JSONL parser under real python3; fence scanner tracks opener length

The parser's turn.completed/turn.failed/disconnect semantics were pinned
by shape only — now the python block is extracted from both RENDERED
sections and run against synthetic event streams (tokens line, FAILED +
not-a-disconnect, silence -> disconnect warning, SESSION_ID echo), with
byte-equivalence safety pins on the bash double-quote extraction. The
fence scanner also gains CommonMark opener-length tracking: a 4-backtick
fence wrapping a 3-backtick example no longer false-positives, with a
self-test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(redact): large report survives a slow piped consumer

Pins the >145KB truncation regression (process.exit before the pipe
drained): 900 MEDIUM findings -> 259KB JSON report through a sleep-first
POSIX consumer that holds the 64KiB kernel buffer full at child exit;
asserts complete parseable JSON with matching counts and exit 2, plus an
--auto-redact mirror (700 redactions, final sentinel byte arrives).
Harness proven red against a copy of the bin with process.exit restored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(dev-setup): update LINK_ROOT source pin to the parsePathFlag shape

Companion to the gen-skill-docs parsePathFlag extraction: the pin still
asserts the same invariant (LINK_ROOT defaults to OUT_DIR, so an in-place
render stays a byte-exact no-op) against the new expression.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): fix-batch properties folded into the v1.78.0.0 entry

Stale-backup refusal + empty-scanned-body guard on the mktemp bullet, the
redact pipe-truncation fix as its own item (a v1.77 bug), and test counts
refreshed to the post-fix-batch suite (8,660).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(upgrade): migration test resolves bash through the parent PATH

A hardcoded /usr/bin:/bin child PATH breaks spawn('bash') on the Windows
curated lane (spawn resolves against the CHILD env's PATH; no bash.exe
lives there). Hermeticity is carried by HOME/GSTACK_* overrides, not PATH.
Found by the cycle-2 review pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory-helpers): per-run cooldown bounds the slow-gitleaks probe cost

Retrying a slow probe per FILE (#2715's slow!=absent split) re-paid up to
probe+retry (12s default) per file — an 887-file ingest on a loaded box
spent hours re-asking the same slow question. After 3 consecutive slow
answers the run stops probing and warns once that remaining files go
unscanned; the availability cache is still never written, so the next
process probes fresh. Slow/absent discrimination is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory-ingest): slug assignments persist across runs via the state consult

First-occurrence-keeps-bare was walk-order-dependent ACROSS runs: a source
that got the suffixed slug once could take the bare slug the next run (its
collider aged out or was skipped as unchanged), leaving gbrain holding the
same transcript under two slugs — and a NEW collider could claim a bare
slug that state shows belongs to an unchanged source, silently overwriting
that page. disambiguateSlugs now consults state.sessions: a recorded slug
stays owned by its source_path, re-ingested sources keep their slug
verbatim, fresh assignments never take another source's slug, and legacy
duplicate records (pre-#2724 overwrites) resolve first-owner-wins and
self-heal on the next state write. Stateless behavior is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): D2/D3 properties folded into the absorbed-PR bullets

Gitleaks per-run probe cooldown on the #2715 credit; cross-run slug
persistence on the #2699/#2724 credit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.78.0.0

README.md: the persistent timeline Stop hook opt-out (#2677) — flag,
env var, and config key with resolution order. BROWSER.md: browse stop
against a dead daemon now reaps the recorded headless Chromium child,
identity-verified (#2709). CLAUDE.md + CONTRIBUTING.md: free-suite test
count ~7,000 → ~8,700 (8,660 as of this wave).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: cross-model doc review fixes for v1.78.0.0

CONTRIBUTING.md: the day-to-day example now edits the .tmpl (SKILL.md
is generated); the OSV row states the explicit --config load and the
reasoned, expiring ignore contract. BROWSER.md: stop row mentions the
identity-checked Chromium reap; env table gains CHROMIUM_PROFILE and
GSTACK_DISABLE_GPU rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): headline claims what the receipts show

"Both red weekly lanes are green again" overclaimed: OSV is verifiably
green (pinned scanner, frozen install, branch dispatch), but the periodic
lane keeps its pre-wave churn (#2756) — what this wave proves is that the
v1.76 regression that silenced plan reviews is dead. Flagged by the
cross-model doc review; headline now leads with the user-visible outcome.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: y$un_ <forrest.sun527@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Lockyer <135391289+Lockyer228@users.noreply.github.com>
Co-authored-by: Udhdhav kheni <udhavkheni12@gmail.com>
Co-authored-by: schienbiz <274676847+schienbiz@users.noreply.github.com>
Co-authored-by: David Park <show@davidani.com>
Co-authored-by: alopes50 <alex@alexlopes.com>
Co-authored-by: Peter van Leeuwen <petervanleeuwen@SB-petervanleeuwen.local>
Co-authored-by: Denis Zjukow <denis.zjukow@gmail.com>
Co-authored-by: Paul Snyman <5826275+snymanpaul@users.noreply.github.com>
Co-authored-by: Adam Badar <badaradam10@gmail.com>
Co-authored-by: loulanyue <260355617@qq.com>
Co-authored-by: Shreshth Kapoor <shreshth@osiflow.com>
Co-authored-by: Ryan Ayers <rayers@dividia.net>
Co-authored-by: Simon Altit <simon.altit@gmail.com>
Co-authored-by: ptt <1928627998@qq.com>
2026-09-01 11:15:11 -07:00

45 KiB

gstack development

Commands

bun install          # install dependencies
bun run test         # run free tests via the strict parallel runner (~90-100s full suite)
bun run test:evals   # run paid evals: LLM judge + E2E (diff-based, ~$4.35/run max)
bun run test:evals:all  # run ALL paid evals regardless of diff
bun run test:gate    # run gate-tier tests only (CI default, blocks merge)
bun run test:periodic  # run periodic-tier tests only (weekly cron / manual)
bun run test:gate:sharded    # gate tier via the sharded paid runner (one Bun process per test file)
bun run test:periodic:sharded  # periodic tier via the sharded paid runner (implies EVALS_ALL=1)
bun run test:e2e     # run E2E tests only (diff-based, ~$4.20/run max)
bun run test:e2e:all # run ALL E2E tests regardless of diff
bun run eval:select  # show which tests would run based on current diff
bun run dev <cmd>    # run CLI in dev mode, e.g. bun run dev goto https://example.com
bun run build        # gen docs + compile binaries
bun run gen:skill-docs  # regenerate SKILL.md files from templates
bun run skill:check  # health dashboard for all skills
bun run dev:skill    # watch mode: auto-regen + validate on change
bun run eval:list    # list all eval runs from ~/.gstack/projects/<slug>/evals/
bun run eval:compare # compare two eval runs (auto-picks most recent)
bun run eval:summary # aggregate stats across all eval runs
bun run eval:flake-rank  # rank tests by flake signal (retried passes first; --json, --dir, --since-days)
bun run slop          # full slop-scan report (all files)
bun run slop:diff     # slop findings in files changed on this branch only

test:evals requires ANTHROPIC_API_KEY. Codex E2E tests (test/codex-e2e.test.ts, test/codex-e2e-sol-scope.test.ts) use Codex's own auth — the hermetic runner copies only auth.json from ${CODEX_HOME:-~/.codex} and pins CODEX_HOME in the child env — no OPENAI_API_KEY env var needed.

Hermetic E2E + env keys: every E2E runner spawns children through test/helpers/hermetic-env.ts (allowlist-scrubbed env, fresh seeded CLAUDE_CONFIG_DIR, temp GSTACK_HOME, --strict-mcp-config); per-test env: overrides merge last onto a COMPLETE hermetic env, so they're safe. A PTY test that types a /skill command must pass seedSkills: true. Debug against real operator state with EVALS_HERMETIC=0. Full detail (env-shim, seeding tripwires, wiring tests): docs/TESTING_INTERNALS.md.

Diff-based test selection: test:evals and test:e2e auto-select tests based on git diff against the base branch. Each test declares its file dependencies in test/helpers/touchfiles.ts. Changes to global touchfiles (session-runner, eval-store, touchfiles.ts itself) trigger all tests. Use EVALS_ALL=1 or the :all script variants to force all tests. Run eval:select to preview which tests would run.

Two-tier system: Tests are classified as gate or periodic in E2E_TIERS (in test/helpers/touchfiles.ts — a facade over touchfiles-data.ts + test-selection.ts). CI runs gate tests per PR via evals.yml's sliced lane (planner manifest → executors → fail-closed report; engine = scripts/test-paid-shards.ts, the same runner as local eval:bg:gate); the free suite runs on every PR via .github/workflows/free-tests.yml (a REQUIRED check, secretless — fork PRs get real signal); ALL periodic tests run weekly via evals-periodic.yml (EVALS_ALL, minus the reasoned exclusions in test/helpers/periodic-exclude-data.ts — reason + tracking required per entry), plus a weekly EVALS_ALL gate census. Use EVALS_TIER=gate or EVALS_TIER=periodic to filter locally. When adding new E2E tests, classify them:

  1. Safety guardrail or deterministic functional test? -> gate
  2. Quality benchmark, Opus model test, or non-deterministic? -> periodic
  3. Requires external service (Codex, Gemini)? -> periodic

Tier declarations are enforced by test/e2e-tier-alignment.test.ts (free, runs in bun test): a skill-e2e-* file named in a touchfiles dep list whose EVALS_TIER self-gate disagrees with its declared tier in E2E_TIERS fails the suite. Files not named in any dep list are reported, not enforced — keep both in sync.

Testing

bun run test         # run before every commit — free, ~90-100s for the full ~8,700-test suite
bun run test:evals   # run before shipping — paid, diff-based (~$4.35/run max)

bun run test routes through scripts/test-free-shards.ts (N concurrent shard processes, serial within each, packed by recorded per-file durations when scripts/free-test-durations.json exists — refresh occasionally with bun run test:free --record-durations; strict-output classification per shard: a shard without bun's terminal summary line FAILS — silent truncation cannot report green). The former trailing serial tree-mutating shard is gone: TREE_MUTATING is empty (gen-skill-docs has a main() guard and --out-dir renders every host, so tests render into mkdtemps — see docs/TESTING_INTERNALS.md). Never type bare bun test for the suite: it walks the whole repo, loading paid eval files and missing the strict classifier. It covers skill validation, gen-skill-docs quality checks, and browse integration tests. bun run test:evals runs LLM-judge quality evals and E2E tests via claude -p. Both must pass before creating a PR.

Project structure

Full annotated tree: docs/PROJECT_STRUCTURE.md. Quick map: browse/ headless-browser CLI, design/ design binary, hosts/ typed host configs, scripts/ build+DX tooling (gen-skill-docs, resolvers), test/ validation+evals, lib/ shared libraries, bin/ CLI utilities, extension/ Chrome extension, one directory per skill (ship/, review/, qa/, ...), .github/ CI, contrib/ contributor tools, docs/designs/ design documents.

SKILL.md workflow

SKILL.md files are generated from .tmpl templates. To update docs:

  1. Edit the .tmpl file (e.g. SKILL.md.tmpl or browse/SKILL.md.tmpl)
  2. Run bun run gen:skill-docs (or bun run build which does it automatically)
  3. Commit both the .tmpl and generated .md files

Generation uses each host's defaultModel (claude for existing hosts, gpt for Codex) unless --model is explicit. Codex installs additionally read the top-level model from ${CODEX_HOME:-~/.codex}/config.toml; rerun ./setup --host codex after changing that model. Note: bun run build and a bare gen:skill-docs --host codex render the host default (gpt) — if your Codex config.toml pins a different model, rerun ./setup --host codex afterwards to restore your profile (single-owner persistence is filed in TODOS.md).

To add a new browse command: add it to browse/src/commands.ts and rebuild. To add a snapshot flag: add it to SNAPSHOT_FLAGS in browse/src/snapshot.ts and rebuild.

Token ceiling: Generated SKILL.md files trip a warning above 160KB (~40K tokens). This is a "watch for feature bloat" guardrail, not a hard gate. Modern flagship models have 200K-1M context windows, so 40K is 4-20% of window, and prompt caching makes the marginal cost of larger skills small. The ceiling exists to catch runaway preamble/resolver growth, not to force compression on carefully-tuned big skills (ship, plan-ceo-review, office-hours legitimately pack 25-35K tokens of behavior). If you blow past 40K, the right fix is usually: (1) look at WHAT grew, (2) if one resolver added 10K+ in a single PR, question whether it belongs inline or as a reference doc, (3) only compress carefully-tuned prose as a last resort — cuts to the coverage audit, review army, or voice directive have real quality cost.

A second, harder ceiling guards the DISCOVERY surface: test/catalog-budget.test.ts caps the aggregate frontmatter name + description across all skills at 1,150 token-equivalents (260-byte per-skill sub-cap), counted through the shared census in test/helpers/skill-census.ts. This one is enforced, not a warning — every host loads the full catalog every session, so growth here taxes every conversation. The failure message carries the re-measure + ratchet protocol. bin/gstack-context-bill shows the full token bill-of-materials for a skills tree (always-on vs per-invocation, --diff, --budget; --exact opts into the real tokenizer and POSTs file text to api.anthropic.com with an egress receipt).

The context-budget ratchet (test/context-budget-ratchet.test.ts, free, runs in bun run test) pins ABSOLUTE ceilings on two more ledgers: the always-on FULL-frontmatter aggregate (catalog-budget counts only name+description) and each skill's per-invocation eager tokens (SKILL.md + forced-read references — size floors and parity ratios guard these relatively, not absolutely), graded against test/fixtures/context-budget.json. A skill that grows past its ceiling fails; a new skill fails until it's consciously budgeted. For legitimate growth or a landed reduction, re-run bun test/helpers/capture-context-budget.ts and commit the refreshed fixture in the same commit, so ceilings ratchet down and every win is locked.

Merge conflicts on SKILL.md files: NEVER resolve conflicts on generated SKILL.md files by accepting either side. Instead: (1) resolve conflicts on the .tmpl templates and scripts/gen-skill-docs.ts (the sources of truth), (2) run bun run gen:skill-docs to regenerate all SKILL.md files, (3) stage the regenerated files. Accepting one side's generated output silently drops the other side's template changes.

Platform-agnostic design

Skills must NEVER hardcode framework-specific commands, file patterns, or directory structures. Instead:

  1. Read CLAUDE.md for project-specific config (test commands, eval commands, etc.)
  2. If missing, AskUserQuestion — let the user tell you or let gstack search the repo
  3. Persist the answer to CLAUDE.md so we never have to ask again

This applies to test commands, eval commands, deploy commands, and any other project-specific behavior. The project owns its config; gstack reads it.

Writing SKILL templates

SKILL.md.tmpl files are prompt templates read by Claude, not bash scripts. Each bash code block runs in a separate shell — variables do not persist between blocks.

Rules:

  • Use natural language for logic and state. Don't use shell variables to pass state between code blocks. Instead, tell Claude what to remember and reference it in prose (e.g., "the base branch detected in Step 0").
  • Don't hardcode branch names. Detect main/master/etc dynamically via gh pr view or gh repo view. Use {{BASE_BRANCH_DETECT}} for PR-targeting skills. Use "the base branch" in prose, <base> in code block placeholders.
  • Keep bash blocks self-contained. Each code block should work independently. If a block needs context from a previous step, restate it in the prose above.
  • Express conditionals as English. Instead of nested if/elif/else in bash, write numbered decision steps: "1. If X, do Y. 2. Otherwise, do Z."

Writing style (V1)

Default output from every tier-≥2 skill follows the Writing Style section in scripts/resolvers/preamble.ts: jargon glossed on first use (curated list in scripts/jargon-list.json, baked at gen-skill-docs time), questions framed in outcome terms ("what breaks for your users if...") not implementation terms, short sentences, decisions close with user impact. Power users who want the tighter V0 prose set gstack-config set explain_level terse (binary switch, no middle mode). See docs/designs/PLAN_TUNING_V1.md for the full design rationale. The review pacing overhaul that originally tried to ride alongside writing-style was extracted to V1.1 — see docs/designs/PACING_UPDATES_V0.md.

Browser interaction

When you need to interact with a browser (QA, dogfooding, cookie setup), use the /browse skill or run the browse binary directly via $B <command>. NEVER use mcp__claude-in-chrome__* tools — they are slow, unreliable, and not what this project uses.

Server / sidebar / extension internals: before editing browse/src/server.ts, extension/, the sidebar PTY, any SSE endpoint, or CDP session code, read docs/BROWSER_INTERNALS.md — sidebar message flow, WebSocket auth, tunnel dual-listener rules, Unicode sanitization at egress, SSE/CDP helpers, setup symlink hardening, and the sidebar security stack all live there, each pinned by a CI tripwire.

Egress receipts at every off-machine sink (v1.63.0.0+). Every gstack-initiated send off the machine MUST write a hash-chained receipt to ~/.gstack/security/egress.jsonl BEFORE the send: TypeScript callers use writeReceipt from lib/egress-receipt.ts; shell scripts source bin/gstack-egress-lib.sh and use _receipted_curl / _receipted_git. Failure polarity is per-class: fail-closed for sensitive sinks (brain-sync, memory-ingest, gbrain-sync, telemetry, ngrok tunnels, mcp-verify, supabase-provision), fail-open

  • stderr warning for user-facing ones (design OpenAI calls, update-check, dashboards, git-class ops). The new-sink scanner in test/egress-receipt-wiring.test.ts fails CI on an unreceipted curl / git push / fetch to a non-loopback host unless the file carries a reasoned entry in its SCANNER_EXEMPT list (user-directed page fetches, reachability probes, instruction strings, skill prose) — if you add a new off-machine sink, wire it through the helpers and add it to the enumerated sink list. Inspect with bin/gstack-egress (list | verify, exit 3 on tamper | grants). Threat model: forensic observability of ATTEMPTED egress, not an exfiltration control.

When developing gstack, .claude/skills/gstack may be a symlink back to this working directory (gitignored). This means skill changes are live immediately, great for rapid iteration, risky during big refactors where half-written skills could break other Claude Code sessions using gstack concurrently.

Check once per session: Run ls -la .claude/skills/gstack to see if it's a symlink or a real copy. If it's a symlink to your working directory, be aware that:

  • Template changes + bun run gen:skill-docs immediately affect all gstack invocations
  • Breaking changes to SKILL.md.tmpl files can break concurrent gstack sessions
  • During large refactors, remove the symlink (rm .claude/skills/gstack) so the global install at ~/.claude/skills/gstack/ is used instead

Prefix setting: Setup creates real directories (not symlinks) at the top level with a SKILL.md symlink inside (e.g., qa/SKILL.md -> gstack/qa/SKILL.md), plus links to each skill's runtime assets (sections/, templates, checklists — everything except SKILL.md, tests, build output, and .tmpl sources). Alias skills (_gstack-command, connect-chrome) install as rewritten copies, never symlinks. This ensures Claude discovers them as top-level skills, not nested under gstack/. Names are either short (qa) or namespaced (gstack-qa), controlled by skill_prefix in ~/.gstack/config.yaml. Pass --no-prefix or --prefix to skip the interactive prompt.

Note: Vendoring gstack into a project's repo is deprecated. Use global install

  • ./setup --team instead. See README.md for team mode instructions.

For plan reviews: When reviewing plans that modify skill templates or the gen-skill-docs pipeline, consider whether the changes should be tested in isolation before going live (especially if the user is actively using gstack in other windows).

Upgrade migrations: When a change modifies on-disk state (directory structure, config format, stale files) in ways that could break existing user installs, add a migration script to gstack-upgrade/migrations/. Read CONTRIBUTING.md's "Upgrade migrations" section for the format and testing requirements. The upgrade skill runs these automatically after ./setup during /gstack-upgrade.

Compiled binaries — never commit browse/dist/, design/dist/, or make-pdf/dist/

The browse/dist/, design/dist/, and make-pdf/dist/ directories contain compiled Bun binaries (browse, find-browse, design, ~62MB each). These are Mach-O arm64 only — they do NOT work on Linux, Windows, or Intel Macs. The ./setup script builds from source for every platform.

These directories are untracked and gitignored (.gitignore:3-6; the browse/dist/ binaries were untracked in 64d5a3e4, v0.11.16.0; the others were never tracked). They will NOT appear in git status. If a dist binary ever does show up in git status, something force-added it (git add -f) — do not commit it; unstage it and find out how it got there.

When staging files, always use specific filenames (git add file1 file2) — never git add . or git add -A, which can sweep in build outputs and junk.

Shared redaction engine catches credentials, PII, and legal/damaging content before it reaches an external sink (codex dispatch, GitHub issue/PR body, pushed commit). It is a guardrail, not airtight enforcementgit push --no-verify, direct gh issue create, and GSTACK_REDACT_PREPUSH=skip all bypass it. It catches accidents and carelessness, the 99% case. Do not claim it stops a determined leaker (a CHANGELOG line that does would fail a hostile screenshotter).

  • Engine + taxonomy: lib/redact-patterns.ts (the single source of truth — 3 tiers; HIGH = genuinely-secret credentials that block, MEDIUM = PII/legal/ internal + high-FP credential shapes that confirm via AskUserQuestion, LOW = FYI) and lib/redact-engine.ts (pure scan() + applyRedactions()). Calibration matters: a gate that cries wolf gets ignored, so context-variable shapes (Stripe pk_live_, Google AIza, JWT, env *_KEY=) sit at MEDIUM.
  • CLI: bin/gstack-redact (exit 0 clean / 2 MEDIUM / 3 HIGH; --json, --auto-redact, --repo-visibility, --from-file). bin/gstack-redact-prepush is the opt-in git hook.
  • Skill docs are generated from scripts/resolvers/redact-doc.ts ({{REDACT_INVOCATION_BLOCK:<sink>}}) so /spec, /cso, /ship, /document-release, /document-generate never drift from the engine.
  • Scan-at-sink: always scan the EXACT bytes that will be sent — write to a temp file, scan that file, pass the SAME file to gh/git. Never scan a string then re-render (that reopens a scan-vs-send gap).
  • Visibility (no tier promotion): resolve once per run, order = local config (gstack-config get redact_repo_visibility, ~/.gstack so never committed) → gh → glab → unknown(=public-strict). Public repos get STERNER per-finding confirmation (no batch-acknowledge, no silent-proceed); MEDIUM is never auto-promoted to HIGH.
  • Tool-attributed fences: wrap Codex/Greptile/eval output in ```codex-review / ```greptile fences so example credentials those tools quote WARN-degrade instead of blocking. A live-format credential inside the fence still blocks.
  • Config keys: redact_repo_visibility (public|private|unknown, local-only override for repos gh/glab can't read), redact_prepush_hook (true|false). There is intentionally NO key to disable HIGH blocking.
  • Audit: the /spec semantic pass appends a content-free record (categories + body sha256, no spec text) to ~/.gstack/security/semantic-reviews.jsonl (0600).

Commit style

Always bisect commits. Every commit should be a single logical change. When you've made multiple changes (e.g., a rename + a rewrite + new tests), split them into separate commits before pushing. Each commit should be independently understandable and revertable.

Examples of good bisection:

  • Rename/move separate from behavior changes
  • Test infrastructure (touchfiles, helpers) separate from test implementations
  • Template changes separate from generated file regeneration
  • Mechanical refactors separate from new features

When the user says "bisect commit" or "bisect and push," split staged/unstaged changes into logical commits and push.

Slop-scan: AI code quality, not AI code hiding

We use slop-scan to catch patterns where AI-generated code is genuinely worse than what a human would write. We are NOT trying to pass as human code. We are AI-coded and proud of it. The goal is code quality.

npx slop-scan scan .          # human-readable report
npx slop-scan scan . --json   # machine-readable for diffing

Config: slop-scan.config.json at repo root (currently excludes **/vendor/**).

Before fixing any finding, read docs/SLOP_SCAN.md: it separates genuine quality fixes (empty catches around file ops → safeUnlink(), process kills → safeKill()) from linter gaming we reject (string-matching error messages, tightening best-effort cleanup). Utilities live in browse/src/error-handling.ts. Don't chase the score.

Community PR guardrails

When reviewing or merging community PRs, always AskUserQuestion before accepting any commit that:

  1. Touches ETHOS.md — this file is Garry's personal builder philosophy. No edits from external contributors or AI agents, period.
  2. Removes or softens promotional material — YC references, founder perspective, and product voice are intentional. PRs that frame these as "unnecessary" or "too promotional" must be rejected.
  3. Changes Garry's voice — the tone, humor, directness, and perspective in skill templates, CHANGELOG, and docs are not generic. PRs that rewrite voice to be more "neutral" or "professional" must be rejected.

Even if the agent strongly believes a change improves the project, these three categories require explicit user approval via AskUserQuestion. No exceptions. No auto-merging. No "I'll just clean this up."

Checking out PRs from garrytan-agents

When the user says "check out " and the PR is from garrytan-agents/gstack (or any other fork that is NOT a collaborator on garrytan/gstack), do NOT just gh pr checkout. Fork PRs don't receive base-repo secrets (ANTHROPIC_API_KEY, OPENAI_API_KEY, etc.), so the eval/E2E CI jobs fail with empty-env auth errors regardless of what's set on the base repo.

Workflow: push the branch to garrytan/gstack (the base repo) and re-target the PR from there.

Concretely, after gh pr checkout <N>:

  1. Note the original PR number and head branch name.
  2. Push the same branch to the base repo: git push origin HEAD:<branch-name> (origin = garrytan/gstack, since the worktree is set up with that remote).
  3. Close the fork PR (gh pr close <N> --comment "moving to base-repo branch for secret access").
  4. Open a new PR from the base-repo branch: gh pr create --base main --head <branch-name>.
  5. New PR's workflows will get secrets automatically.

Why not fix it on the fork side? garrytan-agents isn't a collaborator on garrytan/gstack. Adding it as a collaborator (option A) or flipping the repo-wide "send secrets to fork PRs" toggle (option B) would let secrets reach fork PRs from anyone — broader blast radius than just moving this one branch. Option C (this section) keeps secret-distribution scope tight.

If the user asks you to skip the move (e.g., "just leave it as a fork PR"), respect that — eval CI will fail with empty-env auth, but check-freshness, workflow-lint, and windows-tests will still pass on the fork PR.

CHANGELOG + VERSION style

Versioning invariant (workspace-aware ship). VERSION is a monotonic ordered release identifier, not a strict semver commitment. The bump level (major/minor/patch/micro) expresses intent at ship time. Queue-advancing past a claimed version within the same bump level is explicitly permitted — if branch A claims v1.7.0.0 as a MINOR and branch B is also a MINOR, B lands at v1.8.0.0 (still a MINOR relative to main). Downstream consumers must NOT rely on "MINOR = feature-only, PATCH = fix-only" as a strict contract. This is why bin/gstack-next-version advances within the chosen bump level rather than repicking the level when collisions happen.

package.json carries the npm-valid translation, not VERSION verbatim. VERSION stays the 4-digit source of truth (e.g. 1.67.0.0); package.json and any subdirectory manifests with a version field get the 3-digit npm-valid translation (1.67.0), and lockfile version fields sync only when the lockfile already exists. bin/gstack-version-bump (via lib/version-source.ts) owns the translation and judges drift on translated forms — do NOT "fix" the apparent mismatch by hand, and do not write a 4-digit version into package.json (npm rejects it). Rationale and translation rules live in the lib/version-source.ts header; test/gstack-version-bump.test.ts pins the contract.

Scale-aware bumps — use common sense. When the diff is big, bump MINOR (or MAJOR), not PATCH. PATCH is for bug fixes and small additions; MINOR is for substantial new capability or substantial reduction; MAJOR is for breaking changes. Rough guideposts (don't treat as rules, treat as smell-checks):

  • PATCH (X.Y.Z+1.0): bug fix, doc tweak, small additive change, single test/file added. Net diff under ~500 lines, no new user-facing capability.
  • MINOR (X.Y+1.0.0): new capability shipped (skill, harness, command, big refactor), substantial code reduction (compression, migration), or coordinated multi-file change. Net diff over ~2000 lines added/removed, OR a user-visible feature you'd put in a tweet.
  • MAJOR (X+1.0.0.0): breaking change to public surface (CLI flag rename, skill removed, config format changed), OR a release big enough to be the headline of a blog post.

If you find yourself debating "is 10K added + 24K removed really a PATCH?" — it isn't. Bump MINOR. Same for "this adds a whole new test harness with 6 new E2E tests + helper utilities" — MINOR. The bump level is communication to the user about what kind of release this is; don't undersell it.

When merging origin/main brings a higher VERSION, re-evaluate the bump level against the SCALE of your branch's work, not just whether main moved forward. If main bumped MINOR and your branch is also a substantial change, you bump MINOR again on top (e.g., main at v1.14.0.0, your branch lands v1.15.0.0).

VERSION and CHANGELOG are branch-scoped. Every feature branch that ships gets its own version bump and CHANGELOG entry. The entry describes what THIS branch adds — not what was already on main.

The CHANGELOG entry is the diff between main and the shipping branch — what users get when they upgrade. NOT how the branch got there. A reader landing on the entry should learn what they can do now that they couldn't before; they should not learn about the branch's internal version bumps, the bugs we caught and fixed mid-branch, the plan reviews we ran, or the commits we squashed. That is branch development narrative. It belongs in PR descriptions and commit messages, not CHANGELOG.

Never reference branch-internal versions in a CHANGELOG entry. If your branch bumped VERSION from v1.5.0.0 → v1.5.1.0 → v1.6.0.0 during development and only the final v1.6.0.0 ships to main, the entry must read as if v1.5.1.0 never existed. Concretely, NEVER write:

  • "v1.5.1.0 had a bug that v1.6.0.0 fixes" — readers don't know about v1.5.1.0; it's a branch-internal artifact.
  • "The shipping headline of v1.5.1.0 was broken because..." — same reason. From main's perspective, v1.5.1.0 was never released.
  • "Pre-fix tests encoded the broken behavior" — that's a contributor's victory lap, not a user benefit.
  • "Two surgical edits, both in the dispatch path" — micro-narrative of the patch.

Instead, describe the released system: "Browser-skills run end-to-end with the expected tab-access semantics." If a property of the shipped system is worth calling out (e.g., "skill spawns get permissive tab access; pair-agent tunnel tokens require ownership"), document it as a property, not as a fix. The shipped system is what the user gets; the path to that system is invisible to them.

When to write the CHANGELOG entry:

  • At /ship time (Step 13), not during development or mid-branch.
  • The entry covers ALL commits on this branch vs the base branch.
  • Never fold new work into an existing CHANGELOG entry from a prior version that already landed on main. If main has v0.10.0.0 and your branch adds features, bump to v0.10.1.0 with a new entry — don't edit the v0.10.0.0 entry.

Key questions before writing:

  1. What branch am I on? What did THIS branch change?
  2. Is the base branch version already released? (If yes, bump and create new entry.)
  3. Does an existing entry on this branch already cover earlier work? (If yes, replace it with one unified entry for the final version.)

Merging main does NOT mean adopting main's version. When you merge origin/main into a feature branch, main may bring new CHANGELOG entries and a higher VERSION. Your branch still needs its OWN version bump on top. If main is at v0.13.8.0 and your branch adds features, bump to v0.13.9.0 with a new entry. Never jam your changes into an entry that already landed on main. Your entry goes on top because your branch lands next.

After merging main, always check:

  • Does CHANGELOG have your branch's own entry separate from main's entries?
  • Is VERSION higher than main's VERSION?
  • Is your entry the topmost entry in CHANGELOG (above main's latest)? If any answer is no, fix it before continuing.

After any CHANGELOG edit that moves, adds, or removes entries, immediately run grep "^## \[" CHANGELOG.md to verify no duplicates and a sensible reverse-chronological order. Gaps between version numbers are fine. A branch that ships at v1.6.4.0 without a prior v1.5.2.0 or v1.5.3.0 entry on main is correct — those were branch-internal version numbers that never landed. Do not back-fill gaps with placeholder entries.

Never orphan branch-internal versions. If your branch bumped VERSION several times during development (v1.5.1.0 → v1.5.2.0 → v1.6.4.0, say) and those earlier entries were never released to main, the final ship consolidates ALL of them into a single entry at the final version (v1.6.4.0). Collapse them — delete the old entries and move their content into the final entry, re-version table columns accordingly. Readers see one release, not a branch diary. Gaps are fine (v1.6.3.0 → v1.6.4.0 with no v1.5.x in between on main is correct).

CHANGELOG.md is for users, not contributors. Write it like product release notes:

  • Lead with what the user can now do that they couldn't before. Sell the feature.
  • Use plain language, not implementation details. "You can now..." not "Refactored the..."
  • Never mention TODOS.md, internal tracking, eval infrastructure, or contributor-facing details. These are invisible to users and meaningless to them.
  • Put contributor/internal changes in a separate "For contributors" section at the bottom.
  • Every entry should make someone think "oh nice, I want to try that."
  • No jargon: say "every question now tells you which project and branch you're in" not "AskUserQuestion format standardized across skill templates via preamble resolver."

Only document what shipped between main and this change. Readers do not care how we got here. Keep out of the CHANGELOG, always:

  • Branch resyncs, merge commits with main, rebase activity.
  • Plan approvals, review outcomes (CEO / eng / design / outside-voice / codex findings), AskUserQuestion decisions, scope negotiations.
  • "Work queued," "plan approved," "in-progress," "will ship later" — the CHANGELOG documents what DID ship, not what MIGHT ship.
  • Version-bump housekeeping when no user-facing work actually landed.

If the diff between the base branch version and this version has no user-facing change (only merges, only CHANGELOG edits, only placeholder work), the honest entry is one sentence: "Version bump for branch-ahead discipline. No user-facing changes yet." Stop there. Do not pad. Do not explain the plan that will ship eventually. Do not narrate the branch's history. When real work lands, the entry will replace this at /ship time.

Entry format

Every ## [X.Y.Z] entry starts with a release summary (two-line bold headline, lead paragraph, numbers table, closing paragraph) followed by an ### Itemized changes section. Read docs/CHANGELOG_STYLE.md for the full format spec and voice rules BEFORE writing an entry. Always credit community contributions with Contributed by @username.

AI effort compression

When estimating or discussing effort, always show both human-team and CC+gstack time:

Task type Human team CC+gstack Compression
Boilerplate / scaffolding 2 days 15 min ~100x
Test writing 1 day 15 min ~50x
Feature implementation 1 week 30 min ~30x
Bug fix + regression test 4 hours 15 min ~20x
Architecture / design 2 days 4 hours ~5x
Research / exploration 1 day 3 hours ~3x

Completeness is cheap. Don't recommend shortcuts when the complete implementation is achievable. Boil the ocean — the complete thing is the goal; only genuinely unrelated multi-quarter migrations are separate scope, never an excuse for a shortcut. See the Completeness Principle in the skill preamble for the full philosophy.

Search before building

Before designing any solution that involves concurrency, unfamiliar patterns, infrastructure, or anything where the runtime/framework might have a built-in:

  1. Search for "{runtime} {thing} built-in"
  2. Search for "{thing} best practice {current year}"
  3. Check official runtime/framework docs

Three layers of knowledge: tried-and-true (Layer 1), new-and-popular (Layer 2), first-principles (Layer 3). Prize Layer 3 above all. See ETHOS.md for the full builder philosophy.

Local plans

Contributors can store long-range vision docs and design documents in ~/.gstack-dev/plans/. These are local-only (not checked in). When reviewing TODOS.md, check plans/ for candidates that may be ready to promote to TODOs or implement.

E2E eval failure blame protocol

When an E2E eval fails during /ship or any other workflow, never claim "not related to our changes" without proving it. These systems have invisible couplings — a preamble text change affects agent behavior, a new helper changes timing, a regenerated SKILL.md shifts prompt context.

Required before attributing a failure to "pre-existing":

  1. Run the same eval on main (or base branch) and show it fails there too
  2. If it passes on main but fails on the branch — it IS your change. Trace the blame.
  3. If you can't run on main, say "unverified — may or may not be related" and flag it as a risk in the PR body

"Pre-existing" without receipts is a lazy claim. Prove it or don't say it.

Long-running tasks: don't give up

When running evals, E2E tests, or any long-running background task, poll until completion. Use sleep 180 && echo "ready" + TaskOutput in a loop every 3 minutes. Never switch to blocking mode and give up when the poll times out. Never say "I'll be notified when it completes" and stop checking — keep the loop going until the task finishes or the user tells you to stop.

The full E2E suite can take 30-45 minutes. That's 10-15 polling cycles. Do all of them. Report progress at each check (which tests passed, which are running, any failures so far). The user wants to see the run complete, not a promise that you'll check later.

Running evals as an agent: always detach (SIGTERM-proof)

When you (an agent/harness) launch a long eval/benchmark run, run it through bin/gstack-detach — NEVER as a plain backgrounded Bash task. A plain background task lives in the harness's process group, so a SIGTERM ("polite quit") on a turn boundary, a stopped Monitor, or an interruption kills the run mid-flight (observed: script "test:gate" was terminated by signal SIGTERM ~40 min into a run). On macOS the run can also die to idle-sleep. gstack-detach fixes both: a fresh session (escapes the group SIGTERM) wrapped in caffeinate -i (blocks idle-sleep).

  • Use the eval:bg* scripts (eval:bg, eval:bg:all, eval:bg:gate, eval:bg:periodic) — they wrap the eval command in gstack-detach with the machine-wide gstack-evals lock (concurrent worktrees serialize instead of saturating the shared model API), a per-tier watchdog, and a run-scoped log under ~/.gstack-dev/eval-runs/ (no shared-/tmp collision). Each prints its log path. eval:bg:gate / eval:bg:periodic run their tier through the sharded paid runner (scripts/test-paid-shards.ts, also exposed as test:gate:sharded / test:periodic:sharded): one Bun process per test file, an external wall-clock timeout that kills the shard's process GROUP (stray claude/codex grandchildren included), a per-shard GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/ honored by the EvalCollector constructor, and an aggregate that separates failed vs timed-out vs never-started shards — the detach timeouts (25200s gate / 37800s periodic; floor enforced against the live shard census by test/eval-detach-timeout-floor.test.ts) are sized against worst-case shard wall clock. EVALS_JOBS sets the shard process count (default 8); EVALS_CONCURRENCY is bun's --max-concurrency WITHIN a shard (default 2) — they are deliberately separate knobs. eval:list / eval:compare / eval:summary / eval:flake-rank read the shard dirs too. Or call gstack-detach [--lock NAME] [--timeout SECS] [--label LBL] -- <cmd> directly for any long agent job. Export ANTHROPIC_API_KEY first (never pass keys in argv).
  • Then poll the printed logfile with a death-aware watcher: break on the guaranteed ### gstack-detach EXIT=<code> ### sentinel (success AND failure are both marked, so silence is never mistaken for success). The detached run survives even if your watcher gets reaped, so re-checking the log always works.
  • Why the lock: a shared dev box with several Conductor worktrees will rate-limit the model API if two eval suites run at once (15-way concurrency each), which mass-times-out E2E tests. The lock makes the second run WAIT, not collide.
  • Humans running bun run test:evals foreground in their own terminal don't need this — Ctrl-C is intended there. Detachment is for agent-launched runs only.

E2E test fixtures: extract, don't copy

NEVER copy a full SKILL.md file into an E2E test fixture. SKILL.md files are 1500-2000 lines. When claude -p reads a file that large, context bloat causes timeouts, flaky turn limits, and tests that take 5-10x longer than necessary.

Instead, extract only the section the test actually needs:

// BAD — agent reads 1900 lines, burns tokens on irrelevant sections
fs.copyFileSync(path.join(ROOT, 'ship', 'SKILL.md'), path.join(dir, 'ship-SKILL.md'));

// GOOD — agent reads ~60 lines, finishes in 38s instead of timing out
const full = fs.readFileSync(path.join(ROOT, 'ship', 'SKILL.md'), 'utf-8');
const start = full.indexOf('## Review Readiness Dashboard');
const end = full.indexOf('\n---\n', start);
fs.writeFileSync(path.join(dir, 'ship-SKILL.md'), full.slice(start, end > start ? end : undefined));

Also when running targeted E2E tests to debug failures:

  • Run in foreground (bun test ...), not background with & and tee
  • Never pkill running eval processes and restart — you lose results and waste money
  • One clean run beats three killed-and-restarted runs

Publishing native OpenClaw skills to ClawHub

Native OpenClaw skills live in openclaw/skills/gstack-openclaw-*/SKILL.md. The command is clawhub publish (NOT clawhub skill publish) — full workflow, auth, and verification: docs/OPENCLAW_PUBLISHING.md.

Deploying to the active skill

The active skill lives at ~/.claude/skills/gstack/. After making changes:

  1. Push your branch
  2. Fetch and reset in the skill directory: cd ~/.claude/skills/gstack && git fetch origin && git reset --hard origin/main
  3. Rebuild: cd ~/.claude/skills/gstack && bun run build

If you use gbrain: the git reset --hard in step 2 reverts the brain-aware (GBRAIN_CONTEXT_LOAD / GBRAIN_SAVE_RESULTS) blocks that gstack-config gbrain-refresh renders into the install (those generated blocks differ from main by design). After deploying, re-run gstack-config gbrain-refresh to restore them across all your projects' Claude sessions. It's idempotent.

Or copy the binaries directly:

  • cp browse/dist/browse ~/.claude/skills/gstack/browse/dist/browse
  • cp design/dist/design ~/.claude/skills/gstack/design/dist/design

Skill routing

When the user's request matches an available skill, invoke it via the Skill tool. When in doubt, invoke the skill.

Key routing rules:

  • Product ideas/brainstorming → invoke /office-hours
  • Strategy/scope → invoke /plan-ceo-review
  • Architecture → invoke /plan-eng-review
  • Design system/plan review → invoke /design-consultation or /plan-design-review
  • Full review pipeline → invoke /autoplan
  • Bugs/errors → invoke /investigate
  • QA/testing site behavior → invoke /qa or /qa-only
  • Code review/diff check → invoke /review
  • Visual polish → invoke /design-review
  • Ship/deploy/PR → invoke /ship or /land-and-deploy
  • Save progress → invoke /context-save
  • Resume context → invoke /context-restore

Cross-session decision memory

Durable decisions and their rationale are captured in an append-only, event-sourced store at ~/.gstack/projects/<slug>/decisions.jsonl so neither you nor the user re-litigates a settled call or loses the "why" across sessions. This is the reliable, file-only path: it works with gbrain OFF. (gbrain semantic recall is an optional enhancement layered on top, never a dependency.)

  • Resurface active decisions before re-deciding: bin/gstack-decision-search (--recent N, --scope repo|branch|issue, --query KW, --all, --json). Add --semantic (with --query) to append related hits from gbrain memory when it's up; it degrades silently to the reliable file results when gbrain is off. Session start already surfaces scope-relevant active decisions via Context Recovery. If a decision is listed, treat it as settled with its rationale; if you're about to reverse it, say so explicitly.
  • Capture a DURABLE decision when you or the user make one: bin/gstack-decision-log '{"decision":"...","rationale":"...","scope":"repo|branch|issue","source":"user|skill|agent","confidence":1-10}'. Reverse a prior call with --supersede <id>; expunge an accidental secret with --redact <id>; rewrite the log to the active set with --compact. Non-interactive (never prompts), injection-sanitized, and HIGH-secret-blocking on write.
  • Durable means: architecture choice, scope cut, tool/vendor choice, or a reversal of a prior call. NOT a turn-level edit, a phrasing tweak, or anything trivially re-derivable. Capture is curated at the source — log durable decisions only, or the store becomes noise.

GBrain Search Guidance (configured by /sync-gbrain)

GBrain is set up and synced on this machine. The agent should prefer gbrain over Grep when the question is semantic or when you don't know the exact identifier yet.

This worktree is pinned to a worktree-scoped code source via the .gbrain-source file in the repo root (kubectl-style context). Any gbrain code-def, code-refs, code-callers, code-callees, or query call from anywhere under this worktree routes to that source by default — no --source flag needed. Conductor sibling worktrees of the same repo each have their own pin and their own indexed pages, so semantic results match the actual code on disk in this worktree.

Two indexed corpora available via the gbrain CLI:

  • This worktree's code (auto-pinned via .gbrain-source).
  • ~/.gstack/ curated memory (registered as gstack-brain-<user> source via the existing federation pipeline).

Prefer gbrain when:

  • "Where is X handled?" / semantic intent, no exact string yet: gbrain search "<terms>" or gbrain query "<question>"
  • "Where is symbol Y defined?" / symbol-based code questions: gbrain code-def <symbol> or gbrain code-refs <symbol>
  • "What calls Y?" / "What does Y depend on?": gbrain code-callers <symbol> / gbrain code-callees <symbol>
  • "What did we decide last time?" / past plans, retros, learnings: gbrain search "<terms>" --source gstack-brain-<user>

Grep is still right for known exact strings, regex, multiline patterns, and file globs. Run /sync-gbrain after meaningful code changes; for ongoing auto-sync across all worktrees, run gbrain autopilot --install once per machine — gbrain's daemon handles incremental refresh on a schedule.

Safety: don't run /sync-gbrain while gbrain autopilot is active — the orchestrator refuses destructive source ops when it detects a running autopilot to avoid racing it (#1734). Prefer registering user repos with gbrain sources add --path <dir> (no --url): URL-managed sources can auto-reclone, and the sync code walk for them requires an explicit --allow-reclone opt-in.