* fix(plan-tune): reject never-ask on one-way ids at --write --check already ignored those prefs; --write still stored them and --stats counted them as a working NEVER_ASK. Refuse the write and count leftover on-disk prefs as INERT_ONE_WAY. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix: gstack-config get returns "" with exit 0 for keys that have no default Skill preambles read configuration with VAR=$(gstack-config get <key> 2>/dev/null || echo "<default>") and that fallback only fires on a non-zero exit. lookup_default ended in a catch-all that echoed "" and returned 0, so for any key missing from the table VAR came back empty and the default written right there in the preamble was unreachable. The skill then branched on a value it never specified: "skip entirely if QUESTION_TUNING is false", reached with QUESTION_TUNING="". Four keys that skills actually read had no entry and took that path: question_tuning -> callers assume "false" repo_mode -> callers assume "unknown" team_mode -> callers assume "false" transcript_ingest_mode -> callers assume "off" Each default above is the value the call sites already substitute in their own `|| echo` fallback, so this only makes reachable what was already intended. The catch-all now returns non-zero. That is deliberately scoped to the unknown-key arm alone: keys whose default is intentionally empty still exit 0, because "" is their real answer and their callers depend on it -- cross_project_learnings ("unset triggers the first-time prompt"), redact_repo_visibility ("empty falls through to gh/glab detection"), salience_allowlist, user_slug_at_*. Making every empty answer an error would have broken those. test/gstack-config-defaults.test.ts pins the class rather than the four instances: it parses the case arms and asserts every `gstack-config get <key>` site in the tree is covered, so adding a read without a default fails CI. It also pins the exit-code contract in both directions. Verified failing against the pre-fix script, where it names exactly those four keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(redact): a typo'd subcommand no longer exits 0 having done nothing main() recognised exactly two subcommands and let everything else fall through to the stdin scan. On empty stdin that prints "(no findings)" and exits 0, so: $ gstack-redact install-prepush-hooks # plural typo gstack-redact scan — repo UNKNOWN (no findings) $ echo $? 0 No hook was installed, and the operator has every reason to believe the credential guard is armed. A guard that silently no-ops must never exit 0. Two smaller faults in the same dispatch, both of which lead people here: - There was no --help handler, so `gstack-redact --help` fell through to the scanner. Piping a credential to it scanned the secret and exited 3. - With no piped input and no --from-file, readInput() blocks on readSync(fd 0) until an EOF that an interactive terminal never sends. That prints nothing at all, so it reads as a hang rather than as "this is a filter, feed it". Now: --help/-h/help prints usage and exits 0; an unrecognised positional prints the offender and exits 1; a TTY with nothing piped in prints usage instead of blocking. "scan" stays accepted, because the human output header reads "gstack-redact scan — repo …" and that is what people type. Usage errors exit 1, deliberately not 2 or 3. Those mean MEDIUM and HIGH findings and callers gate dispatch on them, so a usage error exiting 2 would be read as "medium findings — prompt the user". A test pins that. Tests: 4 written failing first, then fixed. Full suite 7,722 pass / 0 fail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(browse): one ambiguous ref no longer kills the whole annotated screenshot `snapshot -a` exits 1 with "Selector matched multiple elements" on most real pages, so /qa, /canary and /land-and-deploy silently produce reports whose screenshots do not exist. Plain `screenshot <path>` is unaffected. Refs are built as getByRole(role, {name}) and disambiguated with .nth() when role+name repeats. That disambiguation cannot fire for a node with NO accessible name: the locator degrades to getByRole(role) with no name filter, and the count driving .nth() is taken from the FILTERED aria snapshot while getByRole matches the unfiltered DOM. Measured on a live page: the tree surfaced 2 unnamed paragraphs, the DOM had 9. Landmarks (banner/main/contentinfo) and paragraphs are correctly unnamed per ARIA, so this is the common case rather than an edge case. boundingBox() then hits Playwright strict mode, and the catch allowlisted only timeout/closed/Target/Execution-context messages — so the strict-mode error was re-thrown and aborted every remaining annotation. Two changes: - `.first()` before boundingBox(), so an ambiguous ref draws a box on its first match instead of aborting. The heatmap path below has always tolerated this via a bare `catch {}`; annotate was the only path that could be killed outright. - the catch no longer re-throws on unrecognised messages. A box we cannot measure is a box we do not draw, never a reason to lose the rest of the page. Set BROWSE_DEBUG to see what was skipped. Also: `-o` passed without `-a`/`-H` was silently ignored (exit 0, no file), which reads as "screenshots are broken" rather than "you forgot a flag". It now warns and points at `browse screenshot <path>`. Verified by rebuilding both ways against the same page with 51 refs present: before — "Selector matched multiple elements", no file written after — exit 0, 229KB PNG Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(version-bump): missing or empty VERSION no longer repairs a fabricated 0.0.0.0 into package.json repair now fails with exit 2 when the VERSION file is absent or empty instead of folding to DEFAULT ("0.0.0.0") — which passed VERSION_RE and regressed package.json below where it started. classify gains an additive versionFileExists field so /ship can tell a real 0.0.0.0 from a fabricated one. Re-derived from PR #2612 under the generated-file screening rule. Fixes #2600 (repair half; the path-configurability half landed in v1.67 via #2531). Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(memory-ingest): --probe counts post-attribution, through the same gate --bulk uses probeMode previously stat'd every walked file, so setup-gbrain gated its silent bulk ingest on pre-filter counts that the write path would never ingest (#2394). The attribution decision now lives in ONE shared gate (sessionIsAttributable — cheap-parse: cwd extraction + memoized resolveGitRemote, never a full page build) used by BOTH probeMode and preparePages, so the two stages' post-attribution counts are structurally identical. ProbeReport gains skipped_unattributed; the probe prints what it excluded and --include-unattributed restores raw counts. The parity is pinned at the prepare stage (probe post-attribution == transcripts reaching import), deliberately NOT == final written. Re-derived from PR #2612 under the generated-file screening rule; the shared-gate design and the remote memo are additions from the plan review. Fixes #2394. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(browse): allow CPU and network throttling for performance measurement Adds Emulation.setCPUThrottlingRate and Network.emulateNetworkConditions to CDP_ALLOWLIST. Motivation: diagnosing a real "uploads take 1-2 minutes" report, the only machine available was a fast developer workstation. Client-side processing measured 1.4s where the user experienced minutes, so the conclusion had to be reached arithmetically rather than observed. Throttling would have let the measurement reproduce the reporter's conditions directly. Both fit the existing posture rather than widening it: - Emulation already allows setDeviceMetricsOverride, clearDeviceMetricsOverride and setUserAgentOverride, which are equally mutating and scoped to the tab. - Neither method reads page content. setCPUThrottlingRate affects only timing; emulateNetworkConditions constrains traffic rather than inspecting it, so no request bodies, headers or cookies are exposed. Both are output: 'trusted' because they return no page-derived data. scope 'tab' for both, matching the surrounding Emulation entries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(session-update): lock pidfile records the live holder; hard TTL bounds every wedge (#2613) echo $$ inside the backgrounded subshell recorded the PARENT hook's PID — which exits immediately — so every subsequent session judged the lock stale and rm -rf'd a LIVE holder's lock, letting concurrent updaters run over each other. The pidfile now records ${BASHPID:-$(sh -c 'echo $PPID')} (macOS bash 3.2 has no BASHPID; the sh child's PPID is exactly this subshell). Staleness is now two independent detectors: PID liveness (as before, but against the real holder), and a 30-minute hard TTL on the heartbeat mtime — reclaimed regardless of kill -0, so a recycled PID or hung holder can't wedge the lock forever. The holder touches the pidfile after the pull and after setup, so a legitimately-slow run keeps itself alive. Empty and missing pidfiles are respected inside the TTL window (the mkdir→echo race) and reclaimed past it. Fixes #2613. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(browse): explicit windowsHide on every Bun.spawn site + census tripwire (#2575 residual) Bun.spawn sites were structurally outside the windowsHide census (it swept child_process bindings only). The runtime was already safe — native Bun hides consoles by default and bun-polyfill.cjs defaults windowsHide !== false since #2523/#2539 — but implicit defaults are exactly what regress silently. Every Bun.spawn/spawnSync in browse/src now carries the explicit flag (harmless on unix-only sites like Xvfb/xattr/open), and a second SWEEP in windows-spawn-hide.test.ts fails CI on any new flagless Bun.spawn site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain): brain worktree advances on the daily sync — no more silently stale brains (#2516) The daily pull refreshed only ~/.gstack itself, never the detached worktree at ~/.gstack-brain-worktree that gbrain actually indexes — so after setup the brain served stale pages forever unless setup-gbrain/sync-gbrain happened to run. brain-sync --once now advances the worktree once per 24h behind an ATTEMPT stamp (.brain-worktree-last-advance — a persistently-failing advance warns once a day, not at every skill boundary), inside the existing run lock and before any ingest step touches the worktree. The new gstack-gbrain-source-wireup --advance-only is built for the unattended cadence: git-only (no gbrain prereqs), pins every operation to the managed worktree (refuses paths that are not worktrees of the artifacts repo), refuses dirty worktrees, and never runs the force-remove recovery — a cron path must not be able to delete local changes. A static pin keeps the force-remove out. docs/gbrain-sync.md stops overclaiming the old cadence. Fixes #2516. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(memory-ingest): honor the per-remote deny/read-only trust policy (#2392) Transcript ingest now respects the same trust store as code import — the gate existed only in gstack-gbrain-sync's runCodeImport, so memory-ingest happily ingested transcripts from deny-listed repos. preparePages filters prepared transcript pages through ONE batch policy lookup (new 'get --batch' verb on bin/gstack-gbrain-repo-policy — the script owns URL normalization; the client adds repoPolicyTierBatch, one spawn for all distinct remotes, so large corpora never pay a 10s-timeout subprocess per remote). Outcomes match code-import semantics: read-only → clean skip (skipped_policy_readonly), deny → counted refusal (skipped_policy_deny), corrupted/unreadable store → HARD ERROR before any write (state, staging, egress receipt, and import all untouched) with the recovery command named — policy corruption must never read as successful ingestion. Artifacts are never policy-filtered (their git_remote is a project slug, not a remote). Fixes #2392. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): repo_mode keeps its empty no-default semantics (#2611 follow-up) The ported defaults table synthesized repo_mode → "unknown", but EMPTY is load-bearing for that key: gstack-repo-mode treats any non-empty answer as a user override and skips its own repo classification — the synthesized default turned the classifier into dead code (REPO_MODE=unknown everywhere; caught by test/gstack-repo-mode.test.ts via the wave's cross-agent blame protocol). repo_mode joins the empty-is-real carve-outs (empty output, exit 0). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): consent before killing a healthy headless daemon The pair-agent headed switch spawned 'connect --force-restart' unconditionally — auto-killing a live headless daemon (open tabs, cookies, logins) in direct contradiction of the iron rule it sits beside ('only an explicit --force-restart may kill a live daemon'). The CLI now captures daemon liveness BEFORE ensureServer (which can itself boot a fresh daemon) and relaunches only when the user passed --force-restart to pair-agent; otherwise it prints the tab count and continues against the existing daemon. The /pair-agent skill gains a matching one-way-door consent question (template half rides the wave's template block). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain-status): MCP scoping is per-project, and project-local beats user scope hasRemoteOnlyGbrainMcp scanned EVERY project's mcpServers in ~/.claude.json, so one project's remote gbrain registration reclassified broken local engines as thin-client machine-wide. It now reads user scope plus only the cwd's nearest-ancestor project key. The precedence itself was verified empirically and hermetically (fake HOME + CLAUDE_CONFIG_DIR fixtures, claude 2.1.233): with both scopes defining gbrain, 'claude mcp get gbrain' reports Scope: Local config — PROJECT-LOCAL WINS. Both in-repo consumers assumed the opposite; brain-cache's endpoint resolution flips to nearest-ancestor-project-first, and the stale user-first pin in brain-cache-roundtrip now pins the verified precedence. (The user-first jq in the brain-sync preamble resolver gets the same swap in the template block.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): gstack-slug matches remote-slug's owner-repo canonical form (live misfile bug) Found live during this wave's CEO review: bin/gstack-slug emitted SLUG=garrytan for this garrytan/gstack worktree while remote-slug correctly gave garrytan-gstack — decisions, timeline, ceo-plans, and learnings were filing into the wrong project store (observed polluting Context Recovery with another repo's decisions). Root cause: a stray empty ~/.git directory made the walk-up crown $HOME as the outermost project root; the remote lookup ran only against that root, failed silently, and the basename fallback cached 'garrytan' sticky. NOT worktree-specific — any strong marker on a non-repo ancestor triggered it. Fix: the walk now finds the outermost ancestor whose .git actually resolves an origin remote and derives owner-repo with remote-slug's byte-identical parse; marker-only ancestors keep anchoring the basename fallback but can no longer shadow a real remote. A new cache self-heal recomputes the poisoned shape (cached == basename of a marker root while a remote-bearing repo exists below), preserving legit #2212 stickiness. Nested-repo walk-up, no-remote and non-git fallbacks, and the SLUG=/BRANCH= eval contract are unchanged, pinned by a 10-case parity suite. Store migration for pre-fix data is tracked in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(brain-sync): per-record spool dir — the enqueue/drain race dies structurally Producers appended lines to .brain-queue.jsonl while the drain re-read and os.replace'd it; the in-code comment admitted a lockless append between the re-read and the replace was lost. Locks and rename-rotation designs were both reviewed and rejected (each retained a tail race); the shipped design is a maildir-style spool: one FILE per record in .brain-queue.d/ (tmp + atomic rename), the drain snapshots filenames, processes, and deletes exactly what it snapshotted. Writer and drainer never share an inode — nothing to race. Semantics: at-least-once (a crash between process and unlink re-drains; downstream content-hash dedup absorbs duplicates); retained (privacy-held) records keep their files; unparseable records are kept + warned, never destroyed. Legacy .brain-queue.jsonl migrates atomically on the next drain (crash-leftover .migrating files recovered too); status/drop-queue count both surfaces; discover-new writes spool records and advances its cursor per-record-written. The preamble's queue-depth line switches to spool count in this wave's template block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin-context): native slug fallback walks up like bash gstack-slug slugFromEnvironment derived the slug from the INNERMOST repo's origin while bash gstack-slug walks to the outermost project root — nested/vendored repos split their stores across the bash/native boundary (win32 hits the native path constantly). The native fallback now ports _outermost_project_root faithfully (strong/weak markers, outermost-strong-wins, 64-depth cap, fixed-point termination) plus the full resolution order: env override → walk-up → sticky cache with the #1125 self-heal → remote get-url → basename. Twelve mirrored scenarios drive BOTH implementations against the same fixtures and pin identical slugs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): git fallback queries the live remote, never mutates, and keeps 3-digit width The degraded path counted every remote-tracking ref on every remote — stale experiment branches and second remotes inflated version allocation, and a failed base read flipped 3-digit repos to 4-digit slots. Now: ls-remote --heads origin first (GIT_TERMINAL_PROMPT=0, 5s timeout, zero local ref mutation); on failure, local refs/remotes/origin ONLY with an explicit stale-refs warning; a failed base read zeroes at the LOCAL version file's width so a 3-digit repo allocates 0.0.1, not 0.0.1.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): hooks register the global-install path and re-point stale ones Registering hooks from a dev worktree baked that worktree's absolute path into settings.json — deleting the worktree left a dead hook erroring on every session stop, and the presence-only dedup (list-sources | grep) could never re-point it. setup's hook paths now route through _hook_install_path (global install preferred, source dir fallback), and the new ensure-event verb on gstack-settings-hook compares the registered command payload against canonical: identical → no write, different → single atomic replacement (never zero or two registrations). The plan-tune hooks had the same stale pattern and get the same fix without re-triggering their consent prompt. Also hardened: bun 1.3.13 turns an uncaught sync fs error in bun -e into a SILENT exit 0 — the registrar's write path now catches, prints, and exits 1, so a failed update can never report fake-green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): learnings capture is unconditional at completion (#2402) 43 of 44 learnings entries came from explicit /learn — the completion-status prose read 'if you discovered a durable project quirk... log it', which models treated as optional. The step now ALWAYS runs: review the session for durable learnings, log each one, and state 'No durable learnings this session' explicitly when the review comes up empty — an empty result, never a skipped step. Re-derived from PR #2612 under the generated-file screening rule. Fixes #2402. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(scrape): untrusted-content warning on the page-fetching skills (#2441) /scrape and /skillify consumed page content with zero injection guidance — the CHANGELOG claimed coverage the skills didn't have. The warning now lives in ONE exported const (UNTRUSTED_CONTENT_WARNING in resolvers/browse.ts), embedded in the browse COMMAND_REFERENCE as before AND injected standalone into both skills via the new {{UNTRUSTED_CONTENT_WARNING}} token — single source, wording can never drift between surfaces. Re-derived from PR #2612 under the generated-file screening rule. (Structural isolation for skillify-generated code is tracked as its own TODO.) Fixes #2441. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): checklist paths resolve from the installed skill root (#2518) /review Step 2 read .claude/skills/review/checklist.md — a path relative to the TARGET repo, which only resolves in gstack's own checkout. Every checklist/greptile-triage/TODOS-format reference (six across five templates — two more than the issue named, same class) now uses the installed-root form ~/.claude/skills/gstack/review/... that the templates' other references already use. The install-root class itself (non-default install dirs) is #1882, deliberately its own PR. Fixes #2518. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): one-way-door consent question before a daemon relaunch (template half) The skill flow now checks daemon liveness before Step 4 and asks an explicit one-way-door question (tabs/cookies/logins are lost) before passing --force-restart — never proceeding on a vague reply. Pairs with the CLI-half commit that stopped pair-agent auto-killing live daemons. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(codex): resume does not amortize the ~21K session prelude (#2387) Measured (#2387): every codex exec call pays Codex's session prelude, and a resumed call came in slightly ABOVE a fresh one — resume buys continuity, never token savings. The skill now says so where the resume flow lives: prefer one codex call per skill, batch questions into it. Fixes #2387. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(upgrade): fast-forward first; reset --hard only behind a proved-safe gate (#2517) /gstack-upgrade went straight to stash + reset --hard origin/main. Now it tries git pull --ff-only --autostash first (the same policy session-update's auto-upgrade uses). The destructive fallback runs unprompted ONLY when both git status --porcelain AND git rev-list origin/main..HEAD are empty — a clean tree with unpushed local commits is NOT safe, reset destroys them. Anything else requires an explicit one-way-door confirmation that lists every dirty file and unpushed commit being discarded. Fixes #2517. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): brain-sync block counts the spool queue and resolves MCP project-first Two resolver halves deferred from earlier wave commits: the queue-depth line counts .brain-queue.d/*.json spool records (plus legacy lines until the drain migrates them), and GBRAIN_MCP_ENTRY_JQ swaps its operands to nearest-ancestor-project-first — matching the empirically verified Claude Code precedence (project-local beats user scope) instead of the backwards user-first assumption. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + golden fixtures (single regen for the template block) Pure generator output for the six template/resolver commits above (learnings capture, untrusted-content warning, review paths, pair-agent consent, codex resume note, upgrade ff-only, brain-sync block) — bun run gen:skill-docs + --host codex + --host factory, with the three ship golden fixtures refreshed per the documented procedure. The three sidecar-path pins in gen-skill-docs.test.ts move to the new installed-root/$GSTACK_ROOT contract (#2518). Restores template freshness; full suite green from here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: TODOS.md — strike the six wave-fixed residuals, add two follow-ups The v1.67 adversarial-review residuals section shrinks to the one item the wave couldn't reach (iOS tap routing — needs real-device verification). New entries: skillify structural isolation (a prose warning is not a boundary for page-derived generated code) and the slug store migration (pre-fix sessions on stray-marker machines filed data under the degraded slug; post-fix reads go to the correct store, so history needs a merge/alias). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align cross-cutting pins with the wave's contracts Three suites pinned pre-wave behavior: browse's gstack-config test asserted the old unknown-key ''/exit-0 shape (#2611 made it exit 1); the Windows-paths suite pinned O_APPEND enqueue atomicity (the spool design satisfies the same invariant via tmp + os.replace, one file per record — pinned in its new form); and nine carve-guard skeleton ceilings absorbed the #2402 unconditional-learnings prose (~450B per skill), bumped with measured values per the guard's own protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: re-anchor the referenced-path scanner self-check to the gstack-rooted review refs The self-check pinned the review checklist as a class-1 alias-relative ref; #2518 moved those refs to the installed gstack root (class 2). The guard now proves the scanner sees them in their new class, so the class-2 assertion can't go vacuous. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: pin the wave's prose-tier behaviors (ship coverage-audit gap closure) The coverage audit found one regression-shaped gap: nothing pinned that the upgrade template's ff-only pull precedes the gated reset --hard (#2517) — a future template edit reverting to reset-first would fail nothing. Pinned: the ordering, the FF_OK gate, and the unpushed-commits check. Also pinned the two minor gaps: the {{UNTRUSTED_CONTENT_WARNING}} injection points in scrape/skillify (#2441) and brain-uninstall's spool-dir cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review round — 8 auto-fixes + 8 accepted findings hardened The ship review army (4 specialists + red-team + checklist, 29 findings) produced 8 mechanical auto-fixes and 11 decisions; the accepted set: - win32 slug parity completed: lib/bin-context.ts gains the remote-first outermost walk + degraded-cache self-heal the bash side got this wave — the two implementations now agree on the stray-marker live-bug shape, pinned by shared fixtures (multi-specialist 9/10 finding). - probe honors the plan's bounded-read decision: 256KB prefix, extraction semantics mirrored from parseTranscriptJsonl so probe/prepare can never diverge on the same file (>1MB transcript test). - policy normalize parity: bash normalize() now matches canonicalizeRemote on .git/-trailing and uppercase-.GIT shapes (7-shape corpus pinned two ways) — a deny for those shapes could previously slip the transcript gate. - session-update reclaim is TOCTOU-safe (atomic mv-aside on both branches). - settings-hook: unparseable settings.json errors instead of being replaced with {}; ensure-event keys on (event, source) so matcher changes update in place — never zero or two registrations. - dot-only slug guard at both parse sites (hostile 'url = ..' can't escape projects/); enqueue tmp-file janitor (1h TTL, inside the drain lock); brain-sync .migrating never clobbered; drop-queue/status count .migrating; snapshot -o warning correct + surfaced in diff mode; version-bump test order-dependence removed; uninstall clears the advance stamp. Deferred with record: slug heal-probe cost sentinel (P3 TODO), FF_OK conflation (noted, misdiagnosis-only). 270 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial round — the P0 finalize fail-safe and 12 hardened findings Three adversarial passes (Claude fresh-context, Codex chaos, Codex structured with P1 gate) on the full wave diff. Multi-source findings, all fixed: - P0: finalize_queue is now explicit-delete-only — a record is unlinked ONLY when classification proves it staged or dropped; a classifier crash, a missing class file, or a malformed pulled .brain-privacy-map.json (which previously nuked the whole snapshotted queue, remotely triggerable) now retains everything, warns, and re-drains next run. load_privacy_map treats corrupt maps as retain-all, never as empty. - next-version cannot silently drop a live claim: unreadable advertised refs get a targeted --depth=1 fetch + retry; still-unreadable claims surface as UNKNOWN warnings instead of duplicate-version silence. - session-update lock: ownership-checked EXIT trap (a TTL-reclaimed holder can no longer delete the new holder's lock) + a 5-min background heartbeat so a legitimately-slow pull/setup is never reclaimed while alive. - ensure-event collapses ALL same-(event,source) duplicates to one canonical entry; unique per-process tmp path; setup call sites surface (not swallow) the hardened refusals. - memory-ingest: --limit counts only policy-permitted pages (denied records no longer starve permitted ones); --probe applies the same policy filter as --bulk (skipped_policy_* fields on the report). - version-bump repair accepts a genuine literal 0.0.0.0 VERSION file. - slug heal restricted to the stray-.git shape — package.json-anchored wrapper roots keep their legit sticky identity (#2212 preserved). - brain-sync: idle fast path sees leftover .migrating records; unparseable spool records quarantine instead of warning forever; migration comment stops overclaiming the transition-window race. - CDP throttling justifications document override persistence (callers own restoration), pinned in the allowlist test. Deferred with record: deny retroactivity for already-ingested pages (P2 TODO, same semantics as the code-import gate); legacy-migration tail race (transition-window, requires pre-spool writers). 288 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + goldens (Windows-separator jq fix) Pure generator output for the brain-sync block's jq ancestor match now accepting backslash-formed Windows project keys — previously project-scoped brains were invisible on Windows while the TS scope resolvers saw them. Golden ship fixtures refreshed per the documented procedure. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: codex verify-pass residuals — chunked cwd read, post-filter partial count, migrating depth The verify re-review passed the P1 gate (0 P1s) and left three residuals, all applied: transcriptCwdFromPrefix reads in chunks until one complete record (4MB cap) so a giant first prompt can't truncate mid-JSON and break probe/bulk parity; partial_pages derives from the FINAL prepared set instead of the whole scanned corpus; the preamble queue-depth line counts leftover .brain-queue.jsonl.migrating records like the status path does (regen + goldens included). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.68.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.68.0.0 BROWSER.md: fix the $B cdp example (positional JSON params, not --json; depth is the real CDP param) and add the new perf-throttling examples (Emulation.setCPUThrottlingRate, Network.emulateNetworkConditions) with their clear-override counterparts. USING_GBRAIN_WITH_GSTACK.md: the state-files table row for the sync queue now names the maildir-style spool dir .brain-queue.d/ that replaced .brain-queue.jsonl this release. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align memory-pipeline probe pins with the #2394 stage-count contract The paid-tier E2E pinned the pre-fix contract (probe headline = raw discovered). Probe now counts post-attribution — the same gate --bulk uses — with an explicit unattributed-skip line. Adds the --include-unattributed companion pin so all 9 fixtures stay accounted for. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): batch missing-tip fetches — one bounded round trip, never a per-branch crawl The targeted-fetch retry for branches whose advertised tip has no local object ran ONE git fetch per branch (10s cap each). On a shallow clone against a busy remote that crawls the network for minutes — CI's shard deadline killed the free suite mid-file. Missing tips now collect into a single batched shallow fetch (15s cap); refs still missing after the batch (one unservable ref fails the whole transfer) get a capped per-branch retry, and anything past the cap warns as an UNKNOWN claim instead of fetching. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the batched fetch + make the offline-contract tests hermetic Two new G2 pins: N unfetched claim branches resolve with exactly ONE fetch spawn (PATH-shimmed git counts invocations), and one unservable ref no longer poisons the batch — live claims resolve via the bounded retry while only the ghost warns UNKNOWN. The #2545 offline-contract tests now run the CLI in a local fixture repo instead of the repo's own checkout: the checkout path did a live ls-remote against the real origin (operator-network-dependent, and the CI shard-deadline hang). The online-contract test gains a succeeding gh stub, so fallback:null is asserted deterministically instead of only when the operator happens to be authed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(redact-cli): derive the synthetic AWS-key fixture — no contiguous credential literal in source The CI quality gate scans every ADDED diff line with the redact engine, so the #2610 port's raw fixture literals failed the very gate they exist to test. The fixture is now assembled at runtime; the scanner still receives the identical bytes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the fixture's host via origin-URL sniff — kills the last environment dependence The hermetic offline-contract fixture had no origin remote, so detectHost() fell through to auth probes: a machine with glab authed passed via the gitlab path while a bare CI runner read host:unknown (offline stays false there) and failed. The fixture now pushes to a local bare origin at a path containing github.com — the URL sniff pins host:github identically everywhere, asserted explicitly in both tests, with every git call still local. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: benjamin beres <benjamin.beres@bienpreter.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Ricky <ricky@kinokostudio.com.hk> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: henbima <henbima@gmail.com>
58 KiB
Review Sections (after scope is agreed)
Anti-skip rule: Never condense, abbreviate, or skip any review section (1-4) regardless of plan type (strategy, spec, code, infra). Every section in this skill exists for a reason. "This is a strategy doc so implementation sections don't apply" is always wrong — implementation details are where strategy breaks down. If a section genuinely has zero findings, say "No issues found" and move on — but you must evaluate it.
Anti-shortcut clause: The plan file is the OUTPUT of the interactive review, not a substitute for it. Writing every finding into one plan write and calling ExitPlanMode without firing AskUserQuestion is the precise failure mode of the May 2026 transcript bug — the model explored, found issues, and dumped them into a deliverable rather than walking the user through them. If you have ANY non-trivial finding in any review section, the path from finding to ExitPlanMode goes THROUGH AskUserQuestion. Zero findings in every section is the only path to ExitPlanMode that bypasses AskUserQuestion. If you find yourself wanting to write a plan with findings before asking, stop and call AskUserQuestion now — that's the bug, recognize it.
Prior Learnings
Search for relevant learnings from previous sessions:
_CROSS_PROJ=$(~/.claude/skills/gstack/bin/gstack-config get cross_project_learnings 2>/dev/null || echo "unset")
echo "CROSS_PROJECT: $_CROSS_PROJ"
if [ "$_CROSS_PROJ" = "true" ]; then
~/.claude/skills/gstack/bin/gstack-learnings-search --limit 10 --cross-project 2>/dev/null || true
else
~/.claude/skills/gstack/bin/gstack-learnings-search --limit 10 2>/dev/null || true
fi
If CROSS_PROJECT is unset (first time): Use AskUserQuestion:
gstack can search learnings from your other projects on this machine to find patterns that might apply here. This stays local (no data leaves your machine). Recommended for solo developers. Skip if you work on multiple client codebases where cross-contamination would be a concern.
Options:
- A) Enable cross-project learnings (recommended)
- B) Keep learnings project-scoped only
If A: run ~/.claude/skills/gstack/bin/gstack-config set cross_project_learnings true
If B: run ~/.claude/skills/gstack/bin/gstack-config set cross_project_learnings false
Then re-run the search with the appropriate flag.
If learnings are found, incorporate them into your analysis. When a review finding matches a past learning, display:
"Prior learning applied: [key] (confidence N/10, from [date])"
This makes the compounding visible. The user should see that gstack is getting smarter on their codebase over time.
1. Architecture review
Evaluate:
- Overall system design and component boundaries.
- Dependency graph and coupling concerns.
- Data flow patterns and potential bottlenecks.
- Scaling characteristics and single points of failure.
- Security architecture (auth, data access, API boundaries).
- Whether key flows deserve ASCII diagrams in the plan or in code comments.
- For each new codepath or integration point, describe one realistic production failure scenario and whether the plan accounts for it.
- Distribution architecture: If this introduces a new artifact (binary, package, container), how does it get built, published, and updated? Is the CI/CD pipeline part of the plan or deferred?
For each issue found in this section, call AskUserQuestion individually. One issue per call. Present options, state your recommendation, explain WHY. Do NOT batch multiple issues into one AskUserQuestion. Use the preamble's AskUserQuestion Format section. The AskUserQuestion call is a tool_use, not prose — call the tool directly.
STOP. Do NOT proceed to the next review section, edit the plan file with the proposed fix, or call ExitPlanMode until the user responds. An issue with an "obvious fix" is still an issue and still needs explicit user approval before it lands in the plan. Loading the AskUserQuestion schema via ToolSearch and then writing the recommendation as chat prose is the failure mode this gate exists to prevent.
Confidence Calibration
Every finding MUST include a confidence score (1-10):
| Score | Meaning | Display rule |
|---|---|---|
| 9-10 | Verified by reading specific code. Concrete bug or exploit demonstrated. | Show normally |
| 7-8 | High confidence pattern match. Very likely correct. | Show normally |
| 5-6 | Moderate. Could be a false positive. | Show with caveat: "Medium confidence, verify this is actually an issue" |
| 3-4 | Low confidence. Pattern is suspicious but may be fine. | Suppress from main report. Include in appendix only. |
| 1-2 | Speculation. | Only report if severity would be P0. |
Finding format:
`[SEVERITY] (confidence: N/10) file:line — description`
Example: `[P1] (confidence: 9/10) app/models/user.rb:42 — SQL injection via string interpolation in where clause` `[P2] (confidence: 5/10) app/controllers/api/v1/users_controller.rb:18 — Possible N+1 query, verify with production logs`
Pre-emit verification gate (#1539 — kills the "field doesn't exist" FP class)
Before any finding is promoted to the report, the gate requires:
-
Quote the specific code line that motivates the finding — file:line plus the verbatim text of the line(s) that triggered it. If the finding is "field X doesn't exist on model Y", quote the lines of class Y where the field would live. If "dict.get() might return None", quote the dict initialization. If "race condition between A and B", quote both A and B.
-
If you cannot quote the motivating line(s), the finding is unverified. Force its confidence to 4-5 (suppressed from the main report). It still goes into the appendix so reviewers can audit calibration, but the user does NOT see it in the critical-pass output. Do not work around this by inventing speculative confidence 7+ — that defeats the gate.
Framework-meta nudge: When the symbol is generated by a framework
metaclass, descriptor, ORM Meta inner-class, or migration history (Django
Meta, Rails has_many/scope, SQLAlchemy relationship/Column,
TypeORM decorators, Sequelize init/belongsTo, Prisma generated client),
quote the meta-construct (the Meta block, the migration, the decorator,
the schema file) instead of expecting the literal name in the class body.
The verification is "I read the source that creates this symbol", not "I
grep'd for the name and didn't find it." Deeper framework-aware verification
(model introspection, migration-history-aware checks, ORM dialect detection)
is deliberately out of scope for the lighter gate — see the deferred
~/.gstack-dev/plans/1539-framework-aware-review.md design doc.
The FP classes the gate kills (measured against Django Sprint 2.5 #1539):
| FP class | Why the gate catches it |
|---|---|
| "field doesn't exist on model" | Requires quoting the model class body or Meta; the field's absence becomes obvious |
| "dict.get() might be None" | Requires quoting the dict initialization (e.g. Django form's cleaned_data is {}-initialized) |
| "save() might lose fields" | Requires quoting the ORM signature or model definition |
| "update_fields might miss X" | Requires quoting the field set; if X doesn't exist, the FP is self-evident |
Calibration learning: If you report a finding with confidence < 7 and the user confirms it IS a real issue, that is a calibration event. Your initial confidence was too low. Log the corrected pattern as a learning so future reviews catch it with higher confidence.
2. Code quality review
Evaluate:
- Code organization and module structure.
- DRY violations—be aggressive here.
- Error handling patterns and missing edge cases (call these out explicitly).
- Technical debt hotspots.
- Areas that are over-engineered or under-engineered relative to my preferences.
- Existing ASCII diagrams in touched files — are they still accurate after this change?
For each issue found in this section, call AskUserQuestion individually. One issue per call. Present options, state your recommendation, explain WHY. Do NOT batch multiple issues into one AskUserQuestion. Use the preamble's AskUserQuestion Format section. The AskUserQuestion call is a tool_use, not prose — call the tool directly.
STOP. Do NOT proceed to the next review section, edit the plan file with the proposed fix, or call ExitPlanMode until the user responds. An issue with an "obvious fix" is still an issue and still needs explicit user approval before it lands in the plan. Loading the AskUserQuestion schema via ToolSearch and then writing the recommendation as chat prose is the failure mode this gate exists to prevent.
3. Test review
100% coverage is the goal. Evaluate every codepath in the plan and ensure the plan includes tests for each one. If the plan is missing tests, add them — the plan should be complete enough that implementation includes full test coverage from the start.
Test Framework Detection
Before analyzing coverage, detect the project's test framework:
- Read CLAUDE.md — look for a
## Testingsection with test command and framework name. If found, use that as the authoritative source. - If CLAUDE.md has no testing section, auto-detect:
setopt +o nomatch 2>/dev/null || true # zsh compat
# Detect project runtime (markers are evidence, not commands to run blind)
[ -f manage.py ] && echo "RUNTIME:python FRAMEWORK:django"
{ [ -f pyproject.toml ] || [ -f pytest.ini ] || [ -f tox.ini ] || [ -f setup.cfg ] || [ -f requirements.txt ]; } && echo "RUNTIME:python"
[ -f Gemfile ] || [ -f Rakefile ] || [ -f .rspec ] && echo "RUNTIME:ruby"
[ -f package.json ] && echo "RUNTIME:node"
[ -f go.mod ] && echo "RUNTIME:go"
[ -f Cargo.toml ] && echo "RUNTIME:rust"
[ -f pom.xml ] && echo "RUNTIME:jvm BUILD:maven"
{ [ -f build.gradle ] || [ -f build.gradle.kts ]; } && echo "RUNTIME:jvm BUILD:gradle"
# Check for existing test infrastructure — config files, scripts, AND test files
ls jest.config.* vitest.config.* playwright.config.* cypress.config.* .rspec pytest.ini tox.ini phpunit.xml 2>/dev/null
[ -f package.json ] && grep -q '"test"[[:space:]]*:' package.json && echo "SCRIPT:package.json test"
[ -f Makefile ] && grep -qE '^(test|check):' Makefile && echo "TARGET:make test"
git ls-files | grep -cE '(^|/)(tests?|spec|__tests__)/|(^|/)tests?\.py$|(^|/)test_[^/]+\.py$|_test\.(go|py|rb|ts|js|exs)$|\.(test|spec)\.[jt]sx?$|_spec\.rb$|Test\.(java|kt)$' | sed 's/^/TESTFILES:/'
- If no framework detected: still produce the coverage diagram, but skip test generation.
Step 1. Trace every codepath in the plan:
Read the plan document. For each new feature, service, endpoint, or component described, trace how data will flow through the code — don't just list planned functions, actually follow the planned execution:
- Read the plan. For each planned component, understand what it does and how it connects to existing code.
- Trace data flow. Starting from each entry point (route handler, exported function, event listener, component render), follow the data through every branch:
- Where does input come from? (request params, props, database, API call)
- What transforms it? (validation, mapping, computation)
- Where does it go? (database write, API response, rendered output, side effect)
- What can go wrong at each step? (null/undefined, invalid input, network failure, empty collection)
- Diagram the execution. For each changed file, draw an ASCII diagram showing:
- Every function/method that was added or modified
- Every conditional branch (if/else, switch, ternary, guard clause, early return)
- Every error path (try/catch, rescue, error boundary, fallback)
- Every call to another function (trace into it — does IT have untested branches?)
- Every edge: what happens with null input? Empty array? Invalid type?
This is the critical step — you're building a map of every line of code that can execute differently based on input. Every branch in this diagram needs a test.
Step 2. Map user flows, interactions, and error states:
Code coverage isn't enough — you need to cover how real users interact with the changed code. For each changed feature, think through:
- User flows: What sequence of actions does a user take that touches this code? Map the full journey (e.g., "user clicks 'Pay' → form validates → API call → success/failure screen"). Each step in the journey needs a test.
- Interaction edge cases: What happens when the user does something unexpected?
- Double-click/rapid resubmit
- Navigate away mid-operation (back button, close tab, click another link)
- Submit with stale data (page sat open for 30 minutes, session expired)
- Slow connection (API takes 10 seconds — what does the user see?)
- Concurrent actions (two tabs, same form)
- Error states the user can see: For every error the code handles, what does the user actually experience?
- Is there a clear error message or a silent failure?
- Can the user recover (retry, go back, fix input) or are they stuck?
- What happens with no network? With a 500 from the API? With invalid data from the server?
- Empty/zero/boundary states: What does the UI show with zero results? With 10,000 results? With a single character input? With maximum-length input?
Add these to your diagram alongside the code branches. A user flow with no test is just as much a gap as an untested if/else.
Step 3. Check each branch against existing tests:
Go through your diagram branch by branch — both code paths AND user flows. For each one, search for a test that exercises it:
- Function
processPayment()→ look forbilling.test.ts,billing.spec.ts,test/billing_test.rb - An if/else → look for tests covering BOTH the true AND false path
- An error handler → look for a test that triggers that specific error condition
- A call to
helperFn()that has its own branches → those branches need tests too - A user flow → look for an integration or E2E test that walks through the journey
- An interaction edge case → look for a test that simulates the unexpected action
Quality scoring rubric:
- ★★★ Tests behavior with edge cases AND error paths
- ★★ Tests correct behavior, happy path only
- ★ Smoke test / existence check / trivial assertion (e.g., "it renders", "it doesn't throw")
E2E Test Decision Matrix
When checking each branch, also determine whether a unit test or E2E/integration test is the right tool:
RECOMMEND E2E (mark as [→E2E] in the diagram):
- Common user flow spanning 3+ components/services (e.g., signup → verify email → first login)
- Integration point where mocking hides real failures (e.g., API → queue → worker → DB)
- Auth/payment/data-destruction flows — too important to trust unit tests alone
RECOMMEND EVAL (mark as [→EVAL] in the diagram):
- Critical LLM call that needs a quality eval (e.g., prompt change → test output still meets quality bar)
- Changes to prompt templates, system instructions, or tool definitions
STICK WITH UNIT TESTS:
- Pure function with clear inputs/outputs
- Internal helper with no side effects
- Edge case of a single function (null input, empty array)
- Obscure/rare flow that isn't customer-facing
REGRESSION RULE (mandatory)
IRON RULE: When the coverage audit identifies a REGRESSION — code that previously worked but the diff broke — a regression test is added to the plan as a critical requirement. No AskUserQuestion. No skipping. Regressions are the highest-priority test because they prove something broke.
A regression is when:
- The diff modifies existing behavior (not new code)
- The existing test suite (if any) doesn't cover the changed path
- The change introduces a new failure mode for existing callers
When uncertain whether a change is a regression, err on the side of writing the test.
Step 4. Output ASCII coverage diagram:
Include BOTH code paths and user flows in the same diagram. Mark E2E-worthy and eval-worthy paths:
CODE PATHS USER FLOWS
[+] src/services/billing.ts [+] Payment checkout
├── processPayment() ├── [★★★ TESTED] Complete purchase — checkout.e2e.ts:15
│ ├── [★★★ TESTED] happy + declined + timeout ├── [GAP] [→E2E] Double-click submit
│ ├── [GAP] Network timeout └── [GAP] Navigate away mid-payment
│ └── [GAP] Invalid currency
└── refundPayment() [+] Error states
├── [★★ TESTED] Full refund — :89 ├── [★★ TESTED] Card declined message
└── [★ TESTED] Partial (non-throw only) — :101 └── [GAP] Network timeout UX
LLM integration: [GAP] [→EVAL] Prompt template change — needs eval test
COVERAGE: 5/13 paths tested (38%) | Code paths: 3/5 (60%) | User flows: 2/8 (25%)
QUALITY: ★★★:2 ★★:2 ★:1 | GAPS: 8 (2 E2E, 1 eval)
Legend: ★★★ behavior + edge + error | ★★ happy path | ★ smoke check [→E2E] = needs integration test | [→EVAL] = needs LLM eval
Fast path: All paths covered → "Test review: All new code paths have test coverage ✓" Continue.
Step 5. Add missing tests to the plan:
For each GAP identified in the diagram, add a test requirement to the plan. Be specific:
- What test file to create (match existing naming conventions)
- What the test should assert (specific inputs → expected outputs/behavior)
- Whether it's a unit test, E2E test, or eval (use the decision matrix)
- For regressions: flag as CRITICAL and explain what broke
The plan should be complete enough that when implementation begins, every test is written alongside the feature code — not deferred to a follow-up.
Test Plan Artifact
After producing the coverage diagram, write a test plan artifact to the project directory so /qa and /qa-only can consume it as primary test input:
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)" && mkdir -p ~/.gstack/projects/$SLUG
USER=$(whoami)
DATETIME=$(date +%Y%m%d-%H%M%S)
Write to ~/.gstack/projects/{slug}/{user}-{branch}-eng-review-test-plan-{datetime}.md:
# Test Plan
Generated by /plan-eng-review on {date}
Branch: {branch}
Repo: {owner/repo}
## Affected Pages/Routes
- {URL path} — {what to test and why}
## Key Interactions to Verify
- {interaction description} on {page}
## Edge Cases
- {edge case} on {page}
## Critical Paths
- {end-to-end flow that must work}
This file is consumed by /qa and /qa-only as primary test input. Include only the information that helps a QA tester know what to test and where — not implementation details.
For LLM/prompt changes: check the "Prompt/LLM changes" file patterns listed in CLAUDE.md. If this plan touches ANY of those patterns, state which eval suites must be run, which cases should be added, and what baselines to compare against. Then use AskUserQuestion to confirm the eval scope with the user.
For each issue found in this section, call AskUserQuestion individually. One issue per call. Present options, state your recommendation, explain WHY. Do NOT batch multiple issues into one AskUserQuestion. Use the preamble's AskUserQuestion Format section. The AskUserQuestion call is a tool_use, not prose — call the tool directly.
STOP. Do NOT proceed to the next review section, edit the plan file with the proposed fix, or call ExitPlanMode until the user responds. An issue with an "obvious fix" is still an issue and still needs explicit user approval before it lands in the plan. Loading the AskUserQuestion schema via ToolSearch and then writing the recommendation as chat prose is the failure mode this gate exists to prevent.
4. Performance review
Evaluate:
- N+1 queries and database access patterns.
- Memory-usage concerns.
- Caching opportunities.
- Slow or high-complexity code paths.
For each issue found in this section, call AskUserQuestion individually. One issue per call. Present options, state your recommendation, explain WHY. Do NOT batch multiple issues into one AskUserQuestion. Use the preamble's AskUserQuestion Format section. The AskUserQuestion call is a tool_use, not prose — call the tool directly.
STOP. Do NOT proceed to the next review section, edit the plan file with the proposed fix, or call ExitPlanMode until the user responds. An issue with an "obvious fix" is still an issue and still needs explicit user approval before it lands in the plan. Loading the AskUserQuestion schema via ToolSearch and then writing the recommendation as chat prose is the failure mode this gate exists to prevent.
Outside Voice — Independent Plan Challenge (default-on)
After all review sections are complete, run an independent second opinion from a
different AI system automatically — it is a standard part of plan review, not an
opt-in. Two models agreeing on a plan is stronger signal than one model's thorough
review. The user turns this off only by asking explicitly
(gstack-config set codex_reviews disabled).
Preflight — decide whether and how the outside voice runs:
# Codex preflight: one block (functions sourced here don't persist to later blocks).
_TEL=$(~/.claude/skills/gstack/bin/gstack-config get telemetry 2>/dev/null || echo off)
_CODEX_CFG=$(~/.claude/skills/gstack/bin/gstack-config get codex_reviews 2>/dev/null || echo enabled)
source ~/.claude/skills/gstack/bin/gstack-codex-probe 2>/dev/null || true
if [ "$_CODEX_CFG" = "disabled" ]; then
_CODEX_MODE="disabled"
# Running-under-Codex presence probe (#2519): a live Codex session exports
# CODEX_THREAD_ID / CODEX_SANDBOX into every shell it spawns (verified
# against a live `codex exec 'env | grep -i codex'` capture, codex 0.147.0).
# Nested codex spawns from inside a Codex host multiply token burn
# (observed: one /review = 15M tokens). GSTACK_FORCE_CODEX_REVIEW=1 forces
# the nested passes anyway.
elif [ "${GSTACK_FORCE_CODEX_REVIEW:-0}" != "1" ] && { [ -n "${CODEX_THREAD_ID:-}" ] || [ -n "${CODEX_SANDBOX:-}" ]; }; then
_CODEX_MODE="under_codex"
elif ! command -v codex >/dev/null 2>&1; then
_CODEX_MODE="not_installed"; _gstack_codex_log_event "codex_cli_missing" 2>/dev/null || true
elif ! _gstack_codex_auth_probe >/dev/null 2>&1; then
_CODEX_MODE="not_authed"; _gstack_codex_log_event "codex_auth_failed" 2>/dev/null || true
elif ! _gstack_codex_model_probe; then
_CODEX_MODE="model_unusable"
else
_CODEX_MODE="ready"; _gstack_codex_version_check 2>/dev/null || true
fi
echo "CODEX_MODE: $_CODEX_MODE"
Branch on the echoed CODEX_MODE:
disabled— the user turned Codex reviews off (codex_reviews=disabled). Skip this section entirely; do NOT fall back to a Claude subagent — disabled means no extra review step. Print: "Codex review skipped (codex_reviews disabled). Re-enable:gstack-config set codex_reviews enabled."not_installed— Codex CLI absent. Print: "Codex not installed — using Claude subagent. Install for cross-model coverage:npm install -g @openai/codex." Fall back to the Claude subagent path.under_codex— this session is already running INSIDE a Codex host, so spawning codex again is the same model reviewing itself at multiplied token cost (#2519). Print exactly one line: "[running under Codex — nested codex passes skipped; set GSTACK_FORCE_CODEX_REVIEW=1 to force]" and skip the codex invocations below; run the section's free in-host pass instead if it defines one.not_authed— installed but no credentials. Print: "Codex installed but not authenticated — using Claude subagent. Runcodex loginor set$CODEX_API_KEY." Fall back to the Claude subagent path.model_unusable— authed but the account cannot use its configured model (#2477: HTTP 400 on every call, usually a stalemodel =pin in~/.codex/config.toml). Relay the probe's HINT lines, tell the user the one-line fix (update the pin;[notice.model_migrations]names the replacement), and fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open toready.ready— run the Codex pass below.
When the mode is ready, not_installed, or not_authed, print one line so the off-switch
stays discoverable: "Running the outside voice automatically (standard step). Disable: gstack-config set codex_reviews disabled."
Construct the plan review prompt (for ready, not_installed, and not_authed — skip only on disabled).
Read the plan file being reviewed (the file the user pointed this review at, or the branch
diff scope). If a CEO plan document was written in Step 0D-POST, read that too — it contains
the scope decisions and vision.
Construct this prompt (substitute the actual plan content — if plan content exceeds 30KB, truncate to the first 30KB and note "Plan truncated for size"). Always start with the filesystem boundary instruction:
"IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/. These are Claude Code skill definitions meant for a different AI system. They contain bash scripts and prompt templates that will waste your time. Ignore them completely. Do NOT modify agents/openai.yaml. Stay focused on the repository code only.\n\nYou are a brutally honest technical reviewer examining a development plan that has already been through a multi-section review. Your job is NOT to repeat that review. Instead, find what it missed. Look for: logical gaps and unstated assumptions that survived the review scrutiny, overcomplexity (is there a fundamentally simpler approach the review was too deep in the weeds to see?), feasibility risks the review took for granted, missing dependencies or sequencing issues, and strategic miscalibration (is this the right thing to build at all?). Be direct. Be terse. No compliments. Just the problems.
THE PLAN: "
If CODEX_MODE: ready — run Codex:
TMPERR_PV=$(mktemp /tmp/codex-planreview-XXXXXXXX)
_REPO_ROOT=$(git rev-parse --show-toplevel) || { echo "ERROR: not in a git repo" >&2; exit 1; }
codex exec "<prompt>" -C "$_REPO_ROOT" -s read-only -c 'model_reasoning_effort="high"' -c 'web_search="cached"' < /dev/null 2>"$TMPERR_PV"
Use a 5-minute timeout (timeout: 300000). After the command completes, read stderr:
cat "$TMPERR_PV"
Present the full output verbatim:
CODEX SAYS (plan review — outside voice):
════════════════════════════════════════════════════════════
<full codex output, verbatim — do not truncate or summarize>
════════════════════════════════════════════════════════════
Error handling: All errors are non-blocking — the outside voice is informational.
- Auth failure (stderr contains "auth", "login", "unauthorized"): "Codex auth failed. Run `codex login` to authenticate." Fall back to the Claude subagent below.
- Timeout: "Codex timed out after 5 minutes." Fall back to the Claude subagent below.
- Empty response: "Codex returned no response." Fall back to the Claude subagent below.
If CODEX_MODE: not_installed or not_authed (or Codex errored at runtime):
Dispatch via the Agent tool. The subagent has fresh context — genuine independence. Bound it the same way as Codex: cap the dispatch at a 5-minute timeout so "never blocking" is also "never hanging."
Subagent prompt: same plan review prompt as above.
Present findings under an OUTSIDE VOICE (Claude subagent): header.
If the subagent fails or times out: "Outside voice unavailable. Continuing to outputs."
(On CODEX_MODE: disabled you already skipped this section per the preflight — do not reach here.)
Cross-model tension:
After presenting the outside voice findings, note any points where the outside voice disagrees with the review findings from earlier sections. Flag these as:
CROSS-MODEL TENSION:
[Topic]: Review said X. Outside voice says Y. [Present both perspectives neutrally.
State what context you might be missing that would change the answer.]
User Sovereignty: Do NOT auto-incorporate outside voice recommendations into the plan. Present each tension point to the user. The user decides. Cross-model agreement is a strong signal — present it as such — but it is NOT permission to act. You may state which argument you find more compelling, but you MUST NOT apply the change without explicit user approval.
For each substantive tension point, use AskUserQuestion:
"Cross-model disagreement on [topic]. The review found [X] but the outside voice argues [Y]. [One sentence on what context you might be missing.]"
RECOMMENDATION: Choose [A or B] because [one-line reason explaining which argument is more compelling and why]. Completeness: A=X/10, B=Y/10.
Options:
- A) Accept the outside voice's recommendation (I'll apply this change)
- B) Keep the current approach (reject the outside voice)
- C) Investigate further before deciding
- D) Add to TODOS.md for later
Wait for the user's response. Do NOT default to accepting because you agree with the outside voice. If the user chooses B, the current approach stands — do not re-argue.
If no tension points exist, note: "No cross-model tension — both reviewers agree."
Persist the result:
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"codex-plan-review","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","status":"STATUS","source":"SOURCE","commit":"'"$(git rev-parse --short HEAD)"'"}'
Substitute: STATUS = "clean" if no findings, "issues_found" if findings exist. SOURCE = "codex" if Codex ran, "claude" if subagent ran.
Cleanup: Run rm -f "$TMPERR_PV" after processing (if Codex was used).
Outside Voice Integration Rule
Outside voice findings are INFORMATIONAL until the user explicitly approves each one. Do NOT incorporate outside voice recommendations into the plan without presenting each finding via AskUserQuestion and getting explicit approval. This applies even when you agree with the outside voice. Cross-model consensus is a strong signal — present it as such — but the user makes the decision.
CRITICAL RULE — How to ask questions
Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:
- One issue = one AskUserQuestion call. Never combine multiple issues into one question.
- Describe the problem concretely, with file and line references.
- Present 2-3 options, including "do nothing" where that's reasonable.
- For each option, specify in one line: effort (human: ~X / CC: ~Y), risk, and maintenance burden. If the complete option is only marginally more effort than the shortcut with CC, recommend the complete option.
- Map the reasoning to my engineering preferences above. One sentence connecting your recommendation to a specific preference (DRY, explicit > clever, minimal diff, etc.).
- Label with issue NUMBER + option LETTER (e.g., "3A", "3B").
- Coverage vs kind: for every per-issue AskUserQuestion you raise in this review, decide whether the options differ in coverage or in kind. If coverage (e.g., more tests vs fewer, complete error handling vs happy-path-only, full edge-case coverage vs shortcut), include
Completeness: N/10on each option. If kind (e.g., architectural choice between two different systems, posture-over-posture, A/B/C where each is a different kind of thing), skip the score and add one line:Note: options differ in kind, not coverage — no completeness score.Do NOT fabricate scores on kind-differentiated questions — filler scores are worse than no score. - Zero findings: if a section has zero findings, state "No issues, moving on" and proceed. Otherwise, use AskUserQuestion for each finding — a finding with an "obvious fix" is still a finding and still needs user approval before any change lands in the plan.
Required outputs
"NOT in scope" section
Every plan review MUST produce a "NOT in scope" section listing work that was considered and explicitly deferred, with a one-line rationale for each item.
"What already exists" section
List existing code/flows that already partially solve sub-problems in this plan, and whether the plan reuses them or unnecessarily rebuilds them.
TODOS.md updates
After all review sections are complete, present each potential TODO as its own individual AskUserQuestion. Never batch TODOs — one per question. Never silently skip this step. Follow the format in ~/.claude/skills/gstack/review/TODOS-format.md.
For each TODO, describe:
- What: One-line description of the work.
- Why: The concrete problem it solves or value it unlocks.
- Pros: What you gain by doing this work.
- Cons: Cost, complexity, or risks of doing it.
- Context: Enough detail that someone picking this up in 3 months understands the motivation, the current state, and where to start.
- Depends on / blocked by: Any prerequisites or ordering constraints.
Then present options: A) Add to TODOS.md B) Skip — not valuable enough C) Build it now in this PR instead of deferring.
Do NOT just append vague bullet points. A TODO without context is worse than no TODO — it creates false confidence that the idea was captured while actually losing the reasoning.
Diagrams
The plan itself should use ASCII diagrams for any non-trivial data flow, state machine, or processing pipeline. Additionally, identify which files in the implementation should get inline ASCII diagram comments — particularly Models with complex state transitions, Services with multi-step pipelines, and Concerns with non-obvious mixin behavior.
Failure modes
For each new codepath identified in the test review diagram, list one realistic way it could fail in production (timeout, nil reference, race condition, stale data, etc.) and whether:
- A test covers that failure
- Error handling exists for it
- The user would see a clear error or a silent failure
If any failure mode has no test AND no error handling AND would be silent, flag it as a critical gap.
Worktree parallelization strategy
Analyze the plan's implementation steps for parallel execution opportunities. This helps the user split work across git worktrees (via Claude Code's Agent tool with isolation: "worktree" or parallel workspaces).
Skip if: all steps touch the same primary module, or the plan has fewer than 2 independent workstreams. In that case, write: "Sequential implementation, no parallelization opportunity."
Otherwise, produce:
- Dependency table — for each implementation step/workstream:
| Step | Modules touched | Depends on |
|---|---|---|
| (step name) | (directories/modules, NOT specific files) | (other steps, or —) |
Work at the module/directory level, not file level. Plans describe intent ("add API endpoints"), not specific files. Module-level ("controllers/, models/") is reliable; file-level is guesswork.
- Parallel lanes — group steps into lanes:
- Steps with no shared modules and no dependency go in separate lanes (parallel)
- Steps sharing a module directory go in the same lane (sequential)
- Steps depending on other steps go in later lanes
Format: Lane A: step1 → step2 (sequential, shared models/) / Lane B: step3 (independent)
-
Execution order — which lanes launch in parallel, which wait. Example: "Launch A + B in parallel worktrees. Merge both. Then C."
-
Conflict flags — if two parallel lanes touch the same module directory, flag it: "Lanes X and Y both touch module/ — potential merge conflict. Consider sequential execution or careful coordination."
Implementation Tasks
Before closing this review, synthesize the findings above into a flat list of
build-actionable tasks. Each task derives from a specific finding — no padding.
Emit the markdown section AND write a JSONL artifact that /autoplan can
aggregate across phases.
Markdown section (always emit)
## Implementation Tasks
Synthesized from this review's findings. Each task derives from a specific
finding above. Run with Claude Code or Codex; checkbox as you ship.
- [ ] **T1 (P1, human: ~2h / CC: ~15min)** — <component> — <imperative title>
- Surfaced by: <section name> — <specific finding text or line reference>
- Files: <paths to touch>
- Verify: <test command or manual check>
- [ ] **T2 (P2, human: ~30min / CC: ~5min)** — ...
Rules:
- P1 blocks ship; P2 should land same branch; P3 is a follow-up TODO.
- If a finding produced no actionable task, do not invent one.
- If a section had zero findings, emit
_No new tasks from <section>._ - Effort uses the AI-compression table from CLAUDE.md.
JSONL artifact (always write, even if zero tasks)
/autoplan reads this file to aggregate across phases. Build each line with
jq -nc so titles and source findings containing quotes, newlines, or
backslashes serialize cleanly — never use hand-rolled echo / printf.
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)"
TASKS_DIR="${HOME}/.gstack/projects/${SLUG:-unknown}"
mkdir -p "$TASKS_DIR"
TASKS_FILE="$TASKS_DIR/tasks-eng-review-$(date +%Y%m%d-%H%M%S).jsonl"
COMMIT=$(git rev-parse HEAD 2>/dev/null || echo unknown)
BRANCH=$(git branch --show-current 2>/dev/null || echo unknown)
RUN_ID="$(date -u +%Y%m%dT%H%M%SZ)-$$"
# Repeat ONE jq invocation per task identified during this review.
# Substitute the placeholders inline with shell variables you set per task:
# TASK_ID (T1, T2, ...), PRIORITY (P1/P2/P3), COMPONENT, TITLE,
# SOURCE_FINDING, EFFORT_HUMAN, EFFORT_CC, FILES_JSON (a JSON array literal
# like '["browse/src/sanitize.ts","browse/src/server.ts"]').
jq -nc \
--arg phase 'eng-review' \
--arg run_id "$RUN_ID" \
--arg branch "$BRANCH" \
--arg commit "$COMMIT" \
--arg id "$TASK_ID" \
--arg priority "$PRIORITY" \
--arg component "$COMPONENT" \
--arg effort_human "$EFFORT_HUMAN" \
--arg effort_cc "$EFFORT_CC" \
--arg title "$TITLE" \
--arg source_finding "$SOURCE_FINDING" \
--argjson files "$FILES_JSON" \
'{phase:$phase, run_id:$run_id, branch:$branch, commit:$commit, id:$id, priority:$priority, component:$component, files:$files, effort_human:$effort_human, effort_cc:$effort_cc, title:$title, source_finding:$source_finding}' \
>> "$TASKS_FILE"
If jq is not installed, fall back to skipping the JSONL write and warn
the user to install jq for autoplan aggregation. Never hand-roll JSONL.
If zero tasks were identified in this review, still touch the JSONL file
(: > "$TASKS_FILE") so the aggregator sees that the phase produced output
this run (an empty file means "ran, no findings" — distinct from "didn't run").
Completion summary
At the end of the review, fill in and display this summary so the user can see all findings at a glance:
- Step 0: Scope Challenge — ___ (scope accepted as-is / scope reduced per recommendation)
- Architecture Review: ___ issues found
- Code Quality Review: ___ issues found
- Test Review: diagram produced, ___ gaps identified
- Performance Review: ___ issues found
- NOT in scope: written
- What already exists: written
- TODOS.md updates: ___ items proposed to user
- Failure modes: ___ critical gaps flagged
- Outside voice: ran (codex/claude) / skipped
- Parallelization: ___ lanes, ___ parallel / ___ sequential
- Lake Score: X/Y recommendations chose complete option
Retrospective learning
Check the git log for this branch. If there are prior commits suggesting a previous review cycle (e.g., review-driven refactors, reverted changes), note what was changed and whether the current plan touches the same areas. Be more aggressive reviewing areas that were previously problematic.
Formatting rules
- NUMBER issues (1, 2, 3...) and LETTERS for options (A, B, C...).
- Label with NUMBER + LETTER (e.g., "3A", "3B").
- One sentence max per option. Pick in under 5 seconds.
- After each review section, pause and ask for feedback before moving on.
Review Log
After producing the Completion Summary above, persist the review result.
PLAN MODE EXCEPTION — ALWAYS RUN: This command writes review metadata to
~/.gstack/ (user config directory, not project files). The skill preamble
already writes to ~/.gstack/sessions/ and ~/.gstack/analytics/ — this is
the same pattern. The review dashboard depends on this data. Skipping this
command breaks the review readiness dashboard in /ship.
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"plan-eng-review","timestamp":"TIMESTAMP","status":"STATUS","unresolved":N,"critical_gaps":N,"issues_found":N,"mode":"MODE","commit":"COMMIT"}'
~/.claude/skills/gstack/bin/gstack-decision-log '{"decision":"Eng review (MODE): ARCH_SUMMARY","rationale":"KEY_DECISION","scope":"branch","source":"skill","confidence":8}' 2>/dev/null || true
The second command records the architecture verdict as a durable cross-session decision (so a future session inherits the chosen approach and what was hardened, not just the count). Same ~/.gstack/ write pattern as review-log, non-interactive, best-effort (|| true). Substitute ARCH_SUMMARY (e.g. "N findings, all folded" or "M unresolved") and KEY_DECISION (the load-bearing architecture call from the report, one line — omit if the review found nothing durable).
Substitute values from the Completion Summary:
- TIMESTAMP: current ISO 8601 datetime
- STATUS: "clean" if 0 unresolved decisions AND 0 critical gaps; otherwise "issues_open"
- unresolved: number from "Unresolved decisions" count
- critical_gaps: number from "Failure modes: ___ critical gaps flagged"
- issues_found: total issues found across all review sections (Architecture + Code Quality + Performance + Test gaps)
- MODE: FULL_REVIEW / SCOPE_REDUCED
- COMMIT: output of
git rev-parse --short HEAD
Review Readiness Dashboard
After completing the review, read the review log and config to display the dashboard.
~/.claude/skills/gstack/bin/gstack-review-read
Parse the output. Find the most recent entry for each skill (plan-ceo-review, plan-eng-review, review, plan-design-review, design-review-lite, adversarial-review, codex-review, codex-plan-review). Ignore entries with timestamps older than 7 days. For the Eng Review row, show whichever is more recent between review (diff-scoped pre-landing review) and plan-eng-review (plan-stage architecture review). Append "(DIFF)" or "(PLAN)" to the status to distinguish. For the Adversarial row, show whichever is more recent between adversarial-review (new auto-scaled) and codex-review (legacy). For Design Review, show whichever is more recent between plan-design-review (full visual audit) and design-review-lite (code-level check). Append "(FULL)" or "(LITE)" to the status to distinguish. For the Outside Voice row, show the most recent codex-plan-review entry — this captures outside voices from both /plan-ceo-review and /plan-eng-review.
Source attribution: If the most recent entry for a skill has a `"via"` field, append it to the status label in parentheses. Examples: plan-eng-review with via:"autoplan" shows as "CLEAR (PLAN via /autoplan)". review with via:"ship" shows as "CLEAR (DIFF via /ship)". Entries without a via field show as "CLEAR (PLAN)" or "CLEAR (DIFF)" as before.
Note: autoplan-voices and design-outside-voices entries are audit-trail-only (forensic data for cross-model consensus analysis). They do not appear in the dashboard and are not checked by any consumer.
Display:
+====================================================================+
| REVIEW READINESS DASHBOARD |
+====================================================================+
| Review | Runs | Last Run | Status | Required |
|-----------------|------|---------------------|-----------|----------|
| Eng Review | 1 | 2026-03-16 15:00 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
| Adversarial | 0 | — | — | no |
| Outside Voice | 0 | — | — | no |
+--------------------------------------------------------------------+
| VERDICT: CLEARED — Eng Review passed |
+====================================================================+
Review tiers:
- Eng Review (required by default): The only review that gates shipping. Covers architecture, code quality, tests, performance. Can be disabled globally with `gstack-config set skip_eng_review true` (the "don't bother me" setting).
- CEO Review (optional): Use your judgment. Recommend it for big product/business changes, new user-facing features, or scope decisions. Skip for bug fixes, refactors, infra, and cleanup.
- Design Review (optional): Use your judgment. Recommend it for UI/UX changes. Skip for backend-only, infra, or prompt-only changes.
- Adversarial Review (automatic): Always-on for every review. Every diff gets both Claude adversarial subagent and Codex adversarial challenge. Large diffs (200+ lines) additionally get Codex structured review with P1 gate. No configuration needed.
- Outside Voice (optional): Independent plan review from a different AI model. Offered after all review sections complete in /plan-ceo-review and /plan-eng-review. Falls back to Claude subagent if Codex is unavailable. Never gates shipping.
Verdict logic:
- CLEARED: Eng Review has >= 1 entry within 7 days from either `review` or `plan-eng-review` with status "clean" (or `skip_eng_review` is `true`)
- NOT CLEARED: Eng Review missing, stale (>7 days), or has open issues
- CEO, Design, and Codex reviews are shown for context but never block shipping
- If `skip_eng_review` config is `true`, Eng Review shows "SKIPPED (global)" and verdict is CLEARED
Staleness detection: After displaying the dashboard, check if any existing reviews may be stale:
- Content-first rule (diff-scoped rows only: `review`, `adversarial-review`, `codex-review`, ship-stage entries). Parse the `---WTREE---` and `---DIRTY---` sections from the bash output. If an entry has a `wtree` field AND it equals the current `---WTREE---` value, the review is CURRENT — identical content, regardless of commit count, rebase, amend, or whether it was committed yet (wtree equality alone proves identical content; that is the keystone property). Skip the commit-count heuristic for that entry and show no staleness note.
- Plan-tier rows (plan-ceo-review, plan-eng-review, plan-design-review) grade a plan file, not the repo tree — never apply the wtree rule to them; they keep the 7-day freshness logic. If such an entry carries a `plan_sha256` field, you MAY compare it against the current plan file's sha256 and note "plan changed since review" on mismatch.
- Fallback (no `wtree` on the entry, or wtree mismatch): parse the `---HEAD---` section to get the current HEAD commit hash. For each review entry that has a `commit` field: compare it against the current HEAD. If different, count elapsed commits: `git rev-list --count STORED_COMMIT..HEAD`. If that command FAILS (the stored commit was rebased away), grade UNKNOWN and treat as stale — do not error. Display: "Note: {skill} review from {date} may be stale — {N} commits since review"
- For entries without a `commit` field (legacy entries): display "Note: {skill} review from {date} has no commit tracking — consider re-running for accurate staleness detection"
- If all reviews grade CURRENT (wtree match or HEAD match), do not display any staleness notes
Plan File Review Report
After displaying the Review Readiness Dashboard in conversation output, also update the plan file itself so review status is visible to anyone reading the plan.
Detect the plan file
- Check if there is an active plan file in this conversation (the host provides plan file paths in system messages — look for plan file references in the conversation context).
- If not found, skip this section silently — not every review runs in plan mode.
Generate the report
Read the review log output you already have from the Review Readiness Dashboard step above. Parse each JSONL entry. Each skill logs different fields:
- plan-ceo-review: `status`, `unresolved`, `critical_gaps`, `mode`, `scope_proposed`, `scope_accepted`, `scope_deferred`, `commit` → Findings: "{scope_proposed} proposals, {scope_accepted} accepted, {scope_deferred} deferred" → If scope fields are 0 or missing (HOLD/REDUCTION mode): "mode: {mode}, {critical_gaps} critical gaps"
- plan-eng-review: `status`, `unresolved`, `critical_gaps`, `issues_found`, `mode`, `commit` → Findings: "{issues_found} issues, {critical_gaps} critical gaps"
- plan-design-review: `status`, `initial_score`, `overall_score`, `unresolved`, `decisions_made`, `commit` → Findings: "score: {initial_score}/10 → {overall_score}/10, {decisions_made} decisions"
- plan-devex-review: `status`, `initial_score`, `overall_score`, `product_type`, `tthw_current`, `tthw_target`, `mode`, `persona`, `competitive_tier`, `unresolved`, `commit` → Findings: "score: {initial_score}/10 → {overall_score}/10, TTHW: {tthw_current} → {tthw_target}"
- devex-review: `status`, `overall_score`, `product_type`, `tthw_measured`, `dimensions_tested`, `dimensions_inferred`, `boomerang`, `commit` → Findings: "score: {overall_score}/10, TTHW: {tthw_measured}, {dimensions_tested} tested/{dimensions_inferred} inferred"
- codex-review: `status`, `gate`, `findings`, `findings_fixed` → Findings: "{findings} findings, {findings_fixed}/{findings} fixed"
All fields needed for the Findings column are now present in the JSONL entries. For the review you just completed, you may use richer details from your own Completion Summary. For prior reviews, use the JSONL fields directly — they contain all required data.
Produce this markdown table:
```markdown
GSTACK REVIEW REPORT
| Review | Trigger | Why | Runs | Status | Findings |
|---|---|---|---|---|---|
| CEO Review | `/plan-ceo-review` | Scope & strategy | {runs} | {status} | {findings} |
| Codex Review | `/codex review` | Independent 2nd opinion | {runs} | {status} | {findings} |
| Eng Review | `/plan-eng-review` | Architecture & tests (required) | {runs} | {status} | {findings} |
| Design Review | `/plan-design-review` | UI/UX gaps | {runs} | {status} | {findings} |
| DX Review | `/plan-devex-review` | Developer experience gaps | {runs} | {status} | {findings} |
| ``` |
Below the table, add these lines. CODEX and CROSS-MODEL are optional (omit when empty); VERDICT is always present:
- CODEX: (only if codex-review ran) — one-line summary of codex fixes
- CROSS-MODEL: (only if both Claude and Codex reviews exist) — overlap analysis
- VERDICT: list reviews that are CLEAR (e.g., "CEO + ENG CLEARED — ready to implement"). If Eng Review is not CLEAR and not skipped globally, append "eng review required".
Unresolved-decisions status (MANDATORY — never omitted; the report's final non-whitespace line). After VERDICT, end the report (content under the `## GSTACK REVIEW REPORT` heading — a bold label, never a new `## ` heading; exempt from the "omit when empty" rule) with exactly one: the exact unbolded line `NO UNRESOLVED DECISIONS` (a bolded one does NOT count), OR a `UNRESOLVED DECISIONS:` header + one bullet per open item (last bullet = final line; add `+ N unresolved from prior reviews` only when N > 0). This avoids double-counting: list THIS review's open items from context; for prior reviews sum `unresolved` over the latest fresh row per skill (dashboard 7-day window) after you DROP the current skill's row; emit the sentinel only when both are zero.
Write to the plan file
PLAN MODE EXCEPTION — ALWAYS RUN: This writes to the plan file, which is the one file you are allowed to edit in plan mode. The plan file review report is part of the plan's living status.
The report must always be the LAST section of the plan file — never mid-file. Use a single delete-then-append flow:
- Read the plan file (Read tool) to see its full current content. Search the read output for a `## GSTACK REVIEW REPORT` heading anywhere in the file.
- If found, use the Edit tool to DELETE the entire existing section. Match from `## GSTACK REVIEW REPORT` through either the next `## ` heading or end of file, whichever comes first. Replace with the empty string. This applies regardless of where the section currently lives — mid-file deletion is intentional, not a special case. If the Edit fails (e.g., concurrent edit changed the content), re-read the plan file and retry once.
- After the delete (or skipped, if no section existed), append the new `## GSTACK REVIEW REPORT` section at the END of the file. Use the Edit tool to match the file's current last paragraph and add the section after it, or use Write to re-emit the whole file with the section at the end.
- Verify with the Read tool that `## GSTACK REVIEW REPORT` is the last `## ` heading in the file before continuing. If it isn't, repeat steps 2-3 once.
Do NOT replace the section in place. The "replace mid-file" path is what allowed prior versions to leave the report mid-file when an older report already lived there — the user then sees a plan whose review report is not at the bottom and (correctly) rejects it.
Capture Learnings
If you discovered a non-obvious pattern, pitfall, or architectural insight during this session, log it for future sessions:
~/.claude/skills/gstack/bin/gstack-learnings-log '{"skill":"plan-eng-review","type":"TYPE","key":"SHORT_KEY","insight":"DESCRIPTION","confidence":N,"source":"SOURCE","files":["path/to/relevant/file"]}'
Types: pattern (reusable approach), pitfall (what NOT to do), preference
(user stated), architecture (structural decision), tool (library/framework insight),
operational (project environment/CLI/workflow knowledge).
Sources: observed (you found this in the code), user-stated (user told you),
inferred (AI deduction), cross-model (both Claude and Codex agree).
Confidence: 1-10. Be honest. An observed pattern you verified in the code is 8-9. An inference you're not sure about is 4-5. A user preference they explicitly stated is 10.
files: Include the specific file paths this learning references. This enables staleness detection: if those files are later deleted, the learning can be flagged.
Only log genuine discoveries. Don't log obvious things. Don't log things the user already knows. A good test: would this insight save time in a future session? If yes, log it.
Brain Calibration Write-Back (Phase 2 / gated)
When the skill makes a typed prediction worth tracking (scope decision,
TTHW target, architectural bet, wedge commitment), it MAY write a
kind=bet take to the brain so a calibration profile builds over time.
Gated on two things:
- Brain trust policy for the active endpoint is
personal(check via~/.claude/skills/gstack/bin/gstack-config get brain_trust_policy@<endpoint-hash>). Shared brains skip write-back to avoid polluting team calibration. - Feature flag
BRAIN_CALIBRATION_WRITEBACKis set (today: false; flips to true when upstream gbrain v0.42+ shipstakes_addMCP op).
When both gates pass, the write-back path uses mcp__gbrain__takes_add
to record a take with weight 0.7 (per SKILL_CALIBRATION_WEIGHTS).
If the MCP op is unavailable, fall back to mcp__gbrain__put_page with
a gstack:takes fence block (documented but uglier path).
Mandatory take frontmatter shape:
kind: bet
holder: <user identity from whoami>
claim: <one-line prediction the skill is making>
weight: 0.7
since_date: <today's date>
expected_resolution: <date in 1-3 months depending on skill>
source_skill: plan-eng-review
After write, invalidate the affected digests so the next preflight reflects the new state:
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)" 2>/dev/null || true
# (no per-skill invalidation targets configured)
Brain Cache Background Refresh
After the skill's work completes (and telemetry has logged), kick a background refresh of any cache digest that's getting close to its TTL. This is non-blocking — the user doesn't wait. Next invocation benefits from the warm cache.
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)" 2>/dev/null || true
(~/.claude/skills/gstack/bin/gstack-brain-cache refresh --project "$SLUG" 2>/dev/null &) || true
Next Steps — Review Chaining
After displaying the Review Readiness Dashboard, check if additional reviews would be valuable. Read the dashboard output to see which reviews have already been run and whether they are stale.
Suggest /plan-design-review if UI changes exist and no design review has been run — detect from the test diagram, architecture review, or any section that touched frontend components, CSS, views, or user-facing interaction flows. If an existing design review's commit hash shows it predates significant changes found in this eng review, note that it may be stale.
Mention /plan-ceo-review if this is a significant product change and no CEO review exists — this is a soft suggestion, not a push. CEO review is optional. Only mention it if the plan introduces new user-facing features, changes product direction, or expands scope substantially.
Note staleness of existing CEO or design reviews if this eng review found assumptions that contradict them, or if the commit hash shows significant drift.
If no additional reviews are needed (or skip_eng_review is true in the dashboard config, meaning this eng review was optional): state "All relevant reviews complete. Run /ship when ready."
Use AskUserQuestion with only the applicable options:
- A) Run /plan-design-review (only if UI scope detected and no design review exists)
- B) Run /plan-ceo-review (only if significant product change and no CEO review exists)
- C) Ready to implement — run /ship when done
Unresolved decisions
If the user does not respond to an AskUserQuestion or interrupts to move on, note which decisions were left unresolved. At the end of the review, list these as "Unresolved decisions that may bite you later" — never silently default to an option.