mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-22 14:07:14 +02:00
* fix(hooks): nest freeze/careful permissionDecision under hookSpecificOutput Claude Code ignores a top-level permissionDecision, so the /freeze deny and /careful ask guards silently allowed everything. Nest both under hookSpecificOutput with permissionDecisionReason, update the shape-blind tests to pin the nested form, and document the constraint in both skill templates (regen included). Closes half of #1459 (freeze enforcement chain). Contributed by @jawadakram20 (PR #2331; team-init hunk deferred to the dedicated team-init fix). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(team-init): required-mode hook blocks with nested schema + exit 2 The generated check-gstack.sh emitted a flat permissionDecision payload and exited 0, which Claude Code ignores — required mode enforced nothing. The generated hook now nests the deny under hookSpecificOutput and exits 2 so the block holds even if the JSON schema drifts again. Adds a temp-repo regression test that runs the generated hook under both installed and missing-gstack homes. Fixes #2413, #2296. Contributed by @Masashi-Ono0611 (PR #2423). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(careful): close three check-careful bypasses via real JSON extraction The grep-based command extractor stopped at the first escaped quote, so any quoted argument truncated the command before the pattern checks ran — `git commit -m "wip" && rm -rf /` was silently allowed. Replace it with a python3/node JSON parse that fails CLOSED on unreadable payloads, add an IFS/base64-to-shell obfuscation tripwire, and stop multi-line commands from riding the single-line safe-exception whitelist (line-based grep would have approved `rm -rf /` when a later line matched node_modules — a hazard the real newline decoding exposed). Contributed by @wtamminga (PR #2426; the -R hunk was dropped — it landed in v1.61.0.0 — and output shapes updated to the nested hookSpecificOutput form). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review,autoplan): require explicit run_in_background: false on specialist agents Claude Code v2.1.198 made subagents run in the background by default, which inverted the old "do not use the flag" guidance: review-army specialists and autoplan dual voices silently launched in the background and the merge step could proceed before they completed — regressing the #497 fix. The generated guidance now instructs an explicit run_in_background: false, and a static tripwire fails the free suite if the inert inverted phrasing ever returns to any generated SKILL.md. Fixes #2440. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(investigate): anchor the scope-lock freeze hook on $HOME, not CLAUDE_SKILL_DIR The investigate skill's PreToolUse hooks and Scope Lock probe resolved check-freeze.sh via ${CLAUDE_SKILL_DIR}, which does not exist when frontmatter hooks run — the || exit 0 tail then failed open, so the debug scope boundary silently never engaged (#1871 follow-up). Anchor all four sites on $HOME/.claude/skills/gstack/ like careful/freeze, and add a static test asserting no frontmatter command: line in the guard-family skills ever references CLAUDE_SKILL_DIR again. Fixes #2469; closes the last live half of #1459 together with the freeze/careful hookSpecificOutput fix. The broader portable-install-root rewrite stays #1882 (its own focused PR per the TODOS.md decision). Reported with a fix by @maxpetrusenkoagent (PR #1873; absorbed narrowly — the cwd-walk rewrite belongs to #1882). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(redact): scan large diffs in line-aligned slices; stop digit-UUIDs matching as cards/phones The prepush guard blocked any push whose added lines exceeded the engine's 1 MiB cap with engine.input_too_large — a size error naming no credential — which trains people onto GSTACK_REDACT_PREPUSH=skip. Scan in 768 KiB line-aligned slices instead (no pattern is multi-line, so a boundary cannot bisect a secret); a single oversized line still goes to the engine intact and fails closed. Also suppress card/phone matches whose span sits ENTIRELY inside a UUID — digit-only UUID fixtures were 14 of 21 MEDIUM findings on an ordinary branch, the noise level that stops people reading MEDIUM at all. Fixes #2304. Contributed by @luckywenapere (PR #2543). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(redact): block Google OAuth client secrets and Telegram bot tokens at HIGH GOCSPX-prefixed client secrets and <bot_id>:<35-char> Telegram tokens are never-publishable credential shapes with unambiguous formats — both now block at HIGH like the other live-format credentials. Contributed by @francis-eye (PR #2357). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(redact-prepush): resolve the real push base instead of EMPTY_TREE whole-repo scans When the remote default branch is not main/master (or origin/HEAD is unset), the merge-base guess failed and the hook fell back to scanning the ENTIRE repository as added lines — re-attributing long-pushed secrets to the current push and, on any real repo, tripping the engine byte cap so the push blocked having scanned nothing. Derive the base from commits reachable from no remote-tracking branch, keep the empty-tree path only for genuinely fresh repos, and split the block message so an unscannable diff is reported as "could not scan (fail closed)" rather than "credential found — rotate it". Contributed by @stormeoio (PR #2398). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(redact-prepush): preserve the trailing newline handed to chained pre-push.local The chaining wrapper captured stdin with $(cat), which strips the trailing newline — a chained shell hook built on `while read` then never entered its loop for the final (usually only) ref line and exited 0, failing OPEN. Use the printf-x sentinel so the byte-exact input reaches the chained hook, with tests covering both the pass-through and the short-circuit paths. Contributed by @francis-eye (PR #2358). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(redact-prepush): close the ext-diff, header-lookalike, and ref-parse bypasses Three ways the pushed diff escaped scanning: (1) a user-level diff.external or textconv driver replaced the diff with its own output — zero '+' lines, so the scan saw nothing (now --no-ext-diff --no-textconv); (2) an added content line whose text begins with "++" renders as "+++…" and the blanket header skip dropped it (now hunk-aware header detection); (3) a pre-push ref line that failed to parse was silently skipped, leaving that ref unscanned (now fails closed with the offending line named). Minimal reimplementation of the two confirmed bypasses from PR #2498 by @lubosxyz (the full PR overlaps the chunked-scan work absorbed separately), plus the unparseable-ref hardening. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): keep the ngrok authtoken out of the transcript and shell argv The not-authed flow told the user to paste their ngrok authtoken into the chat so the agent could run `ngrok config add-authtoken` — putting a live credential in the transcript, tool-call argv, and anything the transcript syncs to. The user now runs the auth command in their own terminal; the agent only verifies via `ngrok config check`, and a pasted token triggers a rotate-and-reauth instruction. A static test pins that no agent-run bash fence ever contains add-authtoken again. Fixes #2335. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(update-check): crash emits CHECK_FAILED instead of reading as up-to-date gstack-update-check signals "up to date" with SILENCE, and it runs under set -e — so any unguarded mid-script failure exited quietly and was indistinguishable from a current install. Observed live as a 45-release silent-staleness incident. An ERR trap (with -E so it propagates into functions) now emits a CHECK_FAILED sentinel naming the line and status, and exits 0 so caller `|| true` guards can't eat it. Behavioral tests cover both the crash and the healthy-silent paths; egress-receipt wiring is untouched and still pinned by test/egress-receipt-wiring.test.ts. Fixes #1974. (#2378's HEAD-SHA staleness half was already fixed on main by the ls-remote + SHA-pinned VERSION resolution — close as already-fixed.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(deps): bump diff 7.0.0 → 9.0.0 (GHSA-73rr-hh4g-fpgx parsePatch DoS) The advisory affects diff 6.x–8.0.2. The only API this repo uses is Diff.diffLines (browse/src/snapshot.ts:571, browse/src/meta-commands.ts:728), which is unchanged across the major hop; snapshot tests pass against 9.0.0. Closes #1588. Contributed by @genisis0x (PR #1599; VERSION collateral stripped, lockfile regenerated fresh). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(evals): skip eval jobs deterministically on fork PRs Fork PRs never receive repository secrets, so every API-calling eval failed at SDK auth — but only when Docker-cache luck let the jobs start at all, making fork PRs randomly red or grey. Skip the eval and report jobs explicitly for fork-origin PRs, keep the image BUILD (validates Dockerfile.ci changes) without the push a fork token can't perform, and leave full coverage for same-repo PRs, pushes, and dispatches. Contributed by @andrey-esipov (PR #2345). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension): deny token/port reads to content-script and foreign senders background.js answered getPort — port, connected state, AND the browse server auth token — to any sender that passed the type allowlist, including content scripts running in web-page context and, behind only the sender.id check, anything without extension-page provenance. The getToken sender.tab restriction covered getToken alone, and only after getPort had already handed out the token. Single decision point now: extension/sender-auth.js classifies each message type; the eight privileged types (getPort, setPort, getServerUrl, getToken, fetchRefs, command, sidebar-command, getTabState) require an own-extension-page sender (chrome-extension://<own id>/ URL, no sender.tab, own sender.id). Denied senders get { error: 'unauthorized' } and nothing else — never the token, never the port. Content-script flows (elementPicked, pickerCancelled, inspectResult, openSidePanel) are untouched, and the sidepanel/popup keep the getPort token field their connect path reads. The policy mirrors the v1.63 server-side model: AUTH_TOKEN is released only to the pinned extension Origin via POST /extension-token, so the extension must not re-leak it to contexts the server would never have trusted. browse/test/extension-sender-auth.test.ts drives the real background.js onMessage listener under a chrome stub with four sender shapes (own extension page, own content script, foreign extension id, missing sender.url) and pins that denied responses carry no token/port fields, that a denied setPort never persists, that a denied command never reaches the network, and that the inspector + tab-state flows keep working. The helper is loaded via importScripts in the classic service worker and require()-able from bun tests. Contributed by @punksterlabs (PR #1822; reimplemented against the v1.63 POST /extension-token pinned-origin model). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(update-check): fixture links gstack-egress-lib.sh — all 38 tests failed on main v1.63.0.0 made bin/gstack-update-check source bin/gstack-egress-lib.sh unconditionally, but the test fixture's GSTACK_DIR only linked gstack-config — every test died at the source line (0/38 pass on pristine main, verified). The suite-truncation bug hid it: the runner was killed by an earlier file's delayed process.exit before this file ran. Link the lib like the real install layout the script assumes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): capture active-tab state before close() — last-tab auto-create raced the close event closeTab checked `tabId === this.activeTabId` AFTER awaiting page.close(), but the page 'close' event handler can fire during that await and reassign activeTabId — losing the race meant the last-tab auto-create never ran, leaving the manager with zero tabs. Capture wasActive before closing, and only reassign activeTabId when it no longer points at a live tab. Part of the test-integrity repairs unmasked by the suite-truncation fix. Contributed by @time-attack (PR #2230, browser-manager hunk). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(browse): delete the orphaned sidebar chat-queue suite; align sidebar-ux/tabs with the PTY-only sidebar browse/test/sidebar-integration.test.ts tested the /sidebar-command queue path ripped in v1.14 (34 references to removed endpoints — 11 permanent failures masked by suite truncation). sidebar-ux.test.ts carried 73 failures pinning the same dead surface (pickSidebarModel, ANALYSIS_WORDS); the trim keeps its 108 live tests, including the background.js token/allowlist gates. sidebar-tabs gets the two matching expectation updates. Closes #2420, #1980. Contributed by @time-attack (PR #2230, sidebar hunks; the security-sidepanel-dom deletion was NOT taken — that suite pins the live sidepanel DOM surface and passes). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(browse): align dual-listener and terminal-agent static guards with the current source Two static-grep guards pinned superseded source shapes and failed once the suite actually ran them: the tunnel dispatch gate is args-aware since the --out disk-write ban (canDispatchOverTunnel takes command AND args), and lazy PTY spawn routes through the maybeSpawnPty helper since v1.44. The updated assertions pin the current, stricter shapes (open() never spawns; the helper is the only spawnClaude caller). Contributed by @time-attack (PR #2230, dual-listener + terminal-agent hunks). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): remove all 8 delayed process.exit teardown bombs — the tier-1 gate can finally fail bun test runs every file in ONE process, so a 500ms setTimeout(process.exit(0)) armed in afterAll fired mid-way through a LATER file and killed the entire suite with exit 0 and no summary — only ~16 of 434 files ran, and every downstream failure was invisible (observed live throughout this wave's enumeration). Changes, all guarded by fault injection: - Replace every delayed-exit teardown with a time-boxed close of the file's own browser (8 files across browse/ and design/); stub the daemon /shutdown timer instead of letting its unconditional process.exit tear the runner down. - test/no-suicide-exit.test.ts: static tripwire — no *.test.ts may schedule a delayed process.exit again. - test/exit-propagation.test.ts + fixtures: fault injection with REAL bun output proves the truncation shape (exit 0, no summary) and that scripts/test-free-shards.ts now detects it: a shard exiting 0 WITHOUT bun's final summary line is treated as FAILED (exit code alone is not evidence of completion). - handoff: the three headed-mode integration tests are darwin-skipped with a pointer to the known macOS headed-launch breakage (#2242/#2554); they keep running on Linux CI. Un-skip in the browse-daemon wave. - feedback-roundtrip: repair the handler call sites unmasked by the fix — handlers take (command, args, session, bm); passing the manager where a session belongs broke all six tests. - user-slug-fallback: HOME isolation makes endpoint_hash deterministic. Fixes #2421, #2435. Contributed by @sneakygriff (PR #2172) with repairs from @time-attack (PR #2230 feedback-roundtrip hunks); supersedes PR #2252 by @whd4 (same defect, credited). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: include design/test/ in the free suite and the sharded runner design/test was absent from both the package.json test globs and TEST_ROOTS in scripts/test-free-shards.ts — its tests (including one of the teardown bombs removed in the previous commit) never ran in any CI or local free run, so design fixes could ship without their unit tests executing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(make-pdf): reject directories when resolving the browse binary access(X_OK) is true for directories (they carry the execute/traverse bit on POSIX and pass the Windows existence check too), so cwd-dependent resolution could pick the ~/.claude/skills/browse alias DIRECTORY as the browse binary. Every browse call then exited 4 with empty stderr, which make-pdf surfaced as "Chromium failed to launch" against a perfectly healthy Chromium (#2156). Guard isExecutable with statSync().isFile() so only regular files qualify. Contributed by @jwilk-hrep (PR #2538). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(make-pdf): write browse-bound temp files under the safe-dirs allowlist os.tmpdir() on macOS resolves to /var/folders/..., which fails browse's safe-dirs validation ([/tmp, cwd]) since the v1.6.0.0 --from-file tightening. Default PDF output (generate with no -o), the preview HTML, tmpFile() scratch files, and setup's smoke-test fixture/output all wrote there, so browse rejected the paths it was asked to read or write. Export PAYLOAD_TMP_DIR from browseClient (the existing TEMP_DIR convention: os.tmpdir() on Windows, /tmp elsewhere) and route orchestrator.ts and setup.ts temp files through it. Contributed by @lvthewah (PR #2505; the browse-binary directory guard from that PR landed separately via PR #2538). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(make-pdf): stop URLs swallowing smartypants placeholders A bare autolinked URL (<a href="X">X</a>) has zero whitespace between the URL text and its own closing tag. TAG_RE carves that </a> into a NUL-delimited SMARTPANTS_PRESERVED placeholder BEFORE the URL pass runs, and URL_RE's \S+ swallowed the adjacent placeholder into the URL match. The restore pass is single-shot, so the inner placeholder never restored: raw "SMARTPANTS_PRESERVED_N" text leaked into the rendered link, the </a> vanished, and link-blue styling bled into the rest of the document (#2084). Excluding the NUL sentinel (\u0000) from the URL character class stops the match from crossing into an already-carved zone. Contributed by @marshaung (PR #2280; PR #2339 by @BrendaB24 covered the same smartypants defect). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(make-pdf): no blank first page when content precedes the first H1 Two paths put invisible content ahead of the first H1 and cost users a blank page 1 (#1904): - A visually-empty preamble (leading <style> block, HTML comment) became its own .chapter. That section took the `.chapter:first-of-type { break-before: auto }` exception, so the first real chapter inherited `break-before: page` and started on page 2. Non-rendering preambles now fold into the first real chapter (markup preserved, no page break); real text preambles keep their own chapter. - Leading YAML frontmatter rendered as a literal paragraph of body text on its own first page (marked has no frontmatter awareness). It is now stripped before parsing; a `---` thematic break elsewhere is untouched. Contributed by @jbetala7 (PR #1913). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): allow about:blank so a restarted daemon can initialise The daemon opens its own first tab on about:blank, so blocking it in validateNavigationUrl meant a restarted daemon could never recreate the blank tab it starts from — and `browse newtab about:blank`, which `make-pdf setup` runs as its Chromium smoke test, failed and surfaced as "Chromium failed to launch" against a healthy browser. Allow about:blank ONLY, never the about: scheme: about:blank has no origin, loads nothing and runs nothing, while about:config and friends are real surfaces. Exact href match (lower-cased, since the URL parser normalises the protocol but not the opaque part), so about:blankfoo stays blocked. Contributed by @jwilk-hrep (PR #2537). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(design): drop gpt-image-2 tool model that 400s under the gpt-4o orchestrator The Responses API rejects pairing a gpt-4o orchestrator with an image_generation tool spec'd as model: "gpt-image-2" (400 invalid_request_error), which took every design image call offline — generate, variants, iterate (both threaded and fresh paths), evolve, and /design-shotgun (#1771). gpt-image-2 is only valid under a gpt-5 orchestrator; with gpt-4o the tool must omit the model field (defaults to gpt-image-1). Remove the model field at all five call sites and add a static-grep tripwire test (design/test/image-gen-pairing.test.ts) that fails CI if any design/src module reintroduces the gpt-4o + gpt-image-2 pairing. Re-enabling gpt-image-2 later requires bumping the orchestrator off gpt-4o in the same diff, which the tripwire permits. Contributed by @Pablosinyores (PR #1773). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(design): variants AbortError message reports the real 240s timeout generateVariant arms its abort at 240_000 ms but the AbortError branch returned "Timeout (120s)" — off by 2x, so a user staring at the failure could not tell whether to bump the timeout, retry, or drop the call. Report the actual configured bound, and pin it with a test that forces the abort path (fast-forwarding only the 240_000 ms timer) and asserts the surfaced string matches. Contributed by @vryahn (PR #1774). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(memory-ingest): stop silently ingesting 0 pages — include gitignored staging, reconcile counts Pages stage into ~/.gstack/.staging-ingest-*/ inside a repo whose .gitignore is `*`, and gbrain import honours .gitignore — so it collected 0 files, imported nothing, and the ingest still reported "written: N" from the STAGED count while advancing state, meaning no future run ever retried. Three layers now: (1) pass --include-gitignored (root cause); (2) if the installed gbrain predates the flag, retry without it (subcommand --help is generic, so the attempt is the only probe) with an upgrade pointer; (3) reconcile gbrain's imported+unchanged accounting against the staged count and REFUSE to advance state on a shortfall, naming the gitignore collision. Fixes #2144, #2104. Contributed by @gawievanblerk (PR #2560) and @Charles-Grant (PR #2486). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): task aggregator returned zero tasks on every run — jq scope bug Inside ($commits | split("|") | ...) the "." context is the split ARRAY, so the filter's bare .commit raised "Cannot index array with string" on every record — and the 2>/dev/null swallowed it, so aggregation silently produced zero tasks no matter how many the reviews emitted. Bind .commit to $c before the pipe. Reproduced live before the fix; regenerated autoplan/SKILL.md. Fixes #2018. Contributed by @kkroo (PR #2416; regenerated against the current template). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(session-update): un-wedge auto-upgrade — autostash over local patches, log the pull's real reason On a normal install the tracked files ARE locally patched (skill-prefix name rewrites, gbrain-refresh blocks), so the bare `git pull --ff-only` refused on every run and auto-upgrade froze forever — observed as 308 consecutive PULL_FAILED entries with the reason discarded by 2>/dev/null. Pull now runs --autostash (local patches ride over the update and pop back), stderr is captured into the log so a genuine failure names its cause, an autostash pop conflict recovers to a clean tree and re-renders the patches (gstack-patch-names + gbrain-refresh, both idempotent), and a successful pull re-renders them as a self-heal. Behavioral tests cover the wedge shape and the reason logging. Fixes #2566. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: raise the free-suite per-test timeout to 30s bun's 5s default is fine for a file run solo, but the monolithic free suite shares one process across 100+ files whose browser instances contend for launch slots — Playwright tests that pass in isolation time out mid-suite. 30s matches the ceiling the enumeration runs used; the sharded runner (test:free) is unaffected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(question-log): parse native AskUserQuestion answers — every native answer logged as __unknown__ Current Claude Code returns AskUserQuestion results as an OBJECT map keyed by question text ({answers: {question: label}}); the hook only handled the legacy array shapes, so 86% of live records carried user_choice __unknown__ — and the bin then scored every one as followed_recommendation false, silently poisoning plan-tune metrics. Adds the object-map extraction (exact + whitespace-normalized + single-question pairing, multiSelect joins, annotations as free_text), strips the (Recommended) suffix from BOTH sides of the comparison, skips the computation entirely on extraction failure, and logs unrecognized shapes to hook-errors.log instead of embedding them in the record. Fixes #2336, #2206. Based on the working patch in #2336 by @yijisoo; suffix comparison fix contributed by @chuchu2781 (PR #2400). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): canonicalize slash branches to dash form — review history stops splitting Branch-name sanitization disagreed across gstack (four incompatible rules), so reviews for the same slash-named branch landed in multiple files and the ship dashboard missed entries. gstack-slug now canonicalizes / to - in one place, and ship's review lookup routes through it; goldens regenerated against the current templates. Fixes #1127, #2550. Contributed by @ShuratCode (PR #2465; duplicate fixes by @xrfael-dev and two others in PRs #1851/#1699/#1621, credited). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): resolve the project root by marker walk-up — subdirectory sessions stop misfiling state gstack-slug derived everything from pwd, so a session in a subdirectory got the subdir's basename as its slug (or an outer monorepo's remote), misfiling reviews/decisions/learnings under a phantom project — and the per-pwd cache made the wrong answer permanent. The resolver now walks up from pwd: outermost STRONG marker wins (.git, package.json, pyproject.toml, Cargo.toml, Gemfile, go.mod, .project.yaml), weak content markers (README, LICENSE) catch non-code project folders, deploy artifacts are deliberately not markers, and GSTACK_PROJECT_SLUG remains the escape hatch. The cache self-heals on mismatch. Main-side invariants preserved on top: the unconditional [a-zA-Z0-9._-] re-sanitize before echo and slash→dash branch canonicalization. Fixes #1125. Contributed by @ajeenkya (PR #1702; rebased over the sanitize and branch-canonicalization work that landed after it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hooks): shared spawn-bin helper — all three AskUserQuestion hooks were inert on Windows The plan-tune hooks resolved bin scripts via new URL(import.meta.url).pathname (which doubles the drive letter on Windows: /C:/C:/...) and spawnSync'd extensionless bash scripts directly (unrunnable without a shell association) — so question logging, preferences, and the error fallback all silently no-op'd on Windows, and /plan-tune collected no data. A single spawn-bin.ts helper now owns bin resolution (fileURLToPath) and win32 bash routing for every hook, with static tripwires so a future hook can't reintroduce the raw pattern. This is the one Windows-spawn idiom for hook code. Fixes #2356. Contributed by @rafassousa (PR #2504; supersedes PR #2399 by @chuchu2781). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(model-overlays): add fable-5, opus-4-8, and sonnet-5 overlays + resolver mappings model-overlays/ had no entry for the current Claude generation, so every session on a Claude 5 family or Opus 4.8 model fell through to the generic claude.md nudges. Adds the three overlays with resolver mappings and per-overlay tests; generated output for the default host is unchanged (overlays activate by detected model). Closes #2509. Contributed by @chrisquorum (PRs #2246, #2243, #2247). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(windows): grant icacls ACEs by *SID, not unqualified username An unqualified username handed to icacls is ambiguous: on a machine whose hostname equals the username (a common Windows setup), it resolves to the MACHINE account instead of the user. Combined with /inheritance:r, that leaves ~/.gstack with a single ACE matching nobody — the process that just "secured" the directory locks itself out, and icacls still reports success. Both icacls sites in the repo (restrictFilePermissions and restrictDirectoryPermissions in browse/src/file-permissions.ts — the only icacls call sites; setup has none) now grant via icacls' literal-SID form `*<SID>`, resolved once per process from System32\whoami.exe (pinned to System32 because a bare `whoami` under a bash-flavoured PATH picks up the MSYS build, which rejects /user). Fallback when the SID can't be resolved is the domain-qualified `USERDOMAIN\username` name, which is unambiguous where the bare username was not. Windows-only regression tests assert the hardened directory stays usable by the calling process (readdir + write), which is exactly the check that a not-toThrow assertion sailed past before. Contributed by @asizux2 (PR #2479); the same defect was independently fixed by @Icandi40, @chiragborse1, @IntegriGit and @voltapix26. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(windows): forward windowsHide through the bun-polyfill spawn shims windowsHide is the one spawn option where Node's default is the opposite of Bun's: Node shows the child's console window, Bun.spawn hides it. The polyfill's spawn and spawnSync shims dropped the option entirely, so the Node fallback path (dist/bun-polyfill.cjs) silently inverted the behavior on the one platform the shim exists to serve — every watchdog respawn of the terminal agent popped a visible bun.exe console window. Three sites fixed: - Bun.spawnSync shim: forwards windowsHide with Bun-matching default true - Bun.spawn shim: same (stdio:'ignore' silences output but does NOT suppress the console window on Windows) - spawnTerminalAgent in terminal-agent-control.ts: explicit windowsHide: true, so the Node fallback path behaves like Bun-native An explicit windowsHide: false is honored at both shims. Three focused tests pin the default-true, default-true-sync, and explicit-false paths by intercepting child_process in a subprocess; the test file's require path now uses forward slashes so it survives interpolation into a JS string literal on Windows. Supersedes PRs #2523, #2294 and #2290, which each covered a subset of these sites. Contributed by @jerrynicholsai (PR #2539); earlier fixes by @jwilk-hrep, @rroojrooj and @WimvandenHeijkant covered subsets of the same sites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(watchdog): signal-0 liveness, tick-scaled respawn guard, windowsHide Three-bug chain behind the Windows terminal-agent leak (console window strobing every 60s, one orphaned agent per watchdog tick until the box ran out of committable memory): 1. isProcessAlive shelled out to `tasklist /FI "PID eq <pid>"` on Windows with a 3s timeout. A Bun.spawnSync that hits its timeout still RETURNS with partial stdout, so the `.includes()` PID match read a LIVE agent as dead — killAgentByRecord skipped the kill, the watchdog respawned around the survivor, and every orphan slowed the next tasklist enough to produce the next false negative. Now: `process.kill(pid, 0)` on every platform (Node and Bun both map signal 0 to an OpenProcess existence check on Windows), with EPERM counted as alive. No subprocess, no timeout, no console window. 2. The respawn circuit-breaker was mathematically unreachable — verified in this tree: RESPAWN_GUARD_WINDOW_MS was a fixed 60_000 against a 60_000ms default tick, and each tick pushes at most one respawn timestamp, so three pushes span ~120s and can never coexist inside a 60s window (eviction is strict `>`, and setInterval drift plus per-tick work always ages the prior entry past the boundary). The guard could not fire at the default tick rate and a steady one-per-tick leak ran unbounded. The window now scales with the tick: max(60_000, tick * (RESPAWN_GUARD_MAX + 2)), so "3 crashes in quick succession → stop" holds at any tick value. 3. The tasklist probe popped a visible console per tick (no windowsHide). Removing the shell-out kills that site; the agent-spawn site itself already passes windowsHide: true (landed with the bun-polyfill windowsHide commit — PR #2414's terminal-agent-control.ts hunk is reconciled there rather than duplicated). New browse/test/process-liveness-windows.test.ts pins all three: no subprocess from the probe, a static tripwire against reintroducing `tasklist` + `PID eq` liveness checks in src/, the spawnTerminalAgent windowsHide + stdio contract, and the window-derived-from-tick arithmetic. terminal-agent-watchdog.test.ts test 4 now pins the window/tick relationship instead of the fixed literal that let this ship. Also converts `new URL(import.meta.url).pathname` to `import.meta.path` across the static-grep tests it touches — the pathname form yields /C:/... on Windows and breaks path.resolve. Contributed by @SYKhayyat (PR #2414). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(terminal-agent): tie agent lifetime to its owning browse server PID The terminal agent is intentionally detached so it survives the short-lived CLI launcher, but its real owner is the persistent browse server. If that server crashed or was killed before running normal shutdown, the agent was adopted by PID 1 and lived forever (#2019). spawnTerminalAgent now requires an ownerPid and exports it to the agent as BROWSE_OWNER_PID; all three spawn sites pass the server PID (cli.ts cold-start, cli.ts supervisor respawn, server.ts watchdog). The agent polls the owner with signal 0 every 15s (GSTACK_TERMINAL_OWNER_WATCHDOG_MS to tune) on an unref'd timer and, when the owner disappears, exits through the SAME cleanup path as an intentional SIGTERM shutdown — now re-entrancy-guarded and also removing the terminal-internal-token file alongside the port file and agent record. Runtime test spawns a real agent tied to a throwaway owner process, kills the owner, and asserts the agent exits and its discovery files (terminal-agent-pid, terminal-port) are gone. Reconciled with the watchdog commit's spawnTerminalAgent contract test (process-liveness-windows.test.ts now passes ownerPid and pins the BROWSE_OWNER_PID env forwarding). Closes #2019. Contributed by @csarigoz (PR #2530). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(windows): give the bun-polyfill spawn shim a real `exited` promise Bun.spawn exposes `proc.exited` as a Promise resolving to the exit code. The Node fallback shim (dist/bun-polyfill.cjs) returned no such field, so every `await proc.exited` on the Windows path resolved instantly to undefined — the Windows cookie picker (cookie-import-browser.ts races proc.exited at three sites) read stdout before the child produced it and silent-failed; browser-skill-commands and terminal-agent hit the same class. The shim now: - drains stdout/stderr eagerly into capped in-memory buffers (Node's Readables are pull-based; without draining, a child writing past the OS pipe buffer blocks in write() and 'exit' never fires), replaying them as fresh single-shot Web ReadableStreams so reads work before or after awaiting exit; - caps the buffer at 16 MB (GSTACK_SPAWN_MAX_BUFFER to override), still draining past the cap so a runaway child can't wedge or OOM; - resolves `exited` with Bun-matching codes (exit code, 128+signal, 1 on spawn error) after both pipes finish, and resolves on 'error' too — Node fires 'error' without 'exit' when the binary is missing, which otherwise hangs the await forever. Six tests pin exit codes, the read-after-exit ordering, spawn-failure resolution, the buffer cap, and the large-output drain. Adapted to the current test file (require path goes through the requirePath variable from the windowsHide commit), and the 1 MB drain test's child now exits in the write callback — on modern Node a pipe write past the OS buffer is async and process.exit() straight after write() truncates at ~64 KB even with a live reader, which fails the test for reasons unrelated to the shim. Contributed by @punksterlabs (PR #1743). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): BROWSE_BIN carries the .exe suffix on Windows On Windows, `bun build --compile` emits browse.exe, but setup's BROWSE_BIN pointed at the suffixless path — so the post-build gate (`[ ! -x "$BROWSE_BIN" ]` → "browse binary missing") could never pass on Windows even after a fully successful build, while the build step itself reported success. Closes #2291. Applied the PR's override after the IS_WINDOWS detection, and also to the second BROWSE_BIN assignment the PR predates: the direct-Codex- install migration path re-derives BROWSE_BIN from the migrated dir and would otherwise drop the suffix again on Windows. Contributed by @rroojrooj (PR #1714). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): link lib/ beside bin/ at all five host-install sites bin/ scripts import shared modules via ../lib (gstack-learnings-log → lib/jsonl-store.ts is the reported case), so any runtime root that exposes bin/ without lib/ breaks 13 bin/ commands — learnings-log, decision-log, telemetry and friends fail with "Cannot find module .../lib/jsonl-store.ts" on every non-Claude install, silently from the skills' perspective. All five host-install sites now carry lib/ next to bin/, each through the existing _link_or_copy helper (never raw ln — the static invariant in test/setup-windows-fallback.test.ts enforces this): - .agents sidecar (create_agents_sidecar asset loop) - Codex runtime root (create_codex_runtime_root) - Factory runtime root (create_factory_runtime_root) - OpenCode runtime root (create_opencode_runtime_root) - Kiro install block New test/setup-runtime-lib-command.test.ts executes the real setup shell for each root in a sandbox (both the symlink branch and the Windows copy branch of _link_or_copy) and runs gstack-learnings-log end-to-end from the installed root, asserting the learning lands in ~/.gstack/projects/<slug>/learnings.jsonl — plus a negative control proving a bin-without-lib root fails exactly the way the bug report did. gen-skill-docs.test.ts's setup-validation block pins the lib link at every site. Cross-checked against PRs #2433, #2410 and #2198: all three cover subsets of these sites; nothing they fix is missing here. Contributed by @fedster99 (PR #2262); overlapping fixes by @gregario, @lsendel and @netkurt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): ship supabase/config.sh with every host runtime root Distinct from the lib/-beside-bin/ defect: gstack-telemetry-sync, gstack-update-check, gstack-security-dashboard and gstack-community-dashboard all source $GSTACK_DIR/supabase/config.sh to resolve GSTACK_SUPABASE_URL, where GSTACK_DIR is the installed root (parent of bin/). The [ -f ... ] guard means a root without the file degrades SILENTLY — telemetry and update checks just stop resolving the project URL on non-Claude installs. Closes #2215. setup now links supabase/config.sh (file-level on purpose — migrations/ and functions/ are dev-only) via _link_or_copy at all five host-install sites: the PR's four (Codex, Factory, OpenCode runtime roots + the Kiro block) plus the .agents sidecar, whose bin/ resolves the same relative path and which the PR predates covering. The runtime-root test now asserts supabase/config.sh is present in every built root, on both the symlink and Windows-copy branches. Contributed by @jizusun (PR #2216). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(windows): curate the fix-wave regression tests into the windows-latest run The windows-free-tests curated set is derived (POSIX-fragility regex scan + explicit deny list), and two of this wave's Windows regression files were auto-excluded on false-positive pattern hits: - browse/test/file-permissions.test.ts tripped the POSIX-mode-bitmask pattern, but every `mode & 0o777` assertion is platform-guarded — and the file carries the win32-only icacls-by-SID regression tests, which can only ever execute on windows-latest. - browse/test/terminal-agent-owner-watchdog.test.ts tripped the spawn(['bun','run',...]) pattern whose reason is the Playwright-bound browse server; it actually spawns terminal-agent.ts (fs/path/crypto + local helpers only, no Playwright at module scope), and the owner-PID orphan leak it pins was reported on Windows (#2019). Adds a KNOWN_WINDOWS_SAFE force-include list (mirror of KNOWN_WINDOWS_INCOMPATIBLE, each entry carrying its false-positive rationale) consulted before the pattern scan, and makes the owner-watchdog test's throwaway owner process Windows-portable (process.execPath instead of `sleep`, which a bare runner may not have). The wave's other new files need no wiring: process-liveness-windows and the bun-polyfill windowsHide/exited tests pass curation automatically; setup-runtime-lib-command self-skips on win32 by design (its Windows branch is exercised by simulating IS_WINDOWS=1 under bash), so force-including it would add a permanently-skipped file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): register the SessionStart hook with a bash prefix on Windows Windows can't execute an extensionless bash script directly — registering the bare gstack-session-update path made the hook pop the "Select an app" dialog on every session start (or silently never run), so team-mode auto-upgrade was dead on Windows installs. Companion to the hooks' spawn-bin routing: same defect class at the registration site. Contributed by @NikhileshNanduri (PR #1813; VERSION/CHANGELOG collateral stripped). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): stop piping gen:skill-docs through tail — generator failures were masked setup piped doc generation through `tail -3`, so a generator crash kept the pipe's exit 0 and installs completed "successfully" with broken or missing SKILL.md files. Capture the real exit status at BOTH sites (the main gen:skill-docs step and the gbrain-detected gen:skill-docs:user regen — the second drifted in after the PR and its own test caught it), print the tail for UX, and fail loudly. Contributed by @DavidMiserak (PR #1898; VERSION/CHANGELOG collateral stripped; extended to the second pipe site). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(mktemp): move the X-run to the end of every temp-file template (BSD/busybox safe) BSD mktemp (macOS) does not substitute an X-run that has a suffix after it: `mktemp "$TMP_ROOT/codex-err-XXXXXX.txt"` creates a LITERAL codex-err-XXXXXX.txt on the first call (exit 0) and every later call fails with `mkstemp failed: File exists` — so /codex breaks from the SECOND run on every Mac, masquerading as a model stall. busybox mktemp (Alpine) rejects the template on the first run. Fixes #2091, #2370. Union of both community fixes, compared at the diff level: - PR #2372: all 11 source sites with a suffix after the X-run — codex SKILL.md.tmpl (5), claude SKILL.md.tmpl (3), bin/gstack-developer-profile (2, suffix folded into the prefix: .json.tmp.XXXXXX), and the office-hours codex pass in scripts/resolvers/review.ts (1). - PR #2103: the second half of #2091 — bin/gstack-paths now strips the trailing slash from TMP_ROOT at the source (macOS $TMPDIR ends in `/`), plus runtime tests pinning that normalization. New repo-wide tripwire in test/regression-issue2091-bsd-mktemp.test.ts: every .tmpl, every SKILL.md, and every scripts/resolvers/*.ts is swept — no mktemp template may carry a suffix after the X-run, with a self-test so the detector can't be quietly blinded. Generated SKILL.md files regenerated via gen:skill-docs in this commit. Contributed by @ShuratCode (PR #2103) and @noron12234 (PR #2372); PR #2285 by @cathrynlavery covered a subset. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(codex,review,ship): scope codex review with an explicit --base flag, never prompt text `codex review` takes its scope ONLY from --base/--commit/--uncommitted. The positional [PROMPT] is mutually exclusive with all three, and a prompt-only `codex review "<text>"` silently falls back to the uncommitted working-tree scope (verified on 0.144.1: it runs `git status --short; git diff` and reviews that) — so the previous prompt-based scoping produced a confidently-worded review of the WRONG changes and read "no changes" on a clean tree. Every diff pass now invokes `codex review --base <base>` with no prompt argument: /codex Step 2A default path, the /review structured pass, and the /ship adversarial-section pass (all via scripts/resolvers/review.ts). Custom review instructions keep their own `codex exec` path (the CLI rejects prompt + scope flag together), with the filesystem boundary preserved there. Two new Error Handling entries teach the failure shapes: the argv-parse error, and the "review says no changes on a branch full of changes" symptom. Tests updated to pin the new invariant instead of banning the fix: the old assertions required the diff range in prompt text and banned the `--base <base> -c '...'` substring, which the correct scoped form contains. Also deletes test/fixtures/golden-ship-claude.md — a 2,565-line orphaned fixture referenced by zero tests (the live goldens are in test/fixtures/golden/, compared by test/host-config.test.ts); the factory golden is refreshed from the regenerated output. Generated SKILL.md files regenerated via gen:skill-docs in this commit. Contributed by @fangearhq-boop (PR #2513). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review,ship): run the codex diff passes under the timeout wrapper (#1036) The `_gstack_codex_timeout_wrapper` added in #1056 was wired into codex/SKILL.md but never into the /review and /ship diff passes, which kept running under a bare 5-minute Bash gate. An unwrapped stall returns no exit code and no output, which downstream reads as "Codex reviewed and found nothing" — a truncated pass silently became a clean bill. Measured on codex-cli 0.145.0: a pass was killed at 287s of a 300s budget mid-tool-call, and the same prompt completed in 336s. Both passes in scripts/resolvers/review.ts (adversarial `codex exec` and the structured `codex review --base` pass) now re-source gstack-codex-probe and run under `_gstack_codex_timeout_wrapper 540`, with the Bash tool gate raised to 600000 ms so the wrapper fires FIRST and a stall surfaces as a diagnosable exit 124. The timeout guidance now says a timed-out pass is MISSING COVERAGE, not a clean result, and points at the run's rollout log under ~/.codex/sessions/ for partial output. The stale "timeout doesn't exist on macOS" claim is gone — the wrapper resolves gtimeout, then timeout, then runs unwrapped, so it is safe without coreutils. Static guards in test/codex-hardening.test.ts pin all three sites (resolver, review/SKILL.md, ship/sections/adversarial.md): both calls wrapped, wrapper budget strictly under the Bash gate, and no reappearance of the macOS claim that steered these call sites away from the wrapper in the first place. The Claude-output path guard in test/gen-skill-docs.test.ts now scrubs ~/.codex/sessions/ (a user-facing Codex CLI path, same class as the ~/.codex/logs/ exemption) before banning Codex host paths. Generated files regenerated via gen:skill-docs; factory golden refreshed. Contributed by @aegixx (PR #2379). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(codex): sandbox the review path, fail the gate closed, order timeouts wrapper-first Closes #2496, #2524, #2477 — three defects in the class "a guard that reports success while doing nothing", all in codex/SKILL.md.tmpl: (a) Review sandbox. The default `codex review` path was the only codex call with no sandbox override, inheriting ~/.codex/config.toml's default — write access on a trusted project — while Important Rules claimed read-only. Top-level `codex review` has no -s/--sandbox flag (verified on 0.147.0), so the invocation now pins `-c 'sandbox_mode="read-only"'`, the same form the consult-resume path already uses. (b) Fail-closed verdict gate. The old rule ("no [P1] found → PASS") could not fail on the default path: native `codex review` output carries no bracketed tags, and a non-zero exit, expired auth, timeout, or empty result also contains no [P1] — all read as PASS. The gate is now an ordered, fail-closed check: non-zero exit → FAIL; empty output → FAIL; [P0]/[P1] (bracketed or codex's native labels) → FAIL with count; NO severity tags at all → FAIL requiring a human read; PASS is only reachable through the explicit tagged-advisory-only branch. [P0] is recognized as blocking, and the review-log findings count includes it. (c) Bash gate above the wrapper. Step 2A instructed `timeout: 300000` under a 330s wrapper, and Challenge's 300s gate sat under a 600s wrapper — the harness killed the call before the wrapper could emit its diagnosable exit-124 message. Every Bash gate now sits strictly ABOVE its wrapper: 360000 over the 330s review wrapper, 660000 over the 600s challenge/consult wrappers, with the ordering rationale stated at each site. Also from #2477/#2524: a new Error Handling entry for the model-entitlement 400 ("The '<model>' model is not supported...") pointing at the `model =` pin and `[notice.model_migrations]` in ~/.codex/config.toml and saying exactly which override to retry with (-m for exec-based modes, `-c model="..."` for review mode, which rejects -m); the Model & Reasoning section no longer documents `-m` for `/codex review`. Static assertions in test/codex-hardening.test.ts pin (a)-(c) across both the .tmpl and the generated SKILL.md: every scoped review invocation carries sandbox_mode="read-only" and never -s; the default-PASS sentence is banned and the fail-closed branches are present; and per-section, every Bash `timeout: N` is strictly greater than every wrapper budget, with 2A/2B/2C all required to be inspected. Generated SKILL.md regenerated via gen:skill-docs in this commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): quoted tilde made Artifacts Sync and telemetry-finalize dead code in 49 skills A tilde inside double quotes never expands, so the generated `_BRAIN_SYNC_BIN="~/..."` assignments resolved to a literal ./~ path and the Artifacts Sync + telemetry-finalize blocks silently no-op'd in every skill that carried them (regression of #785). The preamble resolvers now emit $HOME-based paths; all generated SKILL.md files regenerate identically from the fixed templates, and a static tripwire fails the suite if a quoted-tilde assignment ever reappears in generated output. Fixes #1656, #1715. Contributed by @jawadakram20 (PR #2333). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen-skill-docs): stop the catalog trim chopping descriptions at embedded periods The description-trim regex treated the first period as end-of-sentence, so skill descriptions with embedded periods (e.g. file extensions, version numbers) truncated mid-thought in the generated catalog — the discovery surface every host loads. Trim now respects the full first sentence; diagram's description regenerates to its intended text. Contributed by @sneakygriff (PR #2171). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): update_check:false gates the prose, not just the binary Setting update_check:false stopped the update-check BINARY from running, but every skill preamble still shipped the upgrade-handling instruction prose unconditionally — burning tokens on instructions that could never fire and confusing agents into probing for upgrades anyway. The resolver now suppresses the upgrade-flow prose when the config disables checks. Fixes #2001. Contributed by @jc0d35 (PR #2022). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): sidebar Terminal — drop the duplicate WS subprotocol header, stop doubling CJK IME input The terminal client passed the auth token as the WS subprotocol AND echoed it in a second header, which some Chromium builds reject; and composition events double-sent CJK input (each IME commit arrived once from the composition handler and once from the data handler). One auth path, one input path; also fixes the terminal-agent test that failed on clean main. Contributed by @mindsurf0176 (PR #2515). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): -h/--help prints usage instead of running the installer Asking setup for help RAN the full installer — Playwright download and all. Standard help flags now short-circuit to usage. Contributed by @saen-ai (PR #1219). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hosts): Codex-generated skills reference AGENTS.md, not CLAUDE.md Codex reads AGENTS.md, but its generated skills still told agents to read CLAUDE.md in 8 places — instructions Codex hosts cannot follow. The host config now maps the memory-file name per host; all three ship goldens refreshed from the regenerated output. Contributed by @exGeni (PR #1996). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(retro,ship): count tracked files for the test-file metric, not the working tree The test-file count ran find over the working tree, sweeping untracked build output — a Rails repo reported 623 test files when git tracks 17 (37x), skewing retro narratives and ship dashboards. Count via git ls-files instead; includes the one-line Python-glob widening so non-JS repos stop undercounting. Fixes #2307, #1999. Contributed by @joshRpowell (PR #2308). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(land-and-deploy,gen): auto-merge diagnosis + CRLF-stable generation Two small hardenings: land-and-deploy Step 4 no longer misdiagnoses a failed `gh pr merge --auto` as a permissions problem when the real cause is the merge-method mismatch the command names; and gen-skill-docs normalizes CRLF at the template entry point so Windows checkouts with autocrlf produce byte-identical generated output to CI instead of silently skipping the \n-anchored transforms. Contributed by @Jmeg8r (PR #2437) and @1ncludeSteven (PR #1051). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(land-and-deploy): stop greedy sed from eating the URL scheme in deploy-config parsing The deploy-config bootstrap parsed "Production URL: https://x.com" with sed 's/.*: *//', which cuts at the LAST colon — the one in "https:" — yielding "//x.com". Cut at the first ": " instead (s/^[^:]*: *//). Resolver only; the generated land-and-deploy/SKILL.md regenerates from this source in the docs lane. Contributed by @briascoi (PRs #2555/#2493). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(artifacts-init): honor the provider CLI's git_protocol instead of forcing SSH gstack-artifacts-init unconditionally rewrote the push remote to SSH and hard-failed setup for users whose gh/glab auth is HTTPS-only. Now: - provider-created remotes follow `gh config get git_protocol` / `glab config get git_protocol` (HTTPS when unset — the gh default) - explicit/existing/manual remotes keep their given protocol; unknown URL forms (local bare paths, file://, self-hosted) pass through - new --push-protocol auto|https|ssh flag overrides the inference - the unreachable-remote error names the actual protocol and points at --push-protocol instead of assuming a missing SSH key Closes #1348. Contributed by @time-attack (PR #2225). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): skip the .gitignore append when git already ignores .gstack/ ensureStateDir appended ".gstack/" to a tracked .gitignore even when git already ignored the directory via global excludes, .git/info/exclude, or a parent .gitignore — dirtying the working tree on every daemon start. Run `git check-ignore -q -- .gstack/` first and return early when git says it's covered; git-missing/not-a-repo/timeout all fall through to the existing text-check append (the safe default). Closes #2385. Contributed by @gregario (PR #2430). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): guard browser.process() in resolveDisconnectCause `.process()` only exists on browsers Playwright launched itself; a browser from connectOverCDP() (or a test stub) has no such method, so the blind call threw "browser?.process is not a function" inside the disconnect handler and took down the daemon. Type-check the method before calling it and treat the no-method case as no process handle. Closes #2085. Contributed by @elan2002 (PR #2434). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(lib): narrow the override injection denylist to instruction-shaped phrases The /override[:\s]/i pattern flagged any prose containing "override " or "override:" — CLI flags (--port-override -1), tfvars notes, and plain "you can override the default region" all tripped the injection guard. Require an instruction-shaped continuation: "override (all)? previous | prior | above | the rules/instructions/system prompt". Genuine attempts like "Override: ignore all previous instructions" still block via the ignore-previous pattern. Closes #2401, #1934. Contributed by @Masashi-Ono0611 (PR #2424); same fix independently by @JonasFocus (PR #1940). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(redact): stop the E.164 phone pattern flagging compact timestamps Bare 14-digit runs like 20260727202423 (YYYYMMDDHHMMSS backup/log stamps) matched the phone regex and produced MEDIUM PII findings. Reject a separator-free 14-digit span whose fields parse as a plausible date-time; real numbers carry a + or spacing, so phone coverage is unchanged. Contributed by @abkrim (PR #2428). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(design): create the OpenAI key file owner-only, closing the write-then-chmod race saveApiKey wrote ~/.gstack/openai.json at the default umask and tightened to 0600 afterwards, leaving the API key briefly world-readable between write and chmod (CWE-377/367). Pass mode 0o600 at create; the trailing chmodSync stays as a backstop to tighten a pre-existing loose file. Contributed by @bunlongheng (PR #2468). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): make gstack-config key validation locale-independent POSIX bracket ranges like a-z follow the active collation order; under GNU grep with tr_TR.UTF-8 the range excludes the ASCII letter i, so every key containing i (skill_prefix, explain_level, ...) was rejected as invalid. Pin both get/set validators to LC_ALL=C, with a source-level tripwire test since macOS BSD grep doesn't reproduce the bug. Closes #2494. Contributed by @Math1987 (PR #2506). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(resolvers): stop env-var hosts from doubling $HOME in the binary fallback path The browse/design/make-pdf setup resolvers built the fallback binary path as "$HOME" + dir.replace(/^~/, ''), which is only correct for ~-rooted dirs. Env-var hosts carry an absolute $GSTACK_* dir, so the generated fallback became $HOME$GSTACK_.../browse — a path that never exists. New toShellPath() in scripts/resolvers/types.ts expands ~ to $HOME and passes absolute env-var dirs through untouched; all five call sites route through it. Claude-host generated output is byte-identical, so no SKILL.md regeneration is needed here. Closes #2055. Contributed by @simjak (PR #2056). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings-hook): respect CLAUDE_CONFIG_DIR when resolving settings.json gstack-settings-hook hardcoded $HOME/.claude/settings.json, so users running Claude Code with a relocated CLAUDE_CONFIG_DIR had hooks written to a config file Claude never reads. Resolve ${CLAUDE_CONFIG_DIR:-$HOME/.claude} first; the explicit GSTACK_SETTINGS_FILE override still wins. Partial #349. Contributed by @andrefogelman (PR #2239). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): dispatch a change event after fill for change-only validators Playwright's Locator.fill() dispatches `input` but never `change`, so frameworks that validate on change (AngularJS ng-change, debounced strength/match checks) never saw the filled value — correct in the DOM, failing the framework's own validation. `browse fill` now dispatches `change` after the fill. Failing-first regression test with a change-only password-match fixture included. Contributed by @intelliot (PR #2475). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(safety): unknown question-preference source exits the documented 2, not 1 The --write user-origin gate documents exit 2 as "rejected, do not retry" (profile poisoning defense), but a source outside both the allowed and the explicitly-rejected lists fell through to exit 1 — the generic validation code callers treat as retryable. Unknown sources now exit 2 with the same do-not-retry rejection message as the known non-user-originated ones. Closes #2390. Contributed by @gregario (PR #2429). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pr-title): stop duplicating the version prefix on bare-version titles A title that was nothing but a version ("v1.2.3" — the form ship uses for version-only bumps) matched neither the "v<NEW_VERSION> " literal case nor the trailing-space strip regex, fell through to the prepend path, and came out as "v1.2.3.4 v1.2.3" — which pr-title-sync.yml then wrote back via gh pr edit. Handle the bare form in both the no-change case and the prefix-strip regex, and emit a bare new version when nothing follows. Closes #1886. Contributed by @jbetala7 (PR #1887). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(build): escape literal braces in the bun:sqlite stub regex Perl >= 5.26 treats an unescaped literal `{` in a pattern as fatal ("Unescaped left brace in regex is illegal"), so build-node-server.sh died at the bun:sqlite stub substitution on modern perl. Escape both braces; the replacement output is unchanged. Closes #2300. Contributed by @nuga0718 (PR #2111). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): preserve spaces in gstack-config values get/list read values with awk '{print $2}' | tr -d '[:space:]', which truncated any value containing spaces ("/Users/x/Conductor Workspaces" came back as "/Users/x/Conductor") and set wrote the unfiltered raw value on the append path. New read_config_value() strips only the "key:" prefix and trailing whitespace (cut-style parse), and set appends the same newline-stripped value the in-place edit path uses. Closes #1782. Contributed by @jbetala7 (PR #1783). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): recover a late-healthy detached daemon instead of a false "Server failed to start" startServer spawns the daemon detached + unref'd, then polls health for a fixed budget. On a loaded machine the budget can elapse in the gap between the loop's last tick and the daemon becoming ready — the CLI reported "Server failed to start within Ns" while the very next `browse status` showed a healthy server. Add a final readState()+isServerHealthy() re-check before the timeout throw, and make the budget env-overridable via BROWSE_START_TIMEOUT (BROWSE_* tunable convention). Structural + behavioral tests pin both invariants. Closes #1846. Contributed by @harjothkhara (PR #1847). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): daemon resilience on loaded machines — Bun conn errors, stop/restart flush, startup + git-root budgets Four load-sensitivity fixes in the daemon lifecycle: - sendCommand only recognized Node's ECONNREFUSED/ECONNRESET; the compiled CLI runs on Bun, which reports 'ConnectionRefused'/'ConnectionClosed' ("Unable to connect..."), so daemon crashes leaked the raw error and exited 1 instead of entering the busy-check/restart path. Match both. - stop/restart called shutdown() inline, which exits before the HTTP response flushes — the CLI saw a dropped socket (and would now crash-retry a fresh daemon just to stop it). Defer shutdown ~100ms so the 200 lands first. - Non-CI POSIX startup budget raised 8s -> 15s (cold Chromium measured ~5.7s at load avg 10; load 12+ blew the old budget while the detached daemon was still booting). - getGitRoot's 2s git rev-parse timeout returned null under load (6.3s spikes measured), scattering state files across cwds into split-brain daemons. Raise to 8s, still bounded. Contributed by @mplatts (PR #1732). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(telemetry): ingest keeps error_message/failed_step instead of dropping them The telemetry_events columns exist and bin/gstack-telemetry-log already sends error_message + failed_step, but the Supabase ingest function dropped both fields on insert — every error report arrived with no message and no failing step. Map them through with the same bounded-length sanitization as error_class (500/100 chars). The completion-status resolver now also passes --error-message/--failed-step in the generated skill telemetry block, with instructions to leave them empty on success. Resolver only for the template side; generated SKILL.md files regenerate from this source in the docs lane. Contributed by @sunnnybala (PR #769). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): surface non-EEXIST errors in acquireServerLock instead of masking them acquireServerLock caught every open failure as if the lock were held: EACCES/EROFS/ENOENT surfaced as phantom "another process holds the lock" (null return, no diagnostics), and a failed stale-lock read or unlink was swallowed the same way. Each failure class now logs a coded, pathed diagnostic: non-EEXIST open errors, holder-PID read errors (ENOENT retries the acquire — the holder released between open and read), and stale-lock unlink errors. Four-case unit test included. Closes #1084. Contributed by @jbetala7 (PR #1725); same fix independently by @JiayuuWang (PR #1097). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(paths): shell-quote gstack-paths output so eval round-trips values gstack-paths emitted bare KEY=VALUE lines, so the documented eval "$(gstack-paths)" re-parsed the values: backslashes were eaten as escapes (Windows $TMP C:\Users\... became C:Users...) and a space word-split the assignment, leaving the variable empty. Emit each value with printf %q so eval round-trips byte-for-byte; plain POSIX paths are unchanged. Round-trip regression tests cover backslashes, spaces, and embedded quotes. Closes #2374. Contributed by @fangearhq-boop (PR #2376); same fix independently by @yannickspiess (PR #1580). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * security(browse): drop .svg from the load-html extension allowlist SVG is a script-capable format (inline <script>, event handlers, foreign objects), so allowing it through load-html's HTML allowlist let a local .svg execute script in the browse session context. The allowlist is now .html/.htm/.xhtml only; regression test asserts .svg is rejected. Contributed by @garagon (PR #1153). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(benchmark): validate --timeout-ms as a positive integer gstack-model-benchmark fed --timeout-ms straight through parseInt, so "abc" became NaN and "0"/"-1" passed through — a NaN or non-positive timeout silently disables the per-provider watchdog. Reject anything that isn't a positive (optionally +-prefixed) safe integer with a clear error and exit 1. Closes #1726. Contributed by @jbetala7 (PR #1727). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(fixtures): clean terminology in the security-bench replay fixture Two spots in browse/test/fixtures/security-bench-haiku-responses.json referred to real-world HVAC project naming; replace with the generic "mechanical services" wording. Fixture stays valid JSON; replay tests unchanged. Contributed by @apex-system (PR #2131). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: cancel superseded actionlint and skill-docs runs actionlint.yml and skill-docs.yml trigger on both push and pull_request with no concurrency group, so every push to an active branch left the previous (now-obsolete) runs queued or running — twice per commit on same-repo PR branches. Add the same cancel-in-progress concurrency groups the heavier workflows already use, plus a free static tripwire test that fails CI if a push+pull_request workflow ever ships again without cancel-in-progress. Contributed by @jbetala7 (PR #2053). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(make-pdf): correct CJK rendering — NUL sentinel hardening, SC-first fonts, CJK quote context Three CJK fixes in the PDF pipeline: - smartypants strips stray input NULs up front so document text can never forge the U+0000 placeholder sentinel and leak a preserved-zone marker into the output. - The CJK font stack led with Japanese families, so Simplified-Chinese text rendered han glyphs with JP variants. Lead with PingFang SC / Heiti SC / Noto Sans CJK SC / Source Han Sans SC before the JP fallbacks. - Quote-smartening only recognized ASCII openers as "start of quote" context; the fullwidth colon and CJK brackets now count, so quotes after them curl the right way. Contributed by @rssprivacy-commits (PR #2012). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: regenerate skill output for the quick-win resolver changes Regen for the deploy-config URL-scheme fix (utility resolver), telemetry completion-status resolver, and $HOME-doubling binary-resolver fix; ship goldens refreshed to match. Generated-output-only commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): cached identity is sticky — heal ONLY the provable subdir-cache bug shape The walk-up rewrite recomputed the slug on every run and "healed" the cache toward the fresh value, which broke the #2212 continuity contract: a project that used gstack before adopting a git remote would be silently renamed to the remote-derived slug, orphaning everything under ~/.gstack/projects/. Cached identity now wins, with one precise exception: when the cached value equals THIS pwd's basename while the walk-up proves pwd is not the project root, the entry came from the pre-walk-up subdirectory bug (#1125) and is recomputed. All four slug contracts pass together (repo-mode #2212, walk-up #1125, sanitize, user-slug). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(claude): stop false-blocking macOS keychain subscription auth in host detection The /claude skill's auth probe only recognized env-var/API-key auth, so macOS subscription installs (keychain-backed, where `claude -p` works fine) were told they had no auth. Detection now uses host invocation. Fixes #1890. Contributed by @xing-qnex (PR #2411); PR #2548 by @shawnacalia covered the keychain case. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): Ubuntu 26.04 Playwright platform detect + silence the codesign false alarm Two small setup papercuts: the Playwright platform probe now recognizes Ubuntu 26.04 instead of falling to the generic-Linux path, and macOS installs stop warning about a codesign "failure" that was actually the expected unsigned-adhoc path (the real signature check already gates binary launch). Contributed by @nuga0718 (PR #2113) and @lucascaro (PR #1758). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): land-and-deploy squash readback, next-version paths, embed-flags quoting Three template one-liners: land-and-deploy reads the squash-merge result from the merge commit instead of the stale branch tip; review/landing-report /land-and-deploy templates call bin/gstack-next-version via its installed path instead of a bare repo-relative one; setup-gbrain quotes GBRAIN_EMBED_FLAGS so zsh word-splitting stops silently dropping voyage-code-3 flags. Regenerated output included. Contributed by @stormeoio (PR #2011), @rjmurillo (PR #1820) and @trevorhstandridge (PR #1817). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * release: v1.64.0.0 — fix wave CHANGELOG, VERSION, deferred-wave TODOs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh ship goldens for the telemetry error-field resolver output Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(redact-prepush): assemble the fake AWS key at runtime — the literal blocked our own push The hook's fixtures carried a live-format AKIA literal, and the repo's own pre-push scanner (hardened in this wave) correctly blocked pushing it. The placeholder-suppressed docs key would defeat the detection tests, so the fixtures now concatenate the key at runtime: tests still exercise real detection, and the pushed diff never contains a scannable credential shape. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): terminate the marker walk-up on dirname's fixed point — hung every bin on Windows Under git-bash on Windows a mixed-form path walks C:/Users -> C: -> . -> . forever: dirname's fixed point there is never "/", so the walk-up loop spun and every bin that evals gstack-slug (learnings-log first among them) hung until spawn timeout. Caught by windows-free-tests CI on the wave PR. Break on the fixed point itself with a depth cap for exotic forms; regression tests drive the extracted function with hostile path shapes under a hard timeout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
1158 lines
58 KiB
TypeScript
1158 lines
58 KiB
TypeScript
import { type TemplateContext, toShellPath } from './types';
|
|
import { AI_SLOP_BLACKLIST, OPENAI_HARD_REJECTIONS, OPENAI_LITMUS_CHECKS } from './constants';
|
|
|
|
export function generateDesignReviewLite(ctx: TemplateContext): string {
|
|
const litmusList = OPENAI_LITMUS_CHECKS.map((item, i) => `${i + 1}. ${item}`).join(' ');
|
|
const rejectionList = OPENAI_HARD_REJECTIONS.map((item, i) => `${i + 1}. ${item}`).join(' ');
|
|
// Codex block only for Claude host
|
|
const codexBlock = ctx.host === 'codex' ? '' : `
|
|
|
|
7. **Codex design voice** (optional, automatic if available):
|
|
|
|
\`\`\`bash
|
|
command -v codex >/dev/null 2>&1 && echo "CODEX_AVAILABLE" || echo "CODEX_NOT_AVAILABLE"
|
|
\`\`\`
|
|
|
|
If Codex is available, run a lightweight design check on the diff:
|
|
|
|
\`\`\`bash
|
|
TMPERR_DRL=$(mktemp /tmp/codex-drl-XXXXXXXX)
|
|
_REPO_ROOT=$(git rev-parse --show-toplevel) || { echo "ERROR: not in a git repo" >&2; exit 1; }
|
|
codex exec "Review the git diff on this branch. Run 7 litmus checks (YES/NO each): ${litmusList} Flag any hard rejections: ${rejectionList} 5 most important design findings only. Reference file:line." -C "$_REPO_ROOT" -s read-only -c 'model_reasoning_effort="high"' --enable web_search_cached < /dev/null 2>"$TMPERR_DRL"
|
|
\`\`\`
|
|
|
|
Use a 5-minute timeout (\`timeout: 300000\`). After the command completes, read stderr:
|
|
\`\`\`bash
|
|
cat "$TMPERR_DRL" && rm -f "$TMPERR_DRL"
|
|
\`\`\`
|
|
|
|
**Error handling:** All errors are non-blocking. On auth failure, timeout, or empty response — skip with a brief note and continue.
|
|
|
|
Present Codex output under a \`CODEX (design):\` header, merged with the checklist findings above.`;
|
|
|
|
return `## Design Review (conditional, diff-scoped)
|
|
|
|
Check if the diff touches frontend files using \`gstack-diff-scope\`:
|
|
|
|
\`\`\`bash
|
|
source <(${ctx.paths.binDir}/gstack-diff-scope <base> 2>/dev/null)
|
|
\`\`\`
|
|
|
|
**If \`SCOPE_FRONTEND=false\`:** Skip design review silently. No output.
|
|
|
|
**If \`SCOPE_FRONTEND=true\`:**
|
|
|
|
1. **Check for DESIGN.md.** If \`DESIGN.md\` or \`design-system.md\` exists in the repo root, read it. All design findings are calibrated against it — patterns blessed in DESIGN.md are not flagged. If not found, use universal design principles.
|
|
|
|
2. **Read \`.claude/skills/review/design-checklist.md\`.** If the file cannot be read, skip design review with a note: "Design checklist not found — skipping design review."
|
|
|
|
3. **Read each changed frontend file** (full file, not just diff hunks). Frontend files are identified by the patterns listed in the checklist.
|
|
|
|
4. **Apply the design checklist** against the changed files. For each item:
|
|
- **[HIGH] mechanical CSS fix** (\`outline: none\`, \`!important\`, \`font-size < 16px\`): classify as AUTO-FIX
|
|
- **[HIGH/MEDIUM] design judgment needed**: classify as ASK
|
|
- **[LOW] intent-based detection**: present as "Possible — verify visually or run /design-review"
|
|
|
|
5. **Include findings** in the review output under a "Design Review" header, following the output format in the checklist. Design findings merge with code review findings into the same Fix-First flow.
|
|
|
|
6. **Log the result** for the Review Readiness Dashboard:
|
|
|
|
\`\`\`bash
|
|
${ctx.paths.binDir}/gstack-review-log '{"skill":"design-review-lite","timestamp":"TIMESTAMP","status":"STATUS","findings":N,"auto_fixed":M,"commit":"COMMIT"}'
|
|
\`\`\`
|
|
|
|
Substitute: TIMESTAMP = ISO 8601 datetime, STATUS = "clean" if 0 findings or "issues_found", N = total findings, M = auto-fixed count, COMMIT = output of \`git rev-parse --short HEAD\`.${codexBlock}`;
|
|
}
|
|
|
|
// NOTE: design-checklist.md is a subset of this methodology for code-level detection.
|
|
// When adding items here, also update review/design-checklist.md, and vice versa.
|
|
export function generateDesignMethodology(_ctx: TemplateContext): string {
|
|
return `## Modes
|
|
|
|
### Full (default)
|
|
Systematic review of all pages reachable from homepage. Visit 5-8 pages. Full checklist evaluation, responsive screenshots, interaction flow testing. Produces complete design audit report with letter grades.
|
|
|
|
### Quick (\`--quick\`)
|
|
Homepage + 2 key pages only. First Impression + Design System Extraction + abbreviated checklist. Fastest path to a design score.
|
|
|
|
### Deep (\`--deep\`)
|
|
Comprehensive review: 10-15 pages, every interaction flow, exhaustive checklist. For pre-launch audits or major redesigns.
|
|
|
|
### Diff-aware (automatic when on a feature branch with no URL)
|
|
When on a feature branch, scope to pages affected by the branch changes:
|
|
1. Analyze the branch diff: \`git diff main...HEAD --name-only\`
|
|
2. Map changed files to affected pages/routes
|
|
3. Detect running app on common local ports (3000, 4000, 8080)
|
|
4. Audit only affected pages, compare design quality before/after
|
|
|
|
### Regression (\`--regression\` or previous \`design-baseline.json\` found)
|
|
Run full audit, then load previous \`design-baseline.json\`. Compare: per-category grade deltas, new findings, resolved findings. Output regression table in report.
|
|
|
|
---
|
|
|
|
## Phase 1: First Impression
|
|
|
|
The most uniquely designer-like output. Form a gut reaction before analyzing anything.
|
|
|
|
1. Navigate to the target URL
|
|
2. Take a full-page desktop screenshot: \`$B screenshot "$REPORT_DIR/screenshots/first-impression.png"\`
|
|
3. Write the **First Impression** using this structured critique format:
|
|
- "The site communicates **[what]**." (what it says at a glance — competence? playfulness? confusion?)
|
|
- "I notice **[observation]**." (what stands out, positive or negative — be specific)
|
|
- "The first 3 things my eye goes to are: **[1]**, **[2]**, **[3]**." (hierarchy check — are these the 3 things the designer intended? If not, the visual hierarchy is lying.)
|
|
- "If I had to describe this in one word: **[word]**." (gut verdict)
|
|
|
|
**Narration mode:** Write this section in first person, as if you are a user scanning the page for the first time. "I'm looking at this page... my eye goes to the logo, then a wall of text I skip entirely, then... wait, is that a button?" Name the specific element, its position, its visual weight. If you can't name it specifically, you're not actually scanning, you're generating platitudes.
|
|
|
|
**Page Area Test:** Point at each clearly defined area of the page. Can you instantly name its purpose? ("Things I can buy," "Today's deals," "How to search.") Areas you can't name in 2 seconds are poorly defined. List them.
|
|
|
|
This is the section users read first. Be opinionated. A designer doesn't hedge — they react.
|
|
|
|
---
|
|
|
|
## Phase 2: Design System Extraction
|
|
|
|
Extract the actual design system the site uses (not what a DESIGN.md says, but what's rendered):
|
|
|
|
\`\`\`bash
|
|
# Fonts in use (capped at 500 elements to avoid timeout)
|
|
$B js "JSON.stringify([...new Set([...document.querySelectorAll('*')].slice(0,500).map(e => getComputedStyle(e).fontFamily))])"
|
|
|
|
# Color palette in use
|
|
$B js "JSON.stringify([...new Set([...document.querySelectorAll('*')].slice(0,500).flatMap(e => [getComputedStyle(e).color, getComputedStyle(e).backgroundColor]).filter(c => c !== 'rgba(0, 0, 0, 0)'))])"
|
|
|
|
# Heading hierarchy
|
|
$B js "JSON.stringify([...document.querySelectorAll('h1,h2,h3,h4,h5,h6')].map(h => ({tag:h.tagName, text:h.textContent.trim().slice(0,50), size:getComputedStyle(h).fontSize, weight:getComputedStyle(h).fontWeight})))"
|
|
|
|
# Touch target audit (find undersized interactive elements)
|
|
$B js "JSON.stringify([...document.querySelectorAll('a,button,input,[role=button]')].filter(e => {const r=e.getBoundingClientRect(); return r.width>0 && (r.width<44||r.height<44)}).map(e => ({tag:e.tagName, text:(e.textContent||'').trim().slice(0,30), w:Math.round(e.getBoundingClientRect().width), h:Math.round(e.getBoundingClientRect().height)})).slice(0,20))"
|
|
|
|
# Performance baseline
|
|
$B perf
|
|
\`\`\`
|
|
|
|
Structure findings as an **Inferred Design System**:
|
|
- **Fonts:** list with usage counts. Flag if >3 distinct font families.
|
|
- **Colors:** palette extracted. Flag if >12 unique non-gray colors. Note warm/cool/mixed.
|
|
- **Heading Scale:** h1-h6 sizes. Flag skipped levels, non-systematic size jumps.
|
|
- **Spacing Patterns:** sample padding/margin values. Flag non-scale values.
|
|
|
|
After extraction, offer: *"Want me to save this as your DESIGN.md? I can lock in these observations as your project's design system baseline."*
|
|
|
|
---
|
|
|
|
## Phase 3: Page-by-Page Visual Audit
|
|
|
|
For each page in scope:
|
|
|
|
\`\`\`bash
|
|
$B goto <url>
|
|
$B snapshot -i -a -o "$REPORT_DIR/screenshots/{page}-annotated.png"
|
|
$B responsive "$REPORT_DIR/screenshots/{page}"
|
|
$B console --errors
|
|
$B perf
|
|
\`\`\`
|
|
|
|
### Auth Detection
|
|
|
|
After the first navigation, check if the URL changed to a login-like path:
|
|
\`\`\`bash
|
|
$B url
|
|
\`\`\`
|
|
If URL contains \`/login\`, \`/signin\`, \`/auth\`, or \`/sso\`: the site requires authentication. AskUserQuestion: "This site requires authentication. Want to import cookies from your browser? Run \`/setup-browser-cookies\` first if needed."
|
|
|
|
### Trunk Test (run on every page)
|
|
|
|
Imagine being dropped on this page with no context. Can you immediately answer:
|
|
1. What site is this? (Site ID visible and identifiable)
|
|
2. What page am I on? (Page name prominent, matches what I clicked)
|
|
3. What are the major sections? (Primary nav visible and clear)
|
|
4. What are my options at this level? (Local nav or content choices obvious)
|
|
5. Where am I in the scheme of things? ("You are here" indicator, breadcrumbs)
|
|
6. How can I search? (Search box findable without hunting)
|
|
|
|
Score: PASS (all 6 clear) / PARTIAL (4-5 clear) / FAIL (3 or fewer clear).
|
|
A FAIL on the trunk test is a HIGH-impact finding regardless of how polished the visual design is.
|
|
|
|
### Design Audit Checklist (10 categories, ~80 items)
|
|
|
|
Apply these at each page. Each finding gets an impact rating (high/medium/polish) and category.
|
|
|
|
**1. Visual Hierarchy & Composition** (8 items)
|
|
- Clear focal point? One primary CTA per view?
|
|
- Eye flows naturally top-left to bottom-right?
|
|
- Visual noise — competing elements fighting for attention?
|
|
- Information density appropriate for content type?
|
|
- Z-index clarity — nothing unexpectedly overlapping?
|
|
- Above-the-fold content communicates purpose in 3 seconds?
|
|
- Squint test: hierarchy still visible when blurred?
|
|
- White space is intentional, not leftover?
|
|
|
|
**2. Typography** (15 items)
|
|
- Font count <=3 (flag if more)
|
|
- Scale follows ratio (1.25 major third or 1.333 perfect fourth)
|
|
- Line-height: 1.5x body, 1.15-1.25x headings
|
|
- Measure: 45-75 chars per line (66 ideal)
|
|
- Heading hierarchy: no skipped levels (h1→h3 without h2)
|
|
- Weight contrast: >=2 weights used for hierarchy
|
|
- No blacklisted fonts (Papyrus, Comic Sans, Lobster, Impact, Jokerman)
|
|
- If primary font is Inter/Roboto/Open Sans/Poppins → flag as potentially generic
|
|
- \`text-wrap: balance\` or \`text-pretty\` on headings (check via \`$B css <heading> text-wrap\`)
|
|
- Curly quotes used, not straight quotes
|
|
- Ellipsis character (\`…\`) not three dots (\`...\`)
|
|
- \`font-variant-numeric: tabular-nums\` on number columns
|
|
- Body text >= 16px
|
|
- Caption/label >= 12px
|
|
- No letterspacing on lowercase text
|
|
|
|
**3. Color & Contrast** (10 items)
|
|
- Palette coherent (<=12 unique non-gray colors)
|
|
- WCAG AA: body text 4.5:1, large text (18px+) 3:1, UI components 3:1
|
|
- Semantic colors consistent (success=green, error=red, warning=yellow/amber)
|
|
- No color-only encoding (always add labels, icons, or patterns)
|
|
- Dark mode: surfaces use elevation, not just lightness inversion
|
|
- Dark mode: text off-white (~#E0E0E0), not pure white
|
|
- Primary accent desaturated 10-20% in dark mode
|
|
- \`color-scheme: dark\` on html element (if dark mode present)
|
|
- No red/green only combinations (8% of men have red-green deficiency)
|
|
- Neutral palette is warm or cool consistently — not mixed
|
|
|
|
**4. Spacing & Layout** (12 items)
|
|
- Grid consistent at all breakpoints
|
|
- Spacing uses a scale (4px or 8px base), not arbitrary values
|
|
- Alignment is consistent — nothing floats outside the grid
|
|
- Rhythm: related items closer together, distinct sections further apart
|
|
- Border-radius hierarchy (not uniform bubbly radius on everything)
|
|
- Inner radius = outer radius - gap (nested elements)
|
|
- No horizontal scroll on mobile
|
|
- Max content width set (no full-bleed body text)
|
|
- \`env(safe-area-inset-*)\` for notch devices
|
|
- URL reflects state (filters, tabs, pagination in query params)
|
|
- Flex/grid used for layout (not JS measurement)
|
|
- Breakpoints: mobile (375), tablet (768), desktop (1024), wide (1440)
|
|
|
|
**5. Interaction States** (10 items)
|
|
- Hover state on all interactive elements
|
|
- \`focus-visible\` ring present (never \`outline: none\` without replacement)
|
|
- Active/pressed state with depth effect or color shift
|
|
- Disabled state: reduced opacity + \`cursor: not-allowed\`
|
|
- Loading: skeleton shapes match real content layout
|
|
- Empty states: warm message + primary action + visual (not just "No items.")
|
|
- Error messages: specific + include fix/next step
|
|
- Success: confirmation animation or color, auto-dismiss
|
|
- Touch targets >= 44px on all interactive elements
|
|
- \`cursor: pointer\` on all clickable elements
|
|
- Mindless choice audit: every decision point (button, link, dropdown, modal choice) is a mindless click (obvious what happens). If a click requires thought about whether it's the right choice, flag as HIGH.
|
|
|
|
**6. Responsive Design** (8 items)
|
|
- Mobile layout makes *design* sense (not just stacked desktop columns)
|
|
- Touch targets sufficient on mobile (>= 44px)
|
|
- No horizontal scroll on any viewport
|
|
- Images handle responsive (srcset, sizes, or CSS containment)
|
|
- Text readable without zooming on mobile (>= 16px body)
|
|
- Navigation collapses appropriately (hamburger, bottom nav, etc.)
|
|
- Forms usable on mobile (correct input types, no autoFocus on mobile)
|
|
- No \`user-scalable=no\` or \`maximum-scale=1\` in viewport meta
|
|
|
|
**7. Motion & Animation** (6 items)
|
|
- Easing: ease-out for entering, ease-in for exiting, ease-in-out for moving
|
|
- Duration: 50-700ms range (nothing slower unless page transition)
|
|
- Purpose: every animation communicates something (state change, attention, spatial relationship)
|
|
- \`prefers-reduced-motion\` respected (check: \`$B js "matchMedia('(prefers-reduced-motion: reduce)').matches"\`)
|
|
- No \`transition: all\` — properties listed explicitly
|
|
- Only \`transform\` and \`opacity\` animated (not layout properties like width, height, top, left)
|
|
|
|
**8. Content & Microcopy** (8 items)
|
|
- Empty states designed with warmth (message + action + illustration/icon)
|
|
- Error messages specific: what happened + why + what to do next
|
|
- Button labels specific ("Save API Key" not "Continue" or "Submit")
|
|
- No placeholder/lorem ipsum text visible in production
|
|
- Truncation handled (\`text-overflow: ellipsis\`, \`line-clamp\`, or \`break-words\`)
|
|
- Active voice ("Install the CLI" not "The CLI will be installed")
|
|
- Loading states end with \`…\` ("Saving…" not "Saving...")
|
|
- Destructive actions have confirmation modal or undo window
|
|
- Happy talk detection: scan for introductory paragraphs that start with "Welcome to..." or tell users how great the site is. If you can hear "blah blah blah", it's happy talk. Flag for removal.
|
|
- Instructions detection: any visible instructions longer than one sentence. If users need to read instructions, the design has failed. Flag the instructions AND the interaction they're compensating for.
|
|
- Happy talk word count: count total visible words on the page. Classify each text block as "useful content" vs "happy talk" (welcome paragraphs, self-congratulatory text, instructions nobody reads). Report: "This page has X words. Y (Z%) are happy talk."
|
|
|
|
**9. AI Slop Detection** (10 anti-patterns — the blacklist)
|
|
|
|
The test: would a human designer at a respected studio ever ship this?
|
|
|
|
${AI_SLOP_BLACKLIST.map(item => `- ${item}`).join('\n')}
|
|
|
|
**10. Performance as Design** (6 items)
|
|
- LCP < 2.0s (web apps), < 1.5s (informational sites)
|
|
- CLS < 0.1 (no visible layout shifts during load)
|
|
- Skeleton quality: shapes match real content layout, shimmer animation
|
|
- Images: \`loading="lazy"\`, width/height dimensions set, WebP/AVIF format
|
|
- Fonts: \`font-display: swap\`, preconnect to CDN origins
|
|
- No visible font swap flash (FOUT) — critical fonts preloaded
|
|
|
|
---
|
|
|
|
## Phase 4: Interaction Flow Review
|
|
|
|
Walk 2-3 key user flows and evaluate the *feel*, not just the function:
|
|
|
|
\`\`\`bash
|
|
$B snapshot -i
|
|
$B click @e3 # perform action
|
|
$B snapshot -D # diff to see what changed
|
|
\`\`\`
|
|
|
|
Evaluate:
|
|
- **Response feel:** Does clicking feel responsive? Any delays or missing loading states?
|
|
- **Transition quality:** Are transitions intentional or generic/absent?
|
|
- **Feedback clarity:** Did the action clearly succeed or fail? Is the feedback immediate?
|
|
- **Form polish:** Focus states visible? Validation timing correct? Errors near the source?
|
|
|
|
**Narration mode:** Narrate the flow in first person. "I click 'Sign Up'... spinner appears... 3 seconds pass... still spinning... I'm getting nervous. Finally the dashboard loads, but where am I? The nav doesn't highlight anything." Name the specific element, its position, its visual weight. If you can't name it specifically, you're not actually experiencing the flow, you're generating platitudes.
|
|
|
|
### Goodwill Reservoir (track across the flow)
|
|
|
|
As you walk the user flow, maintain a mental goodwill meter (starts at 70/100).
|
|
These scores are heuristic, not measured. The value is in identifying specific
|
|
drains and fills, not in the final number.
|
|
|
|
Subtract points for:
|
|
- Hidden information the user would want (pricing, contact, shipping): subtract 15
|
|
- Format punishment (rejecting valid input like dashes in phone numbers): subtract 10
|
|
- Unnecessary information requests: subtract 10
|
|
- Interstitials, splash screens, forced tours blocking the task: subtract 15
|
|
- Sloppy or unprofessional appearance: subtract 10
|
|
- Ambiguous choices that require thinking: subtract 5 each
|
|
|
|
Add points for:
|
|
- Top user tasks are obvious and prominent: add 10
|
|
- Upfront about costs and limitations: add 5
|
|
- Saves steps (direct links, smart defaults, autofill): add 5 each
|
|
- Graceful error recovery with specific fix instructions: add 10
|
|
- Apologizes when things go wrong: add 5
|
|
|
|
Report the final goodwill score with a visual dashboard:
|
|
|
|
\`\`\`
|
|
Goodwill: 70 ████████████████████░░░░░░░░░░
|
|
Step 1: Login page 70 → 75 (+5 obvious primary action)
|
|
Step 2: Dashboard 75 → 60 (-15 interstitial tour popup)
|
|
Step 3: Settings 60 → 50 (-10 format punishment on phone)
|
|
Step 4: Billing 50 → 35 (-15 hidden pricing info)
|
|
FINAL: 35/100 ⚠️ CRITICAL UX DEBT
|
|
\`\`\`
|
|
|
|
Below 30 = critical UX debt. 30-60 = needs work. Above 60 = healthy.
|
|
Include the biggest drains and fills as specific findings.
|
|
|
|
---
|
|
|
|
## Phase 5: Cross-Page Consistency
|
|
|
|
Compare screenshots and observations across pages for:
|
|
- Navigation bar consistent across all pages?
|
|
- Footer consistent?
|
|
- Component reuse vs one-off designs (same button styled differently on different pages?)
|
|
- Tone consistency (one page playful while another is corporate?)
|
|
- Spacing rhythm carries across pages?
|
|
|
|
---
|
|
|
|
## Phase 6: Compile Report
|
|
|
|
### Output Locations
|
|
|
|
**Local:** \`.gstack/design-reports/design-audit-{domain}-{YYYY-MM-DD}.md\`
|
|
|
|
**Project-scoped:**
|
|
\`\`\`bash
|
|
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)" && mkdir -p ~/.gstack/projects/$SLUG
|
|
\`\`\`
|
|
Write to: \`~/.gstack/projects/{slug}/{user}-{branch}-design-audit-{datetime}.md\`
|
|
|
|
**Baseline:** Write \`design-baseline.json\` for regression mode:
|
|
\`\`\`json
|
|
{
|
|
"date": "YYYY-MM-DD",
|
|
"url": "<target>",
|
|
"designScore": "B",
|
|
"aiSlopScore": "C",
|
|
"categoryGrades": { "hierarchy": "A", "typography": "B", ... },
|
|
"findings": [{ "id": "FINDING-001", "title": "...", "impact": "high", "category": "typography" }]
|
|
}
|
|
\`\`\`
|
|
|
|
### Scoring System
|
|
|
|
**Dual headline scores:**
|
|
- **Design Score: {A-F}** — weighted average of all 10 categories
|
|
- **AI Slop Score: {A-F}** — standalone grade with pithy verdict
|
|
|
|
**Per-category grades:**
|
|
- **A:** Intentional, polished, delightful. Shows design thinking.
|
|
- **B:** Solid fundamentals, minor inconsistencies. Looks professional.
|
|
- **C:** Functional but generic. No major problems, no design point of view.
|
|
- **D:** Noticeable problems. Feels unfinished or careless.
|
|
- **F:** Actively hurting user experience. Needs significant rework.
|
|
|
|
**Grade computation:** Each category starts at A. Each High-impact finding drops one letter grade. Each Medium-impact finding drops half a letter grade. Polish findings are noted but do not affect grade. Minimum is F.
|
|
|
|
**Category weights for Design Score:**
|
|
| Category | Weight |
|
|
|----------|--------|
|
|
| Visual Hierarchy | 15% |
|
|
| Typography | 15% |
|
|
| Spacing & Layout | 15% |
|
|
| Color & Contrast | 10% |
|
|
| Interaction States | 10% |
|
|
| Responsive | 10% |
|
|
| Content Quality | 10% |
|
|
| AI Slop | 5% |
|
|
| Motion | 5% |
|
|
| Performance Feel | 5% |
|
|
|
|
AI Slop is 5% of Design Score but also graded independently as a headline metric.
|
|
|
|
### Regression Output
|
|
|
|
When previous \`design-baseline.json\` exists or \`--regression\` flag is used:
|
|
- Load baseline grades
|
|
- Compare: per-category deltas, new findings, resolved findings
|
|
- Append regression table to report
|
|
|
|
---
|
|
|
|
## Design Critique Format
|
|
|
|
Use structured feedback, not opinions:
|
|
- "I notice..." — observation (e.g., "I notice the primary CTA competes with the secondary action")
|
|
- "I wonder..." — question (e.g., "I wonder if users will understand what 'Process' means here")
|
|
- "What if..." — suggestion (e.g., "What if we moved search to a more prominent position?")
|
|
- "I think... because..." — reasoned opinion (e.g., "I think the spacing between sections is too uniform because it doesn't create hierarchy")
|
|
|
|
Tie everything to user goals and product objectives. Always suggest specific improvements alongside problems.
|
|
|
|
---
|
|
|
|
## Important Rules
|
|
|
|
1. **Think like a designer, not a QA engineer.** You care whether things feel right, look intentional, and respect the user. You do NOT just care whether things "work."
|
|
2. **Screenshots are evidence.** Every finding needs at least one screenshot. Use annotated screenshots (\`snapshot -a\`) to highlight elements.
|
|
3. **Be specific and actionable.** "Change X to Y because Z" — not "the spacing feels off."
|
|
4. **Never read source code.** Evaluate the rendered site, not the implementation. (Exception: offer to write DESIGN.md from extracted observations.)
|
|
5. **AI Slop detection is your superpower.** Most developers can't evaluate whether their site looks AI-generated. You can. Be direct about it.
|
|
6. **Quick wins matter.** Always include a "Quick Wins" section — the 3-5 highest-impact fixes that take <30 minutes each.
|
|
7. **Use \`snapshot -C\` for tricky UIs.** Finds clickable divs that the accessibility tree misses.
|
|
8. **Responsive is design, not just "not broken."** A stacked desktop layout on mobile is not responsive design — it's lazy. Evaluate whether the mobile layout makes *design* sense.
|
|
9. **Document incrementally.** Write each finding to the report as you find it. Don't batch.
|
|
10. **Depth over breadth.** 5-10 well-documented findings with screenshots and specific suggestions > 20 vague observations.
|
|
11. **Show screenshots to the user.** After every \`$B screenshot\`, \`$B snapshot -a -o\`, or \`$B responsive\` command, use the Read tool on the output file(s) so the user can see them inline. For \`responsive\` (3 files), Read all three. This is critical — without it, screenshots are invisible to the user.`;
|
|
}
|
|
|
|
export function generateDesignSketch(_ctx: TemplateContext): string {
|
|
return `## Visual Sketch (UI ideas only)
|
|
|
|
If the chosen approach involves user-facing UI (screens, pages, forms, dashboards,
|
|
or interactive elements), generate a rough wireframe to help the user visualize it.
|
|
If the idea is backend-only, infrastructure, or has no UI component — skip this
|
|
section silently.
|
|
|
|
**Step 1: Gather design context**
|
|
|
|
1. Check if \`DESIGN.md\` exists in the repo root. If it does, read it for design
|
|
system constraints (colors, typography, spacing, component patterns). Use these
|
|
constraints in the wireframe.
|
|
2. Apply core design principles:
|
|
- **Information hierarchy** — what does the user see first, second, third?
|
|
- **Interaction states** — loading, empty, error, success, partial
|
|
- **Edge case paranoia** — what if the name is 47 chars? Zero results? Network fails?
|
|
- **Subtraction default** — "as little design as possible" (Rams). Every element earns its pixels.
|
|
- **Design for trust** — every interface element builds or erodes user trust.
|
|
|
|
**Step 2: Generate wireframe HTML**
|
|
|
|
Generate a single-page HTML file with these constraints:
|
|
- **Intentionally rough aesthetic** — use system fonts, thin gray borders, no color,
|
|
hand-drawn-style elements. This is a sketch, not a polished mockup.
|
|
- Self-contained — no external dependencies, no CDN links, inline CSS only
|
|
- Show the core interaction flow (1-3 screens/states max)
|
|
- Include realistic placeholder content (not "Lorem ipsum" — use content that
|
|
matches the actual use case)
|
|
- Add HTML comments explaining design decisions
|
|
|
|
Write to a temp file:
|
|
\`\`\`bash
|
|
SKETCH_FILE="/tmp/gstack-sketch-$(date +%s).html"
|
|
\`\`\`
|
|
|
|
**Step 3: Render and capture**
|
|
|
|
\`\`\`bash
|
|
$B goto "file://$SKETCH_FILE"
|
|
$B screenshot /tmp/gstack-sketch.png
|
|
\`\`\`
|
|
|
|
If \`$B\` is not available (browse binary not set up), skip the render step. Tell the
|
|
user: "Visual sketch requires the browse binary. Run the setup script to enable it."
|
|
|
|
**Step 4: Present and iterate**
|
|
|
|
Show the screenshot to the user. Ask: "Does this feel right? Want to iterate on the layout?"
|
|
|
|
If they want changes, regenerate the HTML with their feedback and re-render.
|
|
If they approve or say "good enough," proceed.
|
|
|
|
**Step 5: Include in design doc**
|
|
|
|
Reference the wireframe screenshot in the design doc's "Recommended Approach" section.
|
|
The screenshot file at \`/tmp/gstack-sketch.png\` can be referenced by downstream skills
|
|
(\`/plan-design-review\`, \`/design-review\`) to see what was originally envisioned.
|
|
|
|
**Step 6: Outside design voices** (optional)
|
|
|
|
After the wireframe is approved, offer outside design perspectives:
|
|
|
|
\`\`\`bash
|
|
command -v codex >/dev/null 2>&1 && echo "CODEX_AVAILABLE" || echo "CODEX_NOT_AVAILABLE"
|
|
\`\`\`
|
|
|
|
If Codex is available, use AskUserQuestion:
|
|
> "Want outside design perspectives on the chosen approach? Codex proposes a visual thesis, content plan, and interaction ideas. A Claude subagent proposes an alternative aesthetic direction."
|
|
>
|
|
> A) Yes — get outside design voices
|
|
> B) No — proceed without
|
|
|
|
If user chooses A, launch both voices simultaneously:
|
|
|
|
1. **Codex** (via Bash, \`model_reasoning_effort="medium"\`):
|
|
\`\`\`bash
|
|
TMPERR_SKETCH=$(mktemp /tmp/codex-sketch-XXXXXXXX)
|
|
_REPO_ROOT=$(git rev-parse --show-toplevel) || { echo "ERROR: not in a git repo" >&2; exit 1; }
|
|
codex exec "For this product approach, provide: a visual thesis (one sentence — mood, material, energy), a content plan (hero → support → detail → CTA), and 2 interaction ideas that change page feel. Apply beautiful defaults: composition-first, brand-first, cardless, poster not document. Be opinionated." -C "$_REPO_ROOT" -s read-only -c 'model_reasoning_effort="medium"' --enable web_search_cached < /dev/null 2>"$TMPERR_SKETCH"
|
|
\`\`\`
|
|
Use a 5-minute timeout (\`timeout: 300000\`). After completion: \`cat "$TMPERR_SKETCH" && rm -f "$TMPERR_SKETCH"\`
|
|
|
|
2. **Claude subagent** (via Agent tool):
|
|
"For this product approach, what design direction would you recommend? What aesthetic, typography, and interaction patterns fit? What would make this approach feel inevitable to the user? Be specific — font names, hex colors, spacing values."
|
|
|
|
Present Codex output under \`CODEX SAYS (design sketch):\` and subagent output under \`CLAUDE SUBAGENT (design direction):\`.
|
|
Error handling: all non-blocking. On failure, skip and continue.`;
|
|
}
|
|
|
|
export function generateDesignOutsideVoices(ctx: TemplateContext): string {
|
|
// Codex host: strip entirely — Codex should never invoke itself
|
|
if (ctx.host === 'codex') return '';
|
|
|
|
const rejectionList = OPENAI_HARD_REJECTIONS.map((item, i) => `${i + 1}. ${item}`).join('\n');
|
|
const litmusList = OPENAI_LITMUS_CHECKS.map((item, i) => `${i + 1}. ${item}`).join('\n');
|
|
|
|
// Skill-specific configuration
|
|
const isPlanDesignReview = ctx.skillName === 'plan-design-review';
|
|
const isDesignReview = ctx.skillName === 'design-review';
|
|
const isDesignConsultation = ctx.skillName === 'design-consultation';
|
|
|
|
// Determine opt-in behavior and reasoning effort
|
|
const isAutomatic = isDesignReview; // design-review runs automatically
|
|
const reasoningEffort = isDesignConsultation ? 'medium' : 'high'; // creative vs analytical
|
|
|
|
// Build skill-specific Codex prompt
|
|
let codexPrompt: string;
|
|
let subagentPrompt: string;
|
|
|
|
if (isPlanDesignReview) {
|
|
codexPrompt = `Read the plan file at [plan-file-path]. Evaluate this plan's UI/UX design against these criteria.
|
|
|
|
HARD REJECTION — flag if ANY apply:
|
|
${rejectionList}
|
|
|
|
LITMUS CHECKS — answer YES or NO for each:
|
|
${litmusList}
|
|
|
|
HARD RULES — first classify as MARKETING/LANDING PAGE vs APP UI vs HYBRID, then flag violations of the matching rule set:
|
|
- MARKETING: First viewport as one composition, brand-first hierarchy, full-bleed hero, 2-3 intentional motions, composition-first layout
|
|
- APP UI: Calm surface hierarchy, dense but readable, utility language, minimal chrome
|
|
- UNIVERSAL: CSS variables for colors, no default font stacks, one job per section, cards earn existence
|
|
|
|
For each finding: what's wrong, what will happen if it ships unresolved, and the specific fix. Be opinionated. No hedging.`;
|
|
|
|
subagentPrompt = `Read the plan file at [plan-file-path]. You are an independent senior product designer reviewing this plan. You have NOT seen any prior review. Evaluate:
|
|
|
|
1. Information hierarchy: what does the user see first, second, third? Is it right?
|
|
2. Missing states: loading, empty, error, success, partial — which are unspecified?
|
|
3. User journey: what's the emotional arc? Where does it break?
|
|
4. Specificity: does the plan describe SPECIFIC UI ("48px Söhne Bold header, #1a1a1a on white") or generic patterns ("clean modern card-based layout")?
|
|
5. What design decisions will haunt the implementer if left ambiguous?
|
|
|
|
For each finding: what's wrong, severity (critical/high/medium), and the fix.`;
|
|
} else if (isDesignReview) {
|
|
codexPrompt = `Review the frontend source code in this repo. Evaluate against these design hard rules:
|
|
- Spacing: systematic (design tokens / CSS variables) or magic numbers?
|
|
- Typography: expressive purposeful fonts or default stacks?
|
|
- Color: CSS variables with defined system, or hardcoded hex scattered?
|
|
- Responsive: breakpoints defined? calc(100svh - header) for heroes? Mobile tested?
|
|
- A11y: ARIA landmarks, alt text, contrast ratios, 44px touch targets?
|
|
- Motion: 2-3 intentional animations, or zero / ornamental only?
|
|
- Cards: used only when card IS the interaction? No decorative card grids?
|
|
|
|
First classify as MARKETING/LANDING PAGE vs APP UI vs HYBRID, then apply matching rules.
|
|
|
|
LITMUS CHECKS — answer YES/NO:
|
|
${litmusList}
|
|
|
|
HARD REJECTION — flag if ANY apply:
|
|
${rejectionList}
|
|
|
|
Be specific. Reference file:line for every finding.`;
|
|
|
|
subagentPrompt = `Review the frontend source code in this repo. You are an independent senior product designer doing a source-code design audit. Focus on CONSISTENCY PATTERNS across files rather than individual violations:
|
|
- Are spacing values systematic across the codebase?
|
|
- Is there ONE color system or scattered approaches?
|
|
- Do responsive breakpoints follow a consistent set?
|
|
- Is the accessibility approach consistent or spotty?
|
|
|
|
For each finding: what's wrong, severity (critical/high/medium), and the file:line.`;
|
|
} else if (isDesignConsultation) {
|
|
codexPrompt = `Given this product context, propose a complete design direction:
|
|
- Visual thesis: one sentence describing mood, material, and energy
|
|
- Typography: specific font names (not defaults — no Inter/Roboto/Arial/system) + hex colors
|
|
- Color system: CSS variables for background, surface, primary text, muted text, accent
|
|
- Layout: composition-first, not component-first. First viewport as poster, not document
|
|
- Differentiation: 2 deliberate departures from category norms
|
|
- Anti-slop: no purple gradients, no 3-column icon grids, no centered everything, no decorative blobs
|
|
|
|
Be opinionated. Be specific. Do not hedge. This is YOUR design direction — own it.`;
|
|
|
|
subagentPrompt = `Given this product context, propose a design direction that would SURPRISE. What would the cool indie studio do that the enterprise UI team wouldn't?
|
|
- Propose an aesthetic direction, typography stack (specific font names), color palette (hex values)
|
|
- 2 deliberate departures from category norms
|
|
- What emotional reaction should the user have in the first 3 seconds?
|
|
|
|
Be bold. Be specific. No hedging.`;
|
|
} else {
|
|
// Unknown skill — return empty
|
|
return '';
|
|
}
|
|
|
|
// Build the opt-in section
|
|
const optInSection = isAutomatic ? `
|
|
**Automatic:** Outside voices run automatically when Codex is available. No opt-in needed.` : `
|
|
Use AskUserQuestion:
|
|
> "Want outside design voices${isPlanDesignReview ? ' before the detailed review' : ''}? Codex evaluates against OpenAI's design hard rules + litmus checks; Claude subagent does an independent ${isDesignConsultation ? 'design direction proposal' : 'completeness review'}."
|
|
>
|
|
> A) Yes — run outside design voices
|
|
> B) No — proceed without
|
|
|
|
If user chooses B, skip this step and continue.`;
|
|
|
|
// Build the synthesis section
|
|
const synthesisSection = isPlanDesignReview ? `
|
|
**Synthesis — Litmus scorecard:**
|
|
|
|
\`\`\`
|
|
DESIGN OUTSIDE VOICES — LITMUS SCORECARD:
|
|
═══════════════════════════════════════════════════════════════
|
|
Check Claude Codex Consensus
|
|
─────────────────────────────────────── ─────── ─────── ─────────
|
|
1. Brand unmistakable in first screen? — — —
|
|
2. One strong visual anchor? — — —
|
|
3. Scannable by headlines only? — — —
|
|
4. Each section has one job? — — —
|
|
5. Cards actually necessary? — — —
|
|
6. Motion improves hierarchy? — — —
|
|
7. Premium without decorative shadows? — — —
|
|
─────────────────────────────────────── ─────── ─────── ─────────
|
|
Hard rejections triggered: — — —
|
|
═══════════════════════════════════════════════════════════════
|
|
\`\`\`
|
|
|
|
Fill in each cell from the Codex and subagent outputs. CONFIRMED = both agree. DISAGREE = models differ. NOT SPEC'D = not enough info to evaluate.
|
|
|
|
**Pass integration (respects existing 7-pass contract):**
|
|
- Hard rejections → raised as the FIRST items in Pass 1, tagged \`[HARD REJECTION]\`
|
|
- Litmus DISAGREE items → raised in the relevant pass with both perspectives
|
|
- Litmus CONFIRMED failures → pre-loaded as known issues in the relevant pass
|
|
- Passes can skip discovery and go straight to fixing for pre-identified issues` :
|
|
isDesignConsultation ? `
|
|
**Synthesis:** Claude main references both Codex and subagent proposals in the Phase 3 proposal. Present:
|
|
- Areas of agreement between all three voices (Claude main + Codex + subagent)
|
|
- Genuine divergences as creative alternatives for the user to choose from
|
|
- "Codex and I agree on X. Codex suggested Y where I'm proposing Z — here's why..."` : `
|
|
**Synthesis — Litmus scorecard:**
|
|
|
|
Use the same scorecard format as /plan-design-review (shown above). Fill in from both outputs.
|
|
Merge findings into the triage with \`[codex]\` / \`[subagent]\` / \`[cross-model]\` tags.`;
|
|
|
|
const escapedCodexPrompt = codexPrompt.replace(/`/g, '\\`').replace(/\$/g, '\\$');
|
|
|
|
return `## Design Outside Voices (parallel)
|
|
${optInSection}
|
|
|
|
**Check Codex availability:**
|
|
\`\`\`bash
|
|
command -v codex >/dev/null 2>&1 && echo "CODEX_AVAILABLE" || echo "CODEX_NOT_AVAILABLE"
|
|
\`\`\`
|
|
|
|
**If Codex is available**, launch both voices simultaneously:
|
|
|
|
1. **Codex design voice** (via Bash):
|
|
\`\`\`bash
|
|
TMPERR_DESIGN=$(mktemp /tmp/codex-design-XXXXXXXX)
|
|
_REPO_ROOT=$(git rev-parse --show-toplevel) || { echo "ERROR: not in a git repo" >&2; exit 1; }
|
|
codex exec "${escapedCodexPrompt}" -C "$_REPO_ROOT" -s read-only -c 'model_reasoning_effort="${reasoningEffort}"' --enable web_search_cached < /dev/null 2>"$TMPERR_DESIGN"
|
|
\`\`\`
|
|
Use a 5-minute timeout (\`timeout: 300000\`). After the command completes, read stderr:
|
|
\`\`\`bash
|
|
cat "$TMPERR_DESIGN" && rm -f "$TMPERR_DESIGN"
|
|
\`\`\`
|
|
|
|
2. **Claude design subagent** (via Agent tool):
|
|
Dispatch a subagent with this prompt:
|
|
"${subagentPrompt}"
|
|
|
|
**Error handling (all non-blocking):**
|
|
- **Auth failure:** If stderr contains "auth", "login", "unauthorized", or "API key": "Codex authentication failed. Run \`codex login\` to authenticate."
|
|
- **Timeout:** "Codex timed out after 5 minutes."
|
|
- **Empty response:** "Codex returned no response."
|
|
- On any Codex error: proceed with Claude subagent output only, tagged \`[single-model]\`.
|
|
- If Claude subagent also fails: "Outside voices unavailable — continuing with primary review."
|
|
|
|
Present Codex output under a \`CODEX SAYS (design ${isPlanDesignReview ? 'critique' : isDesignReview ? 'source audit' : 'direction'}):\` header.
|
|
Present subagent output under a \`CLAUDE SUBAGENT (design ${isPlanDesignReview ? 'completeness' : isDesignReview ? 'consistency' : 'direction'}):\` header.
|
|
${synthesisSection}
|
|
|
|
**Log the result:**
|
|
\`\`\`bash
|
|
${ctx.paths.binDir}/gstack-review-log '{"skill":"design-outside-voices","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","status":"STATUS","source":"SOURCE","commit":"'"$(git rev-parse --short HEAD)"'"}'
|
|
\`\`\`
|
|
Replace STATUS with "clean" or "issues_found", SOURCE with "codex+subagent", "codex-only", "subagent-only", or "unavailable".`;
|
|
}
|
|
|
|
// ─── Design Hard Rules (OpenAI framework + gstack slop blacklist) ───
|
|
export function generateDesignHardRules(_ctx: TemplateContext): string {
|
|
const slopItems = AI_SLOP_BLACKLIST.map((item, i) => `${i + 1}. ${item}`).join('\n');
|
|
const rejectionItems = OPENAI_HARD_REJECTIONS.map((item, i) => `${i + 1}. ${item}`).join('\n');
|
|
const litmusItems = OPENAI_LITMUS_CHECKS.map((item, i) => `${i + 1}. ${item}`).join('\n');
|
|
|
|
return `### Design Hard Rules
|
|
|
|
**Classifier — determine rule set before evaluating:**
|
|
- **MARKETING/LANDING PAGE** (hero-driven, brand-forward, conversion-focused) → apply Landing Page Rules
|
|
- **APP UI** (workspace-driven, data-dense, task-focused: dashboards, admin, settings) → apply App UI Rules
|
|
- **HYBRID** (marketing shell with app-like sections) → apply Landing Page Rules to hero/marketing sections, App UI Rules to functional sections
|
|
|
|
**Hard rejection criteria** (instant-fail patterns — flag if ANY apply):
|
|
${rejectionItems}
|
|
|
|
**Litmus checks** (answer YES/NO for each — used for cross-model consensus scoring):
|
|
${litmusItems}
|
|
|
|
**Landing page rules** (apply when classifier = MARKETING/LANDING):
|
|
- First viewport reads as one composition, not a dashboard
|
|
- Brand-first hierarchy: brand > headline > body > CTA
|
|
- Typography: expressive, purposeful — no default stacks (Inter, Roboto, Arial, system)
|
|
- No flat single-color backgrounds — use gradients, images, subtle patterns
|
|
- Hero: full-bleed, edge-to-edge, no inset/tiled/rounded variants
|
|
- Hero budget: brand, one headline, one supporting sentence, one CTA group, one image
|
|
- No cards in hero. Cards only when card IS the interaction
|
|
- One job per section: one purpose, one headline, one short supporting sentence
|
|
- Motion: 2-3 intentional motions minimum (entrance, scroll-linked, hover/reveal)
|
|
- Color: define CSS variables, avoid purple-on-white defaults, one accent color default
|
|
- Copy: product language not design commentary. "If deleting 30% improves it, keep deleting"
|
|
- Beautiful defaults: composition-first, brand as loudest text, two typefaces max, cardless by default, first viewport as poster not document
|
|
|
|
**App UI rules** (apply when classifier = APP UI):
|
|
- Calm surface hierarchy, strong typography, few colors
|
|
- Dense but readable, minimal chrome
|
|
- Organize: primary workspace, navigation, secondary context, one accent
|
|
- Avoid: dashboard-card mosaics, thick borders, decorative gradients, ornamental icons
|
|
- Copy: utility language — orientation, status, action. Not mood/brand/aspiration
|
|
- Cards only when card IS the interaction
|
|
- Section headings state what area is or what user can do ("Selected KPIs", "Plan status")
|
|
|
|
**Universal rules** (apply to ALL types):
|
|
- Define CSS variables for color system
|
|
- No default font stacks (Inter, Roboto, Arial, system)
|
|
- One job per section
|
|
- "If deleting 30% of the copy improves it, keep deleting"
|
|
- Cards earn their existence — no decorative card grids
|
|
- NEVER use small, low-contrast type (body text < 16px or contrast ratio < 4.5:1 on body text)
|
|
- NEVER put labels inside form fields as the only label (placeholder-as-label pattern — labels must be visible when the field has content)
|
|
- ALWAYS preserve visited vs unvisited link distinction (visited links must have a different color)
|
|
- NEVER float headings between paragraphs (heading must be visually closer to the section it introduces than to the preceding section)
|
|
|
|
**AI Slop blacklist** (the 10 patterns that scream "AI-generated"):
|
|
${slopItems}
|
|
|
|
Source: [OpenAI "Designing Delightful Frontends with GPT-5.4"](https://developers.openai.com/blog/designing-delightful-frontends-with-gpt-5-4) (Mar 2026) + gstack design methodology.`;
|
|
}
|
|
|
|
export function generateDesignSetup(ctx: TemplateContext): string {
|
|
return `## DESIGN SETUP (run this check BEFORE any design mockup command)
|
|
|
|
\`\`\`bash
|
|
_ROOT=$(git rev-parse --show-toplevel 2>/dev/null)
|
|
D=""
|
|
[ -n "$_ROOT" ] && [ -x "$_ROOT/${ctx.paths.localSkillRoot}/design/dist/design" ] && D="$_ROOT/${ctx.paths.localSkillRoot}/design/dist/design"
|
|
[ -z "$D" ] && D="${toShellPath(ctx.paths.designDir)}/design"
|
|
if [ -x "$D" ]; then
|
|
echo "DESIGN_READY: $D"
|
|
else
|
|
echo "DESIGN_NOT_AVAILABLE"
|
|
fi
|
|
B=""
|
|
[ -n "$_ROOT" ] && [ -x "$_ROOT/${ctx.paths.localSkillRoot}/browse/dist/browse" ] && B="$_ROOT/${ctx.paths.localSkillRoot}/browse/dist/browse"
|
|
[ -z "$B" ] && B="${toShellPath(ctx.paths.browseDir)}/browse"
|
|
if [ -x "$B" ]; then
|
|
echo "BROWSE_READY: $B"
|
|
else
|
|
echo "BROWSE_NOT_AVAILABLE (will use 'open' to view comparison boards)"
|
|
fi
|
|
\`\`\`
|
|
|
|
If \`DESIGN_NOT_AVAILABLE\`: skip visual mockup generation and fall back to the
|
|
existing HTML wireframe approach (\`DESIGN_SKETCH\`). Design mockups are a
|
|
progressive enhancement, not a hard requirement.
|
|
|
|
If \`BROWSE_NOT_AVAILABLE\`: use \`open file://...\` instead of \`$B goto\` to open
|
|
comparison boards. The user just needs to see the HTML file in any browser.
|
|
|
|
If \`DESIGN_READY\`: the design binary is available for visual mockup generation.
|
|
Commands:
|
|
- \`$D generate --brief "..." --output /path.png\` — generate a single mockup
|
|
- \`$D variants --brief "..." --count 3 --output-dir /path/\` — generate N style variants
|
|
- \`$D compare --images "a.png,b.png,c.png" --output /path/board.html --serve\` — comparison board + HTTP server
|
|
- \`$D serve --html /path/board.html\` — serve comparison board and collect feedback via HTTP
|
|
- \`$D check --image /path.png --brief "..."\` — vision quality gate
|
|
- \`$D iterate --session /path/session.json --feedback "..." --output /path.png\` — iterate
|
|
|
|
**CRITICAL PATH RULE:** All design artifacts (mockups, comparison boards, approved.json)
|
|
MUST be saved to \`~/.gstack/projects/$SLUG/designs/\`, NEVER to \`.context/\`,
|
|
\`docs/designs/\`, \`/tmp/\`, or any project-local directory. Design artifacts are USER
|
|
data, not project files. They persist across branches, conversations, and workspaces.`;
|
|
}
|
|
|
|
export function generateDesignMockup(ctx: TemplateContext): string {
|
|
return `## Visual Design Exploration
|
|
|
|
\`\`\`bash
|
|
_ROOT=$(git rev-parse --show-toplevel 2>/dev/null)
|
|
D=""
|
|
[ -n "$_ROOT" ] && [ -x "$_ROOT/${ctx.paths.localSkillRoot}/design/dist/design" ] && D="$_ROOT/${ctx.paths.localSkillRoot}/design/dist/design"
|
|
[ -z "$D" ] && D="${toShellPath(ctx.paths.designDir)}/design"
|
|
[ -x "$D" ] && echo "DESIGN_READY" || echo "DESIGN_NOT_AVAILABLE"
|
|
\`\`\`
|
|
|
|
**If \`DESIGN_NOT_AVAILABLE\`:** Fall back to the HTML wireframe approach below
|
|
(the existing DESIGN_SKETCH section). Visual mockups require the design binary.
|
|
|
|
**If \`DESIGN_READY\`:** Generate visual mockup explorations for the user.
|
|
|
|
Generating visual mockups of the proposed design... (say "skip" if you don't need visuals)
|
|
|
|
**Step 1: Set up the design directory**
|
|
|
|
\`\`\`bash
|
|
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)"
|
|
_DESIGN_DIR="$HOME/.gstack/projects/$SLUG/designs/mockup-$(date +%Y%m%d)"
|
|
mkdir -p "$_DESIGN_DIR"
|
|
echo "DESIGN_DIR: $_DESIGN_DIR"
|
|
\`\`\`
|
|
|
|
**Step 2: Construct the design brief**
|
|
|
|
Read DESIGN.md if it exists — use it to constrain the visual style. If no DESIGN.md,
|
|
explore wide across diverse directions.
|
|
|
|
**Step 3: Generate 3 variants**
|
|
|
|
\`\`\`bash
|
|
$D variants --brief "<assembled brief>" --count 3 --output-dir "$_DESIGN_DIR/"
|
|
\`\`\`
|
|
|
|
This generates 3 style variations of the same brief (~40 seconds total).
|
|
|
|
**Step 4: Show variants inline, then open comparison board**
|
|
|
|
Show each variant to the user inline first (read the PNGs with Read tool), then
|
|
create and serve the comparison board:
|
|
|
|
\`\`\`bash
|
|
$D compare --images "$_DESIGN_DIR/variant-A.png,$_DESIGN_DIR/variant-B.png,$_DESIGN_DIR/variant-C.png" --output "$_DESIGN_DIR/design-board.html" --serve
|
|
\`\`\`
|
|
|
|
This opens the board in the user's default browser and blocks until feedback is
|
|
received. Read stdout for the structured JSON result. No polling needed.
|
|
|
|
If \`$D serve\` is not available or fails, fall back to AskUserQuestion:
|
|
"I've opened the design board. Which variant do you prefer? Any feedback?"
|
|
|
|
**Step 5: Handle feedback**
|
|
|
|
If the JSON contains \`"regenerated": true\`:
|
|
1. Read \`regenerateAction\` (or \`remixSpec\` for remix requests)
|
|
2. Generate new variants with \`$D iterate\` or \`$D variants\` using updated brief
|
|
3. Create new board with \`$D compare\`
|
|
4. POST the new HTML to the running board. Parse the board URL from stderr
|
|
(\`BOARD_URL: http://127.0.0.1:N/boards/<id>/\` — the daemon path) or fall
|
|
back to the legacy port (\`SERVE_STARTED: port=N\` — only emitted under
|
|
\`--no-daemon\`, hits \`/api/reload\` root). Daemon path:
|
|
\`curl -X POST "\${BOARD_URL}api/reload" -H 'Content-Type: application/json' -d '{"html":"$_DESIGN_DIR/design-board.html"}'\`
|
|
5. Board auto-refreshes in the same tab
|
|
|
|
If \`"regenerated": false\`: proceed with the approved variant.
|
|
|
|
**Step 6: Save approved choice**
|
|
|
|
\`\`\`bash
|
|
echo '{"approved_variant":"<VARIANT>","feedback":"<FEEDBACK>","date":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","screen":"mockup","branch":"'$(git branch --show-current 2>/dev/null)'"}' > "$_DESIGN_DIR/approved.json"
|
|
\`\`\`
|
|
|
|
Reference the saved mockup in the design doc or plan.`;
|
|
}
|
|
|
|
export function generateDesignShotgunLoop(_ctx: TemplateContext): string {
|
|
return `### Comparison Board + Feedback Loop
|
|
|
|
Create the comparison board and serve it over HTTP:
|
|
|
|
\`\`\`bash
|
|
$D compare --images "$_DESIGN_DIR/variant-A.png,$_DESIGN_DIR/variant-B.png,$_DESIGN_DIR/variant-C.png" --output "$_DESIGN_DIR/design-board.html" --serve
|
|
\`\`\`
|
|
|
|
This command generates the board HTML, starts an HTTP server on a random port,
|
|
and opens it in the user's default browser. **Run it in the background** with \`&\`
|
|
because the server needs to stay running while the user interacts with the board.
|
|
|
|
Parse the board URL from stderr output. Default daemon path:
|
|
\`BOARD_URL: http://127.0.0.1:N/boards/<id>/\` (already includes the per-board
|
|
path; use this for the AskUserQuestion URL AND as the base for the reload
|
|
endpoint). Legacy \`--no-daemon\` path emits \`SERVE_STARTED: port=XXXXX\` and
|
|
serves a single board at \`/\`, with reload at \`/api/reload\` — only relevant
|
|
when an external caller explicitly passes \`--no-daemon\`.
|
|
|
|
**PRIMARY WAIT: AskUserQuestion with board URL**
|
|
|
|
After the board is serving, use AskUserQuestion to wait for the user. Include the
|
|
board URL so they can click it if they lost the browser tab:
|
|
|
|
"I've opened a comparison board with the design variants:
|
|
<BOARD_URL> — Rate them, leave comments, remix
|
|
elements you like, and click Submit when you're done. Let me know when you've
|
|
submitted your feedback (or paste your preferences here). If you clicked
|
|
Regenerate or Remix on the board, tell me and I'll generate new variants."
|
|
|
|
Substitute \`<BOARD_URL>\` with the URL parsed from stderr (the daemon path
|
|
emits \`BOARD_URL: http://127.0.0.1:N/boards/<id>/\`).
|
|
|
|
**Do NOT use AskUserQuestion to ask which variant the user prefers.** The comparison
|
|
board IS the chooser. AskUserQuestion is just the blocking wait mechanism.
|
|
|
|
**After the user responds to AskUserQuestion:**
|
|
|
|
Check for feedback files next to the board HTML:
|
|
- \`$_DESIGN_DIR/feedback.json\` — written when user clicks Submit (final choice)
|
|
- \`$_DESIGN_DIR/feedback-pending.json\` — written when user clicks Regenerate/Remix/More Like This
|
|
|
|
\`\`\`bash
|
|
if [ -f "$_DESIGN_DIR/feedback.json" ]; then
|
|
echo "SUBMIT_RECEIVED"
|
|
cat "$_DESIGN_DIR/feedback.json"
|
|
elif [ -f "$_DESIGN_DIR/feedback-pending.json" ]; then
|
|
echo "REGENERATE_RECEIVED"
|
|
cat "$_DESIGN_DIR/feedback-pending.json"
|
|
rm "$_DESIGN_DIR/feedback-pending.json"
|
|
else
|
|
echo "NO_FEEDBACK_FILE"
|
|
fi
|
|
\`\`\`
|
|
|
|
The feedback JSON has this shape:
|
|
\`\`\`json
|
|
{
|
|
"preferred": "A",
|
|
"ratings": { "A": 4, "B": 3, "C": 2 },
|
|
"comments": { "A": "Love the spacing" },
|
|
"overall": "Go with A, bigger CTA",
|
|
"regenerated": false
|
|
}
|
|
\`\`\`
|
|
|
|
**If \`feedback.json\` found:** The user clicked Submit on the board.
|
|
Read \`preferred\`, \`ratings\`, \`comments\`, \`overall\` from the JSON. Proceed with
|
|
the approved variant.
|
|
|
|
**If \`feedback-pending.json\` found:** The user clicked Regenerate/Remix on the board.
|
|
1. Read \`regenerateAction\` from the JSON (\`"different"\`, \`"match"\`, \`"more_like_B"\`,
|
|
\`"remix"\`, or custom text)
|
|
2. If \`regenerateAction\` is \`"remix"\`, read \`remixSpec\` (e.g. \`{"layout":"A","colors":"B"}\`)
|
|
3. Generate new variants with \`$D iterate\` or \`$D variants\` using updated brief
|
|
4. Create new board: \`$D compare --images "..." --output "$_DESIGN_DIR/design-board.html"\`
|
|
5. Reload the board in the user's browser (same tab) — the URL is per-board
|
|
under daemon mode, so use \`<BOARD_URL>\` (from the \`BOARD_URL:\` stderr
|
|
line) as the base:
|
|
\`curl -s -X POST "\${BOARD_URL}api/reload" -H 'Content-Type: application/json' -d '{"html":"$_DESIGN_DIR/design-board.html"}'\`
|
|
Under \`--no-daemon\` the reload endpoint is \`/api/reload\` at the legacy
|
|
port; this path only matters if the caller explicitly opted out of the
|
|
daemon.
|
|
6. The board auto-refreshes. **AskUserQuestion again** with the same board URL to
|
|
wait for the next round of feedback. Repeat until \`feedback.json\` appears.
|
|
|
|
**If \`NO_FEEDBACK_FILE\`:** The user typed their preferences directly in the
|
|
AskUserQuestion response instead of using the board. Use their text response
|
|
as the feedback.
|
|
|
|
**POLLING FALLBACK:** Only use polling if \`$D serve\` fails (no port available).
|
|
In that case, show each variant inline using the Read tool (so the user can see them),
|
|
then use AskUserQuestion:
|
|
"The comparison board server failed to start. I've shown the variants above.
|
|
Which do you prefer? Any feedback?"
|
|
|
|
**After receiving feedback (any path):** Output a clear summary confirming
|
|
what was understood:
|
|
|
|
"Here's what I understood from your feedback:
|
|
PREFERRED: Variant [X]
|
|
RATINGS: [list]
|
|
YOUR NOTES: [comments]
|
|
DIRECTION: [overall]
|
|
|
|
Is this right?"
|
|
|
|
Use AskUserQuestion to verify before proceeding.
|
|
|
|
**Save the approved choice:**
|
|
\`\`\`bash
|
|
echo '{"approved_variant":"<V>","feedback":"<FB>","date":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","screen":"<SCREEN>","branch":"'$(git branch --show-current 2>/dev/null)'"}' > "$_DESIGN_DIR/approved.json"
|
|
\`\`\``;
|
|
}
|
|
|
|
export function generateTasteProfile(ctx: TemplateContext): string {
|
|
return `Read the persistent taste profile if it exists:
|
|
|
|
\`\`\`bash
|
|
_TASTE_PROFILE=~/.gstack/projects/$SLUG/taste-profile.json
|
|
if [ -f "$_TASTE_PROFILE" ]; then
|
|
# Schema v1: { dimensions: { fonts, colors, layouts, aesthetics }, sessions: [] }
|
|
# Each dimension has approved[] and rejected[] entries with
|
|
# { value, confidence, approved_count, rejected_count, last_seen }
|
|
# Confidence decays 5% per week of inactivity — computed at read time.
|
|
cat "$_TASTE_PROFILE" 2>/dev/null | head -200
|
|
echo "TASTE_PROFILE_FOUND"
|
|
else
|
|
echo "NO_TASTE_PROFILE"
|
|
fi
|
|
\`\`\`
|
|
|
|
**If TASTE_PROFILE_FOUND:** Summarize the strongest signals (top 3 approved entries
|
|
per dimension by confidence * approved_count). Include them in the design brief:
|
|
|
|
"Based on ${'\\${SESSION_COUNT}'} prior sessions, this user's taste leans toward:
|
|
fonts [top-3], colors [top-3], layouts [top-3], aesthetics [top-3]. Bias
|
|
generation toward these unless the user explicitly requests a different direction.
|
|
Also avoid their strong rejections: [top-3 rejected per dimension]."
|
|
|
|
**If NO_TASTE_PROFILE:** Fall through to per-session approved.json files (legacy).
|
|
|
|
**Conflict handling:** If the current user request contradicts a strong persistent
|
|
signal (e.g., "make it playful" when taste profile strongly prefers minimal), flag
|
|
it: "Note: your taste profile strongly prefers minimal. You're asking for playful
|
|
this time — I'll proceed, but want me to update the taste profile, or treat this
|
|
as a one-off?"
|
|
|
|
**Decay:** Confidence scores decay 5% per week. A font approved 6 months ago with
|
|
10 approvals has less weight than one approved last week. The decay calculation
|
|
happens at read time, not write time, so the file only grows on change.
|
|
|
|
**Schema migration:** If the file has no \`version\` field or \`version: 0\`, it's
|
|
the legacy approved.json aggregate — \`${ctx.paths.binDir}/gstack-taste-update\`
|
|
will migrate it to schema v1 on the next write.`;
|
|
}
|
|
|
|
// ─── UX Behavioral Foundations (Krug + HCI research) ───
|
|
export function generateUXPrinciples(_ctx: TemplateContext): string {
|
|
return `## UX Principles: How Users Actually Behave
|
|
|
|
These principles govern how real humans interact with interfaces. They are observed
|
|
behavior, not preferences. Apply them before, during, and after every design decision.
|
|
|
|
### The Three Laws of Usability
|
|
|
|
1. **Don't make me think.** Every page should be self-evident. If a user stops
|
|
to think "What do I click?" or "What does this mean?", the design has failed.
|
|
Self-evident > self-explanatory > requires explanation.
|
|
|
|
2. **Clicks don't matter, thinking does.** Three mindless, unambiguous clicks
|
|
beat one click that requires thought. Each step should feel like an obvious
|
|
choice (animal, vegetable, or mineral), not a puzzle.
|
|
|
|
3. **Omit, then omit again.** Get rid of half the words on each page, then get
|
|
rid of half of what's left. Happy talk (self-congratulatory text) must die.
|
|
Instructions must die. If they need reading, the design has failed.
|
|
|
|
### How Users Actually Behave
|
|
|
|
- **Users scan, they don't read.** Design for scanning: visual hierarchy
|
|
(prominence = importance), clearly defined areas, headings and bullet lists,
|
|
highlighted key terms. We're designing billboards going by at 60 mph, not
|
|
product brochures people will study.
|
|
- **Users satisfice.** They pick the first reasonable option, not the best.
|
|
Make the right choice the most visible choice.
|
|
- **Users muddle through.** They don't figure out how things work. They wing
|
|
it. If they accomplish their goal by accident, they won't seek the "right" way.
|
|
Once they find something that works, no matter how badly, they stick to it.
|
|
- **Users don't read instructions.** They dive in. Guidance must be brief,
|
|
timely, and unavoidable, or it won't be seen.
|
|
|
|
### Billboard Design for Interfaces
|
|
|
|
- **Use conventions.** Logo top-left, nav top/left, search = magnifying glass.
|
|
Don't innovate on navigation to be clever. Innovate when you KNOW you have a
|
|
better idea, otherwise use conventions. Even across languages and cultures,
|
|
web conventions let people identify the logo, nav, search, and main content.
|
|
- **Visual hierarchy is everything.** Related things are visually grouped. Nested
|
|
things are visually contained. More important = more prominent. If everything
|
|
shouts, nothing is heard. Start with the assumption everything is visual noise,
|
|
guilty until proven innocent.
|
|
- **Make clickable things obviously clickable.** No relying on hover states for
|
|
discoverability, especially on mobile where hover doesn't exist. Shape, location,
|
|
and formatting (color, underlining) must signal clickability without interaction.
|
|
- **Eliminate noise.** Three sources: too many things shouting for attention
|
|
(shouting), things not organized logically (disorganization), and too much stuff
|
|
(clutter). Fix noise by removal, not addition.
|
|
- **Clarity trumps consistency.** If making something significantly clearer
|
|
requires making it slightly inconsistent, choose clarity every time.
|
|
|
|
### Navigation as Wayfinding
|
|
|
|
Users on the web have no sense of scale, direction, or location. Navigation
|
|
must always answer: What site is this? What page am I on? What are the major
|
|
sections? What are my options at this level? Where am I? How can I search?
|
|
|
|
Persistent navigation on every page. Breadcrumbs for deep hierarchies.
|
|
Current section visually indicated. The "trunk test": cover everything except
|
|
the navigation. You should still know what site this is, what page you're on,
|
|
and what the major sections are. If not, the navigation has failed.
|
|
|
|
### The Goodwill Reservoir
|
|
|
|
Users start with a reservoir of goodwill. Every friction point depletes it.
|
|
|
|
**Deplete faster:** Hiding info users want (pricing, contact, shipping). Punishing
|
|
users for not doing things your way (formatting requirements on phone numbers).
|
|
Asking for unnecessary information. Putting sizzle in their way (splash screens,
|
|
forced tours, interstitials). Unprofessional or sloppy appearance.
|
|
|
|
**Replenish:** Know what users want to do and make it obvious. Tell them what they
|
|
want to know upfront. Save them steps wherever possible. Make it easy to recover
|
|
from errors. When in doubt, apologize.
|
|
|
|
### Mobile: Same Rules, Higher Stakes
|
|
|
|
All the above applies on mobile, just more so. Real estate is scarce, but never
|
|
sacrifice usability for space savings. Affordances must be VISIBLE: no cursor
|
|
means no hover-to-discover. Touch targets must be big enough (44px minimum).
|
|
Flat design can strip away useful visual information that signals interactivity.
|
|
Prioritize ruthlessly: things needed in a hurry go close at hand, everything
|
|
else a few taps away with an obvious path to get there.`;
|
|
}
|
|
|