diff --git a/AGENTS.md b/AGENTS.md index cef8c05f8..7d802490e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -19,7 +19,7 @@ Invoke them by name (e.g., `/office-hours`). | `/plan-design-review` | Rate each design dimension 0-10, explain what a 10 looks like. | | `/plan-devex-review` | DX-mode review: TTHW, magical moments, friction points, persona traces. | | `/plan-tune` | Self-tune AskUserQuestion sensitivity per question. | -| `/autoplan` | One command runs CEO → design → eng → DX review. | +| `/autoplan` | One command runs CEO → design → DX → eng review (eng always last). | | `/design-consultation` | Build a complete design system from scratch. | | `/spec` | Turn vague intent into a precise, executable spec in five phases. Files a GitHub issue, optionally spawns a Claude Code agent in a fresh worktree, and lets `/ship` close the source issue on merge. | diff --git a/CHANGELOG.md b/CHANGELOG.md index 316624c15..6c5b26322 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,67 @@ # Changelog +## [1.75.0.0] - 2026-08-29 + +**Your review now hunts over-built code, not just broken code.** +**And every skill's advice starts with "reuse before you build."** + +This release imports the best of ponytail, the code-minimalism ruleset, without importing its build-less posture. /review gains an eighth lens: a simplification specialist that flags unrequested structure (hand-rolled stdlib, one-implementation abstractions, dead flexibility, dependencies duplicating platform features) in a closed five-tag vocabulary. It is advisory only. It cannot dent your quality score, its fixes are never auto-applied, and on a lean diff it tells you something no reviewer ever says: "lean already — nothing to cut." Every tier-2+ skill also gains the reuse ladder (stop at the first rung that holds: this repo, stdlib, native platform, an installed dependency... then build the complete version of what remains) and a bounded closer, so completion reports stop touring every edit. Accepted shortcuts now leave a durable trail: a decision-ledger entry plus a `gstack-shortcut(dec-)` marker in code, harvested into a debt ledger by /retro. Agent hosts without a full install (Zed, Amp, Cursor side projects) get a 1.8KB rules digest to copy into their own rules file. And /autoplan now runs the Eng review last, always, so the required shipping gate reviews the final amended plan instead of a stale one. + +### The numbers that matter + +Source: this branch's own runs, recorded under `~/.gstack/projects//evals/` and the commit receipts. + +| Metric | Before | After | Δ | +|---|---|---|---| +| /review lenses | 7 | 8 (simplification, advisory) | +1 | +| Simplification lens on this branch's own 5,190-line diff | n/a | `net: -37 lines possible` | dogfooded | +| AskUserQuestion preamble section, per skill | baseline | -236 B (~9.7 KB across 41 skills) | gated by a live A/B: post-cut 7/7 format elements, substance equal to pre-cut | +| Instruction-tier artifact for rules-reading hosts | none | 1,765 B committed digest (2,048 B budget) | new | +| Skills-earn-their-tokens benchmark | none | 3 tasks x 2 arms (with/without skill), judged on the diff left behind | new | +| /autoplan phase order | CEO, Design, Eng, DX | CEO, Design, DX, Eng always last | the gate sees the final plan | + +The A/B receipt is the one to trust: the AskUserQuestion cut only landed because a two-arm eval on real SDK captures showed the shorter render lost nothing (the gate outranked the approval, by design). + +### What this means for your workflow + +Run /review on a branch you suspect is over-built and read the `[ADVISORY]` rows plus the `net: -N lines possible` footer; nothing blocks, nothing auto-applies, you decide. Take a shortcut in an AskUserQuestion and it stops rotting silently: /retro reads the ledger back to you with its upgrade trigger. If you work in a rules-reading editor without a gstack install, copy `agents-digest/gstack-AGENTS.md` into your rules file and the ethos rides along. + +### Itemized changes + +### Added + +- **Simplification review specialist** (`review/specialists/simplification.md`): closed tag vocabulary (`delete:` / `stdlib:` / `native:` / `speculative:` / `shrink:`), dispatched on diffs over 100 lines, `--simplification` force flag. Advisory carve-out end to end: excluded from the quality score and findings-count header, ASK-only in Fix-First, `[ADVISORY]` labels, parent-printed `net: -N lines possible` footer, and a `Simplification: lean already — nothing to cut.` zero-findings line. Precision guarded by a false-flag fixture (a complete-but-lean diff must yield NO FINDINGS) and coverage-vs-structure boundary text (tests, error paths, and edge cases are never deletion targets). +- **Reuse ladder in Search Before Building** (every tier-2+ skill): before writing new code, stop at the first rung that holds — a helper already in the repo, the stdlib, a native platform feature, an already-installed dependency — then build the complete version of what remains. Root-cause rule included: one guard in the shared function beats a guard in every caller. +- **Bounded closer** (every tier-2+ skill): after completing work, report what changed, what was skipped, what to watch — a few short lines, with report-shaped skills (/qa-only, /plan-*-review, /retro, /document-generate) explicitly exempt because their report IS the work. +- **Shortcut debt ledger**: accepting a Completeness ≤ 7 option on a durable-scope call now logs the ceiling and upgrade trigger to the decision ledger and marks each cut corner with `gstack-shortcut(dec-): , upgrade when `; /retro Step 11.5 harvests the markers, joins them on decision ids, and tags `unlinked` and `no-trigger` rot risks. Markers survive the redaction engine (pinned by test). +- **Instruction-only host tier**: `agents-digest/gstack-AGENTS.md`, a committed, generated, budget-capped (2,048 B) digest of the ethos, reuse ladder, and voice rules. Setup's openclaw/hermes explainer arms print its path for copy-in; setup never writes a user's AGENTS.md (tripwire-tested, including laundered write shapes). README's host table now matches what setup actually does. +- **With/without-skill arm benchmark** (periodic eval): three build-shaped tasks (a native-platform overbuild trap, a CRUD endpoint, a bugfix with planted decoys) run through real `claude -p` sessions twice — with the behavioral-layer skill installed and without — and the staged diff each arm leaves is judged for over-engineering on a 0-3 rubric with a full failure taxonomy (judge_error cells excluded from aggregates but surfaced; zero-diff arms are valid cells). Each cell also runs the fixture's own functional oracle (`checks=pass|fail|none`), so a refusal, a broken implementation, and working code stay distinguishable — correctness before LOC. A research instrument, not a gate. +- **/autoplan runs Eng last, always**: mandatory order is now CEO → Design (if UI scope) → DX (if developer-facing scope) → Eng, with a single final approval gate; clearly-wrong premises queue as User-Challenge items instead of stopping mid-run, and accepting one at the gate amends the plan and re-runs Eng on the amended plan. Static test pins the Eng-terminal order. + +### Changed + +- **AskUserQuestion preamble section slimmed** (-236 B per skill, ~9.7 KB across the corpus): duplicate statements of the completeness rule, auto-decide marker, and tool-not-prose rule removed while keeping every verbosity floor and all 14 format pins. Landed only after a live NOT-WORSE A/B (pre-cut vs post-cut render, identical prompt, SDK capture) showed zero format-element loss and equal recommendation substance. +- **Terse-mode label tells the truth**: the claimed savings is now the measured 2.6 KB, not "~3-5 KB". +- `/review` checklist knows Completeness Gaps and Simplification are orthogonal (coverage up, structure down), and a `gstack-shortcut` marker downgrades a would-be gap finding to acknowledged debt — but only when its decision id resolves in the ledger; an orphan marker is reported as a forged suppression (cross-model adversarial catch). + +### Fixed + +- **Version bumps no longer strand the agents digest**: `gstack-version-bump write --regen-digest` reruns the repo's digest generator in the same mutation (explicit opt-in — a plain `write` never executes repo files it merely finds), and both ship's and land-and-deploy's evidence gates allow-list the digest alongside VERSION. Without this, every release commit of this repo would have failed the freshness CI check. +- **The free-suite flaky retry can no longer mask real failures**: the retry pass vetoes on ANY unattributable failure evidence (headerless failures, unhandled errors between tests, truncated runs), an empty shard carries an empty failing-files list instead of crashing the retry, and a dead conditional was removed. +- **Version allocation distrusts laundered git**: `gstack-next-version` consults `ls-remote` only when origin is actually configured, and a configured origin that "successfully" advertises zero heads (the Conductor git shim's failure shape) now falls back to local refs with a loud warning instead of reading an empty queue and reallocating a taken version. Both shim shapes are pinned by regression tests. +- **Browse temp paths are portable**: local file serving accepts both the browse TEMP_DIR and the OS tmpdir (`TEMP_DIRS` allowlist), while remote serving stays pinned to TEMP_DIR only (the security asymmetry is test-enforced). An untrustable TMPDIR (`/`, `$HOME`, an ancestor of the daemon's cwd) is ignored rather than trusted for the daemon's lifetime. +- Setup's instruction-tier pointer prints a script-anchored digest path instead of a cwd-relative one that broke when invoked from another directory. +- /retro's shortcut harvest no longer reports phantom debt from files that merely document the marker convention (placeholder filter + judgment prose). +- The arm benchmark harvests against a recorded seed commit (immune to agents that commit and push), ignores node_modules, survives multi-megabyte patches, names its judge-diff cap and reports truncation loudly, and wraps untrusted diffs in per-call random sentinels so a faked closing marker cannot steer the judge. +- The AUQ A/B eval treats a judge failure on either side as inconclusive instead of coercing it to a fake degradation-or-mask, and reads its pre-cut arm from a vendored fixture instead of a branch-local SHA that dies with the branch. + +### For contributors + +- `bun run test` gains sandbox knobs: `GSTACK_FREE_JOBS` overrides shard count (digits-only, loud on garbage) and `GSTACK_FREE_RETRY_FLAKY=1` opts into one serial retry pass for syscall-supervised sandboxes (default stays OFF — dev boxes should see flakes). `scripts/sandbox-doctor.sh` makes a Vercel/Conductor cloud sandbox run the suite green in one idempotent command (documented in docs/TESTING_INTERNALS.md): atomic git-shim patching with a backup, loud on patch-pattern drift, :99-socket Xvfb detection, dnf-gated installs, and it survives a missing /dev/shm. A failed digest regen now fails `bun run gen:skill-docs` locally instead of deferring the red to CI. +- The arm-benchmark selftest (fixture integrity, judge plumbing, install asymmetry) now runs FREE in `bun run test` on every PR via `test/arm-benchmark-selftest.test.ts` — the paid periodic instrument can no longer burn money on broken fixtures. +- Eval-store schema v2: harvest records carry `{insertions, deletions, net}`; `recordE2E` now populates `tokens_used`. +- Touchfiles dep lists closed gaps (fixtures, judge helper, harness, ship render) and the context-budget ratchet fixture was re-captured, locking the AskUserQuestion reduction so it cannot silently regress. +- New coverage: TEMP_DIRS widening + remote-serving asymmetry, shortcut-marker writer/harvester grammar joint, sandbox-doctor shell guards, GSTACK_FREE_JOBS parsing, laundered-ls-remote shims, digest freshness/budget/writer tripwires, and a real generator round-trip through the version bump. ## [1.74.0.0] - 2026-08-29 **Green now means green: every test runs somewhere, provably.** diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a301a1ff8..f81afc0b5 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -166,6 +166,12 @@ concurrent shard processes under a strict output contract — a shard that exits without bun's own terminal summary line, or a crashed worker, fails the run, so silent truncation can never report green. Pass `--verbose` to forward the full child stream; `--wall-timeout ` overrides the per-shard kill deadline. +`GSTACK_FREE_JOBS=` overrides the shard count (digits only, loud on garbage), +and `GSTACK_FREE_RETRY_FLAKY=1` opts into one serial retry pass for +syscall-supervised sandboxes (off by default — dev boxes should see flakes). +Working in a cloud sandbox? Run `scripts/sandbox-doctor.sh` once per boot to +make the suite run green (details in +[docs/TESTING_INTERNALS.md](docs/TESTING_INTERNALS.md)). Don't type bare `bun test` for the suite: it walks the whole repo, loads paid eval files, and misses the strict classifier. No API keys needed. @@ -353,11 +359,11 @@ worth the review overhead. ## Multi-host development -gstack generates SKILL.md files for 8 hosts from one set of `.tmpl` templates. +gstack generates SKILL.md files for 10 hosts from one set of `.tmpl` templates. Each host is a typed config in `hosts/*.ts`. The generator reads these configs to produce host-appropriate output (different frontmatter, paths, tool names). -**Supported hosts:** Claude (primary), Codex, Factory, Kiro, OpenCode, Slate, Cursor, OpenClaw. +**Supported hosts:** Claude (primary), Codex, Factory, Kiro, OpenCode, Slate, Cursor, OpenClaw, Hermes, GBrain. ### Generating for all hosts @@ -366,7 +372,7 @@ to produce host-appropriate output (different frontmatter, paths, tool names). bun run gen:skill-docs # Claude (default) bun run gen:skill-docs --host codex # Codex bun run gen:skill-docs --host opencode # OpenCode -bun run gen:skill-docs --host all # All 8 hosts +bun run gen:skill-docs --host all # All 10 hosts # Or use build, which does all hosts + compiles binaries bun run build diff --git a/README.md b/README.md index 7b2b9305e..cb76f65a8 100644 --- a/README.md +++ b/README.md @@ -111,16 +111,23 @@ cd ~/gstack && ./setup Or target a specific agent with `./setup --host `: -| Agent | Flag | Skills install to | -|-------|------|-------------------| -| OpenAI Codex CLI | `--host codex` | `${CODEX_HOME:-~/.codex}/skills/gstack-*/` | -| OpenCode | `--host opencode` | `~/.config/opencode/skills/gstack-*/` | -| Cursor | `--host cursor` | `~/.cursor/skills/gstack-*/` | -| Factory Droid | `--host factory` | `~/.factory/skills/gstack-*/` | -| Slate | `--host slate` | `~/.slate/skills/gstack-*/` | -| Kiro | `--host kiro` | `~/.kiro/skills/gstack-*/` | -| Hermes | `--host hermes` | `~/.hermes/skills/gstack-*/` | -| GBrain (mod) | `--host gbrain` | `~/.gbrain/skills/gstack-*/` | +| Agent | Flag | What you get | +|-------|------|--------------| +| OpenAI Codex CLI | `--host codex` | Full install → `${CODEX_HOME:-~/.codex}/skills/gstack-*/` | +| OpenCode | `--host opencode` | Full install → `~/.config/opencode/skills/gstack-*/` | +| Cursor | `--host cursor` | Full install → `~/.cursor/skills/gstack-*/` | +| Factory Droid | `--host factory` | Full install → `~/.factory/skills/gstack-*/` | +| Kiro | `--host kiro` | Full install → `~/.kiro/skills/gstack-*/` | +| Slate | `--host slate` | Pointer to the Claude install (Slate reads `.claude/skills` as a fallback) | +| OpenClaw | `--host openclaw` | ACP spawn pointers + methodology artifacts via `gen:skill-docs --host openclaw` + the instruction-only digest below (full guide: [docs/OPENCLAW.md](docs/OPENCLAW.md)) | +| Hermes | `--host hermes` | Methodology artifacts via `gen:skill-docs --host hermes` + the instruction-only digest below | +| GBrain (mod) | `--host gbrain` | Brain-aware skill variants, shipped from the GBrain repo | + +**Instruction-only tier (any rules-reading agent — Zed, Amp, Jules, side projects):** +copy the 2KB digest at [`agents-digest/gstack-AGENTS.md`](agents-digest/gstack-AGENTS.md) +into a location your agent reads (for example, append it to your project's `AGENTS.md`). +It carries gstack's ethos, reuse ladder, and voice rules — no install required. The +digest's first line shows its gstack version; re-copy it after upgrading. For Codex, setup reads the top-level `model` from `${CODEX_HOME:-~/.codex}/config.toml` and generates the matching behavioral @@ -195,7 +202,7 @@ Each skill feeds into the next. `/office-hours` writes a design doc that `/plan- | `/plan-design-review` | **Senior Designer** | Rates each design dimension 0-10, explains what a 10 looks like, then edits the plan to get there. AI Slop detection. Interactive — one AskUserQuestion per design choice. | | `/plan-devex-review` | **Developer Experience Lead** | Interactive DX review: explores developer personas, benchmarks against competitors' TTHW, designs your magical moment, traces friction points step by step. Three modes: DX EXPANSION, DX POLISH, DX TRIAGE. 20-45 forcing questions. | | `/design-consultation` | **Design Partner** | Build a complete design system from scratch. Researches the landscape, proposes creative risks, generates realistic product mockups. | -| `/review` | **Staff Engineer** | Find the bugs that pass CI but blow up in production. Auto-fixes the obvious ones. Flags completeness gaps. | +| `/review` | **Staff Engineer** | Find the bugs that pass CI but blow up in production. Auto-fixes the obvious ones. Flags completeness gaps. Advisory simplification lens flags over-built code — never blocks, never auto-applies. | | `/investigate` | **Debugger** | Systematic root-cause debugging. Iron Law: no fixes without investigation. Traces data flow, tests hypotheses, stops after 3 failed fixes. | | `/design-review` | **Designer Who Codes** | Same audit as /plan-design-review, then fixes what it finds. Atomic commits, before/after screenshots. | | `/devex-review` | **DX Tester** | Live developer experience audit. Actually tests your onboarding: navigates docs, tries the getting started flow, times TTHW, screenshots errors. Compares against `/plan-devex-review` scores — the boomerang that shows if your plan matched reality. | @@ -214,7 +221,7 @@ Each skill feeds into the next. `/office-hours` writes a design doc that `/plan- | `/retro` | **Eng Manager** | Team-aware weekly retro. Per-person breakdowns, shipping streaks, test health trends, growth opportunities. `/retro global` runs across all your projects and AI tools (Claude Code, Codex, Gemini). | | `/browse` | **QA Engineer** | Give the agent eyes. Real Chromium browser, real clicks, real screenshots. ~100ms per command. `/open-gstack-browser` launches GStack Browser with sidebar, anti-bot stealth, and auto model routing. | | `/setup-browser-cookies` | **Session Manager** | Import cookies from your real browser (Chrome, Arc, Brave, Edge) into the headless session. Test authenticated pages. | -| `/autoplan` | **Review Pipeline** | One command, fully reviewed plan. Runs CEO → design → eng review automatically with encoded decision principles. Surfaces only taste decisions for your approval. | +| `/autoplan` | **Review Pipeline** | One command, fully reviewed plan. Runs CEO → design → DX → eng review automatically (eng always last, so the shipping gate reviews the final amended plan) with encoded decision principles. Surfaces only taste decisions for your approval. | | `/spec` | **Spec Author** | Turn vague intent into a precise, executable spec in five phases (why, scope, technical with mandatory code-reading, draft, file). Codex quality gate before file (blocks below 7/10), fail-closed secret redaction, dedupe against existing issues, archive to `$GSTACK_STATE_ROOT/projects/$SLUG/specs/` for team-corpus recall. `--execute` spawns `claude -p` in a fresh worktree; `/ship` auto-closes the source issue on merge. Plan-mode aware. | | `/learn` | **Memory** | Manage what gstack learned across sessions. Review, search, prune, and export project-specific patterns, pitfalls, and preferences. Learnings compound across sessions so gstack gets smarter on your codebase over time. | | `/make-pdf` | **Publisher** | Markdown in, publication-quality document out. Mermaid and excalidraw fences render as vector diagrams, fully offline. Images scale to the page and never truncate; wide diagrams get their own landscape page. `--to html` emits one self-contained file, `--to docx` a Word doc. | @@ -227,7 +234,7 @@ Each skill feeds into the next. `/office-hours` writes a design doc that `/plan- | **End users** (UI, web app, mobile) | `/plan-design-review` | `/design-review` | | **Developers** (API, CLI, SDK, docs) | `/plan-devex-review` | `/devex-review` | | **Architecture** (data flow, perf, tests) | `/plan-eng-review` | `/review` | -| **All of the above** | `/autoplan` (runs CEO → design → eng → DX, auto-detects which apply) | — | +| **All of the above** | `/autoplan` (runs CEO → design → DX → eng, auto-detects which apply; eng always last) | — | ### Power tools diff --git a/TODOS.md b/TODOS.md index 71e8ab974..7be847dde 100644 --- a/TODOS.md +++ b/TODOS.md @@ -491,6 +491,71 @@ audit trail lives in Aside. ## Test infrastructure +### P1: skillify gate test red — HOME-override sessions never discover project skills (pre-existing) + +**What:** `test/skill-e2e-skillify.test.ts` `skillify-provenance-refusal` fails +on BOTH this branch and origin/main @ b5a951e6 (proven 2026-08-29: identical +2-turn `Unknown skill: skillify` transcripts). Every test in that file passing +`env: { HOME: workDir }` gets ZERO seeded project skills in the session init +(claude CLI 2.1.237); the passing siblings recover by Reading the SKILL.md +directly, the refusal test's agent stops at the Skill error. Fix the harness +(seed skills wherever HOME-overridden discovery looks, or drop the HOME +override and pass the write target another way), or report upstream if +project-scope `.claude/skills` discovery genuinely keys off HOME. + +**Why:** A gate-tier safety test that is red for environmental reasons trains +people to ignore gate reds. + +**Effort:** S-M (harness). **Priority:** P1 (gate hygiene). + +### P2: auq-verbose-vs-carved-ab PRE arm reads a branch-local ref (same fragility class the repetition-cut A/B just fixed) + +**What:** `test/helpers/auq-sdk-capture.ts` `verboseSkill()` defaults to git ref +`ab66193e^`, reachable only from the token-usage-reduction branch — shallow +clones fail today, all clones fail after that branch is pruned. Vendor the +pre-carve render as a fixture the way `auq-pre-cut-plan-ceo-review-SKILL.md` +was vendored for the repetition-cut A/B (v1.75.0.0), or repoint at a +main-reachable commit. + +**Effort:** S. **Priority:** P2 (weekly periodic breaks silently later). + +### P3: eval-store harvest as a discriminated union + +**What:** `EvalTestEntry.harvest` went all-optional in schema v2 (worktree +harvests carry patchPath/isDuplicate, arm-benchmark diff-stats carry +insertions/deletions/net) — compile-time safety for the two writer shapes now +rests on a comment. Model as `{kind:'worktree',...} | {kind:'diff-stat',...}`. +Filed from the v1.73 review army (maintainability); deferred at ship time to +avoid schema churn mid-release. + +**Effort:** S. **Priority:** P3. + +### P2: WS6-2 dead-frontmatter strip — needs a live host, not a sandbox + +**What:** `bin/gstack-context-bill` warns about 14 frontmatter keys "the router +never reads" (ROUTER_KEYS in lib/context-bill.ts is a hand-maintained guess). +The approved ponytail-import plan mandates EMPIRICAL verification before +stripping: remove the keys in a scratch install on a LIVE Claude Code host, +confirm skill discovery/routing/hooks unchanged, then land via the +hosts/claude.ts denylist (keys stay in templates for gen tooling). Deferred at +v1.73 implementation time with a decision-ledger entry (2026-08-28) because the +cloud sandbox cannot exercise live-host discovery. Savings are hundreds of +always-on bytes; growth is already capped by the ratchet regardless. + +**Effort:** S (once on a live host). **Priority:** P2. + +### P3: scope the evidence-gate digest allow-path + +**What:** `agents-digest/gstack-AGENTS.md` rides `--allow-paths` in ship's and +land-and-deploy's evidence checks in EVERY repo, and unlike CHANGELOG/VERSION +it is instruction-bearing for rules-reading hosts. Scope the exemption to +"the bump actually regenerated it" (e.g. gstack-evidence learns a +--allow-if-regenerated flag, or the check compares the digest bytes to a fresh +generator run). Filed from the v1.73 Claude adversarial pass; the gate is +advisory and gstack's freshness CI covers the drift case, so P3. + +**Effort:** S-M. **Priority:** P3. + ### 2026-08-29 test-infra overhaul — follow-ups (filed at implementation) The overhaul landed: green-means-green fixes (make-pdf gates in the required diff --git a/VERSION b/VERSION index e8df5e2eb..2fbbd2a2f 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -1.74.0.0 +1.75.0.0 diff --git a/agents-digest/gstack-AGENTS.md b/agents-digest/gstack-AGENTS.md new file mode 100644 index 000000000..3054cd1ee --- /dev/null +++ b/agents-digest/gstack-AGENTS.md @@ -0,0 +1,35 @@ +# gstack digest v1.75.0.0 — regenerate/re-copy after upgrading gstack + +Behavioral rules from gstack (https://github.com/garrytan/gstack), compressed +for agent hosts without a full skill install. The full skills add workflows, +reviews, and evals on top of these rules. + +## Ethos + +- **Boil the Ocean** — AI makes completeness cheap, so do the complete thing: tests, edge cases, error paths. Shortcuts need an explicit, recorded decision. +- **Search Before Building** — know what exists before deciding what to build. Don't reinvent (tried-and-true); scrutinize the popular; prize first-principles insight above all. +- **User Sovereignty** — models recommend, the user decides. Cross-model agreement is signal, never permission. Ask before changing the user's stated direction. +- **Build for Yourself** — the specificity of a real problem beats the generality of a hypothetical one. + +## The reuse ladder + +Before writing new code, stop at the first rung that holds: +1. A helper, util, or pattern already in this repo. +2. The standard library. +3. A native platform feature (CSS over JS, DB constraint over app code). +4. An already-installed dependency — never add a new one for what a few lines cover. + +Then build the complete version of what remains. Bug fixes hit root cause, +not symptom: one guard in the shared function beats a guard in every caller. + +## Voice + +Direct, concrete, builder-to-builder. Name the file, function, command, and +user-visible impact. Short paragraphs; end with what to do. No filler, no +corporate tone, no AI vocabulary. + +## Full gstack + +Clone https://github.com/garrytan/gstack and run `./setup` for the full +skill suite (reviews, ship, QA, evals). This digest is generated — edit +scripts/gen-agents-digest.ts, not this file. diff --git a/autoplan/SKILL.md b/autoplan/SKILL.md index 0cc6c8dca..6e90fd372 100644 --- a/autoplan/SKILL.md +++ b/autoplan/SKILL.md @@ -79,7 +79,7 @@ If `SKILL_PREFIX` is `"true"`, suggest/invoke `/gstack-*` names. Disk paths stay Branch on the skill-start STATUS lines, in this order: -1. **`CONDUCTOR_SESSION: true` echoed** → do NOT call AskUserQuestion at all (neither native nor any `mcp__*__AskUserQuestion` variant): render EVERY decision brief as the **prose form** below and STOP. Proactive, not a failure reaction — Conductor disables native AUQ and its MCP variant is flaky (`[Tool result missing due to internal error]`). **Auto-decide preferences still apply first:** a surfaced `[plan-tune auto-decide]