mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 18:06:54 +02:00
* refactor(resolvers): split review.ts into MECE resolver modules (pure move) Move every function from scripts/resolvers/review.ts, unchanged, into: - review-dashboard.ts: review dashboard, plan-file review report - plan-gates.ts: approval check, exit-plan-mode gate, plan-file discovery, plan-completion audit/gate (ship + review), plan verification exec - spec-review.ts: both spec review loops, benefits-from, anti-shortcut clause - outside-voice-steps.ts: Codex second opinion, adversarial step, Codex plan review, Codex doc review, disabled-outside record - review-scope.ts: scope drift, cross-review dedup, shared-code reuse review.ts is deleted; index.ts imports the new modules. gen-skill-docs output is byte-identical for every host (--host all). Test imports and source-path references are re-pointed; the two source-text report/gate tests in gen-skill-docs.test.ts become behavioral renders across every consuming skill and host. All 46 touchfile entries that named review.ts now name all five modules, guarded by a recorded selection golden. * test(browse): black-box auth matrix for every server route and both surfaces Drives buildFetchHandler fetchLocal/fetchTunnel with no token, wrong token, root token, scoped token and the SSE cookie for all 33 routes, plus unmatched paths and wrong methods. Denials assert today's exact status, body and content type; allowed credentials assert the handler was reached. Written against the unchanged if-chain server so the W3 route-table refactor must keep it green. * refactor(shard-engine): move scripts/test-strict-output.ts to scripts/lib/shard-engine.ts The shared shard engine grows from the existing strict-output module (runShardChild, killProcessGroup, signal forwarding, strict classifier). scripts/test-strict-output.ts stays as a re-export so existing importers, mock.module paths and the strict-output/run-shard-child tests are unchanged. The engine inherits the global touchfile entry; the free runner's CLI-routing fixture copies the new module. * refactor(resolvers): decompose the three >150-line review resolvers (output-neutral) Split generateAdversarialStep, generateCodexPlanReview and generatePlanCompletionAuditInner into per-section helpers whose template literals are copied verbatim, so every function in the new modules is at or under 150 lines. gen-skill-docs output is byte-identical for every host (--host all, compared against96764e80with a fixed --link-root). * refactor(resolvers): one outside-voice failure policy (deliberate prose unification) outsideVoiceFailurePolicy(ctx, opts) in outside-voice.ts now renders the auth / timeout / empty-response bullets for all four call sites that hand-typed them (Codex second opinion, adversarial step, Codex plan review, design outside voices). Options are explicit per site (timeoutMinutes, onTimeout, stderrOnEmpty, fallback, escape) with no defaults. Deliberate generated-prose changes (every host): - office-hours: 'Fall back to <native> subagent.' becomes 'Fall back to the <native> subagent below.' - plan-devex-review: the plain 'Auth failure (stderr contains ...)' bullets become the canonical bold bullets; auth also triggers on 'API key'; 'auth failed' becomes 'authentication failed'. - review/ship adversarial: 'exceeded 9 minutes and was terminated' becomes 'timed out after 9 minutes and was terminated'; the timeout is still MISSING COVERAGE. - design outside voices: unchanged. Adds ratchet (d) (test/outside-voice-failure-policy.test.ts) with a reasoned allowlist for /codex's own CLI errors, the MISSING COVERAGE retention test, refreshed codex/factory ship goldens, and outside-voice.ts in every touchfile entry of review.ts and design.ts (selection golden extended). * test(pty): fake PTY session driver with an injectable clock through the runner launch seam The three plan-skill runners take an optional PtyDriver (launch, now, monotonic, sleep); omitted, they use the real launcher and clocks exactly as before. test/helpers/pty/fake-session.ts feeds scripted frames through that seam, and claude-pty-runner.runners.unit.test.ts runs observation, counting and floor for success, deadline timeout, permission prompt and plan-ready outcomes with no CLI or real timers. These cases must stay green unchanged through the W4 split and the runPtySession extraction. Touchfiles: every entry that lists claude-pty-runner.ts or pty-screen.ts now also lists test/helpers/pty/**. * refactor(shard-engine): run both lanes on the shared engine; lane policy injected Engine (scripts/lib/shard-engine.ts) gains the W2 primitives: per-shard tmp/Chromium sandbox + async cleanup backstop, log-path allocation and full-stream log capture, one duration-seed reader/writer with a lane predicate, LanePolicy (seed predicate + zero-execution verdict), strictShardStatus, and the shared CLI flag loop. runShardChild takes an optional companion (signal/settle) and waits a bounded 250ms to reap a wall-killed child. Free lane stops spawning shards itself: runFreeShard uses runShardChild with trackShardBrowser as the companion (win32 path unchanged: no process group, no negative-pid kill). Its sync state-dir removal stays lane policy. Paid lane uses the sandbox, log, seed, verdict and flag primitives; the hollow-shard guard applies PAID_LANE_POLICY. Lane outcomes are unchanged (free keeps >= 0 seeds and file-count zero-exec rule; paid keeps > 0 seeds, warning under selection and passed-empty under EVALS_ALL). paid-free-boundary's closure assertion now names the engine module, where the strict classifier lives. * test(shard-engine): engine unit tests, fixture-corpus equivalence, per-lane CLI parity - test/shard-engine.test.ts: failing/unhandled/module-load output fails both lanes, per-lane zero-execution and seed rules, whole-group kill on a wall timeout (both lanes), mocked-win32 path with no negative-pid kill, companion settle order, log capture, sandbox isolation, flag loop. - test/shard-engine-equivalence.test.ts + test/fixtures/shard-equivalence: seven outcome fixtures plus one real shard, run through both lanes and compared with classifications recorded from the base runners (96764e80). - test/shard-cli-parity.test.ts + test/fixtures/shard-cli-parity: flag set, defaults, validation errors and the Unknown argument error per lane match the base runners. * refactor(shard-engine): decompose runFreeShard and runPaidShard to <= 150 lines Output-neutral extraction under the fixture-corpus equivalence and runner tests: captureFreeStream, explainFreeVerdict and logFreeRecovery (free); paidShardCommand, settleShardSpool, settleBootstrapRetention and printLogTail (paid). The bootstrap scope-creation block that bootstrap-retention.test.ts evaluates stays verbatim. * refactor(pty): split claude-pty-runner.ts into test/helpers/pty/* behind a barrel Pure move: every line of the former 5,047-line runner lands verbatim in one module (four private helpers gain `export` for cross-module use): binary, screen (absorbs test/helpers/pty-screen.ts, which now re-exports it), launch, session (PtyDriver), judge, classify, auq, plan-native, boundaries, runners/{observation,counting,floor}. claude-pty-runner.ts re-exports the original public surface by name; pty/ modules import siblings directly. Tests that read the runner's source text: - rewritten as behavioral: the unit test's model-pin tripwire (fake CLI argv: fallback chain, --model before extraArgs, hermetic --strict-mcp-config), pty-skill-seeding-wiring (runners through the fake driver; launcher through a fake CLI reporting CLAUDE_CONFIG_DIR). The "three wrappers forward model" grep is replaced by the runners' fake-driver launch assertions. - pty-screen-session / pty-screen-supervision: stop copying runner source; they mock.module the real pty/screen.ts (and the fixture cleanup) instead. - re-pointed to the owning module (they execute a sliced runner body with injected boundaries; no seam exists for those boundaries yet): eng-seeded-completion-ai, plan-floor-permission, plan-create-prepublication, plan-count-completion; hermetic-wiring's source guard now reads pty/launch.ts and scans every pty/ module for raw process.env spreads. - plan-count-timeout and pty-output-wake mock the viewport at pty/screen.ts. * test(ratchet-c): enforcing module/function size ratchet and moved-code touchfile coverage Ratchet (c) ships enforcing: test/helpers/module-size.ts counts file and top-level function lengths by brace matching over masked source (strings, comments, regex literals and template text masked; ${} expressions kept), covering function declarations, arrow functions assigned to consts and route-table handler properties, with no parser dependency. Its self-test uses template literals and code-fence braces copied from scripts/resolvers/review.ts and design.ts. test/fixtures/module-size-ratchet.json binds scripts/lib/shard-engine.ts (<= 800 lines, <= 150 per function) and records the residual runner sizes (free 2352, paid 1921) as non-growth caps; allowlist entries are keyed on file plus matched text and need a reason. Failure output lists file:line, the rule, Fix: and the allowlist path. touchfiles.test.ts gains the moved-code superset check over test/fixtures/touchfile-move-goldens/ (W2 golden recorded at96764e80: test-strict-output.ts and test-paid-shards.ts global, test-free-shards.ts none). * refactor(browse): declared route table replaces the buildFetchHandler if-chain The ~1,300-line if-chain in buildFetchHandler becomes a route table: each entry declares method, path, auth kind and surfaces, and one auth gate in browse/src/routes/table.ts returns the per-kind denial (root-bearer, scoped, root-or-sse-cookie: 401 Unauthorized; root-token: 403 Root token required; extension-origin: 403 Forbidden). Unmatched requests take the declared fallthrough (root-bearer check, then plain-text 404). Handlers move to browse/src/routes/{core,pairing,pty,tokens,tunnel,activity,commands,files, inspector}.ts and receive a RouteContext with auth checks as functions instead of closing over factory locals. Dispatch order is unchanged: tunnel filter, beforeRoute overlay, gate, handler. TUNNEL_PATHS stays a literal in server.ts. Behavior-preserving: the black-box auth matrix from the previous commit passes unchanged. /memory and /inspector/events are declared root-bearer because the blanket check always ran before their SSE-cookie branch. Source-text route tests are rewritten as behavioral tests through buildFetchHandler or a route's real handler with a stub RouteContext (browse/test/route-test-harness.ts). Checks with no runtime seam are re-pointed to the route modules: Surface type, /inspector/events SSE helper, sanitizeReplacer imports, /pty-inject-scan sidecar-client import, and the ngrok config lookup and startTunnel wiring that stay in server.ts. * test(browse): stubbed-handler auth matrix and route inventory for the route table Every ROUTES entry runs through the real dispatcher and gate with stub handlers on each declared surface and six credentials; denials assert the exact status and body each auth kind returned at96764e8, admitted credentials assert the handler ran (with the gate's TokenInfo for scoped routes). Also pins the reviewed route inventory (method, path, auth kind, surfaces), that every entry declares auth and surfaces, that the table's tunnel paths equal the TUNNEL_PATHS literal with GET /connect admitted, the unmatched fallthrough, and that the root token is rejected on every tunnel route through buildFetchHandler. * test(browse): ratchet (b) keeps route dispatch inside the route table Scans browse/src/server.ts and browse/src/routes/*.ts for pathname comparisons; only the table matcher and the tunnel-surface filter are allowed, listed with reasons in browse/test/fixtures/route-dispatch-allowlist.json (keyed on file plus line text). Also checks every entry declares auth and surfaces and that gstack registers no beforeRoute overlay itself. Self-tests plant a violation and assert the file:line, Fix: and allowlist path in the message, that a shifted line stays allowlisted, and that a reasonless entry is rejected. * test: touchfile superset check for modules moved out of browse/src/server.ts Records the paid evals selected by touching browse/src/server.ts at96764e80(17 E2E, 1 LLM judge) and asserts every browse/src/routes/*.ts module selects a superset. The test reads every golden in test/fixtures/moved-module-selection/ so other moved-code goldens can sit beside it. * test(shard-engine): give non-timeout corpus fixtures CI headroom; keep the POSIX golden off the Windows lane Only the wall-timeout fixture keeps a 3s wall; the rest get 60s so a loaded host cannot turn a pass into a timeout. Base and branch runners still agree on every classification under the new walls. The Windows exclusion entry moves the free runner's ratchet (c) residual cap to 2356 lines. * refactor(pty): one runPtySession loop drives observation, counting and floor test/helpers/pty/session.ts owns launch -> start -> (poll -> tick)* -> timeout and the failure contract the three runners each hand-rolled: the run's own error wins over capture and close errors, close always runs, owned fixture cleanup runs last (also when launch fails). Each runner now supplies a PtySessionPlan: its boot/command step, poll cadence (2s observation/floor sleep; counting's output wake + 250ms coalesce), tick policy (permission handling, native identity, terminal rules stay per runner because they differ) and capture hooks. The runner bodies are decomposed into top-level steps so no function exceeds 150 lines; behavior is unchanged and the fake-driver cases from the first W4 commit pass unmodified. The counting capture step and the native completion-summary predicate are now named functions (countingCapture, isNativeCompletionSummary), so plan-create-prepublication and plan-count-completion call them directly instead of executing sliced source. The two harnesses that still execute a sliced runner body with injected boundaries (eng-seeded-completion-ai, plan-floor-permission) pass the PtyDriver seam instead of overriding Date/Bun.sleep. * test(ratchet-c): register route modules, review resolver modules and server.ts residual cap * refactor(pty): decompose launchClaudePty and engNumberedFindingAUQ under 150 lines launchClaudePty (349 lines) becomes launch preparation (args, hermetic child env, owned state roots), recorder creation, spawn, the trust-dialog watcher, close, and the session handle over one PtyProcess state object. The failure order is unchanged: abort the viewport, dispose any recorders created so far, dispose the viewport, rethrow. The --model / --strict-mcp-config ordering and seedSkills wiring stay pinned by the behavioral fake-CLI tests. engNumberedFindingAUQ (345 lines) keeps its guards and dispatch; each self-contained issue family (declared cache, library retry hooks, cache owner, injected singleton, shared writers, injected export) moves verbatim into its own function. Every pty/ module is now <= 800 lines and every top-level function <= 150 lines. * test(pty): split claude-pty-runner.unit.test.ts along the pty/ module seams The 188 unit tests move verbatim into claude-pty-runner.{screen,classify, auq,launch,plan-native,boundaries}.unit.test.ts (test names unchanged; each file imports only what it uses from the barrel). The five files that no longer read a SKILL.md template join the test-of-test ratchet baseline with a reason. * test(touchfiles): moved PTY modules keep their paid-eval selection test/fixtures/touchfile-selection/w4-pty.json records, at96764e8, the paid evals selected by touching test/helpers/claude-pty-runner.ts (20) and test/helpers/pty-screen.ts (20). touchfiles.test.ts now asserts every .ts file under test/helpers/pty/ (and pty/screen.ts for both sources) selects a superset, reading every golden in that directory so later moves can add one; a planted-violation case pins the report and its Fix line. * fix(browse): unexchanged pair setup keys no longer authenticate bearer requests validateToken accepted a gsk_setup_ key as a bearer on /command, /batch and /file (found while building the W3 auth matrix). A setup key now only authenticates the /connect exchange. * W1: one state-root owner (lib/state-root.ts + bin/gstack-state-root.sh), gstack-paths --explain and fail-stop, parity tests * W1: guarded migration of every executable state-root site; uninstall deletes only ~/.gstack Bins, careful/freeze hooks, setup, upgrade migrations, browse/src, design, ios-qa daemon, lib and scripts resolve the state root through bin/gstack-state-root.sh (bash) or lib/state-root.ts (TS). Bins source the twin and stop with a reinstall message when it is missing; hooks source it and never spawn gstack-paths. browse/src/config.ts and lib/cso/state.ts delegate to resolveStateRoot. Analytics writers and readers move together so the usage log stays one file. gstack-uninstall deletes state only at ~/.gstack, refuses (exit 2) when it resolves to /, $HOME or an ancestor, the checkout or the git root, and leaves any other resolved root in place with the removal command. Fixtures that copy single bins now copy the twin. * W1: privacy keys and trust-policy deny tiers merge across state roots; gstack-config reporting; test hermeticity readConfigKey / gstack_read_config_key return the most restrictive telemetry, memorable_recall, codex_reviews and update_check across the resolved root and ~/.gstack; other keys read the resolved root only. gstack-config set reports an overriding root with the exact override command, list shows the winning root and a root-variable disagreement line. gstack-gbrain-repo-policy get merges deny/read-only tiers. gstack-egress reads through readConfigKey. test-setup.ts strips inherited GSTACK_STATE_ROOT/GSTACK_STATE_DIR and redirects the legacy root. * W1: shared hook logging helper (hosts/claude/hooks/hook-log.ts) One hook-errors.log writer: root from resolveStateRoot, 0600 on every append, opt-in rate limit used only by memorable-user-prompt. The five hooks route through it. * W1: docs/state-root.md and README troubleshooting pointer Precedence table, a real --explain example, the move-your-state recipe, merged privacy keys, the uninstall rule, the resolver-failure fix, and the plugin-mode note (evidence gate: no official plugin distribution). * W1b: template and resolver prose resolve state through guarded gstack-paths; ratchet (a) Every gstack-paths eval in templates and resolvers carries the fail-stop guard; executable ~/.gstack paths in bash blocks (context recovery preamble, eureka log, analytics, project artifacts, upgrade snooze, setup-gbrain lock, retro snapshots, ship consent marker) use $GSTACK_STATE_ROOT, and the writer prose that pairs with them points at the printed PROJECT_DIR / RETRO_FILE. ship drops export GSTACK_STATE_ROOT. SKILL.md regenerated (claude + codex), ship goldens re-pinned, parity and context-budget caps raised to the measured sizes with notes. test/state-root-ratchet.test.ts enforces the rule with a reasoned allowlist; W1 touchfile entries plus a superset golden. * refactor: apply W1 state-root edits in W2/W3/W5-owned files; one moved-code touchfile golden for all workstreams * test: fold the moved-code touchfile golden into touchfiles.test.ts; fix integration fixture closure and caps * v1.91.11.0: CHANGELOG, TODOS, docs and conventions for the refactor wave * test: re-measure plan-ceo/design-consultation caps and ship goldens after the guarded plan-discovery and spec-review blocks; add the state-root twin to the workflow-boundaries fixture * fix(windows): migrations resolve their directory with either path separator; state-root parity compares under the HOME Git Bash actually sees * fix(review,ship): state plan-check timing after smoke expiry and test_stub Skip semantics (review workflow judge clarity) * test(qa-eval): webhook fix eval asks for the fix loop's post-repair probes; eight-scenario coverage stays in the report-only case and the harness recheck * test(qa-eval): re-pin the webhook prompt contract to the fix-loop stage; R29 coverage omissions stay bound by the report-only case * fix(review,ship): plan checks publish a checkpoint before each probe; only the smoke expiry stop is skipped * fix(qa): carry #2999's checkpoint receipt link, report-template line and full-revision placeholder (identical hunks) * test(qa-callers): disable git auto maintenance in the caller fixture Git 2.47+ runs auto maintenance detached after commit; on the CI runner's git 2.55 it rewrote .git/objects fan-out directories while the write observer was running, which surfaced as unauthorized mutations. Same gc.auto=0 / maintenance.auto=false guard the shared-libs fixture already uses. * test(plan-mode-no-op): require prose evidence for the prose-fallback members so a spinner-frame judge verdict cannot end the run as asked * test(ship-docsync): carry #2999's seeded-attempt docsync harness (identical files) The doc-sync fault cases replayed attempt 1 before reaching their gate and ran out of their 285s budget. The fixture now seeds attempt 1 and the parent starts at the gate under test. Taken byte-identical from origin/capy/audit-fix-wave (fb526898,e6ac813d,6ce10ff7,d0c53577,77cce3be). Local: stale-before, recovery and late-result 6/6 PASS (97-164s); the whole file 12/12 PASS.
659 lines
32 KiB
Cheetah
659 lines
32 KiB
Cheetah
---
|
||
name: ship
|
||
preamble-tier: 4
|
||
version: 1.0.0
|
||
description: |
|
||
Ship workflow: detect + merge base branch, run tests, review diff, bump VERSION,
|
||
update CHANGELOG, commit, push, create PR. Use when asked to "ship", "deploy",
|
||
"push to main", "create a PR", "merge and push", or "get it deployed".
|
||
Proactively invoke this skill (do NOT push/PR directly) when the user says code
|
||
is ready, asks about deploying, wants to push code up, or asks to create a PR. (gstack)
|
||
allowed-tools:
|
||
- Bash
|
||
- Read
|
||
- Write
|
||
- Edit
|
||
- Grep
|
||
- Glob
|
||
- Agent
|
||
- AskUserQuestion
|
||
- WebSearch
|
||
sensitive: true
|
||
triggers:
|
||
- ship it
|
||
- create a pr
|
||
- push to main
|
||
- deploy this
|
||
---
|
||
|
||
{{PREAMBLE}}
|
||
|
||
{{THIRD_PARTY_ACTIONS}}
|
||
|
||
# Ship: Fully Automated Ship Workflow
|
||
|
||
STOP blocks advancement until the stated repair/resume route clears; without one, end this attempt.
|
||
Answer each AskUserQuestion before continuing.
|
||
Routine authorization never waives those gates or their required user decisions.
|
||
|
||
**Routine work needs no confirmation:** include uncommitted changes, choose MICRO/PATCH
|
||
under Step 12, draft CHANGELOG and commits, mark completed TODOs and auto-fix findings.
|
||
When Step 7 coverage meets its target, report remaining gaps and verify generated
|
||
tests without another permission question. Step 15 commits those tests.
|
||
|
||
**Route:** integrate (1–3) → test and review (4–11.5) → prepare the release
|
||
(12–15) → verify frozen content (16) → push and publish (17–21).
|
||
Every new invocation repeats Steps 1–16, including both reviews and the docs audit.
|
||
Steps 12, 17 and 19 prevent duplicate bumps, pushes and PRs, never verification.
|
||
|
||
### Keep state between steps
|
||
|
||
Keep one private Markdown **invocation record** outside the product tree and save
|
||
its absolute path. Use these headings so a paused run can resume:
|
||
- **Release:** versions, `BUMP_LEVEL`, reviewed tree and attempt counts.
|
||
- **Decisions:** each approval's finding, files and authorized action. Reuse it only
|
||
for that same scope; a repair never resets approvals or expands them.
|
||
- **Reviews:** handles, original start tokens, terminal states, outputs and queued fixes.
|
||
- **Checks:** command/label, result/counts, timestamp, log and consumed inputs.
|
||
- **Documentation:** candidate/id, attempts used, accepted hashes or named blocked exception.
|
||
- **Next steps:** one ordered work list, with the current step marked.
|
||
|
||
A **receipt** is saved evidence of a check's command, result and consumed content.
|
||
A review's **start token** is the opaque value returned by `gstack-review-log --start`
|
||
before it reads the diff. Keep `REVIEW_START` for Step 9, a separate `PASS_START` for
|
||
each Step 11 attempt, and `DESIGN_START` for design. Finish each pass with its original
|
||
token; `--finish` stamps the binding fields automatically. Never borrow or replace a token.
|
||
`gstack-wtree` prints a Git tree hash covering tracked and non-ignored untracked files,
|
||
not a commit ID. Use `git diff <old-tree> <new-tree>` to compare these snapshots.
|
||
|
||
### Ship control flow
|
||
|
||
You, the **parent** running /ship, own advancement; children return evidence, not
|
||
permission to proceed. Follow the saved work list:
|
||
|
||
1. Start with Steps 1–21 in order, including 11.5 and 14.5. Advance only after
|
||
the current item's gates clear.
|
||
2. Expand a repair into individual steps and insert them before the still-pending
|
||
work. This replaces the current item, whose actual result stays in the record.
|
||
Add its destination only if not already the next pending step.
|
||
3. For another repair, repeat rule 2 without discarding pending work.
|
||
The saved list takes precedence over ordinary next-step
|
||
sentences inside a repair. A range never adds unlisted steps.
|
||
|
||
**Example:** Step 11 fixes insert `9 → 10 → 11` before 11.5. A further Step 9 fix
|
||
affecting 6–8 makes the list `5 → 6 → 7 → 8 → 9 → 10 → 11 → 11.5`.
|
||
The unchanged release steps follow. STOP and AskUserQuestion gates still apply during repairs.
|
||
|
||
Keep the same attempt counts throughout the invocation. A range ending at Step 14
|
||
does not enter Step 14.5. A range that includes Step 14.5 enters its existing audit
|
||
decision, not an unconditional new launch; its initial-plus-ONE limit never resets.
|
||
Permitted repairs continue in this invocation without restarting /ship.
|
||
|
||
---
|
||
|
||
{{SECTION_INDEX:ship}}
|
||
|
||
---
|
||
|
||
{{BASE_BRANCH_DETECT}}
|
||
|
||
`<base>` means the detected branch name for fetch/helper arguments;
|
||
`origin/<base>` is its remote-tracking ref for comparisons. Step 1 fetches it.
|
||
|
||
{{GBRAIN_CONTEXT_LOAD}}
|
||
|
||
## Step 0.9: Apple target detection
|
||
|
||
If the ask is App Store/TestFlight distribution, look for an `.xcodeproj`,
|
||
`.xcworkspace`, or Swift app product. Read `Package.swift` and its entrypoint to
|
||
distinguish an app from a library/CLI. If unclear, use AskUserQuestion to identify
|
||
the target and wait before choosing a release path.
|
||
For a confirmed app, **STOP and Read
|
||
`~/.claude/skills/gstack/ship/sections/apple-release.md` FIRST**. Store distribution proceeds
|
||
through that adapter from the current branch, including a clean base branch.
|
||
The branch gate and repository-landing pipeline below apply ONLY to
|
||
repository-landing asks, including on Apple repos.
|
||
|
||
## Step 1: Pre-flight
|
||
|
||
1. Save the current branch as `<branch-name>`. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
|
||
|
||
2. Run `git status` (never use `-uall`). Uncommitted changes are always included — no need to ask.
|
||
|
||
3. Run `git fetch origin <base>` before inspecting the diff. If fetch fails, STOP:
|
||
report the error and restore access before continuing. Then inspect
|
||
`git diff origin/<base> --stat`, untracked files from status, and
|
||
`git log origin/<base>..HEAD --oneline`.
|
||
|
||
4. Display historical readiness using the dashboard below, then finish Step 1.
|
||
Prior CLEAR reviews or dashboard skips never replace Step 9's gates.
|
||
|
||
{{REVIEW_DASHBOARD}}
|
||
|
||
For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend
|
||
`/plan-eng-review` or `/autoplan` for architecture review.
|
||
|
||
For Design Review: run `source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)`. If `SCOPE_FRONTEND=true` and no design review exists, mention: "Design Review not run — Step 9 includes the lite check; consider /design-review for a full visual audit."
|
||
|
||
Continue to Step 2 without asking; Step 9 applies the review gates.
|
||
|
||
---
|
||
|
||
## Step 2: Distribution Pipeline Check
|
||
|
||
Check distribution for new standalone artifacts (CLI binaries, packages, tools),
|
||
not web services with existing deployment.
|
||
|
||
1. List candidate distribution paths:
|
||
```bash
|
||
git diff origin/<base> --diff-filter=A --name-only | grep -E '(^|/)(cmd/[^/]+/main\.go|bin/[^/]+|Cargo\.toml|setup\.py|package\.json)$' | head -5
|
||
```
|
||
Also inspect matching untracked files from Step 1's status. Read each match:
|
||
a new `package.json` or `Cargo.toml` alone does not establish a publishable
|
||
artifact. Also inspect existing manifests for newly declared binaries or
|
||
package exports. Apply the pipeline gate only when a new distributable is present.
|
||
|
||
2. If new artifact detected, check for a release workflow:
|
||
```bash
|
||
ls .github/workflows/ 2>/dev/null | grep -iE 'release|publish|dist'
|
||
grep -qE 'release|publish|deploy' .gitlab-ci.yml 2>/dev/null && echo "GITLAB_CI_RELEASE"
|
||
```
|
||
|
||
3. **New artifact without a pipeline:** AskUserQuestion: "Users cannot download this
|
||
artifact after merge without a release pipeline."
|
||
- A) Add the platform's release workflow now
|
||
- B) Defer with a P1 distribution TODO in Step 14
|
||
- C) Not needed: internal/web-only, covered by existing deployment
|
||
|
||
4. **If A:** Add packaging/publish configuration using repository CI conventions.
|
||
Ask for unknown targets, registries or access first; never invent credentials.
|
||
Recheck against the artifact and include the workflow in tests and review.
|
||
Do not publish a release during `/ship`.
|
||
5. Otherwise, continue without adding a pipeline.
|
||
|
||
---
|
||
|
||
## Step 3: Merge the base branch (BEFORE tests)
|
||
|
||
Merge the base ref fetched in Step 1 so tests and reviews cover the integrated code:
|
||
|
||
```bash
|
||
git merge origin/<base> --no-edit
|
||
```
|
||
|
||
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). For complex or ambiguous conflicts, **STOP**, show the conflicting choices, use AskUserQuestion for the needed resolution decision, and wait for the answer before editing or continuing.
|
||
|
||
**If already up to date:** Continue silently.
|
||
|
||
If integration changes the artifact or distribution configuration inspected in Step 2,
|
||
repeat Step 2 on the merged content, including its decisions, then continue to Step 4.
|
||
Otherwise continue to Step 4 directly.
|
||
|
||
---
|
||
|
||
{{SECTION:tests}}
|
||
|
||
{{SECTION:test-coverage}}
|
||
|
||
{{SECTION:plan-completion}}
|
||
|
||
{{SECTION:review-army}}
|
||
|
||
{{SECTION:greptile}}
|
||
|
||
{{SECTION:adversarial}}
|
||
|
||
## Step 11.5: Bind the reviews
|
||
|
||
1. **Select the two reviews.** Run `~/.claude/skills/gstack/bin/gstack-review-read`.
|
||
Select this invocation's final Step 9.4 record (`skill:"review"`, `via:"ship"`)
|
||
and Step 11 native record (`skill:"adversarial-review"`). Match each to its saved
|
||
handle, original token and source; reject outside-provider or older invocation records.
|
||
2. **Compare their content.** Require the native record's `review_binding.state`
|
||
to be `verified`. All three snapshots must match: its `wtree`, Step 9.4's
|
||
`review_binding.start_wtree` and `review_binding.end_wtree`. A mismatch or missing
|
||
record/field blocks release preparation: report **Review records missing or mismatched**
|
||
and insert `9 → 10 → 11 → 11.5` before Step 12. Bind the new records at 11.5.
|
||
Never attach new tokens to old work.
|
||
3. **Preserve any QA exception.** A named probe-risk exception may leave Step 9.4's
|
||
root `wtree` absent; item 2 still compares its start/end snapshots. Matching content
|
||
does not mean the failed or unrun probes passed. Keep Step 9.4's incomplete flags
|
||
and the user's exception.
|
||
4. **Save the evidence.** Save both records and matching **reviewed tree** for
|
||
Step 16. Continue to Step 12.
|
||
|
||
## Step 12: Version bump (auto-decide)
|
||
|
||
Item 3 needs `BUMP_LEVEL`: reuse this invocation's saved level. Otherwise FRESH
|
||
chooses it in item 2 and ALREADY_BUMPED derives it in item 1.
|
||
|
||
1. **Classify state** — pure reader, never writes:
|
||
```bash
|
||
bun run ~/.claude/skills/gstack/bin/gstack-version-bump classify --base <base>
|
||
```
|
||
Save the JSON `baseVersion` as `BASE_VERSION`, then read `state` and dispatch:
|
||
- **FRESH** → use the recorded level or choose it in item 2, then check the queue and write.
|
||
- **ALREADY_BUMPED** → keep `NEW_VERSION=currentVersion`. If `BUMP_LEVEL` is missing,
|
||
use the first changed component from `baseVersion` to `currentVersion`
|
||
(major/minor/patch/micro; an absent fourth component is zero). Continue at item 3,
|
||
not another automatic bump.
|
||
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify.
|
||
Success follows ALREADY_BUMPED, including its queue check; failure stops.
|
||
Repair alone never re-bumps.
|
||
- **DRIFT_UNEXPECTED** → STOP: package.json disagrees with VERSION while VERSION
|
||
matches base. Reconcile the manual edit, then reclassify.
|
||
|
||
2. **Decide the bump level** from the diff (agent judgment):
|
||
- **MICRO**: <50 lines, trivial tweaks/config. **PATCH**: 50+ lines, no feature signals.
|
||
- **MINOR**: ask for any feature signal (new route/page, migration, module) or 500+ lines.
|
||
**MAJOR**: ask for milestones or breaking changes. Use AskUserQuestion: recommended
|
||
level with rationale, smaller level, or cancel. Wait; cancel stops before release
|
||
writes or push and preserves existing work.
|
||
Save lowercase `BUMP_LEVEL`. A claimed version may move the next available number
|
||
forward, but cannot change the chosen MICRO/PATCH/MINOR/MAJOR level.
|
||
|
||
3. **Queue-aware pick** (workspace-aware ship):
|
||
```bash
|
||
QUEUE_JSON=$(bun run ~/.claude/skills/gstack/bin/gstack-next-version --base <base> --bump "$BUMP_LEVEL" --current-version "$BASE_VERSION" 2>/dev/null || echo '{"offline":true}')
|
||
CANDIDATE_VERSION=$(echo "$QUEUE_JSON" | jq -r '.version // empty')
|
||
```
|
||
**Qualify first:** require successful utility output and a nonempty valid version.
|
||
`offline:false` qualifies; `offline:true` qualifies only with `fallback:"git"`.
|
||
Offline output without that fallback, failure, malformed output or an empty version
|
||
is unusable, even if it contains a version-looking string.
|
||
|
||
- **Usable candidate:** print warnings and claimed queue. FRESH sets `NEW_VERSION=CANDIDATE_VERSION`.
|
||
ALREADY_BUMPED compares it with `currentVersion`: if different, ask to rebump
|
||
(refresh CHANGELOG/PR title) or keep current (CI rejects a collision).
|
||
Only approval changes the existing version. Check JSON `active_siblings` by
|
||
`branch` and `version`; a sibling holding `>= NEW_VERSION` requires a choice:
|
||
advance past it, or stop this attempt and sync.
|
||
- **No usable candidate:** print queue-unverified. FRESH uses local `BUMP_LEVEL`
|
||
arithmetic; ALREADY_BUMPED keeps `currentVersion`. Never use an empty candidate.
|
||
|
||
4. **Write the bump** (FRESH, or an approved rebump):
|
||
```bash
|
||
bun run ~/.claude/skills/gstack/bin/gstack-version-bump write --version "$NEW_VERSION" --regen-digest
|
||
```
|
||
The CLI validates `MAJOR.MINOR.PATCH.MICRO` (or pinned 3-digit semver) and writes
|
||
VERSION, the manifest and existing `package-lock.json` / `npm-shrinkwrap.json`;
|
||
it never creates lockfiles. Manifest path: `--package-json-path` →
|
||
`.gstack/package-json-path` → `./package.json`. npm files use the 3-digit translation
|
||
(`1.67.0.0` → `1.67.0`); VERSION is authoritative. Exit 3 means a half-write:
|
||
reclassify and `repair` DRIFT_STALE_PKG.
|
||
|
||
`--regen-digest` runs repo code with Step 5's privileges: `scripts/gen-agents-digest.ts`,
|
||
only when it and committed `agents-digest/gstack-AGENTS.md` exist. If `agentsDigest`
|
||
is false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump.
|
||
Before push, verify the committed digest matches generation for the selected VERSION.
|
||
|
||
5. **Record the release decision after a version was actually written**, including
|
||
an approved ALREADY_BUMPED rebump. Skip unchanged versions and manifest-only repairs.
|
||
```bash
|
||
~/.claude/skills/gstack/bin/gstack-decision-log '{"decision":"Ship NEW_VERSION (BUMP_LEVEL)","rationale":"WHY","scope":"repo","source":"skill","confidence":9}' 2>/dev/null || true
|
||
```
|
||
Substitute `NEW_VERSION`, `BUMP_LEVEL`, and one-line `WHY` (scope or breaking-change signal). Best-effort, non-interactive, non-blocking.
|
||
|
||
{{SECTION:changelog}}
|
||
|
||
## Step 14: TODOS.md (auto-update)
|
||
|
||
Read `~/.claude/skills/gstack/review/TODOS-format.md`.
|
||
|
||
**1. Open or create:** Read root `TODOS.md`. An explicit "add TODO" choice authorizes
|
||
creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: A) Create
|
||
a component/priority-organized TODOS.md, B) Skip. Skip goes to item 5.
|
||
|
||
**2. Organization:** Use component headings, `**Priority:**` P0–P4 and `## Completed`
|
||
at the bottom. If disorganized, ask: A) Reorganize preserving all content
|
||
(recommended), B) Leave as-is.
|
||
|
||
**3. Add approved deferrals:**
|
||
- Step 2: add the approved distribution follow-up as P1 with the missing pipeline and affected artifact.
|
||
- Step 8: add each approved P1 plan deferral with `Deferred from plan: {plan file path}` and the missing work.
|
||
- Step 5: retain P0 test-failure entries already written; deduplicate by failure and source, adding missing approved entries with error output and branch.
|
||
Never turn dropped scope into TODOs or invent unapproved follow-ups. Reuse matching existing entries rather than duplicating them.
|
||
|
||
**4. Detect completed TODOs:** Compare titles, files and behavior with
|
||
`git diff origin/<base>`, untracked files and `git log origin/<base>..HEAD --oneline`.
|
||
Move proven completions to `## Completed` with `**Completed:** vX.Y.Z (YYYY-MM-DD)`;
|
||
leave uncertain items open.
|
||
|
||
**5. Save the summary:** Report additions, deferrals, completions, remaining count and
|
||
creation/reorganization. If creation was declined or a write failed, warn and retain
|
||
unsaved follow-ups in Step 19's PR summary. Never claim they were saved;
|
||
TODO write failures are non-blocking.
|
||
|
||
---
|
||
|
||
## Step 14.5: Documentation audit (every ship)
|
||
|
||
**Doc-sync invariant:** Every ship dispatches the /document-release subagent before final
|
||
commit/verification/publication, including reruns, already-pushed branches, existing PRs and docs-only changes.
|
||
No edits means an executed audit, not a skip; report the section's verified outcome.
|
||
|
||
{{SECTION:documentation}}
|
||
|
||
## Step 15: Commit (bisectable chunks)
|
||
|
||
Make bisectable commits; if already committed, continue to Step 16. Never create an empty commit.
|
||
|
||
1. Group changes with their tests, config/routes, views and Step 14.5 docs.
|
||
Migrations may stand alone or accompany their model.
|
||
Under 50 lines across fewer than 4 files may use one commit.
|
||
2. Order dependencies first: infrastructure → models/services → controllers/views.
|
||
Each commit must work independently, without broken imports or missing code.
|
||
Group VERSION + CHANGELOG + TODOS.md after the feature commits.
|
||
3. Use `<type>: <summary>` (feat/fix/chore/refactor/docs) and a brief body.
|
||
Only the final VERSION/CHANGELOG commit gets the release version and co-author
|
||
trailer. Do not create a Git tag:
|
||
|
||
```bash
|
||
git commit -m "$(cat <<'EOF'
|
||
chore: bump version and changelog (vX.Y.Z.W)
|
||
|
||
{{CO_AUTHOR_TRAILER}}
|
||
EOF
|
||
)"
|
||
```
|
||
|
||
---
|
||
|
||
## Step 16: Verification Gate
|
||
|
||
**IRON LAW: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE.**
|
||
|
||
Run stages 1–5 in order. Recovery instructions below name where to resume.
|
||
If content changes during or after verification, restart at stage 1 and complete
|
||
all five stages before Step 17. Content-preserving commits keep valid evidence.
|
||
|
||
### 1. Finish writers and prepare outputs
|
||
|
||
Inspect writer handles, including the docs child. Confirm terminal completion or termination
|
||
before another writer runs. Timeout or cancellation acknowledgment alone means
|
||
STOP until confirmed.
|
||
|
||
Find declared generation/build commands in project instructions, manifests, build
|
||
files and CI. Run them and save results. If none exists, record not applicable and
|
||
the inspected sources. A missing prerequisite or failed build stops shipping:
|
||
report **Build failed or prerequisite missing**, with the command, error and needed
|
||
repair. Never invent a substitute command.
|
||
**If blocked:** Repair the prerequisite or build, then repeat stage 1. After it passes, continue
|
||
to stage 2; treat any content repair as a behavioral change there.
|
||
|
||
### 2. Choose the change route
|
||
|
||
Capture the current tree with `~/.claude/skills/gstack/bin/gstack-wtree`. Inspect
|
||
`git diff <reviewed-tree> <current-tree>` against the snapshot saved before Step 12.
|
||
Missing snapshots block this comparison, regardless of HEAD equality.
|
||
|
||
Classify the comparison in this order:
|
||
|
||
1. **Behavior, tests or build inputs changed:** Prompts/templates count as behavior.
|
||
Insert `5–11.5 → 12–14 → 16` before the pending Step 17, then stop this step.
|
||
This repair excludes Step 14.5 because the rebuild can change generated docs.
|
||
Step 16 restarts at stage 1: rebuild and compare again before stage 3 decides
|
||
documentation freshness. Further repairs use the same work list.
|
||
2. **Only authored docs or release metadata changed:** Keep Step 8's original child
|
||
report and counts. Recheck affected plan items using their recorded verification
|
||
and append current evidence to the invocation record. If a classification is no
|
||
longer supported, run Step 8's audit and decision gates only, then return to
|
||
Step 16 stage 1. Never edit the child's counts yourself.
|
||
3. **No changes, or the docs-only checks still support the plan:** Continue to stage 3
|
||
without a new code review.
|
||
|
||
### 3. Resolve documentation freshness
|
||
|
||
Compare the base and hashes of the selected release paths, generated
|
||
outputs and docs/templates with Step 14.5's saved values. A prior invocation's
|
||
audit or risk decision never qualifies.
|
||
|
||
| Outcome | Action |
|
||
|---|---|
|
||
| This invocation's accepted audit matches all inputs | Continue to stage 4. |
|
||
| User-accepted named documentation risk covers the same approved scope and exact content, and unwaivable gates clear | Continue to stage 4; retain `Documentation: blocked`, its reason and incomplete scope. |
|
||
| Missing, stale or blocked | Use recovery below. Never silently refresh hashes. |
|
||
|
||
Report changed inputs, blockers and attempts used:
|
||
|
||
- **An attempt remains, with changed inputs or an available repair:** insert
|
||
`14.5 → 15 → 16` before Step 17. Use Blocked recovery with the existing count.
|
||
Validate the outcome before Step 15,
|
||
then restart Step 16 stage 1 to regenerate and compare again.
|
||
- **Otherwise:** STOP unless the user accepts
|
||
the specific named documentation risk and all unwaivable gates clear, under
|
||
Step 14.5's Blocked recovery rules. Unchanged approved content goes to stage 4;
|
||
repaired content goes to stage 1.
|
||
|
||
Never run a third audit. Child return is not acceptance.
|
||
|
||
### 4. Verify the frozen candidate
|
||
|
||
Freeze inputs through verification and push. Run declared docs/link/generated-file
|
||
checks; report unavailable checks.
|
||
|
||
**Reuse a check when its inputs match.** Compare hashes or complete bytes of its
|
||
saved and current consumed files, fixtures, dependencies and execution parameters.
|
||
Explain why other changes cannot affect it; changed or unknown dependencies require a rerun.
|
||
For model judges, compare the complete expanded request, rubric, parameters and
|
||
builder/runtime dependencies. Reuse identical passing evidence: cite the original
|
||
command, result/counts, timestamp and log, never resample it. Mandatory reviews still run.
|
||
|
||
**Check each test lane's receipt as well.** Use its actual Step 5 label/command:
|
||
`--label <lane> --expect-cmd '<exact Step 5 command>'`. Inspect changes since the run;
|
||
`--allow-paths` exempts only release metadata. A `package.json` version-only edit
|
||
can qualify; scripts, dependencies and runtime configuration require live tests.
|
||
Uncertain edits cannot be exempted. Docs, TODO edits, new/generated tests and fixes
|
||
make evidence STALE even without a new code review. Use this example only after
|
||
confirming that every allowed edit is release metadata:
|
||
|
||
```bash
|
||
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<tests>' --label vitest --expect-cmd '<vitest>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
|
||
```
|
||
|
||
| Receipt result | Next action |
|
||
|---|---|
|
||
| FRESH (exit 0) | Cite the label, exit, timestamp and log. |
|
||
| STALE/MISSING: changed content, command or age, or no proven run | Run `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`, read the result and recheck once. Handle failures as described below. |
|
||
| Only receipt storage/readback failed | Independently prove unchanged final content, the same command and valid age from the successful run's evidence. Cite its exact command, exit, timestamp and log as **ledger unavailable**, never FRESH. Without that proof, use STALE/MISSING. |
|
||
|
||
No test lanes: require Step 5's explicit untested-scope approval for final content,
|
||
or run Steps 5–15, including the no-tests decision, then return to Step 16 stage 1.
|
||
Report the gap, never FRESH; builds must pass.
|
||
|
||
**New, changed or unwaived test failure:** STOP publication. Run Steps 5–15,
|
||
starting with Step 5's triage, then return to Step 16 stage 1. This recovery also
|
||
applies if a failure appears while reporting in stage 5. Reentry to Step 14.5
|
||
keeps its existing audit count; it does not authorize a third attempt.
|
||
|
||
### 5. Report, then push
|
||
|
||
Commit only approved, verified release changes left uncommitted after Step 15,
|
||
including generated outputs; use its grouping rules and never create an empty commit.
|
||
Preserve unrelated user files.
|
||
|
||
Paste build/docs/test results. Reuse waivers only for the same verified
|
||
pre-existing failures and approved scope; cite the actual approval and failing
|
||
counts, never FRESH or all-green. A new, changed or unwaived test failure uses
|
||
stage 4's recovery before publication. Otherwise continue to Step 17.
|
||
|
||
---
|
||
|
||
## Step 17: Push
|
||
|
||
**Credential pre-push guard (#1946) — run before the push:**
|
||
|
||
```bash
|
||
_REDACT_PREPUSH=$(~/.claude/skills/gstack/bin/gstack-config get redact_prepush_hook 2>/dev/null || echo "false")
|
||
_HOOK_PATH=$(git rev-parse --git-path hooks/pre-push 2>/dev/null || echo "")
|
||
_HOOK_STATE="missing"
|
||
if [ -e "$_HOOK_PATH" ] || [ -L "$_HOOK_PATH" ]; then
|
||
_HOOK_STATE="unmanaged"
|
||
if [ -f "$_HOOK_PATH" ] && [ ! -L "$_HOOK_PATH" ] && grep -Fqx '# gstack-redact pre-push (managed)' "$_HOOK_PATH" 2>/dev/null; then
|
||
_HOOK_STATE="managed"
|
||
fi
|
||
fi
|
||
_HOOKS_DIR=$(git rev-parse --git-path hooks 2>/dev/null || echo "")
|
||
_HOOKS_IN_GIT_DIR="no"
|
||
_HOOKS_CONFIG_STATUS=0
|
||
git config --get core.hooksPath >/dev/null 2>&1 || _HOOKS_CONFIG_STATUS=$?
|
||
if [ -n "$_HOOK_PATH" ] && [ -n "$_HOOKS_DIR" ] && [ "$_HOOKS_CONFIG_STATUS" = "1" ] && [ ! -L "$_HOOKS_DIR" ]; then
|
||
_HOOKS_IN_GIT_DIR="yes"
|
||
fi
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
_PREPUSH_PROMPTED=$([ -f "$GSTACK_STATE_ROOT/.redact-prepush-prompted" ] && echo "yes" || echo "no")
|
||
if [ "$_REDACT_PREPUSH" = "true" ] && [ "$_HOOKS_IN_GIT_DIR" = "yes" ] && [ "$_HOOK_STATE" != "unmanaged" ]; then
|
||
~/.claude/skills/gstack/bin/gstack-redact install-prepush-hook || exit $?
|
||
fi
|
||
echo "REDACT_PREPUSH: $_REDACT_PREPUSH"
|
||
echo "HOOK_STATE: $_HOOK_STATE"
|
||
echo "HOOKS_IN_GIT_DIR: $_HOOKS_IN_GIT_DIR"
|
||
echo "PREPUSH_PROMPTED: $_PREPUSH_PROMPTED"
|
||
```
|
||
|
||
Branch on the echoed values:
|
||
|
||
1. **`REDACT_PREPUSH: true`** — the block installs or refreshes managed
|
||
hooks, preserving `pre-push.local` and complete stdin. On installer
|
||
failure, STOP before pushing. `HOOKS_IN_GIT_DIR: no`: do not install;
|
||
request manual integration. `HOOK_STATE: unmanaged`: ask consent only
|
||
for a regular, non-symlink hook in the default directory without
|
||
`pre-push.local`; otherwise request manual integration. Dangling
|
||
symlinks are unmanaged. Never overwrite either policy.
|
||
2. **`REDACT_PREPUSH` not true AND `PREPUSH_PROMPTED: no`** — one-time
|
||
offer (fires once EVER, machine-wide). AskUserQuestion:
|
||
|
||
> gstack can install a per-repo git pre-push hook that blocks pushes
|
||
> containing credentials (API keys, tokens, private keys). It's a
|
||
> guardrail, not enforcement — `GSTACK_REDACT_PREPUSH=skip` bypasses it.
|
||
> Install it for repos you ship from?
|
||
|
||
Options:
|
||
- A) Yes — install the credential guard (recommended)
|
||
- B) No — never ask again
|
||
|
||
If A: run `~/.claude/skills/gstack/bin/gstack-config set redact_prepush_hook true`
|
||
then re-run the block and apply the same directory and unmanaged-hook rules above.
|
||
If B: run `~/.claude/skills/gstack/bin/gstack-config set redact_prepush_hook false`.
|
||
ALWAYS (after either answer, but NOT if the question itself failed to
|
||
render — a failed AskUserQuestion must re-offer next time):
|
||
```bash
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
touch "$GSTACK_STATE_ROOT/.redact-prepush-prompted"
|
||
```
|
||
3. **Declined earlier** — continue
|
||
without comment.
|
||
|
||
**Idempotency check:** Check if the branch is already pushed and up to date.
|
||
|
||
```bash
|
||
LOCAL=$(git rev-parse HEAD) || exit 1
|
||
REMOTE_REF=$(git ls-remote --heads origin refs/heads/<branch-name>) || {
|
||
echo "STATUS: BLOCKED — cannot verify remote branch; restore access before pushing"
|
||
exit 1
|
||
}
|
||
REMOTE=$(printf '%s\n' "$REMOTE_REF" | awk '{print $1}')
|
||
REMOTE=${REMOTE:-none}
|
||
echo "LOCAL: $LOCAL REMOTE: $REMOTE"
|
||
[ "$LOCAL" = "$REMOTE" ] && echo "ALREADY_PUSHED" || echo "PUSH_NEEDED"
|
||
```
|
||
|
||
If `ALREADY_PUSHED`, skip the push but continue to Step 18. Otherwise push with upstream tracking:
|
||
|
||
```bash
|
||
git push -u origin <branch-name>
|
||
```
|
||
|
||
**If the push fails, STOP.** No Step 19 or publication claim. Report the error:
|
||
- **Non-fast-forward push:** fetch and inspect the remote, then merge under Step 3's
|
||
conflict rules. Run Steps 5–16 before returning to Step 17. Never rewrite history.
|
||
- **Authentication, hook or network failure:** repair the cause, then repeat Step 16
|
||
even if content is unchanged before returning to Step 17. Never bypass failed guards.
|
||
Never force-push.
|
||
Only a successful push or verified `ALREADY_PUSHED` proceeds.
|
||
|
||
Continue to Step 18. No documentation writer runs after push.
|
||
|
||
---
|
||
|
||
## Step 18: Prepare publication metadata
|
||
|
||
First look up open PRs/MRs for `<branch-name>` on the detected platform:
|
||
|
||
- GitHub: `gh pr list --head <branch-name> --state open --json number,title,url`
|
||
- GitLab: `glab mr list --source-branch <branch-name> --output json` (defaults to open).
|
||
|
||
A successful empty array means new; one match supplies the existing title/identity.
|
||
Lookup failure or ambiguous matches **STOP** for resolution, never mean no PR.
|
||
Save the result for Step 19's recheck.
|
||
|
||
Prepare the title from that result; Step 19 scans and publishes it:
|
||
1. For an existing open PR/MR, use the matched title and run
|
||
`~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`.
|
||
2. For a new PR/MR, compose `v<NEW_VERSION> <type>: <summary>`.
|
||
3. Save the result as `NEW_TITLE` for Step 19. Every created or updated title MUST
|
||
start with `v$NEW_VERSION `; never publish an unprefixed title.
|
||
|
||
{{SECTION:pr-body}}
|
||
|
||
## Step 20: Persist ship metrics
|
||
|
||
Log metrics for `/retro` through `gstack-review-log`; it handles project/branch paths,
|
||
JSON validation, storage and sync. It takes **no path argument**; do not build one.
|
||
|
||
```bash
|
||
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"ship","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","coverage_pct":COVERAGE_PCT,"coverage_schema":2,"coverage_pct_value":COVERAGE_PCT_VALUE,"weak_gaps":WEAK_GAPS,"tests_extended":TESTS_EXTENDED,"tests_rejected":TESTS_REJECTED,"regression_proof":REGRESSION_PROOF,"plan_items_total":PLAN_TOTAL,"plan_items_done":PLAN_DONE,"verification_result":"VERIFY_RESULT","version":"VERSION","branch":"'"$(git rev-parse --abbrev-ref HEAD)"'"}'
|
||
```
|
||
|
||
Substitute from earlier steps:
|
||
- **COVERAGE_PCT**: Step 7 diagram's integer percentage; encode null/undetermined as -1
|
||
- **COVERAGE_PCT_VALUE**: Step 7's `coverage_pct_value` (the gate's X) as an integer, or `null` when missing or ignored
|
||
- **WEAK_GAPS**, **TESTS_EXTENDED**, **TESTS_REJECTED**: counts of Step 7's `weak_gaps`, `tests_extended` and `tests_rejected` (0 when the key is missing or ignored)
|
||
- **REGRESSION_PROOF**: `{"red_at_head":N,"base_green":N,"base_unavailable":N}` from Step 7, or `null` when missing
|
||
- **PLAN_TOTAL**: total plan items extracted in Step 8 (0 if no plan file)
|
||
- **PLAN_DONE**: count of DONE + CHANGED items from Step 8 (0 if no plan file)
|
||
- **VERIFY_RESULT**: "pass", "fail", or "skipped", set after Step 9 executes Step 8.1's verification list
|
||
- **VERSION**: from the VERSION file
|
||
|
||
The shell supplies the branch. Run this automatically, without confirmation.
|
||
|
||
---
|
||
|
||
## Step 21: Plan-tune discoverability nudge (first-successful-ship only)
|
||
|
||
After a successful ship, show the non-blocking /plan-tune nudge once per machine:
|
||
|
||
```bash
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
_NUDGE_MARKER="$GSTACK_STATE_ROOT/.plan-tune-nudge-shown"
|
||
_QT=$(~/.claude/skills/gstack/bin/gstack-config get question_tuning 2>/dev/null || echo "false")
|
||
if [ ! -f "$_NUDGE_MARKER" ] && [ "$_QT" = "false" ]; then
|
||
echo ""
|
||
echo "gstack can learn from your AskUserQuestion answers. Run /plan-tune to opt in"
|
||
echo "— it captures which prompts you find valuable vs noisy and (with hooks installed)"
|
||
echo "auto-decides your never-ask preferences."
|
||
mkdir -p "$GSTACK_STATE_ROOT" && touch "$_NUDGE_MARKER"
|
||
fi
|
||
```
|
||
|
||
The marker or enabled question_tuning suppresses it. To re-enable, remove
|
||
`$GSTACK_STATE_ROOT/.plan-tune-nudge-shown` before the next ship.
|
||
|
||
---
|
||
|
||
## Section self-check (before you finish)
|
||
|
||
List the applicable Section index entries and confirm each Read. If you worked from
|
||
memory, STOP, Read the section and redo that step. Use `gstack-version-bump`, never
|
||
hand-roll VERSION/package.json writes.
|
||
|
||
---
|
||
|
||
## Important Rules
|
||
|
||
Follow the numbered gates and their explicit exceptions.
|
||
|
||
- **Never force push.** Use regular `git push` only.
|
||
- **Always use the 4-digit version format** from the VERSION file.
|
||
- **Step 7 generates coverage tests.** They must pass before committing. Never commit failing tests.
|