v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+387 -203
View File
@@ -440,47 +440,62 @@ Some steps require action on a site the user controls: registering an API key, c
# Ship: Fully Automated Ship Workflow
Run `/ship` through to the PR URL. This request authorizes routine work without confirmation; explicit safety and user-decision gates still apply.
STOP blocks advancement until the stated repair/resume route clears; without one, end this attempt.
Answer each AskUserQuestion before continuing.
Routine authorization never waives those gates or their required user decisions.
**Route through the workflow:** detect and merge the base (Steps 1–3), test and
audit the integrated diff (Steps 4–8.2), review and resolve findings (Steps
9–11), prepare the release and commits (Steps 12–15), then verify, push, sync
docs, and open or update the PR (Steps 16–19). A review fix returns to affected
tests and reviews before release preparation; a later code or build-input edit
returns to affected checks and Step 16 before publication. Reuse still-valid
results, but never treat an earlier review or test as covering changed inputs.
**Routine work needs no confirmation:** include uncommitted changes, choose MICRO/PATCH
under Step 12, draft CHANGELOG and commits, mark completed TODOs and auto-fix findings.
When Step 7 coverage meets its target, report remaining gaps and verify generated
tests without another permission question. Step 15 commits those tests.
**Follow every STOP and AskUserQuestion gate**, including:
- On the base branch (abort)
- Merge conflicts that can't be auto-resolved (stop, show conflicts)
- In-branch test failures (pre-existing failures are triaged, not auto-blocking)
- Pre-landing review finds ASK items that need user judgment
- Prior Learnings needs its first-time cross-project setting (Step 8)
- MINOR or MAJOR version bump needed (ask — see Step 12)
- Greptile review comments that need user decision (complex fixes, false positives)
- AI-assessed coverage below target (see Step 7 for minimum/target decisions)
- Plan items NOT DONE or UNVERIFIABLE (see Step 8)
- Plan verification failures (see Step 8.1)
- TODOS.md missing and user wants to create one (ask — see Step 14)
- TODOS.md disorganized and user wants to reorganize (ask — see Step 14)
**Route:** integrate (1–3) → test and review (4–11.5) → prepare the release
(12–15) → verify frozen content (16) → push and publish (17–21).
Every new invocation repeats Steps 1–16, including both reviews and the docs audit.
Steps 12, 17 and 19 prevent duplicate bumps, pushes and PRs, never verification.
**Never stop for:**
- Uncommitted changes (always include them)
- Version bump choice (auto-pick MICRO or PATCH — see Step 12)
- CHANGELOG content (auto-generate from diff)
- Commit message approval (auto-commit)
- Multi-file changesets (auto-split into bisectable commits)
- TODOS.md completed-item detection (auto-mark)
- Auto-fixable review findings (dead code, N+1, stale comments — fixed automatically)
- Test coverage gaps within target threshold (generate, verify, then commit with Step 15; flag any remaining gaps in the PR body)
### Keep state between steps
**Re-run behavior (idempotency):**
Every invocation repeats verification: tests, coverage, plan completion, both
reviews, VERSION/CHANGELOG, TODOS and doc-sync. Only *actions* are idempotent:
- Step 12: If VERSION already bumped, skip the bump but still read the version
- Step 17: If already pushed, skip the push command
- Step 19: If PR exists, update the body instead of creating a new PR
Prior execution never exempts verification.
Keep one private Markdown **invocation record** outside the product tree and save
its absolute path. Use these headings so a paused run can resume:
- **Release:** versions, `BUMP_LEVEL`, reviewed tree and attempt counts.
- **Decisions:** each approval's finding, files and authorized action. Reuse it only
for that same scope; a repair never resets approvals or expands them.
- **Reviews:** handles, original start tokens, terminal states, outputs and queued fixes.
- **Checks:** command/label, result/counts, timestamp, log and consumed inputs.
- **Documentation:** candidate/id, attempts used, accepted hashes or named blocked exception.
- **Next steps:** one ordered work list, with the current step marked.
A **receipt** is saved evidence of a check's command, result and consumed content.
A review's **start token** is the opaque value returned by `gstack-review-log --start`
before it reads the diff. Keep `REVIEW_START` for Step 9, a separate `PASS_START` for
each Step 11 attempt, and `DESIGN_START` for design. Finish each pass with its original
token; `--finish` stamps the binding fields automatically. Never borrow or replace a token.
`gstack-wtree` prints a Git tree hash covering tracked and non-ignored untracked files,
not a commit ID. Use `git diff <old-tree> <new-tree>` to compare these snapshots.
### Ship control flow
You, the **parent** running /ship, own advancement; children return evidence, not
permission to proceed. Follow the saved work list:
1. Start with Steps 1–21 in order, including 11.5 and 14.5. Advance only after
the current item's gates clear.
2. Expand a repair into individual steps and insert them before the still-pending
work. This replaces the current item, whose actual result stays in the record.
Add its destination only if not already the next pending step.
3. For another repair, repeat rule 2 without discarding pending work.
The saved list takes precedence over ordinary next-step
sentences inside a repair. A range never adds unlisted steps.
**Example:** Step 11 fixes insert `9 → 10 → 11` before 11.5. A further Step 9 fix
affecting 6–8 makes the list `5 → 6 → 7 → 8 → 9 → 10 → 11 → 11.5`.
The unchanged release steps follow. STOP and AskUserQuestion gates still apply during repairs.
Keep the same attempt counts throughout the invocation. A range ending at Step 14
does not enter Step 14.5. A range that includes Step 14.5 enters its existing audit
decision, not an unconditional new launch; its initial-plus-ONE limit never resets.
Permitted repairs continue in this invocation without restarting /ship.
---
@@ -491,15 +506,18 @@ sections. Read a section in full before doing its step; do not work from memory.
| When | Read this section |
|------|-------------------|
| the ship target is an Apple platform app (.xcodeproj, .xcworkspace, or an app-product Swift package) — read BEFORE Step 1's branch gate and any preflight; store distribution never routes through the branch/PR ceremony | `sections/apple-release.md` |
| App Store/TestFlight distribution is requested for an Apple app (.xcodeproj, .xcworkspace, or an app-product Swift package) — read at Step 0.9 before the branch gate; an Apple repository-landing request follows the normal pipeline | `sections/apple-release.md` |
| running the test suites and (if prompt files changed) the eval suites (Steps 4-6) | `sections/tests.md` |
| auditing test coverage of the diff (Step 7) | `sections/test-coverage.md` |
| auditing plan completion, verification, and scope drift (Step 8) | `sections/plan-completion.md` |
| the pre-landing review and specialist dispatch (Step 9) | `sections/review-army.md` |
| exploratory QA before Fix-First (Step 9.2.1) | Use the QA Read directive in `sections/review-army.md` |
| reusing explicitly skipped shared-code advice (Step 9.3) | `sections/shared-code-reuse.md` |
| addressing Greptile review comments when a PR exists (Step 10) | `sections/greptile.md` |
| the adversarial review and learnings capture (Step 11) | `sections/adversarial.md` |
| writing the CHANGELOG entry (Step 13) | `sections/changelog.md` |
| dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19) | `sections/pr-body.md` |
| auditing docs before final commit/verification (Step 14.5), on every ship | `sections/documentation.md` |
| creating or updating the PR/MR with the verified documentation outcome (Step 19) | `sections/pr-body.md` |
---
@@ -549,8 +567,11 @@ branch name wherever the instructions say "the base branch" or `<default>`.
## Step 0.9: Apple target detection
If the repo has an `.xcodeproj`, `.xcworkspace`, or Swift app package AND the ask
is App Store/TestFlight distribution, **STOP and Read
If the ask is App Store/TestFlight distribution, look for an `.xcodeproj`,
`.xcworkspace`, or Swift app product. Read `Package.swift` and its entrypoint to
distinguish an app from a library/CLI. If unclear, use AskUserQuestion to identify
the target and wait before choosing a release path.
For a confirmed app, **STOP and Read
`~/.claude/skills/gstack/ship/sections/apple-release.md` FIRST**. Store distribution proceeds
through that adapter from the current branch, including a clean base branch.
The branch gate and repository-landing pipeline below apply ONLY to
@@ -558,7 +579,7 @@ repository-landing asks, including on Apple repos.
## Step 1: Pre-flight
1. Check the current branch. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
1. Save the current branch as `<branch-name>`. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
2. Run `git status` (never use `-uall`). Uncommitted changes are always included — no need to ask.
@@ -567,9 +588,8 @@ repository-landing asks, including on Apple repos.
`git diff origin/<base> --stat`, untracked files from status, and
`git log origin/<base>..HEAD --oneline`.
4. Display historical review readiness. This preflight snapshot does not replace
Step 9's mandatory review or its blocker, ASK, and convergence gates — even
when prior reviews are CLEAR or the dashboard's global skip is enabled.
4. Display historical readiness using the dashboard below, then finish Step 1.
Prior CLEAR reviews or dashboard skips never replace Step 9's gates.
## Review Readiness Dashboard
@@ -579,68 +599,95 @@ During pre-flight, read the existing review log and config to display readiness;
~/.claude/skills/gstack/bin/gstack-review-read
```
Render each record using its recorded host, source, outside_provider, outside_status, and phase. Historical source "claude" means a native Claude subagent; source "claude-code" means the external CLI. Never infer a historical provider from the current harness. Unknown model identity remains unknown. Missing/disabled/skipped outside coverage is distinct from native completion.
**1. Choose the records to display.** Use the latest record for each row below.
Do not use a record older than 7 days to clear a row, and never substitute an older
success for a newer failure. Ship metrics are not review records.
Parse the output. Find the most recent entry for each skill (plan-ceo-review, plan-eng-review, review, plan-design-review, design-review-lite, adversarial-review, codex-review, codex-plan-review). Ignore entries with timestamps older than 7 days. For the Eng Review row, show whichever is more recent between `review` (diff-scoped pre-landing review) and `plan-eng-review` (plan-stage architecture review). Append "(DIFF)" or "(PLAN)" to the status to distinguish. For the Adversarial row, show whichever is more recent between `adversarial-review` (new auto-scaled) and `codex-review` (legacy). For Design Review, show whichever is more recent between `plan-design-review` (full visual audit) and `design-review-lite` (code-level check). Append "(FULL)" or "(LITE)" to the status to distinguish. For the Outside Voice row, show the most recent `codex-plan-review` entry — this captures outside voices from both /plan-ceo-review and /plan-eng-review.
| Row | Choose the latest of | Status suffix |
|---|---|---|
| Eng Review | `review` or `plan-eng-review` | (DIFF) or (PLAN) |
| CEO Review | `plan-ceo-review` | — |
| Design Review | `plan-design-review` or `design-review-lite` | (FULL) or (LITE) |
| Adversarial | `adversarial-review` or legacy `codex-review` | — |
| Outside Voice | `codex-plan-review` from CEO or Eng review | — |
**Source attribution:** If the most recent entry for a skill has a \`"via"\` field, append it to the status label in parentheses. Examples: `plan-eng-review` with `via:"autoplan"` shows as "CLEAR (PLAN via /autoplan)". `review` with `via:"ship"` shows as "CLEAR (DIFF via /ship)". Entries without a `via` field show as "CLEAR (PLAN)" or "CLEAR (DIFF)" as before.
Keep each record's host, source, outside_provider, outside_status and phase.
Historical source "claude" is a native subagent; "claude-code" is the external CLI.
Do not infer old providers or unknown models from today's harness. A native result
does not fill missing, disabled or skipped outside coverage.
From gstack-review-read output, use entries whose skill is `autoplan-voices` or `design-outside-voices` for the coverage detail below the dashboard. Group by workflow run and phase, not merely skill. Show each phase’s recorded provider and outside_status; partial coverage must remain partial. These records do not change the engineering gate.
**Source attribution:** Append a recorded `via` to the suffix, for example
"CLEAR (PLAN via /autoplan)" or "CLEAR (DIFF via /ship)". Without `via`, keep
"CLEAR (PLAN)" or "CLEAR (DIFF)". Below the dashboard, group `autoplan-voices`
and `design-outside-voices` by workflow run and phase. Show each phase's provider
and outside_status; retain partial coverage. These details do not clear Eng Review.
Display:
**2. Check freshness before choosing a verdict.**
```
+====================================================================+
| REVIEW READINESS DASHBOARD |
+====================================================================+
| Review | Runs | Last Run | Status | Required |
|-----------------|------|---------------------|-----------|----------|
| Eng Review | 1 | 2026-03-16 15:00 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
| Adversarial | 0 | — | — | no |
| Outside Voice | 0 | — | — | no |
+--------------------------------------------------------------------+
| VERDICT: CLEARED — Eng Review passed |
+====================================================================+
```
- **Content-first rule:** For `review`, `adversarial-review`, `codex-review`,
ship-stage reviews and `design-review-lite`, use `review_freshness.status`
and show its `reason`. CURRENT means a completed clean review whose start and
end content fingerprints equal the current `---WTREE---` fingerprint. This
fingerprint covers working-tree content, not just the commit.
STALE or UNVERIFIED cannot clear Eng Review. Missing `review_freshness`,
including legacy log-only records, means UNVERIFIED. Never fall back to HEAD
equality or commit distance for diff evidence, even at zero commits.
Show recorded cycles, completed/converged fields and missing source/phase
coverage. Unknown coverage is not a pass.
- **Plan records** (plan-ceo-review, plan-eng-review, plan-design-review and
codex-plan-review) use the 7-day window, not the working-tree fingerprint.
If `plan_sha256` is present, you may compare the plan file and report a mismatch.
For plan records only, compare the recorded commit with `---HEAD---`.
If different, run `git rev-list --count STORED_COMMIT..HEAD` and report
"Note: {skill} review from {date} may be stale — {N} commits since review".
A failed command means UNKNOWN, treated as stale. Without commit tracking,
retain the note to consider re-running. Omit staleness notes when all reviews
are current.
**Review tiers:**
- **Eng Review (historical readiness):** Required for a CLEARED dashboard, not for continuing Step 1. Step 9 remains mandatory, with its finding, approval and convergence gates. The skip_eng_review setting changes this dashboard only.
- **CEO Review (optional):** Use your judgment. Recommend it for big product/business changes, new user-facing features, or scope decisions. Skip for bug fixes, refactors, infra, and cleanup.
- **Design Review (optional):** Use your judgment. Recommend it for UI/UX changes. Skip for backend-only, infra, or prompt-only changes.
- **Adversarial Review (automatic):** Always-on for every review. Every diff gets a native adversarial pass and, when enabled and available, a host-selected outside challenge. Large diffs (200+ lines) additionally get a structured outside review with P1 gate.
- **Outside Voice (default-on):** Independent plan review through the host-selected provider after /plan-ceo-review and /plan-eng-review. The codex_reviews switch disables the entire extra step. Provider failure uses the existing native fallback and reports missing outside coverage. Never gates shipping.
**3. Choose the historical verdict.** CLEARED requires the selected Eng Review
to be `clean`, within 7 days and fresh under step 2. Otherwise report NOT CLEARED
and its missing, stale or open-issue reason. If `skip_eng_review` is true, show
"SKIPPED (global)" for Eng Review and CLEARED for this dashboard.
This verdict never skips Step 9 or its finding, approval and convergence gates. Continue Step 1 even when history is NOT CLEARED.
**Verdict logic:**
- **CLEARED**: Eng Review has >= 1 entry within 7 days from either \`review\` or \`plan-eng-review\` with status "clean"; diff review must also grade CURRENT below (or \`skip_eng_review\` is \`true\`)
- **NOT CLEARED**: Eng Review missing, stale (>7 days), or has open issues
- CEO, Design, and outside reviews are shown for context but never block shipping
- If \`skip_eng_review\` config is \`true\`, Eng Review shows "SKIPPED (global)" and verdict is CLEARED
Other rows provide context, not a substitute for Eng Review:
- Recommend CEO Review for product/business or scope decisions, not routine fixes or cleanup.
- Recommend Design Review for UI/UX work, not backend, infrastructure or prompt-only work.
- Adversarial review always includes a native pass. Available, enabled outside
challenges supplement it; diffs of 200+ lines also get the structured P1 gate.
- Outside Voice is the default-on plan review after CEO/Eng review. `codex_reviews`
disables that extra step. Provider failure uses native fallback and records
missing outside coverage; this dashboard row never gates shipping.
**Staleness detection:** Grade before deciding CLEARED:
- Ship telemetry reports metrics, not review coverage; it never satisfies a review row.
- **Content-first rule (diff-scoped rows only: `review`, `adversarial-review`, `codex-review`, ship-stage entries, `design-review-lite`).** Use the helper's computed `review_freshness.status` and show its `reason`. CURRENT requires a completed clean pass with captured start/end wtree equal to the current `---WTREE---`. STALE or UNVERIFIED never clears Eng Review. Missing `review_freshness` is UNVERIFIED, including legacy log-only rows. Never fall back to HEAD equality or commit distance for diff evidence, even at 0 commits. Show recorded cycles, completed/converged state, and missing per-source/phase coverage; unknown is not a pass.
- Plan-tier rows (plan-ceo-review, plan-eng-review, plan-design-review, codex-plan-review) grade a plan file, not the repo tree — never apply the wtree rule to them; they keep the 7-day freshness logic. If an entry carries `plan_sha256`, you MAY compare it with the plan file and note "plan changed since review" on mismatch.
- Plan-tier fallback only: parse `---HEAD---`. For entries with a different `commit`, count elapsed commits: `git rev-list --count STORED_COMMIT..HEAD`. If that command FAILS, grade UNKNOWN and treat as stale. Display: "Note: {skill} review from {date} may be stale — {N} commits since review". Missing commit tracking retains the legacy note to consider re-running.
- If all reviews grade CURRENT, do not display staleness notes
**4. Display the dashboard.** Show missing, stale, disabled or unavailable results
explicitly, never as CLEAR. Display a fresh `clean` result as CLEAR and
`issues_open` as ISSUES OPEN without changing the stored status.
If Eng Review is not CLEAR, print its actual status and reason: "Eng Review: {status} — {reason}. Ship will run its pre-landing review in Step 9." For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend `/plan-eng-review` or `/autoplan` for architecture review.
**REVIEW READINESS DASHBOARD**
If CEO Review is missing, mention as informational ("CEO Review not run — recommended for product changes") but do NOT block.
Use one row for each entry in step 1. Only Eng Review is marked required.
| Review | Runs | Last run | Status | Required |
|---|---:|---|---|---|
| {row and suffix} | {count} | {timestamp or —} | {actual status and reason} | {yes/no} |
VERDICT: {CLEARED or NOT CLEARED} — {reason}
For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend
`/plan-eng-review` or `/autoplan` for architecture review.
For Design Review: run `source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)`. If `SCOPE_FRONTEND=true` and no design review exists, mention: "Design Review not run — Step 9 includes the lite check; consider /design-review for a full visual audit."
Continue to Step 2 without a preflight approval question. Apply the review gates when Step 9 runs.
Continue to Step 2 without asking; Step 9 applies the review gates.
---
## Step 2: Distribution Pipeline Check
If the diff introduces a new standalone artifact (CLI binary, library package, tool) — not a web
service with existing deployment — verify that a distribution pipeline exists.
Check distribution for new standalone artifacts (CLI binaries, packages, tools),
not web services with existing deployment.
1. Check for newly added distribution entry points and package manifests:
1. List candidate distribution paths:
```bash
git diff origin/<base> --diff-filter=A --name-only | grep -E '(^|/)(cmd/[^/]+/main\.go|bin/[^/]+|Cargo\.toml|setup\.py|package\.json)$' | head -5
```
@@ -655,16 +702,17 @@ service with existing deployment — verify that a distribution pipeline exists.
grep -qE 'release|publish|deploy' .gitlab-ci.yml 2>/dev/null && echo "GITLAB_CI_RELEASE"
```
3. **If no release pipeline exists and a new artifact was added:** Use AskUserQuestion:
- "This PR adds a new binary/tool but there's no CI/CD pipeline to build and publish it.
Users won't be able to download the artifact after merge."
- A) Add a release workflow now (CI/CD release pipeline — GitHub Actions or GitLab CI depending on platform)
- B) Defer — add a P1 distribution TODO in Step 14
- C) Not needed — this is internal/web-only, existing deployment covers it
3. **New artifact without a pipeline:** AskUserQuestion: "Users cannot download this
artifact after merge without a release pipeline."
- A) Add the platform's release workflow now
- B) Defer with a P1 distribution TODO in Step 14
- C) Not needed: internal/web-only, covered by existing deployment
4. **If the user chooses A:** Add packaging and publish configuration using this repository's CI conventions. Ask for the intended distribution target if it is unknown; do not invent a registry or credentials. Include the new workflow in the tests and review below. Do not publish a release during `/ship`.
5. **If release pipeline exists:** Continue silently.
6. **If no new artifact detected:** Skip silently.
4. **If A:** Add packaging/publish configuration using repository CI conventions.
Ask for unknown targets, registries or access first; never invent credentials.
Recheck against the artifact and include the workflow in tests and review.
Do not publish a release during `/ship`.
5. Otherwise, continue without adding a pipeline.
---
@@ -676,10 +724,14 @@ Merge the base ref fetched in Step 1 so tests and reviews cover the integrated c
git merge origin/<base> --no-edit
```
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). If conflicts are complex or ambiguous, **STOP** and show them.
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). For complex or ambiguous conflicts, **STOP**, show the conflicting choices, use AskUserQuestion for the needed resolution decision, and wait for the answer before editing or continuing.
**If already up to date:** Continue silently.
If integration changes the artifact or distribution configuration inspected in Step 2,
repeat Step 2 on the merged content, including its decisions, then continue to Step 4.
Otherwise continue to Step 4 directly.
---
> **STOP.** Before running the test suites and (if prompt files changed) the eval suites (Steps 4-6), Read `~/.claude/skills/gstack/ship/sections/tests.md` and execute it
@@ -700,43 +752,92 @@ git merge origin/<base> --no-edit
> **STOP.** Before the adversarial review and learnings capture (Step 11), Read `~/.claude/skills/gstack/ship/sections/adversarial.md` and execute it
> in full. Do not work from memory — that section is the source of truth for this step.
## Step 11.5: Bind the reviews
1. **Select the two reviews.** Run `~/.claude/skills/gstack/bin/gstack-review-read`.
Select this invocation's final Step 9.4 record (`skill:"review"`, `via:"ship"`)
and Step 11 native record (`skill:"adversarial-review"`). Match each to its saved
handle, original token and source; reject outside-provider or older invocation records.
2. **Compare their content.** Require the native record's `review_binding.state`
to be `verified`. All three snapshots must match: its `wtree`, Step 9.4's
`review_binding.start_wtree` and `review_binding.end_wtree`. A mismatch or missing
record/field blocks release preparation: report **Review records missing or mismatched**
and insert `9 → 10 → 11 → 11.5` before Step 12. Bind the new records at 11.5.
Never attach new tokens to old work.
3. **Preserve any QA exception.** A named probe-risk exception may leave Step 9.4's
root `wtree` absent; item 2 still compares its start/end snapshots. Matching content
does not mean the failed or unrun probes passed. Keep Step 9.4's incomplete flags
and the user's exception.
4. **Save the evidence.** Save both records and matching **reviewed tree** for
Step 16. Continue to Step 12.
## Step 12: Version bump (auto-decide)
Use **`gstack-version-bump`** for classify/write/repair and `gstack-next-version`
for slot selection. Bump level and queue collisions remain agent decisions.
Item 3 needs `BUMP_LEVEL`: reuse this invocation's saved level. Otherwise FRESH
chooses it in item 2 and ALREADY_BUMPED derives it in item 1.
1. **Classify state** — pure reader, never writes:
```bash
bun run ~/.claude/skills/gstack/bin/gstack-version-bump classify --base <base>
```
Save the JSON `baseVersion` as `BASE_VERSION`, then read `state` and dispatch:
- **FRESH** → do the bump (steps 2-4).
- **ALREADY_BUMPED** → keep `NEW_VERSION` at `currentVersion`. Reuse this branch's earlier ship decision for `BUMP_LEVEL` if recorded; otherwise compare `baseVersion` and `currentVersion` left to right: the first changed major/minor/patch/micro component supplies `BUMP_LEVEL` (a missing fourth component is zero). Then run step 3's queue check. This recovers the level, not permission to bump again.
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify. On success, follow **ALREADY_BUMPED**, including its queue check; on failure, STOP. Repair alone never re-bumps.
- **DRIFT_UNEXPECTED** → **STOP**. package.json disagrees with VERSION while VERSION matches base — a manual edit bypassed /ship. Reconcile manually, then re-run.
- **FRESH** → use the recorded level or choose it in item 2, then check the queue and write.
- **ALREADY_BUMPED** → keep `NEW_VERSION=currentVersion`. If `BUMP_LEVEL` is missing,
use the first changed component from `baseVersion` to `currentVersion`
(major/minor/patch/micro; an absent fourth component is zero). Continue at item 3,
not another automatic bump.
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify.
Success follows ALREADY_BUMPED, including its queue check; failure stops.
Repair alone never re-bumps.
- **DRIFT_UNEXPECTED** → STOP: package.json disagrees with VERSION while VERSION
matches base. Reconcile the manual edit, then reclassify.
2. **Decide the bump level** from the diff (agent judgment):
- **MICRO**: <50 lines, trivial tweaks/config. **PATCH**: 50+ lines, no feature signals.
- **MINOR**: AskUserQuestion for any feature signal (new route/page, migration, new module), OR 500+ lines. **MAJOR**: AskUserQuestion for milestones or breaking changes. Offer the recommended level with rationale, a smaller level, or cancel; wait for the answer. Cancel ends this ship attempt before release writes or push; preserve existing work.
Save `BUMP_LEVEL` as lowercase `micro`, `patch`, `minor`, or `major`. Queue placement may advance the slot without changing the intended level.
- **MINOR**: ask for any feature signal (new route/page, migration, module) or 500+ lines.
**MAJOR**: ask for milestones or breaking changes. Use AskUserQuestion: recommended
level with rationale, smaller level, or cancel. Wait; cancel stops before release
writes or push and preserves existing work.
Save lowercase `BUMP_LEVEL`. A claimed version may move the next available number
forward, but cannot change the chosen MICRO/PATCH/MINOR/MAJOR level.
3. **Queue-aware pick** (workspace-aware ship):
```bash
QUEUE_JSON=$(bun run ~/.claude/skills/gstack/bin/gstack-next-version --base <base> --bump "$BUMP_LEVEL" --current-version "$BASE_VERSION" 2>/dev/null || echo '{"offline":true}')
CANDIDATE_VERSION=$(echo "$QUEUE_JSON" | jq -r '.version // empty')
```
- **Usable candidate** (including `offline:true` with `fallback:"git"`): print warnings and any claimed queue. FRESH sets `NEW_VERSION` to `CANDIDATE_VERSION`. ALREADY_BUMPED compares it with `currentVersion`; if different, ask to rebump (refresh CHANGELOG/PR title) or keep current (CI rejects a collision). Only approval changes the existing version. An active sibling is a workspace listed in JSON `active_siblings`; use its `branch` and `version`. If one holds `>= NEW_VERSION`, ask to advance past it or stop this attempt and sync.
- **No usable candidate** (utility failure or empty result): print queue-unverified; FRESH sets `NEW_VERSION` using local `BUMP_LEVEL` arithmetic, while ALREADY_BUMPED keeps `currentVersion`. Do not follow the usable-candidate instructions above.
**Qualify first:** require successful utility output and a nonempty valid version.
`offline:false` qualifies; `offline:true` qualifies only with `fallback:"git"`.
Offline output without that fallback, failure, malformed output or an empty version
is unusable, even if it contains a version-looking string.
- **Usable candidate:** print warnings and claimed queue. FRESH sets `NEW_VERSION=CANDIDATE_VERSION`.
ALREADY_BUMPED compares it with `currentVersion`: if different, ask to rebump
(refresh CHANGELOG/PR title) or keep current (CI rejects a collision).
Only approval changes the existing version. Check JSON `active_siblings` by
`branch` and `version`; a sibling holding `>= NEW_VERSION` requires a choice:
advance past it, or stop this attempt and sync.
- **No usable candidate:** print queue-unverified. FRESH uses local `BUMP_LEVEL`
arithmetic; ALREADY_BUMPED keeps `currentVersion`. Never use an empty candidate.
4. **Write the bump** (FRESH, or an approved rebump):
```bash
bun run ~/.claude/skills/gstack/bin/gstack-version-bump write --version "$NEW_VERSION" --regen-digest
```
The CLI validates 4-digit `MAJOR.MINOR.PATCH.MICRO` (or 3-digit pinned semver), then writes VERSION, the manifest, and existing `package-lock.json` / `npm-shrinkwrap.json` files; it never creates lockfiles. Manifest resolution: `--package-json-path` → `.gstack/package-json-path` → `./package.json` (supports subdirectory packages). npm manifests/locks use the 3-digit translation (`1.67.0.0` → `1.67.0`); VERSION remains authoritative. Exit 3 means a half-write: reclassify and use `repair` for DRIFT_STALE_PKG.
The CLI validates `MAJOR.MINOR.PATCH.MICRO` (or pinned 3-digit semver) and writes
VERSION, the manifest and existing `package-lock.json` / `npm-shrinkwrap.json`;
it never creates lockfiles. Manifest path: `--package-json-path` →
`.gstack/package-json-path` → `./package.json`. npm files use the 3-digit translation
(`1.67.0.0` → `1.67.0`); VERSION is authoritative. Exit 3 means a half-write:
reclassify and `repair` DRIFT_STALE_PKG.
`--regen-digest` executes repo code with the same privileges as Step 5: `scripts/gen-agents-digest.ts`, only when it and committed `agents-digest/gstack-AGENTS.md` both exist. Check `agentsDigest`: if false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump before continuing. Its VERSION stamp is freshness-gated.
`--regen-digest` runs repo code with Step 5's privileges: `scripts/gen-agents-digest.ts`,
only when it and committed `agents-digest/gstack-AGENTS.md` exist. If `agentsDigest`
is false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump.
Before push, verify the committed digest matches generation for the selected VERSION.
5. **Record the release decision** (skip if ALREADY_BUMPED):
5. **Record the release decision after a version was actually written**, including
an approved ALREADY_BUMPED rebump. Skip unchanged versions and manifest-only repairs.
```bash
~/.claude/skills/gstack/bin/gstack-decision-log '{"decision":"Ship NEW_VERSION (BUMP_LEVEL)","rationale":"WHY","scope":"repo","source":"skill","confidence":9}' 2>/dev/null || true
```
@@ -747,13 +848,15 @@ for slot selection. Bump level and queue collisions remain agent decisions.
## Step 14: TODOS.md (auto-update)
Persist approved follow-ups, then conservatively mark completed work.
Read `~/.claude/skills/gstack/review/TODOS-format.md`.
Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format reference (or `review/TODOS-format.md` in a gstack checkout).
**1. Open or create:** Read root `TODOS.md`. An explicit "add TODO" choice authorizes
creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: A) Create
a component/priority-organized TODOS.md, B) Skip. Skip goes to item 5.
**1. Open or create:** Read root `TODOS.md`. An earlier explicit "add TODO" choice authorizes its creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: "Create a component/priority-organized TODOS.md?" Options: A) Create now, B) Skip. If B, continue to Step 15 with the outcome in the summary below.
**2. Organization:** Expect component headings, `**Priority:**` P0–P4 fields, and `## Completed` at the bottom. If disorganized, ask: A) Reorganize (recommended), B) Leave as-is. A preserves all content; B continues without restructuring.
**2. Organization:** Use component headings, `**Priority:**` P0–P4 and `## Completed`
at the bottom. If disorganized, ask: A) Reorganize preserving all content
(recommended), B) Leave as-is.
**3. Add approved deferrals:**
- Step 2: add the approved distribution follow-up as P1 with the missing pipeline and affected artifact.
@@ -761,25 +864,40 @@ Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format r
- Step 5: retain P0 test-failure entries already written; deduplicate by failure and source, adding missing approved entries with error output and branch.
Never turn dropped scope into TODOs or invent unapproved follow-ups. Reuse matching existing entries rather than duplicating them.
**4. Detect completed TODOs:** Match titles, files, and behavior against `git diff origin/<base>`, untracked files from status, and `git log origin/<base>..HEAD --oneline`. Only clear evidence earns completion; leave uncertain items open. Move completed items to `## Completed` and append `**Completed:** vX.Y.Z (YYYY-MM-DD)`.
**4. Detect completed TODOs:** Compare titles, files and behavior with
`git diff origin/<base>`, untracked files and `git log origin/<base>..HEAD --oneline`.
Move proven completions to `## Completed` with `**Completed:** vX.Y.Z (YYYY-MM-DD)`;
leave uncertain items open.
**5. Save the summary:** Report added/deferred items, items marked complete, remaining count, and any creation/reorganization. If creation was declined or a write fails, warn and retain the unpersisted follow-ups in the Step 19 PR summary; never claim they were saved. A TODO write failure remains non-blocking.
**5. Save the summary:** Report additions, deferrals, completions, remaining count and
creation/reorganization. If creation was declined or a write failed, warn and retain
unsaved follow-ups in Step 19's PR summary. Never claim they were saved;
TODO write failures are non-blocking.
---
## Step 14.5: Documentation audit (every ship)
**Doc-sync invariant:** Every ship dispatches the /document-release subagent before final
commit/verification/publication, including reruns, already-pushed branches, existing PRs and docs-only changes.
No edits means an executed audit, not a skip; report the section's verified outcome.
> **STOP.** Before auditing docs before final commit/verification (Step 14.5), on every ship, Read `~/.claude/skills/gstack/ship/sections/documentation.md` and execute it
> in full. Do not work from memory — that section is the source of truth for this step.
## Step 15: Commit (bisectable chunks)
Create small, logical commits for `git bisect`. If all changes are already committed, continue to Step 16; never create an empty commit.
Make bisectable commits; if already committed, continue to Step 16. Never create an empty commit.
1. Group by coherent change. Keep each model/service/controller with its tests;
keep controller views together. Migrations may stand alone or accompany their
model; config/routes may accompany the feature they enable. A diff under
50 lines across fewer than 4 files may use one commit.
1. Group changes with their tests, config/routes, views and Step 14.5 docs.
Migrations may stand alone or accompany their model.
Under 50 lines across fewer than 4 files may use one commit.
2. Order dependencies first: infrastructure → models/services → controllers/views.
Each commit must work independently, without broken imports or missing code.
VERSION + CHANGELOG + TODOS.md belong in the final commit.
Group VERSION + CHANGELOG + TODOS.md after the feature commits.
3. Use `<type>: <summary>` (feat/fix/chore/refactor/docs) and a brief body.
Only the final VERSION/CHANGELOG commit gets the version tag and co-author trailer:
Only the final VERSION/CHANGELOG commit gets the release version and co-author
trailer. Do not create a Git tag:
```bash
git commit -m "$(cat <<'EOF'
@@ -796,53 +914,119 @@ EOF
**IRON LAW: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE.**
Find generation/build commands in CLAUDE.md/AGENTS.md, package scripts, and build
configuration; run them first, skipping only when none are defined. A failed build blocks push. If it changes tracked files, inspect the
changes, run affected checks from Steps 6–11, refresh release facts, and commit
under Step 15 before returning here. Reuse unchanged results and actual approvals.
Run stages 1–5 in order. Recovery instructions below name where to resume.
If content changes during or after verification, restart at stage 1 and complete
all five stages before Step 17. Content-preserving commits keep valid evidence.
Then check test evidence against the final content:
### 1. Finish writers and prepare outputs
Inspect writer handles, including the docs child. Confirm terminal completion or termination
before another writer runs. Timeout or cancellation acknowledgment alone means
STOP until confirmed.
Find declared generation/build commands in project instructions, manifests, build
files and CI. Run them and save results. If none exists, record not applicable and
the inspected sources. A missing prerequisite or failed build stops shipping:
report **Build failed or prerequisite missing**, with the command, error and needed
repair. Never invent a substitute command.
**If blocked:** Repair the prerequisite or build, then repeat stage 1. After it passes, continue
to stage 2; treat any content repair as a behavioral change there.
### 2. Choose the change route
Capture the current tree with `~/.claude/skills/gstack/bin/gstack-wtree`. Inspect
`git diff <reviewed-tree> <current-tree>` against the snapshot saved before Step 12.
Missing snapshots block this comparison, regardless of HEAD equality.
Classify the comparison in this order:
1. **Behavior, tests or build inputs changed:** Prompts/templates count as behavior.
Insert `5–11.5 → 12–14 → 16` before the pending Step 17, then stop this step.
This repair excludes Step 14.5 because the rebuild can change generated docs.
Step 16 restarts at stage 1: rebuild and compare again before stage 3 decides
documentation freshness. Further repairs use the same work list.
2. **Only authored docs or release metadata changed:** Keep Step 8's original child
report and counts. Recheck affected plan items using their recorded verification
and append current evidence to the invocation record. If a classification is no
longer supported, run Step 8's audit and decision gates only, then return to
Step 16 stage 1. Never edit the child's counts yourself.
3. **No changes, or the docs-only checks still support the plan:** Continue to stage 3
without a new code review.
### 3. Resolve documentation freshness
Compare the base and hashes of the selected release paths, generated
outputs and docs/templates with Step 14.5's saved values. A prior invocation's
audit or risk decision never qualifies.
| Outcome | Action |
|---|---|
| This invocation's accepted audit matches all inputs | Continue to stage 4. |
| User-accepted named documentation risk covers the same approved scope and exact content, and unwaivable gates clear | Continue to stage 4; retain `Documentation: blocked`, its reason and incomplete scope. |
| Missing, stale or blocked | Use recovery below. Never silently refresh hashes. |
Report changed inputs, blockers and attempts used:
- **An attempt remains, with changed inputs or an available repair:** insert
`14.5 → 15 → 16` before Step 17. Use Blocked recovery with the existing count.
Validate the outcome before Step 15,
then restart Step 16 stage 1 to regenerate and compare again.
- **Otherwise:** STOP unless the user accepts
the specific named documentation risk and all unwaivable gates clear, under
Step 14.5's Blocked recovery rules. Unchanged approved content goes to stage 4;
repaired content goes to stage 1.
Never run a third audit. Child return is not acceptance.
### 4. Verify the frozen candidate
Freeze inputs through verification and push. Run declared docs/link/generated-file
checks; report unavailable checks.
**Reuse a check when its inputs match.** Compare hashes or complete bytes of its
saved and current consumed files, fixtures, dependencies and execution parameters.
Explain why other changes cannot affect it; changed or unknown dependencies require a rerun.
For model judges, compare the complete expanded request, rubric, parameters and
builder/runtime dependencies. Reuse identical passing evidence: cite the original
command, result/counts, timestamp and log, never resample it. Mandatory reviews still run.
**Check each test lane's receipt as well.** Use its actual Step 5 label/command:
`--label <lane> --expect-cmd '<exact Step 5 command>'`. Inspect changes since the run;
`--allow-paths` exempts only release metadata. A `package.json` version-only edit
can qualify; scripts, dependencies and runtime configuration require live tests.
Uncertain edits cannot be exempted. Docs, TODO edits, new/generated tests and fixes
make evidence STALE even without a new code review. Use this example only after
confirming that every allowed edit is release metadata:
```bash
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<exact tests-lane command from Step 5>' --label vitest --expect-cmd '<exact vitest-lane command from Step 5>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<tests>' --label vitest --expect-cmd '<vitest>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
```
Use only Step 5's actual lane labels and exact commands; `vitest` is an example.
If Step 4 explicitly declined testing and no lanes exist, report that gap instead
of inventing FRESH evidence. Build verification still applies.
| Receipt result | Next action |
|---|---|
| FRESH (exit 0) | Cite the label, exit, timestamp and log. |
| STALE/MISSING: changed content, command or age, or no proven run | Run `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`, read the result and recheck once. Handle failures as described below. |
| Only receipt storage/readback failed | Independently prove unchanged final content, the same command and valid age from the successful run's evidence. Cite its exact command, exit, timestamp and log as **ledger unavailable**, never FRESH. Without that proof, use STALE/MISSING. |
The allow-list covers release bookkeeping, including Step 12's package/digest
version stamps. Behavioral package.json edits still require live tests despite
the path exemption. Do not add `TODOS.md` or generated tests to the allow-list:
Step 7 tests, review fixes, and Step 14 TODO edits intentionally make evidence STALE.
No test lanes: require Step 5's explicit untested-scope approval for final content,
or run Steps 5–15, including the no-tests decision, then return to Step 16 stage 1.
Report the gap, never FRESH; builds must pass.
- **Every line FRESH (exit 0):** recorded runs passed on identical content except
the listed release files. Cite label, exit, timestamp, and log path; continue.
- **Any STALE/MISSING (exit non-zero):** inspect the reason before choosing recovery:
- **Content, command or age mismatch, or no passing live evidence:** rerun the
affected lanes on final content, wrapped as `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`.
Read results and recheck once. TODO edits and generated tests are content
changes, not ledger-only bookkeeping.
- **Ledger read/write failure only:** if a successful live run already covers
the unchanged final content, exact command and permitted age, cite its exit,
timestamp and log directly. Report ledger unavailable and continue, never
ledger FRESH. Do not rerun green suites solely because the ledger cannot save
or read its record. If unchanged content cannot be confirmed, STOP.
**New, changed or unwaived test failure:** STOP publication. Run Steps 5–15,
starting with Step 5's triage, then return to Step 16 stage 1. This recovery also
applies if a failure appears while reporting in stage 5. Reentry to Step 14.5
keeps its existing audit count; it does not authorize a third attempt.
A failed CHECK identifies evidence to repair; it is not a test failure. The
required live RUN must pass, except for the explicit triage waiver below.
### 5. Report, then push
Paste build and rerun results. Later code, test, or build-input changes return
through this gate before pushing. Step 18 owns validation of its post-push
docs-only edits; follow repository-required checks there too. Do not claim an
earlier test run covered changed inputs.
Commit only approved, verified release changes left uncommitted after Step 15,
including generated outputs; use its grouping rules and never create an empty commit.
Preserve unrelated user files.
**If tests fail here:** apply Step 5's triage. A prior explicit waiver remains valid
only for the same verified pre-existing failures and approved scope; cite that
approval and actual failing counts, never FRESH or all-green evidence. New,
changed, or unwaived failures STOP publication and return to Step 5.
Claiming work is complete without verification is dishonesty, not efficiency.
Paste build/docs/test results. Reuse waivers only for the same verified
pre-existing failures and approved scope; cite the actual approval and failing
counts, never FRESH or all-green. A new, changed or unwaived test failure uses
stage 4's recovery before publication. Otherwise continue to Step 17.
---
@@ -929,94 +1113,94 @@ If `ALREADY_PUSHED`, skip the push but continue to Step 18. Otherwise push with
git push -u origin <branch-name>
```
**If the push fails, STOP.** Report its error; do not run Steps 18–19 or claim
publication. For a non-fast-forward rejection, fetch and inspect the remote branch,
merge its changes without rewriting history, and return to Step 5 through Step 16
before retrying. Resolve ambiguous conflicts with the user; never force-push.
For authentication, hook, or network failures, fix that cause, rerun affected checks
if content changed, then recheck Step 16 before retrying. Never bypass a failed guard.
**If the push fails, STOP.** No Step 19 or publication claim. Report the error:
- **Non-fast-forward push:** fetch and inspect the remote, then merge under Step 3's
conflict rules. Run Steps 5–16 before returning to Step 17. Never rewrite history.
- **Authentication, hook or network failure:** repair the cause, then repeat Step 16
even if content is unchanged before returning to Step 17. Never bypass failed guards.
Never force-push.
Only a successful push or verified `ALREADY_PUSHED` proceeds.
Continue to mandatory Step 18 (dispatch /document-release), then Step 19 (create/update PR/MR). A push alone does not complete /ship.
Continue to Step 18. No documentation writer runs after push.
---
**PR/MR title invariant (always applies — do not skip even if you don't open the section below):** Any PR or MR you create OR update in the next step MUST have a title that starts with `v$NEW_VERSION` (the version bumped in Step 12), in the format `v<NEW_VERSION> <type>: <summary>`. Never create or edit a PR/MR title without this prefix. Compute the correct title with the single source of truth helper: `~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`. The full create/update procedure (idempotency, redaction scan, self-check) is in the section below.
## Step 18: Prepare publication metadata
**Doc-sync invariant (always applies — do not skip even if you don't open the section below):** Step 18 dispatches the /document-release subagent BEFORE the PR/MR is created or updated in Step 19. Never skip the dispatch itself; only a failed subagent is non-blocking (proceed to Step 19 without a `## Documentation` section).
First look up open PRs/MRs for `<branch-name>` on the detected platform:
> **STOP.** Before dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19), Read `~/.claude/skills/gstack/ship/sections/pr-body.md` and execute it
- GitHub: `gh pr list --head <branch-name> --state open --json number,title,url`
- GitLab: `glab mr list --source-branch <branch-name> --output json` (defaults to open).
A successful empty array means new; one match supplies the existing title/identity.
Lookup failure or ambiguous matches **STOP** for resolution, never mean no PR.
Save the result for Step 19's recheck.
Prepare the title from that result; Step 19 scans and publishes it:
1. For an existing open PR/MR, use the matched title and run
`~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`.
2. For a new PR/MR, compose `v<NEW_VERSION> <type>: <summary>`.
3. Save the result as `NEW_TITLE` for Step 19. Every created or updated title MUST
start with `v$NEW_VERSION `; never publish an unprefixed title.
> **STOP.** Before creating or updating the PR/MR with the verified documentation outcome (Step 19), Read `~/.claude/skills/gstack/ship/sections/pr-body.md` and execute it
> in full. Do not work from memory — that section is the source of truth for this step.
## Step 20: Persist ship metrics
Log coverage and plan completion for `/retro` through `gstack-review-log`.
It resolves the project/branch, validates JSON, creates storage and queues sync.
It takes **no path argument**: hand-built `<branch>-reviews.jsonl` paths break
branches containing `/`.
Log metrics for `/retro` through `gstack-review-log`; it handles project/branch paths,
JSON validation, storage and sync. It takes **no path argument**; do not build one.
```bash
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"ship","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","coverage_pct":COVERAGE_PCT,"plan_items_total":PLAN_TOTAL,"plan_items_done":PLAN_DONE,"verification_result":"VERIFY_RESULT","version":"VERSION","branch":"'"$(git rev-parse --abbrev-ref HEAD)"'"}'
```
Substitute from earlier steps:
- **COVERAGE_PCT**: coverage percentage from Step 7 diagram (integer, or -1 if undetermined)
- **COVERAGE_PCT**: Step 7 diagram's integer percentage; encode null/undetermined as -1
- **PLAN_TOTAL**: total plan items extracted in Step 8 (0 if no plan file)
- **PLAN_DONE**: count of DONE + CHANGED items from Step 8 (0 if no plan file)
- **VERIFY_RESULT**: "pass", "fail", or "skipped" from Step 8.1
- **VERIFY_RESULT**: "pass", "fail", or "skipped", set after Step 9 executes Step 8.1's verification list
- **VERSION**: from the VERSION file
The branch name is filled in by the shell — there is no `BRANCH` placeholder to
substitute.
This step is automatic — never skip it, never ask for confirmation.
The shell supplies the branch. Run this automatically, without confirmation.
---
## Step 21: Plan-tune discoverability nudge (first-successful-ship only)
Plan-tune cathedral T15. After a successful ship, surface /plan-tune once
per machine. Single line, non-blocking, marker-gated so it never re-fires.
After a successful ship, show the non-blocking /plan-tune nudge once per machine:
```bash
_NUDGE_MARKER="$HOME/.gstack/.plan-tune-nudge-shown"
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"
export GSTACK_STATE_ROOT
_NUDGE_MARKER="$GSTACK_STATE_ROOT/.plan-tune-nudge-shown"
_QT=$(~/.claude/skills/gstack/bin/gstack-config get question_tuning 2>/dev/null || echo "false")
if [ ! -f "$_NUDGE_MARKER" ] && [ "$_QT" = "false" ]; then
echo ""
echo "gstack can learn from your AskUserQuestion answers. Run /plan-tune to opt in"
echo "— it captures which prompts you find valuable vs noisy and (with hooks installed)"
echo "auto-decides your never-ask preferences."
touch "$_NUDGE_MARKER"
mkdir -p "$GSTACK_STATE_ROOT" && touch "$_NUDGE_MARKER"
fi
```
If the marker exists, OR question_tuning is already on, the nudge is a
no-op. The marker guarantees at-most-once per machine. To re-enable:
`rm ~/.gstack/.plan-tune-nudge-shown` before next ship.
The marker or enabled question_tuning suppresses it. To re-enable, remove
`$GSTACK_STATE_ROOT/.plan-tune-nudge-shown` before the next ship.
---
## Section self-check (before you finish)
You ran a carved skill. For your situation, list every section the Section index
named as applying, and confirm you issued a Read for each one. If you executed any
of those steps from memory without reading its section, you skipped the source of
truth — STOP, Read it now, and redo that step. Deterministic version work goes
through `gstack-version-bump`; never hand-roll the VERSION/package.json write.
List the applicable Section index entries and confirm each Read. If you worked from
memory, STOP, Read the section and redo that step. Use `gstack-version-bump`, never
hand-roll VERSION/package.json writes.
---
## Important Rules
- **Never skip tests.** If tests fail, stop.
- **Never skip the pre-landing review.** If checklist.md is unreadable, stop.
Follow the numbered gates and their explicit exceptions.
- **Never force push.** Use regular `git push` only.
- **Never ask for trivial confirmations** (e.g., "ready to push?", "create PR?"). DO stop for: version bumps (MINOR/MAJOR), pre-landing review findings (ASK items), and Codex structured review [P1] findings (large diffs only).
- **Always use the 4-digit version format** from the VERSION file.
- **Date format in CHANGELOG:** `YYYY-MM-DD`
- **Split commits for bisectability** — each commit = one logical change.
- **TODOS.md completion detection must be conservative.** Only mark items as completed when the diff clearly shows the work is done.
- **Use Greptile reply templates from greptile-triage.md.** Every reply includes evidence (inline diff, code references, re-rank suggestion). Never post vague replies.
- **Never push without fresh verification evidence.** If code changed after Step 5 tests, re-run before pushing.
- **Step 7 generates coverage tests.** They must pass before committing. Never commit failing tests.
- **The goal is: user says `/ship`, next thing they see is the review + PR URL + auto-synced docs.**