v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+316 -164
View File
@@ -32,47 +32,62 @@ triggers:
# Ship: Fully Automated Ship Workflow
Run `/ship` through to the PR URL. This request authorizes routine work without confirmation; explicit safety and user-decision gates still apply.
STOP blocks advancement until the stated repair/resume route clears; without one, end this attempt.
Answer each AskUserQuestion before continuing.
Routine authorization never waives those gates or their required user decisions.
**Route through the workflow:** detect and merge the base (Steps 1–3), test and
audit the integrated diff (Steps 4–8.2), review and resolve findings (Steps
9–11), prepare the release and commits (Steps 12–15), then verify, push, sync
docs, and open or update the PR (Steps 16–19). A review fix returns to affected
tests and reviews before release preparation; a later code or build-input edit
returns to affected checks and Step 16 before publication. Reuse still-valid
results, but never treat an earlier review or test as covering changed inputs.
**Routine work needs no confirmation:** include uncommitted changes, choose MICRO/PATCH
under Step 12, draft CHANGELOG and commits, mark completed TODOs and auto-fix findings.
When Step 7 coverage meets its target, report remaining gaps and verify generated
tests without another permission question. Step 15 commits those tests.
**Follow every STOP and AskUserQuestion gate**, including:
- On the base branch (abort)
- Merge conflicts that can't be auto-resolved (stop, show conflicts)
- In-branch test failures (pre-existing failures are triaged, not auto-blocking)
- Pre-landing review finds ASK items that need user judgment
- Prior Learnings needs its first-time cross-project setting (Step 8)
- MINOR or MAJOR version bump needed (ask — see Step 12)
- Greptile review comments that need user decision (complex fixes, false positives)
- AI-assessed coverage below target (see Step 7 for minimum/target decisions)
- Plan items NOT DONE or UNVERIFIABLE (see Step 8)
- Plan verification failures (see Step 8.1)
- TODOS.md missing and user wants to create one (ask — see Step 14)
- TODOS.md disorganized and user wants to reorganize (ask — see Step 14)
**Route:** integrate (1–3) → test and review (4–11.5) → prepare the release
(12–15) → verify frozen content (16) → push and publish (17–21).
Every new invocation repeats Steps 1–16, including both reviews and the docs audit.
Steps 12, 17 and 19 prevent duplicate bumps, pushes and PRs, never verification.
**Never stop for:**
- Uncommitted changes (always include them)
- Version bump choice (auto-pick MICRO or PATCH — see Step 12)
- CHANGELOG content (auto-generate from diff)
- Commit message approval (auto-commit)
- Multi-file changesets (auto-split into bisectable commits)
- TODOS.md completed-item detection (auto-mark)
- Auto-fixable review findings (dead code, N+1, stale comments — fixed automatically)
- Test coverage gaps within target threshold (generate, verify, then commit with Step 15; flag any remaining gaps in the PR body)
### Keep state between steps
**Re-run behavior (idempotency):**
Every invocation repeats verification: tests, coverage, plan completion, both
reviews, VERSION/CHANGELOG, TODOS and doc-sync. Only *actions* are idempotent:
- Step 12: If VERSION already bumped, skip the bump but still read the version
- Step 17: If already pushed, skip the push command
- Step 19: If PR exists, update the body instead of creating a new PR
Prior execution never exempts verification.
Keep one private Markdown **invocation record** outside the product tree and save
its absolute path. Use these headings so a paused run can resume:
- **Release:** versions, `BUMP_LEVEL`, reviewed tree and attempt counts.
- **Decisions:** each approval's finding, files and authorized action. Reuse it only
for that same scope; a repair never resets approvals or expands them.
- **Reviews:** handles, original start tokens, terminal states, outputs and queued fixes.
- **Checks:** command/label, result/counts, timestamp, log and consumed inputs.
- **Documentation:** candidate/id, attempts used, accepted hashes or named blocked exception.
- **Next steps:** one ordered work list, with the current step marked.
A **receipt** is saved evidence of a check's command, result and consumed content.
A review's **start token** is the opaque value returned by `gstack-review-log --start`
before it reads the diff. Keep `REVIEW_START` for Step 9, a separate `PASS_START` for
each Step 11 attempt, and `DESIGN_START` for design. Finish each pass with its original
token; `--finish` stamps the binding fields automatically. Never borrow or replace a token.
`gstack-wtree` prints a Git tree hash covering tracked and non-ignored untracked files,
not a commit ID. Use `git diff <old-tree> <new-tree>` to compare these snapshots.
### Ship control flow
You, the **parent** running /ship, own advancement; children return evidence, not
permission to proceed. Follow the saved work list:
1. Start with Steps 1–21 in order, including 11.5 and 14.5. Advance only after
the current item's gates clear.
2. Expand a repair into individual steps and insert them before the still-pending
work. This replaces the current item, whose actual result stays in the record.
Add its destination only if not already the next pending step.
3. For another repair, repeat rule 2 without discarding pending work.
The saved list takes precedence over ordinary next-step
sentences inside a repair. A range never adds unlisted steps.
**Example:** Step 11 fixes insert `9 → 10 → 11` before 11.5. A further Step 9 fix
affecting 6–8 makes the list `5 → 6 → 7 → 8 → 9 → 10 → 11 → 11.5`.
The unchanged release steps follow. STOP and AskUserQuestion gates still apply during repairs.
Keep the same attempt counts throughout the invocation. A range ending at Step 14
does not enter Step 14.5. A range that includes Step 14.5 enters its existing audit
decision, not an unconditional new launch; its initial-plus-ONE limit never resets.
Permitted repairs continue in this invocation without restarting /ship.
---
@@ -89,8 +104,11 @@ Prior execution never exempts verification.
## Step 0.9: Apple target detection
If the repo has an `.xcodeproj`, `.xcworkspace`, or Swift app package AND the ask
is App Store/TestFlight distribution, **STOP and Read
If the ask is App Store/TestFlight distribution, look for an `.xcodeproj`,
`.xcworkspace`, or Swift app product. Read `Package.swift` and its entrypoint to
distinguish an app from a library/CLI. If unclear, use AskUserQuestion to identify
the target and wait before choosing a release path.
For a confirmed app, **STOP and Read
`~/.claude/skills/gstack/ship/sections/apple-release.md` FIRST**. Store distribution proceeds
through that adapter from the current branch, including a clean base branch.
The branch gate and repository-landing pipeline below apply ONLY to
@@ -98,7 +116,7 @@ repository-landing asks, including on Apple repos.
## Step 1: Pre-flight
1. Check the current branch. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
1. Save the current branch as `<branch-name>`. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
2. Run `git status` (never use `-uall`). Uncommitted changes are always included — no need to ask.
@@ -107,28 +125,26 @@ repository-landing asks, including on Apple repos.
`git diff origin/<base> --stat`, untracked files from status, and
`git log origin/<base>..HEAD --oneline`.
4. Display historical review readiness. This preflight snapshot does not replace
Step 9's mandatory review or its blocker, ASK, and convergence gates — even
when prior reviews are CLEAR or the dashboard's global skip is enabled.
4. Display historical readiness using the dashboard below, then finish Step 1.
Prior CLEAR reviews or dashboard skips never replace Step 9's gates.
{{REVIEW_DASHBOARD}}
If Eng Review is not CLEAR, print its actual status and reason: "Eng Review: {status} — {reason}. Ship will run its pre-landing review in Step 9." For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend `/plan-eng-review` or `/autoplan` for architecture review.
If CEO Review is missing, mention as informational ("CEO Review not run — recommended for product changes") but do NOT block.
For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend
`/plan-eng-review` or `/autoplan` for architecture review.
For Design Review: run `source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)`. If `SCOPE_FRONTEND=true` and no design review exists, mention: "Design Review not run — Step 9 includes the lite check; consider /design-review for a full visual audit."
Continue to Step 2 without a preflight approval question. Apply the review gates when Step 9 runs.
Continue to Step 2 without asking; Step 9 applies the review gates.
---
## Step 2: Distribution Pipeline Check
If the diff introduces a new standalone artifact (CLI binary, library package, tool) — not a web
service with existing deployment — verify that a distribution pipeline exists.
Check distribution for new standalone artifacts (CLI binaries, packages, tools),
not web services with existing deployment.
1. Check for newly added distribution entry points and package manifests:
1. List candidate distribution paths:
```bash
git diff origin/<base> --diff-filter=A --name-only | grep -E '(^|/)(cmd/[^/]+/main\.go|bin/[^/]+|Cargo\.toml|setup\.py|package\.json)$' | head -5
```
@@ -143,16 +159,17 @@ service with existing deployment — verify that a distribution pipeline exists.
grep -qE 'release|publish|deploy' .gitlab-ci.yml 2>/dev/null && echo "GITLAB_CI_RELEASE"
```
3. **If no release pipeline exists and a new artifact was added:** Use AskUserQuestion:
- "This PR adds a new binary/tool but there's no CI/CD pipeline to build and publish it.
Users won't be able to download the artifact after merge."
- A) Add a release workflow now (CI/CD release pipeline — GitHub Actions or GitLab CI depending on platform)
- B) Defer — add a P1 distribution TODO in Step 14
- C) Not needed — this is internal/web-only, existing deployment covers it
3. **New artifact without a pipeline:** AskUserQuestion: "Users cannot download this
artifact after merge without a release pipeline."
- A) Add the platform's release workflow now
- B) Defer with a P1 distribution TODO in Step 14
- C) Not needed: internal/web-only, covered by existing deployment
4. **If the user chooses A:** Add packaging and publish configuration using this repository's CI conventions. Ask for the intended distribution target if it is unknown; do not invent a registry or credentials. Include the new workflow in the tests and review below. Do not publish a release during `/ship`.
5. **If release pipeline exists:** Continue silently.
6. **If no new artifact detected:** Skip silently.
4. **If A:** Add packaging/publish configuration using repository CI conventions.
Ask for unknown targets, registries or access first; never invent credentials.
Recheck against the artifact and include the workflow in tests and review.
Do not publish a release during `/ship`.
5. Otherwise, continue without adding a pipeline.
---
@@ -164,10 +181,14 @@ Merge the base ref fetched in Step 1 so tests and reviews cover the integrated c
git merge origin/<base> --no-edit
```
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). If conflicts are complex or ambiguous, **STOP** and show them.
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). For complex or ambiguous conflicts, **STOP**, show the conflicting choices, use AskUserQuestion for the needed resolution decision, and wait for the answer before editing or continuing.
**If already up to date:** Continue silently.
If integration changes the artifact or distribution configuration inspected in Step 2,
repeat Step 2 on the merged content, including its decisions, then continue to Step 4.
Otherwise continue to Step 4 directly.
---
{{SECTION:tests}}
@@ -182,43 +203,92 @@ git merge origin/<base> --no-edit
{{SECTION:adversarial}}
## Step 11.5: Bind the reviews
1. **Select the two reviews.** Run `~/.claude/skills/gstack/bin/gstack-review-read`.
Select this invocation's final Step 9.4 record (`skill:"review"`, `via:"ship"`)
and Step 11 native record (`skill:"adversarial-review"`). Match each to its saved
handle, original token and source; reject outside-provider or older invocation records.
2. **Compare their content.** Require the native record's `review_binding.state`
to be `verified`. All three snapshots must match: its `wtree`, Step 9.4's
`review_binding.start_wtree` and `review_binding.end_wtree`. A mismatch or missing
record/field blocks release preparation: report **Review records missing or mismatched**
and insert `9 → 10 → 11 → 11.5` before Step 12. Bind the new records at 11.5.
Never attach new tokens to old work.
3. **Preserve any QA exception.** A named probe-risk exception may leave Step 9.4's
root `wtree` absent; item 2 still compares its start/end snapshots. Matching content
does not mean the failed or unrun probes passed. Keep Step 9.4's incomplete flags
and the user's exception.
4. **Save the evidence.** Save both records and matching **reviewed tree** for
Step 16. Continue to Step 12.
## Step 12: Version bump (auto-decide)
Use **`gstack-version-bump`** for classify/write/repair and `gstack-next-version`
for slot selection. Bump level and queue collisions remain agent decisions.
Item 3 needs `BUMP_LEVEL`: reuse this invocation's saved level. Otherwise FRESH
chooses it in item 2 and ALREADY_BUMPED derives it in item 1.
1. **Classify state** — pure reader, never writes:
```bash
bun run ~/.claude/skills/gstack/bin/gstack-version-bump classify --base <base>
```
Save the JSON `baseVersion` as `BASE_VERSION`, then read `state` and dispatch:
- **FRESH** → do the bump (steps 2-4).
- **ALREADY_BUMPED** → keep `NEW_VERSION` at `currentVersion`. Reuse this branch's earlier ship decision for `BUMP_LEVEL` if recorded; otherwise compare `baseVersion` and `currentVersion` left to right: the first changed major/minor/patch/micro component supplies `BUMP_LEVEL` (a missing fourth component is zero). Then run step 3's queue check. This recovers the level, not permission to bump again.
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify. On success, follow **ALREADY_BUMPED**, including its queue check; on failure, STOP. Repair alone never re-bumps.
- **DRIFT_UNEXPECTED** → **STOP**. package.json disagrees with VERSION while VERSION matches base — a manual edit bypassed /ship. Reconcile manually, then re-run.
- **FRESH** → use the recorded level or choose it in item 2, then check the queue and write.
- **ALREADY_BUMPED** → keep `NEW_VERSION=currentVersion`. If `BUMP_LEVEL` is missing,
use the first changed component from `baseVersion` to `currentVersion`
(major/minor/patch/micro; an absent fourth component is zero). Continue at item 3,
not another automatic bump.
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify.
Success follows ALREADY_BUMPED, including its queue check; failure stops.
Repair alone never re-bumps.
- **DRIFT_UNEXPECTED** → STOP: package.json disagrees with VERSION while VERSION
matches base. Reconcile the manual edit, then reclassify.
2. **Decide the bump level** from the diff (agent judgment):
- **MICRO**: <50 lines, trivial tweaks/config. **PATCH**: 50+ lines, no feature signals.
- **MINOR**: AskUserQuestion for any feature signal (new route/page, migration, new module), OR 500+ lines. **MAJOR**: AskUserQuestion for milestones or breaking changes. Offer the recommended level with rationale, a smaller level, or cancel; wait for the answer. Cancel ends this ship attempt before release writes or push; preserve existing work.
Save `BUMP_LEVEL` as lowercase `micro`, `patch`, `minor`, or `major`. Queue placement may advance the slot without changing the intended level.
- **MINOR**: ask for any feature signal (new route/page, migration, module) or 500+ lines.
**MAJOR**: ask for milestones or breaking changes. Use AskUserQuestion: recommended
level with rationale, smaller level, or cancel. Wait; cancel stops before release
writes or push and preserves existing work.
Save lowercase `BUMP_LEVEL`. A claimed version may move the next available number
forward, but cannot change the chosen MICRO/PATCH/MINOR/MAJOR level.
3. **Queue-aware pick** (workspace-aware ship):
```bash
QUEUE_JSON=$(bun run ~/.claude/skills/gstack/bin/gstack-next-version --base <base> --bump "$BUMP_LEVEL" --current-version "$BASE_VERSION" 2>/dev/null || echo '{"offline":true}')
CANDIDATE_VERSION=$(echo "$QUEUE_JSON" | jq -r '.version // empty')
```
- **Usable candidate** (including `offline:true` with `fallback:"git"`): print warnings and any claimed queue. FRESH sets `NEW_VERSION` to `CANDIDATE_VERSION`. ALREADY_BUMPED compares it with `currentVersion`; if different, ask to rebump (refresh CHANGELOG/PR title) or keep current (CI rejects a collision). Only approval changes the existing version. An active sibling is a workspace listed in JSON `active_siblings`; use its `branch` and `version`. If one holds `>= NEW_VERSION`, ask to advance past it or stop this attempt and sync.
- **No usable candidate** (utility failure or empty result): print queue-unverified; FRESH sets `NEW_VERSION` using local `BUMP_LEVEL` arithmetic, while ALREADY_BUMPED keeps `currentVersion`. Do not follow the usable-candidate instructions above.
**Qualify first:** require successful utility output and a nonempty valid version.
`offline:false` qualifies; `offline:true` qualifies only with `fallback:"git"`.
Offline output without that fallback, failure, malformed output or an empty version
is unusable, even if it contains a version-looking string.
- **Usable candidate:** print warnings and claimed queue. FRESH sets `NEW_VERSION=CANDIDATE_VERSION`.
ALREADY_BUMPED compares it with `currentVersion`: if different, ask to rebump
(refresh CHANGELOG/PR title) or keep current (CI rejects a collision).
Only approval changes the existing version. Check JSON `active_siblings` by
`branch` and `version`; a sibling holding `>= NEW_VERSION` requires a choice:
advance past it, or stop this attempt and sync.
- **No usable candidate:** print queue-unverified. FRESH uses local `BUMP_LEVEL`
arithmetic; ALREADY_BUMPED keeps `currentVersion`. Never use an empty candidate.
4. **Write the bump** (FRESH, or an approved rebump):
```bash
bun run ~/.claude/skills/gstack/bin/gstack-version-bump write --version "$NEW_VERSION" --regen-digest
```
The CLI validates 4-digit `MAJOR.MINOR.PATCH.MICRO` (or 3-digit pinned semver), then writes VERSION, the manifest, and existing `package-lock.json` / `npm-shrinkwrap.json` files; it never creates lockfiles. Manifest resolution: `--package-json-path` → `.gstack/package-json-path` → `./package.json` (supports subdirectory packages). npm manifests/locks use the 3-digit translation (`1.67.0.0` → `1.67.0`); VERSION remains authoritative. Exit 3 means a half-write: reclassify and use `repair` for DRIFT_STALE_PKG.
The CLI validates `MAJOR.MINOR.PATCH.MICRO` (or pinned 3-digit semver) and writes
VERSION, the manifest and existing `package-lock.json` / `npm-shrinkwrap.json`;
it never creates lockfiles. Manifest path: `--package-json-path` →
`.gstack/package-json-path` → `./package.json`. npm files use the 3-digit translation
(`1.67.0.0` → `1.67.0`); VERSION is authoritative. Exit 3 means a half-write:
reclassify and `repair` DRIFT_STALE_PKG.
`--regen-digest` executes repo code with the same privileges as Step 5: `scripts/gen-agents-digest.ts`, only when it and committed `agents-digest/gstack-AGENTS.md` both exist. Check `agentsDigest`: if false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump before continuing. Its VERSION stamp is freshness-gated.
`--regen-digest` runs repo code with Step 5's privileges: `scripts/gen-agents-digest.ts`,
only when it and committed `agents-digest/gstack-AGENTS.md` exist. If `agentsDigest`
is false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump.
Before push, verify the committed digest matches generation for the selected VERSION.
5. **Record the release decision** (skip if ALREADY_BUMPED):
5. **Record the release decision after a version was actually written**, including
an approved ALREADY_BUMPED rebump. Skip unchanged versions and manifest-only repairs.
```bash
~/.claude/skills/gstack/bin/gstack-decision-log '{"decision":"Ship NEW_VERSION (BUMP_LEVEL)","rationale":"WHY","scope":"repo","source":"skill","confidence":9}' 2>/dev/null || true
```
@@ -228,13 +298,15 @@ for slot selection. Bump level and queue collisions remain agent decisions.
## Step 14: TODOS.md (auto-update)
Persist approved follow-ups, then conservatively mark completed work.
Read `~/.claude/skills/gstack/review/TODOS-format.md`.
Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format reference (or `review/TODOS-format.md` in a gstack checkout).
**1. Open or create:** Read root `TODOS.md`. An explicit "add TODO" choice authorizes
creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: A) Create
a component/priority-organized TODOS.md, B) Skip. Skip goes to item 5.
**1. Open or create:** Read root `TODOS.md`. An earlier explicit "add TODO" choice authorizes its creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: "Create a component/priority-organized TODOS.md?" Options: A) Create now, B) Skip. If B, continue to Step 15 with the outcome in the summary below.
**2. Organization:** Expect component headings, `**Priority:**` P0–P4 fields, and `## Completed` at the bottom. If disorganized, ask: A) Reorganize (recommended), B) Leave as-is. A preserves all content; B continues without restructuring.
**2. Organization:** Use component headings, `**Priority:**` P0–P4 and `## Completed`
at the bottom. If disorganized, ask: A) Reorganize preserving all content
(recommended), B) Leave as-is.
**3. Add approved deferrals:**
- Step 2: add the approved distribution follow-up as P1 with the missing pipeline and affected artifact.
@@ -242,25 +314,39 @@ Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format r
- Step 5: retain P0 test-failure entries already written; deduplicate by failure and source, adding missing approved entries with error output and branch.
Never turn dropped scope into TODOs or invent unapproved follow-ups. Reuse matching existing entries rather than duplicating them.
**4. Detect completed TODOs:** Match titles, files, and behavior against `git diff origin/<base>`, untracked files from status, and `git log origin/<base>..HEAD --oneline`. Only clear evidence earns completion; leave uncertain items open. Move completed items to `## Completed` and append `**Completed:** vX.Y.Z (YYYY-MM-DD)`.
**4. Detect completed TODOs:** Compare titles, files and behavior with
`git diff origin/<base>`, untracked files and `git log origin/<base>..HEAD --oneline`.
Move proven completions to `## Completed` with `**Completed:** vX.Y.Z (YYYY-MM-DD)`;
leave uncertain items open.
**5. Save the summary:** Report added/deferred items, items marked complete, remaining count, and any creation/reorganization. If creation was declined or a write fails, warn and retain the unpersisted follow-ups in the Step 19 PR summary; never claim they were saved. A TODO write failure remains non-blocking.
**5. Save the summary:** Report additions, deferrals, completions, remaining count and
creation/reorganization. If creation was declined or a write failed, warn and retain
unsaved follow-ups in Step 19's PR summary. Never claim they were saved;
TODO write failures are non-blocking.
---
## Step 14.5: Documentation audit (every ship)
**Doc-sync invariant:** Every ship dispatches the /document-release subagent before final
commit/verification/publication, including reruns, already-pushed branches, existing PRs and docs-only changes.
No edits means an executed audit, not a skip; report the section's verified outcome.
{{SECTION:documentation}}
## Step 15: Commit (bisectable chunks)
Create small, logical commits for `git bisect`. If all changes are already committed, continue to Step 16; never create an empty commit.
Make bisectable commits; if already committed, continue to Step 16. Never create an empty commit.
1. Group by coherent change. Keep each model/service/controller with its tests;
keep controller views together. Migrations may stand alone or accompany their
model; config/routes may accompany the feature they enable. A diff under
50 lines across fewer than 4 files may use one commit.
1. Group changes with their tests, config/routes, views and Step 14.5 docs.
Migrations may stand alone or accompany their model.
Under 50 lines across fewer than 4 files may use one commit.
2. Order dependencies first: infrastructure → models/services → controllers/views.
Each commit must work independently, without broken imports or missing code.
VERSION + CHANGELOG + TODOS.md belong in the final commit.
Group VERSION + CHANGELOG + TODOS.md after the feature commits.
3. Use `<type>: <summary>` (feat/fix/chore/refactor/docs) and a brief body.
Only the final VERSION/CHANGELOG commit gets the version tag and co-author trailer:
Only the final VERSION/CHANGELOG commit gets the release version and co-author
trailer. Do not create a Git tag:
```bash
git commit -m "$(cat <<'EOF'
@@ -277,53 +363,119 @@ EOF
**IRON LAW: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE.**
Find generation/build commands in CLAUDE.md/AGENTS.md, package scripts, and build
configuration; run them first, skipping only when none are defined. A failed build blocks push. If it changes tracked files, inspect the
changes, run affected checks from Steps 6–11, refresh release facts, and commit
under Step 15 before returning here. Reuse unchanged results and actual approvals.
Run stages 1–5 in order. Recovery instructions below name where to resume.
If content changes during or after verification, restart at stage 1 and complete
all five stages before Step 17. Content-preserving commits keep valid evidence.
Then check test evidence against the final content:
### 1. Finish writers and prepare outputs
Inspect writer handles, including the docs child. Confirm terminal completion or termination
before another writer runs. Timeout or cancellation acknowledgment alone means
STOP until confirmed.
Find declared generation/build commands in project instructions, manifests, build
files and CI. Run them and save results. If none exists, record not applicable and
the inspected sources. A missing prerequisite or failed build stops shipping:
report **Build failed or prerequisite missing**, with the command, error and needed
repair. Never invent a substitute command.
**If blocked:** Repair the prerequisite or build, then repeat stage 1. After it passes, continue
to stage 2; treat any content repair as a behavioral change there.
### 2. Choose the change route
Capture the current tree with `~/.claude/skills/gstack/bin/gstack-wtree`. Inspect
`git diff <reviewed-tree> <current-tree>` against the snapshot saved before Step 12.
Missing snapshots block this comparison, regardless of HEAD equality.
Classify the comparison in this order:
1. **Behavior, tests or build inputs changed:** Prompts/templates count as behavior.
Insert `5–11.5 → 12–14 → 16` before the pending Step 17, then stop this step.
This repair excludes Step 14.5 because the rebuild can change generated docs.
Step 16 restarts at stage 1: rebuild and compare again before stage 3 decides
documentation freshness. Further repairs use the same work list.
2. **Only authored docs or release metadata changed:** Keep Step 8's original child
report and counts. Recheck affected plan items using their recorded verification
and append current evidence to the invocation record. If a classification is no
longer supported, run Step 8's audit and decision gates only, then return to
Step 16 stage 1. Never edit the child's counts yourself.
3. **No changes, or the docs-only checks still support the plan:** Continue to stage 3
without a new code review.
### 3. Resolve documentation freshness
Compare the base and hashes of the selected release paths, generated
outputs and docs/templates with Step 14.5's saved values. A prior invocation's
audit or risk decision never qualifies.
| Outcome | Action |
|---|---|
| This invocation's accepted audit matches all inputs | Continue to stage 4. |
| User-accepted named documentation risk covers the same approved scope and exact content, and unwaivable gates clear | Continue to stage 4; retain `Documentation: blocked`, its reason and incomplete scope. |
| Missing, stale or blocked | Use recovery below. Never silently refresh hashes. |
Report changed inputs, blockers and attempts used:
- **An attempt remains, with changed inputs or an available repair:** insert
`14.5 → 15 → 16` before Step 17. Use Blocked recovery with the existing count.
Validate the outcome before Step 15,
then restart Step 16 stage 1 to regenerate and compare again.
- **Otherwise:** STOP unless the user accepts
the specific named documentation risk and all unwaivable gates clear, under
Step 14.5's Blocked recovery rules. Unchanged approved content goes to stage 4;
repaired content goes to stage 1.
Never run a third audit. Child return is not acceptance.
### 4. Verify the frozen candidate
Freeze inputs through verification and push. Run declared docs/link/generated-file
checks; report unavailable checks.
**Reuse a check when its inputs match.** Compare hashes or complete bytes of its
saved and current consumed files, fixtures, dependencies and execution parameters.
Explain why other changes cannot affect it; changed or unknown dependencies require a rerun.
For model judges, compare the complete expanded request, rubric, parameters and
builder/runtime dependencies. Reuse identical passing evidence: cite the original
command, result/counts, timestamp and log, never resample it. Mandatory reviews still run.
**Check each test lane's receipt as well.** Use its actual Step 5 label/command:
`--label <lane> --expect-cmd '<exact Step 5 command>'`. Inspect changes since the run;
`--allow-paths` exempts only release metadata. A `package.json` version-only edit
can qualify; scripts, dependencies and runtime configuration require live tests.
Uncertain edits cannot be exempted. Docs, TODO edits, new/generated tests and fixes
make evidence STALE even without a new code review. Use this example only after
confirming that every allowed edit is release metadata:
```bash
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<exact tests-lane command from Step 5>' --label vitest --expect-cmd '<exact vitest-lane command from Step 5>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<tests>' --label vitest --expect-cmd '<vitest>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
```
Use only Step 5's actual lane labels and exact commands; `vitest` is an example.
If Step 4 explicitly declined testing and no lanes exist, report that gap instead
of inventing FRESH evidence. Build verification still applies.
| Receipt result | Next action |
|---|---|
| FRESH (exit 0) | Cite the label, exit, timestamp and log. |
| STALE/MISSING: changed content, command or age, or no proven run | Run `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`, read the result and recheck once. Handle failures as described below. |
| Only receipt storage/readback failed | Independently prove unchanged final content, the same command and valid age from the successful run's evidence. Cite its exact command, exit, timestamp and log as **ledger unavailable**, never FRESH. Without that proof, use STALE/MISSING. |
The allow-list covers release bookkeeping, including Step 12's package/digest
version stamps. Behavioral package.json edits still require live tests despite
the path exemption. Do not add `TODOS.md` or generated tests to the allow-list:
Step 7 tests, review fixes, and Step 14 TODO edits intentionally make evidence STALE.
No test lanes: require Step 5's explicit untested-scope approval for final content,
or run Steps 5–15, including the no-tests decision, then return to Step 16 stage 1.
Report the gap, never FRESH; builds must pass.
- **Every line FRESH (exit 0):** recorded runs passed on identical content except
the listed release files. Cite label, exit, timestamp, and log path; continue.
- **Any STALE/MISSING (exit non-zero):** inspect the reason before choosing recovery:
- **Content, command or age mismatch, or no passing live evidence:** rerun the
affected lanes on final content, wrapped as `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`.
Read results and recheck once. TODO edits and generated tests are content
changes, not ledger-only bookkeeping.
- **Ledger read/write failure only:** if a successful live run already covers
the unchanged final content, exact command and permitted age, cite its exit,
timestamp and log directly. Report ledger unavailable and continue, never
ledger FRESH. Do not rerun green suites solely because the ledger cannot save
or read its record. If unchanged content cannot be confirmed, STOP.
**New, changed or unwaived test failure:** STOP publication. Run Steps 5–15,
starting with Step 5's triage, then return to Step 16 stage 1. This recovery also
applies if a failure appears while reporting in stage 5. Reentry to Step 14.5
keeps its existing audit count; it does not authorize a third attempt.
A failed CHECK identifies evidence to repair; it is not a test failure. The
required live RUN must pass, except for the explicit triage waiver below.
### 5. Report, then push
Paste build and rerun results. Later code, test, or build-input changes return
through this gate before pushing. Step 18 owns validation of its post-push
docs-only edits; follow repository-required checks there too. Do not claim an
earlier test run covered changed inputs.
Commit only approved, verified release changes left uncommitted after Step 15,
including generated outputs; use its grouping rules and never create an empty commit.
Preserve unrelated user files.
**If tests fail here:** apply Step 5's triage. A prior explicit waiver remains valid
only for the same verified pre-existing failures and approved scope; cite that
approval and actual failing counts, never FRESH or all-green evidence. New,
changed, or unwaived failures STOP publication and return to Step 5.
Claiming work is complete without verification is dishonesty, not efficiency.
Paste build/docs/test results. Reuse waivers only for the same verified
pre-existing failures and approved scope; cite the actual approval and failing
counts, never FRESH or all-green. A new, changed or unwaived test failure uses
stage 4's recovery before publication. Otherwise continue to Step 17.
---
@@ -410,93 +562,93 @@ If `ALREADY_PUSHED`, skip the push but continue to Step 18. Otherwise push with
git push -u origin <branch-name>
```
**If the push fails, STOP.** Report its error; do not run Steps 18–19 or claim
publication. For a non-fast-forward rejection, fetch and inspect the remote branch,
merge its changes without rewriting history, and return to Step 5 through Step 16
before retrying. Resolve ambiguous conflicts with the user; never force-push.
For authentication, hook, or network failures, fix that cause, rerun affected checks
if content changed, then recheck Step 16 before retrying. Never bypass a failed guard.
**If the push fails, STOP.** No Step 19 or publication claim. Report the error:
- **Non-fast-forward push:** fetch and inspect the remote, then merge under Step 3's
conflict rules. Run Steps 5–16 before returning to Step 17. Never rewrite history.
- **Authentication, hook or network failure:** repair the cause, then repeat Step 16
even if content is unchanged before returning to Step 17. Never bypass failed guards.
Never force-push.
Only a successful push or verified `ALREADY_PUSHED` proceeds.
Continue to mandatory Step 18 (dispatch /document-release), then Step 19 (create/update PR/MR). A push alone does not complete /ship.
Continue to Step 18. No documentation writer runs after push.
---
**PR/MR title invariant (always applies — do not skip even if you don't open the section below):** Any PR or MR you create OR update in the next step MUST have a title that starts with `v$NEW_VERSION` (the version bumped in Step 12), in the format `v<NEW_VERSION> <type>: <summary>`. Never create or edit a PR/MR title without this prefix. Compute the correct title with the single source of truth helper: `~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`. The full create/update procedure (idempotency, redaction scan, self-check) is in the section below.
## Step 18: Prepare publication metadata
**Doc-sync invariant (always applies — do not skip even if you don't open the section below):** Step 18 dispatches the /document-release subagent BEFORE the PR/MR is created or updated in Step 19. Never skip the dispatch itself; only a failed subagent is non-blocking (proceed to Step 19 without a `## Documentation` section).
First look up open PRs/MRs for `<branch-name>` on the detected platform:
- GitHub: `gh pr list --head <branch-name> --state open --json number,title,url`
- GitLab: `glab mr list --source-branch <branch-name> --output json` (defaults to open).
A successful empty array means new; one match supplies the existing title/identity.
Lookup failure or ambiguous matches **STOP** for resolution, never mean no PR.
Save the result for Step 19's recheck.
Prepare the title from that result; Step 19 scans and publishes it:
1. For an existing open PR/MR, use the matched title and run
`~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`.
2. For a new PR/MR, compose `v<NEW_VERSION> <type>: <summary>`.
3. Save the result as `NEW_TITLE` for Step 19. Every created or updated title MUST
start with `v$NEW_VERSION `; never publish an unprefixed title.
{{SECTION:pr-body}}
## Step 20: Persist ship metrics
Log coverage and plan completion for `/retro` through `gstack-review-log`.
It resolves the project/branch, validates JSON, creates storage and queues sync.
It takes **no path argument**: hand-built `<branch>-reviews.jsonl` paths break
branches containing `/`.
Log metrics for `/retro` through `gstack-review-log`; it handles project/branch paths,
JSON validation, storage and sync. It takes **no path argument**; do not build one.
```bash
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"ship","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","coverage_pct":COVERAGE_PCT,"plan_items_total":PLAN_TOTAL,"plan_items_done":PLAN_DONE,"verification_result":"VERIFY_RESULT","version":"VERSION","branch":"'"$(git rev-parse --abbrev-ref HEAD)"'"}'
```
Substitute from earlier steps:
- **COVERAGE_PCT**: coverage percentage from Step 7 diagram (integer, or -1 if undetermined)
- **COVERAGE_PCT**: Step 7 diagram's integer percentage; encode null/undetermined as -1
- **PLAN_TOTAL**: total plan items extracted in Step 8 (0 if no plan file)
- **PLAN_DONE**: count of DONE + CHANGED items from Step 8 (0 if no plan file)
- **VERIFY_RESULT**: "pass", "fail", or "skipped" from Step 8.1
- **VERIFY_RESULT**: "pass", "fail", or "skipped", set after Step 9 executes Step 8.1's verification list
- **VERSION**: from the VERSION file
The branch name is filled in by the shell — there is no `BRANCH` placeholder to
substitute.
This step is automatic — never skip it, never ask for confirmation.
The shell supplies the branch. Run this automatically, without confirmation.
---
## Step 21: Plan-tune discoverability nudge (first-successful-ship only)
Plan-tune cathedral T15. After a successful ship, surface /plan-tune once
per machine. Single line, non-blocking, marker-gated so it never re-fires.
After a successful ship, show the non-blocking /plan-tune nudge once per machine:
```bash
_NUDGE_MARKER="$HOME/.gstack/.plan-tune-nudge-shown"
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"
export GSTACK_STATE_ROOT
_NUDGE_MARKER="$GSTACK_STATE_ROOT/.plan-tune-nudge-shown"
_QT=$(~/.claude/skills/gstack/bin/gstack-config get question_tuning 2>/dev/null || echo "false")
if [ ! -f "$_NUDGE_MARKER" ] && [ "$_QT" = "false" ]; then
echo ""
echo "gstack can learn from your AskUserQuestion answers. Run /plan-tune to opt in"
echo "— it captures which prompts you find valuable vs noisy and (with hooks installed)"
echo "auto-decides your never-ask preferences."
touch "$_NUDGE_MARKER"
mkdir -p "$GSTACK_STATE_ROOT" && touch "$_NUDGE_MARKER"
fi
```
If the marker exists, OR question_tuning is already on, the nudge is a
no-op. The marker guarantees at-most-once per machine. To re-enable:
`rm ~/.gstack/.plan-tune-nudge-shown` before next ship.
The marker or enabled question_tuning suppresses it. To re-enable, remove
`$GSTACK_STATE_ROOT/.plan-tune-nudge-shown` before the next ship.
---
## Section self-check (before you finish)
You ran a carved skill. For your situation, list every section the Section index
named as applying, and confirm you issued a Read for each one. If you executed any
of those steps from memory without reading its section, you skipped the source of
truth — STOP, Read it now, and redo that step. Deterministic version work goes
through `gstack-version-bump`; never hand-roll the VERSION/package.json write.
List the applicable Section index entries and confirm each Read. If you worked from
memory, STOP, Read the section and redo that step. Use `gstack-version-bump`, never
hand-roll VERSION/package.json writes.
---
## Important Rules
- **Never skip tests.** If tests fail, stop.
- **Never skip the pre-landing review.** If checklist.md is unreadable, stop.
Follow the numbered gates and their explicit exceptions.
- **Never force push.** Use regular `git push` only.
- **Never ask for trivial confirmations** (e.g., "ready to push?", "create PR?"). DO stop for: version bumps (MINOR/MAJOR), pre-landing review findings (ASK items), and Codex structured review [P1] findings (large diffs only).
- **Always use the 4-digit version format** from the VERSION file.
- **Date format in CHANGELOG:** `YYYY-MM-DD`
- **Split commits for bisectability** — each commit = one logical change.
- **TODOS.md completion detection must be conservative.** Only mark items as completed when the diff clearly shows the work is done.
- **Use Greptile reply templates from greptile-triage.md.** Every reply includes evidence (inline diff, code references, re-rank suggestion). Never post vague replies.
- **Never push without fresh verification evidence.** If code changed after Step 5 tests, re-run before pushing.
- **Step 7 generates coverage tests.** They must pass before committing. Never commit failing tests.
- **The goal is: user says `/ship`, next thing they see is the review + PR URL + auto-synced docs.**