mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-07 20:07:21 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
+387
-203
@@ -440,47 +440,62 @@ Some steps require action on a site the user controls: registering an API key, c
|
||||
|
||||
# Ship: Fully Automated Ship Workflow
|
||||
|
||||
Run `/ship` through to the PR URL. This request authorizes routine work without confirmation; explicit safety and user-decision gates still apply.
|
||||
STOP blocks advancement until the stated repair/resume route clears; without one, end this attempt.
|
||||
Answer each AskUserQuestion before continuing.
|
||||
Routine authorization never waives those gates or their required user decisions.
|
||||
|
||||
**Route through the workflow:** detect and merge the base (Steps 1–3), test and
|
||||
audit the integrated diff (Steps 4–8.2), review and resolve findings (Steps
|
||||
9–11), prepare the release and commits (Steps 12–15), then verify, push, sync
|
||||
docs, and open or update the PR (Steps 16–19). A review fix returns to affected
|
||||
tests and reviews before release preparation; a later code or build-input edit
|
||||
returns to affected checks and Step 16 before publication. Reuse still-valid
|
||||
results, but never treat an earlier review or test as covering changed inputs.
|
||||
**Routine work needs no confirmation:** include uncommitted changes, choose MICRO/PATCH
|
||||
under Step 12, draft CHANGELOG and commits, mark completed TODOs and auto-fix findings.
|
||||
When Step 7 coverage meets its target, report remaining gaps and verify generated
|
||||
tests without another permission question. Step 15 commits those tests.
|
||||
|
||||
**Follow every STOP and AskUserQuestion gate**, including:
|
||||
- On the base branch (abort)
|
||||
- Merge conflicts that can't be auto-resolved (stop, show conflicts)
|
||||
- In-branch test failures (pre-existing failures are triaged, not auto-blocking)
|
||||
- Pre-landing review finds ASK items that need user judgment
|
||||
- Prior Learnings needs its first-time cross-project setting (Step 8)
|
||||
- MINOR or MAJOR version bump needed (ask — see Step 12)
|
||||
- Greptile review comments that need user decision (complex fixes, false positives)
|
||||
- AI-assessed coverage below target (see Step 7 for minimum/target decisions)
|
||||
- Plan items NOT DONE or UNVERIFIABLE (see Step 8)
|
||||
- Plan verification failures (see Step 8.1)
|
||||
- TODOS.md missing and user wants to create one (ask — see Step 14)
|
||||
- TODOS.md disorganized and user wants to reorganize (ask — see Step 14)
|
||||
**Route:** integrate (1–3) → test and review (4–11.5) → prepare the release
|
||||
(12–15) → verify frozen content (16) → push and publish (17–21).
|
||||
Every new invocation repeats Steps 1–16, including both reviews and the docs audit.
|
||||
Steps 12, 17 and 19 prevent duplicate bumps, pushes and PRs, never verification.
|
||||
|
||||
**Never stop for:**
|
||||
- Uncommitted changes (always include them)
|
||||
- Version bump choice (auto-pick MICRO or PATCH — see Step 12)
|
||||
- CHANGELOG content (auto-generate from diff)
|
||||
- Commit message approval (auto-commit)
|
||||
- Multi-file changesets (auto-split into bisectable commits)
|
||||
- TODOS.md completed-item detection (auto-mark)
|
||||
- Auto-fixable review findings (dead code, N+1, stale comments — fixed automatically)
|
||||
- Test coverage gaps within target threshold (generate, verify, then commit with Step 15; flag any remaining gaps in the PR body)
|
||||
### Keep state between steps
|
||||
|
||||
**Re-run behavior (idempotency):**
|
||||
Every invocation repeats verification: tests, coverage, plan completion, both
|
||||
reviews, VERSION/CHANGELOG, TODOS and doc-sync. Only *actions* are idempotent:
|
||||
- Step 12: If VERSION already bumped, skip the bump but still read the version
|
||||
- Step 17: If already pushed, skip the push command
|
||||
- Step 19: If PR exists, update the body instead of creating a new PR
|
||||
Prior execution never exempts verification.
|
||||
Keep one private Markdown **invocation record** outside the product tree and save
|
||||
its absolute path. Use these headings so a paused run can resume:
|
||||
- **Release:** versions, `BUMP_LEVEL`, reviewed tree and attempt counts.
|
||||
- **Decisions:** each approval's finding, files and authorized action. Reuse it only
|
||||
for that same scope; a repair never resets approvals or expands them.
|
||||
- **Reviews:** handles, original start tokens, terminal states, outputs and queued fixes.
|
||||
- **Checks:** command/label, result/counts, timestamp, log and consumed inputs.
|
||||
- **Documentation:** candidate/id, attempts used, accepted hashes or named blocked exception.
|
||||
- **Next steps:** one ordered work list, with the current step marked.
|
||||
|
||||
A **receipt** is saved evidence of a check's command, result and consumed content.
|
||||
A review's **start token** is the opaque value returned by `gstack-review-log --start`
|
||||
before it reads the diff. Keep `REVIEW_START` for Step 9, a separate `PASS_START` for
|
||||
each Step 11 attempt, and `DESIGN_START` for design. Finish each pass with its original
|
||||
token; `--finish` stamps the binding fields automatically. Never borrow or replace a token.
|
||||
`gstack-wtree` prints a Git tree hash covering tracked and non-ignored untracked files,
|
||||
not a commit ID. Use `git diff <old-tree> <new-tree>` to compare these snapshots.
|
||||
|
||||
### Ship control flow
|
||||
|
||||
You, the **parent** running /ship, own advancement; children return evidence, not
|
||||
permission to proceed. Follow the saved work list:
|
||||
|
||||
1. Start with Steps 1–21 in order, including 11.5 and 14.5. Advance only after
|
||||
the current item's gates clear.
|
||||
2. Expand a repair into individual steps and insert them before the still-pending
|
||||
work. This replaces the current item, whose actual result stays in the record.
|
||||
Add its destination only if not already the next pending step.
|
||||
3. For another repair, repeat rule 2 without discarding pending work.
|
||||
The saved list takes precedence over ordinary next-step
|
||||
sentences inside a repair. A range never adds unlisted steps.
|
||||
|
||||
**Example:** Step 11 fixes insert `9 → 10 → 11` before 11.5. A further Step 9 fix
|
||||
affecting 6–8 makes the list `5 → 6 → 7 → 8 → 9 → 10 → 11 → 11.5`.
|
||||
The unchanged release steps follow. STOP and AskUserQuestion gates still apply during repairs.
|
||||
|
||||
Keep the same attempt counts throughout the invocation. A range ending at Step 14
|
||||
does not enter Step 14.5. A range that includes Step 14.5 enters its existing audit
|
||||
decision, not an unconditional new launch; its initial-plus-ONE limit never resets.
|
||||
Permitted repairs continue in this invocation without restarting /ship.
|
||||
|
||||
---
|
||||
|
||||
@@ -491,15 +506,18 @@ sections. Read a section in full before doing its step; do not work from memory.
|
||||
|
||||
| When | Read this section |
|
||||
|------|-------------------|
|
||||
| the ship target is an Apple platform app (.xcodeproj, .xcworkspace, or an app-product Swift package) — read BEFORE Step 1's branch gate and any preflight; store distribution never routes through the branch/PR ceremony | `sections/apple-release.md` |
|
||||
| App Store/TestFlight distribution is requested for an Apple app (.xcodeproj, .xcworkspace, or an app-product Swift package) — read at Step 0.9 before the branch gate; an Apple repository-landing request follows the normal pipeline | `sections/apple-release.md` |
|
||||
| running the test suites and (if prompt files changed) the eval suites (Steps 4-6) | `sections/tests.md` |
|
||||
| auditing test coverage of the diff (Step 7) | `sections/test-coverage.md` |
|
||||
| auditing plan completion, verification, and scope drift (Step 8) | `sections/plan-completion.md` |
|
||||
| the pre-landing review and specialist dispatch (Step 9) | `sections/review-army.md` |
|
||||
| exploratory QA before Fix-First (Step 9.2.1) | Use the QA Read directive in `sections/review-army.md` |
|
||||
| reusing explicitly skipped shared-code advice (Step 9.3) | `sections/shared-code-reuse.md` |
|
||||
| addressing Greptile review comments when a PR exists (Step 10) | `sections/greptile.md` |
|
||||
| the adversarial review and learnings capture (Step 11) | `sections/adversarial.md` |
|
||||
| writing the CHANGELOG entry (Step 13) | `sections/changelog.md` |
|
||||
| dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19) | `sections/pr-body.md` |
|
||||
| auditing docs before final commit/verification (Step 14.5), on every ship | `sections/documentation.md` |
|
||||
| creating or updating the PR/MR with the verified documentation outcome (Step 19) | `sections/pr-body.md` |
|
||||
|
||||
---
|
||||
|
||||
@@ -549,8 +567,11 @@ branch name wherever the instructions say "the base branch" or `<default>`.
|
||||
|
||||
## Step 0.9: Apple target detection
|
||||
|
||||
If the repo has an `.xcodeproj`, `.xcworkspace`, or Swift app package AND the ask
|
||||
is App Store/TestFlight distribution, **STOP and Read
|
||||
If the ask is App Store/TestFlight distribution, look for an `.xcodeproj`,
|
||||
`.xcworkspace`, or Swift app product. Read `Package.swift` and its entrypoint to
|
||||
distinguish an app from a library/CLI. If unclear, use AskUserQuestion to identify
|
||||
the target and wait before choosing a release path.
|
||||
For a confirmed app, **STOP and Read
|
||||
`~/.claude/skills/gstack/ship/sections/apple-release.md` FIRST**. Store distribution proceeds
|
||||
through that adapter from the current branch, including a clean base branch.
|
||||
The branch gate and repository-landing pipeline below apply ONLY to
|
||||
@@ -558,7 +579,7 @@ repository-landing asks, including on Apple repos.
|
||||
|
||||
## Step 1: Pre-flight
|
||||
|
||||
1. Check the current branch. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
|
||||
1. Save the current branch as `<branch-name>`. If on the base branch or the repo's default branch, **abort**: "You're on the base branch. Ship from a feature branch."
|
||||
|
||||
2. Run `git status` (never use `-uall`). Uncommitted changes are always included — no need to ask.
|
||||
|
||||
@@ -567,9 +588,8 @@ repository-landing asks, including on Apple repos.
|
||||
`git diff origin/<base> --stat`, untracked files from status, and
|
||||
`git log origin/<base>..HEAD --oneline`.
|
||||
|
||||
4. Display historical review readiness. This preflight snapshot does not replace
|
||||
Step 9's mandatory review or its blocker, ASK, and convergence gates — even
|
||||
when prior reviews are CLEAR or the dashboard's global skip is enabled.
|
||||
4. Display historical readiness using the dashboard below, then finish Step 1.
|
||||
Prior CLEAR reviews or dashboard skips never replace Step 9's gates.
|
||||
|
||||
## Review Readiness Dashboard
|
||||
|
||||
@@ -579,68 +599,95 @@ During pre-flight, read the existing review log and config to display readiness;
|
||||
~/.claude/skills/gstack/bin/gstack-review-read
|
||||
```
|
||||
|
||||
Render each record using its recorded host, source, outside_provider, outside_status, and phase. Historical source "claude" means a native Claude subagent; source "claude-code" means the external CLI. Never infer a historical provider from the current harness. Unknown model identity remains unknown. Missing/disabled/skipped outside coverage is distinct from native completion.
|
||||
**1. Choose the records to display.** Use the latest record for each row below.
|
||||
Do not use a record older than 7 days to clear a row, and never substitute an older
|
||||
success for a newer failure. Ship metrics are not review records.
|
||||
|
||||
Parse the output. Find the most recent entry for each skill (plan-ceo-review, plan-eng-review, review, plan-design-review, design-review-lite, adversarial-review, codex-review, codex-plan-review). Ignore entries with timestamps older than 7 days. For the Eng Review row, show whichever is more recent between `review` (diff-scoped pre-landing review) and `plan-eng-review` (plan-stage architecture review). Append "(DIFF)" or "(PLAN)" to the status to distinguish. For the Adversarial row, show whichever is more recent between `adversarial-review` (new auto-scaled) and `codex-review` (legacy). For Design Review, show whichever is more recent between `plan-design-review` (full visual audit) and `design-review-lite` (code-level check). Append "(FULL)" or "(LITE)" to the status to distinguish. For the Outside Voice row, show the most recent `codex-plan-review` entry — this captures outside voices from both /plan-ceo-review and /plan-eng-review.
|
||||
| Row | Choose the latest of | Status suffix |
|
||||
|---|---|---|
|
||||
| Eng Review | `review` or `plan-eng-review` | (DIFF) or (PLAN) |
|
||||
| CEO Review | `plan-ceo-review` | — |
|
||||
| Design Review | `plan-design-review` or `design-review-lite` | (FULL) or (LITE) |
|
||||
| Adversarial | `adversarial-review` or legacy `codex-review` | — |
|
||||
| Outside Voice | `codex-plan-review` from CEO or Eng review | — |
|
||||
|
||||
**Source attribution:** If the most recent entry for a skill has a \`"via"\` field, append it to the status label in parentheses. Examples: `plan-eng-review` with `via:"autoplan"` shows as "CLEAR (PLAN via /autoplan)". `review` with `via:"ship"` shows as "CLEAR (DIFF via /ship)". Entries without a `via` field show as "CLEAR (PLAN)" or "CLEAR (DIFF)" as before.
|
||||
Keep each record's host, source, outside_provider, outside_status and phase.
|
||||
Historical source "claude" is a native subagent; "claude-code" is the external CLI.
|
||||
Do not infer old providers or unknown models from today's harness. A native result
|
||||
does not fill missing, disabled or skipped outside coverage.
|
||||
|
||||
From gstack-review-read output, use entries whose skill is `autoplan-voices` or `design-outside-voices` for the coverage detail below the dashboard. Group by workflow run and phase, not merely skill. Show each phase’s recorded provider and outside_status; partial coverage must remain partial. These records do not change the engineering gate.
|
||||
**Source attribution:** Append a recorded `via` to the suffix, for example
|
||||
"CLEAR (PLAN via /autoplan)" or "CLEAR (DIFF via /ship)". Without `via`, keep
|
||||
"CLEAR (PLAN)" or "CLEAR (DIFF)". Below the dashboard, group `autoplan-voices`
|
||||
and `design-outside-voices` by workflow run and phase. Show each phase's provider
|
||||
and outside_status; retain partial coverage. These details do not clear Eng Review.
|
||||
|
||||
Display:
|
||||
**2. Check freshness before choosing a verdict.**
|
||||
|
||||
```
|
||||
+====================================================================+
|
||||
| REVIEW READINESS DASHBOARD |
|
||||
+====================================================================+
|
||||
| Review | Runs | Last Run | Status | Required |
|
||||
|-----------------|------|---------------------|-----------|----------|
|
||||
| Eng Review | 1 | 2026-03-16 15:00 | CLEAR | YES |
|
||||
| CEO Review | 0 | — | — | no |
|
||||
| Design Review | 0 | — | — | no |
|
||||
| Adversarial | 0 | — | — | no |
|
||||
| Outside Voice | 0 | — | — | no |
|
||||
+--------------------------------------------------------------------+
|
||||
| VERDICT: CLEARED — Eng Review passed |
|
||||
+====================================================================+
|
||||
```
|
||||
- **Content-first rule:** For `review`, `adversarial-review`, `codex-review`,
|
||||
ship-stage reviews and `design-review-lite`, use `review_freshness.status`
|
||||
and show its `reason`. CURRENT means a completed clean review whose start and
|
||||
end content fingerprints equal the current `---WTREE---` fingerprint. This
|
||||
fingerprint covers working-tree content, not just the commit.
|
||||
STALE or UNVERIFIED cannot clear Eng Review. Missing `review_freshness`,
|
||||
including legacy log-only records, means UNVERIFIED. Never fall back to HEAD
|
||||
equality or commit distance for diff evidence, even at zero commits.
|
||||
Show recorded cycles, completed/converged fields and missing source/phase
|
||||
coverage. Unknown coverage is not a pass.
|
||||
- **Plan records** (plan-ceo-review, plan-eng-review, plan-design-review and
|
||||
codex-plan-review) use the 7-day window, not the working-tree fingerprint.
|
||||
If `plan_sha256` is present, you may compare the plan file and report a mismatch.
|
||||
For plan records only, compare the recorded commit with `---HEAD---`.
|
||||
If different, run `git rev-list --count STORED_COMMIT..HEAD` and report
|
||||
"Note: {skill} review from {date} may be stale — {N} commits since review".
|
||||
A failed command means UNKNOWN, treated as stale. Without commit tracking,
|
||||
retain the note to consider re-running. Omit staleness notes when all reviews
|
||||
are current.
|
||||
|
||||
**Review tiers:**
|
||||
- **Eng Review (historical readiness):** Required for a CLEARED dashboard, not for continuing Step 1. Step 9 remains mandatory, with its finding, approval and convergence gates. The skip_eng_review setting changes this dashboard only.
|
||||
- **CEO Review (optional):** Use your judgment. Recommend it for big product/business changes, new user-facing features, or scope decisions. Skip for bug fixes, refactors, infra, and cleanup.
|
||||
- **Design Review (optional):** Use your judgment. Recommend it for UI/UX changes. Skip for backend-only, infra, or prompt-only changes.
|
||||
- **Adversarial Review (automatic):** Always-on for every review. Every diff gets a native adversarial pass and, when enabled and available, a host-selected outside challenge. Large diffs (200+ lines) additionally get a structured outside review with P1 gate.
|
||||
- **Outside Voice (default-on):** Independent plan review through the host-selected provider after /plan-ceo-review and /plan-eng-review. The codex_reviews switch disables the entire extra step. Provider failure uses the existing native fallback and reports missing outside coverage. Never gates shipping.
|
||||
**3. Choose the historical verdict.** CLEARED requires the selected Eng Review
|
||||
to be `clean`, within 7 days and fresh under step 2. Otherwise report NOT CLEARED
|
||||
and its missing, stale or open-issue reason. If `skip_eng_review` is true, show
|
||||
"SKIPPED (global)" for Eng Review and CLEARED for this dashboard.
|
||||
This verdict never skips Step 9 or its finding, approval and convergence gates. Continue Step 1 even when history is NOT CLEARED.
|
||||
|
||||
**Verdict logic:**
|
||||
- **CLEARED**: Eng Review has >= 1 entry within 7 days from either \`review\` or \`plan-eng-review\` with status "clean"; diff review must also grade CURRENT below (or \`skip_eng_review\` is \`true\`)
|
||||
- **NOT CLEARED**: Eng Review missing, stale (>7 days), or has open issues
|
||||
- CEO, Design, and outside reviews are shown for context but never block shipping
|
||||
- If \`skip_eng_review\` config is \`true\`, Eng Review shows "SKIPPED (global)" and verdict is CLEARED
|
||||
Other rows provide context, not a substitute for Eng Review:
|
||||
- Recommend CEO Review for product/business or scope decisions, not routine fixes or cleanup.
|
||||
- Recommend Design Review for UI/UX work, not backend, infrastructure or prompt-only work.
|
||||
- Adversarial review always includes a native pass. Available, enabled outside
|
||||
challenges supplement it; diffs of 200+ lines also get the structured P1 gate.
|
||||
- Outside Voice is the default-on plan review after CEO/Eng review. `codex_reviews`
|
||||
disables that extra step. Provider failure uses native fallback and records
|
||||
missing outside coverage; this dashboard row never gates shipping.
|
||||
|
||||
**Staleness detection:** Grade before deciding CLEARED:
|
||||
- Ship telemetry reports metrics, not review coverage; it never satisfies a review row.
|
||||
- **Content-first rule (diff-scoped rows only: `review`, `adversarial-review`, `codex-review`, ship-stage entries, `design-review-lite`).** Use the helper's computed `review_freshness.status` and show its `reason`. CURRENT requires a completed clean pass with captured start/end wtree equal to the current `---WTREE---`. STALE or UNVERIFIED never clears Eng Review. Missing `review_freshness` is UNVERIFIED, including legacy log-only rows. Never fall back to HEAD equality or commit distance for diff evidence, even at 0 commits. Show recorded cycles, completed/converged state, and missing per-source/phase coverage; unknown is not a pass.
|
||||
- Plan-tier rows (plan-ceo-review, plan-eng-review, plan-design-review, codex-plan-review) grade a plan file, not the repo tree — never apply the wtree rule to them; they keep the 7-day freshness logic. If an entry carries `plan_sha256`, you MAY compare it with the plan file and note "plan changed since review" on mismatch.
|
||||
- Plan-tier fallback only: parse `---HEAD---`. For entries with a different `commit`, count elapsed commits: `git rev-list --count STORED_COMMIT..HEAD`. If that command FAILS, grade UNKNOWN and treat as stale. Display: "Note: {skill} review from {date} may be stale — {N} commits since review". Missing commit tracking retains the legacy note to consider re-running.
|
||||
- If all reviews grade CURRENT, do not display staleness notes
|
||||
**4. Display the dashboard.** Show missing, stale, disabled or unavailable results
|
||||
explicitly, never as CLEAR. Display a fresh `clean` result as CLEAR and
|
||||
`issues_open` as ISSUES OPEN without changing the stored status.
|
||||
|
||||
If Eng Review is not CLEAR, print its actual status and reason: "Eng Review: {status} — {reason}. Ship will run its pre-landing review in Step 9." For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend `/plan-eng-review` or `/autoplan` for architecture review.
|
||||
**REVIEW READINESS DASHBOARD**
|
||||
|
||||
If CEO Review is missing, mention as informational ("CEO Review not run — recommended for product changes") but do NOT block.
|
||||
Use one row for each entry in step 1. Only Eng Review is marked required.
|
||||
|
||||
| Review | Runs | Last run | Status | Required |
|
||||
|---|---:|---|---|---|
|
||||
| {row and suffix} | {count} | {timestamp or —} | {actual status and reason} | {yes/no} |
|
||||
|
||||
VERDICT: {CLEARED or NOT CLEARED} — {reason}
|
||||
|
||||
For diffs >200 lines (`git diff origin/<base> --stat | tail -1`), recommend
|
||||
`/plan-eng-review` or `/autoplan` for architecture review.
|
||||
|
||||
For Design Review: run `source <(~/.claude/skills/gstack/bin/gstack-diff-scope <base> 2>/dev/null)`. If `SCOPE_FRONTEND=true` and no design review exists, mention: "Design Review not run — Step 9 includes the lite check; consider /design-review for a full visual audit."
|
||||
|
||||
Continue to Step 2 without a preflight approval question. Apply the review gates when Step 9 runs.
|
||||
Continue to Step 2 without asking; Step 9 applies the review gates.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Distribution Pipeline Check
|
||||
|
||||
If the diff introduces a new standalone artifact (CLI binary, library package, tool) — not a web
|
||||
service with existing deployment — verify that a distribution pipeline exists.
|
||||
Check distribution for new standalone artifacts (CLI binaries, packages, tools),
|
||||
not web services with existing deployment.
|
||||
|
||||
1. Check for newly added distribution entry points and package manifests:
|
||||
1. List candidate distribution paths:
|
||||
```bash
|
||||
git diff origin/<base> --diff-filter=A --name-only | grep -E '(^|/)(cmd/[^/]+/main\.go|bin/[^/]+|Cargo\.toml|setup\.py|package\.json)$' | head -5
|
||||
```
|
||||
@@ -655,16 +702,17 @@ service with existing deployment — verify that a distribution pipeline exists.
|
||||
grep -qE 'release|publish|deploy' .gitlab-ci.yml 2>/dev/null && echo "GITLAB_CI_RELEASE"
|
||||
```
|
||||
|
||||
3. **If no release pipeline exists and a new artifact was added:** Use AskUserQuestion:
|
||||
- "This PR adds a new binary/tool but there's no CI/CD pipeline to build and publish it.
|
||||
Users won't be able to download the artifact after merge."
|
||||
- A) Add a release workflow now (CI/CD release pipeline — GitHub Actions or GitLab CI depending on platform)
|
||||
- B) Defer — add a P1 distribution TODO in Step 14
|
||||
- C) Not needed — this is internal/web-only, existing deployment covers it
|
||||
3. **New artifact without a pipeline:** AskUserQuestion: "Users cannot download this
|
||||
artifact after merge without a release pipeline."
|
||||
- A) Add the platform's release workflow now
|
||||
- B) Defer with a P1 distribution TODO in Step 14
|
||||
- C) Not needed: internal/web-only, covered by existing deployment
|
||||
|
||||
4. **If the user chooses A:** Add packaging and publish configuration using this repository's CI conventions. Ask for the intended distribution target if it is unknown; do not invent a registry or credentials. Include the new workflow in the tests and review below. Do not publish a release during `/ship`.
|
||||
5. **If release pipeline exists:** Continue silently.
|
||||
6. **If no new artifact detected:** Skip silently.
|
||||
4. **If A:** Add packaging/publish configuration using repository CI conventions.
|
||||
Ask for unknown targets, registries or access first; never invent credentials.
|
||||
Recheck against the artifact and include the workflow in tests and review.
|
||||
Do not publish a release during `/ship`.
|
||||
5. Otherwise, continue without adding a pipeline.
|
||||
|
||||
---
|
||||
|
||||
@@ -676,10 +724,14 @@ Merge the base ref fetched in Step 1 so tests and reviews cover the integrated c
|
||||
git merge origin/<base> --no-edit
|
||||
```
|
||||
|
||||
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). If conflicts are complex or ambiguous, **STOP** and show them.
|
||||
**If there are merge conflicts:** Try to auto-resolve if they are simple (VERSION, schema.rb, CHANGELOG ordering). For complex or ambiguous conflicts, **STOP**, show the conflicting choices, use AskUserQuestion for the needed resolution decision, and wait for the answer before editing or continuing.
|
||||
|
||||
**If already up to date:** Continue silently.
|
||||
|
||||
If integration changes the artifact or distribution configuration inspected in Step 2,
|
||||
repeat Step 2 on the merged content, including its decisions, then continue to Step 4.
|
||||
Otherwise continue to Step 4 directly.
|
||||
|
||||
---
|
||||
|
||||
> **STOP.** Before running the test suites and (if prompt files changed) the eval suites (Steps 4-6), Read `~/.claude/skills/gstack/ship/sections/tests.md` and execute it
|
||||
@@ -700,43 +752,92 @@ git merge origin/<base> --no-edit
|
||||
> **STOP.** Before the adversarial review and learnings capture (Step 11), Read `~/.claude/skills/gstack/ship/sections/adversarial.md` and execute it
|
||||
> in full. Do not work from memory — that section is the source of truth for this step.
|
||||
|
||||
## Step 11.5: Bind the reviews
|
||||
|
||||
1. **Select the two reviews.** Run `~/.claude/skills/gstack/bin/gstack-review-read`.
|
||||
Select this invocation's final Step 9.4 record (`skill:"review"`, `via:"ship"`)
|
||||
and Step 11 native record (`skill:"adversarial-review"`). Match each to its saved
|
||||
handle, original token and source; reject outside-provider or older invocation records.
|
||||
2. **Compare their content.** Require the native record's `review_binding.state`
|
||||
to be `verified`. All three snapshots must match: its `wtree`, Step 9.4's
|
||||
`review_binding.start_wtree` and `review_binding.end_wtree`. A mismatch or missing
|
||||
record/field blocks release preparation: report **Review records missing or mismatched**
|
||||
and insert `9 → 10 → 11 → 11.5` before Step 12. Bind the new records at 11.5.
|
||||
Never attach new tokens to old work.
|
||||
3. **Preserve any QA exception.** A named probe-risk exception may leave Step 9.4's
|
||||
root `wtree` absent; item 2 still compares its start/end snapshots. Matching content
|
||||
does not mean the failed or unrun probes passed. Keep Step 9.4's incomplete flags
|
||||
and the user's exception.
|
||||
4. **Save the evidence.** Save both records and matching **reviewed tree** for
|
||||
Step 16. Continue to Step 12.
|
||||
|
||||
## Step 12: Version bump (auto-decide)
|
||||
|
||||
Use **`gstack-version-bump`** for classify/write/repair and `gstack-next-version`
|
||||
for slot selection. Bump level and queue collisions remain agent decisions.
|
||||
Item 3 needs `BUMP_LEVEL`: reuse this invocation's saved level. Otherwise FRESH
|
||||
chooses it in item 2 and ALREADY_BUMPED derives it in item 1.
|
||||
|
||||
1. **Classify state** — pure reader, never writes:
|
||||
```bash
|
||||
bun run ~/.claude/skills/gstack/bin/gstack-version-bump classify --base <base>
|
||||
```
|
||||
Save the JSON `baseVersion` as `BASE_VERSION`, then read `state` and dispatch:
|
||||
- **FRESH** → do the bump (steps 2-4).
|
||||
- **ALREADY_BUMPED** → keep `NEW_VERSION` at `currentVersion`. Reuse this branch's earlier ship decision for `BUMP_LEVEL` if recorded; otherwise compare `baseVersion` and `currentVersion` left to right: the first changed major/minor/patch/micro component supplies `BUMP_LEVEL` (a missing fourth component is zero). Then run step 3's queue check. This recovers the level, not permission to bump again.
|
||||
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify. On success, follow **ALREADY_BUMPED**, including its queue check; on failure, STOP. Repair alone never re-bumps.
|
||||
- **DRIFT_UNEXPECTED** → **STOP**. package.json disagrees with VERSION while VERSION matches base — a manual edit bypassed /ship. Reconcile manually, then re-run.
|
||||
- **FRESH** → use the recorded level or choose it in item 2, then check the queue and write.
|
||||
- **ALREADY_BUMPED** → keep `NEW_VERSION=currentVersion`. If `BUMP_LEVEL` is missing,
|
||||
use the first changed component from `baseVersion` to `currentVersion`
|
||||
(major/minor/patch/micro; an absent fourth component is zero). Continue at item 3,
|
||||
not another automatic bump.
|
||||
- **DRIFT_STALE_PKG** → run `gstack-version-bump repair`, then reclassify.
|
||||
Success follows ALREADY_BUMPED, including its queue check; failure stops.
|
||||
Repair alone never re-bumps.
|
||||
- **DRIFT_UNEXPECTED** → STOP: package.json disagrees with VERSION while VERSION
|
||||
matches base. Reconcile the manual edit, then reclassify.
|
||||
|
||||
2. **Decide the bump level** from the diff (agent judgment):
|
||||
- **MICRO**: <50 lines, trivial tweaks/config. **PATCH**: 50+ lines, no feature signals.
|
||||
- **MINOR**: AskUserQuestion for any feature signal (new route/page, migration, new module), OR 500+ lines. **MAJOR**: AskUserQuestion for milestones or breaking changes. Offer the recommended level with rationale, a smaller level, or cancel; wait for the answer. Cancel ends this ship attempt before release writes or push; preserve existing work.
|
||||
Save `BUMP_LEVEL` as lowercase `micro`, `patch`, `minor`, or `major`. Queue placement may advance the slot without changing the intended level.
|
||||
- **MINOR**: ask for any feature signal (new route/page, migration, module) or 500+ lines.
|
||||
**MAJOR**: ask for milestones or breaking changes. Use AskUserQuestion: recommended
|
||||
level with rationale, smaller level, or cancel. Wait; cancel stops before release
|
||||
writes or push and preserves existing work.
|
||||
Save lowercase `BUMP_LEVEL`. A claimed version may move the next available number
|
||||
forward, but cannot change the chosen MICRO/PATCH/MINOR/MAJOR level.
|
||||
|
||||
3. **Queue-aware pick** (workspace-aware ship):
|
||||
```bash
|
||||
QUEUE_JSON=$(bun run ~/.claude/skills/gstack/bin/gstack-next-version --base <base> --bump "$BUMP_LEVEL" --current-version "$BASE_VERSION" 2>/dev/null || echo '{"offline":true}')
|
||||
CANDIDATE_VERSION=$(echo "$QUEUE_JSON" | jq -r '.version // empty')
|
||||
```
|
||||
- **Usable candidate** (including `offline:true` with `fallback:"git"`): print warnings and any claimed queue. FRESH sets `NEW_VERSION` to `CANDIDATE_VERSION`. ALREADY_BUMPED compares it with `currentVersion`; if different, ask to rebump (refresh CHANGELOG/PR title) or keep current (CI rejects a collision). Only approval changes the existing version. An active sibling is a workspace listed in JSON `active_siblings`; use its `branch` and `version`. If one holds `>= NEW_VERSION`, ask to advance past it or stop this attempt and sync.
|
||||
- **No usable candidate** (utility failure or empty result): print queue-unverified; FRESH sets `NEW_VERSION` using local `BUMP_LEVEL` arithmetic, while ALREADY_BUMPED keeps `currentVersion`. Do not follow the usable-candidate instructions above.
|
||||
**Qualify first:** require successful utility output and a nonempty valid version.
|
||||
`offline:false` qualifies; `offline:true` qualifies only with `fallback:"git"`.
|
||||
Offline output without that fallback, failure, malformed output or an empty version
|
||||
is unusable, even if it contains a version-looking string.
|
||||
|
||||
- **Usable candidate:** print warnings and claimed queue. FRESH sets `NEW_VERSION=CANDIDATE_VERSION`.
|
||||
ALREADY_BUMPED compares it with `currentVersion`: if different, ask to rebump
|
||||
(refresh CHANGELOG/PR title) or keep current (CI rejects a collision).
|
||||
Only approval changes the existing version. Check JSON `active_siblings` by
|
||||
`branch` and `version`; a sibling holding `>= NEW_VERSION` requires a choice:
|
||||
advance past it, or stop this attempt and sync.
|
||||
- **No usable candidate:** print queue-unverified. FRESH uses local `BUMP_LEVEL`
|
||||
arithmetic; ALREADY_BUMPED keeps `currentVersion`. Never use an empty candidate.
|
||||
|
||||
4. **Write the bump** (FRESH, or an approved rebump):
|
||||
```bash
|
||||
bun run ~/.claude/skills/gstack/bin/gstack-version-bump write --version "$NEW_VERSION" --regen-digest
|
||||
```
|
||||
The CLI validates 4-digit `MAJOR.MINOR.PATCH.MICRO` (or 3-digit pinned semver), then writes VERSION, the manifest, and existing `package-lock.json` / `npm-shrinkwrap.json` files; it never creates lockfiles. Manifest resolution: `--package-json-path` → `.gstack/package-json-path` → `./package.json` (supports subdirectory packages). npm manifests/locks use the 3-digit translation (`1.67.0.0` → `1.67.0`); VERSION remains authoritative. Exit 3 means a half-write: reclassify and use `repair` for DRIFT_STALE_PKG.
|
||||
The CLI validates `MAJOR.MINOR.PATCH.MICRO` (or pinned 3-digit semver) and writes
|
||||
VERSION, the manifest and existing `package-lock.json` / `npm-shrinkwrap.json`;
|
||||
it never creates lockfiles. Manifest path: `--package-json-path` →
|
||||
`.gstack/package-json-path` → `./package.json`. npm files use the 3-digit translation
|
||||
(`1.67.0.0` → `1.67.0`); VERSION is authoritative. Exit 3 means a half-write:
|
||||
reclassify and `repair` DRIFT_STALE_PKG.
|
||||
|
||||
`--regen-digest` executes repo code with the same privileges as Step 5: `scripts/gen-agents-digest.ts`, only when it and committed `agents-digest/gstack-AGENTS.md` both exist. Check `agentsDigest`: if false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump before continuing. Its VERSION stamp is freshness-gated.
|
||||
`--regen-digest` runs repo code with Step 5's privileges: `scripts/gen-agents-digest.ts`,
|
||||
only when it and committed `agents-digest/gstack-AGENTS.md` exist. If `agentsDigest`
|
||||
is false, run `bun scripts/gen-agents-digest.ts` and stage the digest with the bump.
|
||||
Before push, verify the committed digest matches generation for the selected VERSION.
|
||||
|
||||
5. **Record the release decision** (skip if ALREADY_BUMPED):
|
||||
5. **Record the release decision after a version was actually written**, including
|
||||
an approved ALREADY_BUMPED rebump. Skip unchanged versions and manifest-only repairs.
|
||||
```bash
|
||||
~/.claude/skills/gstack/bin/gstack-decision-log '{"decision":"Ship NEW_VERSION (BUMP_LEVEL)","rationale":"WHY","scope":"repo","source":"skill","confidence":9}' 2>/dev/null || true
|
||||
```
|
||||
@@ -747,13 +848,15 @@ for slot selection. Bump level and queue collisions remain agent decisions.
|
||||
|
||||
## Step 14: TODOS.md (auto-update)
|
||||
|
||||
Persist approved follow-ups, then conservatively mark completed work.
|
||||
Read `~/.claude/skills/gstack/review/TODOS-format.md`.
|
||||
|
||||
Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format reference (or `review/TODOS-format.md` in a gstack checkout).
|
||||
**1. Open or create:** Read root `TODOS.md`. An explicit "add TODO" choice authorizes
|
||||
creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: A) Create
|
||||
a component/priority-organized TODOS.md, B) Skip. Skip goes to item 5.
|
||||
|
||||
**1. Open or create:** Read root `TODOS.md`. An earlier explicit "add TODO" choice authorizes its creation with `# TODOS` and `## Completed`. Otherwise, if missing, ask: "Create a component/priority-organized TODOS.md?" Options: A) Create now, B) Skip. If B, continue to Step 15 with the outcome in the summary below.
|
||||
|
||||
**2. Organization:** Expect component headings, `**Priority:**` P0–P4 fields, and `## Completed` at the bottom. If disorganized, ask: A) Reorganize (recommended), B) Leave as-is. A preserves all content; B continues without restructuring.
|
||||
**2. Organization:** Use component headings, `**Priority:**` P0–P4 and `## Completed`
|
||||
at the bottom. If disorganized, ask: A) Reorganize preserving all content
|
||||
(recommended), B) Leave as-is.
|
||||
|
||||
**3. Add approved deferrals:**
|
||||
- Step 2: add the approved distribution follow-up as P1 with the missing pipeline and affected artifact.
|
||||
@@ -761,25 +864,40 @@ Read `~/.claude/skills/gstack/review/TODOS-format.md` for the canonical format r
|
||||
- Step 5: retain P0 test-failure entries already written; deduplicate by failure and source, adding missing approved entries with error output and branch.
|
||||
Never turn dropped scope into TODOs or invent unapproved follow-ups. Reuse matching existing entries rather than duplicating them.
|
||||
|
||||
**4. Detect completed TODOs:** Match titles, files, and behavior against `git diff origin/<base>`, untracked files from status, and `git log origin/<base>..HEAD --oneline`. Only clear evidence earns completion; leave uncertain items open. Move completed items to `## Completed` and append `**Completed:** vX.Y.Z (YYYY-MM-DD)`.
|
||||
**4. Detect completed TODOs:** Compare titles, files and behavior with
|
||||
`git diff origin/<base>`, untracked files and `git log origin/<base>..HEAD --oneline`.
|
||||
Move proven completions to `## Completed` with `**Completed:** vX.Y.Z (YYYY-MM-DD)`;
|
||||
leave uncertain items open.
|
||||
|
||||
**5. Save the summary:** Report added/deferred items, items marked complete, remaining count, and any creation/reorganization. If creation was declined or a write fails, warn and retain the unpersisted follow-ups in the Step 19 PR summary; never claim they were saved. A TODO write failure remains non-blocking.
|
||||
**5. Save the summary:** Report additions, deferrals, completions, remaining count and
|
||||
creation/reorganization. If creation was declined or a write failed, warn and retain
|
||||
unsaved follow-ups in Step 19's PR summary. Never claim they were saved;
|
||||
TODO write failures are non-blocking.
|
||||
|
||||
---
|
||||
|
||||
## Step 14.5: Documentation audit (every ship)
|
||||
|
||||
**Doc-sync invariant:** Every ship dispatches the /document-release subagent before final
|
||||
commit/verification/publication, including reruns, already-pushed branches, existing PRs and docs-only changes.
|
||||
No edits means an executed audit, not a skip; report the section's verified outcome.
|
||||
|
||||
> **STOP.** Before auditing docs before final commit/verification (Step 14.5), on every ship, Read `~/.claude/skills/gstack/ship/sections/documentation.md` and execute it
|
||||
> in full. Do not work from memory — that section is the source of truth for this step.
|
||||
|
||||
## Step 15: Commit (bisectable chunks)
|
||||
|
||||
Create small, logical commits for `git bisect`. If all changes are already committed, continue to Step 16; never create an empty commit.
|
||||
Make bisectable commits; if already committed, continue to Step 16. Never create an empty commit.
|
||||
|
||||
1. Group by coherent change. Keep each model/service/controller with its tests;
|
||||
keep controller views together. Migrations may stand alone or accompany their
|
||||
model; config/routes may accompany the feature they enable. A diff under
|
||||
50 lines across fewer than 4 files may use one commit.
|
||||
1. Group changes with their tests, config/routes, views and Step 14.5 docs.
|
||||
Migrations may stand alone or accompany their model.
|
||||
Under 50 lines across fewer than 4 files may use one commit.
|
||||
2. Order dependencies first: infrastructure → models/services → controllers/views.
|
||||
Each commit must work independently, without broken imports or missing code.
|
||||
VERSION + CHANGELOG + TODOS.md belong in the final commit.
|
||||
Group VERSION + CHANGELOG + TODOS.md after the feature commits.
|
||||
3. Use `<type>: <summary>` (feat/fix/chore/refactor/docs) and a brief body.
|
||||
Only the final VERSION/CHANGELOG commit gets the version tag and co-author trailer:
|
||||
Only the final VERSION/CHANGELOG commit gets the release version and co-author
|
||||
trailer. Do not create a Git tag:
|
||||
|
||||
```bash
|
||||
git commit -m "$(cat <<'EOF'
|
||||
@@ -796,53 +914,119 @@ EOF
|
||||
|
||||
**IRON LAW: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE.**
|
||||
|
||||
Find generation/build commands in CLAUDE.md/AGENTS.md, package scripts, and build
|
||||
configuration; run them first, skipping only when none are defined. A failed build blocks push. If it changes tracked files, inspect the
|
||||
changes, run affected checks from Steps 6–11, refresh release facts, and commit
|
||||
under Step 15 before returning here. Reuse unchanged results and actual approvals.
|
||||
Run stages 1–5 in order. Recovery instructions below name where to resume.
|
||||
If content changes during or after verification, restart at stage 1 and complete
|
||||
all five stages before Step 17. Content-preserving commits keep valid evidence.
|
||||
|
||||
Then check test evidence against the final content:
|
||||
### 1. Finish writers and prepare outputs
|
||||
|
||||
Inspect writer handles, including the docs child. Confirm terminal completion or termination
|
||||
before another writer runs. Timeout or cancellation acknowledgment alone means
|
||||
STOP until confirmed.
|
||||
|
||||
Find declared generation/build commands in project instructions, manifests, build
|
||||
files and CI. Run them and save results. If none exists, record not applicable and
|
||||
the inspected sources. A missing prerequisite or failed build stops shipping:
|
||||
report **Build failed or prerequisite missing**, with the command, error and needed
|
||||
repair. Never invent a substitute command.
|
||||
**If blocked:** Repair the prerequisite or build, then repeat stage 1. After it passes, continue
|
||||
to stage 2; treat any content repair as a behavioral change there.
|
||||
|
||||
### 2. Choose the change route
|
||||
|
||||
Capture the current tree with `~/.claude/skills/gstack/bin/gstack-wtree`. Inspect
|
||||
`git diff <reviewed-tree> <current-tree>` against the snapshot saved before Step 12.
|
||||
Missing snapshots block this comparison, regardless of HEAD equality.
|
||||
|
||||
Classify the comparison in this order:
|
||||
|
||||
1. **Behavior, tests or build inputs changed:** Prompts/templates count as behavior.
|
||||
Insert `5–11.5 → 12–14 → 16` before the pending Step 17, then stop this step.
|
||||
This repair excludes Step 14.5 because the rebuild can change generated docs.
|
||||
Step 16 restarts at stage 1: rebuild and compare again before stage 3 decides
|
||||
documentation freshness. Further repairs use the same work list.
|
||||
2. **Only authored docs or release metadata changed:** Keep Step 8's original child
|
||||
report and counts. Recheck affected plan items using their recorded verification
|
||||
and append current evidence to the invocation record. If a classification is no
|
||||
longer supported, run Step 8's audit and decision gates only, then return to
|
||||
Step 16 stage 1. Never edit the child's counts yourself.
|
||||
3. **No changes, or the docs-only checks still support the plan:** Continue to stage 3
|
||||
without a new code review.
|
||||
|
||||
### 3. Resolve documentation freshness
|
||||
|
||||
Compare the base and hashes of the selected release paths, generated
|
||||
outputs and docs/templates with Step 14.5's saved values. A prior invocation's
|
||||
audit or risk decision never qualifies.
|
||||
|
||||
| Outcome | Action |
|
||||
|---|---|
|
||||
| This invocation's accepted audit matches all inputs | Continue to stage 4. |
|
||||
| User-accepted named documentation risk covers the same approved scope and exact content, and unwaivable gates clear | Continue to stage 4; retain `Documentation: blocked`, its reason and incomplete scope. |
|
||||
| Missing, stale or blocked | Use recovery below. Never silently refresh hashes. |
|
||||
|
||||
Report changed inputs, blockers and attempts used:
|
||||
|
||||
- **An attempt remains, with changed inputs or an available repair:** insert
|
||||
`14.5 → 15 → 16` before Step 17. Use Blocked recovery with the existing count.
|
||||
Validate the outcome before Step 15,
|
||||
then restart Step 16 stage 1 to regenerate and compare again.
|
||||
- **Otherwise:** STOP unless the user accepts
|
||||
the specific named documentation risk and all unwaivable gates clear, under
|
||||
Step 14.5's Blocked recovery rules. Unchanged approved content goes to stage 4;
|
||||
repaired content goes to stage 1.
|
||||
|
||||
Never run a third audit. Child return is not acceptance.
|
||||
|
||||
### 4. Verify the frozen candidate
|
||||
|
||||
Freeze inputs through verification and push. Run declared docs/link/generated-file
|
||||
checks; report unavailable checks.
|
||||
|
||||
**Reuse a check when its inputs match.** Compare hashes or complete bytes of its
|
||||
saved and current consumed files, fixtures, dependencies and execution parameters.
|
||||
Explain why other changes cannot affect it; changed or unknown dependencies require a rerun.
|
||||
For model judges, compare the complete expanded request, rubric, parameters and
|
||||
builder/runtime dependencies. Reuse identical passing evidence: cite the original
|
||||
command, result/counts, timestamp and log, never resample it. Mandatory reviews still run.
|
||||
|
||||
**Check each test lane's receipt as well.** Use its actual Step 5 label/command:
|
||||
`--label <lane> --expect-cmd '<exact Step 5 command>'`. Inspect changes since the run;
|
||||
`--allow-paths` exempts only release metadata. A `package.json` version-only edit
|
||||
can qualify; scripts, dependencies and runtime configuration require live tests.
|
||||
Uncertain edits cannot be exempted. Docs, TODO edits, new/generated tests and fixes
|
||||
make evidence STALE even without a new code review. Use this example only after
|
||||
confirming that every allowed edit is release metadata:
|
||||
|
||||
```bash
|
||||
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<exact tests-lane command from Step 5>' --label vitest --expect-cmd '<exact vitest-lane command from Step 5>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
|
||||
~/.claude/skills/gstack/bin/gstack-evidence check --label tests --expect-cmd '<tests>' --label vitest --expect-cmd '<vitest>' --max-age 24 --allow-paths CHANGELOG.md,VERSION,package.json,agents-digest/gstack-AGENTS.md
|
||||
```
|
||||
|
||||
Use only Step 5's actual lane labels and exact commands; `vitest` is an example.
|
||||
If Step 4 explicitly declined testing and no lanes exist, report that gap instead
|
||||
of inventing FRESH evidence. Build verification still applies.
|
||||
| Receipt result | Next action |
|
||||
|---|---|
|
||||
| FRESH (exit 0) | Cite the label, exit, timestamp and log. |
|
||||
| STALE/MISSING: changed content, command or age, or no proven run | Run `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`, read the result and recheck once. Handle failures as described below. |
|
||||
| Only receipt storage/readback failed | Independently prove unchanged final content, the same command and valid age from the successful run's evidence. Cite its exact command, exit, timestamp and log as **ledger unavailable**, never FRESH. Without that proof, use STALE/MISSING. |
|
||||
|
||||
The allow-list covers release bookkeeping, including Step 12's package/digest
|
||||
version stamps. Behavioral package.json edits still require live tests despite
|
||||
the path exemption. Do not add `TODOS.md` or generated tests to the allow-list:
|
||||
Step 7 tests, review fixes, and Step 14 TODO edits intentionally make evidence STALE.
|
||||
No test lanes: require Step 5's explicit untested-scope approval for final content,
|
||||
or run Steps 5–15, including the no-tests decision, then return to Step 16 stage 1.
|
||||
Report the gap, never FRESH; builds must pass.
|
||||
|
||||
- **Every line FRESH (exit 0):** recorded runs passed on identical content except
|
||||
the listed release files. Cite label, exit, timestamp, and log path; continue.
|
||||
- **Any STALE/MISSING (exit non-zero):** inspect the reason before choosing recovery:
|
||||
- **Content, command or age mismatch, or no passing live evidence:** rerun the
|
||||
affected lanes on final content, wrapped as `~/.claude/skills/gstack/bin/gstack-evidence run --label <lane> -- '<command>'`.
|
||||
Read results and recheck once. TODO edits and generated tests are content
|
||||
changes, not ledger-only bookkeeping.
|
||||
- **Ledger read/write failure only:** if a successful live run already covers
|
||||
the unchanged final content, exact command and permitted age, cite its exit,
|
||||
timestamp and log directly. Report ledger unavailable and continue, never
|
||||
ledger FRESH. Do not rerun green suites solely because the ledger cannot save
|
||||
or read its record. If unchanged content cannot be confirmed, STOP.
|
||||
**New, changed or unwaived test failure:** STOP publication. Run Steps 5–15,
|
||||
starting with Step 5's triage, then return to Step 16 stage 1. This recovery also
|
||||
applies if a failure appears while reporting in stage 5. Reentry to Step 14.5
|
||||
keeps its existing audit count; it does not authorize a third attempt.
|
||||
|
||||
A failed CHECK identifies evidence to repair; it is not a test failure. The
|
||||
required live RUN must pass, except for the explicit triage waiver below.
|
||||
### 5. Report, then push
|
||||
|
||||
Paste build and rerun results. Later code, test, or build-input changes return
|
||||
through this gate before pushing. Step 18 owns validation of its post-push
|
||||
docs-only edits; follow repository-required checks there too. Do not claim an
|
||||
earlier test run covered changed inputs.
|
||||
Commit only approved, verified release changes left uncommitted after Step 15,
|
||||
including generated outputs; use its grouping rules and never create an empty commit.
|
||||
Preserve unrelated user files.
|
||||
|
||||
**If tests fail here:** apply Step 5's triage. A prior explicit waiver remains valid
|
||||
only for the same verified pre-existing failures and approved scope; cite that
|
||||
approval and actual failing counts, never FRESH or all-green evidence. New,
|
||||
changed, or unwaived failures STOP publication and return to Step 5.
|
||||
|
||||
Claiming work is complete without verification is dishonesty, not efficiency.
|
||||
Paste build/docs/test results. Reuse waivers only for the same verified
|
||||
pre-existing failures and approved scope; cite the actual approval and failing
|
||||
counts, never FRESH or all-green. A new, changed or unwaived test failure uses
|
||||
stage 4's recovery before publication. Otherwise continue to Step 17.
|
||||
|
||||
---
|
||||
|
||||
@@ -929,94 +1113,94 @@ If `ALREADY_PUSHED`, skip the push but continue to Step 18. Otherwise push with
|
||||
git push -u origin <branch-name>
|
||||
```
|
||||
|
||||
**If the push fails, STOP.** Report its error; do not run Steps 18–19 or claim
|
||||
publication. For a non-fast-forward rejection, fetch and inspect the remote branch,
|
||||
merge its changes without rewriting history, and return to Step 5 through Step 16
|
||||
before retrying. Resolve ambiguous conflicts with the user; never force-push.
|
||||
For authentication, hook, or network failures, fix that cause, rerun affected checks
|
||||
if content changed, then recheck Step 16 before retrying. Never bypass a failed guard.
|
||||
**If the push fails, STOP.** No Step 19 or publication claim. Report the error:
|
||||
- **Non-fast-forward push:** fetch and inspect the remote, then merge under Step 3's
|
||||
conflict rules. Run Steps 5–16 before returning to Step 17. Never rewrite history.
|
||||
- **Authentication, hook or network failure:** repair the cause, then repeat Step 16
|
||||
even if content is unchanged before returning to Step 17. Never bypass failed guards.
|
||||
Never force-push.
|
||||
Only a successful push or verified `ALREADY_PUSHED` proceeds.
|
||||
|
||||
Continue to mandatory Step 18 (dispatch /document-release), then Step 19 (create/update PR/MR). A push alone does not complete /ship.
|
||||
Continue to Step 18. No documentation writer runs after push.
|
||||
|
||||
---
|
||||
|
||||
**PR/MR title invariant (always applies — do not skip even if you don't open the section below):** Any PR or MR you create OR update in the next step MUST have a title that starts with `v$NEW_VERSION` (the version bumped in Step 12), in the format `v<NEW_VERSION> <type>: <summary>`. Never create or edit a PR/MR title without this prefix. Compute the correct title with the single source of truth helper: `~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`. The full create/update procedure (idempotency, redaction scan, self-check) is in the section below.
|
||||
## Step 18: Prepare publication metadata
|
||||
|
||||
**Doc-sync invariant (always applies — do not skip even if you don't open the section below):** Step 18 dispatches the /document-release subagent BEFORE the PR/MR is created or updated in Step 19. Never skip the dispatch itself; only a failed subagent is non-blocking (proceed to Step 19 without a `## Documentation` section).
|
||||
First look up open PRs/MRs for `<branch-name>` on the detected platform:
|
||||
|
||||
> **STOP.** Before dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19), Read `~/.claude/skills/gstack/ship/sections/pr-body.md` and execute it
|
||||
- GitHub: `gh pr list --head <branch-name> --state open --json number,title,url`
|
||||
- GitLab: `glab mr list --source-branch <branch-name> --output json` (defaults to open).
|
||||
|
||||
A successful empty array means new; one match supplies the existing title/identity.
|
||||
Lookup failure or ambiguous matches **STOP** for resolution, never mean no PR.
|
||||
Save the result for Step 19's recheck.
|
||||
|
||||
Prepare the title from that result; Step 19 scans and publishes it:
|
||||
1. For an existing open PR/MR, use the matched title and run
|
||||
`~/.claude/skills/gstack/bin/gstack-pr-title-rewrite.sh "$NEW_VERSION" "<current title>"`.
|
||||
2. For a new PR/MR, compose `v<NEW_VERSION> <type>: <summary>`.
|
||||
3. Save the result as `NEW_TITLE` for Step 19. Every created or updated title MUST
|
||||
start with `v$NEW_VERSION `; never publish an unprefixed title.
|
||||
|
||||
> **STOP.** Before creating or updating the PR/MR with the verified documentation outcome (Step 19), Read `~/.claude/skills/gstack/ship/sections/pr-body.md` and execute it
|
||||
> in full. Do not work from memory — that section is the source of truth for this step.
|
||||
|
||||
## Step 20: Persist ship metrics
|
||||
|
||||
Log coverage and plan completion for `/retro` through `gstack-review-log`.
|
||||
It resolves the project/branch, validates JSON, creates storage and queues sync.
|
||||
It takes **no path argument**: hand-built `<branch>-reviews.jsonl` paths break
|
||||
branches containing `/`.
|
||||
Log metrics for `/retro` through `gstack-review-log`; it handles project/branch paths,
|
||||
JSON validation, storage and sync. It takes **no path argument**; do not build one.
|
||||
|
||||
```bash
|
||||
~/.claude/skills/gstack/bin/gstack-review-log '{"skill":"ship","timestamp":"'"$(date -u +%Y-%m-%dT%H:%M:%SZ)"'","coverage_pct":COVERAGE_PCT,"plan_items_total":PLAN_TOTAL,"plan_items_done":PLAN_DONE,"verification_result":"VERIFY_RESULT","version":"VERSION","branch":"'"$(git rev-parse --abbrev-ref HEAD)"'"}'
|
||||
```
|
||||
|
||||
Substitute from earlier steps:
|
||||
- **COVERAGE_PCT**: coverage percentage from Step 7 diagram (integer, or -1 if undetermined)
|
||||
- **COVERAGE_PCT**: Step 7 diagram's integer percentage; encode null/undetermined as -1
|
||||
- **PLAN_TOTAL**: total plan items extracted in Step 8 (0 if no plan file)
|
||||
- **PLAN_DONE**: count of DONE + CHANGED items from Step 8 (0 if no plan file)
|
||||
- **VERIFY_RESULT**: "pass", "fail", or "skipped" from Step 8.1
|
||||
- **VERIFY_RESULT**: "pass", "fail", or "skipped", set after Step 9 executes Step 8.1's verification list
|
||||
- **VERSION**: from the VERSION file
|
||||
|
||||
The branch name is filled in by the shell — there is no `BRANCH` placeholder to
|
||||
substitute.
|
||||
|
||||
This step is automatic — never skip it, never ask for confirmation.
|
||||
The shell supplies the branch. Run this automatically, without confirmation.
|
||||
|
||||
---
|
||||
|
||||
## Step 21: Plan-tune discoverability nudge (first-successful-ship only)
|
||||
|
||||
Plan-tune cathedral T15. After a successful ship, surface /plan-tune once
|
||||
per machine. Single line, non-blocking, marker-gated so it never re-fires.
|
||||
After a successful ship, show the non-blocking /plan-tune nudge once per machine:
|
||||
|
||||
```bash
|
||||
_NUDGE_MARKER="$HOME/.gstack/.plan-tune-nudge-shown"
|
||||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"
|
||||
export GSTACK_STATE_ROOT
|
||||
_NUDGE_MARKER="$GSTACK_STATE_ROOT/.plan-tune-nudge-shown"
|
||||
_QT=$(~/.claude/skills/gstack/bin/gstack-config get question_tuning 2>/dev/null || echo "false")
|
||||
if [ ! -f "$_NUDGE_MARKER" ] && [ "$_QT" = "false" ]; then
|
||||
echo ""
|
||||
echo "gstack can learn from your AskUserQuestion answers. Run /plan-tune to opt in"
|
||||
echo "— it captures which prompts you find valuable vs noisy and (with hooks installed)"
|
||||
echo "auto-decides your never-ask preferences."
|
||||
touch "$_NUDGE_MARKER"
|
||||
mkdir -p "$GSTACK_STATE_ROOT" && touch "$_NUDGE_MARKER"
|
||||
fi
|
||||
```
|
||||
|
||||
If the marker exists, OR question_tuning is already on, the nudge is a
|
||||
no-op. The marker guarantees at-most-once per machine. To re-enable:
|
||||
`rm ~/.gstack/.plan-tune-nudge-shown` before next ship.
|
||||
The marker or enabled question_tuning suppresses it. To re-enable, remove
|
||||
`$GSTACK_STATE_ROOT/.plan-tune-nudge-shown` before the next ship.
|
||||
|
||||
---
|
||||
|
||||
## Section self-check (before you finish)
|
||||
|
||||
You ran a carved skill. For your situation, list every section the Section index
|
||||
named as applying, and confirm you issued a Read for each one. If you executed any
|
||||
of those steps from memory without reading its section, you skipped the source of
|
||||
truth — STOP, Read it now, and redo that step. Deterministic version work goes
|
||||
through `gstack-version-bump`; never hand-roll the VERSION/package.json write.
|
||||
List the applicable Section index entries and confirm each Read. If you worked from
|
||||
memory, STOP, Read the section and redo that step. Use `gstack-version-bump`, never
|
||||
hand-roll VERSION/package.json writes.
|
||||
|
||||
---
|
||||
|
||||
## Important Rules
|
||||
|
||||
- **Never skip tests.** If tests fail, stop.
|
||||
- **Never skip the pre-landing review.** If checklist.md is unreadable, stop.
|
||||
Follow the numbered gates and their explicit exceptions.
|
||||
|
||||
- **Never force push.** Use regular `git push` only.
|
||||
- **Never ask for trivial confirmations** (e.g., "ready to push?", "create PR?"). DO stop for: version bumps (MINOR/MAJOR), pre-landing review findings (ASK items), and Codex structured review [P1] findings (large diffs only).
|
||||
- **Always use the 4-digit version format** from the VERSION file.
|
||||
- **Date format in CHANGELOG:** `YYYY-MM-DD`
|
||||
- **Split commits for bisectability** — each commit = one logical change.
|
||||
- **TODOS.md completion detection must be conservative.** Only mark items as completed when the diff clearly shows the work is done.
|
||||
- **Use Greptile reply templates from greptile-triage.md.** Every reply includes evidence (inline diff, code references, re-rank suggestion). Never post vague replies.
|
||||
- **Never push without fresh verification evidence.** If code changed after Step 5 tests, re-run before pushing.
|
||||
- **Step 7 generates coverage tests.** They must pass before committing. Never commit failing tests.
|
||||
- **The goal is: user says `/ship`, next thing they see is the review + PR URL + auto-synced docs.**
|
||||
Reference in new issue
Block a user