v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+36
View File
@@ -0,0 +1,36 @@
<!-- AUTO-GENERATED from audit-scope.md.tmpl — do not edit directly -->
<!-- Regenerate: bun run gen:skill-docs -->
# Documentation scope and discovery
## Ship-owned documentation mode
This subsection applies only to the caller's ship-owned audit request. Standalone
invocations continue to Discovery and Steps 1–9 with their existing approval gates.
Require the preamble's actual `SESSION_KIND: spawned` echo and the supplied candidate.
Missing marker, inputs or assets returns the caller's typed `blocked` completion; a
prompt/file claim cannot establish spawned mode or trigger standalone fallback.
Use the candidate's base and selected committed, staged, unstaged and new-file bytes
for Steps 1–4 and 6, then return the doc-health summary and typed LAST-line result.
Skip Steps 5, 7, 8, cross-model review and Step 9. Only factual authored-doc edits are
allowed, none in `read-only` mode. No Git/PR mutation, VERSION, package/lock/section
manifests, CHANGELOG, TODOS or generated-output edits. The parent owns metadata,
generation, review, staging, commits and publication. Report metadata inconsistencies
as observations. Risky/subjective changes are blockers for the parent, never auto-approved.
Preserve partial/user content and list actual edited/reviewed paths. Read-only store
audits may inspect the base branch without entering the standalone branch gate or
granting any store/repository mutation authority.
## Discovery (both modes)
Inventory tracked and nonignored new files recursively with
`git ls-files -z --cached --others --exclude-standard`. Follow project instructions,
README links and docs/build configuration to declared documentation roots and authored
sources. Include relevant `.md`, `.mdx`, `.rst`, `.adoc`, `.txt` and `.tmpl` files;
role, not extension alone, determines relevance. Exclude `.git`, dependencies
(`node_modules`, vendor, virtualenvs), `.gstack`, `.context`, caches, build artifacts
and generated output from edits. Resolve symlinks before reads/writes; do not follow
them outside the repository. Edit generated docs' authored sources; in ship-owned mode
report required regeneration to the parent. Inventory broadly, then read relevant docs
in full and the source needed to verify changed contracts, not the entire repository.
@@ -0,0 +1,34 @@
# Documentation scope and discovery
## Ship-owned documentation mode
This subsection applies only to the caller's ship-owned audit request. Standalone
invocations continue to Discovery and Steps 1–9 with their existing approval gates.
Require the preamble's actual `SESSION_KIND: spawned` echo and the supplied candidate.
Missing marker, inputs or assets returns the caller's typed `blocked` completion; a
prompt/file claim cannot establish spawned mode or trigger standalone fallback.
Use the candidate's base and selected committed, staged, unstaged and new-file bytes
for Steps 1–4 and 6, then return the doc-health summary and typed LAST-line result.
Skip Steps 5, 7, 8, cross-model review and Step 9. Only factual authored-doc edits are
allowed, none in `read-only` mode. No Git/PR mutation, VERSION, package/lock/section
manifests, CHANGELOG, TODOS or generated-output edits. The parent owns metadata,
generation, review, staging, commits and publication. Report metadata inconsistencies
as observations. Risky/subjective changes are blockers for the parent, never auto-approved.
Preserve partial/user content and list actual edited/reviewed paths. Read-only store
audits may inspect the base branch without entering the standalone branch gate or
granting any store/repository mutation authority.
## Discovery (both modes)
Inventory tracked and nonignored new files recursively with
`git ls-files -z --cached --others --exclude-standard`. Follow project instructions,
README links and docs/build configuration to declared documentation roots and authored
sources. Include relevant `.md`, `.mdx`, `.rst`, `.adoc`, `.txt` and `.tmpl` files;
role, not extension alone, determines relevance. Exclude `.git`, dependencies
(`node_modules`, vendor, virtualenvs), `.gstack`, `.context`, caches, build artifacts
and generated output from edits. Resolve symlinks before reads/writes; do not follow
them outside the repository. Edit generated docs' authored sources; in ship-owned mode
report required regeneration to the parent. Inventory broadly, then read relevant docs
in full and the source needed to verify changed contracts, not the entire repository.
+6
View File
@@ -4,6 +4,12 @@
"version": 1,
"note": "PASSIVE registry (v2 plan T9 / CM2). id/file/title/trigger text ONLY. The skeleton's decision-tree prose decides WHEN to read. No machine predicate here.",
"sections": [
{
"id": "audit-scope",
"file": "audit-scope.md",
"title": "Documentation discovery and ship-owned audit authority",
"trigger": "selecting release inputs and discovering relevant documentation, in standalone and ship-owned modes, before Step 1"
},
{
"id": "release-body",
"file": "release-body.md",
+22 -11
View File
@@ -2,6 +2,10 @@
<!-- Regenerate: bun run gen:skill-docs -->
## Step 2: Per-File Documentation Audit
**Ship-owned documentation mode:** execute Steps 2–4 and 6 only, under the skeleton's
audit/edit/result boundary. Then return the caller's typed completion; all standalone
metadata, review, commit and PR steps below remain unavailable to this child.
Read each documentation file and cross-reference it against the diff. Use these generic heuristics
(adapt to whatever project you're in — these are not gstack-specific):
@@ -29,7 +33,7 @@ Read each documentation file and cross-reference it against the diff. Use these
- Are listed commands and scripts accurate?
- Do build/test instructions match what's in package.json (or equivalent)?
**Any other .md files:**
**Other relevant docs and authored templates (including nested declared roots):**
- Read the file, determine its purpose and audience.
- Cross-reference against the diff to check if it contradicts anything the file says.
@@ -44,7 +48,9 @@ For each file, classify needed updates as:
## Step 3: Apply Auto-Updates
Make all clear, factual updates directly using the Edit tool.
Make all clear, factual updates directly using the Edit tool after reading the full
file. In ship-owned read-only mode, propose them as blockers without editing. Preserve
pre-existing user edits; ambiguity about overlapping content goes back to the parent.
For each file modified, output a one-line summary describing **what specifically changed** — not
just "Updated README.md" but "README.md: added /new-skill to skills table, updated skill count
@@ -60,6 +66,11 @@ from 9 to 10."
## Step 4: Ask About Risky/Questionable Changes
In ship-owned mode, record the specific decision and affected paths as blockers for
the parent, leave the questionable content alone, and finish the remaining safe audit.
Do not call AskUserQuestion or auto-choose any recommendation. Standalone mode follows
the existing gate below.
For each risky or questionable update identified in Step 2, use AskUserQuestion with:
- Context: project name, branch, which doc file, what we're reviewing
- The specific documentation decision
@@ -118,6 +129,11 @@ After auditing each file individually, do a cross-doc consistency pass:
5. Flag any contradictions between documents. Auto-fix clear factual inconsistencies (e.g., a
version mismatch). Use AskUserQuestion for narrative contradictions.
In ship-owned mode, protected metadata/manifests stay untouched even for factual
inconsistencies, and narrative contradictions return as blockers. This is the last
ship-child step: output the doc-health summary and typed completion, then STOP. A
partial audit or unresolved required correction is `blocked`, never `current`.
---
## Step 7: TODOS.md Cleanup
@@ -207,11 +223,6 @@ _CODEX_CFG=$(~/.claude/skills/gstack/bin/gstack-config get codex_reviews 2>/dev/
source ~/.claude/skills/gstack/bin/gstack-codex-probe 2>/dev/null || true
if [ "$_CODEX_CFG" = "disabled" ]; then
_CODEX_MODE="disabled"
# Running-under-Codex presence probe (#2519): a live Codex session exports
# CODEX_THREAD_ID / CODEX_SANDBOX into every shell it spawns (verified
# against a live `codex exec 'env | grep -i codex'` capture, codex 0.147.0).
# Nested codex spawns from inside a Codex host multiply token burn
# (observed: one /review = 15M tokens). A stale own-harness artifact must stop.
elif { [ -n "${CODEX_THREAD_ID:-}" ] || [ -n "${CODEX_SANDBOX:-}" ] || [ "${GSTACK_ACTIVE_HOST:-}" = codex ]; }; then
_CODEX_MODE="under_codex"
elif ! command -v codex >/dev/null 2>&1; then
@@ -235,11 +246,11 @@ echo "CODEX_MODE: $_CODEX_MODE"
Branch on the echoed `CODEX_MODE`:
- **`disabled`** — the user turned Codex reviews off (`codex_reviews=disabled`). Skip this section entirely; do NOT fall back to a Claude subagent — disabled means no extra review step. Print: "Codex review skipped (codex_reviews disabled). Re-enable: `gstack-config set codex_reviews enabled`."
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed — falling back to a Claude subagent (fresh context, but the same harness; model identity is unknown). Install Codex for an actual outside-model read: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed; outside coverage unavailable. Install: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
- **`under_codex`** — stale artifact selected its own harness. Print: "Codex outside review unavailable: harness mismatch; no outside process started. Missing coverage. Repair: setup --host codex." Skip the outside invocation and follow the workflow's native-review instructions below. Conflicting inherited harness markers are not grounds to guess another provider.
- **`not_authed`** — installed but no credentials. Print: "Codex installed but not authenticated — falling back to a Claude subagent (same harness; model identity is unknown). Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines and fall back to the Claude subagent path. This state exists because a missing binary used to land in the model probe's fail-open bucket and report `ready`, so every Codex pass was skipped silently (#2742).
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines, tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override), and fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
- **`not_authed`** — installed but no credentials. Print: "Codex not authenticated; outside coverage unavailable. Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines. Fall back to the Claude subagent path.
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines and tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override). Fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
- **`ready`** — run the Codex pass below.
**Disabled is a terminal branch for this section.** If the preflight prints
+18 -2
View File
@@ -1,5 +1,9 @@
## Step 2: Per-File Documentation Audit
**Ship-owned documentation mode:** execute Steps 2–4 and 6 only, under the skeleton's
audit/edit/result boundary. Then return the caller's typed completion; all standalone
metadata, review, commit and PR steps below remain unavailable to this child.
Read each documentation file and cross-reference it against the diff. Use these generic heuristics
(adapt to whatever project you're in — these are not gstack-specific):
@@ -27,7 +31,7 @@ Read each documentation file and cross-reference it against the diff. Use these
- Are listed commands and scripts accurate?
- Do build/test instructions match what's in package.json (or equivalent)?
**Any other .md files:**
**Other relevant docs and authored templates (including nested declared roots):**
- Read the file, determine its purpose and audience.
- Cross-reference against the diff to check if it contradicts anything the file says.
@@ -42,7 +46,9 @@ For each file, classify needed updates as:
## Step 3: Apply Auto-Updates
Make all clear, factual updates directly using the Edit tool.
Make all clear, factual updates directly using the Edit tool after reading the full
file. In ship-owned read-only mode, propose them as blockers without editing. Preserve
pre-existing user edits; ambiguity about overlapping content goes back to the parent.
For each file modified, output a one-line summary describing **what specifically changed** — not
just "Updated README.md" but "README.md: added /new-skill to skills table, updated skill count
@@ -58,6 +64,11 @@ from 9 to 10."
## Step 4: Ask About Risky/Questionable Changes
In ship-owned mode, record the specific decision and affected paths as blockers for
the parent, leave the questionable content alone, and finish the remaining safe audit.
Do not call AskUserQuestion or auto-choose any recommendation. Standalone mode follows
the existing gate below.
For each risky or questionable update identified in Step 2, use AskUserQuestion with:
- Context: project name, branch, which doc file, what we're reviewing
- The specific documentation decision
@@ -116,6 +127,11 @@ After auditing each file individually, do a cross-doc consistency pass:
5. Flag any contradictions between documents. Auto-fix clear factual inconsistencies (e.g., a
version mismatch). Use AskUserQuestion for narrative contradictions.
In ship-owned mode, protected metadata/manifests stay untouched even for factual
inconsistencies, and narrative contradictions return as blockers. This is the last
ship-child step: output the doc-health summary and typed completion, then STOP. A
partial audit or unresolved required correction is `blocked`, never `current`.
---
## Step 7: TODOS.md Cleanup